Brands

MiniMax

MiniMax H3 and Hailuo video generation plus speech synthesis and voice cloning.

Models

Video

Hailuo Video

View details

Hailuo 02 and Hailuo 2.3 support text-to-video and image-to-video workflows, while 2.3 Fast is intended for faster image-driven iteration. The models accept prompts for action and camera direction; Hailuo's API also supports bracketed camera commands for compatible image-to-video variants.

MiniMax Hailuo 02

#Available

Hailuo 02 generates video from text, a first frame, or paired first and last frames.

MiniMax Hailuo 2.3

#Available

Hailuo 2.3 generates video from text or a first-frame image, with controls for camera direction, duration, and resolution.

MiniMax Hailuo 2.3 Fast

#Available

Hailuo 2.3 Fast for quicker image-to-video iteration from a supplied first frame.

MiniMax H3 is a multimodal video model for building a shot from text, animating a first frame, targeting a last frame, interpolating between both endpoints, or guiding the result with up to nine images, three videos, and three audio clips. It supports explicit 4–15 second durations and 768P or 2K output; text-to-video uses a selected aspect ratio, while frame-driven generation follows the supplied image through adaptive sizing.

MiniMax H3

#Available

Generate 4-15 second MiniMax H3 videos from text, endpoint frames, or multimodal references at 768P or 2K.

Models

Audio

MiniMax Speech

View details

MiniMax Speech 2.8 HD turns text into speech using a selected voice ID. Omnipix exposes emotion, speed, volume, pitch, language assistance, channel layout, sample rate, bitrate, file format, and optional subtitles. It is suited to workflows that need more explicit control over both performance and the delivered audio file.

MiniMax Speech 2.8 HD

#Available

Synthesize controlled multilingual speech with preset or compatible cloned voices and production-oriented audio settings.

MiniMax Voice Cloning

View details

MiniMax Voice Cloning creates a reusable private voice from a reference audio file. Name the voice, enter a short preview script, and optionally enable noise reduction or volume normalization. Clean, single-speaker source audio makes evaluation easier; use only recordings you are authorized to submit and reproduce.

MiniMax Voice Cloning

#Available

Create a named MiniMax voice ID from an authorized reference recording for later speech synthesis.