MiniMax

MiniMax H3

MiniMax H3 is a multimodal video model for building a shot from text, animating a first frame, targeting a last frame, interpolating between both endpoints, or guiding the result with up to nine images, three videos, and three audio clips. It supports explicit 4–15 second durations and 768P or 2K output; text-to-video uses a selected aspect ratio, while frame-driven generation follows the supplied image through adaptive sizing.

How to use this model

  1. 1

    Choose text-to-video, a first or last frame, both endpoint frames, or a multimodal reference workflow before attaching media.

  2. 2

    Describe the subject, action, camera movement, style, sound, and editing rhythm in one coherent prompt.

  3. 3

    Use a concrete aspect ratio for text-to-video and select adaptive whenever a first or last frame is supplied.

  4. 4

    Add only the image, video, and audio references that have a clear role in the intended shot.

  5. 5

    Select duration and resolution, then review continuity, identity, motion, and synchronization across the full result.

H3Available

MiniMax H3

Multimodal MiniMax H3 video generation from text, first or last frames, and image, video, or audio references, with 4-15 second output up to 2K.

Properties

  • Supports text to video workflows.
  • Supports image to video workflows.

Best for

  • Reference-guided character, motion, camera, style, voice, and editing-rhythm control
  • First-frame, last-frame, and first-and-last-frame video generation
  • Flexible 4-15 second clips at 768P or 2K

Avoid for

  • Output longer than 15 seconds or resolutions other than 768P and 2K

Tips

  • Select adaptive aspect ratio whenever a first or last frame is supplied.

How to use this model

  1. 1

    Describe the scene, camera movement, subject motion, and timing.

  2. 2

    Add the supported starting image, ending image, or video reference when needed.

  3. 3

    Select duration, resolution, and aspect ratio before starting the generation.

Parameters

Prompt

Required

Describe the scene, action, camera, style, sound, and editing rhythm for the generated video.

First Frame

Optional

Optional opening frame. The output uses its aspect ratio; select adaptive aspect ratio when provided.

Last Frame

Optional

Optional final frame. It can be used alone or together with a first frame; select adaptive aspect ratio when provided.

Reference Images

Optional

Up to nine images for character, composition, content, or style guidance.

Reference Videos

Optional

Up to three videos for motion, camera, content, style, or editing guidance.

Reference Audio

Optional

Up to three audio clips for voice, sound, timing, or editing-rhythm guidance.

Duration

Optional

Output duration from 4 to 15 seconds.

  • 4 seconds

  • 5 seconds

  • 6 seconds

  • 7 seconds

  • 8 seconds

  • 9 seconds

  • 10 seconds

  • 11 seconds

  • 12 seconds

  • 13 seconds

  • 14 seconds

  • 15 seconds

Resolution

Optional

Output resolution: 768P or 2K.

  • 768P

  • 2K

Aspect Ratio

Optional

Choose a concrete ratio for text-to-video or adaptive for first- or last-frame generation.

  • 21:9

  • 16:9

  • 4:3

  • 1:1

  • 3:4

  • 9:16

  • Adaptive