MiniMax H3
Multimodal MiniMax H3 video generation from text, first or last frames, and image, video, or audio references, with 4-15 second output up to 2K.
Properties
- Supports text to video workflows.
- Supports image to video workflows.
Best for
- Reference-guided character, motion, camera, style, voice, and editing-rhythm control
- First-frame, last-frame, and first-and-last-frame video generation
- Flexible 4-15 second clips at 768P or 2K
Avoid for
- Output longer than 15 seconds or resolutions other than 768P and 2K
Tips
- Select adaptive aspect ratio whenever a first or last frame is supplied.
How to use this model
- 1
Describe the scene, camera movement, subject motion, and timing.
- 2
Add the supported starting image, ending image, or video reference when needed.
- 3
Select duration, resolution, and aspect ratio before starting the generation.
Parameters
Prompt
promptDescribe the scene, action, camera, style, sound, and editing rhythm for the generated video.
First Frame
first_frame_imageOptional opening frame. The output uses its aspect ratio; select adaptive aspect ratio when provided.
Last Frame
last_frame_imageOptional final frame. It can be used alone or together with a first frame; select adaptive aspect ratio when provided.
Reference Images
reference_imagesUp to nine images for character, composition, content, or style guidance.
Reference Videos
reference_videosUp to three videos for motion, camera, content, style, or editing guidance.
Reference Audio
reference_audiosUp to three audio clips for voice, sound, timing, or editing-rhythm guidance.
Duration
durationOutput duration from 4 to 15 seconds.
4 seconds
5 seconds
6 seconds
7 seconds
8 seconds
9 seconds
10 seconds
11 seconds
12 seconds
13 seconds
14 seconds
15 seconds
Resolution
resolutionOutput resolution: 768P or 2K.
768P
2K
Aspect Ratio
aspect_ratioChoose a concrete ratio for text-to-video or adaptive for first- or last-frame generation.
21:9
16:9
4:3
1:1
3:4
9:16
Adaptive