Unified Multimodal Context
Tell H3 how each reference should shape the result in plain language: hold a character’s identity from an image, borrow camera movement from a video, and match vocals or sound from audio in the same generation.
Try it NowCreate AI video from a prompt or guide it with images, video, and audio in the same request. MiniMax H3 generates up to 15 seconds at 2K with native stereo sound and precise multimodal control.
Connect text, image, video, and audio in one creative context, generate native stereo video at up to 2K, and keep motion and subject detail under tighter control.
Tell H3 how each reference should shape the result in plain language: hold a character’s identity from an image, borrow camera movement from a video, and match vocals or sound from audio in the same generation.
Try it NowGenerate visuals and stereo audio together in one pass instead of adding sound afterward. Direct dialogue, music, ambience, and effects alongside the shot for more coherent audiovisual timing.
Try it NowWrite the camera move, lighting, atmosphere, and pacing into the brief and have the shot follow it. Use video references for motion transfer or targeted changes while preserving the rest of the scene.
Try it NowThree steps to your first multimodal AI video.
Step 01
Open Spira’s video generator and select MiniMax H3. Pick your duration, aspect ratio, and either 768p or 2K output.

Step 02
Attach reference images, a reference video, or a reference audio track, then describe the subject, action, camera movement, lighting, and pacing.

Step 03
Generate the clip, check how closely it followed your references, refine the direction if needed, then download the strongest result.

| Pricing | |||||
| Starting price | From $9/month | Pay as you go | From $110/month | From $19/month | From $18/seat/month |
| GPT Image 2 | $0.03 | Not supported | $0.55 | $0.35 | $0.2635 |
| Nano banana image | $0.01 | $0.04 | $0.55 | $0.1 | $0.047 |
| Seedance2.0 Video, 720P | $0.15/s | $0.30/s | $0.66/s | $0.59/s | $0.21/s |
| Veo3.1, 720p | $0.06/s | $0.20/s | $1.37/s | $0.07/s | $0.48/s |
| Gemini Omni Video, 720 P | $0.06/s | Not supported | Not supported | Not supported | Not supported |
| Kling 3.0 | $0.06/s | $0.084/s | $0.8/s | $0.086/s | $0.105/s |
H3 is a general-purpose multimodal model that connects text, images, video, and audio in one context. It can generate video and native stereo sound together, follow detailed creative instructions, transfer motion from video, and make reference-driven changes without splitting those jobs across separate models.
Up to 9 reference images, 3 reference videos, and 3 reference audio clips, with a combined limit of 12 items. Reference videos and audio clips must each be 2 to 15 seconds long, and audio cannot be the only reference.
Yes. One image can define the first frame, two images can define the first and last frames, and a reference video can guide motion, camera language, timing, or a targeted transformation. In multimodal reference mode, images can also define identity, products, or visual style instead of acting as frames.
Clips run from 4 to 15 seconds at 768p or 2K, in 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. When you generate from references, the output can also follow the references’ own aspect ratio automatically.
Cost is based on the output resolution and the length of the clip, and reference videos add their own duration to the total. Spira shows the credit cost for your exact settings before you generate.