MiniMax H3

Create AI video from a prompt or guide it with images, video, and audio in the same request. MiniMax H3 generates up to 15 seconds at 2K with native stereo sound and precise multimodal control.

Key Features of MiniMax H3

Connect text, image, video, and audio in one creative context, generate native stereo video at up to 2K, and keep motion and subject detail under tighter control.

Unified Multimodal Context

Tell H3 how each reference should shape the result in plain language: hold a character’s identity from an image, borrow camera movement from a video, and match vocals or sound from audio in the same generation.

Try it Now

Native 2K Video with Stereo Sound

Generate visuals and stereo audio together in one pass instead of adding sound afterward. Direct dialogue, music, ambience, and effects alongside the shot for more coherent audiovisual timing.

Try it Now

Commercial-Grade Creative Control

Write the camera move, lighting, atmosphere, and pacing into the brief and have the shot follow it. Use video references for motion transfer or targeted changes while preserving the rest of the scene.

Try it Now

How to Use MiniMax H3 on Spira

Three steps to your first multimodal AI video.

  1. Step 01

    Choose Your Preferred Model

    Open Spira’s video generator and select MiniMax H3. Pick your duration, aspect ratio, and either 768p or 2K output.

    Spira video generator with MiniMax H3 selected
  2. Step 02

    Add References and Direct the Shot

    Attach reference images, a reference video, or a reference audio track, then describe the subject, action, camera movement, lighting, and pacing.

    Prompt and reference media controls in the Spira generator
  3. Step 03

    Generate, Review, and Export

    Generate the clip, check how closely it followed your references, refine the direction if needed, then download the strongest result.

    Finished MiniMax H3 clip previewed in Spira

Compare across

Spira Plans
Pricing
Starting priceFrom $9/monthPay as you goFrom $110/monthFrom $19/monthFrom $18/seat/month
GPT Image 2$0.03Not supported$0.55$0.35$0.2635
Nano banana image$0.01$0.04$0.55$0.1$0.047
Seedance2.0 Video, 720P$0.15/s$0.30/s$0.66/s$0.59/s$0.21/s
Veo3.1, 720p$0.06/s$0.20/s$1.37/s$0.07/s$0.48/s
Gemini Omni Video, 720 P$0.06/sNot supportedNot supportedNot supportedNot supported
Kling 3.0$0.06/s$0.084/s$0.8/s$0.086/s$0.105/s

Frequently asked questions about MiniMax H3

  • H3 is a general-purpose multimodal model that connects text, images, video, and audio in one context. It can generate video and native stereo sound together, follow detailed creative instructions, transfer motion from video, and make reference-driven changes without splitting those jobs across separate models.

  • Up to 9 reference images, 3 reference videos, and 3 reference audio clips, with a combined limit of 12 items. Reference videos and audio clips must each be 2 to 15 seconds long, and audio cannot be the only reference.

  • Yes. One image can define the first frame, two images can define the first and last frames, and a reference video can guide motion, camera language, timing, or a targeted transformation. In multimodal reference mode, images can also define identity, products, or visual style instead of acting as frames.

  • Clips run from 4 to 15 seconds at 768p or 2K, in 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16. When you generate from references, the output can also follow the references’ own aspect ratio automatically.

  • Cost is based on the output resolution and the length of the clip, and reference videos add their own duration to the total. Spira shows the credit cost for your exact settings before you generate.