Feedback
AI Ad Video Example
Loading...
minimax h3 video model
Produce 2K video with built-in stereo audio in one request using the minimax h3 video model. It handles text, images, clips, and sound in up to 15 seconds.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.
MiniMax H3 video model: create 2K video with sound from any input
The minimax h3 video model is an open-weight, general-purpose omni-modal generation model hosted on fal.ai from day one. It processes text, stills, motion, and audio in a single context, producing up to 15 seconds of 2K output with native stereo audio, precise localized edits, clean on-screen text, and up to 12 reference inputs.
- All modalities in one shared contextFeed it up to 9 images, 3 video clips, and 3 audio tracks in one pass; the minimax h3 video model blends subject, performance, camera motion, and sound into a single cohesive result.
- Stereo sound composed with the pictureEvery output carries original music, dialogue, foley, and ambience locked to the edit through the minimax h3 video model, plus voice transfer or cloning from any reference recording.
- Targeted edits without touching the restSwap objects, change signs, replace dialog, or switch day to night. The minimax h3 video model modifies only the region you specify, while the surrounding frame remains unchanged.
Kick off your minimax h3 video model workflow in three steps
Make one API call to the minimax h3 video model and receive 2K footage with matched audio in minutes.
Production-ready capabilities of the minimax h3 video model
Three API endpoints, one omni-modal context, native stereo sound, local edits, crisp text rendering, and usage-based pricing — the minimax h3 video model covers the full 2K video workflow on fal.ai.
Three paths to generate a video
Use text-to-video, image-to-video with first/last-frame control, or reference-to-video through the minimax h3 video model. Each route fits a different production style.
Twelve reference inputs per request
The minimax h3 video model accepts 9 reference images, 3 video clips, and 3 audio tracks, reading identity, performance, camera work, composition, and cut rhythm from them.
Legible text and living interfaces
Generate clear titles, end cards, captions, and logos, or animate UI elements like landing pages, menus, HUDs, and kinetic typography with the minimax h3 video model.
Long prompts for complete scenes
Pack an entire shot list into one request. The minimax h3 video model supports up to 7,000 characters of prompt text, giving you granular control over every frame.
2K resolution and fluid 24fps playback
Receive 1440px-short-edge video at 24fps for up to 15 seconds, with six aspect ratios plus adaptive mode from the minimax h3 video model.
Serverless pricing per generation
Use the minimax h3 video model through a pay-as-you-go serverless API with no minimum or subscription, and own commercial rights to everything it generates.
Questions creators ask about the minimax h3 video model
Quick answers for working with the minimax h3 video model on fal.ai — from endpoints and resolution to audio and licensing.
Can you explain the minimax h3 video model?
It is MiniMax's open-weight, omni-modal model and a Day 0 launch partner on fal.ai. In a single context it handles text, images, video, and audio, and produces 2K clips with native stereo audio up to 15 seconds.
Which API endpoints are exposed?
The minimax h3 video model exposes three routes: text-to-video, image-to-video with optional first/last-frame control, and reference-to-video. The reference mode keeps subjects, style, motion, camera direction, and voices from your source assets.
What resolutions, durations, and aspect ratios are available?
Renders reach 2K (1440px short edge) at 24fps. Request anywhere from 5 to 15 seconds and pick 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16, or choose adaptive.
Does the output really include audio?
Yes, every generation from the minimax h3 video model includes true stereo sound — original music, dialogue, foley, and ambience synced to the cut — with optional voice transfer or cloning from reference tracks.
How many references can I upload at once?
Use up to 12 references: 9 images, 3 video clips, and 3 audio tracks, each clip or track lasting 2-15 seconds. Audio must always be paired with at least one image or video when calling the minimax h3 video model.
Can I use the generated clips commercially?
Yes. Clips generated through the fal.ai API with the minimax h3 video model can be used in commercial projects, subject to fal.ai's terms of service.
Bring your next video idea to life with the minimax h3 video model
Produce 2K clips with true stereo audio in a single multimodal request. The minimax h3 video model gives you precise editing, flexible inputs, and serverless pricing on fal.ai.
