COMING SOON TO SYLLABY
Get notified when it’s ready
MiniMax H3
As Featured In




















03 — WHY MINIMAX H3
Six things H3 does that the others do not
This is not the full feature list. It is the six things no other video model has together.
It hears, not just sees
Hand it a track and it reads the audio the way it reads your images and video, in one context.
VIDEO + STEREO
Native stereo sound
Not mono, not dubbed. Left and right arrive with the frames.
TEXT
It gets the text right
Titles, packaging and product type come out clean.
2K
Fifteen seconds at 2K
Fifteen seconds of video at 2K, sound included, in one generation.
Precise editing
Regenerate part of a clip. The rest never moves.
Motion transfer
Point at a clip you like and put its motion on your own subject.
Made with MiniMax H3
Made in Syllaby with MiniMax H3. Every one came out of the model in a single pass, with its sound.
05 — RESOLUTION
Fifteen seconds, straight out at 2K
Up to fifteen seconds at 2K with the sound made in the same pass. A cheaper 768p tier is listed too.
A full 15 second generation at native 2K.
2K
06 — AUDIO
Native stereo sound, made with the picture
Not mono, and not laid over the top afterwards. Voice, effects and music are modelled together with the picture, so what you hear sits where the action is.
STEREO
L / R
07 — INPUT
One prompt. Four kinds of input.
Text, images, video and audio arrive in one context. MiniMax’s own example: “Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matching Audio 3.”
08 — TEXT
Accurate text and brand rendering
Type is where most video models fall apart. MiniMax lists accurate text and brand rendering among the things H3 is built for.
Opening title card
Product page hero
Animated poster
Product end card
09 — EDITING
Precise video editing by reference
H3 edits by reference rather than starting over, so you can change a detail and everything you already approved stays put.
MiniMax H3 questions
What it is, when it lands, and what it will cost you.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
MiniMax H3, also called Hailuo 3.0, is MiniMax’s general purpose multimodal generative model, announced on 31 Jul 2026. It understands text, images, video and audio in one context, and generates video up to fifteen seconds at 2K with stereo sound made in the same pass.
Still have questions?
Contact our support teamCOMING TO SYLLABY
Get MiniMax H3 on day one
It is days away. Leave your email and we will tell you the moment it is live inside Syllaby.
One email the day it goes live. Nothing else.