MiniMax H3 Prompt Guide for Text, Image, Audio, and References
Write better MiniMax H3 prompts for Hailuo 03 video, including shots, camera motion, dialogue, stereo sound, first and last frames, and multimodal references.
MiniMax H3 prompts work best when they describe a video as a sequence that changes over time. A list of attractive visual keywords may define a look, but it does not tell H3 who moves, how the camera responds, when dialogue begins, or how the scene should sound.
This guide turns the official H3 prompting format into a practical workflow for text-to-video, image-to-video, first-and-last-frame, and reference-to-video. Hailuo 03 and Hailuo 3 are common names for the same H3 workflow on this site.
The reliable MiniMax H3 prompt structure
Start with six decisions:
- Format: live action, animation, product film, motion design, or another clear visual language.
- Opening composition: subject, environment, framing, and lighting.
- Observable action: what changes from the first moment to the last.
- Camera: one motivated movement or a deliberately static shot.
- Sound: ambience, physical sounds, dialogue, and optional score.
- Ending state: the composition or action the clip should resolve on.
For a short single-shot clip, a natural-language prompt is often enough:
Live-action cinematic product film. A translucent orange perfume bottle stands on wet black stone at blue hour. Condensation travels down the glass as the camera makes a slow, small clockwise arc. A narrow beam of warm light moves across the label, then the camera settles on a clean front-facing hero composition. Stereo sound: light rain, a quiet glass resonance, and no background music.
The prompt gives the model a subject, motion, camera path, light change, final frame, and sound plan without stacking incompatible ideas.
Structure longer clips as shots
Use multiple shots only when each cut reveals new information. Number them in playback order and give later shots an increasing cut time. The official guide does not require a timestamp on the first shot.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a wide
shot frames a baker opening a small shop before sunrise. The camera slowly
pushes toward the counter as she places a warm loaf beneath a pendant light.
[Shot 2] At 00:04.000, the shot cuts to a close-up of steam rising from the
bread while her final words continue across the cut.
overall_soundscape: Wooden shutters scrape open, trays clink softly, and the
doorbell rings once above a quiet street.
non_diegetic_music: Sparse upright-bass notes at a slow tempo, fading at the
end.
For a four-to-six-second clip, one strong shot usually gives the model more room than three rushed cuts. Add a cut only when the story needs a different viewpoint, place, time, or subject state.
Direct camera motion precisely
Camera direction has three useful dimensions: movement type, range, and speed. Write them as part of the action rather than as a keyword list.
| Goal | Useful direction |
|---|---|
| Move physically closer | slow push in |
| Change focal length | controlled zoom in |
| Follow a moving subject | low tracking shot |
| Reveal space horizontally | fast pan right or truck right |
| Move around a subject | slow arc shot |
| Preserve exact composition | locked static shot |
Avoid combinations such as “zoom, drone orbit, handheld tracking, whip pan” in one short shot. Choose the movement that best reveals the action.
Prompt native dialogue and stereo sound
MiniMax H3 can create native stereo sound with the video. Treat sound as part of the timeline rather than an afterthought.
Use stable speaker IDs when more than one person talks. Keep the spoken words inside a dialogue block and describe the voice outside it:
The woman with a quiet, breathy voice (S1) turns toward the train window and
says: <d>[English] I get off at the next station.</d>
If the line is voiceover, say that it is off-screen and keep the visible
character's lips closed. Put room tone, footsteps, fabric, weather, impacts,
and other physical sound in overall_soundscape. Put audience-only background
music in non_diegetic_music, or use N/A when you want no score.
Image-to-video: describe the change, not the picture
The first frame already defines appearance. Repeat only the identity details that must remain stable, then explain what happens next.
For the target video, at 0.00 seconds into the target video, <Picture 1>
(from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Preserve the woman's face,
clothing, seat position, and the rain-covered train window from <Picture 1>.
The camera slowly trucks right as she lifts her gaze toward the passing city
lights and folds the letter along its existing crease.
overall_soundscape: Train wheels form a steady metallic rhythm beneath low
ventilation and rain against glass.
non_diegetic_music: N/A
An effective image-to-video prompt answers: what stays fixed, what starts moving, how far it moves, what the camera does, and where the shot ends.
First and last frame: write the path between them
Do not describe two isolated pictures. Explain the visible transition that can connect them. A single continuous shot is normally the clearest option.
Picture 1 aligns with the opening frame and Picture 2 aligns with the final
frame. Preserve the cyclist and street throughout. She releases the bicycle
handle, lifts the closed umbrella, opens it above her shoulder, and steps
beneath it while the camera slowly pulls back. Water rolls from the expanding
fabric. End on the exact pose, spacing, and composition in Picture 2.
If the requested endpoint requires an impossible jump in identity, geometry, or camera position, simplify the action or use an intermediate reference.
Reference-to-video: assign every asset one role
Reference mode can combine images, video, and audio. Ambiguous references compete with one another, so identify what each source controls:
- Image 1: character identity and wardrobe.
- Image 2: product shape, materials, and colors.
- Video 1: movement and camera rhythm only.
- Audio 1: voice timbre and delivery only.
Then state what must not transfer. For example: “Use Video 1 for the walking rhythm and camera path; do not copy its actor, background, clothing, or lighting.” This is more reliable than asking the model to “use the style of all references.”
Common prompt failures
Too many events for the duration
Five transformations, three cuts, dialogue, and a final logo reveal rarely fit inside six seconds. Reduce the idea to one subject, one action, one camera move, and one ending.
Conflicting camera directions
A camera cannot remain locked while orbiting and tracking. Decide whether the subject or camera supplies the motion.
Identity without preservation rules
For people and products, name the stable traits: face, hair, clothing, silhouette, label, material, color, and relative proportions.
Generic sound language
“Epic audio” gives little timing information. Name the source and moment: footsteps cross left to right, a glass click occurs on the reveal, or a low room tone continues throughout.
Negative constraints without positive direction
“No flicker, no morphing, no cuts” does not define the desired shot. Write the positive sequence first, then add a short preservation list.
A reusable checklist
Before generating, confirm that the prompt contains:
- one clear visual style and opening composition;
- an action that fits the selected duration;
- one coherent camera plan;
- an explicit ending state;
- separate ambience, dialogue, and music instructions;
- stable identities and a role for every reference;
- only the constraints that protect the intended result.
For current parameter limits and reference counts, see the MiniMax H3 review. Browse sourced outputs in MiniMax H3 examples, or create with MiniMax H3 when your prompt is ready.