All posts
Guide2026/08/05Updated 2026/08/05

MiniMax H3 Image-to-Video Guide — Hailuo 03 First and Last Frames

Animate one image or connect first and last frames with MiniMax H3. Learn preservation, motion, camera, sound, aspect-ratio, and reference techniques.

MiniMax H3 image-to-video starts from a still image and builds motion, camera movement, lighting changes, and native stereo sound around it. You can use one first frame or provide both a first and last frame when the shot must land on a specific composition.

The best source image is not necessarily the most detailed. It is the one that makes the subject, depth, and available movement easy to understand.

Choose a source image with room to move

Use a sharp image with a clear subject and visible separation between foreground and background. Check these details before upload:

  • hands, faces, product labels, and important edges are not cropped;
  • the subject has physical room to perform the requested action;
  • the aspect ratio matches the intended output;
  • lighting direction and reflections are internally consistent;
  • small text is avoided unless preserving it is essential.

A tightly cropped portrait may work for expression and subtle camera motion, but it gives the model little room for a full-body action. A wide product shot may support an orbit, while a flat front view may be better for a controlled push-in.

Tell H3 what must stay unchanged

Start the prompt with a short preservation contract. For a portrait, preserve face shape, hairstyle, age, clothing, and accessories. For a product, preserve silhouette, proportions, material, color, packaging, and readable brand details.

Then describe only the new motion:

Preserve the exact bottle shape, amber glass, cream label, black cap, and wet stone surface from the first frame. Condensation slowly travels down the glass while the camera makes a small clockwise arc. A warm reflection crosses the label and the shot ends on a stable front three-quarter view.

This is more useful than spending half the prompt redescribing what the image already shows.

Separate subject motion from camera motion

Decide which layer should move:

Desired resultSubject motionCamera motion
Portrait comes aliveblink, breath, small head turnlocked or slow push
Product reveallight, particles, restrained mechanismslow arc
Landscape atmosphereclouds, leaves, water, mistgentle pull back
Poster animationtypography or graphic layersmostly static

Large subject action combined with a fast camera move can make identity drift harder to diagnose. Begin with one controlled movement, then increase ambition only if the key details remain stable.

Use a last frame for a defined destination

First-and-last-frame generation is useful for product rotations, before/after states, pose transitions, interface reveals, and controlled scene changes. The prompt must describe a plausible path rather than two disconnected images.

Write the transition in four parts:

  1. establish the opening state;
  2. name the observable intermediate action;
  3. narrow the difference between the frames;
  4. land on the exact final composition.

If the first frame shows a closed package and the last frame shows it open, describe the lid lifting, hinges moving, contents appearing, and the camera settling. Do not ask for unrelated scene changes at the same time.

Add sound that follows the visible action

H3 can generate native stereo sound alongside the video. Tie each sound to a source: fabric movement for a turning person, hinge clicks for a package, wind and leaves for a landscape, or a restrained spatial whoosh for a camera move.

Keep audience-only music separate. Use non_diegetic_music: N/A when the clip should contain only scene sound. This prevents a generic soundtrack from competing with product or dialogue audio.

Match duration to the action

The current H3 workflow supports clips from 4 to 15 seconds. Short durations fit one direct action. Longer durations can support a slower transition or a small shot progression, but more time does not automatically justify more events.

Use 4–6 seconds for a blink, light sweep, subtle orbit, or single reveal. Use 8–15 seconds when the subject must cross space, perform multiple connected steps, or reach a substantially different final frame.

Troubleshooting image-to-video results

The subject changes identity

Reduce action range and camera speed. Put the preservation contract first and remove style words that imply a redesign.

The last frame appears only as a hard cut

Describe more intermediate states and remove extra cuts. Confirm that the two frames can be connected through continuous motion.

The image looks like a moving photograph

Add one concrete environmental or subject action: breathing, fabric response, light movement, reflections, drifting mist, or a clear interaction.

Product geometry melts or duplicates

Use a slower camera move, explicitly preserve part count and proportions, and avoid transformations that require hidden geometry the first frame does not show.

The crop changes unexpectedly

Choose an output ratio that matches the image or supply enough space around the subject. A wide image forced into portrait framing will lose content.

Image-to-video prompt template

Preserve [identity, materials, colors, proportions, clothing, and key details]
from the first frame.

[Subject] performs [one observable action] while [environment response]. The
camera [one movement with speed and range]. Lighting [one controlled change].
End on [specific stable composition or the supplied last frame].

overall_soundscape: [ambient bed and synchronized physical sounds].
non_diegetic_music: [instrumentation and pacing, or N/A].

Do not change [critical identity details]. Avoid [short list of likely errors].

Review more formats in the MiniMax H3 prompt guide, see sourced MiniMax H3 examples, or animate an image with H3.

Sources