Home Guides MiniMax H3 Max reference to video
MiniMax H3 Max reference to video
MiniMax H3 Max reference to video
H3 Max Studio · last updated 3 September 2026

MiniMax H3 Max reference to video lets you attach images, short video, and audio so the model can borrow texture, motion, or a voice. You can send up to 12 files. Audio cannot be the only reference. Adaptive aspect is available, plus the six official ratios.
When reference to video is the right mode
MiniMax H3 Max reference to video is for jobs that should inherit something you already shot or designed. A fabric, a camera rhythm, a palette, a 4-second product turn, a room tone. It is not a dumping ground for every file on your desktop. Each extra file is a constraint. Too many constraints and the clip becomes a muddle.
Use text to video when you have no files. Use image to video when you have one opening still and maybe an end still. Use reference to video when those two modes cannot see the material you need the model to notice.
Official R2V rules on this studio
- Up to 12 files total.
- File kinds: images, videos, audio.
- Audio cannot be the only reference. Pair it with at least one image or video.
- Reference videos should be 2 to 15 seconds each. Combined reference video is capped at 15 seconds.
- Aspect can be adaptive, or one of 21:9, 16:9, 4:3, 1:1, 3:4, 9:16.
- Prompt, duration 5–15s, 480P or 768P, expansion, seed, and safety still apply.
Those are the official fields we send. If a blog tells you to “drop a 2-minute master,” it is not describing this API.
Credits that people miss
Duration is still 5 credits per second. Then:
- Extra images after the first four cost 10 credits each.
- Each reference video costs 45 credits.
A 5-second R2V job with one image is 25 credits. The same job with two reference videos is 25 + 90 = 115 credits, before any extra stills. Failed jobs refund the reserved amount, including those add-ons, when the failure is a platform error.
This is why R2V is the wrong mode for a first experiment. Learn the prompt on T2V or I2V, then attach references when you know what they are for.

Pay duration first. Then pay for extra stills and each reference video.
How to choose files
Pick references that agree with each other. Three stills of the same bottle under the same light beat six stills from six campaigns. A 3-second handheld push-in beats a 15-second montage if you want that push-in.
Do not attach a competitor’s commercial and ask Max to “make it ours.” That is a rights problem and a taste problem. Do not attach a celebrity still you do not control. Do not attach only a voiceover and hope the model invents a picture — the API will not take audio alone.
Name files in the prompt when it helps. “Follow the motion of Video 1. Keep the label from Image 1. Use the room tone from Audio 1.” The studio labels files Image 1, Image 2, and so on as you add them. Use those names.
Adaptive versus a locked ratio
Adaptive lets the model pick a frame that fits the references. Use it when the stills and clips already share a shape. Lock a ratio when the delivery spec is fixed: a 9:16 story, a 16:9 site hero, a 21:9 pass.
If your references are mixed — a 9:16 phone clip and a 16:9 still — either crop before upload or accept that adaptive will compromise. The model is not an editor with a timeline. It is a generator that saw those files.
Prompting on top of references
The prompt still needs subject, camera, light, and sound. References do not replace those four parts. They bias them. A reference video of a scooter against giant type will not produce a scooter spot if the prompt describes a detective in fog. You will get a fight.
Write the shot you want. Then say what each file is for. Keep it short. A 200-word essay plus 12 files is how you spend 115 credits on noise.
A conservative R2V sequence
- Get a 5-second T2V or I2V draft you like.
- Switch to Reference. Attach one still and, if you truly need it, one short video.
- Keep duration at 5 seconds and 768P.
- Name the files in the prompt.
- Unmute. If the motion from the video took over the still, drop the video or shorten it.
- Only then add a second video or more stills.
That sequence is slower to describe than “upload everything.” It is cheaper to run.
What R2V will not do
It will not extend a 5-second output into a minute. It will not replace an editor. It will not clear a copyright. It will not lip-sync to a 3-minute interview you trimmed badly and uploaded as the only file. It will not run base MiniMax H3.
It will, on a good day, let a 768p clip inherit a fabric, a move, or a room so you stop fighting the empty frame.
Rights and a clean bin
Every file you attach is a rights statement. A moodboard scraped from search is not a reference library. A trailer you do not own is not a motion brief. If legal would not let you cut that file into a paid ad, do not send it to fal through this studio.
Keep a small bin per job: one hero still, one motion clip if you need it, one room tone if you recorded it. Name them before you upload so the prompt can say Image 1 and Video 1 without you guessing. Twelve unlabeled files is how R2V turns into a collage.
Commercial use of your output still requires that you had rights to the inputs. Credits do not wash a stolen still.
FAQ
Can audio be the only MiniMax H3 Max reference?
No. Pair audio with at least one image or video.
How many files can I send?
At most 12. Extra images after four cost 10 credits each. Each reference video costs 45 credits.
What is adaptive aspect?
A reference-to-video option that lets the model pick a frame that fits the files. Text to video cannot use it.
Should I start a project in reference to video?
Usually no. Prove the prompt in text or image mode, then attach references on purpose.