Anyone who has handed a storyboard to a crew knows two kinds of collaborator. One shoots the boards. The other reads the boards as a suggestion and comes back with something you did not ask for. AI video models turn out to split along the same line, and a recent test makes the difference unusually easy to see.
One brief, two models, fifteen pairs
The team at Everypixel put two current video models, Seedance 2.0 and MiniMax H3, through the same set of briefs and published a shot-by-shot comparison of two AI video models in August 2026. The five scenarios read like a small slate of commercial and genre work: a handheld UGC selfie on a city street, a night car chase on wet neon roads, a stylized 3D cartoon about a small robot and a cookie jar, a cel-shaded 2D anime fight, and a 3D broadcast news opener with chrome typography. Each one was broken down beat by beat over ten seconds, with every camera position described.
Each scenario was then fed to the models in three ways: text only, text plus a pre-generated opening frame that locks characters, palette and framing, and a numbered six-panel storyboard sheet the model was asked to shoot in order. That makes fifteen head-to-head pairs; four are discussed in detail, and the authors say the other eleven were closer together.
They did not judge on pretty frames. Their criteria were the ones a 1st AD would care about: did the model shoot what the brief said, did anything appear that nobody asked for, did objects hold their shape in motion, did the shots read as one continuous event, and how many takes it took before a clip could go to a client.
The titan with two heads
The headline example is real and comes straight from the storyboard test. The six-panel anime sequence showed a lone hero on a broken rooftop, a stone titan as tall as the surrounding towers, and a single white slash through the titan’s wrist. According to the authors, MiniMax H3 shot the panels in the order they were drawn. Seedance 2.0 treated the first panel as a starting point, developed it, and returned a clip in which the titan had two heads, plus other elements that appear nowhere in the reference.
This was not a one-off. Across the storyboard runs, Seedance treated panels as artistic direction rather than specification, reordering beats and adding shots.
What each model got right and wrong
|
Scenario |
MiniMax H3 |
Seedance 2.0 |
|
UGC selfie, single take (text only) |
Handled it |
Handled it; differences not perceptible in ordinary viewing |
|
Night car chase, six setups (text only) |
Car stays intact, but the six setups never assemble into one chase |
The car loses its structure mid-clip; still closer to usable overall |
|
3D robot cartoon, six timed beats |
Respects timing; camera moves read as deliberate |
Mostly correct, then adds unrequested keys, disembodied hands and stray objects |
|
Six-panel anime storyboard |
Returns the sequence as drawn |
Reinterprets; titan gains a second head |
The chase is the useful counterweight. Both models needed roughly three attempts before either produced something usable, and the authors call it Seedance’s clearest advantage in the whole test. Their explanation: a ten-second action scene holds far more information than any prompt encodes, so a model that invents plausibly gets further on a thin brief.
On cost, the authors report that, at the API settings they used in production, ten seconds of Seedance output cost them three times as much as ten seconds on MiniMax. They flag that the Seedance figure includes a markup from the route they call it through, so it is not a clean rate-card comparison.
What this means for indie previs
The article’s key point is that improvisation and execution are a trade-off, not a quality gap. If you are using AI video to explore, to find a look or rough out an action beat you have not designed yet, a model that improvises is doing you a favour. Once you have boarded a scene and the creative decisions are made, the same habit turns every generation into an inspection pass: you have to find what changed before you can approve anything.
For small crews, pick the model by project stage, not by showreel. Exploration and concept work can tolerate invention. Previs that will be shown to a DP, a stunt coordinator or an investor needs a model that shoots the plan.
How to brief an AI video model like a DP
Give it a timed shot list, not a mood. Break the clip into beats with seconds attached, as the test did.
Specify every camera setup: size, angle, movement and where it starts and ends.
Lock the look with an opening frame before you ask for motion, so characters, palette and framing are fixed.
Number your storyboard panels and state the order explicitly.
Name what must not appear, especially if your model tends to add props or extra figures.
Budget for takes. On complex action, plan on roughly three generations per usable clip.
Review the periphery of the frame, not just the subject, before anything leaves the room.
What this comparison does not tell you
It is one set of briefs from one team, not a benchmark. Fifteen pairs, with four examined in detail.
The cost ratio reflects the specific endpoint the authors use, including a markup; they did not separate model price from route premium.
Seedance 2.5, which accepts up to thirty reference images and generates up to thirty seconds in one pass, was deliberately left out, so the results say nothing about long-form or reference-heavy work.
Self-hosting MiniMax H3 might change the economics further, but the authors say they have no credible estimate yet.
Before you choose a model, decide whether you want a collaborator who improvises or an operator who follows the boards. In this test, those were two different models.

















