
We Ran the Same 12 Prompts Through 6 Video Models: Here Is Exactly Where Each AI Video Generator Breaks
Model comparisons in this space are mostly just marketing. A vendor publishes its best three clips, a reviewer publishes a highlight reel, and nobody shows the render where the subject grew a sixth finger or the poured coffee flowed upward. That leaves anyone choosing a tool to guess at the thing that actually determines cost, which is not how good the best output looks but how often the average output is unusable.
So we built a stress test designed to fail. Twelve prompts, each targeting a known weak point in video diffusion, run identically through six leading models. Twelve generations per prompt per model, scored blind. That is 864 clips, and roughly 400 were unusable by any reasonable standard.
What follows is the failure map. It is not a ranking, because no model won outright. It is a guide to which problems you can prompt around and which to solve in the edit.
The twelve prompts each targeted one documented weakness: hands manipulating an object, water interacting with a solid, a dense crowd, a mirror reflection, legible in-frame text, a collision, spoken dialogue, a two-shot continuity match, a falling object, an animal gait cycle, a subject passing behind an obstacle, and a scale relationship between two everyday objects. We didn't choose anything to flatter a model, and we didn't choose anything impossible.
Each prompt was written once and never adapted per model, which matters because vendors optimize for their own syntax, and adapting prompts would have measured our tuning rather than their capability. Every prompt was five to eight seconds, 16:9, at the highest resolution each system offered without additional cost.
Three reviewers scored each clip blind on a simple usable or unusable basis, plus a written note on the failure mode. Usable meant a working editor could cut the clip into a paid project without an apology. Disagreements went to a fourth reviewer. We ran generations through the ImagineArt AI Video Generator wherever a model was available there, which removed interface differences as a confounding variable and kept queue times comparable.
We are naming failure categories rather than vendors throughout. Model versions turn over every few weeks, and a named leaderboard would be wrong by the time you read it. The failure patterns, by contrast, have been stable for over a year.
The oldest problem is no longer the worst one, but it has not gone away. Across six models, a close-up of hands performing a deliberate task, tying a shoelace, produced a usable clip 41 percent of the time. Static hands are now basically solved. Hands that manipulate an object are not, because the model must track occlusion between fingers frame to frame.
The reliable workaround is distance. The same action framed as a medium shot rather than a close-up jumped to 79 percent usable. If the hands must be close, keep the clip under four seconds and avoid crossing fingers over one another. Gloves helped more than we expected, and so did any action where the object stays visible rather than disappearing into a grip.
Fluid simulation was the biggest positive surprise. Pouring, splashing, and steam all rendered convincingly in five of six models, with an 83 percent usable rate. Water in motion has enough training data and enough visual noise that small errors read as natural variation.
The exception is water interacting with a solid in a specific way. A hand dipping into a bowl produced correct ripples but frequently lost the waterline on the wrist. Steam rising from a cup was excellent. Steam that had to obscure and then reveal a face was not.

This is the worst category in the test, and it is barely discussed. A street scene with fifteen or more visible people produced a usable clip 22 percent of the time. Background figures merge, walk through one another, gain and lose limbs, and reverse direction mid-stride. The foreground subject is usually fine, which is exactly why the failures slip past a quick review.
The workaround is population control. Specify a number, name the depth of field, and keep the background soft: "six people in the background, shallow depth of field, background heavily blurred." That pushed the usable rate to 64 percent. Nothing we tried made a sharp, dense crowd reliable in any AI video generator we tested.
Reflections failed in a consistent and instructive way. A subject walking past a shop window produced a reflection in every model, which is impressive, but the reflection matched the subject's actual motion only 35 percent of the time. Reflected figures lagged, faced the wrong direction, or showed different clothing.
Mirrors were worse than windows, because a mirror puts the error in the center of the frame at full contrast. If your shot needs a mirror, frame it so the reflection is partial or angled away. Full-face mirror shots remain a coin flip.
Every model can now render short text that looks like text. Legibility across a five-second clip is another matter. A shop sign reading "OPEN" held its spelling for the full clip in 58 percent of generations. Longer strings collapsed quickly, and any text on a moving surface degraded within two seconds.
Treat in-frame text as a compositing job. Generate the plate without the sign, then add the type in an editor. Every hour spent re-rolling a prompt to fix a misspelled sign is an hour that compositing would have solved on the first attempt.
Speed is handled better than expected, but impact is not. A sprinter crossing the frame was usable 71 percent of the time. A ball striking a wall and rebounding was usable 29 percent of the time, because the model must decide on a moment of contact and reverse a vector precisely.
Anything with a collision, a catch, a punch, or a bounce should be generated as two clips cut at the moment of impact. That is standard practice in live action too, and it turns the hardest frame in the shot into an edit point.
Models with native audio have changed this category entirely, and it is now the clearest point of separation between systems. Two of six produced convincing single-speaker dialogue with accurate lip sync at 74 percent usable. The others produced mouth movement unconnected to any specific sound.
Two-speaker scenes remain unsolved everywhere. Turn-taking broke in almost every attempt, with both characters speaking at once or one continuing to move their mouth in silence. Generate conversations one speaker at a time and cut between them. Shorter lines also scored far better than longer ones, and anything past roughly fifteen spoken words drifted out of sync before the clip ended.
We tested whether two clips generated from the same description would match well enough to cut together. They did 33 percent of the time. Wardrobe details drifted most, followed by hair, then lighting direction.
Reference images fixed much of this. Supplying the same still to both generations raised the match rate to 68 percent. Anyone assembling multi-shot sequences should treat a reference image as mandatory rather than optional, and should check a Free Online Video Generator for reference support before committing to it.
Objects at rest behave correctly. Objects in transition frequently do not. A falling book landed correctly in 44 percent of generations; the rest saw it drift, hover, or pass partly into the floor. Heavy objects lifted by a person were worse, because the model rarely shows the strain that sells the weight.
Describe the physical consequence rather than the physics: "her shoulders drop as she takes the weight" outperformed any description of mass or gravity by a wide margin.
Quadruped movement is unreliable across the board. A trotting dog was usable 38 percent of the time, with leg count and gait phase the common failures. Cats fared better than dogs; birds in flight were poor, and horses were the worst of everything we tested.
Slower gaits work better. A walking animal beat a running one in every model, and a seated or standing animal was nearly always fine.
When a subject passes behind an object and comes back out, it often comes back changed. Our test had a woman walk behind a pillar and reappear. She reappeared as the same person 47 percent of the time. Shorter occlusions were safer, and anything over one second was close to a re-roll.

This is the failure that most often survives casual review, because each half of the clip looks correct in isolation. Watch the whole clip before you approve it.
Relative size is handled loosely. A coffee cup next to a laptop was correctly proportioned in 61 percent of generations, and the errors were subtle enough to pass unless the shot lingered. Naming both objects with an explicit size relationship in the prompt helped a little, but no AI video generator we tested held scale reliably across a camera move, and the drift always grew as the move continued.
Averaged across all twelve categories, the best model reached 61 percent usable, and the worst reached 44 percent. That spread is narrower than the marketing implies, and the more useful finding is that the models fail in different places. One system led on dialogue and trailed badly on crowds. Another handled physics well and produced the weakest hands.
The practical consequence is that a single model is the wrong unit of choice. You want access to several through one interface, which is why we ran most of this test inside a Free Online Video Generator rather than maintaining six separate accounts, and why ImagineArt was our working environment throughout.
It also means the question worth asking a vendor is not which model is best but which shots the model is best at. A system that is unbeatable at dialogue and mediocre at crowds is an excellent choice for a talking-head campaign and a poor one for a busy street scene. Match the tool to the shot list rather than to the leaderboard.
Four rules came out of this test. Keep clips short, because failure rates climb sharply after six seconds in every category. Move the camera back, because distance hides the errors that close-ups expose. Cut at the hard moments, turning impacts, occlusions, and speaker changes into edit points instead of generation problems. And composite anything precise, especially text and logos, rather than asking a model to render it. A fifth rule emerged late: review at full speed and at quarter speed, because the errors that survive a fast viewing are the ones that clients notice in the final cut.
None of these are workarounds you outgrow. They are how the work is done today, and the crews producing consistent output at volume already treat them as defaults rather than concessions.
The honest summary is that roughly half of what you generate will be unusable, and the skill worth developing is designing shots that fail cheaply. Choose the prompts that avoid the weak categories, budget for re-rolls in the ones you cannot avoid, and pick an AI video generator that gives you more than one model to fall back on when a shot refuses to cooperate.
Run the twelve prompts yourself. Any Free Online Video Generator will let you reproduce this test in an afternoon, and your own failure map will be more useful than ours, because it will be built on the shots you actually need. ImagineArt keeps the whole comparison in one place, which is the only reason a test this size was practical for us at all.