Ask a language model for "a script about the Roman Empire" and you get an essay: paragraphs, transitions, a conclusion. None of it is renderable. A short video script is a different artefact, and the difference is structural.
A script is a scene list, not prose
The unit of a short video is the scene: one visual, one line of narration. That constraint drives everything. Each scene needs a line short enough to be spoken while its image is on screen — generally 3-6 seconds, which is 7-14 words.
When the model returns prose instead, you have to break it into scenes yourself and discover that sentences do not divide cleanly at image boundaries. Asking for the scene structure up front avoids the whole problem.
Word count is a hard constraint, not a suggestion
Narration runs at roughly 2.2 words per second at a natural pace. So:
- 30-second short → about 65 words
- 45-second short → about 100 words
- 60-second short → about 130 words
Most first drafts come back two to three times too long. If the script is not written against a word budget, the narration overruns the visuals and the video either rushes or runs long.
Hook first, context later
Written prose builds context then delivers the payoff. Short video inverts this: the payoff goes first, and context earns its place only if the viewer stayed. A model told to "write engagingly" will default to the essay ordering — it has to be told explicitly that the surprising specific goes in line one.
Every scene needs an image prompt, not just narration
A scene is only half-written when you have the line. It also needs a description of what is on screen — and that description has to be renderable as a single still. "The decline of the Roman economy" is not an image. "A worn Roman coin, heavily clipped at the edges, on dark stone" is.
This is where a lot of AI video output looks generic: the visual prompts are abstractions, so the generator returns stock-looking filler.
Story type changes the structure
A listicle, a narrative and an explainer have genuinely different shapes. A listicle needs parallel beats and a countdown rhythm. A narrative needs escalation and a turn. An explainer builds one idea across scenes. Telling the model which shape to use produces markedly better structure than a generic "write a script" request.
What this looks like in practice
In AutoShortsX these constraints are the prompt: you pick a story type and a target length, and the model returns a scene list — each with narration and an image prompt, written against the word budget, hook first. Then every scene is editable, so you can rewrite one line without regenerating the whole video.