Back to blog
Prompt engineeringJul 31, 2026 · 6 min read

Prompt Engineering for Video: The Words That Actually Move the Needle

Copy a prompt formula from an image generator into a video tool and it mostly falls flat. Video models weigh different signals, respond to different words, and punish a few habits that image prompting actively rewards.

Subject, action, setting, camera, in that order

A strong video prompt answers four questions, roughly in this order: who or what is in frame, what are they doing, where is it happening, and how is the camera behaving. Skip the order and a model tends to weight whatever comes first. Neon-lit alley, a fox darts between puddles, low tracking shot reads differently to the model than a fox darts between puddles in a neon-lit alley, low tracking shot. Same words, different emphasis in the result.

Subject

a fox

Action

darts between puddles

Setting

neon-lit alley

Camera

low tracking shot

Style words are a garnish, not the meal

8K, trending on artstation, masterpiece come from image-generation communities and carry almost no weight in a video model, since they were shorthand for still-image aesthetics tagged across art platforms. Video models train mostly on real footage and film, so the style cues that land are actual techniques: shot on 16mm, overexposed daylight, shallow depth of field. Name a real visual choice instead of a superlative.

"8K, trending on artstation, masterpiece"Before

"8K, trending on artstation, masterpiece"

"shot on 16mm, shallow depth of field"After

"shot on 16mm, shallow depth of field"

Verbs carry more weight than adjectives

Swap an adjective for a stronger verb and the output usually shifts more than swapping five adjectives at once. Beautiful ocean does very little. Ocean crashes over the rock and drags foam back with it gives the model a specific motion to render, and everything about pacing and camera timing organizes around that verb. When a shot feels flat, the fix is rarely a better adjective. It’s a more specific verb.

Skip this

a beautiful, dramatic, epic ocean

Try this

the ocean crashes over the rock and drags foam back with it

Borrow the vocabulary a cinematographer actually uses

Cinematic is a review-site word, not a production word. Real film vocabulary is more specific, and models respond to it more reliably: 35mm for a wider, slightly distorted closeness, 85mm for flattering compression on a face, rack focus for the background sliding into focus as the foreground blurs, Dutch angle for a tilted, off-balance frame, rim light for a thin edge of light outlining a subject against a darker background. None of it requires a film degree. It’s just more specific than the adjectives most people reach for first.

The establishing beat

Prompts that open mid-action sometimes render a first frame that doesn’t match what the rest of the clip is doing, since the model has to guess a starting pose from a description that’s already in motion. Give it half a beat of stillness first. A runner crouched at the starting line, then bursts forward at the gun tends to hold together better than a runner bursts forward at the gun alone, because the model gets a clear resting state to launch the motion from instead of inventing one.

Test one variable, not five

Change the lighting, the camera move, and the wardrobe in the same prompt, and there’s no way to know which change caused the shift in the output. Hold a shot mostly fixed and change one thing between generations: lighting only, then camera only, then wardrobe only. It takes a few extra tries and builds a real sense of which words in your vocabulary carry weight.

Negative instructions rarely do what you expect

Typing no text, no watermark, no blur doesn’t remove those things, and can even make them more likely to show up, since the words are still sitting in the prompt for the model to latch onto. Describe the frame you want instead. Swap no crowd in the background for empty street, one person crossing.

Match length to complexity

One clear action in a simple setting needs one sentence. A specific subject, a distinct action, a defined setting, and a camera direction need more, but padding a prompt with extra adjectives past that point gets ignored or diluted. Long enough to specify what has to happen, short enough that no word is dead weight.

Say it, and Zo makes it.

Chat with Zo