A director's recipe for Sora 2, Veo 3, Kling and Runway. The seven parts of a video prompt, a camera-move cheat sheet, and free copy-paste examples that actually move.
Think of a video prompt as a shot note. A photo prompt describes one frozen moment. A video prompt has to describe change over time, so it needs two ingredients a still image never uses: motion and a camera. Leave those out and the model guesses, and its guesses are usually stiff and boring. Here are the seven parts that turn a vague idea into a shot the model can actually film.
Here is the same idea written two ways. First as a labeled checklist, then folded into the natural sentence you would actually send.
Prompt recipe[Subject] a lone astronaut in a weathered white suit
[Action] slowly turns to face the camera, then raises one hand
[Camera] slow dolly-in from a low angle
[Lighting] harsh side light from a distant sun, deep shadows
[Mood] tense, awe-struck, cinematic
[Style] shot on 35mm film, shallow depth of field
[Length] 6 seconds, one unbroken shot
Same prompt, ready to send
A lone astronaut in a weathered white suit slowly turns to
face the camera and raises one hand. Slow dolly-in from a low
angle. Harsh side light from a distant sun casts deep shadows.
Tense, awe-struck, cinematic. Shot on 35mm film with shallow
depth of field. One unbroken 6-second shot.
You do not need the brackets when you send it. They are a checklist to make sure nothing is missing. Most models read best when you fold the seven parts into one or two natural sentences, in roughly that order.
The single biggest upgrade to any video prompt is naming the camera move with real film words. "Cool angle" tells the model nothing. "Slow dolly-in from a low angle" tells it exactly what to do. Here is the vocabulary worth memorising.
| Camera move | What the camera does | Reach for it when |
|---|---|---|
| Dolly in / out | The whole camera glides toward or away from the subject | Building intimacy, or revealing the space around someone |
| Tracking / follow | Camera travels alongside a moving subject | Walking, running, chases, any forward action |
| Pan | Camera pivots left or right from a fixed spot | Revealing a scene, or following motion across the frame |
| Tilt | Camera pivots up or down from a fixed spot | Revealing height, looking up at something huge |
| Crane / boom | Camera rises or lowers through the air | Epic reveals and sweeping establishing shots |
| Orbit / arc | Camera circles around the subject | Hero moments and showing all sides of a thing |
| Handheld | Subtle natural shake, a human holding the camera | Documentary feel, tension, raw realism |
| Push in | A slow, steady dolly straight toward the subject | Rising emotion, a dawning realization |
| Zoom | The lens magnifies without the camera moving | Quick emphasis (note: this differs from a dolly) |
| Static / locked off | No movement at all, a fixed frame | Calm, composed shots that let the action carry itself |
| Whip pan | A very fast pan that blurs the frame | Energetic transitions between two moments |
Rule of thumb: one camera move per clip. If you want a dolly-in and then an orbit, that is two shots. Generate two clips and cut them together, rather than asking one short clip to do both.
Video models live and die on verbs. If your prompt reads like a photo caption ("a red sports car on a mountain road"), the model has no idea what should happen, so it barely animates. Give it a verb and a direction ("a red sports car speeds around the bend, spraying gravel") and it comes alive. Describe how things move, not only what they are.
Pacing is the second half. Speed words steer the whole feel of the shot. "Slowly", "gradually" and "drifting" give calm, floaty motion; "suddenly", "in a burst" and "snaps" give sharp, energetic motion. You can also call the tempo directly with "slow motion" or "real-time speed". Keep it to one main action per short clip. A five-second shot that tries to have a person sit, stand, wave, and walk out will look rushed and broken. Pick the one beat that matters and let it breathe.
Weak vs strong motionWeak: a waterfall in a forest, beautiful, 4k
Strong: water tumbles over mossy rocks in slow motion, fine
mist drifting up into a shaft of morning light, ferns
swaying gently at the edges. Camera slowly pushes in.
The seven-part recipe works everywhere, but each model has a personality. Lean into its strength and your hit rate jumps.
| Model | Signature strength | Prompt tip |
|---|---|---|
| Sora 2 | Physical realism and synchronized audio | Describe the sound and the physics out loud; keep to one clear action so the realism has room to shine |
| Veo 3 | Native audio, including spoken dialogue | Put dialogue in quotes and name the speaker; add an ambient-sound line for the room |
| Kling | Smooth motion and strong image-to-video | Feed it a start image and describe only the motion to add, not the whole scene |
| Runway Gen-4 | Consistent characters and creative control | Use reference images and keep the character description identical word-for-word across shots |
Because Sora 2 and Veo 3 generate audio along with the picture, silence in your prompt is a wasted opportunity. Name the sound effects, the ambience, and any spoken lines. Veo 3 in particular is strong at lip-synced dialogue, so give it real lines. Put each line in quotes and label who says it, then add a line for the background sound.
Veo 3 dialogue exampleTwo friends sit in a diner booth at night, warm neon glow
through the rain-streaked window. Handheld, shallow focus.
Woman (smiling): "You actually did it."
Man (laughing): "I told you I would."
Ambient: quiet diner chatter, rain on glass, a coffee cup
clinks down. Warm, hopeful mood. 8 seconds.
Kling is one of the best at taking a still image and moving it convincingly, so it rewards image-to-video prompts (see the next section). Runway Gen-4 is built for creators stitching several shots together, so its superpower is keeping the same character and world across clips. If you are making a sequence in Runway, copy the exact character description into every prompt so the model does not redraw a new face each time.
Image-to-video is the most reliable way to get a specific look. You make a strong still first (in Midjourney, FLUX, Nano Banana, whatever you like), then hand it to the video model as the opening frame. The rules flip here: because the model can already see the scene, you should not re-describe it. Describe only what should move and how the camera should travel. Re-listing the subject just confuses the model and can make it redraw things you already got right.
Image-to-video example (Kling)[Start image: a knight in silver armor standing in a snowy field]
Snow begins to fall, the knight's cape flutters in a slow wind,
she turns her head to look off toward the horizon. Camera slowly
orbits to the left. Everything else stays still. Realistic motion,
5 seconds.
Notice the prompt never says "silver armor" or "snowy field" again. The image carries all of that. Some models, Kling included, also let you set a start frame and an end frame; when you have both, describe the journey between them ("the door swings from closed to fully open") instead of the objects.
The fastest way to learn the rhythm is to start from prompts that already work, then tweak them. We keep a free, copy-paste library for each of the big video models:
Hundreds of shot-ready video and image prompts, organised by model and use case, so you skip the blank box and start from something that already looks pro.
Open The Vault → Rate my prompt