AI Image-to-Video Prompt Builder

Turn a still image into natural-looking AI video. Fill in each part of the formula and copy one fact-checked, ready-to-use prompt.

The secret: a good source image isn't enough โ€” a plain "make her move" style prompt is what makes AI video look stiff or unnatural. Tell the model the action, body movement, camera movement, environment movement, facial movement, and what should stay consistent, and it has far less left to guess.

How it works

Each of the six fields below covers one thing current AI image-to-video models actually need spelled out. Action and camera movement carry the most weight โ€” get those two specific and you're most of the way there. Body, environment, and facial movement are refinements: add them when a first pass looks flat, rather than by default. Consistency is different in kind from the other five โ€” it isn't motion at all, it's an instruction about what must not change, which is exactly why identity or outfit tends to drift without it.

One thing to skip: don't use any of these fields to redescribe what your source image already shows. The model can already see the subject, the setting, and the lighting โ€” repeating that back wastes prompt space that should be describing motion instead.

Click Load example to see the formula filled in, or Copy prompt once you've written your own.

FAQ

Why does my AI-generated video look stiff, weird, or unnatural?

Usually because the prompt only describes what the image looks like, not what should move. Image-to-video models can already see the subject and setting from your source image โ€” what they need from your prompt is the action, camera movement, and small ambient motion, described specifically rather than with a vague instruction like "make her move."

Do I need to fill in all six parts every time?

Not necessarily. The full formula gives you the most control, but current guidance across Runway, Kling, and Amazon Nova's prompting docs is to keep the actual motion cues โ€” action, camera, environment, face โ€” to one or two types for the cleanest results. Action and camera movement matter most; treat body, environment, and facial movement as detail to add if the first pass looks flat. Consistency isn't a motion cue at all โ€” it's a lock, telling the model what should not change.

How many camera movements should one clip use?

One. Pick a single movement โ€” dolly in, push in, tracking shot, slow orbit, zoom out, or static โ€” and describe it clearly. Stacking multiple camera movements in one short clip is a common cause of warped or inconsistent motion.

Do I need to redescribe what's already in my image?

No โ€” and doing so can actively hurt your result. The model already sees the subject, setting, and lighting from your source image. Spend your prompt describing what changes (the motion), not what's already visible. The one exception is the consistency section, which explicitly locks details in rather than redescribing them.

Why does my AI video's face or outfit change partway through the clip?

This is usually an identity-drift issue, not a motion issue. Add an explicit consistency instruction naming exactly what should stay identical โ€” facial identity, hairstyle, clothing, body proportions โ€” for the entire duration, rather than assuming the model will infer it from the source image alone.