Google · model workflow guide

Veo 3.1 Reference Image Guide

A reference image should tell the model what already exists. Your text should then explain what changes over time instead of redundantly redescribing every visible detail.

How GenVideoKit structures Veo 3.1 prompts

Use coherent natural-language direction with explicit temporal flow from start state to end state. Describe subject action, camera behavior, lighting and environmental motion. When relevant, integrate dialogue, ambience and sound effects because this model family supports native audiovisual generation. Avoid conflicting camera moves and keep continuity physically plausible.

Text → videoImage → videoNative audioFirst / last frameExtension

Production shortcut

Let the image define the starting appearance. Use the text to define motion, camera behavior and the elements that must remain locked.

Reference pricing snapshot
$0.4 / second

Good reference frame

  • Subject is readable and not heavily occluded.
  • Pose leaves room for the intended motion.
  • Important product or identity details are visible.
  • Lighting direction and environment are coherent.
  • The frame already resembles the intended starting composition.

Prompt what changes

Do not spend the prompt redescribing everything already visible. State the motion, camera move, final state and the details that must not drift.

Preserve the reference subject, wardrobe and room. The subject slowly stands, turns toward the window and takes one step forward. Camera makes a gentle lateral track only. Keep furniture geometry and light direction unchanged.