Generation modes and drafting
Text-to-video, image-to-video, multi-reference, first/last frame, and extend do not all need the same kind of information.
| Mode | What the prompt should prioritize | Common failure |
|---|---|---|
| Text to video | Complete but compact shot layers | Too many events packed into one clip |
| Image to video | Preserve @Image1 and only add motion plus necessary light change | Re-describing the still until face or product drifts |
| Video to video | Transfer motion, camera, and rhythm | Copying a likeness that should not be copied |
| Multi-reference | One clear role per asset | One image controlling identity, pose, and style at once |
| First / last frame | A continuous path from the first frame to the last | Writing the last frame as a vague mood instead of an end state |
| Extend | Continue only from an accepted real ending | Blind-extending the original prompt |
A common fidelity line for image-to-video
@Image1 is the product reference; preserve label, logo, shape, and color exactly. Only motion changes: …