Text, First Frame, and First-to-Last Frame
The official wan3.0-video model accepts a text-only brief, a first-frame image, or both first and last frames. Text-to-video is useful when the whole shot can be described in language. A first frame anchors subject appearance, composition, color, and camera position. First-and-last-frame guidance adds a defined destination, which is useful for reveals, transformations, product rotations, or scenes that must finish on a deliberate composition. In every mode, describe the motion between states instead of treating the prompt as a still-image caption.
