Multiple Reference Types
Combine text, images, video, and audio instead of forcing one input to carry the full brief.
Task mode
Not supported by this modelCreate a new video from reference media
Upload Media *
Supports images, videos, and audio. Mix and match.
Single clips now run up to 30 seconds and can be extended twice, creating complete and compelling stories. Motion is more stable and visuals feel more lifelike.
Prompt:An old neighborhood barbershop is about to close. Outside, the sky has grown dark and the street is tinted blue, while warm yellow light fills the shop. Old photographs, vintage barber certificates, and faded customer portraits hang on the walls. Characters: a barber and a customer whose relationship feels natural and familiar, like longtime neighborhood friends. Avoid exaggerated acting. …[approximately 1,250 characters omitted]
Create a reference-driven AI video from a combination of text, images, video clips, and audio. Use each source for a specific job: define identity and composition with images, guide motion with video, set pacing or sound with audio, and connect the direction with a prompt. This workflow is best when a single prompt or image is not enough to preserve a brand, character, product, performance, or editing rhythm.
Step 1
Upload the images, video clips, or audio that should guide the generated result.
Step 2
Decide which reference controls identity, composition, motion, timing, or sound.
Step 3
Describe how the references should work together in the final scene.
Step 4
Preview the output, adjust the strongest reference or prompt constraint, and generate again.
Combine text, images, video, and audio instead of forcing one input to carry the full brief.
Use visual references to help preserve a person, character, product, or brand direction.
Guide camera language, performance, and scene movement with an example clip.
Use sound as a pacing, mood, dialogue, or synchronization reference where supported.
Refine the result by strengthening the asset or instruction that matters most.
Keep the brief stable while changing one reference or constraint at a time.
| Workflow | Primary input | Best for | Main control |
|---|---|---|---|
| Media to video | Text, images, video, audio | Reference-led production | Identity, motion, timing, and sound |
| Image to video | One or more images | Animating existing visuals | Composition and subject continuity |
| Text to video | Written prompt | Creating a new scene | Concept, action, camera, and style |
Input support and reference limits vary by model. The workspace shows the media types accepted by the selected model before generation.