Multiple Reference Types
Combine text, images, video, and audio instead of forcing one input to carry the full brief.
Upload Media *
Supports images, videos, and audio. Mix and match.
Generate Audio
Generate audio for the video
A crystal ball remains perfectly locked in focus while eight worlds transform around it through beat-matched cinematic cuts.
Create a reference-driven AI video from a combination of text, images, video clips, and audio. Use each source for a specific job: define identity and composition with images, guide motion with video, set pacing or sound with audio, and connect the direction with a prompt. This workflow is best when a single prompt or image is not enough to preserve a brand, character, product, performance, or editing rhythm.
Step 1
Upload the images, video clips, or audio that should guide the generated result.
Step 2
Decide which reference controls identity, composition, motion, timing, or sound.
Step 3
Describe how the references should work together in the final scene.
Step 4
Preview the output, adjust the strongest reference or prompt constraint, and generate again.
Combine text, images, video, and audio instead of forcing one input to carry the full brief.
Use visual references to help preserve a person, character, product, or brand direction.
Guide camera language, performance, and scene movement with an example clip.
Use sound as a pacing, mood, dialogue, or synchronization reference where supported.
Refine the result by strengthening the asset or instruction that matters most.
Keep the brief stable while changing one reference or constraint at a time.
| Workflow | Primary input | Best for | Main control |
|---|---|---|---|
| Media to video | Text, images, video, audio | Reference-led production | Identity, motion, timing, and sound |
| Image to video | One or more images | Animating existing visuals | Composition and subject continuity |
| Text to video | Written prompt | Creating a new scene | Concept, action, camera, and style |
Input support and reference limits vary by model. The workspace shows the media types accepted by the selected model before generation.