Seedance 2
    Home
    • Text to Video

      Generate video from text prompts

    • Image to Video

      Transform static images into dynamic videos

    • Media to Video

      Multimodal (text, image, video, audio) video generation

    • Text to Image

      Generate images from text descriptions

    • Image to Image

      Transform and edit images with AI

  • My Creations
    Examples
    BlogPricing
Home
Text to Video
Image to Video
Media to Video
Text to Image
Image to Image
My Creations
TelegramDiscordUpgrade
Seedance 2 logoSeedance 2

High-quality AI video generation powered by Seedance 2. Turn a prompt, image, or simple script into consistent, well-paced, stable videos.

Product
GenerateSeedance 2.5Seedance 2.0 MiniExamplesPricing
Resources
UpdatesBlog
Company
About UsTerms of ServicePrivacy PolicySupport
© 2024 Seedance 2, All rights reserved
Privacy PolicyTerms of Service
Email
Home
Text to Video
Image to Video
Media to Video
Text to Image
Image to Image
My Creations
TelegramDiscordUpgrade
4 available

Upload Media *

Supports images, videos, and audio. Mix and match.

0 / 2500

Generate Audio

Generate audio for the video

Members only
Become a Member >

30-second extended storytelling with expanded multimodal references

A crystal ball remains perfectly locked in focus while eight worlds transform around it through beat-matched cinematic cuts.

Multimodal AI Video Generator — Create Video from Media

Create a reference-driven AI video from a combination of text, images, video clips, and audio. Use each source for a specific job: define identity and composition with images, guide motion with video, set pacing or sound with audio, and connect the direction with a prompt. This workflow is best when a single prompt or image is not enough to preserve a brand, character, product, performance, or editing rhythm.

How to create a video from multiple media references

  1. Step 1

    Add Your References

    Upload the images, video clips, or audio that should guide the generated result.

  2. Step 2

    Assign a Role to Each Asset

    Decide which reference controls identity, composition, motion, timing, or sound.

  3. Step 3

    Write the Connecting Prompt

    Describe how the references should work together in the final scene.

  4. Step 4

    Generate, Review, and Refine

    Preview the output, adjust the strongest reference or prompt constraint, and generate again.

Multimodal controls for reference-driven video

Multiple Reference Types

Combine text, images, video, and audio instead of forcing one input to carry the full brief.

Identity Guidance

Use visual references to help preserve a person, character, product, or brand direction.

Motion References

Guide camera language, performance, and scene movement with an example clip.

Audio Direction

Use sound as a pacing, mood, dialogue, or synchronization reference where supported.

Reference Weighting

Refine the result by strengthening the asset or instruction that matters most.

Controlled Iteration

Keep the brief stable while changing one reference or constraint at a time.

Media to Video FAQ

Choose the right AI video input workflow

WorkflowPrimary inputBest forMain control
Media to videoText, images, video, audioReference-led productionIdentity, motion, timing, and sound
Image to videoOne or more imagesAnimating existing visualsComposition and subject continuity
Text to videoWritten promptCreating a new sceneConcept, action, camera, and style

Input support and reference limits vary by model. The workspace shows the media types accepted by the selected model before generation.

Continue with a related Seedance workflow

Explore Seedance 2.5Review multimodal capabilities, model guidance, and example directions.Animate one imageUse a simpler workflow when one visual reference is enough.Compare pricingSee plan limits, credits, generation access, and commercial-use options.