Short answer: Treat an audio file as a timing and atmosphere reference, then describe the visual action that should land on each important beat. Audio can guide rhythm, pauses, and sound continuity, but it does not replace a visual prompt for the subject, camera, setting, or final frame.
Audio reference versus sound effects
An audio reference has a job beyond “add music.” It can provide a beat pattern, a spoken cadence, a room tone, or a transition cue that helps structure the clip. A sound-effect prompt, by contrast, usually asks for a particular sound inside an already planned scene.
| Need | What to provide | What the visual prompt still needs |
|---|---|---|
| Beat-led movement | a track or rhythm with clear accents | subject, action, camera, and which accents matter |
| Dialogue timing | a clean voice track or timing guide | speaker, eyeline, mouth movement, and shot boundaries |
| Atmosphere | room tone, rain, traffic, or natural ambience | location, lighting, and visible environmental motion |
| Transition cue | a rise, hit, pause, or impact | the exact visual change that happens at the cue |
Plan the audio-to-video handoff
Before writing a prompt, mark only the beats that affect the picture. For a short product shot, three markers are enough:
- Opening: the product is already visible and the camera is stable.
- Accent: one camera move or product action lands on the first strong beat.
- End: the motion eases to a hold while the audio resolves.
This avoids asking a short clip to follow every tiny waveform change. If the source audio is complex, use the main pulse and describe the rest as atmosphere.
Open the reference workflow
Open Media to Video, select Seedance 2.5, and choose Standard reference. Upload your audio and image, use the asset labels shown in the interface, then review the output settings and estimated credits before generating.
The available controls can depend on the selected model, mode, account, and plan. The Seedance 2.5 model overview explains the broader model capabilities, but the creator interface is the final check for the inputs and settings available to your account.
Self-contained prompt example
This is an illustrative prompt, not a tested result. Prepare an audio excerpt with one clear accent and a sustained ending, plus a reference image of a travel mug. Attach the same references again for each new generation and use the actual asset labels shown in the creator center; @Audio1 and @Image1 below are example labels.
Use @Audio1 as a timing and atmosphere reference, and use @Image1 for the
travel mug's appearance. Do not copy any unrelated sound source or visual from
another reference.
Create a short product-demo video on a light wooden table beside a window. Keep
the mug's shape, handle, lid, color, and printed markings consistent with
@Image1. At the opening, hold a stable medium shot. On the first clear accent in
@Audio1, make one slow left-to-right camera move while the mug remains still.
Let the camera ease to a stop during the sustained ending and hold the mug in
frame. Keep the soft morning-light mood. Preserve any markings already visible
in @Image1, but add no new text or logos. Use @Audio1 to guide pacing and mood;
exact reproduction of the source audio is not guaranteed. No hands, extra
products, zoom, orbit, or cuts.
The prompt assigns each input one role: audio guides timing and mood, while the image guides product appearance. That boundary makes a failure easier to diagnose.
Three common failure modes
The video follows the mood but misses the beat
Use fewer timing cues and identify the exact beat that matters: “on the first clear accent” is more actionable than “match every beat.” State whether the camera, subject, or edit should change at that point.
The audio is present but the picture has no matching action
Name the visible event: a hand enters, a product turns, a camera reaches a close-up, or a subject stops. Avoid relying on “make it dynamic” as a visual instruction.
The subject changes when the rhythm changes
Keep identity and movement separate. Attach the product or character reference, state the features that must remain stable, and use the audio only for timing. If a reference image matters, reattach it rather than assuming the system remembers a previous generation.
Review the result in an external editor
Import the generated video and the source audio into an external editor that displays an audio waveform. Compare the intended cue with the visible action.
An audio reference does not mean the source audio track will be preserved exactly in the generated clip. If you need to keep the original music or narration, align the final video and audio in the external editor. Check:
- does the named action land near the intended accent;
- does the camera move in one consistent direction and stop at the stated point;
- does the subject remain recognizable while the rhythm changes;
- does the ending give the editor a usable hold;
- are music, speech, room tone, and effects distinct enough to review;
- did the model add text, logos, people, or unrelated objects?
If the timing is close but not exact, adjust one cue or align the audio and video externally. Do not add several new camera moves and more reference files at the same time. External alignment can make the final cut precise, but it does not prove that the generation itself followed the audio.
For visual prompt structure, use the Seedance 2.5 prompt examples. For diagnosing unwanted sound or music, read Seedance 2.5 prompt troubleshooting.
FAQ
Can audio alone create the visual story?
An audio reference alone does not reliably specify the visual story you want. Add a visual prompt to define the subject, setting, action, camera and ending.
How many beats should I specify?
Start with one to three meaningful markers for a short clip. More markers can make the brief harder to follow, especially when the audio has dense percussion.
Can I use a song as a reference?
Only use audio you have the right to use. Also check the service's current terms and your plan before publishing a commercial result.
What if the timing is slightly off?
Review the output in an external editor, change one timing instruction, and try again. If exact synchronization is required, align the generated video and audio there rather than promising frame-perfect generation.