

How to Write an AI Audio Brief for a 30-Second Scene
A convincing sound scene starts before anyone speaks. You need to know what the listener should notice first, what changes, and which sounds deserve space. Writing “cinematic audio with dialogue and music” leaves those decisions to chance. A brief with a timeline gives the model—and your own ears—a clearer target.
Start with five production decisions
Before opening a generator, write down the duration, setting, speakers, emotional turn, and delivery format. For a 30-second scene, one location and two voices are usually enough. Give each voice a role and describe the ambience as a continuous bed, then reserve only a few moments for foreground effects.
For example: a rainy teahouse at dusk; a customer and the owner; the customer begins impatient and leaves reassured. Rain stays outside. A sliding door, footsteps, and a porcelain cup are the only distinct effects. The dialogue must remain understandable over a restrained music bed.
SeedAudio prompt editor: keep the scene description, character direction, and sound cues together.
Give the scene a simple timeline
0–5 seconds: establish rain outside and a quiet interior. A door slides open; wet footsteps stop near the counter.5–12 seconds: the customer asks, “Is it too late?” Leave a beat before the reply.12–21 seconds: the owner sets down a cup and answers, “Only if you wanted the last train.” Let the cup sound land between the lines.21–30 seconds: the customer exhales, both voices soften, and the music lifts slightly before a clean ending.
The timing is a creative guide. Listen to the generated result and adjust it by ear; do not assume every sound will land on an exact frame in the current studio.
A prompt you can adapt
Generate a 30-second audio scene in a small teahouse at dusk. Two distinct adult voices: a hurried customer and a calm owner. Keep the dialogue clear and close. Soft rain remains outside throughout; the room itself is quiet. At the opening, a wooden sliding door moves and two wet footsteps approach. The customer asks, “Is it too late?” After a short pause, a porcelain cup touches the counter. The owner replies, “Only if you wanted the last train.” The customer exhales and the tension eases. Use a subtle, warm music bed that never masks speech. End naturally at 30 seconds. Avoid extra voices, dramatic impacts, and loud percussion.
Replace the setting, emotional turn, and dialogue with your own story. If you have reference audio, use only material you have permission to upload.
Revise one layer at a time
On the first listen, check three things: Can you understand every word? Does the cup sound happen at the dramatic turn? Does the music support the mood without announcing it? If speech is buried, lower or simplify the music direction. If the scene feels crowded, remove an effect. If the ending feels abrupt, shorten a line before asking for a longer fade.
Change one instruction per new take. That makes it easier to hear what actually improved.
Try it in the current studio
You can practice this scene-first workflow at SeedAudio 2.0 AI. The generator currently runs Seed Audio 1.0. Seed Audio 2.0 is presented on the site as an upcoming model with video-aware inputs, longer scenes, more reference audio, and separate dialogue, music, ambience, and effects stems. Those 2.0 capabilities are not yet live in this studio, so use the prompt above as a practical starting point and refine the audio you can hear now.
