Short answer
Write a direction such as "sparse, no lead melody, low under narration". Turn on Preserve speech before you generate, so the dialogue and narration stay under the new music. Leave Auto Duck on so the music lowers while someone speaks, then check the balance on the Music track before you export.
01
Write the brief for a voice
The loudest problem in a voiceover mix is usually the arrangement, not the level. A busy melody in the same range as a voice competes with it at any volume. So the brief starts with what the music should leave out.
- No lead melody while someone talks. Pads, soft chords and a light pulse sit under words; a hook does not.
- Steady energy. Sudden builds and drops pull attention from the sentence that lands on them.
- Soft, low or high instruments. Felt piano, warm pads, soft plucks and light percussion leave the middle of the range, where speech lives, mostly free.
Soft minimal electronic with plucked synths and a gentle kick at 100 BPM, sparse, no lead melody, low in the mix under narration.
02
Keep the voices with Preserve speech
When you generate music for a video in Sonilo, the Preserve speech switch decides what happens to the voices already in the clip. It starts off.
- On: Sonilo separates the dialogue and narration from the clip's audio and keeps them under the new music. The music is generated with that voice track as part of its input.
- Off: the finished video has only the new music. The dialogue in the clip will not be in it.
For any video with a voiceover, turn it on before you generate. It costs no extra credits. On a phone the same switch is called Keep speech & vocals.
03
Let Auto Duck lower the music while someone talks
Ducking lowers one sound while another plays: here, the music while a voice speaks. In the Studio result, Auto Duck sits on the Music track and is on by default whenever speech was found in the mix. It lowers the music while someone is speaking and lets it back up in between, and the exported video keeps it.
Two things to know:
- Auto Duck follows the Original track, the kept voices. If Original is muted, there is nothing to duck under.
- Sound effects never duck. A whoosh or an impact stays at its own level.
The Music track also has its own level. 0 dB is the level the music was matched to, and the slider moves it 12 dB either way. A common starting point is to leave it at 0 with Auto Duck on, then lower it a few dB if the words still feel crowded.
04
Check the mix before export
Listen once at a low volume and once on a phone speaker. If you can follow every sentence on a phone at low volume, the mix will hold up almost everywhere.
- Play the busiest spoken section and the quietest one.
- If a word disappears, lower the Music track or ask for a sparser direction and generate again.
- Check the gaps between sentences: the music should come back up there, not stay buried.
- Export the video (MP4): it combines every track you can hear, with the music level and ducking you set.
The Sonilo API has the same idea as an endpoint: POST /v1/audio-ducking takes one voice input and one music input and returns a mix with the music lowered under the voice. It runs as a task you poll. See the API docs.
05
When the voiceover is not recorded yet
Sonilo can write the voice too. In Video translation, switch to Voiceover, write the script one line per cue, and choose whether each line keeps its natural length or is fitted to the clip. The voiceover replaces the clip's audio track, so make it first, then give the voiced video to Video to Music with Preserve speech on. Doing it in that order means the music is generated around the final words.
06
Three videos, three directions
Tutorial with constant narration
Gentle lo-fi beat with soft keys, quiet drums and warm bass, kept low and even under a voice that talks the whole time.
Interview documentary
Sparse felt piano with slow string pads and long silences, under a reflective interview.
Product walkthrough
Clean corporate tech music with plucked synths, airy pads and a light four on the floor kick, with space for a voiceover.
Podcast video
Warm lo-fi bed at 80 BPM with soft keys and a muted kick, very low and even under a two person conversation.
More in the tutorial and podcast groups of the prompt library.
FAQ
Frequently asked questions
Related
Keep reading
Music under the words
