Skip to content

Mixing guide

How to keep music under a voiceover

Music that fights a voice loses the viewer. Keep it under the words in three places: what the music is asked to be, whether the voices are kept, and how loud it sits while someone talks.

Ask for less, keep the voices, then let the music dip. A voiceover needs music that is sparse where the words are, a mix that keeps the dialogue in the finished video, and a level that drops while someone is talking. In Sonilo those are the direction, the Preserve speech switch and Auto Duck.
By
Sonilo Editorial Team
Published
Reading time
7 minutes
Sonilo Create interface with a music prompt and generated versions

Short answer

Write a direction such as "sparse, no lead melody, low under narration". Turn on Preserve speech before you generate, so the dialogue and narration stay under the new music. Leave Auto Duck on so the music lowers while someone speaks, then check the balance on the Music track before you export.

01

Write the brief for a voice

The loudest problem in a voiceover mix is usually the arrangement, not the level. A busy melody in the same range as a voice competes with it at any volume. So the brief starts with what the music should leave out.

  • No lead melody while someone talks. Pads, soft chords and a light pulse sit under words; a hook does not.
  • Steady energy. Sudden builds and drops pull attention from the sentence that lands on them.
  • Soft, low or high instruments. Felt piano, warm pads, soft plucks and light percussion leave the middle of the range, where speech lives, mostly free.
Direction

Soft minimal electronic with plucked synths and a gentle kick at 100 BPM, sparse, no lead melody, low in the mix under narration.

02

Keep the voices with Preserve speech

When you generate music for a video in Sonilo, the Preserve speech switch decides what happens to the voices already in the clip. It starts off.

  • On: Sonilo separates the dialogue and narration from the clip's audio and keeps them under the new music. The music is generated with that voice track as part of its input.
  • Off: the finished video has only the new music. The dialogue in the clip will not be in it.

For any video with a voiceover, turn it on before you generate. It costs no extra credits. On a phone the same switch is called Keep speech & vocals.

03

Let Auto Duck lower the music while someone talks

Ducking lowers one sound while another plays: here, the music while a voice speaks. In the Studio result, Auto Duck sits on the Music track and is on by default whenever speech was found in the mix. It lowers the music while someone is speaking and lets it back up in between, and the exported video keeps it.

Two things to know:

  • Auto Duck follows the Original track, the kept voices. If Original is muted, there is nothing to duck under.
  • Sound effects never duck. A whoosh or an impact stays at its own level.

The Music track also has its own level. 0 dB is the level the music was matched to, and the slider moves it 12 dB either way. A common starting point is to leave it at 0 with Auto Duck on, then lower it a few dB if the words still feel crowded.

04

Check the mix before export

Listen once at a low volume and once on a phone speaker. If you can follow every sentence on a phone at low volume, the mix will hold up almost everywhere.

  1. Play the busiest spoken section and the quietest one.
  2. If a word disappears, lower the Music track or ask for a sparser direction and generate again.
  3. Check the gaps between sentences: the music should come back up there, not stay buried.
  4. Export the video (MP4): it combines every track you can hear, with the music level and ducking you set.
For developers

The Sonilo API has the same idea as an endpoint: POST /v1/audio-ducking takes one voice input and one music input and returns a mix with the music lowered under the voice. It runs as a task you poll. See the API docs.

05

When the voiceover is not recorded yet

Sonilo can write the voice too. In Video translation, switch to Voiceover, write the script one line per cue, and choose whether each line keeps its natural length or is fitted to the clip. The voiceover replaces the clip's audio track, so make it first, then give the voiced video to Video to Music with Preserve speech on. Doing it in that order means the music is generated around the final words.

06

Three videos, three directions

Tutorial with constant narration

Gentle lo-fi beat with soft keys, quiet drums and warm bass, kept low and even under a voice that talks the whole time.

Interview documentary

Sparse felt piano with slow string pads and long silences, under a reflective interview.

Product walkthrough

Clean corporate tech music with plucked synths, airy pads and a light four on the floor kick, with space for a voiceover.

Podcast video

Warm lo-fi bed at 80 BPM with soft keys and a muted kick, very low and even under a two person conversation.

More in the tutorial and podcast groups of the prompt library.

FAQ

Frequently asked questions

Music under the words

Generate music for your narrated video

Upload the video with its voiceover, turn on Preserve speech, and compare two versions with Auto Duck on.