Noticias
Introducing Sonilo Sound Effects 1.0
A step toward a Sound World Model for video
- Escrito por
- Equipo de Sonilo
- Publicado
A generated sound has one pass-fail test: you believe it, or you don’t. A sports car either sounds like that engine, at that speed, from that distance, or the whole shot feels fake.
We built Sonilo Sound Effects 1.0 around that standard. It is another step toward a Sound World Model for video: a model that understands what is happening on screen, when it happens, how fast it moves, how far away it is, and what the moment should feel like.
Upload a video and Sonilo returns original music and sound effects, placed on the action, as one finished track. Or start with a text prompt and generate original sound effects directly. Every track is licensed.
Benchmarks
Leading the benchmarks
We evaluated Sonilo Sound Effects 1.0 and Sonilo Music 1.1 with audio-eval, an open benchmarking toolkit for generated audio. It scores each model on standard perceptual and distributional metrics, including CLAP similarity, Fréchet Distance, KL divergence, Audiobox Aesthetics, DeSync, and ImageBind alignment, across five established evaluation sets.
Sonilo Sound Effects 1.0 swept 15 of 15 metrics on CineBench, our internal video-to-sound-effects eval set, against Mirelo v1.6.
- 27 of 30
- video-to-sound-effects metrics led across Movie Gen Audio Bench and CineBench, versus Mirelo v1.6
- 16 of 24
- text-to-sound-effects metrics led across Clotho and AudioCaps, versus ElevenLabs SFX v2 and Mirelo v1.6
- 10 of 13
- MusicCaps metrics where Sonilo Music 1.1 leads Suno v5.5
Video-to-sound-effects benchmark
Movie Gen Audio Bench and CineBench
Sonilo Sound Effects 1.0 against Mirelo v1.6, scored on DeSync (how far the sound lands from the action), ImageBind alignment (the right sound for what is on screen), KL divergence, Inception Score, Audiobox Aesthetics, and Fréchet Distance.

Text-to-sound-effects benchmark
Clotho and AudioCaps
Sonilo Sound Effects 1.0 against ElevenLabs SFX v2 and Mirelo v1.6, on the same metric set plus CLAP similarity for prompt matching.

Text-to-music benchmark
MusicCaps
Sonilo Music 1.1 against Suno v5.5, on the same metric set used for text-to-music generation.

Capability 01
The video decides the placement, not your timeline
The decision we want to talk about is the one you can’t see in the demo. We could have shipped loose sound clips and left the placement to you. That version demos fine. It also leaves you searching libraries, cutting clips, and lining up sound by hand, which is the exact work we set out to remove.
So the model reads the footage before it generates anything: the timing, pacing, action, and emotion of every scene. It has to know when the engine revs before it makes the rev.
Capability 02
Music and effects, one finished track
Music works the same way. It follows the story instead of looping under it, and it comes from the same upload as the effects. One finished track, balanced, not a stack of clips left for you to mix.
The longer arc is a Sound World Model: one model that reads what happens on screen and generates soundtracks for it end to end. Music carries emotion. Sound Effects creates presence.
Tell us if you believe the engine.
Availability