Product

Text-to-Music API: Generate Licensed Music from a Prompt

Written by
Sonilo Team
Published
Text-to-Music API: Generate Licensed Music from a Prompt cover image

The Text-to-Music API generates original, licensed background music from a single text prompt — no source video required. Here's the full request format, response modes, and working code examples.

Quick answer

Sonilo's Text-to-Music API (POST /v1/text-to-music) generates original background music from a text prompt alone, with no source video required, unlike Sonilo's video-to-music endpoint. Send a prompt describing the mood, instrumentation, and tempo you want, for example "warm lo-fi piano with soft rain ambience, 80 BPM", along with a target duration, and the API returns fully licensed, commercial-safe music either as a live stream of audio chunks or as an async task you poll for a downloadable file. Output is licensed through Shutterstock, and in independent benchmarking on the open-source audio-eval toolkit, Sonilo Music 1.1 outperformed Suno v5.5 on 10 of 13 MusicCaps metrics.

FieldTypeNotes
promptstring, required1-2000 characters. Describes mood, instrumentation, and tempo, e.g. "warm lo-fi piano with soft rain ambience, 80 BPM".
durationinteger, requiredTarget length in seconds, 5-360.
segmentsstring, optionalJSON-encoded string of per-segment prompts, for tracks whose mood or instrumentation should change partway through.
modestring, optional"stream" (default) or "async". Stream returns audio chunks as an NDJSON event stream; async returns a task_id to poll.
output_formatstring, optional"m4a" (default) or "wav". WAV output requires mode=async.

Example request

The following example generates 30 seconds of upbeat corporate background music asynchronously and polls for the result:

curl -X POST https://api.sonilo.com/v1/text-to-music -H "Authorization: Bearer $SONILO_API_KEY" -F prompt="upbeat corporate background track, 90 BPM" -F duration=30 -F mode=async

Async mode returns 202 Accepted immediately with {"task_id": "...", "status": "processing"}. Poll GET /v1/tasks/{task_id} until the response status is "succeeded", at which point the result includes a music object with url, content_type, and file_size, or until it's "failed".

Omit mode (or set it to "stream", the default) to receive a 200 NDJSON event stream of audio chunks as they're generated, which suits low-latency playback. Async mode is the better fit for batch jobs, WAV output, or any workflow where polling for a finished file is preferable to handling a live stream.

curl -X POST https://api.sonilo.com/v1/text-to-music -H "Authorization: Bearer $SONILO_API_KEY" -F prompt="warm lo-fi piano with soft rain ambience, 80 BPM" -F duration=60

FAQ

What's the maximum duration for text-to-music?

360 seconds (6 minutes) per request, with a 5-second minimum. Use the optional segments field to script per-segment prompts if you need the mood or instrumentation to change over the course of a longer track.

Can I get WAV output instead of M4A?

Yes. Set output_format to "wav". WAV output is only available with mode=async; stream mode always returns M4A chunks.

Is the output licensed for commercial use?

Yes. All Text-to-Music output is fully licensed through Shutterstock and cleared for commercial use, so tracks can go into client work, ads, and monetized video without additional licensing.

Is there a Python or JavaScript SDK?

Yes. Install the official SDK with pip install sonilo (Python) or npm install sonilo (JavaScript/TypeScript) to call the endpoint without hand-building multipart requests.

How is this different from the video-to-music endpoint?

Text-to-Music generates standalone background music from a prompt alone, with no video input. Sonilo's video-to-music endpoint is video-aware: it analyzes an uploaded video and generates a soundtrack timed to its content. Use Text-to-Music when you need a track that isn't tied to a specific video.

Related Sonilo pages

  • Full API reference: https://platform.sonilo.com/docs/api/text-to-music
  • Pricing: https://sonilo.com/pricing
  • Open-source benchmark toolkit (audio-eval): https://github.com/shwj114114/audio_eval