# Mix audio tracks

Layer voices and music into one file, with the music dipping under speech.

Tool: Mix tracks (`mix_tracks`). Price: 1 credit a minute of input for whole files, charged only when a job succeeds.

## Free API

```bash
curl -F file=@voice.wav -F file=@music.mp3 -F 'tracks=[{"role":"voice"},{"role":"music","volume_db":-6}]' -F duck=true https://audo.ai/api/try/mix -o voice-mix.wav
```

Free for the first 3 minutes of each file, 3 times a day per person, with no key or account. Uploads up to 100 MB, 2 at once from one address. The response is the result itself; a job that runs longer than about 30 seconds returns 202 with a link to check on it. With an API key (`-H "Authorization: Bearer $AUDO_API_KEY"`), the same request processes the whole file with credits. Details: https://audo.ai/docs/api.md

## Options

Form fields with the same names as the MCP tool's inputs:

- `tracks`: The tracks, such as [{"file_id": "file_8Kc2QmP4", "role": "voice"}, {"file_id": "file_3Jd9RtY2", "role": "music", "volume_db": -6}]. 2 to 20.
- `align`: Lines up tracks recorded at the same time (the voice tracks, or every track when none has a role) by their sound, and corrects clock drift. Default false.
- `duck`: Lowers music tracks while a voice track speaks. Needs a voice and a music track. Default false.
- `duck_gap_lu`: With duck, how far under the speech the music sits, in LU (6 to 30), such as 12. Default 12.
- `loudness`: A loudness for the mix: none (the tracks as set, lowered only if it would clip), podcast (-16 LUFS), streaming (-14), or broadcast (-23). Default none.
- `output_format`: The mix's format: same (as the first track's audio), wav, mp3, m4a, or flac. Default same.

## Languages

Any language: it works on the sound, not the words.

## From an assistant (MCP)

Add https://audo.ai/mcp and sign in (setup for each assistant: https://audo.ai/setup.md). The tool is `mix_tracks`:

> Layers 2 to 20 audio files into one mix, each with its own volume, start time, and role (voice or music). align lines up tracks recorded at the same time by their sound and corrects clock drift; duck lowers music under speech. Use it for multitrack podcasts or a voice over a music bed. Not for playing files one after another (join_audio). Set free to true for the first 3 minutes free, 3 times a day; otherwise 1 credit a minute of the mix. Returns job_id; the result has each track's offset and drift, the ducking, and the loudness.

Example: "Use Audo to mix host.wav, guest.wav, and music.mp3, with the music lowered under the voices."

## More

- Every tool: https://audo.ai/llms.txt
- OpenAPI: https://audo.ai/api/openapi.json
- This page for people: https://audo.ai/mix-audio
