# Create an SRT file from video or audio

Turn your recording into an SRT file with timed captions.

## What's inside an SRT file

An SRT file is a list of captions in plain text. Each one has a number, a start and an end time, and the words to show. The example above is a scientist describing a mammoth trackway outdoors, and its SRT file is exactly what came back: four captions, each with its number and times.

A minute of audio is usually ready in about 7 seconds. Start from a clear recording for the best captions; for a noisy one, run it through Remove noise first. When several people talk, no caption mixes two of them, and the SRT names no one, so it stays clean on screen. You also get VTT captions, the plain text, and the time of every word.

Tool: Transcribe (`transcribe_audio`). The same tool's other pages: https://audo.ai/transcribe.md. Price: 2 credits a minute of input for whole files, charged only when a job succeeds.

## Free API

```bash
curl -F file=@vlog.mp4 -F format=srt https://audo.ai/api/try/transcribe -D response.headers -o response.body
```

Read response.headers and response.body before naming media. Handle 202, 200 JSON and errors with the complete response client at https://audo.ai/docs/api#responses.

Free for the first 2 minutes of each file, 3 times a day per person, with no key or account. Uploads up to 100 MB, 2 at once from one address. The response is the result itself; a job that runs longer than about 30 seconds returns 202 with a link to check on it. With an API key (`-H "Authorization: Bearer $AUDO_API_KEY"`), omitting free runs the whole file with credits if the balance covers it, otherwise the free portion if quota remains. Set free=true to prohibit spending; for paid whole files use free=false with an approved positive max_credits. A free result cut short comes with finishUrl (and the Audo-Finish-Url header when the response is the file itself): give it to the person, who pays there (from $5, no sign-up), and the whole file runs; keep checking the status link, whose finishedBy names the whole-file job and its own status link. Details: https://audo.ai/docs/api.md

## Options

Form fields with the same names as the MCP tool's inputs:

- `language`: The spoken language, as a code such as en or ja. Default: detected.
- `verbatim`: Keeps filler words such as um and uh, so they can be cut later. Default true.
- `speaker_labels`: Labels who speaks when: Speaker 1, Speaker 2, and so on, in the order they first speak, up to 8 people. No names. Default true.
- `format`: what the free API returns: json (default), srt, vtt, or txt.

## Speakers

Transcription labels who speaks when, unless `speaker_labels` is false: Speaker 1, Speaker 2, and so on, numbered in the order people first speak, without real names. It finds at most 8 speakers.

- JSON: every word and sentence has `speaker` (a number from 1), and the transcript has `speaker_count`. A sentence never spans two speakers.
- TXT: with two or more speakers, one paragraph per turn, starting "Speaker 1: ".
- VTT: with two or more speakers, each cue names its speaker in a voice span, such as `<v Speaker 1>`.
- SRT: no names, so it stays clean for YouTube.
- No caption spans two speakers.
- The job's result has `speaker_labels` (`on`, `off`, or `failed`) and `speaker_count`.
- Reading the transcript: MCP `get_transcript` gives `speaker_count` and each sentence's `speaker`; the REST transcript endpoint gives `speakerCount` and `speaker`.

If labeling fails, the job still succeeds with every word, just without speakers: `speaker_labels` is `failed`, and the result has a `warning`.

## Languages

60 languages, detected automatically. To name one, set `language` to its code: `af` Afrikaans, `ar` Arabic, `hy` Armenian, `as` Assamese, `az` Azerbaijani, `bn` Bengali, `bs` Bosnian, `bg` Bulgarian, `yue` Cantonese, `ca` Catalan, `zh` Chinese (Simplified), `cs` Czech, `da` Danish, `nl` Dutch, `en` English, `et` Estonian, `fil` Filipino, `fi` Finnish, `fr` French, `gl` Galician, `de` German, `el` Greek, `gu` Gujarati, `he` Hebrew, `hi` Hindi, `hu` Hungarian, `is` Icelandic, `id` Indonesian, `it` Italian, `ja` Japanese, `kn` Kannada, `kk` Kazakh, `ko` Korean, `lv` Latvian, `lt` Lithuanian, `mk` Macedonian, `ms` Malay, `ml` Malayalam, `mr` Marathi, `ne` Nepali, `nb` Norwegian Bokmål, `or` Odia, `fa` Persian, `pl` Polish, `pt` Portuguese, `pa` Punjabi, `ro` Romanian, `ru` Russian, `sk` Slovak, `sl` Slovenian, `es` Spanish, `sw` Swahili, `sv` Swedish, `ta` Tamil, `te` Telugu, `th` Thai, `tr` Turkish, `uk` Ukrainian, `ur` Urdu, `vi` Vietnamese.

## From an assistant (MCP)

Add https://audo.ai/mcp and sign in (setup for each assistant: https://audo.ai/setup.md). The tool is `transcribe_audio`:

> Transcribes speech in an audio or video file, with a start and end time for every word and who speaks when (Speaker 1, Speaker 2, and so on, in the order they first speak). Use it when the person wants text or captions, or wants to cut by words. Not for lining up a script they already have (align_script). Set free to true to do the first 2 minutes at no cost, up to 3 times a day; otherwise it costs 2 credits a minute. Returns job_id; when done, the job lists TXT, SRT, VTT, and JSON files, and get_transcript reads them.

Example: "Use Audo to make an SRT file from vlog.mp4."

## More

- Every tool: https://audo.ai/llms.txt
- OpenAPI: https://audo.ai/api/openapi.json
- This page for people: https://audo.ai/srt-file-generator
