ParseJet

Audio to Text — Transcribe Any Recording

Upload a recording and get a clean, punctuated transcript in seconds. ParseJet transcribes MP3, WAV, M4A, OGG, FLAC and WebM with speech recognition that runs on our own servers — dozens of languages, automatic language detection, timed segments for captions. Free to try, no signup; the same transcription is available through the API.

Drop a file here or browse

Accepts MP3,WAV,M4A,OGG,FLAC,WEBM,AAC files

Free — 3 requests/day, no signup. for 300 credits/month free.

How it works

1

Upload your audio

Drop any recording above — voice memo, interview, lecture, call. Large files are compressed in your browser first, so even a 40 MB WAV uploads in seconds.

2

Speech recognition

The language is detected automatically and the speech is transcribed on ParseJet’s servers — nothing is sent to a third-party AI service.

3

Copy your transcript

Copy the text into notes, documents or a prompt. Developers get the same result as JSON with segment and word timestamps.

Key features

What makes this audio to text stand out.

Every common audio format

MP3, WAV, M4A, OGG, FLAC, WebM and AAC — straight from a phone, a recorder, a DAW or a meeting tool.

Dozens of languages, auto-detected

English, Spanish, Portuguese, French, German, Japanese, Chinese, Korean, Arabic, Russian and many more. Pass a language hint for short clips.

Readable, punctuated text

Sentences, punctuation and paragraph breaks at pauses — not a wall of lowercase words.

Timestamps for captions

Every response includes timed segments; the API can add word-level timings for SRT or VTT subtitles.

Private by design

Recognition runs on servers ParseJet operates. Audio is processed in memory and discarded once the transcript is returned.

Same engine via API

One POST with the file, JSON back. Python and JavaScript SDKs, 300 free credits a month.

Use cases

Common scenarios where this tool saves you time.

Interviews and podcasts

Turn an hour of conversation into searchable, quotable text for articles, show notes and research.

Meetings and calls

Transcribe the recording, then paste it into Claude or ChatGPT for minutes and action items.

Lectures and voice memos

Recordings from class or ideas captured on the go become notes you can search and edit.

Captions and accessibility

Use the timed segments to produce subtitles, or publish the transcript alongside the audio.

Automate with the API

Use the same tool programmatically. Works with any language — just HTTP.

cURL
# Transcribe a recording (MP3, WAV, M4A, OGG, FLAC, WebM)
curl -X POST https://api.parsejet.com/v1/parse/audio \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]" \
  -F "language=en" \
  -F "with_timestamps=true"

# Response: text + metadata.segments [{start, end, text}] + metadata.words
Python
import httpx

resp = httpx.post(
    "https://api.parsejet.com/v1/parse/audio",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    files={"file": open("interview.mp3", "rb")},
    data={"with_timestamps": "true"},
    timeout=600,
)
data = resp.json()
print(data["text"])                       # the transcript
for seg in data["metadata"]["segments"]:  # timed segments for captions
    print(f'{seg["start"]:.1f}s  {seg["text"]}')
JavaScript
const formData = new FormData();
formData.append("file", audioFile); // MP3, WAV, M4A, OGG, FLAC, WebM
formData.append("with_timestamps", "true");

const res = await fetch("https://api.parsejet.com/v1/parse/audio", {
  method: "POST",
  headers: { Authorization: "Bearer YOUR_API_KEY" },
  body: formData,
});
const { text, metadata } = await res.json();
console.log(text, metadata.segments);

Want to automate this?

ParseJet API gives you the same parsing power via a single HTTP endpoint. No ffmpeg, no poppler, no tesseract — just one API call.

curl -X POST https://api.parsejet.com/v1/parse/auto/url \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com"}'
Read API Docs

Frequently asked questions

How do I transcribe an audio file to text?

Drop the recording above and wait a moment; the transcript appears below the box and you can copy it. No account is needed for the first three transcriptions a day. For batches, POST the files to /v1/parse/audio.

How long can the recording be?

Up to 5 minutes of audio per request and 25 MB per file. Longer recordings: split them (most editors can cut on silence) and transcribe the parts — the API returns timestamps, so the pieces line up.

Which languages are supported?

Dozens, including English, Spanish, Portuguese, French, German, Italian, Dutch, Polish, Russian, Turkish, Arabic, Japanese, Korean and Chinese. The language is detected automatically; a hint (language=es) helps on very short clips.

How accurate is the transcription?

Clear speech from one or two speakers transcribes with high accuracy, including punctuation. Heavy background noise, crosstalk and thick accents lower it. Uncompressed or high-bitrate audio does not help much — clarity of the speech matters more than the file format.

Can I get subtitles (SRT) from the transcript?

The response includes timed segments, and with_timestamps=true adds word timings, which is everything an SRT or VTT file needs. A one-click subtitle download is on the roadmap.

Is my recording uploaded to a third party?

No. Speech recognition runs on ParseJet’s own servers with open-source models. Large files are compressed in your browser before upload, and the audio is discarded once the transcript is returned.

Is it free?

Yes. Three free transcriptions a day with no signup. A transcription costs 3 credits (up to 5 minutes of audio); the free account includes 300 credits per month. Paid plans start at $19/month.

Start extracting text for free

No signup required. Parse your first file in seconds.

View Pricing