ParseJet

PDF OCR — Text Recognition for Scanned PDFs

A scanned PDF is just pictures of pages: you cannot search it, select text, or copy from it. ParseJet’s PDF OCR renders each image-only page, runs text recognition on our own servers, and returns the document as text you can search, copy, and reuse. Pages that already have a text layer are extracted as-is, so mixed documents come out complete.

Drop a file here or browse

Accepts PDF files

Free — 3 requests/day, no signup. for 300 credits/month free.

How it works

1

Upload the scanned PDF

Drop the file above — from a scanner, a phone scanning app, or a fax. Digital PDFs with a text layer work too.

2

Text recognition, page by page

ParseJet checks every page. Pages with real text are read directly; image-only pages are rendered and passed through OCR.

3

Copy the text

Copy the whole document, or fetch it as Markdown or plain text through the API to feed search, spreadsheets, or an LLM.

Key features

What makes this pdf ocr stand out.

Detects which pages need OCR

Nothing to configure. A page with fewer than a couple of dozen extractable characters is treated as a scan and recognized; every other page keeps its original text.

20+ languages

Latin-script languages, Simplified and Traditional Chinese, and Japanese need no settings. Korean, Cyrillic, Arabic, Hindi, Thai, and Greek scans work with a language hint.

Mixed documents handled

Contracts with a scanned signature page, reports with inserted scans — the text comes back in page order with the recognized pages spliced in.

Nothing leaves our servers

Pages are rendered and recognized on ParseJet’s own infrastructure with open-source models; no third-party AI service sees your document.

Markdown or plain text

Ask for output_format=markdown to keep headings and tables from the digital pages, or raw for plain text.

Up to 30 scanned pages per request

Enough for most letters, forms, and reports. Longer scans: split the file, or process it in batches through the API.

Use cases

Common scenarios where this tool saves you time.

Archived contracts and invoices

Make years of scanned paperwork searchable — extract the text once and index it.

Phone-scanned documents

Scanner apps save image-only PDFs. Run them through OCR to copy an address, a clause, or a total.

Old books and reports

Turn scanned pages into text for quoting, translation, or accessibility tools.

Feeding scans to AI

Claude and ChatGPT cannot read a picture of a page reliably. OCR first, then paste or pipe the text in.

Automate with the API

Use the same tool programmatically. Works with any language — just HTTP.

cURL
# Scanned pages are detected and OCR'd automatically
curl -X POST "https://api.parsejet.com/v1/parse/auto/file?output_format=markdown" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]"

# Non-Latin scripts: add a language hint (ko, ru, ar, hi, th, el ...)
curl -X POST "https://api.parsejet.com/v1/parse/auto/file?language=ru" \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "[email protected]"

# Response metadata tells you which pages were recognized:
# "metadata": { "ocr_pages": [1, 2, 3], ... }
Python
import httpx

resp = httpx.post(
    "https://api.parsejet.com/v1/parse/auto/file",
    headers={"Authorization": "Bearer YOUR_API_KEY"},
    files={"file": open("scan.pdf", "rb")},
    params={"output_format": "markdown"},
    timeout=120,
)
data = resp.json()
print(data["text"])
print("OCR'd pages:", data["metadata"].get("ocr_pages", []))
JavaScript
const formData = new FormData();
formData.append("file", pdfFile);

const res = await fetch(
  "https://api.parsejet.com/v1/parse/auto/file?output_format=markdown",
  { method: "POST", headers: { Authorization: "Bearer YOUR_API_KEY" }, body: formData },
);
const { text, metadata } = await res.json();
console.log(text);
console.log("OCR'd pages:", metadata.ocr_pages ?? []);

Want to automate this?

ParseJet API gives you the same parsing power via a single HTTP endpoint. No ffmpeg, no poppler, no tesseract — just one API call.

curl -X POST https://api.parsejet.com/v1/parse/auto/url \ -H "Content-Type: application/json" \ -d '{"url":"https://example.com"}'
Read API Docs

Frequently asked questions

How do I OCR a PDF?

Upload the PDF above. ParseJet looks at every page, recognizes the ones that are only images, and shows you the full text — usually within a few seconds for a typical letter or form. Copy it, or call the API to do the same for many files.

How do I know if my PDF needs OCR?

Open it and try to select a word. If nothing highlights, or zooming in makes the letters look pixelated like a photo, the page is an image and needs OCR. PDFs from scanners, fax machines, and phone scanning apps are almost always like this.

Do I get a searchable PDF back?

You get the text, not a new PDF file. That is what you need for copying, searching, indexing, translation, or feeding an AI model. If you specifically need the original PDF with an invisible text layer added, a desktop tool such as OCRmyPDF does that.

Is there a page limit?

Up to 30 scanned pages are recognized per request; pages that already contain text do not count. For longer scans, split the file or send it in parts via the API. The response metadata lists the pages that were recognized and any that were skipped.

Which languages are supported?

English and all Latin-script languages, Simplified and Traditional Chinese, and Japanese work automatically. For Korean, Russian and other Cyrillic languages, Arabic, Hindi, Thai, or Greek, pass language=ko, ru, ar, hi, th, or el (the localized versions of this page do it for you).

Is it free?

Yes. Three free conversions a day with no signup, 300 credits a month with a free account, and paid plans from $19/month for bigger files and higher quotas.

Start extracting text for free

No signup required. Parse your first file in seconds.

View Pricing