Convert scanned PDFs to Markdown with on-device OCR

A scanned PDF is just photographs of pages — invisible to search, to screen readers, and to AI tools. Milldown detects scans automatically and recognizes the text on-device — with Apple's Vision framework on your Macwith the Windows OCR engine on your PC. No cloud OCR service ever sees your documents.

How it works

  1. Drop the scanned PDF into Milldown like any other file.No settings needed — if a PDF has little or no embedded text, OCR kicks in automatically.
  2. Apple's Vision framework reads every page on-device.The Windows OCR engine reads every page on-device.Each page is rendered at roughly 216 dpi first — high enough for accurate recognition without producing enormous bitmaps. It is the same recognition behind Live Text in Photos.
  3. You get ordered, structured text with per-page markers.An "On-device OCR" badge shows the result came from recognition, not extraction.

Why on-device OCR matters

How Milldown decides a PDF is scanned

There is no setting to get wrong. After converting a PDF, Milldown looks at how much text came out per page: fewer than 200 characters per page and the file is treated as image-only, so recognition runs instead. A PDF with a real text layer is never re-OCR’d — extraction is both faster and more accurate than recognition, so it always wins where it is available.

That threshold is why mixed PDFs behave the way they do. A document with twenty text pages and one scanned insert converts from its text layer; the inserted page contributes little. Documents that are scanned throughout are the ones that cross the line. If you have a mixed file whose scanned pages matter, split it and convert those pages separately.

What you get back

Where it struggles — and what it refuses

Recognition is not extraction, and the limits matter:

Languages

The language list comes from macOS, not from Milldown, so it grows when your Mac does. On current macOS the accurate recogniser handles 30 language variants:

English · French · Italian · German · Spanish · Portuguese (Brazil) · Chinese (Simplified and Traditional) · Cantonese · Korean · Japanese · Russian · Ukrainian · Thai · Vietnamese · Arabic · Turkish · Indonesian · Czech · Danish · Dutch · Norwegian (Bokmål and Nynorsk) · Malay · Polish · Romanian · Swedish

Language correction is on, which improves accuracy on ordinary prose and occasionally “corrects” an unusual proper noun or part number. Check identifiers against the scan before relying on them.

Images, too

The same on-device recognition applies to JPEG and PNG files — a photographed whiteboard, a screenshot, a page snapped on a phone. An image with no readable text is refused for the same reason a blank scan is.

Further reading: Converting documents to Markdown offline · Private document conversion for Claude

Try it on your own files

Milldown converts PDFs, Office documents, web pages, and scans into clean Markdown — entirely on your computer. Free for 7 days.

Download for macOS

macOS 14+ · Apple Silicon · $39 one-time after trial (excl. VAT) · no subscription