Convert scanned PDFs to Markdown with on-device OCR
A scanned PDF is just photographs of pages — invisible to search, to screen readers, and to AI tools. Milldown detects scans automatically and recognizes the text on-device — with Apple's Vision framework on your Macwith the Windows OCR engine on your PC. No cloud OCR service ever sees your documents.
How it works
- Drop the scanned PDF into Milldown like any other file.No settings needed — if a PDF has little or no embedded text, OCR kicks in automatically.
- Apple's Vision framework reads every page on-device.The Windows OCR engine reads every page on-device.Each page is rendered at roughly 216 dpi first — high enough for accurate recognition without producing enormous bitmaps. It is the same recognition behind Live Text in Photos.
- You get ordered, structured text with per-page markers.An "On-device OCR" badge shows the result came from recognition, not extraction.
Why on-device OCR matters
- Privacy — scanned documents are usually the sensitive ones: contracts, invoices, medical records, IDs. Cloud OCR means uploading them to someone's server. Milldown never does.
- Accessibility — recognized text is navigable by VoiceOverNarrator and any screen reader; the original scan wasn't.
- Searchability — text in Markdown is findable in Spotlight, Obsidian, grep, anything.
- Works offline — recognition needs no connection at all.
How Milldown decides a PDF is scanned
There is no setting to get wrong. After converting a PDF, Milldown looks at how much text came out per page: fewer than 200 characters per page and the file is treated as image-only, so recognition runs instead. A PDF with a real text layer is never re-OCR’d — extraction is both faster and more accurate than recognition, so it always wins where it is available.
That threshold is why mixed PDFs behave the way they do. A document with twenty text pages and one scanned insert converts from its text layer; the inserted page contributes little. Documents that are scanned throughout are the ones that cross the line. If you have a mixed file whose scanned pages matter, split it and convert those pages separately.
What you get back
- One page marker per page — multi-page scans come back with
#### Page 1,#### Page 2and so on, so a quotation can be traced to a page. Single-page scans have no marker. - Reading order, not layout — recognition returns lines of text. Straightforward single-column pages come out in the order you would read them.
- An “On-device OCR” badge on the converted document, so you can tell recognition from extraction at a glance — they are not equally trustworthy, and you should know which you have.
- A token count, like any other conversion. The PDF savings apply to the recognised text the same way.
Where it struggles — and what it refuses
Recognition is not extraction, and the limits matter:
- Tables come back as text, not tables. The recogniser returns lines; it does not reconstruct rows and columns into a Markdown table. Figures inside a scanned table are read correctly, but you will be re-assembling the structure yourself.
- Multi-column layouts can interleave. Newspaper and journal pages sometimes come back reading across columns rather than down them. Worth a glance before you trust the order.
- Handwriting is unreliable. This is built for printed text. Neat block capitals sometimes work; cursive generally does not.
- Quality follows the scan. A clean 300-dpi scan recognises nearly perfectly. A phone photo of a crumpled page, a heavy skew, or a faint fax will not.
- A scan with nothing readable is refused, with a message saying so. It never hands you an empty file labelled “converted” — a silent empty result is worse than an error, because you might paste it somewhere.
Languages
The language list comes from macOS, not from Milldown, so it grows when your Mac does. On current macOS the accurate recogniser handles 30 language variants:
English · French · Italian · German · Spanish · Portuguese (Brazil) · Chinese (Simplified and Traditional) · Cantonese · Korean · Japanese · Russian · Ukrainian · Thai · Vietnamese · Arabic · Turkish · Indonesian · Czech · Danish · Dutch · Norwegian (Bokmål and Nynorsk) · Malay · Polish · Romanian · Swedish
Language correction is on, which improves accuracy on ordinary prose and occasionally “corrects” an unusual proper noun or part number. Check identifiers against the scan before relying on them.
Images, too
The same on-device recognition applies to JPEG and PNG files — a photographed whiteboard, a screenshot, a page snapped on a phone. An image with no readable text is refused for the same reason a blank scan is.
Further reading: Converting documents to Markdown offline · Private document conversion for Claude
Try it on your own files
Milldown converts PDFs, Office documents, web pages, and scans into clean Markdown — entirely on your computer. Free for 7 days.
Download for macOSmacOS 14+ · Apple Silicon · $39 one-time after trial (excl. VAT) · no subscription