All tools › Optimize › OCR PDF
OCR PDF
Make scanned text searchable. Free, no sign-up.
or drop it here
OCR PDF runs inside your browser — your file never reaches our servers.
How to OCR PDF
OCR reads the words in a scanned page and puts an invisible text layer behind the image, so the page looks identical but is searchable.
- Select your scanned PDF. Photos of documents work too.
- Choose the language. English and Korean need no download; other languages fetch a pack once, with your permission.
- Download. The page looks the same, but you can now select and search the text.
What the searchable PDF is made of
Each page is rendered to a JPEG about 1100 pixels wide at quality 0.8, recognised, and rebuilt as that image with the recognised words written invisibly over it; the single-page results are then merged into one document. So what you download is a new file built from page images, and the original's links, bookmarks, annotations and form fields are not carried across. On A4, 1100 pixels is about 133 dpi — ample for recognition and for reading on screen, and well below a 300 dpi archival scan.
Because every page is re-encoded, a large scan usually comes back noticeably smaller. That is welcome for email, and a reason to keep the original when the resolution itself matters, as it does for evidence, signed contracts and anything going into an archive. The whole document is processed, too: there is no way to recognise only pages 4 to 9.
Choosing a language, and when to skip OCR
One language per run. English and Korean ship with the site and need no download, and the English + Korean option loads both at once — the only mixed-language combination available. The other thirteen fetch about 12 MB of recognition data once, after you agree to it, and it stays cached in this browser afterwards; your document still never moves. For a document that mixes languages, choose the one most of the words are in.
Before running it at all, try selecting a word in your PDF. If the text selects, there is nothing to recognise, and OCR would replace a crisp vector page with a 1100-pixel image for no gain. Where recognition genuinely helps, some tools already do it for you: PDF to Word and PDF to Excel recognise text-less pages during the conversion, though that built-in pass covers English and Korean only. Compare PDF, by contrast, refuses to run on two scans — the text layer this tool adds is exactly what it needs.
Questions
Does my document get sent away to be recognised?
No. Recognition runs on your own device. For languages other than English and Korean we fetch a public recognition model once — and we ask permission before doing it — but your document itself never moves.
Why is OCR slow?
Every page has to be rendered and then analysed character by character. It is the most demanding tool on the site, so a modern computer finishes in seconds where an old phone may take a minute per page.
How accurate is it?
Very good on clean, straight, reasonably high-resolution scans. Accuracy drops on faint, skewed or handwritten pages — run Deskew pages first if your scan is crooked, which measurably improves results.
Can I recognise two languages in one document?
Only English and Korean together, using the English + Korean option; every other language runs on its own. For a document that mixes languages, pick the one most of the text is in — the wrong pack costs accuracy across the whole page, not only on the foreign words.
Where does the language pack come from, and does it work offline afterwards?
English and Korean ship with the site. The other thirteen fetch a public, open-source recognition model — about 12 MB, once, and only after you agree — which is then cached in this browser, so that language works offline from then on. Clearing your browser's site data removes it, and the next run fetches it again.
Will the file get bigger or smaller after OCR?
A scan usually gets smaller, because each page is re-encoded as a JPEG around 1100 pixels wide. A PDF that already had text will get bigger and lose its sharpness, since its vector pages are replaced by images — the clearest reason not to run OCR on a document whose text already selects.
Are my files uploaded to a server?
No. Every tool on this site runs inside your browser, using the same engine as the Folia desktop app. Your file is read from your disk into your browser’s memory, processed there, and handed back as a download. It never crosses the internet, so there is nothing for us to store, leak or hand over.
Is there a file size limit?
We impose none. The only ceiling is your own device’s memory, which differs from one computer to another — so we read what your browser reports about this machine and show the resulting limit on the page. Light tools like merge and rotate handle far more than heavy ones like OCR, which must render every page. For anything larger, Folia for Desktop has no limit at all.
Is it really free? Do I need an account?
Yes, and no. Every tool is free and unlimited, with no watermark and no sign-up. Ads pay for the site, and your own computer does the work, which costs us nothing. Signing in with a paid Folia plan simply removes the ads.