How OCR PDF works
A scanned PDF is a stack of pictures. To make it searchable, each page goes through three steps, one page at a time so even long scans fit in memory:
- Render. pdf.js draws the page at 300 DPI, the resolution the Tesseract documentation recommends. Phones get a lower cap (5 megapixels, about 230 DPI for a Letter page) so the browser tab doesn't run out of memory.
- Recognise. Tesseract.js, the WebAssembly build of the open-source Tesseract engine, finds lines and words and reports a box and a confidence score for each word.
- Write. Each word is written into the page as invisible text, stretched to cover its box exactly. This is the same method Tesseract's own PDF output and OCRmyPDF use, with a font whose letters are blank, so any script (including Arabic, Hindi and Chinese) can be searched and copied.
Pixels at 300 DPI = (page width in inches × 300) × (page height in inches × 300). Text size in points = distance from the top of the word's line to its baseline, in pixels, × 72 ÷ 300.
Worked example: a 12-page signed lease
A Letter page is 8.5 × 11 in, so at 300 DPI it renders at 2,550 × 3,300 = 8.4 megapixels, inside the 25-megapixel cap for laptops. 10 pt body text is about 42 pixels per em (10 ÷ 72 × 300), large enough for Tesseract to read reliably. At 2 to 6 seconds a page, 12 pages take roughly 24 to 72 seconds. When it finishes, "lease-searchable.pdf" downloads; open it, press Ctrl+F, type "security deposit", and the reader highlights the clause on page 7.
Getting the best results
| Problem | What to do |
|---|---|
| Low scan resolution | Rescan at 300 DPI. Text in 150 DPI scans and faxes is often too small to read reliably. |
| Pages scanned sideways | Turn them upright first with Organize PDF. This tool does not detect page orientation. |
| Document in another language | Choose it in the language list. English data misreads accented letters and can't read Hindi, Arabic or Chinese at all. |
| Handwriting | Expect poor results: Tesseract is trained on printed text. |
| Some pages already searchable | Leave "Skip pages that already have text" ticked, so those pages aren't given a second, duplicate layer. |
Privacy
Scans are often the most sensitive files people have: signed contracts, IDs, medical records, tax returns. Online OCR services upload them. Here the engine and language data are downloaded from this site once and cached by your browser, and every page is recognised on your own device. If you only need the text of a photo or screenshot rather than a searchable PDF, use Image to text. If your PDF already has selectable text, PDF to text copies it out instantly, without OCR.
Questions people ask
Will the PDF look any different after OCR?
No. The scanned pictures are left exactly as they were. The recognised words are added as an invisible layer on top, positioned over the words in the picture, so you can search, select and copy them. The file grows only slightly, because the hidden text is tiny next to the scanned pictures.
How long does it take?
About 2 to 6 seconds per page on a recent laptop, after a one-time download of the OCR engine (about 3.7 MB) and the language data (about 1.9 MB for English). A 20-page scan usually takes one to two minutes. Phones are slower. The progress bar shows which page is being read, and Cancel stops at once.
Why are some words wrong when I search or copy?
OCR guesses each letter from the picture, so its accuracy depends on the scan. Sharp, straight, 300 DPI scans of printed text come out nearly perfect. Faxes, skewed phone photos, small print, stamps and handwriting cause mistakes. The ticket shows an average confidence score; under 70% means you should expect errors.
What happens to pages that already have text?
By default they are skipped, so a PDF that mixes typed pages and scanned pages doesn't end up with every word twice. A page counts as having text if it already holds at least 20 letters or digits that can be selected. Untick the option to OCR every page anyway.
Can it OCR a password-protected PDF?
Yes, if you know the password: you are asked for it when you add the file. The searchable copy is saved without the password, and without any print or copy restrictions the original had, so lock it again with Password protect PDF if it needs protecting.
Is my scan uploaded?
No. The OCR engine (Tesseract) is downloaded from this site and runs in your browser. Your PDF is read, recognised and saved on your device; the page never sends it anywhere.
