What is OCR?
OCR, optical character recognition, is software that looks at an image of a page and works out which letters and words are in it. Applied to a scanned PDF, it adds an invisible text layer so you can search, select and copy the text. This viewer runs OCR in your browser with tesseract.js, so the scan is never uploaded.
How OCR works, step by step
An OCR engine does not "read" the way a person does; it runs a pipeline of image processing and pattern recognition:
- Preprocessing. The page image is converted to black and white, straightened (deskewed) if the scan is tilted, and cleaned of speckles.
- Layout analysis. The engine finds blocks of text, separates them from pictures and tables, and splits blocks into lines and words.
- Character recognition. Each line is passed to a model (modern engines such as Tesseract 4 and 5 use an LSTM neural network) that outputs the most likely sequence of characters.
- Language correction. A dictionary and word statistics for the chosen language fix ambiguous shapes: "rn" versus "m", "0" versus "O", "1" versus "l".
- Output. Recognized words come back with their position on the page and a confidence score.
The position data is what makes PDF OCR useful: the words can be written back into the file exactly where they appear in the image.
What a text layer is
After OCR, the scanned image stays untouched and a second, invisible layer of text is placed on top of it, each word positioned over its printed counterpart. Readers render the image and ignore the invisible glyphs, but Ctrl+F, text selection, copy and screen readers all use the text layer. This is often called a "searchable PDF" or "image over text" PDF.
Because the picture is still the visible page, OCR mistakes never change what you see; they only affect search and copy. That also means you can run OCR again later with a better engine without losing anything. For what an image-only file looks like before this step, see what a scanned PDF is.
A text layer is different from converting a scan to Word. Conversion tries to rebuild paragraphs, fonts and tables as editable objects; a text layer only makes the existing page searchable. The conversion route is covered in PDF to Word.
OCR accuracy and languages
On a clean 300 dpi scan of printed English text, current engines reach roughly 98 to 99 percent character accuracy, which is one or two wrong characters per page. Accuracy falls quickly as the source gets worse: 150 dpi, faxes, photocopies of photocopies, colored backgrounds, or photos taken with a phone at an angle.
Language matters because the dictionary step relies on it. Running an English model over a Spanish or German page produces plausible-looking but wrong words. Tesseract has trained data for over 100 languages and scripts, including Arabic, Cyrillic, Chinese, Japanese and Hindi; right-to-left and connected scripts generally score lower than Latin text.
In this viewer the OCR language is picked from your browser language, with English as the default. The language data is downloaded once from a CDN; the document itself never leaves the browser.
When OCR fails or gives bad results
Recognition is only as good as the pixels it gets. The usual causes of garbage output:
- Handwriting. Standard OCR is trained on printed type. Cursive and most handwritten notes are not recognized; that needs dedicated handwriting recognition.
- Low resolution. Letters smaller than about 20 pixels tall lose their shape. Rescan at 300 dpi if you can.
- Skew and curvature. Phone photos of a curved book page confuse line detection.
- Decorative fonts, stamps, watermarks and underlines that touch the letters.
- Tables and multi-column layouts, where the reading order may come out scrambled even if the words are right.
- Wrong language selected, as described above.
If a scan is also locked with permissions, OCR tools may refuse to write a new file; remove the restrictions first.
How this viewer does OCR in the browser
Open the scanned PDF and click OCR. The viewer uses tesseract.js, a port of the Tesseract engine compiled to run inside a web page, so recognition happens on your own CPU rather than on a server. When it finishes, an invisible text layer is added over each page; Ctrl+F and copy start working immediately, and Download writes a new PDF with that layer embedded.
The trade-off is speed. A desktop OCR package uses native code; in a browser expect tens of seconds per page on a laptop, longer on a phone, so a 200-page scan is a coffee-break job. Very large scans can also run out of memory on mobile devices. For a quick start see how to make a scanned PDF searchable, or open the OCR PDF tool directly.
Frequently asked questions
What does OCR stand for?
Optical character recognition: software that identifies the letters and words in an image of a page and produces machine-readable text.
What is OCR in a PDF?
It is the process of adding a hidden text layer over scanned page images so the PDF becomes searchable and its text can be selected and copied.
How accurate is OCR?
About 98 to 99 percent per character on clean, 300 dpi printed text; noticeably worse on faxes, low-resolution scans, phone photos and unusual fonts. Handwriting is generally not recognized.
Can OCR read handwriting?
Standard OCR engines, including Tesseract, are trained on printed type and do poorly on handwriting. Neat block capitals sometimes work; cursive almost never does.
Is online OCR safe for confidential documents?
Most online OCR services upload the file to their servers. This viewer runs the engine inside your browser, so the document stays on your device.
Does OCR change how the PDF looks?
No. The scanned image remains the visible page; the recognized text is invisible and only used for search, selection and copy.