Open a PDF
Home › PDF explained › What is a scanned PDF?

What is a scanned PDF?

A scanned PDF is a file in which every page is an image, usually produced by a scanner, a copier or a phone camera app, with no text underneath. It looks like a normal document but search, copy and screen readers find nothing. OCR adds the missing text layer, and this viewer can do it in the browser.

Image-only PDF: what is actually in the file

When a scanner or a multifunction printer saves to PDF, it takes a photo of the page (typically JPEG for color, CCITT or JBIG2 for black and white), and writes a PDF where each page's content is one drawing command: paint this image to fill the page. There are no fonts and no characters, so to a computer the page is as meaningful as a photo of a wall.

The same thing happens when you "print" a document to an image-based driver, when a fax arrives as PDF, or when a phone app stitches camera shots into a document. Some people call these "image PDFs", "flat PDFs" or "non-searchable PDFs"; they all mean the same thing. The opposite is a "native" or "text" PDF exported from an application, which is explained in what a PDF file is.

Try it on your PDFDrop a PDF here to open it in the viewer. Nothing is uploaded.Free, no account, no watermark. The file stays in this browser tab.

How to tell if a PDF is scanned

Three quick checks that work in any reader, including this one:

  1. Search. Press Ctrl+F (Cmd+F on Mac) and type a word that is clearly visible on the page. If it is a scan, there are no matches.
  2. Select. Drag across a sentence. In a text PDF the words highlight one by one; in a scan you get a selection rectangle, or nothing.
  3. Zoom. Zoom to 400 percent. Real text stays crisp at any size because it is vector; scanned text turns into visible pixels and jagged edges.

Other hints: the file is unusually large for its page count, the pages are slightly tilted, there are gray shadows at the edges, or Document properties shows a scanner model as the producer.

Why search, copy and accessibility fail

Everything that treats a PDF as text relies on the text layer. Without one:

  • Ctrl+F returns nothing, and desktop search tools skip the file.
  • Copy and paste gives you an image, or nothing, instead of words.
  • Screen readers announce an empty page, which is why image-only PDFs fail accessibility requirements such as Section 508 and WCAG.
  • Converters that turn PDF into Word or Excel have no text to work with and either refuse or produce a single big picture.
  • Indexing by email servers, document management systems and Google Drive search does not see the content.

Highlighting also behaves differently. Text highlighting snaps to words; on a scan there are no words, so you draw a box or a freehand mark over the area instead, which the viewer supports with its Draw tool.

Why scanned PDFs are so large

A page of text stored as characters takes a few kilobytes. The same page stored as a 300 dpi color image is roughly 8.7 million pixels before compression, and even after JPEG compression it lands between 200 KB and 1 MB. Multiply by a 100-page contract and you have a 50 MB file where a text version would be under 1 MB.

Common size reducers: scan in black and white or grayscale rather than color, use 300 dpi (enough for OCR, half the pixels of 400), and rescan rather than photograph. After the fact, compression tools re-encode the images at lower quality; this viewer does not compress, but the compress PDF guide lists what does. Large scans (hundreds of pages) can be slow or run out of memory in a browser, especially on phones.

How to fix a scanned PDF with OCR

OCR recognizes the characters in each page image and writes them into the file as an invisible layer, positioned over the printed words. The image remains the visible page, so nothing changes visually; search, copy and screen readers start working. The mechanism is described in what OCR is.

In this viewer, open the scan, click OCR and wait; recognition uses tesseract.js inside the browser, with the language taken from your browser settings (English by default), and the file is never uploaded. It is slow on long documents, tens of seconds per page on a laptop, but it is private and free. Then Download saves a searchable copy. Step-by-step instructions are in how to make a scanned PDF searchable, or go straight to the OCR tool.

If you need the scan as an editable Word document rather than a searchable PDF, run OCR first and then convert; see PDF to Word.

Frequently asked questions

How do I know if a PDF is scanned?

Press Ctrl+F and search for a word you can see. No result means the page is an image. Trying to select text or zooming in to see pixels confirms it.

Why can't I search my PDF?

The file has no text layer, which is typical for scans and faxes. Run OCR to add one; afterwards search and copy work normally.

Why can't I copy text from a PDF?

Either the PDF is image-only (a scan) or copying is blocked by permissions. OCR fixes the first case; removing restrictions fixes the second.

How do I make a scanned PDF searchable for free?

Open it in this viewer, click OCR, and download the result. Recognition runs in your browser with no upload.

Why is my scanned PDF so big?

Each page is stored as a photo. Color scans at high resolution take hundreds of kilobytes per page; black-and-white at 300 dpi is far smaller.

Can I convert a scanned PDF to Word?

Yes, but only after OCR has recognized the text. Word and most converters then rebuild the layout from the text layer, with mixed results on complex pages.