Convert PDF to text
To convert PDF to text, copy it straight out of a PDF viewer, save it as plain text from Acrobat or Preview, or run pdftotext for whole folders; a scanned PDF needs OCR first. PDF Viewer on this site does not write a TXT file, but it does the two things that matter most: it OCRs scans in the browser, and then you can select and copy the text without uploading anything.
Text PDF or scan: the only question that matters
Extracting text only works when the PDF has text to extract. Open the file in the viewer and try to select a line with the mouse, or press Ctrl+F and search for a word you see. If selection works, every method below applies. If the cursor selects nothing, the page is an image.
For a scan, run OCR on the PDF. The viewer recognizes the text with tesseract.js inside the browser (language follows your browser setting, English by default) and adds an invisible text layer; it takes tens of seconds per page on a laptop, so start with the pages you need. After that, select the text on the page, copy it, and paste it into any editor. The scan never leaves your device. The scanned PDF guide covers the details.
Which PDF to text method for which case
| Case | Method | Notes |
|---|---|---|
| A few paragraphs | Select in the viewer, copy, paste | Fastest, keeps nothing but the words |
| Whole document to a TXT file | Acrobat: File, Export To, Text; Preview: File, Export, Plain Text (via Automator) or select all, copy | Acrobat keeps reading order best |
| Many PDFs, scriptable | pdftotext (Poppler) on the command line | Add -layout to keep columns |
| Need paragraphs, not lines | Open in Google Docs or Word, then save as TXT | Reflows text, see PDF to Word |
| Scanned pages | OCR in the viewer, then copy | Proofread names and numbers |
| Tables of numbers | pdftotext -layout, or PDF to Excel | Columns stay aligned |
Main method: copy from the viewer, pdftotext for bulk
Copy from the viewer:
- Open the PDF in the viewer by dragging it onto the page or with the Open a PDF button.
- If it is a scan, click OCR in the toolbar and wait for the pages to finish.
- Drag across the text to select it, press Ctrl+C (Cmd+C on Mac).
- Paste into Notepad, TextEdit, an email or a code editor.
pdftotext for whole files and folders:
- Install Poppler (Homebrew on Mac, package manager on Linux, a Windows build added to PATH).
- Run
pdftotext -layout input.pdf output.txt. Drop-layoutif you want flowing text rather than aligned columns. - Add
-f 5 -l 9to extract pages 5 to 9 only; use a shell loop for a folder.
Acrobat Pro: File, Export To, Text (Plain). Google Docs: upload, Open with Google Docs, File, Download, Plain Text. Online extractors exist, but they require uploading the document, which is unnecessary when every OS already has a free way.
What gets lost and how to fix it
- Line breaks. PDFs store lines, not paragraphs, so pasted text has a hard return after every line. In Word, Find and Replace ^p with a space, then restore paragraph breaks; in a code editor, join lines with a regex.
- Column order. Two-column layouts can come out interleaved. pdftotext without -layout usually detects columns; the viewer copies in reading order for most text PDFs.
- Hyphenation. Words split at line ends keep their hyphens.
- Ligatures and symbols. fi and fl ligatures, bullets and math symbols may paste as odd characters depending on the fonts.
- Headers, footers and page numbers are mixed into the text; strip them with a search.
- Formatting. Bold, headings and tables are gone by definition; use PDF to Word if you need them.
Text from a PDF on a phone and in bulk
On an iPhone, open the PDF in Files or Safari and long-press the text to select it; for a scanned PDF, Live Text recognizes the words on screen after you tap the text-select button in the viewer. On Android, Google Lens in the Files or Drive preview does the same; both send nothing anywhere for a text PDF, while Lens may use Google's servers for recognition. The viewer on this site also works on phones for normal-size files; large scans with hundreds of pages are slow or run out of memory there, so OCR those on a computer.
For thousands of files, a script is the only sane route. With pypdf in Python: text = ''.join(p.extract_text() for p in PdfReader('f.pdf').pages). pdftotext is faster and better at reading order, so most pipelines call it in a loop and only fall back to a library when they need per-page positions. Scans in the batch need an OCR step; command-line Tesseract or ocrmypdf adds the text layer before extraction.
Check the result
Count something: the number of pages or section headings in the PDF should match what you see in the text. Keep the PDF open in the online PDF viewer and use search inside the PDF to compare a few sentences from the middle and the end, where column mix-ups show up first. For OCR output, read every number and proper name against the page image.
Frequently asked questions
How do I extract text from a PDF?
Open it in a PDF viewer, select the text and copy it. For a whole file, use Acrobat's Export To Text or the free pdftotext command.
How to copy text from a scanned PDF?
Run OCR first. The viewer on this site OCRs in the browser and adds a text layer; after that you can select and copy as usual.
How to convert PDF to text without Adobe?
pdftotext from the Poppler tools, Google Docs (File, Download, Plain Text) or simply select all and copy in any viewer. None of them needs Acrobat.
Why can't I copy text from a PDF?
Either the PDF is a scan with no text layer, or copying is locked by owner permissions. OCR fixes the first; removing restrictions fixes the second.
How to keep the layout when converting PDF to text?
Use pdftotext -layout, which preserves column positions with spaces. Plain copy and paste keeps reading order but not alignment.
Does this PDF viewer save a PDF as a TXT file?
No. It OCRs and lets you copy the text in the browser; paste it into any editor and save as .txt yourself.