Free online PDF to text converter that pulls the plain text out of a PDF into a .txt file, dropping formatting, images and layout so only the words remain. A scanned PDF has no text layer, so Smart Auto-Detect runs OCR on it.
Document format work by OnlinePCApps since 2013
The words in a PDF are often the part worth keeping. This converter lifts them out as plain text with nothing else attached.
Some PDFs block selecting or copying the words on the page. Pulling the text layer out to a plain file gives the content back to work with, without wrestling the reader.
A report or a contract holds wording to cite exactly. Pulling it to text keeps the phrasing right for a quote, a footnote or a search, with no retyping and no typos slipping in.
A script, a search index or an analysis pipeline wants raw text, not a page layout. A clean .txt file drops straight in, so the content can be processed without any formatting in the way.
Drop a document onto the panel above, or click to browse. The file card shows its name and size, with Change File beside it for a different document.
Set Which pages to extract to All pages, First page, First 3 or a list like 1, 3-5, 8. Leave Smart Auto-Detect on to run OCR when a scanned page turns up, with a Document Language and a Scan Recognition Quality of 120 to 360 DPI under the advanced settings, or select Fast Digital Text Only for a clean digital PDF.
Click Extract Text. When the extraction completes, Download TXT File saves a .txt named after the PDF, and Reset clears the panel. The original PDF stays untouched, and the plain text opens in any editor or drops into any tool.
To convert PDF to text is to lift the words off a page and leave everything else behind. This page does that with two methods and a page range.
The converter takes the text layer a PDF carries and gathers every word in reading order. The result is the raw wording of the document, ready to copy, search or edit anywhere.
Fonts, colours, columns, pictures and page design are all dropped. Plain text is the entire point, a lightweight file with nothing to reformat and nothing to get in the way of the words.
Smart Auto-Detect takes the digital text first and runs OCR on its own when a scanned or non-selectable page turns up, with a language code and a 120 to 360 DPI setting. Fast Digital Text Only skips the scan check for a clean PDF from Word or Canva.
The words come down as one .txt file named after the PDF, from all pages or the range typed in. Plain text opens in any editor and on any device, with no reader or licence needed to use it.
To convert PDF to text, to extract text from a PDF and to save a PDF as TXT all name the same step. The .txt is written new and the original PDF is left as it stands.
Extraction takes a text layer and drops everything else. These points decide what lands in the file.
| The case | Result | What happens and why |
|---|---|---|
| Text-based PDF | clean | A PDF with a text layer comes straight out, with the words kept exactly, as how a PDF stores text explains. |
| Scanned PDF | OCR by itself | A scan is a picture with no text layer. With Smart Auto-Detect on, OCR recognizes the characters first, using the Document Language and Scan Recognition Quality set under the advanced settings. |
| Formatting | dropped | Fonts, colours, images and layout are left out, since plain text holds only words. |
| Copy-protected | extracted anyway | A file that blocks copying in a reader can still have its text layer pulled out here. |
| Multi-column | reading order | The text follows the page order, though a busy multi-column page can come out of sequence. |
| Encoding | UTF-8 | The file is saved in Unicode UTF-8, so accents and other scripts carry across intact. Character encoding explained covers why that matters. |
| Output | one .txt | The words go into a plain .txt file named after the PDF, from every page or the range typed in. |
| Formatted need | use Word | For headings and styling, PDF to Word rebuilds a document instead of stripping it. |
| Page range | all or a list | Which pages to extract takes All pages, First page, First 3 or a list like 1, 3-5, 8, so one chapter comes out without the rest. |
The Words Only, With the Layout Stripped Away
A PDF pins its text to fixed spots and dresses it in fonts, columns and images. Converting to text throws all of that out and keeps only the words, which a PDF stores by position rather than as lines, so they are reconstructed into reading order and written to a plain .txt file that any editor can open. A text-based PDF carries a text layer the converter takes straight out, while a scanned page is only a picture and needs OCR to recognize the characters first, which Smart Auto-Detect runs on its own. Because the result is plain text, a formatted document or a table has no place to live, so PDF to Word suits a formatted rebuild and PDF to Excel suits a table of figures. The file is written in Unicode UTF-8 so accents and non-Latin scripts survive. The standards below define the text and the page it comes from.
Both pull the text out of a PDF. This converter differs in price, in how a scan is handled and in how many files go through per run.
| Point of comparison | This tool This PDF to text converter On the device | Hosted suites Converters bundled with an editor and a plan |
|---|---|---|
| Where the document goes | Never leaves the device | Uploaded to a server |
| Price | Free with no cap | Often a couple of tasks a day |
| OCR for scanned PDFs | Smart Auto-Detect, built in | Often an OCR option |
| Order on a busy layout | Simple pages come out best | Heavier engine sorts more |
| Many files at once | One PDF per run | Batch on their hardware |
The last two rows lean to the hosted suites, which sort a tangled reading order with a heavier engine and clear many files in a queue. The price and OCR rows go to this converter, which is free with no daily cap and switches OCR on by itself. Browser-based file tools weighs the rest of the trade.
The words are read out of the PDF inside the browser, so the document and its text stay on the machine that opened it. Nothing is sent off to be read elsewhere.
The words on a page are frequently the sensitive part, a contract clause, a medical note or a private letter. A hosted converter uploads the whole file to read it, which is more exposure than pulling out text should ever call for.
Every one of these gets words out of a PDF. They differ in effort, in reach and in what happens to a scanned page.
Copy and paste, Word and the command line all need a reader, a licence or a terminal. None of them runs OCR on a scan by itself. The online PDF to text converter above needs none of those and does.
A few notes set the right expectation for the text file that comes back.
If the words in the PDF can be selected, it has a text layer and comes out cleanly on Fast Digital Text Only. If they cannot, the file is a scan with only a picture of text, so leave Smart Auto-Detect on and OCR recognizes it first.
The result is words alone, with no headings, tables or images. For a formatted document reach for PDF to Word, and for a table of figures reach for PDF to Excel instead.
A plain page comes out top to bottom as expected. A multi-column or heavily boxed page can come out in a different order, since plain text has no columns to keep the flow apart.
The online PDF to text converter handles one document per run, which suits the file in hand. A folder of reports to turn into text, or a book of many hundreds of pages, is work for the desktop edition, which reads from disk and writes a text file for every document in one pass.
Aim it at a directory and every PDF gives up a matching text file in a run, every one saved next to the document it came from.
Scanned pages go through optical character recognition as they process, turning a picture of text into words that can be saved.
A long report or a book of many hundreds of pages comes off disk without holding the entire file in a browser tab.
Free, with no sign-up and no watermark. Pull the plain text out of a PDF into one .txt, OCR included for a scan.