How to make a scanned PDF searchable
Add a hidden, searchable text layer to a scanned English PDF with on-device OCR, check the result, and understand what OCR can get wrong.
Your PDFs are processed in your browser and are not uploaded to any server.
A scanned PDF is a stack of photos of paper. It looks like text, but your PDF viewer sees only pictures, so Ctrl+F finds nothing and you cannot copy a sentence. OCR (optical character recognition) reads the letters in those pictures. The OCR PDF tool puts the recognized words into the PDF as an invisible text layer that sits exactly over the scan. The page looks the same as before, but you can now search it, select words, and copy them.
Before you start
- English only. The tool recognizes English text. There is no language setting, and text in other languages will be missed or come out wrong.
- A one-time download. The first time you press start, your browser loads the OCR engine and the English model, about 7 MB, from this website. Your PDF is not part of that download and is never sent anywhere.
- Time and memory. Each page takes a few seconds. While a page is read, it is turned into a large image, and the whole PDF stays in the tab's memory. Very large documents, with hundreds of pages or very large files, can slow the browser down or make it stop the tab, especially on phones.
Make the PDF searchable step by step
- Open the OCR PDF tool.
- Choose your scanned PDF with the file button, or drop it on the box. The tool works on one PDF at a time.
- Press Start OCR. Progress is shown page by page. If it is taking too long, press Cancel; nothing is saved, and you can start again with a smaller file.
- Read the summary. It says how many pages had text recognized and how many words were recognized, which pages already had text, and which pages had nothing readable.
- Press the download button. The new file is named after the original with
-searchableadded, socontract.pdfbecomescontract-searchable.pdf. Your original file is not changed.
Check that it worked
Open the new PDF in your usual viewer and search for a word you can see on a page, using Ctrl+F or Cmd+F on a Mac. Then drag across a line to select it and paste it somewhere. If the search finds the word and the pasted text reads correctly, the text layer is in place. Pay extra attention to numbers, names, and dates: OCR can make mistakes, and a single wrong digit is easy to miss.
What the tool does with each kind of page
- Scanned or image-only pages: the English text is recognized and added as an invisible layer.
- Pages that already have selectable text: kept exactly as they are. If every page already has text, the summary says so and the PDF is unchanged.
- Blank pages, or pictures without readable English text: nothing is added, and the summary lists those pages. If no page gives any text, you get a message instead of a file.
In every case the page count and order stay the same, and the scan itself is not redrawn, cleaned up, or straightened.
Getting better results
- Start from the best scan you can. Clear, dark print on a light background, scanned straight, gives far better results than a blurry phone photo taken at an angle.
- Expect trouble with handwriting, very small print, unusual fonts, stamps over text, tables, and pages with many columns. Recognition is not guaranteed to be accurate for any document, and these cases are the hardest.
- For a long document, split it into parts with the Split PDF tool, run OCR on each part, and join the searchable parts again with Merge PDF. Closing other tabs also frees memory.
What to do next
A searchable PDF has a real text layer, so you can now use tools that need one. To get the recognized text as an editable file, convert the searchable PDF with PDF to Markdown; the PDF to Markdown guide explains how to check that file. Any OCR mistakes carry over into it, so read it before you rely on it.
Limits to keep in mind
- English only. No other OCR languages are available.
- OCR can make mistakes. Check important text against the scan.
- Very large documents can reach the browser's memory limit, especially on phones.
- Password-protected PDFs cannot be opened. Remove the password in another app first.