How to turn a PDF into Markdown for notes or an AI assistant
Step-by-step: convert a text-based PDF to a Markdown file in your browser, check the result, and fix the usual problems with scans, columns, and tables.
Your PDFs are processed in your browser and are not uploaded to any server.
Copying text out of a PDF viewer often gives you broken lines, lost headings, and stray hyphens. Markdown is a plain-text format with a few simple marks: # for a heading and - for a list item. It opens in any text editor, works in most notes apps and documentation tools, and is easy to paste into an AI assistant as context. This guide shows how to get a Markdown file from a PDF with the PDF to Markdown tool and how to check it before you use it.
First, check that your PDF has real text
The tool reads the PDF's text layer, which is the selectable text that a PDF viewer lets you highlight and copy. Open your PDF and try to select a sentence.
- You can select the words: the PDF has a text layer, and this guide is for you.
- Nothing gets selected, or the whole page highlights like a picture: the page is probably a scan. The PDF to Markdown tool does not recognize text in images. Make the PDF searchable with the OCR PDF tool first, then convert the searchable PDF. The guide to making a scanned PDF searchable explains how.
Convert the PDF step by step
- Open the PDF to Markdown tool.
- Choose your PDF with the file button, or drop it on the box. The tool works on one PDF at a time, and files larger than 300 MB are refused because a browser tab cannot handle them reliably.
- Press Convert to Markdown. The tool does nothing until you press it. Progress is shown page by page, and you can press Cancel at any point.
- When it finishes, press Download report.md (the button shows your own file name). The file has the same name as your PDF with a
.mdending, soreport.pdfbecomesreport.md. It is UTF-8 text, so accented letters and other non-English characters in the text layer come through. - Read the short message above the download button. If some pages had no text, it tells you which ones were skipped and whether they looked like scanned images.
What the Markdown file looks like
The tool guesses the structure from how the text looks on the page. Text in a larger font becomes a heading: the largest size becomes #, the next size ##, and so on. Lines that start with a bullet character become - list items, and lines that start with a number such as 1. or 2) stay numbered. Lines that wrap inside one paragraph are joined back together, and a gap between lines starts a new paragraph. A short one-page update could come out like this:
# Quarterly update
## What changed
Shipping times went down after the new warehouse opened.
- Orders now leave within two days
- Returns are handled at the same site
1. Review the new rates
2. Update the price listThe pages follow one another in a single file, with no page-break markers between them. Images, tables as tables, links, colors, and fonts are not kept: the file contains text only.
Check the result before you rely on it
The structure is a best guess, so spend a minute reading the file, especially before you hand it to an AI assistant that will treat it as accurate.
- Reading order: text is taken in the order it is stored inside the PDF, which is not always the order you read it in. Pages with two or more columns, sidebars, or text boxes can come out interleaved or with sections in an unexpected place. Move those paragraphs back where they belong.
- Tables: a table becomes ordinary lines of text, often one cell after another. If the numbers matter, copy them from the PDF by hand or rebuild the table.
- Repeated page text: running headers, footers, and page numbers are text too, so they usually appear in the file on every page. Delete them if they get in the way.
- Headings: a bold line in body-sized text is not a heading to the tool, and a large pull quote can become one. Fix the
#marks where the guess is wrong.
Using the file with an AI assistant or notes app
Open the .md file in a text editor and paste the part you need, or attach the file if your app accepts text files. Markdown headings help an assistant tell sections apart, and a shorter, cleaned-up file leaves more room for your question. Remove anything private before you paste it. The PDF to Markdown tool itself never sends your PDF or its text anywhere, but whatever you paste into another service is handled by that service.
If something goes wrong
- “No text could be extracted”: the PDF has no text layer. If it says the pages contain images, it is a scan; run it through OCR PDF first. Print-ready files where the letters were turned into shapes (outlined fonts) also have no text to read.
- “Password-protected”: remove the password in the app you normally use to open the PDF, save a copy, and try that copy.
- “Looks damaged or is not a valid PDF”: download or export the PDF again, then retry.
Limits to keep in mind
- Scanned or image-only pages are skipped: the tool does not recognize text in images.
- The reading order can be imperfect, especially with columns, tables, footnotes, and sidebars.
- Headings, lists, and paragraphs are detected from font size, spacing, and leading marks, so they can be wrong.
- Only text is kept. Images, links, and table layout are dropped.