miinideckmiinideck
PricingUse casesBlog
Sign in
Tools

PDF to Markdown, for feeding an AI

Drop a PDF and get Markdown back. It is read in this browser — the file is never uploaded.

Drop a .pdf here — it is read in this browser and never uploaded.

How it works

Drop a PDF and the text comes back as Markdown: headings as headings, bullets as bullets, and paragraphs rejoined into paragraphs. That last one is the difference between output a model reads well and output it does not — a PDF stores one line at a time, so a naive extraction returns a sentence broken across six lines, and words hyphenated at the margin come back as "gener- ated".

Running headers and page numbers are stripped by default. On a forty-page report that is forty copies of the company name and forty "Page n of 40" lines, which is pure noise in a prompt and, on a metered model, noise you pay for. The panel says how many lines it removed, and a checkbox puts them back if you wanted them.

The reason to convert at all is that Markdown is much cheaper to read than a PDF. A PDF handed to a model either goes through the model's own extraction — which you cannot see or correct — or is processed as page images, which costs far more for the same words. Markdown is the format you can look at before you send it, which means you can tell when it went wrong.

Everything happens in this tab. The file is read with your browser's own file API, parsed here by pdf.js, and turned into Markdown here. There is no upload, no server round-trip, and nothing written down — which matters mostly when the PDF is not yours to hand to a converter site in the first place.

The output is yours to take: copy it, or download a .md file. If you want to read it as a page rather than raw text, the Markdown viewer opens it, and if the thing you build from it needs to reach a client, that is a separate step you take on purpose.

What this tool doesn’t do

  • Tables come out as lines of text, not as Markdown tables. Column structure lives in the page geometry, not in the file, and reconstructing it from coordinates is wrong often enough that a confident-looking wrong table is worse than obviously-plain text. If tables are the point of your document, Marker (datalab-to/marker) and Microsoft's MarkItDown both do a serious job of them locally.
  • Multi-column layouts read in the order the file stores them, which is usually correct and occasionally is not. Academic papers and newspaper-style layouts are where it shows.
  • Scanned PDFs have no text at all — the page is a picture, so the words are pixels. There is nothing to extract and no OCR here; the tool says so rather than handing you an empty file. Acrobat, ABBYY FineReader, or opening the PDF in Google Docs will OCR it and give you selectable text.
  • Images, charts and equations are not carried over. What comes back is the text layer, which is what a language model reads anyway.
  • Password-protected PDFs will not open. Remove the password in a PDF reader first.

Questions

Why convert a PDF to Markdown before giving it to an AI?

Two reasons, and the second is the one people find out about later. First, cost and length: the same words cost far less as text than as a PDF a model has to process as page images. Second, control: when you hand over a PDF, whatever extraction happens is invisible to you, so if it reads a two-column page in the wrong order you have no way of knowing — you just get a confident answer built on scrambled text. Markdown is the version you can read before you send it.

Does my PDF get uploaded anywhere?

No. Your browser reads the file, pdf.js parses it in this tab, and the Markdown is assembled here. Your document is never sent anywhere, never stored, and never logged, and closing the tab leaves nothing behind. To be exact about what the page does do: like any web page it loads its own code from miinideck, and for a document in Chinese, Japanese or Korean it fetches the character-map files needed to decode those glyphs — so the fact that you opened a CJK document is visible in that request, while nothing of its contents ever is. That is the whole reason this tool has the shape it has: the documents people most want to convert — a client's contract, an internal report — are exactly the ones they should not be dropping into a converter site.

Will tables survive the conversion?

Not as tables. A PDF does not record "this is a table" — it records where each piece of text sits on the page, and a grid has to be inferred from that. Inference gets it wrong often enough that we would rather give you plain text you can see is plain than a table that looks authoritative and has two columns swapped. For table-heavy documents, Marker and MarkItDown are built for exactly that job and run locally too.

My PDF produced nothing. What happened?

It is almost certainly a scan — a photograph or scanner image of a page, wrapped in a PDF. There are no characters in the file, only pixels that look like characters, so there is nothing for any text extractor to find. Turning that into text needs OCR: Acrobat and ABBYY FineReader do it, and uploading the PDF to Google Drive and opening it with Google Docs does it for free. Bring the result back here if you still want Markdown.

Why are the headers and page numbers missing?

They are removed on purpose. A line that appears at the top or bottom of most pages is chrome rather than content, and on a long report it repeats dozens of times — noise in a prompt, and length you pay for on a metered model. The panel tells you how many lines it took out, and the "keep headers and page numbers" checkbox restores them.

Does it work for Chinese, Japanese or Korean PDFs?

Yes. CJK text needs the character maps that tell the reader which glyph is which character, and those ship with this page — without them a Chinese PDF extracts as blank or as rows of identical boxes, which is the failure most browser-based converters have.

How large a PDF can it handle?

Up to 30 MB, and the limit exists because the work happens in your browser rather than on a server. A several-hundred-page document will take a moment and show you which page it is on. Beyond that ceiling the honest answer is to split the file, because a tab that runs out of memory fails in a much less useful way than a message.

Is Markdown better than plain text for this?

Usually, yes, and for a specific reason: headings and lists tell a model how the document is organised, so it can tell a section title from a sentence and an enumerated list from prose. Flattening everything to plain text throws that away, and it is information the PDF actually contained. It costs almost nothing in length — a few hash marks and hyphens.

Can I share the result with someone?

Not from here — this page converts and hands the file back, and that is where it stops. When what you build from it needs to reach a client, miinideck publishes a page at an unguessable link that stays out of search by default. Without an account the link lasts 7 days and has no password; sign in (free) to add a password or set your own expiry. That involves a real upload, so it only happens when you ask for it.

Related

MD viewerRead the Markdown you just got, as a page.PDF vs Markdown for AI tokensWhat the same document costs in each format.How many tokens is an image PDF?Why a scan costs so much more than text.Markdown vs HTML vs PDFWhich format for which job.Embedding a PDF in a page you plan to shareWhen you want the PDF kept, not converted.
miinideck

HTML files, finally as links — for AI builders, agencies, and consultants. Default-noindex, default-private, default-yours.

Product

  • Pricing
  • Use cases
  • Try it free

Resources

  • Blog
  • Free tools
  • Featured on
  • Report abuse

Legal

  • Privacy
  • Terms
Listed onmiinideck listed on Product Huntmiinideck listed on Faziermiinideck listed on TheSaaSDirmiinideck listed on AIToolHuntmiinideck listed on LaunchNest
© 2026 miinideckMade for people who don't want their work indexed.