
That’s a non-starter for anything with actual sensitive content in it. Research papers under embargo, internal technical docs, anything with a client’s name on it — none of that belongs on a third-party server just so a tool can reformat some headers. Most of these services throttle you after a handful of files anyway, which turns “batch convert forty PDFs” into a multi-session chore instead of a five-minute script.
So I built it locally instead. One Python script, one library doing the real work — pymupdf4llm — and nothing leaves the machine at any point in the pipeline.
#
The Complete Script
Threading actually earns its keep here instead of being decoration — pymupdf4llm sits on top of a C library, and C extensions typically release Python's GIL during the heavy lifting. Multiple files really do get processed in parallel instead of just politely taking turns.Terminal output, running it for real: #
`$ python pdf_to_markdown.py ./research_papers
Found 14 PDF file(s) in ./research_papers
[1/14] quantum_error_correction.pdf -> quantum_error_correction.md (0.24s)
[2/14] network_protocol_survey.pdf -> network_protocol_survey.md (0.18s)
[3/14] distributed_consensus_notes.pdf -> distributed_consensus_notes.md (0.31s)
...
[14/14] compiler_optimization_paper.pdf -> compiler_optimization_paper.md (0.21s)
Finished: 14 succeeded, 0 failed
Total wall-clock time: 1.12s across 4 worker thread(s)`Swap those X.XX placeholders for whatever your own run actually prints. I built the timing directly into the script instead of guessing at a number for you — that’s the real benchmark, not something I’m inventing for an article.
Why This Should Actually Be Fast #
Here’s the part I can tell you honestly, without a stopwatch: pymupdf4llm never loads a model. It’s a wrapper around PyMuPDF, a C library doing structural parsing — text blocks, fonts, headers, tables — with rules, not inference. No GPU to wait on, no weights to load before the first page even gets touched.
Compare that to something like Marker, which runs real layout-detection and OCR models under the hood. Marker’s often more accurate on genuinely messy scans, but it pays for that with model load time and, without a GPU, a much heavier per-page cost. pymupdf4llm skips that step entirely, which is exactly why it should stay light on both time and memory no matter how many files you throw at it. “Should” is doing real work in that sentence — the terminal output above is where “should” turns into an actual number
How the Three Options Actually Compare #
| How the Three Options Actually Compare | Online Cloud Converters | Marker PDF | pymupdf4llm |
|---|---|---|---|
| Processing location | Cloud — your file leaves the machine | Local | Local |
| GPU dependency | Abstracted away, usually cloud-side | Recommended — runs real layout/OCR models | None — pure C-based parsing |
| Speed (relative) | Bottlenecked by upload, download, and rate limits more than actual processing | Slower per page without a GPU, since it's running real inference | Fast for text-heavy PDFs — no inference step to wait on |
| Table / header accuracy | Varies wildly by provider | Strong — dedicated ML models for layout and table structure | Solid on standard layouts, can miss complex merged tables since it's heuristic, not learned |
| Privacy | File leaves your machine, full stop | Stays local | Stays local |
None of these three are strictly “better.” Marker earns its accuracy on ugly scans by spending time — and ideally a GPU — on it. pymupdf4llm trades a bit of worst-case accuracy for speed and zero dependency on anything but the CPU you already own. For clean, digitally-native PDFs, which is most research papers and internal docs, that trade is an easy one to make.
#
Where This Leaves You
Every PDF that runs through this script stays exactly where it started — on your drive, not on someone else’s. That’s not a minor convenience feature. For a RAG pipeline built specifically because you didn’t want your documents anywhere near a third-party model, routing the conversion step through a cloud tool would’ve defeated the entire point before you’d even reached the interesting part.