What actually happens when you hand a PDF to a model
I spent three weeks debugging a pipeline that pulled training data from annual reports. The reports were fine, the OCR was fine, the tokenizer was fine. The problem was the machine learning pdf format itself — specifically, how layout gets preserved inside the file and how badly that breaks when you assume text flows top-to-bottom the way a Word document does. A PDF doesn't store paragraphs. It stores drawing commands: place this glyph at x=142, y=687; place this one at x=142, y=681. Columns, footnotes, sidebars, tables that span pages — none of that is semantic. It's visual. When you extract text with a naive reader you get the right words in the wrong order, and your model learns that "see Figure 3" appears before the figure caption, which is backwards from how a human reads it.
Building a machine learning pdf extraction pipeline
The first thing I learned the hard way is that pdfminer.six and PyMuPDF give you different answers for the same file, and neither is automatically better. PyMuPDF is faster and handles scanned images well when you pair it with Tesseract. pdfminer preserves layout structure — column detection, bounding boxes, reading order — but it's slower and chokes on malformed files. For a production pipeline I ended up running both and falling back to whichever one produced a valid parse tree. Here's the actual sequence that works for me now:
First, classify the PDF. Is it text-based or scanned? A quick heuristic: open it with fitz (PyMuPDF) and count non-empty text blocks. If fewer than 30% of pages return text, it's scanned. For text-based PDFs, use pdfminer to extract structured content with bounding boxes. For scanned PDFs, run OCR page-by-page with Tesseract, but set --psm 6 (assume a single uniform block of text) for reports and --psm 3 for mixed layouts with columns and captions. Then clean the output. This is where most people stop and wonder why their model's BLEU score is terrible. You need to reassemble sentences that PDF layout broke across columns. A column break looks like a normal space in the raw text, but semantically it's a newline. I write a post-processor that uses bounding box Y-coordinates to detect when text wraps to the next column, then joins those fragments with a space instead of a newline. Tables are worse. They need to become markdown or CSV before the model sees them, because a table rendered as plain text loses all structure and the model treats column headers as random words.
Chunking comes after cleanup. Don't chunk by character count. Chunk by semantic units: paragraphs, then tables, then figure captions separately from the text that references them. A 500-token chunk that splits a sentence in half teaches the model that sentence boundaries are arbitrary, which makes retrieval garbage. I usually aim for 300 to 400 tokens per chunk with a 50-token overlap, but only when the overlap doesn't cut through a paragraph.
👉 Clique no botão abaixo para saber mais sobre o assunto!
Edge cases that aren't edge cases
The one that cost me the most time: PDFs with embedded fonts that map characters to wrong Unicode codepoints. I had a German annual report where every "ß" (eszett) was encoded as two separate "s" characters because the font used a legacy encoding. The text looked correct on screen. The raw bytes were wrong. My model learned that "Fluß" was a word. It took me two days to realize the source was the font encoding, not the OCR, and the fix was to force pdfminer to use the TOUnicode CMap when available, which reconstructs the correct Unicode mapping from the font's glyph table. Another one: PDFs with transparent text layers. Some generators put a visual copy of the text as vector shapes and a second invisible copy as real text for accessibility. If you extract from the visible layer you get garbage characters. If you extract from the invisible layer you get the right text but lose formatting. The workaround is to check for the RolePlane and ActualText attributes that pdfminer exposes — when they exist, they tell you which text stream is the semantic one.
What this approach doesn't solve
Extraction is only the first step. A cleaned PDF can still contain things that break downstream models. Mathematical notation in LaTeX-rendered PDFs becomes unreadable soup when converted to plain text. Chemical structures, circuit diagrams, architectural floor plans — none of that transfers to text. If your use case involves those, you need a multimodal model that accepts images, not a text pipeline. Retrieval quality depends heavily on your embedding model. Most general-purpose embeddings perform badly on technical PDFs because the training data is dominated by web pages and Wikipedia. I've seen domain-specific fine-tuning improve retrieval recall by 15 to 20 percentage points on scientific papers, but that requires labeled query-document pairs, which most teams don't have. A cheaper alternative is to add metadata — section headings, page numbers, document type — to your chunks and weight the metadata during retrieval. It's not as good as fine-tuning, but it usually recovers enough signal to make the system usable.
The biggest limitation is cost. OCR on a 500-page report at 300 DPI with Tesseract takes roughly 8 to 12 minutes on a modern CPU, and that's without parallelization. PDFMiner on the same file takes about 3 minutes. If you're processing hundreds of documents, you need a queue and you need to accept that some files will fail. I keep a failure log with the error type and retry with adjusted parameters — usually raising the DPI or switching parsers — before marking a file as permanently broken. About 5% of PDFs I encounter fall into that category, and they're usually files generated by old accounting software with non-standard compression.
Where to find existing implementations
If you're building a machine learning pdf pipeline from scratch, start with LangChain's PDFLoader or LlamaIndex's SimpleDirectoryReader with the PDFReader extractor. Neither is perfect out of the box, but they handle the common cases and let you swap in custom post-processors for column repair and table conversion. The custom post-processor is where the actual work happens — the extractors are just starting points. For production scale, I recommend writing your own wrapper around PyMuPDF and pdfminer rather than trying to make a generic framework do everything. The framework abstractions leak when you hit the edge cases I described above, and you end up fighting the API instead of solving the problem. A 200-line wrapper that tries pdfminer first, falls back to PyMuPDF, runs the column-join post-processor, and logs failures is easier to maintain than a 2000-line chain of adapters.