What a bad extraction costs, in points and in sections
Extraction failures are invisible on your screen and expensive everywhere else. Here are the costs, with the numbers attached.
— read the counts first, then the score.

An extraction failure is the most expensive thing that can happen to a résumé, because nothing about it is visible to you. The page looks right. The reading is wrong. Here is what that costs, priced.
The three numbers that reveal it
Words, lines and recognised headings. A specimen document came back as 162 words, 18 lines and 4 recognised headings. A weak one came back as 48 words, 10 lines and 2 headings. If your two-page résumé reports numbers like the second set, the score is a distraction and the extraction is the whole story.
The parseability deductions, in points
| Detected | Cost, out of 100 |
|---|---|
| A table or multi-column pattern | 22 |
| No standard section with readable content | 24 (8 if only one is missing) |
| Three or more unreadable glyphs | 18 |
| Decorative characters | 8 |
| Shortened years | 8 |
| Mixed date formats | 4 |
Parseability is 30% of the published score, so those deductions are scaled accordingly. It is also capped at 25 when the text is under 25 words, and at 45 when no section has readable content at all.
The second, larger cost
A failed extraction takes completeness with it. Experience is 30 points of the completeness 100, education 15, skills 15. A layout that hides your experience section does not cost you 22 points of parseability; it costs those and then the 30, and then the content dimension has nothing to grade. That compounding is what produces the arithmetic on a weak document: 92 × 30% + 40 × 30% + 0 × 40% = 40.
Then the ceilings
After the weighted sum, four caps apply. Under 25 readable words, the score is held at 15. With no experience, education or skills found, 25. With no experience and no education, 40. With no experience and content under 30, 55. Those exist so a document nobody can read cannot look respectable, and they are explained in evidence ceilings explained.
The causes, in order of frequency
Columns, tables, contact details in a header region, decorative bullets from an icon font, and a scan with no text layer. Each has a fix and none is cosmetic. The catalogue is in why résumés get scrambled, the font half in fonts that survive extraction, and the scanned case in checking a scanned résumé.
How to check yours in a minute
Run the ATS checker on the file you actually send, not on the text you pasted from your notes. The file is the thing an employer receives, and the difference between the two is exactly where extraction failures live. Then read the counts before you read the number.
Why this failure is uniquely expensive
Because it is silent on both sides. You cannot see it, since your document renders correctly on your screen. The employer cannot see it either, since what arrives is a record with fields in it rather than a warning. Every other weakness in a résumé is at least visible to somebody: a thin summary can be read, a vague line can be judged, a gap can be asked about. An extraction failure produces a plausible-looking record of a different person, and nobody in the process has any reason to suspect it. Ten seconds of checking removes a whole category of invisible loss.
What this page does not claim
That our extraction is what an employer's system will do. Different products use different libraries. What generalises is the principle rather than the point values: a document whose text has one clear order survives every reader, and one whose order depends on where things sit on the page does not.
There is a version of this problem that catches careful people, and it is worth knowing about. A document can extract perfectly and still lose the content dimension, because the achievement lines did not survive as lines. Until recently our own exported PDFs drew bullets as layout rather than writing them into the text, so the same résumé measured seven achievement lines pasted and four exported. That was found, measured and repaired, and the whole episode is in pulling text back out of a PDF.
A final note about when to check. Run it on the file after every change that touches layout: a new design, a new section, a margin adjustment that moved the page break. Content edits rarely break extraction and structural ones sometimes do, and the failure is silent in both directions. Two minutes after a design change is a cheap insurance premium against a document that quietly stopped being readable halfway through a job search.
See the counts before you see the score
The report prints the words, lines and recognised headings it read. Those three numbers are the extraction, in public.
Open the ATS checkerFonts that survive extractionQuestions
- How do I know my extraction failed
- Compare the report's reading counts with your page. A two-page document that reports 48 words has not been read, whatever the score says.
- What does a multi-column layout cost
- 22 points of parseability, which is the largest single deduction in that dimension. Parseability is 30% of the score.
- What happens when a section is not found
- It fails its completeness check, and parseability deducts separately for a heading with no readable content. Missing experience alone costs 30 of the completeness 100.
- What is a ceiling
- A cap applied after the weighted sum, so a document with almost nothing readable in it cannot look ready. Under 25 words, the score is held at 15.
- Does a scan extract at all
- Sometimes, through recognition, and unevenly. Recognition recovers words more reliably than it recovers achievement lines, so a clean scan still loses points.
Conxfolio is a free set of four career tools: a résumé builder with 37 rendered layouts, an ATS résumé checker that prints its own arithmetic, a cover letter builder that traces every proof paragraph back to the line it came from, and a portfolio builder with 20 authored designs. There is no account to create, nothing is held back for a paid plan, and no language model is used anywhere in the product, so the readers, the score and the letter are deterministic code you can check.