Skip to the article
Conxfolio
Blog · Walkthroughs

A scanned résumé, measured: 75 where the original scored 97

We rasterised our own PDF at 150 dpi, uploaded the image, and wrote down what the recognition path recovered.

— the words came back; the lines did not.

The arithmetic line on a Conxfolio report reading 100 times 30 per cent plus 85 times 30 per cent plus 49 times 40 per cent equals 75
The scan's arithmetic. Parseability survived intact; the content dimension did not.

A scanned résumé is a picture of a document. Nothing in it is text until something recognises it, and recognition is very good at words and much worse at structure. We measured exactly how much worse, on a file we control.

The experiment

Our own PDF export, rasterised at 150 dpi into an image, uploaded to the ATS checker as a picture. The same document had already been measured as text, at 97, so the difference is attributable to the scan rather than to the writing.

What came back

The arithmetic line reading 100 times 30 per cent plus 85 times 30 per cent plus 49 times 40 per cent equals 75
The scan's sum, photographed. 75, against 97 for the same document as text.

100 × 30% + 85 × 30% + 49 × 40% = 75. The reading was 157 words, 19 lines, 4 recognised headings, 2 roles and 9 skills. Parseability came back at 100 with 5 of 5 checks passed, which is the surprising part: the page structure survived.

Where the 22 points went

The three weighted diagnostics on the scanned document
Parseability 100. Completeness 85, 6 of 7 checks. Content 49, with none of its four rules at full marks.

Almost all of it in the content dimension, which fell from 92 to 49. Its four rules depend on recognising an achievement line as a line: action-led writing, measurable evidence, depth across roles, readable volume. Recognition returns the words without reliably returning the line boundaries, so the rules that grade lines have far less to work with.

Why rescanning does not fix it

Because the loss is structural rather than a matter of legibility. The scan was clean, at a resolution most office scanners exceed, and the words were recovered. Raising the resolution recovers words that were already recovered.

The actual repair

The résumé builder's start step
Import what recognition recovered, then correct it. The review shows every field it could not fill.

Import the scan into the résumé builder, let the reading recover what it can, and correct the rest in the review. The reading never invents: a field it could not read stays empty and is reported as not found, which is described in dropped fields and why they are counted. Then export, and the new file has a real text layer.

Proving the repair

Run the check on the exported PDF rather than on the editor's text. Our exports now write bullet glyphs into the text layer, so a pasted document and an exported one score the same on the specimen, which is the repair described in pulling text back out of a PDF.

What to do if the scan is all you have

Import it anyway, because recognition recovered the words reliably and words are the expensive part to retype. The measured reading came back with 157 words, two roles and nine skills from an image, which is most of a document. What you then supply by hand is the structure: which lines belong to which role, where the achievement lines start and stop, and the dates. That is fifteen minutes of correcting rather than an hour of typing, and the export at the end is a real document rather than a picture of one.

What this page does not claim

That 75 is what every scan scores. One document, one resolution, one recognition engine. What generalises is the shape of the loss rather than its size: the words survive, the achievement lines do not, and the dimension worth 40% is the one that pays for it. It is also worth saying plainly that this is a place our own product is weaker than it should be, and it is recorded as such rather than hidden.

If somebody has sent you a scan and you are checking it for them, the same measurement applies and the same caution goes with it. A 75 on a scanned document does not mean the writing is mediocre; it means a third of the evidence never reached the scorer. Read the reading counts first, and if they are plausible while the content row is low, the scan is the explanation. Checking someone else's résumé covers how to say that helpfully, and checking a scanned résumé is the step-by-step version.

One preventative note, since scans usually arrive from a specific habit. If your only copy of a résumé is a printed page or a photograph of one, rebuild it now rather than at the moment you need to apply. The recognition path recovers enough to make that a short job today and a rushed one on a deadline. Once a real document exists in the builder, every export from it has a written text layer and the problem does not recur.

Rebuild it as a real document

Import what recognition recovered, correct it in the review, and export a file with a real text layer.

Open the résumé builderChecking a scanned résumé

Questions

What did the scan actually score
100 × 30% + 85 × 30% + 49 × 40% = 75, against 97 for the same document as text. The reading was 157 words, 19 lines, 4 recognised headings, 2 roles and 9 skills.
So recognition worked
For the words, yes. The word count came back within five of the original. What it lost was the structure of the achievement lines, so the content dimension fell from 92 to 49.
Why does that matter more than the words
Because the content dimension is 40% of the score, and its four rules all depend on recognising a line as an achievement rather than as a run of text.
Can I fix a scan
Not by rescanning at a higher resolution. Rebuild the document so it has a real text layer, which is what the builder is for.
Is a photograph of a printout the same problem
Yes, with added distortion. Anything without a text layer goes down the recognition path.

Conxfolio is a free set of four career tools: a résumé builder with 37 rendered layouts, an ATS résumé checker that prints its own arithmetic, a cover letter builder that traces every proof paragraph back to the line it came from, and a portfolio builder with 20 authored designs. There is no account to create, nothing is held back for a paid plan, and no language model is used anywhere in the product, so the readers, the score and the letter are deterministic code you can check.