Do the evaluator roll-ups match the human labels?
Bottom line
The evaluator prose contains real signal, but it is not a complete or stable measurement of resume quality. On the two source judges of interest—GPT-5.4 high and Opus 4.6 adaptive medium—the semantic checker finds 135 of 230 core label opportunities (58.7%), 26 of 80 soft opportunities (32.5%), and 2 of 100 prohibited-claim opportunities (2.0%). Those recall figures are optimistic: the mistargeted-case audit below shows that the checker sometimes counts an adjacent criticism as the specific human label.
The second stage is mostly a compressor, not a rescue mechanism. Flash-Lite returns valid, source-linked output for 218/230 core-label opportunities (94.8%) and preserves 92/135 source-expressed core claims end to end (68.1%). Haiku returns valid output for 146/230 (63.5%) and preserves 76/135 (56.3%). Conditional on a valid output, Haiku is the more faithful reader—86.4% core preservation versus 71.9% for Flash-Lite—but its quotation/schema failures erase that advantage in an unretried pipeline.
That means the useful experimental object is a transition, not another mean:
human label → evaluator prose → accepted roll-up → artifact verification
The roll-up makes the evaluator's analysis legible. It does not establish that the analysis is true, and it cannot recover a percept that the evaluator never expressed.
What the evaluators are saying
The strongest repeatable findings are not subtle writing preferences. Across all five prompts, both GPT-5.4 and Opus consistently catch the unsupported skills in the skills-spam trap, that case's unaddressed UX-research requirement, the salience pair's in-progress work presented as complete, the pitch pair's seniority inflation, and the bolding pair's gradient- checkpointing upgrade. They almost always catch the truncation in the floor-check case and the credential-trap case's implementation-without-measurement problem.
They are much weaker on linguistic and editorial defects. Across the ten GPT/Opus opportunities per label, neither judge ever identifies the long-and-spammy case's length-driven keyword spam, the bolding pair's inline-bolding spam, the irrelevant credential in the credential trap, or the grounding pair's awkward phrasing as the corresponding human finding. The voice pair's imperative-summary label and the long-and-spammy case's unfocused-writing label are each recovered only once. The prose evaluators are therefore better at grounding, omission, and calibration failures than at whole-document voice, emphasis, and rhetorical-shape failures.
Opus recovers more core labels than GPT-5.4 (70/115 vs 65/115), while GPT-5.4 recovers far more soft labels (19/40 vs 7/40). Both trigger only one of fifty prohibited-label checks. This does not make either judge globally better; it describes their operating points on this small label catalog.
Prompt treatment changes what gets surfaced
For GPT-5.4 and Opus together, the core-label match rate rises from 54.3% under the control to 60.9% under mixed-neutral and 67.4% under mixed-harsh. Human-reading is 56.5%; bias-priming returns to 54.3%. This is evidence that the prompt changes the content of the prose even when a small numeric scale moves little. It is not evidence that harshness is automatically better: more criticism can include adjacent or over-broad claims, and this label set is too small to estimate that tradeoff tightly.
What the roll-up changes
The two quality questions separate cleanly:
- Did the synthesizer return an admissible artifact? Flash-Lite wins by a wide margin.
- When it returned one, did it preserve the evaluator's substantive point? Haiku wins on core and soft claims.
For prohibited claims, preservation is not desirable. GPT and Opus together make two prohibited assertions; both synthesizers filter them from every valid roll-up. Flash-Lite then adds two prohibited label-like claims in different runs, while Haiku adds none in this two-judge focus set. The additions are rare, but they show why a normalized report still needs artifact-level verification.
The mistargeted-case audit: the human label is sharper than the machinery
The core mistargeted-case judgment is stable in the source prose. All ten GPT-5.4/Opus reviews say, in substance, that the resume leans too technical and under-argues the business-partner, coaching, change, or adoption side. All five Gemini reviews miss that human label. Each synthesizer carries the hybrid-role point through in seven of the ten GPT/Opus runs. Haiku loses two to invalid output and drops one valid source claim; Flash-Lite loses one to invalid output and drops two.
GPT-5.4 is uniquely good on the concrete adoption evidence: all five GPT reviews call out the unused 20% portal-adoption growth and 18% support-call reduction. Flash-Lite preserves all five; Haiku preserves four and loses one to invalid output. Opus talks about adoption and change, but it does not surface those two quantified outcomes. The automated checker incorrectly promotes one generic Opus adoption comment into the quantified label.
The road-mapping/collaboration label is the diagnostic failure. The checker reports it in nine of fifteen source reviews. Manual reading finds zero reviews that name initiative road-mapping and only two that explicitly say unused cross-functional or cross-organization evidence should have been connected. The roll-up checker then calls three Haiku outputs and one Flash-Lite output matches, but none states the full road-mapping/collaboration inference; most discuss adjacent adoption, coaching, or business-positioning evidence.
So the mistargeted-case answer is not “the roll-up matches the labels.” It is:
- the evaluators reliably perceive the high-level hybrid-role mismatch;
- GPT-5.4 reliably sees the omitted quantified adoption evidence;
- neither stage expresses the owner's road-mapping inference;
- a cheap semantic checker can conceal that miss by accepting nearby language.
| Human label (mistargeted case) | Automated source verdict | Manual reading | Second-stage result |
|---|---|---|---|
| Too technical for the hybrid role | 10/15 | 10/10 GPT+Opus; 0/5 Gemini | 7/10 GPT+Opus for each synthesizer |
| Missed road-mapping / collaboration inference | 9/15 | 0 name road-mapping; 2 partial collaboration matches | 0 full matches; automated adjacent matches are overstated |
| Omitted 20% / 18% adoption outcomes | 6/15 | 5/5 GPT; 0/5 Opus; 0/5 Gemini | Flash 5/5 GPT; Haiku 4/5 GPT plus one invalid |
Explore every label transition
This table contains the 1,230 source-review-to-roll-up opportunities. A
missing roll-up verdict means the synthesizer output failed structural or
verbatim-citation validation; it is not counted as a semantic “no.” Source
citations were deliberately excluded from roll-up checking so copied evidence
could not create a match by itself.
Method and limitations
The current human catalog contains 41 single-run claims across 15 of the 16 cases: 23 core required findings, 8 soft required findings, 8 core prohibited claims, and 2 soft prohibited claims. Repeating them across five evaluator prompts and three judges creates 615 source-review checks. Valid Haiku roll-ups add 451 checks and valid Flash-Lite roll-ups add 593, for 1,659 checks total. The grounding-fabricated case is present in the full 240-review matrix but absent here because its current catalog contains pair-only, rather than single-run, labels.
Each check used claude-haiku-4-5-20251001, thinking off, a 1,024-token maximum,
strict JSON, and Anthropic's Batch API. The run consumed 1,778,312 input tokens
and 57,355 output tokens for an estimated $1.0325 at batch rates. There were
no API or parse failures.
The exact checker instruction was:
You are verifying whether a resume-review actually makes a specific point.
You will receive one claim and the full text of one review. Decide only
whether THIS review makes THIS point. Judge substance, not wording: a
paraphrase that raises the same specific issue counts; a vague adjacent
remark, or a passing use of similar words about a different issue, does not.
The claim arrives with a kind:
- kind "should_mention": answer yes only if the review surfaces the same
substantive problem or observation the claim describes.
- kind "should_not_assert": answer yes only if the review affirmatively
asserts the claim as its own finding. A review that raises the idea and
rejects it, or hedges it into a non-finding, is a no.
Return JSON only:
- `verdict`: "yes" or "no".
- `confidence`: "high" or "low".
- `quote`: the shortest excerpt from the review (up to 200 characters) that
supports your verdict, or "" when the verdict is no because nothing in the
review relates to the claim.
Only 135/196 source “yes” excerpts and 179/231 roll-up “yes” excerpts are literal substrings. Many of the rest are shortened or stitched with ellipses, but the field must not be mistaken for verified provenance. More importantly, the mistargeted case's road-mapping result shows that even a well-instructed semantic checker can be too liberal. These outputs are an index for human reading, not a new truth layer.
How this should be used for generation graphs
Do not turn the table above directly into one graph-quality number. Use it as a scorecard with separate axes:
- artifact validity: did the graph produce a usable resume and a usable evaluation artifact?
- verified core recovery: after reading the artifact, which owner-ratified defects were actually present and surfaced?
- prohibited assertion rate: which human-rejected criticisms did the graph or evaluator introduce?
- editorial quality: blinded human preference on voice, hierarchy, relevance, and whole-document argument—the areas these evaluators miss most;
- stability: how much do findings change across prompt treatment and judge?
For each candidate generation graph, run the same controlled cases, preserve the evaluator prose, use the second stage to normalize it, and manually verify the final claim set against the resume and source data. Compare graphs on that vector and inspect the disagreements. A weighted scalar can be added later for a specific decision, but it should be derived from the verified claim ledger, not from evaluator scores or unreviewed semantic-check verdicts.
The companion full-matrix output explorer shows every source review, accepted roll-up, rejected raw output, exact roll-up prompt, settings, usage, and measured synthesis cost.
.observablehq table {
font-size: 0.82rem;
}
.observablehq td {
vertical-align: top;
max-width: 34rem;
}
.observablehq blockquote {
border-left: 0.25rem solid color-mix(in srgb, currentColor 24%, transparent);
margin-left: 0;
padding-left: 1rem;
}