Do the evaluator roll-ups match the human labels?

Bottom line

The evaluator prose contains real signal, but it is not a complete or stable measurement of resume quality. On the two source judges of interest—GPT-5.4 high and Opus 4.6 adaptive medium—the semantic checker finds 135 of 230 core label opportunities (58.7%), 26 of 80 soft opportunities (32.5%), and 2 of 100 prohibited-claim opportunities (2.0%). Those recall figures are optimistic: the mistargeted-case audit below shows that the checker sometimes counts an adjacent criticism as the specific human label.

The second stage is mostly a compressor, not a rescue mechanism. Flash-Lite returns valid, source-linked output for 218/230 core-label opportunities (94.8%) and preserves 92/135 source-expressed core claims end to end (68.1%). Haiku returns valid output for 146/230 (63.5%) and preserves 76/135 (56.3%). Conditional on a valid output, Haiku is the more faithful reader—86.4% core preservation versus 71.9% for Flash-Lite—but its quotation/schema failures erase that advantage in an unretried pipeline.

That means the useful experimental object is a transition, not another mean:

human label → evaluator prose → accepted roll-up → artifact verification

The roll-up makes the evaluator's analysis legible. It does not establish that the analysis is true, and it cannot recover a percept that the evaluator never expressed.

What the evaluators are saying

The strongest repeatable findings are not subtle writing preferences. Across all five prompts, both GPT-5.4 and Opus consistently catch the unsupported skills in the skills-spam trap, that case's unaddressed UX-research requirement, the salience pair's in-progress work presented as complete, the pitch pair's seniority inflation, and the bolding pair's gradient- checkpointing upgrade. They almost always catch the truncation in the floor-check case and the credential-trap case's implementation-without-measurement problem.

They are much weaker on linguistic and editorial defects. Across the ten GPT/Opus opportunities per label, neither judge ever identifies the long-and-spammy case's length-driven keyword spam, the bolding pair's inline-bolding spam, the irrelevant credential in the credential trap, or the grounding pair's awkward phrasing as the corresponding human finding. The voice pair's imperative-summary label and the long-and-spammy case's unfocused-writing label are each recovered only once. The prose evaluators are therefore better at grounding, omission, and calibration failures than at whole-document voice, emphasis, and rhetorical-shape failures.

Opus recovers more core labels than GPT-5.4 (70/115 vs 65/115), while GPT-5.4 recovers far more soft labels (19/40 vs 7/40). Both trigger only one of fifty prohibited-label checks. This does not make either judge globally better; it describes their operating points on this small label catalog.

Prompt treatment changes what gets surfaced

For GPT-5.4 and Opus together, the core-label match rate rises from 54.3% under the control to 60.9% under mixed-neutral and 67.4% under mixed-harsh. Human-reading is 56.5%; bias-priming returns to 54.3%. This is evidence that the prompt changes the content of the prose even when a small numeric scale moves little. It is not evidence that harshness is automatically better: more criticism can include adjacent or over-broad claims, and this label set is too small to estimate that tradeoff tightly.

What the roll-up changes

The two quality questions separate cleanly:

  1. Did the synthesizer return an admissible artifact? Flash-Lite wins by a wide margin.
  2. When it returned one, did it preserve the evaluator's substantive point? Haiku wins on core and soft claims.

For prohibited claims, preservation is not desirable. GPT and Opus together make two prohibited assertions; both synthesizers filter them from every valid roll-up. Flash-Lite then adds two prohibited label-like claims in different runs, while Haiku adds none in this two-judge focus set. The additions are rare, but they show why a normalized report still needs artifact-level verification.

The mistargeted-case audit: the human label is sharper than the machinery

The core mistargeted-case judgment is stable in the source prose. All ten GPT-5.4/Opus reviews say, in substance, that the resume leans too technical and under-argues the business-partner, coaching, change, or adoption side. All five Gemini reviews miss that human label. Each synthesizer carries the hybrid-role point through in seven of the ten GPT/Opus runs. Haiku loses two to invalid output and drops one valid source claim; Flash-Lite loses one to invalid output and drops two.

GPT-5.4 is uniquely good on the concrete adoption evidence: all five GPT reviews call out the unused 20% portal-adoption growth and 18% support-call reduction. Flash-Lite preserves all five; Haiku preserves four and loses one to invalid output. Opus talks about adoption and change, but it does not surface those two quantified outcomes. The automated checker incorrectly promotes one generic Opus adoption comment into the quantified label.

The road-mapping/collaboration label is the diagnostic failure. The checker reports it in nine of fifteen source reviews. Manual reading finds zero reviews that name initiative road-mapping and only two that explicitly say unused cross-functional or cross-organization evidence should have been connected. The roll-up checker then calls three Haiku outputs and one Flash-Lite output matches, but none states the full road-mapping/collaboration inference; most discuss adjacent adoption, coaching, or business-positioning evidence.

So the mistargeted-case answer is not “the roll-up matches the labels.” It is:

Human label (mistargeted case) Automated source verdict Manual reading Second-stage result
Too technical for the hybrid role 10/15 10/10 GPT+Opus; 0/5 Gemini 7/10 GPT+Opus for each synthesizer
Missed road-mapping / collaboration inference 9/15 0 name road-mapping; 2 partial collaboration matches 0 full matches; automated adjacent matches are overstated
Omitted 20% / 18% adoption outcomes 6/15 5/5 GPT; 0/5 Opus; 0/5 Gemini Flash 5/5 GPT; Haiku 4/5 GPT plus one invalid

Explore every label transition

This table contains the 1,230 source-review-to-roll-up opportunities. A missing roll-up verdict means the synthesizer output failed structural or verbatim-citation validation; it is not counted as a semantic “no.” Source citations were deliberately excluded from roll-up checking so copied evidence could not create a match by itself.

Method and limitations

The current human catalog contains 41 single-run claims across 15 of the 16 cases: 23 core required findings, 8 soft required findings, 8 core prohibited claims, and 2 soft prohibited claims. Repeating them across five evaluator prompts and three judges creates 615 source-review checks. Valid Haiku roll-ups add 451 checks and valid Flash-Lite roll-ups add 593, for 1,659 checks total. The grounding-fabricated case is present in the full 240-review matrix but absent here because its current catalog contains pair-only, rather than single-run, labels.

Each check used claude-haiku-4-5-20251001, thinking off, a 1,024-token maximum, strict JSON, and Anthropic's Batch API. The run consumed 1,778,312 input tokens and 57,355 output tokens for an estimated $1.0325 at batch rates. There were no API or parse failures.

The exact checker instruction was:

You are verifying whether a resume-review actually makes a specific point.
You will receive one claim and the full text of one review. Decide only
whether THIS review makes THIS point. Judge substance, not wording: a
paraphrase that raises the same specific issue counts; a vague adjacent
remark, or a passing use of similar words about a different issue, does not.

The claim arrives with a kind:

- kind "should_mention": answer yes only if the review surfaces the same
  substantive problem or observation the claim describes.
- kind "should_not_assert": answer yes only if the review affirmatively
  asserts the claim as its own finding. A review that raises the idea and
  rejects it, or hedges it into a non-finding, is a no.

Return JSON only:

- `verdict`: "yes" or "no".
- `confidence`: "high" or "low".
- `quote`: the shortest excerpt from the review (up to 200 characters) that
  supports your verdict, or "" when the verdict is no because nothing in the
  review relates to the claim.

Only 135/196 source “yes” excerpts and 179/231 roll-up “yes” excerpts are literal substrings. Many of the rest are shortened or stitched with ellipses, but the field must not be mistaken for verified provenance. More importantly, the mistargeted case's road-mapping result shows that even a well-instructed semantic checker can be too liberal. These outputs are an index for human reading, not a new truth layer.

How this should be used for generation graphs

Do not turn the table above directly into one graph-quality number. Use it as a scorecard with separate axes:

For each candidate generation graph, run the same controlled cases, preserve the evaluator prose, use the second stage to normalize it, and manually verify the final claim set against the resume and source data. Compare graphs on that vector and inspect the disagreements. A weighted scalar can be added later for a specific decision, but it should be derived from the verified claim ledger, not from evaluator scores or unreviewed semantic-check verdicts.

The companion full-matrix output explorer shows every source review, accepted roll-up, rejected raw output, exact roll-up prompt, settings, usage, and measured synthesis cost.

.observablehq table {
  font-size: 0.82rem;
}
.observablehq td {
  vertical-align: top;
  max-width: 34rem;
}
.observablehq blockquote {
  border-left: 0.25rem solid color-mix(in srgb, currentColor 24%, transparent);
  margin-left: 0;
  padding-left: 1rem;
}