Parityhealth-plan operations benchmark

Every response, and the thinking behind it.

All 17,631 responses from the 28 ranked models are on this page in full: what was sent, what the model thought, what it answered, and how the grader scored it.

17,631
Responses
12,586
With a reasoning trace
210
Tasks
28
Models
95.9%
Judge agreement
822
Re-graded by a second judge

Read the run

Filter by model, family or outcome, then open any row for what was sent, what the model thought, what it answered and how the grader scored it. Press / to search. Narrow to one model or one family and full-text search over the responses themselves turns on.

Mean score
Format valid
Models
Items
Reasoning tokens
Loading the response index.

What the field found easy, hard, and impossible

Solved by every model
56
of 210 tasks. These carry no ranking information.
Solved by nobody
0
An task nobody passes is as likely to be a defective task as a hard one.
Re-graded by a second judge
822
95.9% exact agreement, mean absolute difference 0.017 on the 0–1 scale

Tasks no model scored on

Every task was solved by at least one model.

Tasks the field disagrees about most

Standard deviation of the score across all models. These carry the most ranking information and are the most interesting to read side by side.

ItemTitleMeanSD
LDG-012Twenty-four claims on an HDHP from a warm start51.20.500
LDG-002Fourteen claims from a warm start, an adjustment and a void47.60.499
LDG-004Copays that credit the deductible, fifteen claims53.60.499
LDG-011Twenty-four claims, five members, four edits56.00.496
LDG-008Twenty claims, mixed network, three edits57.10.495
BEN-022Allowed below the copay58.30.493
LDG-005Five members, eighteen claims, three edits40.50.491
LDG-006HDHP from a warm start with the family ceiling in reach60.70.488