Every response, and the thinking behind it.
All 17,631 responses from the 28 ranked models are on this page in full: what was sent, what the model thought, what it answered, and how the grader scored it.
Read the run
Filter by model, family or outcome, then open any row for what was sent, what the model thought, what it answered and how the grader scored it. Press / to search. Narrow to one model or one family and full-text search over the responses themselves turns on.
Mean score
Format valid
Models
Items
Reasoning tokens
Loading the response index.
What the field found easy, hard, and impossible
Solved by every model
56
of 210 tasks. These carry no ranking information.
Solved by nobody
0
An task nobody passes is as likely to be a defective task as a hard one.
Re-graded by a second judge
822
95.9% exact agreement, mean absolute difference 0.017 on the 0–1 scale
Tasks no model scored on
| Every task was solved by at least one model. |
Tasks the field disagrees about most
Standard deviation of the score across all models. These carry the most ranking information and are the most interesting to read side by side.
| Item | Title | Mean | SD |
|---|---|---|---|
| LDG-012 | Twenty-four claims on an HDHP from a warm start | 51.2 | 0.500 |
| LDG-002 | Fourteen claims from a warm start, an adjustment and a void | 47.6 | 0.499 |
| LDG-004 | Copays that credit the deductible, fifteen claims | 53.6 | 0.499 |
| LDG-011 | Twenty-four claims, five members, four edits | 56.0 | 0.496 |
| LDG-008 | Twenty claims, mixed network, three edits | 57.1 | 0.495 |
| BEN-022 | Allowed below the copay | 58.3 | 0.493 |
| LDG-005 | Five members, eighteen claims, three edits | 40.5 | 0.491 |
| LDG-006 | HDHP from a warm start with the family ceiling in reach | 60.7 | 0.488 |