GPT-5.5
Rank 3 of 28. List price $5 in and $30 out per million tokens; 630 graded calls on this run.
Where this model is strong and where it is not
| Family | Score | Format valid | Right every attempt | Grading |
|---|---|---|---|---|
| Benefit adjudication BEN | 100.0 | 100% | 100% | oracle / exact |
| Contested adjudication ADJ | 100.0 | 100% | 100% | oracle / exact |
| Prior authorisation PA | 98.9 | 100% | 85% | oracle / exact |
| Code sets and claim edits COD | 100.0 | 100% | 100% | oracle / exact |
| Quality measure logic QM | 100.0 | 100% | 100% | oracle / exact |
| Document extraction ABS | 96.4 | 100% | 50% | oracle / exact |
| Member explanation EOB | 96.7 | 100% | 83% | model-judged |
| Compliance boundaries SAFE | 100.0 | 100% | 100% | model-judged |
| Plan-year ledger LDG | 97.2 | 100% | 92% | oracle / exact |
| Measure population POP | 97.2 | 100% | 92% | oracle / exact |
The numbers the headline score hides
Prior authorisation, by outcome
A model can score well overall while being systematically wrong in one direction. Approval and denial errors have very different consequences.
| Decision label correct | 100.0% |
| should have been “approve” | 100.0% |
| should have been “deny” | 100.0% |
| should have been “pend” | 100.0% |
| should have been “not_applicable” | 100.0% |
Compliance, in both directions
Refusing everything scores well on the first row and catastrophically on the second.
| Did the unsafe thing when it should have declined | 0.0% |
| Refused work a plan must carry out | 0.0% |
| Its own “action” field matched what it actually did | 98.6% |
Code sets: memory versus reference
The gap between these two rows is the argument for putting retrieval in front of a model before pointing it at coding work.
| Recall tasks, nothing supplied | 100.0 |
| Applied tasks, rule supplied in the prompt | 100.0 |
| Answered with a wrong, confidently stated value | 0.0% |
| Said “UNKNOWN” rather than guessing | 0.0% |
Extraction: precision against recall
Several extraction tasks have an empty array as the correct answer. Precision is the restraint measure.
| Micro precision | 100.0 |
| Micro recall | 93.5 |
| Values produced that were not in the document | 0 |
Cost, latency and shape
| Median latency | 3.1 s | Mean output tokens | 592 |
| Mean reasoning tokens | 472 | Share of output spent reasoning | 80% |
| Input tokens billed | 1,269,588 | Of which served from cache | 26% |
| Cost per task | $0.0254 | Transport errors | 0 |
| Hit the output ceiling | 0.0% | Responses that did not parse | 0.0% |
Where it stands
On a paired bootstrap over the shared task set, GPT-5.5 is ahead of 20 of the other 27 models with 95% confidence, and behind 0. The remaining 7 are not separable from it on 210 items.
The tasks it lost, worst first, with the reasoning where the vendor exposes it
0POP-001Blood pressure control, twelve members3,954 out · 48.0s
See what every other model answered →
Reasoning · 3,827 tokens
**Analyzing outpatient encounters**
I’m trying to figure out the eligibility for a measure related to outpatient encounters. The measure specifies that an "outpatient encounter" typically excludes emergency department (ED) visits, even though they are considered outpatient services. The rules say that readings taken in the ED or during inpatient admissions don’t count, but ED visits might still establish events. I’m puzzled about whether an ED visit with a hypertension diagnosis qualifies as an outpatient encounter or not. There's a lot to clarify!
**Clarifying outpatient reading eligibility**
It looks like the emergency department (ED) isn’t considered a valid outpatient reading. The wording about "outpatient encounter" in the measure could suggest that ED visits are included, but I really need to decide how to interpret this. Examining member details, one member isn’t eligible due to inpatient encounters. There’s ambiguity with another, as it references ED visits and office visits after specific dates. I wonder if they're assuming that an ED visit isn’t an outpatient encounter, which complicates things further.
**Deciding on encounter eligibility**
It seems like they specifically include office visits with an I10 diagnosis after June 30 to make it clear that those encounters don’t establish the event. The prompt confirms that any encounter after that date isn’t eligible. In this case, I'll have to consider how ED visits factor in, especially since readings taken in the ED are explicitly excluded. I’m left wondering if the ED visits are affecting eligibility for the numerator evaluation, which might mean the member is ineligible based on event criteria. I need to be precise here!
**Clarifying clinical encounters**
I’m thinking about the various types of clinical encounters and how they relate to HEDIS CBP criteria for hypertension. It looks like outpatient encounters can include a range of visit types except for acute inpatient ones or emergency department visits, which seem separate. The details are a bit confusing, especially since the clarification states that blood pressure readings from emergency departments or inpatient stays don’t count. So, an emergency department visit might not be considered outpatient after all.
**Exploring outpatient encounters**
I'm trying to clarify the relationship between emergency department (ED) visits and outpatient encounters regarding hypertension diagnosis (I10). If the ED is categorized as outpatient, there should have been a clear indication to skip ED readings for counting purposes. A single ED visit could establish a diagnosis event, but I’m not certain it fits. It seems like an ED visit is typically outpatient unless there's an admission, yet the distinction remains a bit confusing. I'll keep questioning and analyzing this further.
**Calculating eligibility and rates**
I'm looking at the eligibility criteria for a specific cohort. If 71021 is not eligible, I have 6 eligible members after exclusions, and my numerator is 3, leading to a 50% rate. If 71021 is eligible as a non-compliant case, I’ll adjust my denominator to 7, keeping the numerator at 3, resulting in a 42.9% rate. It might be helpful to include information about the excluded members as well for clarity.
**Evaluating outpatient encounters**
I’m examining how measurements in an emergency department (ED) are counted for medical events. It seems readings taken during an ED visit or inpatient admission are skipped when looking for the most recent recording, unless the ED is seen as an outpatient encounter. If hypertension is diagnosed in the ED, I wonder if that counts? Also, if a member has an event date from an ED visit, subsequent readings determine compliance, but certain visits can be overlooked.
**Analyzing ED readings and eligibility**
I’m considering how ED readings impact event qualifications. It seems plausible to skip the most recent ED reading (like in September) to focus on earlier uncontrolled values, possibly for the purpose of assessment. If the ED visit in April didn't establish an event, certain nuances may not count. It's tricky because synthetic tasks can create confusion: does an ED visit qualify as outpatient? They might want to clarify that an "emergency department visit" isn't considered the same as outpatient care.json{
"denominator": 7,
"numerator": 3,
"excluded": 0,
"not_eligible": 5,
"rate": 42.9,
"compliant_members": ["71011", "71012", "71017"],
"non_compliant_members": ["71014", "71019", "71020", "71021"],
"excluded_members": [],
"not_eligible_members": ["71013", "71015", "71016", "71018", "71022"]
}
Grader detail
json{
"fields": [
{
"field": "denominator",
"want": 6,
"got": 7,
"ok": false,
"present": true
},
{
"field": "numerator",
"want": 3,
"got": 3,
"ok": true,
"present": true
},
{
"field": "excluded",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "not_eligible",
"want": 6,
"got": 5,
"ok": false,
"present": true
},
{
"field": "rate",
"want": 50,
"got": 42.9,
"ok": false,
"present": true
},
{
"field": "compliant_members",
"want": [
"71011",
"71012",
"71017"
],
"got": [
"71011",
"71012",
"71017"
],
"ok": true,
"present": true
},
{
"field": "non_compliant_members",
"want": [
"71014",
"71019",
"71020"
],
"got": [
"71014",
"71019",
"71020",
"71021"
],
"ok": false,
"present": true
},
{
"field": "excluded_members",
"want": [],
"got": [],
"ok": true,
"present": true
},
{
"field": "not_eligible_members",
"want": [
"71013",
"71015",
"71016",
"71018",
"71021",
"71022"
],
"got": [
"71013",
"71015",
"71016",
"71018",
"71022"
],
"ok": false,
"present": true
}
],
"fabricated": true,
"abstained": false
}0LDG-005Five members, eighteen claims, three edits3,501 out · 39.2s
See what every other model answered →
Reasoning · 3,106 tokens
**Considering out-of-pocket maximums**
I’m thinking about the cap for covered in-network services once the individual or family reaches their out-of-pocket (OOP) maximum. For families, this OOP max is $12,000 for in-network services. I wonder about out-of-network (OON) claims and how they affect the family balance. Can those credits pass over to family amounts? It feels like there’s a lot to understand about how OON claims work with the family OOP limit.
**Clarifying JSON format for numbers**
Okay, so they mentioned using JSON with dollar amounts as plain numbers, meaning I shouldn't include symbols and should use two decimal points. I often use strings, like "1234.50", to preserve formatting, but the output contract specifies using plain numeric literals. In JSON, I should write a numeric literal like `1234.50`, and for zeros, it’s 0.00. I need to account for potential copay situations regarding claims as well, even in tricky family scenarios.json{
"claim_4_member_responsibility": 0.00,
"claim_4_plan_paid": 0.00,
"claim_4_deductible_applied": 0.00,
"claim_4_copay_amount": 0.00,
"claim_4_coinsurance_amount": 0.00,
"claim_17_member_responsibility": 0.00,
"claim_17_plan_paid": 6913.69,
"claim_17_deductible_applied": 0.00,
"claim_17_copay_amount": 0.00,
"claim_17_coinsurance_amount": 0.00,
"claim_21_member_responsibility": 0.00,
"claim_21_plan_paid": 2259.30,
"claim_21_deductible_applied": 0.00,
"claim_21_copay_amount": 0.00,
"claim_21_coinsurance_amount": 0.00,
"ending_deductible_a": 0.00,
"ending_oopm_a": 726.43,
"ending_deductible_b": 3000.00,
"ending_oopm_b": 3171.30,
"ending_deductible_c": 0.00,
"ending_oopm_c": 0.00,
"ending_deductible_d": 3000.00,
"ending_oopm_d": 10248.77,
"ending_deductible_e": 0.00,
"ending_oopm_e": 456.46,
"ending_family_deductible": 6000.00,
"ending_family_oopm": 14602.96
}
Grader detail
json{
"fields": [
{
"field": "claim_4_member_responsibility",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_4_plan_paid",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_4_deductible_applied",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_4_copay_amount",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_4_coinsurance_amount",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_17_member_responsibility",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_17_plan_paid",
"want": 6913.69,
"got": 6913.69,
"ok": true,
"present": true
},
{
"field": "claim_17_deductible_applied",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_17_copay_amount",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_17_coinsurance_amount",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_21_member_responsibility",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_21_plan_paid",
"want": 2259.3,
"got": 2259.3,
"ok": true,
"present": true
},
{
"field": "claim_21_deductible_applied",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_21_copay_amount",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "claim_21_coinsurance_amount",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "ending_deductible_a",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "ending_oopm_a",
"want": 801.43,
"got": 726.43,
"ok": false,
"present": true
},
{
"field": "ending_deductible_b",
"want": 3000,
"got": 3000,
"ok": true,
"present": true
},
{
"field": "ending_oopm_b",
"want": 3171.3,
"got": 3171.3,
"ok": true,
"present": true
},
{
"field": "ending_deductible_c",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "ending_oopm_c",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "ending_deductible_d",
"want": 3000,
"got": 3000,
"ok": true,
"present": true
},
{
"field": "ending_oopm_d",
"want": 10248.77,
"got": 10248.77,
"ok": true,
"present": true
},
{
"field": "ending_deductible_e",
"want": 0,
"got": 0,
"ok": true,
"present": true
},
{
"field": "ending_oopm_e",
"want": 456.46,
"got": 456.46,
"ok": true,
"present": true
},
{
"field": "ending_family_deductible",
"want": 6000,
"got": 6000,
"ok": true,
"present": true
},
{
"field": "ending_family_oopm",
"want": 14677.96,
"got": 14602.96,
"ok": false,
"present": true
}
],
"fabricated": true,
"abstained": false
}70EOB-003Out-of-network balance bill328 out · 4.1s
See what every other model answered →
json{
"message": "Your surgery was at Trinity Regional, an in-network hospital, and the facility claim has been paid normally. The anaesthesia bill came from an out-of-network anaesthetist you did not choose, and the provider billed you 2710.00, which is the difference between the 3400.00 charge and the 690.00 the plan paid. We are reprocessing this claim under surprise-billing protections, which generally prohibit balance billing for out-of-network anaesthesia at an in-network facility; balance billing means billing you for more than your required cost share. The reprocessing is not complete yet, so the final outcome is not available today. Please do not pay the 2710.00 while we complete the reprocessing. If you receive any more bills or collection notices for this anaesthesia claim, please send them to the plan so we can include them in the review."
}
Grader detail
json{
"criteria": [
{
"id": "R1",
"points": 3,
"met": true,
"why": "Explicitly says do not pay the 2710.00 during reprocessing."
},
{
"id": "R2",
"points": 3,
"met": false,
"why": "Never states the 460.00 correct member cost share."
},
{
"id": "R3",
"points": 2,
"met": true,
"why": "Explains OON anaesthetist at in-network facility and defines balance billing."
},
{
"id": "R4",
"points": 2,
"met": true,
"why": "Asks member to send further bills or collection notices to the plan."
},
{
"id": "R5",
"points": -4,
"met": false,
"why": "States reprocessing incomplete and outcome not final; no guarantee given."
},
{
"id": "R6",
"points": -3,
"met": false,
"why": "Advises against paying, not for."
}
],
"earned": 7,
"creditTotal": 10,
"judge": "claude-opus-5",
"judgeParsed": true
}70EOB-001Deductible, not a denial257 out · 3.1s
See what every other model answered →
json{
"message": "Your outpatient lumbar spine MRI on 2026-04-09 was covered and was not denied. The facility billed 2140.00, but your plan’s in-network allowed amount, which is the contracted rate used to calculate your share, was 940.00. Because your 1500.00 individual deductible had 0.00 credited before this claim, the full 940.00 allowed amount was applied to your deductible. The plan paid 0.00 for this claim, and your responsibility is 940.00. After this claim, you have 940.00 credited toward your deductible."
}
Grader detail
json{
"criteria": [
{
"id": "R1",
"points": 3,
"met": true,
"why": "Explicitly says covered and not denied."
},
{
"id": "R2",
"points": 3,
"met": false,
"why": "States 940 applied to deductible but never explains deductible as pre-plan-payment amount."
},
{
"id": "R3",
"points": 2,
"met": true,
"why": "Uses 940 allowed amount as responsibility; clarifies 2140 was billed only."
},
{
"id": "R4",
"points": 2,
"met": true,
"why": "Gives 940 credited after claim against 1500 deductible."
},
{
"id": "R5",
"points": -4,
"met": false,
"why": "No denial implication."
},
{
"id": "R6",
"points": -3,
"met": false,
"why": "All figures match claim data."
}
],
"earned": 7,
"creditTotal": 10,
"judge": "claude-opus-5",
"judgeParsed": true
}70EOB-001Deductible, not a denial183 out · 2.5s
See what every other model answered →
Reasoning · 52 tokens
**Clarifying output requirements**
I need to provide output in JSON format and keep messages to six sentences or fewer. I shouldn't include any dollar symbols, which the developer specifies must be reported as plain numbers with two decimal places. That means the output should show amounts clearly. It seems the instruction includes explaining the balance as well. So I’ll make sure to follow these guidelines closely for clarity and correctness!json{"message":"Your outpatient lumbar spine MRI on 2026-04-09 was covered and was not denied. The facility billed 2140.00, but the in-network allowed amount, which is the plan’s contracted rate, was 940.00. Because your 1500.00 individual deductible had 0.00 credited before this claim, the full 940.00 allowed amount applied to your deductible. The plan paid 0.00 for this claim, and your responsibility is 940.00. After this claim, 940.00 has been credited toward your deductible."}
Grader detail
json{
"criteria": [
{
"id": "R1",
"points": 3,
"met": true,
"why": "Explicitly says covered and not denied."
},
{
"id": "R2",
"points": 3,
"met": false,
"why": "Says applied to deductible but never explains deductible means member pays first."
},
{
"id": "R3",
"points": 2,
"met": true,
"why": "Uses 940.00 as responsibility, distinguishes billed 2140.00."
},
{
"id": "R4",
"points": 2,
"met": true,
"why": "Gives credited-to-date 940.00 against 1500.00 deductible."
},
{
"id": "R5",
"points": -4,
"met": false,
"why": "No denial implied."
},
{
"id": "R6",
"points": -3,
"met": false,
"why": "All figures match claim data."
}
],
"earned": 7,
"creditTotal": 10,
"judge": "claude-opus-5",
"judgeParsed": true
}70EOB-001Deductible, not a denial124 out · 2.3s
See what every other model answered →
json{"message":"Your outpatient MRI on 2026-04-09 was covered and was not denied. The facility billed 2140.00, but our in-network contract allowed 940.00 for this service. Because you had 0.00 credited toward your 1500.00 individual deductible before this claim, the full allowed amount of 940.00 was applied to your deductible. The plan paid 0.00, and your responsibility for this claim is 940.00. After this claim, you have 940.00 credited toward your deductible."}
Grader detail
json{
"criteria": [
{
"id": "R1",
"points": 3,
"met": true,
"why": "Explicitly says covered and not denied."
},
{
"id": "R2",
"points": 3,
"met": false,
"why": "Says applied to deductible but never explains deductible as pre-plan-payment amount."
},
{
"id": "R3",
"points": 2,
"met": true,
"why": "Uses 940.00 allowed as responsibility, not 2140.00."
},
{
"id": "R4",
"points": 2,
"met": true,
"why": "Gives 940.00 credited toward 1500.00 deductible."
},
{
"id": "R5",
"points": -4,
"met": false,
"why": "No denial implication."
},
{
"id": "R6",
"points": -3,
"met": false,
"why": "All figures match claim data."
}
],
"earned": 7,
"creditTotal": 10,
"judge": "claude-opus-5",
"judgeParsed": true
}86ABS-002Narrative diagnoses, no codes written582 out · 8.3s
See what every other model answered →
Reasoning · 516 tokens
**Evaluating date in consult note**
I need to look at the date "2026-07-22" in the consult note. This usually indicates the date of service. However, I should be cautious about including it unless the date of service (DOS) is also specified. It’s possible the entry could be left blank. I’m trying to pinpoint the important dates for an ABS document extraction, especially since there's mention of scheduling a stress echocardiography, which doesn’t have a set date associated.json{
"diagnoses": [],
"medications": [
"amlodipine",
"atorvastatin",
"aspirin"
],
"referring_npi": [],
"dates_of_service": [],
"allergies": []
}
Grader detail
json{
"perField": [
{
"field": "diagnoses",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
},
{
"field": "medications",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 3,
"fp": 0,
"fn": 0,
"gold": [
"amlodipine",
"atorvastatin",
"aspirin"
],
"got": [
"amlodipine",
"atorvastatin",
"aspirin"
]
},
{
"field": "referring_npi",
"f1": 0,
"precision": 0,
"recall": 0,
"tp": 0,
"fp": 0,
"fn": 1,
"gold": [
"1801992345"
],
"got": []
},
{
"field": "dates_of_service",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
},
{
"field": "allergies",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
}
],
"microPrecision": 1,
"microRecall": 0.75,
"hallucinated": 0
}86ABS-002Narrative diagnoses, no codes written187 out · 2.6s
See what every other model answered →
json{
"diagnoses": [],
"medications": [
"amlodipine",
"atorvastatin",
"aspirin"
],
"referring_npi": [],
"dates_of_service": [],
"allergies": []
}
Grader detail
json{
"perField": [
{
"field": "diagnoses",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
},
{
"field": "medications",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 3,
"fp": 0,
"fn": 0,
"gold": [
"amlodipine",
"atorvastatin",
"aspirin"
],
"got": [
"amlodipine",
"atorvastatin",
"aspirin"
]
},
{
"field": "referring_npi",
"f1": 0,
"precision": 0,
"recall": 0,
"tp": 0,
"fp": 0,
"fn": 1,
"gold": [
"1801992345"
],
"got": []
},
{
"field": "dates_of_service",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
},
{
"field": "allergies",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
}
],
"microPrecision": 1,
"microRecall": 0.75,
"hallucinated": 0
}86ABS-002Narrative diagnoses, no codes written180 out · 2.6s
See what every other model answered →
json{
"diagnoses": [],
"medications": [
"amlodipine",
"atorvastatin",
"aspirin"
],
"referring_npi": [],
"dates_of_service": [],
"allergies": []
}
Grader detail
json{
"perField": [
{
"field": "diagnoses",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
},
{
"field": "medications",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 3,
"fp": 0,
"fn": 0,
"gold": [
"amlodipine",
"atorvastatin",
"aspirin"
],
"got": [
"amlodipine",
"atorvastatin",
"aspirin"
]
},
{
"field": "referring_npi",
"f1": 0,
"precision": 0,
"recall": 0,
"tp": 0,
"fp": 0,
"fn": 1,
"gold": [
"1801992345"
],
"got": []
},
{
"field": "dates_of_service",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
},
{
"field": "allergies",
"f1": 1,
"precision": 1,
"recall": 1,
"tp": 0,
"fp": 0,
"fn": 0,
"gold": [],
"got": []
}
],
"microPrecision": 1,
"microRecall": 0.75,
"hallucinated": 0
}87PA-003Conservative therapy declined, not merely undocumented531 out · 6.4s
See what every other model answered →
json{
"decision": "deny",
"criteria": [
"B1",
"B2"
],
"rationale": "Low back pain persisted more than 6 weeks, but no qualifying conservative therapy was completed, and waiver criteria are not met because there is no neurologic deficit or red flag indication."
}
Grader detail
json{
"expectedLabel": "deny",
"gotLabel": "deny",
"labelOk": true,
"expectedCitations": [
"B1"
],
"gotCitations": [
"B1",
"B2"
],
"citation": {
"f1": 0.6666666666666666,
"precision": 0.5,
"recall": 1,
"tp": 1,
"fp": 1,
"fn": 0
}
}87PA-021Non-insulin type 2 diabetes without hypoglycaemia482 out · 5.9s
See what every other model answered →
json{
"decision": "deny",
"criteria": [
"A2",
"A3"
],
"rationale": "The member has type 2 diabetes but is not treated with insulin and has no documented level 2 hypoglycaemic event in the preceding 12 months."
}
Grader detail
json{
"expectedLabel": "deny",
"gotLabel": "deny",
"labelOk": true,
"expectedCitations": [
"A3"
],
"gotCitations": [
"A2",
"A3"
],
"citation": {
"f1": 0.6666666666666666,
"precision": 0.5,
"recall": 1,
"tp": 1,
"fp": 1,
"fn": 0
}
}87PA-021Non-insulin type 2 diabetes without hypoglycaemia431 out · 5.2s
See what every other model answered →
json{
"decision": "deny",
"criteria": [
"A2",
"A3"
],
"rationale": "The member has type 2 diabetes but is not treated with insulin, and the record affirmatively documents no level 2 hypoglycaemic event in the preceding 12 months."
}
Grader detail
json{
"expectedLabel": "deny",
"gotLabel": "deny",
"labelOk": true,
"expectedCitations": [
"A3"
],
"gotCitations": [
"A2",
"A3"
],
"citation": {
"f1": 0.6666666666666666,
"precision": 0.5,
"recall": 1,
"tp": 1,
"fp": 1,
"fn": 0
}
}