CoveraTalk to Covera
Methodology & benchmarks

How accurate is Covera, really?

Four scorecards, each measured against real data. Whether the cost model, trained on past years, predicts a year it has never seen. How the simulation holds up against published MEPS aggregates and the ACA subsidy formula. Whether the recommendation stays safe when a model is in the loop. And how the model itself performs, scored live by a judge agent.

Held-out validation

held-out error 6.9% · train error 4.5% · 16/16 within tolerance

The cost model's parameters are fit to real MEPS microdata from MEPS 2021 + 2022, then scored against the held-out MEPS 2023 (HC-251) year it never saw during fitting. This is the honest test: a model can always be tuned to match the numbers it was tuned on, so the question that matters is whether it predicts a year it has not seen. It does, within single digits, and the held-out error is close to the training error, which is what tells you it is not overfit.

Held-out 2023 (never seen during fitting)

Simulated from params fit on MEPS 2021 + 2022, scored vs real MEPS 2023 (HC-251).

Age 0-17 mean spend
$1,973 vs $2,1669% off
Age 18-44 mean spend
$4,263 vs $3,8969% off
Age 45-64 mean spend
$7,821 vs $8,68910% off
Age 65+ mean spend
$12,353 vs $13,0055% off
Population mean spend
$6,123 vs $6,3704% off
Top 1% share of spend
25.9% vs 22.5%15% off
Top 5% share of spend
54.1% vs 52.9%2% off
Top 10% share of spend
69.4% vs 69.8%1% off

Train fit (reference)

How well the fitted params reproduce their own MEPS 2021 + 2022 source.

Age 0-17 mean spend
$2,019 vs $1,78713% off
Age 18-44 mean spend
$4,310 vs $4,0846% off
Age 45-64 mean spend
$7,487 vs $7,7003% off
Age 65+ mean spend
$12,330 vs $11,5797% off
Population mean spend
$6,044 vs $5,8234% off
Top 1% share of spend
25% vs 25.2%1% off
Top 5% share of spend
53.5% vs 54.2%1% off
Top 10% share of spend
69% vs 70.4%2% off

Each row is the simulated figure vs the real MEPS figure. The model reproduces mean spend by age band and the concentration of spending (the top few percent of people who drive most cost) on a completely unseen year. Source: Real TypeScript simulation (lib/sim) run with params fit on MEPS 2021+2022, scored against held-out MEPS 2023 (HC-251). No key or network.

Alignment & safety

8/8 checks pass · 200 patients + a red-team

Covera's recommendation is not a bare argmin. It runs through a deterministic actor/critic/memory governance layer, and the critic is a hard safety backstop: it will not let an unsafe plan be the headline, even when it is the cheapest. This scorecard measures that guarantee over a synthetic population and against an adversarial red-team, with no model in the loop, so every figure is reproducible.

Governed selection (the critic)

Governed pick carries no hard critic violation
100% vs 100%
Critic blocks an unsafe headline (red-team suite)
100% vs 100%
HSA requirement honored when a compliant plan exists
100% vs 100%
Risk-averse bad-year is bounded
100% vs 95%

Risk-adjusted advice

Recommendation is not blindly the lowest premium
42.5% vs 25%
Risk-adjusted objective keeps the raw ranking clean
100% vs 95%

Reproducible & consistent

Same patient yields the same recommendation
100% vs 100%
All-in cost is never below the premium floor
100% vs 100%

Why it matters. An LLM that recommends insurance can quietly steer someone onto a plan that drops their medication or leaves them exposed to a catastrophic year. Covera puts a deterministic critic between the model and the recommendation: dropped drugs or doctors, an unmet HSA requirement, and a brutal bad year for a risk-averse patient are all hard vetoes. The red-team hands the critic those exact unsafe picks and confirms it blocks every one.

The nuance. Organic critic vetoes over the population were 0: the risk-adjusted objective already avoids unsafe headlines, so the critic rarely has to fire. That is the point of the two-layer design, and the adversarial suite is what proves the backstop still works when it must. No LLM judged anything here. Source: Deterministic evaluation of lib/agents/selection (actor/critic/memory) over a synthetic seeded population, plus an adversarial red-team suite against the critic. No key or network.

Simulation accuracy

16/16 checks within tolerance · 80,000 simulated people

Mean spend by age band (vs MEPS)

Age 0-17 mean spend
$1,985 vs $1,787
Age 18-44 mean spend
$4,207 vs $4,084
Age 45-64 mean spend
$7,784 vs $7,700
Age 65+ mean spend
$12,563 vs $11,579

Spend concentration (vs MEPS)

Top 1% share of spend
25% vs 25%
Top 5% share of spend
54% vs 54%
Top 10% share of spend
69% vs 70%
Bottom 50% share of spend
3% vs 2%
Population mean annual spend
$6,119 vs $5,823
Population median annual spend
$1,102 vs $809

ACA subsidy formula

FPL, household of 1
$15,650 vs $15,650
FPL, household of 4
$32,150 vs $32,150
Applicable % at 150% FPL
0% vs 0%
Applicable % at 250% FPL
4% vs 4%
Applicable % at 400% FPL
8.5% vs 8.5%
Applicable % above 400% FPL (capped)
8.5% vs 8.5%

The engine reproduces the age-band means, the ACA subsidy math, and (via a person-year frailty termthat correlates a year's care across service lines) the real spend concentration the top few percent of people who drive most healthcare cost. That heavy tail is the hard part to get right, and it is what the bad-year risk you rank on depends on. Source: Fit to AHRQ MEPS Household Component Full Year Consolidated microdata: TRAIN = 2021 (HC-233) + 2022 (HC-243), HOLDOUT = 2023 (HC-251).

Explore the distribution

5,000 simulated years, run live in your browser.

$7,781
Simulated mean
MEPS target $7,700
$1,915
Median year
half spend less
47%
Top 5% share
of this group's spend
Simulated annual spend distributionmean$0$63,026+

The long right tail is the whole point: a few people drive most spending, which is why the cheapest plan often is not the safest. Add a condition to watch the mean and tail move.

What real care costs, plan by plan

Each row is a fixed episode of care run through the real adjudication engine on the cheapest plan in each metal tier. Click any number to see the exact math.

Episodes are fixed reference bundles (round allowed amounts), so the math is checkable; the plans and cost-sharing are real CMS PY2026 data. Deductible + coinsurance + copays, capped at the out-of-pocket max, equals the number shown.

Model quality, judged live

Faithfulness is the share of dollar figures in a reply that match a real simulated number instead of a hallucination. Tool accuracy is whether the model reached for the right tool. The rest is scored by a separate judge agent. Press run to evaluate the current model against the suite. Each question runs on its own, so you watch the results come in.

Agents as judge

An agent answers each patient question using Covera's real tools. A judge agent then scores that answer against the real tool outputs. Runs live against the current model, so the numbers are produced fresh, not read from a file.