Coverage Cat AI Insurance Benchmark

Which AI models understand insurance work?

We evaluate frontier models on two separate insurance tasks: price estimation for anonymized umbrella quote rows, and brokerage/agent task reasoning against benchmark reference answers for underwriting, eligibility, and coverage questions.

Leaderboard

Brokerage/agent task performance

Brokerage/agent task results are judged against benchmark reference answers and ranked separately from price estimation, because they measure answer quality rather than premium calibration.

Best task Elo DeepSeek 3.2 1,701 Elo
Best judge avg GLM 5 1.83 / 4
Scored task answers 8,967 1,281 unique task questions

Elo Rating

Pairwise model strength on the same judged brokerage/agent task questions. Higher scores mean a model more often beat comparable models on answer quality.

Higher is better
1,701
DeepSeek logo
DeepSeek 3.2
1,531
Z.ai logo
GLM 5
1,530
Mistral logo
Mistral Large
1,514
Claude logo
Claude Opus 4.7
1,452
OpenAI logo
ChatGPT 5.5
1,434
Moonshot AI logo
Kimi K2.5
1,338
xAI logo
Grok 4.3

Win Rate

The share of pairwise brokerage/agent task battles won, with ties counted as half a win. Higher is better.

Higher is better
49.5%
DeepSeek logo
DeepSeek 3.2
56.1%
Z.ai logo
GLM 5
42.1%
Mistral logo
Mistral Large
47.5%
Claude logo
Claude Opus 4.7
52.5%
OpenAI logo
ChatGPT 5.5
48.4%
Moonshot AI logo
Kimi K2.5
53.9%
xAI logo
Grok 4.3

Task Judge Avg

Average primary judge score on a 0-4 rubric for factual correctness, completeness, and match to the reference answer. Higher is better.

Higher is better
1.65
DeepSeek logo
DeepSeek 3.2
1.83
Z.ai logo
GLM 5
1.18
Mistral logo
Mistral Large
1.39
Claude logo
Claude Opus 4.7
1.60
OpenAI logo
ChatGPT 5.5
1.42
Moonshot AI logo
Kimi K2.5
1.66
xAI logo
Grok 4.3

Model comparison

Brokerage/agent task ranking

Rank Model Elo Win rate Judge avg Scored answers Record
1
DeepSeek logo
DeepSeek 3.2 DeepSeek
1,701 49.5% 1.65 1,281 2046-2124-3516
2
Z.ai logo
GLM 5 Z.ai
1,531 56.1% 1.83 1,281 2288-1354-4044
3
Mistral logo
Mistral Large Mistral
1,530 42.1% 1.18 1,281 1323-2532-3831
4
Claude logo
Claude Opus 4.7 Claude
1,514 47.5% 1.39 1,281 1533-1916-4237
5
OpenAI logo
ChatGPT 5.5 OpenAI
1,452 52.5% 1.60 1,281 2036-1649-4001
6
Moonshot AI logo
Kimi K2.5 Moonshot AI
1,434 48.4% 1.42 1,281 1541-1792-4353
7
xAI logo
Grok 4.3 xAI
1,338 53.9% 1.66 1,281 2232-1632-3822

Eval examples

Two different benchmark tasks

Price-estimation rows are scored against actual quote outcomes. Brokerage/agent task rows are scored against reference answers and judged separately, so their leaderboard should be read as answer-quality performance rather than premium-estimation performance.

Price-estimation examples

  • Estimate the annual premium and uncertainty range for an anonymized $1M California umbrella quote from a specific carrier.
  • Given state, carrier, coverage limit, and anonymized risk features, return calibrated P10/P50/P90 premium estimates.
  • Predict a quote range that contains the actual annualized premium without making the interval unnecessarily wide.

Brokerage/agent task examples

  • A household has a listed underwriting profile. Are they likely to be eligible with a specific umbrella carrier?
  • In Texas, how much more does moving from $1M to $2M of umbrella coverage typically cost with a named carrier?
  • A customer asks about coverage requirements or eligibility constraints. What should an assistant say, using the benchmark reference answer?

Methodology

Domain-specific, aggregate-only benchmarking

General AI benchmarks rarely measure whether a model can reason through the details that matter in insurance: liability limits, carrier constraints, premium ranges, eligibility rules, and uncertainty. This benchmark focuses on those workflows.

Price scoring

Quote rows compare each model's estimated annual premium and range against the actual quote outcome. Coverage rewards calibrated ranges; MAPE rewards accurate point estimates; Winkler loss penalizes ranges that miss the actual quote or are too wide.

Brokerage/agent task scoring

Brokerage/agent task rows compare model answers to benchmark reference answers with AI judging. The public task view reports aggregate judge scores, pairwise Elo, and win rate only.

How Elo works

For each shared scenario or question, every pair of model outputs is compared. Better outputs win the local battle, ties split credit, and Elo updates model strength within that benchmark section.

Data protection

Public results are aggregate-only. The page does not expose raw prompts, row identifiers, model responses, judge reasoning, or any operational eval artifacts. The evals use anonymized data on no-retention and no-logging platforms, so customer data is never exposed even to model providers.