NeurIPS 2026 Poster

From Uncertain Judgments to Calibrated Rankings: Conformal Elo Estimation for LLM Evaluation

1 ELLIS Institute Tübingen   2 OpenEuroLLM
ELLIS Institute Tübingen OpenEuroLLM

Estimating Elo ratings with an AI judge

Language models are often compared by asking people which of two answers they prefer. These votes become Elo ratings: scores that order models by how often they are preferred. Collecting enough votes for each new model is expensive.

An AI model can judge the answers instead. But the usual rating calculation counts only wins, ties, and losses—so a close decision counts just as much as a clear win. The resulting scores can put models in roughly the right order while exaggerating the gaps between them.

Soft-Elo uses the judge’s score difference to estimate human preferences, then feeds those probabilities into the same rating calculation.

Paper Figure 1b: Soft-Elo estimates lie closer to the diagonal matching human Elo; Hard-Elo stretches the rating scale.
“Hard-Elo” denotes the usual win/tie/loss calculation; “Soft-Elo” is our method. Each point is a model, and the diagonal marks agreement with human Elo. Figure 1b.

Abstract

Evaluating new large language models typically requires costly human annotation campaigns at scale. LLM-as-a-judge offers a cheaper alternative, but judge scores carry systematic errors—such as position bias, self-preference, or intransitivity—that can strongly miscalibrate the resulting rankings. We quantify the resulting judge–human disagreement at two complementary levels. At the local level, we estimate per-battle uncertainty from the judge’s own score differences by propagating calibrated win probabilities rather than hard labels into the Bradley–Terry procedure. This estimator, Soft-Elo, brings LLM-derived ratings within 17.9 Elo mean absolute error of human-derived ones when averaged over 55 held-out models on LMArena.

At the global level, we apply split conformal prediction to the residual gap between LLM-derived and human-derived Elo ratings across held-out models, producing prediction intervals with distribution-free marginal coverage guarantees under exchangeability that account for irreducible LLM–human disagreement. Together, these two layers yield a low-cost evaluation tool that provides developers with calibrated Elo estimates and honest uncertainty bounds, without requiring large-scale human annotations for each new model.

Score gap → human agreement

When a judge gives two answers almost the same score, its preferred answer is less likely to match the human’s choice. Larger score gaps are more reliable. This relationship is information that a win/loss decision alone discards.

Larger gaps, more human agreement

Human agreement rises with the absolute gap between the judge’s answer scores. Six gap bins, averaged across eight judges.
Eight judges, comparisons where both judge and human choose a winner. Bars show mean agreement ±1 standard deviation.

Match confidence to that agreement

For Qwen3-32B, fitting beta moves predicted confidence closer to observed judge–human agreement along the perfect-calibration diagonal.
Calibration brings predicted confidence closer to observed agreement. Qwen3-32B; the diagonal marks perfect calibration.

From score gaps to Soft-Elo

The score gap therefore tells us more than who won: it indicates how much to trust that decision. Soft-Elo learns this link from human votes, using β to convert each gap into a preference probability. It then fits Elo to these probabilities instead of win/loss labels.

Explore the calculation

One new model, three rated opponents.
01

Score gap

For each opponent, the judge scores both answers to the same question (1–10).

Opponent New answer Their answer
Model AHuman Elo: 1,000 6.3 4.1
Model BHuman Elo: 1,200 6.3 6.1
Model CHuman Elo: 1,400 6.3 8.1

Against B, the score gap is Δs = 6.3 − 6.1 = 0.2.

02

Human preference

pj estimates how likely a human is to prefer the new answer over opponent j’s answer.

pj= 11+e−βΔsj
vs A 75.0%Win
vs B 52.5%Win
vs C 28.9%Loss

With β = 0.50, that 0.2-point gap gives a preference probability of 52.5%.

03

Elo rating

We now find a rating whose predicted win rates match these probabilities. Bradley–Terry links Elo to win rates:

qj(R)= 11+10Rj−R400

R: new model’s Elo · Rj: opponent’s Elo

1,218Fitted Elo · R*
vs A77.8% vs B52.6% vs C26.0%

Its 52.6% predicted win rate against B closely matches the 52.5% from calibration.

The rating fit

R⋆=arg maxR∑j [pjlogqj(R) +(1−pj)log(1−qj(R))]

We search for the Elo rating R that maximizes this sum, keeping opponents’ ratings fixed. The usual calculation uses 1 for a win, ½ for a tie, and 0 for a loss in place of pj. Only these comparison targets change.

Conformal prediction intervals

This gives one estimated Elo rating. AI judges can still disagree with people, so we measure the remaining Elo errors on models with human ratings and use them to set an interval around each new estimate.

Interval example

Learn from the remaining errors

Illustrative errors on calibration models Nineteen model errors, each divided by its bootstrap standard error. The selected quantile determines the prediction interval’s width. 0 1 2 3 Error / bootstrap SE
18th of 19 errors → q = 2.0

Apply that margin to the new model

Illustrative conformal interval for the new model The interval is centered on the fitted Elo. With a bootstrap standard error of 25 Elo, q of 2 gives a margin of 50 Elo. 1,118 1,218 1,318 New model’s Elo 1,218 ± 50 Elo
90% prediction interval · width 100 Elo

Higher coverage → wider intervals.

A 90% coverage target means including the human rating for at least 90% of new models on average, under the same sampling conditions as calibration (exchangeability).

More accurate ratings, tighter uncertainty intervals

Across 55 held-out LMArena models, using calibrated score gaps reduces Elo error for every judge. With less disagreement left between estimated and human ratings, conformal intervals become narrower too.

61% lower rating error

45.9 → 17.9 Elo mean absolute error, averaged across judges

39–70% narrower intervals

Across all eight judges, at the same 90% coverage target

Eight judges, two improvements: Soft-Elo reduces mean absolute rating error by 39–73% and prediction interval width by 39–70%. For DeepSeek-V3.2, error falls from 63.4 to 17.1 Elo and interval width from 261 to 78 Elo. All eight judges improve on both measures.
Tables 2–3. Interval widths are median widths averaged over five calibration/test splits. Lower is better in both panels.

92.1–96.4% observed coverage across the eight judges, against a 90% target.

Model order remains largely unchanged. Soft-Elo makes the gaps between models closer to those measured by human votes, with tighter intervals around each rating.

Cite this work

@inproceedings{kargi2026softelo,
  title     = {From Uncertain Judgments to Calibrated Rankings:
               Conformal Elo Estimation for LLM Evaluation},
  author    = {Kargi, Bora and Salinas, David},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026},
  url       = {https://arxiv.org/abs/2606.13221}
}