Problem 06 · Calibration
Calibrated on average.
Prediction markets are defended as well calibrated. Calibration varies enormously by domain, horizon, and liquidity. The average is doing a lot of hiding.
Simple assumes nothing
The standard defence of prediction markets is that when they say 70%, the thing happens about 70% of the time. Averaged over everything, that is roughly true.
Averages hide a lot. A heavily traded sports market with unambiguous rules behaves nothing like a thin market on a geopolitical event a year out. The first is close to a fair price. The second can be badly wrong and still sit inside the overall average, because the good markets carry it.
For anyone who wants to use a price to make a decision, the average is useless. You need to know whether this market, on this kind of question, at this liquidity, is worth trusting. Nobody can currently tell you.
Moderate assumes you know what a market is
Aggregate calibration is the wrong statistic to publish. It describes a portfolio of markets and gets quoted as though it described each one.
Le (2026), "Decomposing Crowd Wisdom", is the state of the art. Across 292 million Kalshi and Polymarket trades it breaks calibration into four components that together explain 87% of the variance. That is the shape of the answer, and what it establishes is that calibration is predictable from observable market features. A per-market trust score is therefore buildable, and nobody has built one.
The missing pieces are practical. Time to resolution interacts with everything, since long-horizon markets have unresolved information and thin books. Favourite-longshot bias shows up in some categories and not others. And nobody has produced the thing a decision maker actually needs, which is a function from market features to an expected error bar on the price.
Technical state of the art and the gap
Calibration should be estimated conditionally. Instead of
E[outcome | price] = price globally, estimate it per
stratum: category, liquidity depth, time to resolution, participant
composition, resolution-source type. Le's four-component decomposition is
the reference point, and the open work is turning a variance
decomposition into a predictive model with out-of-sample validation.
Two specific gaps. Nobody has separated calibration error caused by thin
liquidity from error caused by resolution ambiguity, and those call for
different fixes, one financial and one editorial. And no published work
conditions on agent share, which connects straight to the correlation
problem: if ρ is rising, calibration in agent-dense
categories should be degrading measurably, and that is a testable
prediction sitting in public data right now.
Deliverable shape: a trust score served per market, of the form "expected absolute deviation from the true probability, given these features". Any buyer of this signal needs that number before they can price the signal at all, which puts this problem underneath the revenue problem rather than beside it.
Where I would start
- Read Le (2026) and reproduce the decomposition on a smaller sample. Reproduction tells you quickly whether the data is workable.
- Build the crudest possible trust score. Regress absolute price error on liquidity, time to resolution, and category. Validate out of sample. Crude and validated beats sophisticated and unvalidated.
- Add agent share as a feature. That is the link to problem 04 and nobody has tested it.
- Ship it as an API before writing it up. A score people can query gets used, and use produces feedback that a paper never will.
What counts as a result
A per-market expected error, validated out of sample. That is the missing prerequisite for anyone paying for prediction market data, and it is the most directly useful thing on this list.
Related
- 02 Nobody buys the answer the problem this unblocks
- 03 The price moves the thing it prices what calibration cannot detect
- 05 Most of the volume is decoration clean the inputs first