Problem 03 · Reflexivity
The price moves the thing it prices.
A market says the bank fails with 90% probability. Depositors see it and withdraw. The bank fails. The forecast was correct and it was also the cause.
Simple assumes nothing
Prediction markets get treated as thermometers. They are closer to thermostats. A thermometer reads the temperature. A thermostat reads it and then changes it.
Once a probability is public, people act on it. Donors abandon a candidate sitting at 5%. Depositors pull money out of a bank the market says is failing. Journalists write the story the price implies, and the story moves the price again. The forecast becomes part of the machinery producing the outcome.
Sometimes that is good. A fire alarm makes fire less likely, and a market warning of a shortage can cause the shortage to be prevented. Sometimes it is a self-fulfilling collapse. Nobody can currently tell you which one a given market is, in advance. That is the problem.
Moderate assumes you know what a market is
The mechanism is a loop. Price feeds belief, belief feeds action, action feeds outcome, outcome feeds resolution, resolution feeds price. Standard market theory assumes the link from price back to outcome is absent. In prediction markets it is often present and occasionally strong.
The sign matters more than the size. Negative feedback damps, so the market converges and looks like a good forecaster. Positive feedback amplifies, so the market can drive the outcome and still score as perfectly calibrated, because it caused what it predicted. A market that causes 100% of its own outcomes has flawless calibration.
That breaks the standard defence of the whole field. Accuracy is offered as proof of value, and accuracy is exactly what a self-fulfilling market would display. Nobody has separated the two empirically. The design question of how to damp the loop deliberately is untouched.
Technical state of the art and the gap
Let p be the real outcome probability, q the
market price, and f a response function capturing how
agents' actions change the real probability, so p = f(q).
Equilibrium requires q = f(q). Existence follows from
continuity on the unit interval by Brouwer. Uniqueness does not. Where
f' > 1, multiple equilibria exist and the market's
selection among them is path dependent, which means it is set by early
liquidity and order flow rather than by information.
Three pieces are open. Estimating f' empirically for real
categories, which needs an instrument for exogenous price movement. Bank
runs, candidate viability, and protocol governance are where
f' is plausibly above 1. Designing contracts that damp the
loop on purpose, for instance resolving on a measurement window that
closes before the price is published, so actions taken in response
cannot enter resolution. And the identification problem of separating
"predicted well" from "caused it", which needs something like the
Rasooly and Rozzi manipulation design applied to causal rather than
manipulative price moves.
The adversarial version is already documented. Prediction Laundering (arXiv 2602.05181) sets out four stages by which a messy bet gets scrubbed into authoritative truth: structural sanitization, probabilistic flattening, architectural masking, epistemic hardening. That is reflexivity with an interface on top, and almost nobody in this space has engaged with the paper.
Where I would start
- Pick one category where the loop is plausibly strong. Protocol governance votes are the most accessible, since the price, the vote, and the outcome are all on-chain.
- Find an instrument for exogenous price movement. Large uninformed orders, liquidity shocks, or resolution of a correlated but unrelated market. Without one you cannot separate prediction from causation and the study is not worth running.
-
Estimate the sign and rough magnitude of
f'on that one category. Even a crude estimate is new. - On the design side, write down one contract that damps the loop and check whether it is still tradeable. A damped contract nobody wants to trade is not a solution.
What counts as a result
One category with an estimated response slope and an error bar. If
f' exceeds 1 anywhere, that has immediate consequences for
how large platforms should let those markets get.
Related
- 06 Calibrated on average why calibration cannot detect this
- 04 Thousands of traders, one opinion what removes the damping
- 08 The thermometer costs $34,000 the same loop through a different channel