Adam Szymański

Problem 03 · Reflexivity

The price moves the thing it prices.

A market says the bank fails with 90% probability. Depositors see it and withdraw. The bank fails. The forecast was correct and it was also the cause.

Theory Empirical Hard identification

Simple assumes nothing

Prediction markets get treated as thermometers. They are closer to thermostats. A thermometer reads the temperature. A thermostat reads it and then changes it.

Once a probability is public, people act on it. Donors abandon a candidate sitting at 5%. Depositors pull money out of a bank the market says is failing. Journalists write the story the price implies, and the story moves the price again. The forecast becomes part of the machinery producing the outcome.

Sometimes that is good. A fire alarm makes fire less likely, and a market warning of a shortage can cause the shortage to be prevented. Sometimes it is a self-fulfilling collapse. Nobody can currently tell you which one a given market is, in advance. That is the problem.

Moderate assumes you know what a market is

The mechanism is a loop. Price feeds belief, belief feeds action, action feeds outcome, outcome feeds resolution, resolution feeds price. Standard market theory assumes the link from price back to outcome is absent. In prediction markets it is often present and occasionally strong.

The sign matters more than the size. Negative feedback damps, so the market converges and looks like a good forecaster. Positive feedback amplifies, so the market can drive the outcome and still score as perfectly calibrated, because it caused what it predicted. A market that causes 100% of its own outcomes has flawless calibration.

That breaks the standard defence of the whole field. Accuracy is offered as proof of value, and accuracy is exactly what a self-fulfilling market would display. Nobody has separated the two empirically. The design question of how to damp the loop deliberately is untouched.

Technical state of the art and the gap

Let p be the real outcome probability, q the market price, and f a response function capturing how agents' actions change the real probability, so p = f(q). Equilibrium requires q = f(q). Existence follows from continuity on the unit interval by Brouwer. Uniqueness does not. Where f' > 1, multiple equilibria exist and the market's selection among them is path dependent, which means it is set by early liquidity and order flow rather than by information.

Three pieces are open. Estimating f' empirically for real categories, which needs an instrument for exogenous price movement. Bank runs, candidate viability, and protocol governance are where f' is plausibly above 1. Designing contracts that damp the loop on purpose, for instance resolving on a measurement window that closes before the price is published, so actions taken in response cannot enter resolution. And the identification problem of separating "predicted well" from "caused it", which needs something like the Rasooly and Rozzi manipulation design applied to causal rather than manipulative price moves.

The adversarial version is already documented. Prediction Laundering (arXiv 2602.05181) sets out four stages by which a messy bet gets scrubbed into authoritative truth: structural sanitization, probabilistic flattening, architectural masking, epistemic hardening. That is reflexivity with an interface on top, and almost nobody in this space has engaged with the paper.

Where I would start

  1. Pick one category where the loop is plausibly strong. Protocol governance votes are the most accessible, since the price, the vote, and the outcome are all on-chain.
  2. Find an instrument for exogenous price movement. Large uninformed orders, liquidity shocks, or resolution of a correlated but unrelated market. Without one you cannot separate prediction from causation and the study is not worth running.
  3. Estimate the sign and rough magnitude of f' on that one category. Even a crude estimate is new.
  4. On the design side, write down one contract that damps the loop and check whether it is still tradeable. A damped contract nobody wants to trade is not a solution.

What counts as a result

One category with an estimated response slope and an error bar. If f' exceeds 1 anywhere, that has immediate consequences for how large platforms should let those markets get.

Related