Calibration In Prediction Markets – Theory and Evidence
Published August 2026
Abstract
Prediction markets are commonly described as producing prices that admit a literal probabilistic interpretation, but this empirical claim has yet to be evaluated at scale on any U.S.-regulated exchange. This paper represents the first such exercise, using the complete history of resolved markets on Kalshi – the world’s largest prediction market – comprising 2,243,741 markets across eleven categories from the platform’s launch in 2021 through mid-2026, to evaluate calibration and accuracy as functions of time-to-resolution, trading volume, and the number of participating traders. We document that Kalshi’s prices are, in aggregate, extremely well calibrated as markets approach resolution. Assessing all markets (excluding Exotics), we record a decline in Brier score from approximately 0.08–0.09 at a 3-Month horizon to roughly 0.02 at Close. Further, reliability diagrams track the 45-degree line closely across nearly every category examined. In parallel, we note that naive accuracy rises from 88.3% at a 3-Month horizon to 97.2% at Close.
Beneath these strong aggregate results lie more granular conclusions. First, calibration quality improves near-monotonically with both event-level trading volume and the number of unique traders, which we interpret as preliminary evidence that deeper participation may sharpen prices rather than distort them. Second, calibration is stronger and accuracy higher for high-attention categories listed well in advance, such as Economics, than for categories for which the information discovery period is limited and outcomes widely dispersed, such as Mentions. Third, though it is a natural assumption that calibration is a linear function of time-to-resolution, we find that anchoring to the true timing of the underlying event’s occurrence rather than to raw closing timestamps, which may be subject to significant delay, unexpectedly improves calibration in Sports and de-biases results in Elections. To our knowledge, this represents the first attempt to re-scale expiration times in the study of prediction markets. We interpret these findings, taken together, as variation around a strong baseline rather than as a qualification of it: Kalshi’s prices behave like genuine probabilities, and increasingly so as resolution approaches.
