Kalshi Research
Kalshi Research
Topics
Mission
Indices
PublicationsTeamKalshi Markets
Contact usresearch@kalshi.com

Beyond Consensus

Prediction Markets and the Forecasting of Inflation Shocks

Published January 2025

0.1000.0500.2000.100Moderate ShockMajor Shock0.000.020.040.060.080.100.120.140.160.180.20Mean Absolute ErrorEvent TypesConsensusKalshi Market (1 Week Prior)

Figure 1: Market Advantage Across Event Types (1 Week Prior) YOY CPI

Figure 1: Market advantage comparison across shock regimes and forecast horizons.
00

About

Approximately a week prior to the publication of significant economic statistics, analysts at large financial institutions and senior economists will produce estimates of expected figures. Aggregated together, these forecasts, known as "consensus estimates", provide a highly-regarded lens into market expectations and participant positioning. In this study, we compare the predictive performance of consensus estimates to implied pricing from Kalshi prediction markets in estimating true values of a core macroeconomic signal - year-over-year headline inflation (YOY CPI). Across the study period ranging from February 2023 to mid-2025, we find that Kalshi forecasts exhibit a 40.1% lower mean absolute error (MAE) than consensus forecasts across all regimes (normal and shock environments), and document significant outperformance during shocks (significant deviations of actual values from consensus that we term "shock alpha") - recording 50% lower MAE in shocks of > 0.2pp and in shocks of 0.1pp ≤ shock < 0.2pp on a week-ahead timeline. On the same timeline, we find in cases where markets disagree with consensus, market predictions prove superior in 75% of cases. Further, we find preliminary evidence of the ability to predict shocks, with market-consensus deviation showing a positive correlation with forecast surprises and threshold analysis identifying that deviations exceeding 0.1pp predict a ~81.2% shock rate, moving to ~82.4% a day prior to release. While the sample size is naturally modest, our findings suggest the power of the "wisdom of the crowd", implying that prediction market signals are derived heterogeneously from consensus estimates, and thereby offering a valuable complement to traditional forecasts.

Overall Accuracy Advantage

Kalshi forecasts exhibit a 40.1% lower mean absolute error (MAE) than consensus forecasts across all regimes (normal and shock environments).

Shock Alpha

During major shocks (> 0.2pp), Kalshi forecasts record 50% lower MAE than consensus on a week-ahead timeline, or 60% lower MAE on the day prior to release. During moderate shocks (0.1pp ≤ shock < 0.2pp), the advantage is 50% lower MAE on a week-ahead timeline, moving to 56.2% lower one day prior.

Predictive Signal

Market-consensus deviations exceeding 0.1pp predict a ~81.2% shock rate, moving to ~82.4% a day prior to release. When markets disagree with consensus, market predictions prove superior in 75% of cases.

01

Introduction

Macroeconomic forecasters face an inherent challenge: the moments when accurate predictions matter most - times of market dislocation, policy shifts, and structural breaks - are also precisely when historical models tend to fail. Financial market participants routinely publish consensus forecasts for key economic releases days in advance, aggregating expert opinion into a market expectation. Yet these consensus views, while valuable, share common methodological lineages and information sources.

For institutional investors, risk managers, and policymakers, the stakes of forecast accuracy are asymmetric. A marginally better forecast during uncontroversial periods offers only modest value. But superior accuracy during market dislocations - when volatility spikes, correlations break down, or historical relationships fail - can lead to significant alpha generation and limit drawdowns.

Consequently, understanding parameter behavior during periods of market volatility is of paramount importance. In this study, we present the case of a key macroeconomic indicator: year-over-year inflation (YOY CPI) - a central input to future interest rate decisions and signal of economic health. We compare and assess forecast accuracy across multiple time horizons leading up to official data publication. Our central finding is the existence of what we would term "shock alpha" - the incremental forecast accuracy achieved by market-based estimates during tail events relative to consensus benchmarks. This outperformance is not merely academic; it represents improved signal quality at precisely the moments when forecast errors carry the highest economic costs. In this context, the relevant question is not whether prediction market forecasts are always right, but whether they provide a valuable differentiated signal worth incorporating into traditional decision frameworks.

2

Methodology

2.1 Data

We analyze daily implied values from prediction market traders on our platform at three time horizons: one week prior to release (matching consensus timing), one day prior, and the morning of release. Each market utilized is (or was) a live tradeable market that reflects real-money positions at various levels of liquidity. For consensus estimates, we collected institutional consensus estimates for YoY CPI, typically published approximately one week prior to official Bureau of Labor Statistics releases.

Sample period: February 2023 to mid-2025, covering over 25 months of monthly CPI releases across varied macroeconomic regimes.

2.2 Shock Classification

We classify events into three categories based on the magnitude of surprise relative to history. These shocks are measured as the absolute difference between consensus estimates and the actual data print:

  • Normal events: Forecast errors of below 0.1 pp for YOY CPI
  • Moderate shocks: Forecast errors of between 0.1pp and 0.2pp for YOY CPI
  • Major shocks:Forecast errors of > 0.2pp for YOY CPI

This classification allows us to examine whether forecasting advantages vary systematically with the difficulty of the prediction problem.

2.3 Performance Metrics

In order to understand performance, we assess the following performance metrics:

  • Mean Absolute Error (MAE): Primary accuracy measure, calculated as the average absolute difference between forecast and realized value.
  • Win Rate: When consensus and market estimates differ by at least 0.1 percentage points (rounded to one decimal), we record which estimate was closer to the realized outcome.
  • Forecast Horizon Analysis: We track how market estimate accuracy evolves from one week out to release day, revealing the value of continuous information incorporation.
3

Results

3.1 CPI Forecasting Performance

Overall Accuracy Advantage

Across all market conditions, market-based CPI estimates demonstrate a 40.1% lower MAE compared to consensus forecasts. Across all time periods, market-based CPI estimates demonstrate lower MAE of between 40.1% (one week out) and 42.3% (one day out). Moreover, where the estimates differ between consensus and implied values, Kalshi market-based estimates record statistically significant win rates ranging from 75.0% at one-week out (p=0.0384, binomial test), to 81.2% on the day of release (p=0.0106, binomial test). When including ties with consensus (to one decimal place), market-based estimates match-or-outperform consensus ~85% of the time, measured one week in advance. This high degree of directional accuracy implies that market disagreement from consensus is significantly informative of the potential of a shock event.

Shock Alpha

The accuracy differential becomes most pronounced during shock events. Across moderate shock events of between 0.1 and 0.2pp, Kalshi trader estimates taken at the same time as consensus report a 50% lower MAE, moving to 56.2% lower one day prior. Across major shock events of 0.2pp or above, Kalshi trader estimates taken at the same time as consensus report also report a 50% lower MAE a week in advance of release, or a 60% lower MAE on the day prior to release. During a normal regime undefined by shocks, trader estimates perform similarly to consensus. Though the sample size of shocks is small (as it should be in a world where they are largely unexpected), the pattern is clear - when the forecasting environment becomes most challenging, the information aggregation advantage of markets becomes most valuable.

However, it is not only useful to note that Kalshi trader estimates outperform during such times of shock, but also that the disagreement between Kalshi trader estimates and consensus may be indicative of a shock. Compared in disagreement, implied trader estimates report a 75% win-rate relative to consensus figures at comparable timeframes. Moreover, threshold analysis identifies that deviations exceeding 0.1pp predict a ~81.2% shock rate, moving up to an ~84.2% shock rate the day prior to data release. This practically significant difference suggests that prediction market forecasts can be used not only as competing estimates, but as meta-signals about forecast uncertainty itself, transforming market-consensus disagreement into a quantifiable early warning indicator for potential surprises.

4

Discussion

The obvious question that arises here is the following: why do markets outperform consensus during crises? We propose three complementary mechanisms:

4.1 Market participant heterogeneity and the wisdom of the crowd

Traditional consensus forecasts, while incorporating multiple institutional views, tend to share methodological assumptions and information sources. Econometric models, Wall Street research, and government data releases form a common knowledge base. Prediction markets, by contrast, aggregate positions from participants with diverse informational foundations: proprietary models, sector-specific insights, alternative data sources, and informed intuition. This participant diversity finds theoretical foundation in the wisdom of crowds literature, which demonstrates that aggregating independent forecasts from diverse sources can yield superior estimates when participants possess relevant information and forecast errors exhibit less than perfect correlation, and this informational diversity becomes most valuable during regime changes as those that possess fragmentary information combine to produce a collective signal.

4.2 Participant incentive variation

Institutional consensus forecasters operate within complex organizational and reputational systems that induce systematic deviations from pure accuracy maximization. The career concerns of professional forecasters create asymmetric payoff structures where large forecast errors impose significant reputational costs while exceptional accuracy, particularly when achieved through substantial deviation from peer forecasts, may not generate proportionate professional benefits. This asymmetry induces herding behavior whereby forecasters cluster their estimates near consensus values even when private information or model outputs suggest divergent predictions, as the professional cost of being incorrect in isolation frequently exceeds the professional benefit of being correct in isolation.

By contrast, prediction market participants face direct alignment between forecast accuracy and financial outcomes where forecast precision translates to profit and forecast error translates to loss. Reputational considerations are absent, as the sole penalty for deviation from market consensus is financial and depends exclusively on forecast accuracy. This structure creates stronger selection pressure for predictive accuracy whereby participants who systematically identify consensus forecast errors accumulate capital and increase their market influence through position sizing, while those who mechanically follow consensus incur losses when consensus proves incorrect.

During periods of elevated uncertainty, when professional costs of deviating from expert consensus reach their maximum for institutional forecasters, this incentive divergence may be most pronounced and economically significant.

4.3 Aggregation efficiency

The empirical observation that market-based accuracy advantages exist even one week prior to releases, matching the characteristic temporal horizon of consensus forecast publication, suggests a mechanism beyond simple information speed advantages oft-cited for prediction market participants. Rather, markets may achieve more efficient aggregation of dispersed information fragments that are too scattered, too sector-specific, or too ambiguous to incorporate formally into traditional econometric forecasting frameworks. The comparative advantage may derive less from temporal primacy in accessing common information than from the capacity to synthesize heterogeneous information that survey-based consensus mechanisms process inefficiently even over equivalent time horizons.

5

Limitations and Caveats

Our results warrant an important qualification. Of course, given that our overall sample spans ~30 months, major shock events are definitionally rare. This means that statistical power for larger tail events remains limited. Longer times series will strengthen the capacity for future inference, though current results are highly suggestive of outperformance and signal differentiation.

6

Conclusion

We document systematic and economically meaningful outperformance of prediction market forecasts relative to expert consensus, particularly during shock events when forecast accuracy matters most. Market-based CPI estimates show ~40% lower overall error and up to ~60% lower error in periods of major structural change.

Based on these results, a number of areas for future research become salient: (1) the predictability of shock alpha events themselves through volatility and disagreement measures over larger samples and across a basket of macroeconomic indicators, (2) the liquidity thresholds at which prediction markets achieve higher accuracy than traditional forecasting, and (3) the relationship between prediction market forecasts and implied forecasts derived from higher-frequency traded financial instruments.

In environments where consensus forecasts reflect correlated model assumptions and shared information sets, prediction markets offer an alternative aggregation mechanism that may detect regime changes earlier and process heterogeneous information more efficiently. For decision-makers operating in economic environments characterized by increasing structural uncertainty and tail event frequency, shock alpha may therefore constitute not merely an incremental forecasting improvement but a fundamental component of robust risk management infrastructure.