Superforecasting
Also known as: Good Judgment Project
Philip Tetlock's framework characterizing the practices and dispositions associated with consistently accurate forecasting under uncertainty.
Superforecasting is the framework articulated by Philip Tetlock characterizing the practices and dispositions associated with consistently accurate forecasting of geopolitical, economic, and other complex events, developed through the Good Judgment Project (GJP) — an Intelligence Advanced Research Projects Activity (IARPA) forecasting tournament run from 2011 to 2015 — and consolidated in Tetlock and Dan Gardner's Superforecasting: The Art and Science of Prediction (2015). Building on Tetlock's earlier Expert Political Judgment (2005) which had documented widespread failure of expert prediction with the famous 'hedgehog/fox' distinction (after Isaiah Berlin's essay), the GJP demonstrated that a top-performing minority — superforecasters — substantially outperformed both intelligence-community baseline and most other forecasters through identifiable practices. The framework's substantial empirical foundation in randomized assignment of forecasters to questions, calibrated scoring (Brier scores), and tournament structure distinguishes it from earlier theoretical-forecasting literature and provides empirical grounding for specific dispositions and practices.
Core components
- Tournament-structure foundation: randomized forecaster assignment to questions, time-stamped probability forecasts, calibrated scoring (Brier score) enabling comparison
- Superforecaster identification: top 2 percent of forecasters by accuracy over the tournament's questions, characterized by both better calibration and resolution
- The Ten Commandments of Superforecasting (Tetlock-Gardner book formulation): (1) triage — focus on questions in the 'Goldilocks zone' where deliberate effort can make a difference
- (2) break seemingly intractable problems into tractable sub-problems
- (3) strike the right balance between inside view (situation-specific factors) and outside view (base rates of similar situations)
- (4) strike the right balance between under- and over-reacting to evidence
- (5) look for the clashing causal forces at work in each problem
- (6) strive to distinguish as many degrees of doubt as the problem permits
- (7) strike the right balance between under- and over-confidence
- (8) look for the errors behind your mistakes but beware rear-view-mirror hindsight
- (9) bring out the best in others and let others bring out the best in you
- (10) master the error-balancing bicycle
- Active open-mindedness: the disposition to consider alternative viewpoints and update beliefs in response to evidence
- Fermi-style decomposition: breaking complex questions into estimable components
- Reference-class forecasting: using base rates from analogous historical situations
- Calibration: probabilities should match observed frequencies — saying 70 percent should mean being right 70 percent of the time across many predictions
- Resolution: distinguishing situations with different actual probabilities rather than always predicting middle probabilities
- Aggregation effects: combining forecasts from multiple superforecasters substantially improves accuracy beyond any individual
- The hedgehog-fox distinction: hedgehogs commit to one big idea, foxes integrate from multiple sources — foxes outperformed hedgehogs in Tetlock's earlier research and superforecasters exhibit fox-like cognitive style
- Practice and feedback loops: superforecasting skills improve with deliberate practice on questions with prompt, accurate feedback
Primary use case
Foundational framework for empirically-grounded forecasting practice; applied principally in: intelligence community analytic practice (the IARPA tournament's original sponsor), commercial-forecasting consulting (Good Judgment Inc., Metaculus, prediction-market platforms), policy analysis and scenario planning, individual decision-making practice in finance, business, and personal life, academic research in judgment-and-decision-making, public-engagement contexts about expert prediction (the COVID-19 pandemic generated substantial superforecasting commentary), AI-forecasting and existential-risk literature (where superforecasting practices have been argued to improve calibration on long-term technological questions); standard reference in forecasting-practice and judgment-research curricula.
Common criticisms
- Superforecasting has substantial empirical foundation but specific critiques exist — Nassim Taleb has argued that the framework, while valid for the medium-term political-event questions on which it was tested, may not generalize to the genuinely fat-tailed domains where forecasting matters most (financial crises, pandemics, geopolitical-conflict outbreaks) — the tournament questions tend to involve relatively well-bounded events rather than the rare-high-impact events Taleb has emphasized
- the questions used in the GJP tournament have been argued to be selected for forecaster-tractable structure, with selection effects that may overstate the framework's applicability to questions outside that selection
- critics have argued that the difference between average forecasters and superforecasters partly reflects investment of effort, with superforecasters spending substantially more time per question — the framework rewards effortful forecasting that may not generalize to lower-effort settings
- some empirical literature has documented superforecaster performance regression to the mean — top performers in one season substantially decline in subsequent seasons, complicating the framework's identification of stable individual differences
- the framework's applicability to longer time horizons (multi-year, decade-scale) and to questions about technological transformation (AI capabilities, climate impacts) is methodologically harder to evaluate because few resolved questions of those types have been forecasted
- critics have argued the framework documents what superforecasters do but provides limited guidance for how to become one — many of the dispositions are personality-correlates rather than trainable practices
- the Brier-score scoring system has been argued to under-incentivize calibrated extreme forecasts in some structures
- the relationship between superforecasting and broader expertise traditions (subject-matter expertise, professional judgment) is contested — Tetlock's framework can be read as anti-expertise (foxes-over-hedgehogs) or as expertise-redirecting (toward general forecasting capability), with implications for institutional design
- specific high-profile predictions by named superforecasters have been critiqued in retrospect with mixed outcomes
- commercialization of superforecasting through Good Judgment Inc. has produced criticism about gap between research findings and consulting practice.
Lineage
- Siblings
- Prediction Markets, Reference Class Forecasting