Delphi Method
RAND-developed structured method for eliciting expert consensus through iterated anonymous rounds with controlled feedback.
The Delphi Method is the structured forecasting and expert-judgment technique using anonymous iterative rounds with controlled feedback to produce group estimates that benefit from expert knowledge while mitigating dominant-personality, conformity, and groupthink effects. Developed at the RAND Corporation in the 1950s and early 1960s by Olaf Helmer, Norman Dalkey, and Theodore Gordon for US Air Force national-security forecasting (originally Project DELPHI on Soviet bombing-target prioritization), the method was publicly described in the foundational Helmer-Dalkey 1963 RAND Memorandum 'An Experimental Application of the DELPHI Method to the Use of Experts.' The technique involves recruiting a panel of experts, posing structured questions, collecting anonymous individual responses, providing controlled feedback (typically statistical summary of group responses with rationale from outliers), then iterating until convergence or stable disagreement. Linstone and Turoff's 1975 The Delphi Method: Techniques and Applications consolidated the methodology. Delphi has been applied broadly across technology forecasting, healthcare-quality consensus, policy planning, and qualitative research.
Core components
- Expert panel: typically 10-100 domain experts selected for relevant knowledge, often with geographic or institutional diversity
- Anonymity: panel members do not know one another's identities (or specific responses) — distinguishes Delphi from face-to-face expert panels and reduces dominant-personality and conformity effects
- Iteration: typically 2-4 rounds (rarely more) of structured-question response and feedback
- Controlled feedback: between rounds, panel receives statistical summary of group responses (median, interquartile range, distribution) and reasoning from outliers without individual-attribution information
- Stopping criterion: convergence (responses stabilize), stable disagreement (responses do not converge but stop changing), or pre-specified round limit
- Question types: forecasts (when will X occur?), rankings (rank-order these technologies by importance), policy preferences (which interventions should government prioritize?), event-likelihood estimates
- Statistical aggregation: median for central tendency, interquartile range for consensus measure
- Variants: Policy Delphi (Turoff, exploring policy-option viewpoints), Real-Time Delphi (online rounds without distinct round breaks), Argument Delphi (emphasis on argumentation rather than convergence), Estimate-Talk-Estimate (variant minimizing the anonymous-iteration structure)
- Application contexts: technology forecasting (notably the 1971 RAND Report on Long-Range Forecasting), Japanese national technology-foresight programs (NISTEP Delphi surveys), healthcare-quality consensus development, healthcare clinical-guideline development, policy planning
Primary use case
Structured expert-judgment elicitation for forecasting and consensus development; applied principally in: technology forecasting (NISTEP Japanese Delphi surveys, German BMBF Delphi programs, South Korean foresight programs); healthcare-quality consensus development (substantial use in nursing research, clinical-guideline development, RAND/UCLA appropriateness method); policy planning; research agenda-setting; market research and competitive-intelligence consensus; risk assessment in low-data domains; academic and professional reference in futures studies, operations research, healthcare research methodology, and qualitative research methods literature; intellectual foundation for broader expert-elicitation tradition and for Reference Class Forecasting (parallel structured-empirical forecasting methodology); modern computer-mediated implementations through specialized Delphi software (eDelphi, Welphi); input to broader forecasting-tournament tradition (Tetlock's Good Judgment Project) which has produced alternative empirical evidence on expert-aggregation methods.
Common criticisms
- Delphi has substantial empirical literature documenting both supportive and skeptical findings — supportive studies show Delphi can produce more accurate forecasts than individual experts and can develop consensus across diverse expert groups
- skeptical studies (Sniezek 1992, Rowe and Wright 1999 comprehensive review, Tetlock's expert-judgment research) show Delphi's accuracy advantages over alternative aggregation methods are often modest, and the convergence Delphi produces may reflect group-pressure-toward-conformity (in modified form) rather than genuine knowledge integration
- expert-panel composition substantially affects results — different expert groups applying Delphi to the same question produce different consensus, raising reproducibility concerns
- the iteration-toward-convergence dynamic may produce false consensus when underlying uncertainty is genuine
- technology-forecasting applications have documented mixed accuracy — the 1971 RAND Long-Range Forecasting Delphi predictions have been retrospectively analyzed (with substantial predictions proving wrong, particularly for longer-horizon forecasts)
- for healthcare-consensus applications, Delphi-derived clinical recommendations may reflect expert-panel-composition bias more than underlying clinical evidence
- the anonymous-iteration structure that distinguishes Delphi from face-to-face panels has been argued by some methodologists to be less important than panel composition and question design
- modern forecasting-tournament evidence (Tetlock's Superforecasting research) suggests alternative aggregation methods (probability elicitation with aggregation algorithms) may outperform Delphi in calibration
- Real-Time Delphi and online Delphi variants have produced mixed evidence about whether removing the discrete-round structure improves or degrades accuracy
- commercial-Delphi-consulting practices have produced varying quality, with the 'Delphi' brand applied to exercises lacking the methodological discipline the original method specifies.
Lineage
- Siblings
- Reference Class Forecasting, Scenario Planning (Wack), Superforecasting