Regression Analysis

framework · mathematics · formal-scientific

Statistical modeling of relationships between dependent and independent variables.

Regression Analysis is the family of statistical techniques for modeling the relationship between a dependent variable (response, outcome) and one or more independent variables (predictors, covariates), estimating the conditional expectation of the dependent given the independents and quantifying the uncertainty of the estimates. The technique's name comes from Francis Galton's 1886 observation of 'regression toward mean stature' — children of unusually tall parents tend to be tall but less so than their parents, and vice versa. Galton's collaborator Karl Pearson developed the mathematical framework for correlation and the early formalization of linear regression in the 1890s-1900s, with substantial subsequent development through R.A. Fisher (maximum likelihood, ANOVA), the Yule, Wald, and Cochran traditions. The basic model — linear regression with normally distributed errors — extends through enormous generalizations: multiple regression (multiple predictors), generalized linear models (Nelder-Wedderburn 1972 — Poisson regression for counts, logistic regression for binary outcomes, etc.), nonlinear regression, robust regression (Huber 1964), regularized regression (ridge — Hoerl-Kennard 1970, lasso — Tibshirani 1996, elastic net — Zou-Hastie 2005), nonparametric regression (kernel methods, splines, additive models), and machine-learning regression (decision trees, random forests, gradient boosting, neural networks). Regression analysis is foundational across essentially every quantitative scientific discipline and is the workhorse method in epidemiology, economics, social sciences, and applied machine learning. The framework's apparent simplicity masks substantial sophistication required for valid inference (assumption checking, confounding, multicollinearity, heteroscedasticity, omitted variable bias, causal identification).

Originators

Francis Galton (concept of regression, 1886); Karl Pearson (mathematical formalization, 1890s-1900s); R.A. Fisher (substantial development, 1920s-1950s); subsequent foundational figures including John Nelder, Robert Wedderburn, Robert Tibshirani, Trevor Hastie high

Year / Decade

1886 (Galton concept); 1890s-1900s (Pearson formalization); 1920s onward (Fisher development); ongoing extensions high

Primary sources

Galton, F. (1886). 'Regression towards Mediocrity in Hereditary Stature', Fisher, R.A. (1925). Statistical Methods for Research Workers, Nelder, J.A. & Wedderburn, R.W.M. (1972). 'Generalized Linear Models', JRSS Series A, Hastie, T., Tibshirani, R. & Friedman, J. (2009, 2nd ed.). The Elements of Statistical Learning high

Core components

Primary use case

Foundational statistical method across essentially every quantitative scientific discipline; epidemiology (risk factor analysis); economics (econometrics); social sciences (sociology, political science, education); biostatistics; psychometrics; finance (asset pricing, factor models); machine learning (linear/logistic regression as baselines, regularized regression for high-dimensional data); foundation for substantial commercial software (SPSS, SAS, R packages, scikit-learn); pedagogical foundation in essentially every applied statistics curriculum.

Common criticisms

Lineage

Child of
Frequentist Statistics
Siblings
Frequentist Statistics, Causal Inference
Derived from
Frequentist Statistics