Glossary

statistical model

An equation fitted to data that splits a measured outcome into a part explained by other recorded quantities and a part left to chance, estimating how each quantity relates to the outcome.

A statistical model is an equation fitted to measurements. The NIST handbook of statistical methods describes it as two components added together: a deterministic part, which is a mathematical function of one or more other measured quantities, and a random part that follows a probability distribution. The quantity being explained is the response variable. The quantities doing the explaining are the predictor or explanatory variables, also called independent variables or regressors. Between them sit the parameters, which the handbook defines as the quantities estimated during the modelling process, and whose true values it calls unknown and unknowable outside a simulation. Fitting the model to data is the act of estimating those parameters, which divides the variation in the response into the part the predictors account for and the part they do not.

The form most health research uses carries many predictors at once. In linear least squares regression, the most widely used method in the handbook's account, each explanatory variable is multiplied by its own unknown parameter and the terms are added together, so age, smoking, income and any other recorded characteristic can all sit in the same equation. This is what a paper means when it reports that its models adjusted for a list of factors: those factors were included as predictors, so the parameter estimated for the exposure under study is the relation left once the others have been fitted alongside it. The same fit is what produces the effect measures a study reports, and the intervals around them.

Two limits travel with every model. The first is that it can only account for what was measured and entered into it, so an influence nobody recorded is not in the equation and its effect stays folded into the result: this is why adjustment narrows the problem of confounding without settling it. The second is the random component itself. The model describes a tendency across a group, with the leftover variation summarised as a distribution of errors rather than explained, so a fitted parameter says what the data support on average and not what happened to any one person in them.

Sources

  1. NIST/SEMATECH, e-Handbook of Statistical Methods, 4.1.1.1 What is process modeling? Primary
  2. NIST/SEMATECH, e-Handbook of Statistical Methods, 4.1.4.1 Linear Least Squares Regression

Checked 17 September 2026