Glossary
multiple comparisons
Testing several outcomes in one study, which raises the chance that at least one difference looks real when nothing is there, unless the analysis is planned to allow for it.
A single statistical test carries a fixed risk of a false positive, which statisticians call a Type I error. The Food and Drug Administration's 2022 guidance for industry on multiple endpoints in clinical trials states the conventional size of that risk plainly: a two-sided test at the 0.05 level means the probability of falsely concluding that the drug differs from the control, when no difference exists, is no more than 5 percent, or one chance in 20. Run the test once and that is the risk being carried. Run tests on many outcomes and the chance that at least one of them comes back falsely positive is larger than the risk attached to any single one.
The guidance gives the arithmetic. For three independent endpoints the overall Type I error rate is about 7 percent, and for ten independent endpoints it is about 22 percent. It names this higher than intended overall rate, arising when multiple tests are conducted without adjustment, the multiplicity problem. The problem is not confined to separate outcomes: the guidance notes that analysing multiple dose groups, multiple time points or multiple subgroups of participants inflates the rate in the same way, and that correlated endpoints inflate it too, though potentially by a lesser degree. Inflating the rate, it says, makes conclusions about whether an effect has been demonstrated unreliable.
The remedy is planned in advance rather than applied afterwards. The guidance's stated principle for controlling multiplicity is to prospectively specify all planned endpoints, time points, analysis populations, doses and analyses, and only then select and prespecify the adjustment. It also groups endpoints into a hierarchy of families: primary endpoints, which establish the effect the study is meant to show, secondary endpoints, which support the primary result or demonstrate additional effects, and exploratory endpoints, which exist for research purposes or to generate new hypotheses. That hierarchy is the reason a result on an outcome tested with no multiplicity adjustment is a weaker thing than a result on a prespecified, adjusted primary endpoint, even when the two carry equally impressive numbers.
Sources
Checked 4 September 2026