Glossary
p-value
The probability of getting a result at least as extreme as the one observed, if the treatment truly made no difference. It is not the probability that a finding is true.
A p-value answers one narrow question. The American Statistical Association, in the statement its board issued in 2016 to correct how the number was being used, defines it informally as the probability, under a specified statistical model, that a statistical summary of the data, such as the difference between two group means, would be equal to or more extreme than the value actually observed. The model it is calculated under almost always contains a null hypothesis, which usually says there is no difference between the groups. The smaller the p-value, the more awkwardly the data sit with the idea that nothing is going on, provided the assumptions behind the calculation hold.
What it is not is the part that goes wrong most often. The association's second principle is blunt: p-values do not measure the probability that the studied hypothesis is true, or the probability that the data were produced by random chance alone. It is a statement about data in relation to a hypothetical explanation, not a statement about the explanation itself. The fifth principle adds that a p-value does not measure the size of an effect or the importance of a result, because any effect, however tiny, will produce a small p-value if the sample is large enough, and a real effect can produce an unimpressive one in a small study.
The threshold is a convention rather than a fact of nature. The NIST/SEMATECH e-Handbook of Statistical Methods describes the significance level as the risk of rejecting the null hypothesis when it is in fact true, notes that a level of 0.05 means that happens 5 percent of the time when the null hypothesis holds, and says the common choices of 0.1, 0.05 and 0.01 are somewhat arbitrary. The association's third principle is that scientific conclusions should not be based only on whether a p-value passes a specific threshold, since a conclusion does not become true on one side of the divide and false on the other. That is why a study reporting a difference it could not distinguish from chance has not thereby shown the difference is absent.
Sources
- American Statistical Association, The ASA Statement on p-Values: Context, Process, and Purpose, Wasserstein and Lazar, The American Statistician 70(2), 129 to 133 Primary
- National Institute of Standards and Technology, NIST/SEMATECH e-Handbook of Statistical Methods, section 7.1.3, What are statistical tests?
Checked 4 September 2026