You've probably seen the words “statistically significant” in newspaper headlines about scientific studies. "Scientists find statistically significant link between X and Y." While it sounds like concrete proof has been identified, “statistically significant” doesn’t truly mean there has been a causal claim, making it one of the most widely misunderstood concepts in science communication.
What it technically means
Statistical significance is a mathematical statement about probability. When researchers say a result is statistically significant, they typically mean that the probability of observing a result as extreme as the one they found is below a threshold called the p-value cutoff, usually set at 0.05.
A p-value of 0.05 means there is a 5% probability that you'd see results this extreme by chance alone if the null hypothesis (the assumption that there is no real effect) were true. When p < 0.05, researchers conventionally conclude the result is statistically significant, meaning unlikely enough to have occurred by chance that they feel comfortable claiming the effect is real.
What it does NOT mean
Here's where the confusion comes in. Statistical significance does not mean that the effect is large, important, or meaningful in the real world nor does it mean that the result will replicate nor does it mean that the probability is 95% that the researchers are right.
A study with 1,000,000 participants can find a statistically significant difference in blood pressure between two groups that amounts to 0.1 mmHg, which is far too small to have any clinical relevance whatsoever. Large sample sizes make it easier to detect even trivially small effects, and those effects can be highly statistically significant while being relatively meaningless in practice.
Conversely, a study with a small sample size might fail to detect a real, clinically important effect (a "false negative" result) because the study didn't have enough participants to have sufficient statistical power to find it.
The difference between statistical and practical significance
Practical significance (also called clinical significance or effect size) is the question of whether the effect is large enough to actually matter. Effect size measures, like Cohen's d, odds ratios, or relative risk, quantify how large an effect actually is, independent of sample size. These are often more informative than p-values, and increasingly, journals require authors to report them alongside significance tests.
The next time you see a headline that says a study found a statistically significant result, you should definitely question it: how large was the effect? How large was the sample? Has this been replicated?