Over the past decade, a rather uncomfortable finding has emerged from across multiple scientific disciplines: a startling number of published research findings cannot be replicated. If you run the same study again with a different group of participants, you get different results. Sometimes the results are drastically different. This is what researchers call the replication crisis, and it is damaging confidence in some branches of science.
Where it started
Acknowledgement of the replication crisis began in 2011, when a series of high-protile failures to replicate important findings began accumulating. In psychology, a massive coordinated effort called the Reproducibility Project attempted to replicate 100 published studies and found that only about 36-39% produced results consistent with the original. Studies on priming (the idea that subtle cues unconsciously influence behavior), ego depletion (the idea that willpower is a limited resource that gets depleted), and social contagion of emotions all failed to replicate consistently.
Biomedical research was also implicated: a report from pharmaceutical company Bayer found that only about 25% of published preclinical findings could be reproduced internally. In cancer biology, Amgen scientists reported being unable to reproduce 47 of 53 landmark papers they attempted to replicate.
What causes it?
Multiple factors contribute to the replication crisis. For one, small sample sizes mean that studies don't have enough participants to reliably detect the effects they're claiming to find, and the results they do find are more likely to be due to chance. P-hacking (also called data dredging) refers to the practice of running many statistical analyses on a dataset until one reaches a p-value below 0.05 (a threshold used to declare statistical significance) without correcting for the fact that by chance alone, 1 in 20 analyses will produce a significant result.
Publication bias compounds the problem: journals prefer positive results, so researchers are incentivized to present their data in the most favorable light, which can shade into selective reporting. And the "file drawer problem" (discussed here!) means that studies that don't find significant effects often go unpublished, leaving the literature full of positive findings that aren't representative of the full picture.
What's being done
The scientific community has responded with a range of reforms. Preregistration, for one, involves publicly registering a study's hypotheses and analysis plan before collecting data. This makes p-hacking much harder. Open data and open code requirements from journals mean that other researchers can verify analyses. Registered Reports, a publication format in which journals agree to publish a study based on the quality of the design before results are known, remove the publication bias incentive entirely. And many fields have launched large-scale replication efforts to systematically test which findings hold up.
What this should and shouldn't mean to you
The replication crisis should not make you distrust science wholesale. It should make you a more discerning reader of science. Single studies, especially small, novel, surprising studies, are ones you should be skeptical of. Meta-analyses and systematic reviews, which synthesize evidence across many studies, are generally more reliable than any individual paper. It is worth noting, though, that the fact that the field is working to address its own methodological failures is evidence that progress is being made.