Library

Psychology · Metascience

Two hundred scientists tried to replicate 100 psychology studies. A third held up.

August 17, 2026article
"A result nobody can repeat was never really a result."

Picture a study that has already done its work: published, cited, absorbed into a field’s working knowledge, quoted in textbooks and talks. What happens if someone opens it back up and simply tries it again — same materials, more statistical power, no changed question? In 2015 a coalition calling itself the Open Science Collaboration did exactly that at a scale no one had attempted before, and the answer it returned unsettled the discipline that produced it.

About 270 researchers, coordinated by the psychologist Brian Nosek, picked 100 studies published in 2008 in three respected journals — Psychological Science, the Journal of Personality and Social Psychology, and Cognition — and tried to replicate every one of them from scratch. Ninety-seven of the originals had reported a statistically significant result. When the team reran all 100 with more statistical power than the originals ever had, only 36 percent came back significant in the same direction. Effect sizes told the same story from another angle: the average effect shrank from about 0.4 in the original papers to about 0.2 once replicated, roughly cut in half. Even by the most generous standard the team used — a holistic judgment rather than a strict significance cutoff — only 39 of the 100 studies were judged to have genuinely held up. The split wasn’t even across subfields: findings in cognitive psychology, about attention, memory, perception, replicated far more reliably than findings in social psychology, which proved easier to nudge and harder to pin down a second time.

None of this happened by accident of design. The team built the project deliberately not to fail: rather than hunting for papers they suspected were shaky, they took a broad, representative slice of ordinary published science, all from a single year. Wherever possible they used the original materials, wrote to the original authors to get procedural details right, let those authors review the replication plans before they ran, and registered every analysis in advance so the question couldn’t quietly shift after the data came in. The samples were large enough that if an original effect was real and as strong as reported, the team had roughly a 92 percent chance of catching it. The results ran in Science in 2015.

What made the finding land so hard was the decade of quiet warnings it seemed to confirm. In 2011 a well-known psychology journal had published a paper, using entirely standard methods, that appeared to show people could sense future events — evidence that ordinary, accepted procedures could manufacture strong support for something impossible. Around the same time, researchers showed that a handful of small, individually reasonable choices — testing a few extra participants, measuring several outcomes and reporting only the one that worked, checking results early and deciding then whether to keep collecting data — could pull a statistically significant finding out of pure noise without anyone intending to cheat. Years earlier, the physician-researcher John Ioannidis had argued, in a paper whose title said the quiet part out loud, that most published research findings might be false, simply from the arithmetic of small samples and flexible analysis in a system that rewards exciting results. The 2015 project put empirical weight behind that argument, and the pattern turned out not to be a psychology problem alone: other fields that went looking found similarly low reproduction rates in corners of cancer biology and experimental economics, and once-celebrated ideas, like willpower as a depletable resource, met the same wall when large teams tried to reproduce them.

The team was careful about what a failed replication does and doesn’t prove. It doesn’t mean the original finding was false: an effect can depend on some hidden condition — the sample, the culture, the year, even the room — that shifted between the first study and the second, or the replication itself could be flawed in ways no one can see. It’s just as possible the original result was a genuine statistical fluke, amplified by a publishing system that favors surprising results over null ones. Reproducibility, the team insisted, isn’t a single number you can stamp on an entire field; it measures how often this particular set of studies, replicated in this particular way, came back. The field’s response wasn’t to bury the number but to build around it — preregistering hypotheses and analysis plans before collecting data went from a fringe habit to a mainstream expectation, journals began offering Registered Reports that accept studies on the strength of their design rather than their outcome, and large multi-lab collaborations formed to test findings openly, across dozens of labs at once.