Science

The NIH Plan To Strengthen 'Replication and Reproducibility' Will Not Check Anything

The world’s largest research funder asked its grantees whether it should audit them. Guess what they said.

|


The National Institutes of Health (NIH) will spend $174 million over five years on "R3PEATS," a new initiative aimed at "strengthening replication and reproducibility" of NIH-funded research. The program, which NIH Director Jay Bhattacharya announced in a Science editorial and a press briefing in September, is the agency's answer to the 20-year-old observation that most published research findings are false.

The initiative's name stands for "Rigor, Replicability, and Reproducibility to Promote Excellence, Accuracy, and Translation in Science." But just as Oreos are filled with "creme" because there is no cream in them and Cheez Whiz is only distantly related to cheese, R3PEATS will not repeat anything, thanks to resistance from NIH grantees.

The $174 million allocated to R3PEATS amounts to about $35 million a year, or 0.07 percent of the NIH's $47.2 billion annual budget. "We should be giving at least 20 percent of the NIH budgets to replication," Health and Human Services (HHS) Secretary Robert F. Kennedy Jr., whose department includes the NIH, declared during a confirmation hearing in January 2025. That would be about $9.4 billion a year—270 times what the NIH actually put up.

We know what reproducibility research costs because private groups have been doing it for a decade with foundation or government money. In 2013, the Reproducibility Project: Cancer Biology got a $1.3 million grant from the Laura and John Arnold Foundation for 50 high-impact replication experiments, and it eventually spent about $50,000 per paper. This year, psychologist Brian Nosek's Center for Open Science published replications of 164 social science papers, 49 percent of which held up. That project was funded by a $7.6 million grant from the Pentagon's Defense Advanced Research Projects Agency, which amounts to less than $50,000 per paper, even assuming all that money was spent on replication.

At those prices, $174 million would buy a few thousand replications, which might cover a random sample of a few hundred NIH-funded findings each year. But the NIH is not actually spending that budget on replication.

The money is split five ways: $44.75 million for "network and community," $42.4 million for "grassroots rigor projects," $45 million for "multi-team triangulation," $38.25 million for "rigor scholar cohort awards," and $3.75 million for NIH staff and workshops. The amount earmarked for replicating published findings is zero. 

Why is that? In January, NIH staff proposed three to five replication centers that would redo studies chosen by a committee. The NIH ran the idea past its Council of Councils, an advisory body made up largely of researchers the NIH funds, and the council objected. Members worried that a small group would "choose which studies to replicate," that replications might "re-adjudicate established science," and that they could "disincentivize innovative research." Science noted the additional concern that redoing studies "could unfairly cast doubt on some scientists' research."

Casting doubt on findings that don't hold up is what replication is for. The NIH asked its grantees whether it should check their work, the grantees said no, and the checking was excised from the plan.

Replication was replaced by a list of things that are good, popular, and unmeasurable. The project's "anticipated outcomes" are evidence of effective culture change approaches, new knowledge about methodological and biological variability, effective mentorship and incentive strategies, and a self-sufficient U.S. reproducibility network. The R3PEATS plan includes no baseline replication rate, no target, no metric, and no date by which anyone will know whether it worked. The plan does not even estimate the size of the problem that R3PEATS is supposed to address.

Bhattacharya does take a stab at such an estimate in his Science editorial, but he gets crucial details wrong. "Among studies that reexamined established medical practices, roughly 40% found that the practice did not work better than an earlier approach or no treatment at all," he writes, citing a 2013 article in Mayo Clinic Proceedings. So far, so good. But then Bhattacharya adds that "preclinical cancer research also yields troubling rates of failure," linking to a 2016 PLOS Biology audit by epidemiologist Shareen A. Iqbal and four other scientists.

That study does not look at preclinical cancer research, and it does not estimate a failure rate. It is a transparency survey of 441 randomly selected PubMed articles. Of 268 papers that included data, none provided complete raw data and only one described a protocol.

Although the PLOS Biology report does not say what Bhattacharya claims, it does highlight an ongoing problem. People like me would check NIH-funded studies for free if the NIH would only require the data and protocol information it has said it "expects" from researchers since 2003.

Every item in the R3PEATS agenda is something a nongovernmental group could do, and most are things nongovernmental groups already do. Nosek runs replications. Psychologist Dorothy Bishop built the U.K. Reproducibility Network with no budget worth mentioning. Journals can run replication sections, universities can train postdocs, and anyone with a laptop can build a PubMed browser.

Last month, the NIH unveiled one such browser, Linked Discoveries, that color-codes 200 related papers by citation count and retraction status. Bhattacharya says it "doesn't determine whether a finding is correct," and he is right.

The action items that Bhattacharya mentions in his Science editorial are addressed to other people: "research institutions should," "publishers should," "funders should," "scientific leaders should." Bhattacharya, who runs the world's largest research funder, says "funders should recognize" instead of "the NIH will require."

The NIH could make its money conditional. It could require preregistration of each study's confirmatory analysis and reporting of its raw data and protocol. It could make grant renewal contingent on publication of data from the previous award. It could audit a random sample of funded findings every year and publish the results under the grantee's name. It could score future grant requests on whether the applicant's past findings held up.

None of this costs $174 million. All of it is unpopular with grantees, which is why only a funder can impose it.

The NIH does not seem inclined to do so. Data sharing has been "expected" since 2003. A 2007 federal law actually requires that results of registered clinical trials be posted within a year, threatening penalties of more than $10,000 a day. But a 2020 Lancet study found that just 31 percent of clinical trials funded by the U.S. government (mostly the NIH) met the deadline, even worse than the 41 percent rate for all such studies.

The Food and Drug Administration (FDA) did not send its first notice of noncompliance with the data mandate until April 2021 and, as far as I can find, has never collected a fine. The FDA's big enforcement push was a March 2026 letter to 2,200 research sponsors reminding them that the law exists. 

About half of NIH-funded trials with data due in 2019 and 2020 reported on time, the HHS inspector general found. The Government Accountability Office reports that the NIH never suspended grant funding for noncompliance before October 2021, 14 years into the mandate.

In January 2016, the NIH added mandatory "rigor and transparency" sections to every grant application. A decade later, the result is a paragraph of boilerplate in each proposal. 

Fraud cases underline the ineffectiveness of such half-hearted safeguards. Charles Piller, the Science reporter who documented the trial-reporting failure, also wrote the book on the scandals involving manipulation of Alzheimer's disease images. The doctored papers were caught by unpaid data sleuths working nights, not by any funded program. One of the researchers named in Piller's book ran the neuroscience division of the NIH's own National Institute on Aging. A serious, well-funded audit program is necessary to change the culture that Piller describes, in which everyone is afraid to speak out, institutions protect their stars, and journals take years to retract bogus studies. 

It is important to note that the ideal replication rate is not 100 percent. Researchers should not wait to publish until they have proved their findings beyond a reasonable doubt, and even the best studies will often be superseded, at least in part, by future work. The problem is not useful work that gets modified or overturned when other researchers try to build on it. The problem is entirely useless work by researchers who conceal data, obfuscate methodology, and misrepresent results.

There are real challenges in rooting out bad research without treating scientists like criminals, discouraging innovation, and unfairly ruining careers. Some fraction of audit findings will be wrong, so some good results will be discredited and too much attention will go to litigating the past rather than building the future. But R3PEATS does not rise to these challenges. It is content to remain on the breakfast-meeting sideboard with Froot Loops and Krispy Kreme.