LLMs in Peer Review: Research Aid or Threat to Science?

NC
Nacho Conesa
calendar_today January 23, 2026 schedule 6 min read Artificial Intelligence
Visual representation of a language model analyzing a scientific paper

A large-scale Nature study examines LLM feedback in peer review. Do AI models improve scientific quality or undermine research integrity?

Peer review is one of the cornerstones of modern science. Yet the system has been under strain for decades: overloaded reviewers, impossible deadlines, and subtle biases that slip through the cracks. Now, a large-scale study published in Nature puts a question on the table that many researchers had been reluctant to ask out loud: can large language models (LLMs) participate meaningfully and reliably in scientific peer review?

The Study That Could Not Be Ignored

Published in January 2026, the research is one of the most ambitious experiments conducted in this area to date. It was a large-scale randomized trial in which LLM-generated feedback — specifically from models in the GPT-4 family — was introduced into real peer-review processes at academic conferences in computer science and artificial intelligence.

The findings are nuanced, which is precisely what makes them compelling. On one hand, human reviewers who received prior LLM feedback tended to write longer, more technically comprehensive reviews. On the other, a troubling convergence of opinions was detected: when reviewers first saw the model's assessment, their own scores migrated closer to the automated system's rating — a phenomenon the authors call "algorithmic anchoring."

The Anchoring Problem and the Risk of Homogenization

Algorithmic anchoring is not a trivial issue. Peer review functions, in theory, because different experts bring independent perspectives. If all reviewers start from the same model-generated analysis, the diversity of judgment collapses. That may sound like efficiency, but in science, disagreement and debate are engines of progress.

Consider a paper on a novel federated learning technique that receives an initial LLM rating labeling it "incremental and low-impact." If three human reviewers read that assessment before forming their own opinion, there is a real probability that genuinely valuable work gets rejected. The model becomes an invisible gatekeeper that no one elected.

Where LLMs Actually Add Value

The study also identifies areas where AI contributes meaningfully. LLMs prove particularly useful for:

  • Detecting internal inconsistencies in manuscripts — contradictory claims, formula typos, malformed citations.
  • Assessing expository clarity and suggesting structural improvements.
  • Surfacing related work the authors may have overlooked.
  • Reducing the time a reviewer needs to complete an initial evaluation.

These capabilities are genuinely valuable, especially in fields where submission volumes have grown exponentially. NeurIPS 2024 alone received over 15,000 papers — a figure that makes it nearly impossible to find qualified reviewers for every submission.

Scientific Integrity: The Line That Must Not Be Crossed

The underlying debate is not whether LLMs can help, but how their use should be governed. Several high-impact journals — including some in the Nature portfolio and the IEEE family — have already published explicit policies prohibiting generative AI from serving as a primary reviewer or as a substitute for human judgment. Reality, however, is messier: many reviewers are already using these models informally, without disclosure.

That opacity may be the most urgent problem. If a reviewer uses ChatGPT to draft sections of their review without declaring it, a systemic and invisible bias is introduced into the process. The authors of the Nature study advocate for mandatory transparency frameworks: reviewers should disclose whether they used AI assistance, much as they currently declare conflicts of interest.

A Practical Path Forward

The most sensible proposal to emerge from this work is neither outright prohibition nor uncritical adoption, but rather designing hybrid workflows with clear safeguards. Specific examples the authors suggest include showing LLM feedback only after a reviewer has already submitted their preliminary assessment, restricting its use to formal checks such as citation completeness and statistical reporting, or deploying it as a training tool for early-career reviewers who are still developing their critical eye.

Science has always absorbed new tools. The microscope, Bayesian statistics, simulation software — each generated initial resistance and eventually became indispensable. LLMs will likely follow the same arc. The critical variable is whether the scientific community shapes the terms of that integration before the integration shapes them.

More articles