Myth Busters: Artificial Intelligence Edition

Three things scientists get wrong about AI in the research workflow

 

Myth #1: If an AI gives a real citation, the claim it supports is accurate.

 

Not necessarily. AI tools frequently attach genuine, verifiable references to statements those references never actually make, a phenomenon sometimes called citation misdirection. The paper exists, the authors are real, and the DOI is real, but the source does not support the point being cited. This failure mode is more insidious than an invented reference, because the usual check to see if the paper exists passes cleanly. A reviewer or reader who spot-checks only the reference list will never catch it; the error surfaces only when someone reads the cited paper and finds the claim is not there. Recent analyses of AI-generated references describe an entire taxonomy of these failures, including semantic hallucination and identifier hijacking, in which a real DOI or PMID is attached to the wrong work entirely [1]. In one review of hallucinated citations across dozens of conference papers, every case involved multiple bibliographic fields being wrong at once [1].

  BUSTED: A real citation is not the same as a suggested one. Read the source before you cite it.

 

Myth #2: AI pulls its citations from a database of published papers.

 

It does not. Most large language models generate text by predicting plausible sequences of words rather than retrieving verified bibliographic records, so they can invent references outright. These fabrications are convincing by design, complete with realistic author names, journal titles, volume numbers, and DOIs. The scale is not trivial. In one controlled study, GPT-4o generated 176 citations across six literature reviews; 35 (19.9%) were entirely fabricated, and among the 141 real citations, 45.4% contained errors, most often invalid DOIs [2]. Fabrication rates were highest for specialized, less visible topics. These fabrications are also leaking into the published literature: a Columbia University team audited 2.5 million papers and 97.1 million references in PubMed Central and identified 4,046 fabricated citations across 2,810 papers, with the rate rising more than 12-fold since 2023 and climbing sharply from mid-2024 onward [3].

  BUSTED: Verify every AI-supplied reference in PubMed before it goes near a manuscript, grant, or review.

 

Myth #3: Anything I type into an AI chatbot stays private.

 

Often it does not. Depending on the tool and its settings, text you enter may be stored on external servers, reviewed by humans, or used to train future models. Importantly, once submitted, you have no practical control over where it goes. For scientists, that is a serious exposure when the input includes unpublished data, a draft R01, or a manuscript under confidential peer review. This is not hypothetical policy hand-wringing. NIH explicitly prohibits peer reviewers from using generative AI to analyze applications or formulate critiques, and states that uploading content or original concepts from an application, proposal, or critique to an online AI tool violates NIH peer review confidentiality and integrity requirements [4]. The stated rationale is control: NIH notes there is no guarantee of where data submitted to an AI tool are sent, saved, viewed, or later reused, and a confidentiality breach can end a reviewer’s service or trigger suspension or debarment [4,5]. Many journals apply similar restrictions to manuscript review, and authors must now state if they have used AI for any part of the manuscript preparation process (even just for grammar corrections).

  BUSTED: Check a tool’s data-retention and training policies first — and keep confidential material out of it.

 
References

  1. BibTeX citation hallucinations in scientific publishing agents: evaluation and mitigation. arXiv preprint, 2026 (see also: Detecting and correcting reference hallucinations in commercial LLMs and deep research agents. arXiv preprint, 2026). https://arxiv.org/pdf/2604.03159

  2. Linardon J, Jarman HK, McClure Z, Anderson C, Liu C, Messer M. Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models: experimental study. JMIR Ment Health. 2025. https://pmc.ncbi.nlm.nih.gov/articles/PMC12658395/

  3. Topaz M, Roguin N, Gupta P, et al. Fabricated citations: an audit across 2·5 million biomedical papers. Lancet. 2026 May 7. https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(26)00603-3/fulltext

  4. NIH. The use of generative artificial intelligence technologies is prohibited for the NIH peer review process. Notice NOT-OD-23-149. June 23, 2023. https://grants.nih.gov/grants/guide/notice-files/NOT-OD-23-149.html

  5. Lauer M, Constant S, Wernimont A. Using AI in peer review is a breach of confidentiality. NIH Open Mike blog. June 23, 2023. https://www.era.nih.gov/news/era-enhancement-new-language-confidentiality-agreement-iar-prohibiting-use-generative