Let's Use ChatGPT for Research: Fictitious Citations, Who’s Responsible?
- Authors: Abdullah S.1, Anwar F.2, Khalid Z.3, Siddique A.4
-
Affiliations:
- Lady Reading Hospital, Peshawar, Pakistan
- College of Computer and Information Sciences, Imam Ibn Saud Islamic University (IMSIU), Riyadh, Saudi Arabia
- Shahida Islam Medical Complex, Lodhran , Pakistan
- FMH College of Medicine and Dentistry, Lahore, Pakistan
- Section: LETTER TO THE EDITOR
- Submitted: 10.05.2026
- Accepted: 31.07.2026
- Published: 31.07.2026
- URL: https://consortium-psy.com/jour/article/view/15879
- DOI: https://doi.org/10.17816/CP15879
- ID: 15879
Cite item
Full Text
Full Text
The integrity of biomedical literature rests on the verifiability of its references. The Lancet published a landmark audit that revealed a major crisis. Among 97.1 million verified references which researchers extracted from 2.5 million PubMed Central papers contained 4,046 fake citations which were found in 2,810 academic works; they discovered that these academic works showed an increase in fabrication rate which had grown from 4 per 10,000 papers during 2023 to 57 per 10,000 papers by early 2026. The temporal inflection point in mid-2024, consistent with a publication lag following widespread adoption of large language models (LLMs), implicates artificial intelligence (AI) writing tools as a principal driver of this acceleration [1].
The scale of AI use in biomedical writing is creating a concerning pattern that needs immediate attention. Vocabulary analysis of over 15 million PubMed abstracts published between 2010 and 2024 estimated that at least 13.5% of biomedical abstracts published in 2024 were processed with LLMs, with some subcorpora reaching 40% [2]. LLMs generate references that create the appearance of authentic academic work because they produce citations that follow correct formatting rules, include actual authorship, and present believable publication dates, but do not refer to any real scholarly material [1].
Controlled evaluations illustrate the extent of this failure because ChatGPT-3.5 and GPT-4 generated 636 citations, including 84 multidisciplinary literature reviews that contained 55% ChatGPT-3.5-generated citations and 18% from GPT-4 were found to be completely invented because most of their remaining genuine citations contained major errors [3]. The latest study compared five LLM platforms and evaluated their ability to find references in articles from British Medical Journal, Journal of the American Medical Association (JAMA), and New England Journal of Medicine and discovered that these systems failed to obtain correct reference data in 47.8% of cases, which demonstrates that this issue continues to exist in present-day model versions [4].
The results of these investigations have broader implications than the individual studies. Systematic reviews, which serve as the evidence base for clinical guidelines and treatment choices, now face increasing contamination. A cross-sectional study of 200,000 systematic reviews published between 2013 and 2024, indexed in Web of Science, found that 0.15% incorporated retracted paper mill articles into their evidence synthesis, with 124 postretraction citations documented, including 13 occurring more than 500 days after retraction [5]. The presence of fabricated or fraudulent citations in systematic reviews poses a critical risk to clinical practice because it disrupts the evidence chain. At the time of analysis, 98.4% of the
2,810 papers identified in The Lancet audit had received no publisher response according to the audit results [1].
Every part of the publication process needs structural changes to effectively handle this threat. Pre-review submission workflows need automated reference verification systems because publishers already have this technology available, but institutional resistance prevents its implementation [1]. Article records should include integrity metadata from indexing services, which will help subsequent processes identify problem references. All academic institutions need to implement mandatory AI-use disclosure policies and establish citation verification training programs for researchers, which they should enforce through their conduct frameworks. Review articles, which have a fabrication rate 57% higher than other article types, deserve particular scrutiny [1]. Without protective measures, the scientific literature risks becoming increasingly contaminated with fictional content as researchers increasingly use LLMgenerated material to produce peer reviewed work without protective measures.
Article history
Submitted: 10 May 2026
Accepted: 8 Jul. 2026
Published Online: 31 Jul. 2026
Authors’ contribution: All the authors made a significant contribution to the article, checked and approved its final version prior to publication.
Funding: The research was carried out without additional funding.
Conflict of interests: The authors declare no conflicts of interest.
About the authors
Sadaf Abdullah
Lady Reading Hospital, Peshawar, Pakistan
Email: abdullahsadaf2@gmail.com
ORCID iD: 0009-0003-0335-7144
Fareeha Anwar
College of Computer and Information Sciences, Imam Ibn Saud Islamic University (IMSIU), Riyadh, Saudi Arabia
Email: fejaz@imamu.edu.sa
ORCID iD: 0000-0002-6993-7761
Zunaib Khalid
Shahida Islam Medical Complex, Lodhran , Pakistan
Email: zunaibkhalid34@gmail.com
ORCID iD: 0009-0005-5630-2838
Abubakar Siddique
FMH College of Medicine and Dentistry, Lahore, Pakistan
Author for correspondence.
Email: drabubakar561@gmail.com
ORCID iD: 0009-0006-8926-6543
Pakistan
References
- Topaz M, Roguin N, Gupta P, et al. Fabricated citations: an audit across 2.5 million biomedical papers. Lancet. 2026;407(10541):1779–1781. doi: 10.1016/S0140-6736(26)00603-3
- Kobak D, González-Márquez R, Horvát EÁ, Lause J. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Sci Adv. 2025;11(27):eadt3813. doi: 10.1126/sciadv.adt3813
- Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Sci Rep. 2023;13(1):14045. doi: 10.1038/s41598-023-41032-5
- Gao J, Zhang Y, Disis ML, Zhang L. Errors in AI-Assisted Retrieval of Medical Literature: A Comparative Study. arXiv:2603.22344 [Preprint]. 2026 Mar 21. doi: 10.48550/arXiv.2603.22344
- Tang G, Cai H. Citation contamination by paper mill articles in systematic reviews of the life sciences. JAMA Netw Open. 2025;8(6):e2515160. doi: 10.1001/jamanetworkopen.2025.15160
Supplementary files




