I flagged two research papers for fake authors and both were accepted as orals Between the two of us, we reviewed 22 paper submissions this summer, spread across NeurIPS, WACV, and TerraBytes (a geos
By Coderz Club · 2026-07-31 · Tags: ai
I flagged two research papers for fake authors and both were accepted as orals
Between the two of us, we reviewed 22 paper submissions this summer, spread across NeurIPS, WACV, and TerraBytes (a geospatial workshop at ECCV). Fifteen of the 22 (68%) contained entirely fabricated citations, fabricated author lists for existing papers, and/or were clearly LLM-generated (e.g. hallucinated technical jargon, nonsensical writing, irrelevant citations). This is called being in the “slop trenches” (i.e. dealing with the output of AI slop cannons). Here we complain about this being mostly a waste of time, do a small lit review on the state of LLMs in scientific writing and reviewing, get Claude to generate questions for a Q&A, and release a bib-audit skill that Isaac cooked up. The table below shows the number of review assignments with fabricated citations, fabricated authors, or writing that was unmistakably LLM-generated, out of the total assignments for that venue. Venue Caleb Isaac NeurIPS (Datasets and Benchmarks track) 2 of 5 NeurIPS (Position Paper track) 2 of 2 TerraBytes (ECCV workshop) 1 of 2 4 of 5 WACV 3 of 4 3 of 4 Total 6 of 11 (55%) 9 of 11 (82%) We are definitely not the only ones in the slop trenches. The scale of this problem has been measured in various ways over the past year. A Nature analysis from April found at least tens of thousands of 2025 publications “probably” contain invalid AI-generated references. An audit of arXiv, bioRxiv, SSRN, and PubMed Central by Zhao et al. estimated roughly 146,900 hallucinated citations in 2025 alone, spread thinly across many papers rather than concentrated in a few bad actors, with early-career researchers and small teams the most likely to include them. They also found that reviews are not catching these – they traced bioRxiv preprints that contained hallucinated references to their published versions and found 85.3% of the hallucinations remained. An audit in The Lancet covering 2.5 million biomedical papers found that the share of papers with at least one fabricated reference rose six-fold in two years, from one in 2828 papers in 2023 to one in 458 in 2025, reaching one in 277 in early 2026. Ansari (2026) did an analysis of 100 hallucinated citations drawn from papers that were actually published at NeurIPS 2025. Every one of those citations made it past three to five expert reviewers, and the 53 papers carrying them, about 1% of acceptances, sit in the proceedings today. The slop also flows both ways. Pangram (a company that has an AI writing detector product) did an analysis of ICLR 2026 reviews and found that 21% of the reviews (15,899 of them!) were fully AI-generated, and over half had some form of AI involvement. Gartenberg et al. (2026) measured a 42% post-ChatGPT surge in submissions to the journal (Organization Science) and found that over 30% of its peer reviews now use some degree of AI. ICML 2026 hid prompt-injection stings in submissions and found “795 reviews (~1% of all reviews) written by 506 unique reviewers who were assigned Policy A (no LLMs) were detected to have used LLMs in their review”. Now that NeurIPS reviews are out and rebuttals are coming in, we’ve seen an obvious AI-generated ethics review, and several obvious AI-generated rebuttals. This is a problem for a whole lot of reasons, one of which is that AI reviewers can be gamed directly – Li et al. (2026) found that adversarially rewritten abstracts improve AI-generated review outcomes “without changing the underlying scientific content and communication of the paper, and even without knowledge of the reviewing model.” Their strongest attack succeeded about 38% of the time, inflating acceptance ratings by +1.31 for Gemini 3 Flash reviewers and +0.88 for GPT 5.4 Mini reviewers on a 10-point scale. This all is very annoying from inside the review queue. Peer review is unpaid work that we do (often on nights and weekends) because peer review on our own work is so valuable. Spending hours going through a submission and then realizing that there are hallucinated citations is infuriating as it is a waste of our time! If you haven’t spent enough time with your work to even get the references correct, then a.) why should we spend time reviewing it for you, and b.) what are you hoping to accomplish with the submission in the first place? Learning from reviews requires reflecting on your work, and you need to spend time with your work in order to do this. Q&A How many papers did you review this summer, and how many would you have desk rejected on citations alone? Caleb: Eleven (see table above). Five of these had hallucinated authors and/or entire papers in their bibliography. In two of the WACV submissions, reference [1], the literal first entry in the bibliography, listed hallucinated authors for real papers. One of the NeurIPS papers had so much hallucinated/nonsensical jargon (53 pages of it) that there wasn’t any point in going through the bibliography. Isaac: Eleven as well: two N
Between the two of us, we reviewed 22 paper submissions this summer, spread across NeurIPS, WACV, and TerraBytes (a geospatial workshop at ECCV). Fifteen of the 22 (68%) contained entirely fabricated citations, fabricated author lists for existing papers, and/or were clearly LLM-generated (e.g. hallucinated technical jargon, nonsensical writing, irrelevant citations). This is called being in the “slop trenches” (i.e. dealing with the output of AI slop cannons). Here we complain about this being mostly a waste of time, do a small lit review on the state of LLMs in scientific writing and reviewing, get Claude to generate questions for a Q&A, and release a bib-audit skill that Isaac cooked up. The table below shows the number of review assignments with fabricated citations, fabricated authors, or writing that was unmistakably LLM-generated, out of the total assignments for that venue. Venue Caleb Isaac NeurIPS (Datasets and Benchmarks track) 2 of 5 NeurIPS (Position Paper track) 2 of 2 TerraBytes (ECCV workshop) 1 of 2 4 of 5 WACV 3 of 4 3 of 4 Total 6 of 11 (55%) 9 of 11 (82%) We are definitely not the only ones in the slop trenches. The scale of this problem has been measured in various ways over the past year. A Nature analysis from April found at least tens of thousands of 2025 publications “probably” contain invalid AI-generated references. An audit of arXiv, bioRxiv, SSRN, and PubMed Central by Zhao et al. estimated roughly 146,900 hallucinated citations in 2025 alone, spread thinly across many papers rather than concentrated in a few bad actors, with early-career researchers and small teams the most likely to include them. They also found that reviews are not catching these – they traced bioRxiv preprints that contained hallucinated references to their published versions and found 85.3% of the hallucinations remained. An audit in The Lancet covering 2.5 million biomedical papers found that the share of papers with at least one fabricated reference rose six-fold in two years, from one in 2828 papers in 2023 to one in 458 in 2025, reaching one in 277 in early 2026. Ansari (2026) did an analysis of 100 hallucinated citations drawn from papers that were actually published at NeurIPS 2025. Every one of those citations made it past three to five expert reviewers, and the 53 papers carrying them, about 1% of acceptances, sit in the proceedings today. The slop also flows both ways. Pangram (a company that has an AI writing detector product) did an analysis of ICLR 2026 reviews and found that 21% of the reviews (15,899 of them!) were fully AI-generated, and over half had some form of AI involvement. Gartenberg et al. (2026) measured a 42% post-ChatGPT surge in submissions to the journal (Organization Science) and found that over 30% of its peer reviews now use some degree of AI. ICML 2026 hid prompt-injection stings in submissions and found “795 reviews (~1% of all reviews) written by 506 unique reviewers who were assigned Policy A (no LLMs) were detected to have used LLMs in their review”. Now that NeurIPS reviews are out and rebuttals are coming in, we’ve seen an obvious AI-generated ethics review, and several obvious AI-generated rebuttals. This is a problem for a whole lot of reasons, one of which is that AI reviewers can be gamed directly – Li et al. (2026) found that adversarially rewritten abstracts improve AI-generated review outcomes “without changing the underlying scientific content and communication of the paper, and even without knowledge of the reviewing model.” Their strongest attack succeeded about 38% of the time, inflating acceptance ratings by +1.31 for Gemini 3 Flash reviewers and +0.88 for GPT 5.4 Mini reviewers on a 10-point scale. This all is very annoying from inside the review queue. Peer review is unpaid work that we do (often on nights and weekends) because peer review on our own work is so valuable. Spending hours going through a submission and then realizing that there are hallucinated citations is infuriating as it is a waste of our time! If you haven’t spent enough time with your work to even get the references correct, then a.) why should we spend time reviewing it for you, and b.) what are you hoping to accomplish with the submission in the first place? Learning from reviews requires reflecting on your work, and you need to spend time with your work in order to do this. Q&A How many papers did you review this summer, and how many would you have desk rejected on citations alone? Caleb: Eleven (see table above). Five of these had hallucinated authors and/or entire papers in their bibliography. In two of the WACV submissions, reference [1], the literal first entry in the bibliography, listed hallucinated authors for real papers. One of the NeurIPS papers had so much hallucinated/nonsensical jargon (53 pages of it) that there wasn’t any point in going through the bibliography. Isaac: Eleven as well: two N