As I have mentioned in a previous post, I spent a lot of time this summer battling bots on Reddit. I would like to say I'm pretty good by now at identifying the cadences of LLM-generated prose, especially in Reddit posts—LLM-generated Reddit posts tend to have some very specific tics. (If you want to see them in action, a very useful thing to do is to ask ChatGPT to make you a Reddit post, and you'll see how much it is like all the posts you've seen elsewhere.)
But my sense of how LLM-generated something is is not definitive proof, so in my Reddit bot takedowns, I tend also to link to AI-detector results. Fairly often, someone pops up to claim AI detectors are unreliable. I'm not so sure. I mean, I am sure there are some unreliable ones out there (anyone can create a site and claim it's an AI detector), but that doesn't mean they all are.
My poking around has turned up that it seems like the consensus is the top three are GPTZero, Pangram, and Turnitin. Obviously Turnitin is only useful to people in the academic ecosystem; I've settled on using Pangram in my Reddit takedowns, because you can link to your results page and you get a decent number of tests per day on the free account.There are two ways to think about the reliability of these tools. First is the false-negative rate. How often do they identify something that is LLM-generated as not LLM-generated? And then there's the false-positive rate. How often do they identify something that is not LLM-generated as LLM-generated? For this post, I'm focusing on the latter. This is important in the context of calling out people on Reddit and also on calling out students: if you make an accusation, you want to know that the positive result of the AI detector is reliable. (In the context of identifying academic misconduct, you also don't want a high false-negative rate: if a ton of students using LLMs are not being flagged, that's annoying... but that's a different issue than the one I'm discussing here.)
Pangram claims a false-positive rate of 0.0041% in their own materials. But, of course, they are incentivized to claim to have really low ones! What I have been curious about, then, is what independent studies turn up. Specifically, I was looking for sources where the writers used Pangram themselves on a corpus of known human-written work, and used that to calculate a false-positive rate. (Or, where I could use their data to calculate it, if they didn't do it themselves.) I excluded anything that was directly published by Pangram itself, or anyone working for Pangram. (Fun fact, though: Adam et al. is published by Pangram competitor GPTZero!) Here's what I have been able to find.* I have sorted the results by denominator, i.e., I start with the study that had the largest corpus of human-generated texts and work my way down to the smallest. To make this table legible, I just included author names in it; full citations are in my Works Cited at the end of this post.
| Source | Peer Review? | Specific Citation | False-Positive Rate |
|---|---|---|---|
| Evanko and Di Natale | Conference paper | Calibrated on abstracts from 2021 to 2022 (slide 4). | 0.22% (40/18,467) |
| Adam et al. | Preprint | Tables 1-2 on p. 7 give the accuracy (all texts classified correctly) and recall (true-positive rate) across five domains, each of which consists of 1,000 AI texts and 1,000 human texts; only in one domain did Pangram have 2 false positives. | ~0.04% (~2/5,000) |
| Elazar and Antoniak | Preprint | “We test [...] Pangram on CS papers from 2020- 2022. We find that [...] Pangram did not predict any false positives” (p. 3); they sampled 100 papers per month for the three pre-LLM years of 2020-22 (p. 4). |
0.0% (0/3,600) |
| Saha et al. | Peer reviewed | “For reviews that are fully human written, Pangram flags none of them as ‘AI’ or ‘Mixed’” (p. 3); number of human reviews given on p. 5. | 0.0% (0/3,499) |
| Jabarian and Imas | No | “FPR is essentially zero across all thresholds 0.5 and greater” (p. 10); number of human texts given on p. 5. | 0.1% (2/1,992) |
| Day | No | “~1,000 random abstracts were retrieved from articles published in 2010”; “1 positive detection in 2010 data.” | 0.1% (1/974) |
| Shah and Levy | Preprint | “The pre-AI period (2019–2022) yields 1 detection across 800 documents, a rate of 0.1%” (p. 28). | 0.1% (1/800) |
| Lee | No | “Pangram and GPTZero correctly flagged no human writing as AI.” | 0.0% (0/495) |
| Karr et al. | Conference paper | FPR for 2013-15 control period given in Table 2 on p. 3; number of human texts given in Table 1 on p. 2. | 0.0% (0/306) |
| Saha and Feizi | Conference paper | FPR given in Table 6 on p. 25426. | 0.0% (0/300) |
| Allaham and Diakopoulos | Preprint | “In our evaluation of both Pangram and GPTZero on a curated dataset of 200 human-authored texts [...], we find that Pangram and GPTZero have 100% accuracy in correctly detecting human-authored and AI-generated texts, respectively” (p. 3). |
0.0% (0/200) |
| Lu et al. | No | No AI-generated papers found in the 2022 control group (Table 2). | 0.0% (0/159) |
| Russell et al. | Peer reviewed | FPR given in Table 2.B on p. 5346; “Each experiment consists of 60 articles, 30 human-written and 30 LLM-generated” (p. 5343), and there were five experiments. | 2.0% (3/150) |
| Chakrabarty et al. | Preprint | “Human written text was never misclassified” (fig. 2E, p. 5). | 0.0% (0/150) |
| Ren et al. | Preprint | Ran a control on abstracts posted to arXiv in 2010: “Pangram also perfectly classifies the human abstracts” (p. 42). | 0.0% (0/100) |
| Guppenberger | Preprint | FPR given in Table 2 on p. 10. | 0.0% (0/50) |
| Van Vlasselaer et al. | Peer reviewed | “The results for fully human-written texts show that all four tools classified 100% of human texts correctly as ‘True Negatives’ [...]. This indicates that none of the detectors are prone to incorrectly flagging human writing as AI-generated for this collection of papers” (p. 12). | 0.0% (0/40) |
You'll see that the majority of independent tests give Pangram a 0% false-positive rate, including one where the corpus included 3,600 texts! If you just crudely average all of the false-positive rates from all seventeen studies, the mean false-positive rate is 0.15%. If you pool all the numerators and denominators, you get 49 / 36,282, which is a false positive rate of 0.14%.†
This would suggest that it would take you 658 submissions to find one false positive according to the crude mean, and 740 submissions to find one false positive according to the pooled average. So overall, Pangram has a very low false-positive rate, and thus you should feel entitled to make LLM-use accusations with a high degree of confidence if you have a Pangram score as support.
Note that many of these studies had other goals in mind than just assessing the false-positive rate of Pangram, so in some cases, it's a very small part of a wider topic. Often, this other stuff is quite interesting: Ren et al. and Karr et al. both discuss, for example, the susceptibility of LLM detectors like Pangram to attempts to evade AI detection, thus resulting in a higher false-negative rate. (It's worth pointing out that some other of these studies, such as Russell et al., Jabarian and Imas, and Van Vlasselaer et al. tried out evasion techniques that Pangram was able to still detect, however.)
One that I found particularly interesting was Russell et al., whose purpose was to compare automated detection tools to human beings' own ability to detect LLM-generated prose. They hired human beings to annotate their pool of texts, and found that
annotators who rarely or never use LLMs are poor detectors of AI-generated text; in fact, they overestimate their own ability to perform the task by providing high confidence scores for their decisions. We identify a subset of five high-performing annotators who frequently use LLMs for writing tasks (e.g., editing, copywriting, creative writing). The majority vote of this subset of “expert” annotators fails to predict the correct label on only one out of 300 articles. (5342–3)
In fact, only Pangram outperformed their group of "expert" annotators. (This might account for my own ability to detect LLM-generated prose; while I have certainly never used ChatGPT to generate content for this blog or academic writing, I do mess with it a lot in other ways, including helping me devise images and Tasks for my Star Trek Adventures scenarios and perform analyses on my iTunes library, so I have seen a lot of its prose in action by now.)
The other study I found particularly interesting was Lu et al. This was done by the chairs of a track at the conference NeurIPS, who "made the decision to require that all papers be substantially human-written, with AI used for only copy-editing or similar peripheral changes..." They partnered with Pangram to enforce this policy, and used papers from 2022 as a control. 0% of those papers were identified by Pangram as AI-written. When they used it on the 2026 batch, they had to outright desk-reject 18.4% of submissions for AI use, and ask for "evidence of substantial human engagement" on another 12.7%. A full 30% of people used AI to write their position papers knowing that this track of NeurIPS had a no-AI policy! It's not just undergraduate students who can't be bothered to obey academic integrity rules, and who don't think about the implications of using LLMs for their writing tasks.
* Full disclosure: this is a combination of sources I found myself and supplemental ones found for me by ChatGPT (and one by the world's most underachieving LLM, Google Gemini). ChatGPT also helped me do the HTML formatting for my table and Works Cited.
† There were two other pieces I chose not to include in my table above, partially because they were just published by individuals on Substack, but they are both also interesting reading. @VeryFinePrint OCRed forty-six undigitized old books, which came out to 8,025 of what Pangram calls "segments," and got false positives on just 4 segments, so 0.05%... but it turned out that in all four segments, their AI-based OCR program had hallucinated text, so Pangram was accurately detecting AI-written segments! The other, by Kelly, states that he tested 96,468 essays and got just 10 false negatives and 1 false positive, but doesn't actually state how many of the 96,468 essays were human-written, so we don't know what the denominator should be. I have included both sources on the Works Cited below.
Works Cited
Adam, George Alexandru, et al. “GPTZero: Robust Detection of LLM-Generated Texts.” arXiv, 13 Feb. 2026, doi: 10.
Allaham, Mowafak, and Nicholas Diakopoulos. “Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources.” arXiv, 22 May 2026, doi: 10.
Chakrabarty, Tuhin, et al. “Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers.” SSRN, 19 Aug. 2026, doi: 10.
Day, Adam. “GenAI Detection That Actually Works.” Clear Skies, Medium, 18 Aug. 2025, clearskiesadam.
Elazar, Yanai, and Maria Antoniak. “LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on arXiv.” arXiv, 19 Jan. 2026, doi: 10.
Evanko, Daniel S., and Michael Di Natale. “Quantifying and Assessing the Use of Generative AI by Authors and Reviewers in the Cancer Research Field.” International Conference on Peer Review and Scientific Publication, 3 Sept. 2025, peerreviewcongress.
Guppenberger, James. “Vernacular Remote Control: Adversarial Robustness Testing of Six Commercial AI Text Detectors on Provenance-Guaranteed Pre-AI Prose.” Research Square, 12 June 2026, doi: 10.
Jabarian, Brian, and Alex Imas. “Artificial Writing and Automated Detection.” National Bureau of Economic Research, Sept. 2025, doi: 10.
Karr, Jonathan A., Jr., et al. “Why AI Detection Fails for Academic Integrity.” ACM AI Leadership Summit, 20 Aug. 2026. arXiv, doi: 10.
Kelly, Trent. “Practical Attacks on AI Text Classifiers with RL.” Substack, 8 July 2025, trentmkelly.
Lee, Jaeho. “AI Detectors Rarely Flag Human Writing, but Sometimes Miss AI Text Imitating Real Authors.” Epoch AI, 15 July 2026, epoch.
Lu, Alex, et al. “AI-Generated Papers in the NeurIPS 2026 Position Paper Track.” NeurIPS, 2 June 2026, blog.
Ren, Kevin, et al. “Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift.” arXiv, 23 June 2026, doi: 10.
Russell, Jenna, et al. “People Who Frequently Use ChatGPT for Writing Tasks Are Accurate and Robust Detectors of AI-Generated Text.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, July 2025, pp. 5342–73. ACL Anthology, doi: 10.
Saha, Rounak, et al. “Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable.” International Conference on Machine Learning, 23 June 2026. arXiv, doi: 10.
Saha, Shoumik, and Soheil Feizi. “Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing.” Findings of the Association for Computational Linguistics, July 2025, pp. 25414–31. ACL Anthology, doi: 10.
Shah, Anand, and Joshua Levy. “Access to Justice in the Age of AI: Evidence from U.S. Federal Courts.” SSRN, 21 May 2026, doi: 10.
Van Vlasselaer, Marijke, et al. “Who Wrote This? Evaluating the Reliability of AI Detection Tools in Higher Education.” International Journal for Educational Integrity, vol. 22, 29 June 2026. Springer Nature Link, doi: 10.
@VeryFinePrint. “Scanning for Pangram Errors.” Substack, 15 July 2026, veryfineprint.

No comments:
Post a Comment