Showing posts with label topic: llms. Show all posts
Showing posts with label topic: llms. Show all posts

25 September 2026

Are the Writers Upset about Substack's Pangram Integration Just the Ones Using AI to Write?

A point that I came across in writing my previous post about the false-positive rate of the Pangram AI detection software was that the blogging platform Substack recently integrated Pangram, with the position that while using AI to write is fine, undisclosed AI writing is not. (This seems like a reasonable position to me, to be honest.) Many Substack writers are now outraged, claiming their purely human-written and -polished prose is getting flagged by Pangram as AI-written.

This made me remember when, a few months ago, a facebook friend shared an essay on Tolkien by the Substack writer Genny Harrison, and reading it immediately tripped my internal AI detector:

"Not because the claim was obviously absurd, but because it was obviously compressed."
"That question is not radical. It is basic historiography."
"Recognizing that possibility is not an attack on Tolkien. It is what historical reading looks like."
"Tolkien is more than an author. He is a formative experience."
"it is no longer a myth. It has become a shrine."
"Middle-earth is not a crime scene. It is an inheritance."

If you've spent any time reading AI-generated prose, it couldn't be any more obvious. And indeed, if you copy a chunk of it into Pangram, it scores it as 100% AI-generated.

I was curious if she admitted to AI use or not, so I searched Harrison's facebook profile to see if she mentioned it, and found a post where she decried the integration of Substack's AI detector: 

The classifier is also only as reliable as the moment its labels were collected. Models change constantly. Humanizer tools exist for the sole purpose of defeating detection, and many of them work well enough to make the entire exercise resemble airport security for adjectives.

Think about who that leaves. The person running an essay mill through a laundering tool walks through clean. The writer most likely to be flagged is the one with a consistent voice, careful structure, and a habit of revising until the sentences behave.

That is the familiar genius of automated enforcement. The sophisticated learn how to evade it. The honest remain available for inspection.
Again, if you've read a lot of LLM-generated stuff, this sounds exactly like it. And indeed, well, I leave you to guess what the Pangram result is. One might note with a careful read of Harrison's piece that she never actually says, "I do not use AI to write my posts." She says she puts a lot of work into them, she frets about a lot of harms that AI detection might do in the abstract. I feel like this non-denial is pretty telling.
 
(One also notes that Harrison claims of Pangram's training data, "Someone wrote all of it. Somebody's novels, somebody's blog posts, somebody's dissertation, somebody's newsletter about mushrooms, all of it harvested and processed into a classifier that now stands at the door of my essay, deciding whether I am real. The exact grievance the anti-machine crowd has been shouting for three years, that these systems were built out of human labor nobody agreed to donate, applies with full force to the referee they just installed. They did not defeat the thing they hate. They gave it a badge and a percentage sign." This in fact totally untrue. Pangram states its datasets of human writing are "commercially licensed," which they actually paid for it, they didn't just rip it off the Internet as so many LLM creators did.)
 
a robot writing
(image courtesy NBC News)
Jump over to the Substack subreddit, and the posters there are outraged. Someone created a post called "Running list of Pangram’s 'AI Detection' errors and other related Substack errors." The comments are filled with people saying they didn't use AI, they've never even read anything written by AI, yet somehow their posts are being flagged!
 
But... can we trust them? I mean, Genny Harrison very strongly implies she doesn't use AI, enough that I would feel free to call her purposefully dishonest, and yet her essays score very highly. Anyone on Reddit can say, "oh, AI detectors are so unreliable," but what proves that they didn't use an LLM? Hence why the Pangram tests I covered in my previous post used pre-2022 material as a control when calibrating Pangram.
 
My favorite commenter on that reddit posts is the one who says,
I've got an open offer for anyone reading this. If you can find a piece of pre-2022 writing, with an accompanying web.archive.org link proving it's pre-AI writing, that is at least 100 words in length, and which Pangram flags as AI-generated, I will donate $100 to the charity of your choice. Everyone says they're getting tons of false positives, so this should be pretty easy, but for some reason every time I've posted this offer before, nobody has been able to do it. 
As far as I could tell, no one ever took them up on this offer.
 
Case closed?
 
Well, maybe not. In the course of researching this post, I came across another Substack post by the blogger Freddie deBoer (I read his work occasionally maybe a decade ago): "I Wouldn't Say Pangram is Broken, But I Would Say That It's Brittle." DeBoer says a reader accused him of using LLM to write part of one post, and linked to a Pangram result; deBoer actually paid for a Pangram account so he could mess around with it a bit and was able to reproduce this result. I am familiar enough with deBoer to believe that he is telling the truth when he says he did not use AI.
 
His post is worth reading in full, but points out the issues Pangram has with 1) very short texts, and 2) texts that switch back and forth between AI- and LLM-generated writing. One of the cofounders of Pangram actually pops up in the comments to concede Pangram needs to be better about handling these things. So though it seems that while Pangram might be very good... it also has its vulnerabilities. I do think, going by deBoer's post, that the frequency with which Pangram rates even longer texts as either 0% or 100% is a bit suspect, never in between, and the fact that it doesn't better foreground its confidence probability is also something that could use some improvement.
 
If you want to know, incidentally, I put a chunk of my previous post about Pangram into Pangram, and it said I am 100% human. 

11 September 2026

What Is the Pangram AI Detector's False-Positive Rate for LLM-Generated Prose?

As I have mentioned in a previous post, I spent a lot of time this summer battling bots on Reddit. I would like to say I'm pretty good by now at identifying the cadences of LLM-generated prose, especially in Reddit posts—LLM-generated Reddit posts tend to have some very specific tics. (If you want to see them in action, a very useful thing to do is to ask ChatGPT to make you a Reddit post, and you'll see how much it is like all the posts you've seen elsewhere.)

But my sense of how LLM-generated something is is not definitive proof, so in my Reddit bot takedowns, I tend also to link to AI-detector results. Fairly often, someone pops up to claim AI detectors are unreliable. I'm not so sure. I mean, I am sure there are some unreliable ones out there (anyone can create a site and claim it's an AI detector), but that doesn't mean they all are.

My poking around has turned up that it seems like the consensus is the top three are GPTZero, Pangram, and Turnitin. Obviously Turnitin is only useful to people in the academic ecosystem; I've settled on using Pangram in my Reddit takedowns, because you can link to your results page and you get a decent number of tests per day on the free account.

There are two ways to think about the reliability of these tools. First is the false-negative rate. How often do they identify something that is LLM-generated as not LLM-generated? And then there's the false-positive rate. How often do they identify something that is not LLM-generated as LLM-generated? For this post, I'm focusing on the latter. This is important in the context of calling out people on Reddit and also on calling out students: if you make an accusation, you want to know that the positive result of the AI detector is reliable. (In the context of identifying academic misconduct, you also don't want a high false-negative rate: if a ton of students using LLMs are not being flagged, that's annoying... but that's a different issue than the one I'm discussing here.)

Pangram claims a false-positive rate of 0.0041% in their own materials. But, of course, they are incentivized to claim to have really low ones! What I have been curious about, then, is what independent studies turn up. Specifically, I was looking for sources where the writers used Pangram themselves on a corpus of known human-written work, and used that to calculate a false-positive rate. (Or, where I could use their data to calculate it, if they didn't do it themselves.) I excluded anything that was directly published by Pangram itself, or anyone working for Pangram. (Fun fact, though: Adam et al. is published by Pangram competitor GPTZero!) Here's what I have been able to find.* I have sorted the results by denominator, i.e., I start with the study that had the largest corpus of human-generated texts and work my way down to the smallest. To make this table legible, I just included author names in it; full citations are in my Works Cited at the end of this post.

Source Peer Review?  Specific Citation False-Positive Rate
Evanko and Di Natale Conference paper Calibrated on abstracts from 2021 to 2022 (slide 4). 0.22% (40/18,467)
Adam et al. Preprint Tables 1-2 on p. 7 give the accuracy (all texts classified correctly) and recall (true-positive rate) across five domains, each of which consists of 1,000 AI texts and 1,000 human texts; only in one domain did Pangram have 2 false positives. ~0.04% (~2/5,000)
Elazar and Antoniak Preprint “We test [...] Pangram on CS papers from 2020-
2022. We find that [...] Pangram did not predict
any false positives” (p. 3); they sampled 100 papers per month for the three pre-LLM years of 2020-22 (p. 4).
0.0% (0/3,600)
Saha et al. Peer reviewed “For reviews that are fully human written, Pangram flags none of them as ‘AI’ or ‘Mixed’” (p. 3); number of human reviews given on p. 5. 0.0% (0/3,499)
Jabarian and Imas No “FPR is essentially zero across all thresholds 0.5 and greater” (p. 10); number of human texts given on p. 5. 0.1% (2/1,992)
Day No “~1,000 random abstracts were retrieved from articles published in 2010”; “1 positive detection in 2010 data.” 0.1% (1/974)
Shah and Levy Preprint “The pre-AI period (2019–2022) yields 1 detection across 800 documents, a rate of 0.1%” (p. 28). 0.1% (1/800)
Lee No “Pangram and GPTZero correctly flagged no human writing as AI.” 0.0% (0/495)
Karr et al. Conference paper FPR for 2013-15 control period given in Table 2 on p. 3; number of human texts given in Table 1 on p. 2. 0.0% (0/306)
Saha and Feizi Conference paper FPR given in Table 6 on p. 25426. 0.0% (0/300)
Allaham and Diakopoulos Preprint “In our evaluation of both Pangram and GPTZero on
a curated dataset of 200 human-authored texts [...], we find that Pangram and GPTZero have 100% accuracy in correctly detecting human-authored and AI-generated texts, respectively” (p. 3).
0.0% (0/200)
Lu et al. No No AI-generated papers found in the 2022 control group (Table 2). 0.0% (0/159)
Russell et al. Peer reviewed FPR given in Table 2.B on p. 5346; “Each experiment consists of 60 articles, 30 human-written and 30 LLM-generated” (p. 5343), and there were five experiments. 2.0% (3/150)
Chakrabarty et al. Preprint “Human written text was never misclassified” (fig. 2E, p. 5). 0.0% (0/150)
Ren et al. Preprint Ran a control on abstracts posted to arXiv in 2010: “Pangram also perfectly classifies the human abstracts” (p. 42). 0.0% (0/100)
Guppenberger Preprint FPR given in Table 2 on p. 10. 0.0% (0/50)
Van Vlasselaer et al. Peer reviewed “The results for fully human-written texts show that all four tools classified 100% of human texts correctly as ‘True Negatives’ [...]. This indicates that none of the detectors are prone to incorrectly flagging human writing as AI-generated for this collection of papers” (p. 12). 0.0% (0/40)

You'll see that the majority of independent tests give Pangram a 0% false-positive rate, including one where the corpus included 3,600 texts! If you just crudely average all of the false-positive rates from all seventeen studies, the mean false-positive rate is 0.15%. If you pool all the numerators and denominators, you get 49 / 36,282, which is a false positive rate of 0.14%.†

This would suggest that it would take you 658 submissions to find one false positive according to the crude mean, and 740 submissions to find one false positive according to the pooled average. So overall, Pangram has a very low false-positive rate, and thus you should feel entitled to make LLM-use accusations with a high degree of confidence if you have a Pangram score as support.

Note that many of these studies had other goals in mind than just assessing the false-positive rate of Pangram, so in some cases, it's a very small part of a wider topic. Often, this other stuff is quite interesting: Ren et al. and Karr et al. both discuss, for example, the susceptibility of LLM detectors like Pangram to attempts to evade AI detection, thus resulting in a higher false-negative rate. (It's worth pointing out that some other of these studies, such as Russell et al., Jabarian and Imas, and Van Vlasselaer et al. tried out evasion techniques that Pangram was able to still detect, however.)

One that I found particularly interesting was Russell et al., whose purpose was to compare automated detection tools to human beings' own ability to detect LLM-generated prose. They hired human beings to annotate their pool of texts, and found that

annotators who rarely or never use LLMs are poor detectors of AI-generated text; in fact, they overestimate their own ability to perform the task by providing high confidence scores for their decisions. We identify a subset of five high-performing annotators who frequently use LLMs for writing tasks (e.g., editing, copywriting, creative writing). The majority vote of this subset of “expert” annotators fails to predict the correct label on only one out of 300 articles. (5342–3)

In fact, only Pangram outperformed their group of "expert" annotators. (This might account for my own ability to detect LLM-generated prose; while I have certainly never used ChatGPT to generate content for this blog or academic writing, I do mess with it a lot in other ways, including helping me devise images and Tasks for my Star Trek Adventures scenarios and perform analyses on my iTunes library, so I have seen a lot of its prose in action by now.)

The other study I found particularly interesting was Lu et al. This was done by the chairs of a track at the conference NeurIPS, who "made the decision to require that all papers be substantially human-written, with AI used for only copy-editing or similar peripheral changes..." They partnered with Pangram to enforce this policy, and used papers from 2022 as a control. 0% of those papers were identified by Pangram as AI-written. When they used it on the 2026 batch, they had to outright desk-reject 18.4% of submissions for AI use, and ask for "evidence of substantial human engagement" on another 12.7%. A full 30% of people used AI to write their position papers knowing that this track of NeurIPS had a no-AI policy! It's not just undergraduate students who can't be bothered to obey academic integrity rules, and who don't think about the implications of using LLMs for their writing tasks.

* Full disclosure: this is a combination of sources I found myself and supplemental ones found for me by ChatGPT (and one by the world's most underachieving LLM, Google Gemini). ChatGPT also helped me do the HTML formatting for my table and Works Cited.

† There were two other pieces I chose not to include in my table above, partially because they were just published by individuals on Substack, but they are both also interesting reading. @VeryFinePrint OCRed forty-six undigitized old books, which came out to 8,025 of what Pangram calls "segments," and got false positives on just 4 segments, so 0.05%... but it turned out that in all four segments, their AI-based OCR program had hallucinated text, so Pangram was accurately detecting AI-written segments! The other, by Kelly, states that he tested 96,468 essays and got just 10 false negatives and 1 false positive, but doesn't actually state how many of the 96,468 essays were human-written, so we don't know what the denominator should be. I have included both sources on the Works Cited below.

Works Cited

Adam, George Alexandru, et al. “GPTZero: Robust Detection of LLM-Generated Texts.” arXiv, 13 Feb. 2026, doi: 10.48550/arXiv.2602.13042.

Allaham, Mowafak, and Nicholas Diakopoulos. “Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources.” arXiv, 22 May 2026, doi: 10.48550/arXiv.2605.23684.

Chakrabarty, Tuhin, et al. “Readers Prefer Outputs of AI Trained on Copyrighted Books over Expert Human Writers.” SSRN, 19 Aug. 2026, doi: 10.2139/ssrn.5606570.

Day, Adam. “GenAI Detection That Actually Works.” Clear Skies, Medium, 18 Aug. 2025, clearskiesadam.medium.com/genai-detection-that-actually-works-edc562581fed.

Elazar, Yanai, and Maria Antoniak. “LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on arXiv.” arXiv, 19 Jan. 2026, doi: 10.48550/arXiv.2601.17036.

Evanko, Daniel S., and Michael Di Natale. “Quantifying and Assessing the Use of Generative AI by Authors and Reviewers in the Cancer Research Field.” International Conference on Peer Review and Scientific Publication, 3 Sept. 2025, peerreviewcongress.org/abstract/quantifying-and-assessing-the-use-of-generative-ai-by-authors-and-reviewers-in-the-cancer-research-field.

Guppenberger, James. “Vernacular Remote Control: Adversarial Robustness Testing of Six Commercial AI Text Detectors on Provenance-Guaranteed Pre-AI Prose.” Research Square, 12 June 2026, doi: 10.21203/rs.3.rs-9596240/v1.

Jabarian, Brian, and Alex Imas. “Artificial Writing and Automated Detection.” National Bureau of Economic Research, Sept. 2025, doi: 10.3386/w34223.

Karr, Jonathan A., Jr., et al. “Why AI Detection Fails for Academic Integrity.” ACM AI Leadership Summit, 20 Aug. 2026. arXiv, doi: 10.48550/arXiv.2608.11256.

Kelly, Trent. “Practical Attacks on AI Text Classifiers with RL.” Substack, 8 July 2025, trentmkelly.substack.com/p/practical-attacks-on-ai-text-classifiers.

Lee, Jaeho. “AI Detectors Rarely Flag Human Writing, but Sometimes Miss AI Text Imitating Real Authors.” Epoch AI, 15 July 2026, epoch.ai/data-insights/ai-detectors-false-negatives.

Lu, Alex, et al. “AI-Generated Papers in the NeurIPS 2026 Position Paper Track.” NeurIPS, 2 June 2026, blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track.

Ren, Kevin, et al. “Hitting a Moving Target: Test-Time Adaptation for AI Text Detection under Continual Distribution Shift.” arXiv, 23 June 2026, doi: 10.48550/arXiv.2606.25152.

Russell, Jenna, et al. “People Who Frequently Use ChatGPT for Writing Tasks Are Accurate and Robust Detectors of AI-Generated Text.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, July 2025, pp. 5342–73. ACL Anthology, doi: 10.18653/v1/2025.acl-long.267.

Saha, Rounak, et al. “Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable.” International Conference on Machine Learning, 23 June 2026. arXiv, doi: 10.48550/arXiv.2603.20450.

Saha, Shoumik, and Soheil Feizi. “Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing.” Findings of the Association for Computational Linguistics, July 2025, pp. 25414–31. ACL Anthology, doi: 10.18653/v1/2025.findings-acl.1303.

Shah, Anand, and Joshua Levy. “Access to Justice in the Age of AI: Evidence from U.S. Federal Courts.” SSRN, 21 May 2026, doi: 10.2139/ssrn.6766859.

Van Vlasselaer, Marijke, et al. “Who Wrote This? Evaluating the Reliability of AI Detection Tools in Higher Education.” International Journal for Educational Integrity, vol. 22, 29 June 2026. Springer Nature Link, doi: 10.1007/s40979-026-00226-w.

@VeryFinePrint. “Scanning for Pangram Errors.” Substack, 15 July 2026, veryfineprint.substack.com/p/scanning-for-pangram-errors.

28 August 2026

The Internet Is Dying

You may be familiar with "Dead Internet Theory," which states that "activity and content on the internet, including social media accounts, are predominantly being created and automated by artificial intelligence agents" and that "[m]any of the accounts that engage with such content also appear to be managed by artificial intelligence agents."

I don't know if the Internet actually is dead yet, but it certainly seems to be dying. Facebook is always pushing pages on me when I just want to see what my friends are up to; a lot of those are clearly AI-written prose. But even some of my friends who should know better often share AI-written posts!

image courtesy Stackery
Reddit, however, is even worse. So many subs have, of late, been taken over by bots that post LLM-generated content. The problem I have run into is that many posters don't seem to recognize this—these posts often get a ton of engagement and upvotes. I am pretty good at recognizing it, because of course I read hundred of pages of LLM-generated prose every semester. But a stranger on the Internet only has my word for it. Though one of the things I am coming to learn from all of this is that I think people who are bad at reading don't understand that other people can actually be good at it.

I wouldn't say I am perfect at it, but I am pretty well calibrated. I usually check by using Pangram; Pangram claims to have a false-positive rate below 1%. Of course they're incentivized to say that, but if you poke around on Google Scholar, the trend of independent analyses of Pangram seems to put the false-positive rate below 3%. But even if you link people to the Pangram result, they'll still be all, "But how do you know?" There are some bad LLM detectors out there which have given the good ones a bad rap. (And, in any case, I suspect most of the Ai DeTeCtOrS aRe So UnReLiAbLe people are probably the ones themselves using LLMs to write.)

It's infuriating to me that I have to deal with this on Reddit, because if I want to read crappy LLM-generated prose, I can just do it at work! Reddit is supposedly where I go to relax and have fun. On the other hand, there is a bit of satisfaction I get from getting a post deleted—I am certainly the number one bot hunter on a couple of subs.

I am getting better at making my case, though. The key is often to be able to link to a pattern of behavior—and this is also the key to understand what's going on, because that's probably the question I get the most. What is a bot's incentive to post a heartwarming parenting story on Daddit or a broadly worded book recommendation request on r/printSF? As the Conversation article I linked to above says, "This creates a vicious cycle of artificial engagement, one that has no clear agenda and no longer involves humans at all." That article discusses propaganda, but there are other reasons at play when it comes to the Dead Internet.

At least when it comes to Reddit, it's about slipping in product recommendations. Take this user, for example. Aside from all being LLM-generated the posts are pretty innocuous, but read every single one of them, and a pattern begins to emerge:

  • A post in r/PostConcussion is supposedly about memory issues but links to an online concussion test they took.
  • A post in r/eldercare is about their parent getting scammed by AI, and includes a link to a for-profit service that helps people get their money back. 
  • A post in r/Entrepreneurs is supposedly about getting recommendations on how to fight fake reviews, and links to one specific service that will help. 
  • A post in r/PaintByNumbers asks if other posters have painted places they want to go to, but slips in a mention of a specific kit from a specific company. 
  • A post in r/family supposedly asks for advice about telling her sister she could have bought her engagement ring for $6,000 less online, and makes sure to mention the site. 

Note that of these five, three include a link, though two do not. 

(These bots make it tricky by hiding their posting history, but you can usually find the posts via Google, or the very helpful Arctic Shift tool. Arctic Shift reveals another one by this poster that got deleted on r/OldHomeRepair, where the poster asks about a quote they got for modernization.)

Obviously links help SEO, but what's up with the nonlinked posts? And some are even critical of the services mentioned, like the one on r/OldHomeRepair?

The answer turns out to be... other LLMs! There's a good 404 Media article on it (here's a link to an archived version, because it's paywalled). Lots of people ask LLMs for product recommendations these days. What's the best place for people to get production recommendations? Well, Reddit, because it is of course a compilation of real human experiences! Reddit posts get scraped and searched by LLMs and absorbed into their training data. So if you ask ChatGPT to recommend good paint-by-number kits, it's going to recommend one that lots of people on Reddit mentioned that they liked!

The posts that don't mention products, in between them, are to build up karma—hence heartwarming and/or ragebait stories on Daddit, for example. In fact, I've observed a pattern: often the bots build up comment karma with short, random comments in popular subs, then move on to making posts to farm karma, then move on to product recs.

Of course there's a tragedy-of-the-commons problem here. If everyone markets their products on Reddit via bots, then Reddit will cease to be the repository of real human recommendations that made the bots want to post there to begin with.

Although... I'm becoming increasingly skeptical that real human experience is what people want, anyway. Isn't it just easier to talk to bots? Maybe it's not the Internet that's dying, but the whole human social experience. 

20 June 2025

We Are Now Firmly in the ChatGPT Era of Academic Writing

It seems unlikely that if you teach college writing, that you haven't been dealing with the impact of ChatGPT and other LLMs. I of course have had students using this next technology since Spring 2023. But the impact it's had on me and and my teaching is probably most starkly represented by this chart:


According to my syllabus policies, as well as our program policies, use of ChatGPT or similar technology is prohibited. Though I guess I can see how it might be useful in other courses, I don't see how it's useful in a writing course; my class (I would argue) is trying to use writing as a mode of thinking. We don't write to record what we already know, we write to come into the knowing of something.

The thing about ChatGPT use is that it's difficult to "prove" in some kind of "objective" sense—and this is the kind of thing academic integrity panels supposedly want. But the whole reason I am a professor of academic writing is that I supposedly have some kind of expertise in academic writing, and I do. It's an expertise honed by reading student writing for years, almost decades. I first taught an academic writing course in Fall 2008. Without methodically counting it all up, I would estimate that in that time I've taught sixty-five sections of first-year writing course, which means I've read the work of approximately thirteen hundred students, each of whom ought to write at a minimum of ten pages per semester, if we're going to be conservative. That means I've read at least (and certainly much more than) thirteen thousand pages of college-student papers in my life.

image generated, of course, by ChatGPT
You read a genre that much, you get to know it. I know what college students write like. What I'm reading now isn't it. But frustratingly, I don't know how much the academic integrity panel goes for "it just doesn't sound right" as evidence. A colleague of mine has suggested we should file more charges on this basis, though, and see what happens. Another colleague of mine has suggested that how college students write may be shifting as a result of ChatGPT, in that even if they're not actually using it your class, they read its output so much that it's influencing how they write. Now that's scary.

But anyway, clearly lots of them are using it. (Though, admittedly, some are clearly not! I have a lot of fondness for the crappy C paper now.)

The thing you can actually prove, though, is the veracity of sources. That is the basis on which all of my charges were filed this semester. Of my fourteen cases, I think three were about sources that did not exist. This, for me, means failure of the course. If you don't see that a basic part of research writing is that the sources you cite have to actually exist, then I don't know that you really belong in my class at all. I can't teach you this. Sure, use ChatGPT to find sources (in my personal experiments I have not found it to be very good at this, but I am sure it could be), but then make sure they are real! And of course, if you're summarizing fabricated sources on, say, an annotated bibliography, you are lying, because how could you have read a source that doesn't exist?

All of my other charges were about fabricated quotations from real sources, students claiming that direct quotations existed that did not actually exist. My institution requires that all academic integrity filings be accompanied by a formal meeting that is witness by a neutral faculty member; to make up for all the colleagues that had to witness mine, I witness a lot of other people's. This was the thing most of them saw as well. Frustratingly, a lot of students had this weird defense: that they didn't know quotation marks were reserved for direct quotations. One might tempted to believe this, except that very few students were making this mistake over a year ago, and though I think that the teaching of writing at the pre-college level has probably got worse since I started teaching in 2008, I don't believe that just a couple years ago, teachers stopped explaining what quotations marks are for.

So, the only explanation is that they are using ChatGPT to "find quotations"... but again, not confirming the existence of the quotations. Some of my students have admitted this when confronted, others have come up with not very compelling explanations. I have seen this across the board, in both my research writing class and in my text-based humanities course. (I can see why students might think I might not know that they made up a quotation from a source they found that I've never read; I don't see why they don't realize I won't catch them when they make up quotes from stories we all read together!)

Thankfully (I guess) it doesn't matter. At my institution, fabricated quotations or sources fall under the category of Deception and misrepresentation; I don't have to prove the fabrications come from anywhere in particular, I just have to prove they don't exist.

In my research writing course, fabrications in the final paper mean failing the final paper. It seems to me there's a basic parameter of the assignment you've failed to reach. And failing the final paper means failing the course. It seems to me there's something this course was supposed to teach you that you just didn't learn, and thus you need to take it again, if you can't write at least a D-level research paper by the end of an entire semester. In the past, I haven't had to enforce this policy very much, but I did at least three or four times this semester.

On top of all this, one has to remember—these are just the students I caught fabricating. There are also all the students using LLMs that I recognized but couldn't "prove," and thus just dinged them on points (I have become fond of marking assignments with a "0" and writing, "This is so vague an AI could have written it"), and then all the students who are actually "good" at using LLMs to write and turning in work that seems human-written... even when it's not. 

Even though ChatGPT has been around a couple years, it's clear that something is shifting in the way students use it. You can see that even just last semester, I only filed one charge (though I didn't teach research writing that time.) Many of my colleagues had similar experiences this semester in particular (though I don't think anyone filed as many charges as I did). As one of my colleagues has said, it seems like there's a shared ethos we used to assume the existence of that just doesn't exist anymore. And how can you teach people if that's the case? I can't teach someone that they want to be ethical.

So anyway, it's been a depressing semester. I even had two students file appeals that persisted beyond the end of the semester (and another just not respond to my communication attempts, which is for some reason an automatic appeal), which meant that it didn't stop even when the semester was over! But finally about two weeks ago, I heard about my last outstanding appeal.

I can't just keep doing everything the same way, clearly. But more on that in another post.