A
point that I came across in writing my previous post about the false-positive rate of the Pangram AI detection software was that the blogging platform
Substack recently integrated Pangram, with the position that while
using AI to write is fine, undisclosed AI writing is not. (This
seems like a reasonable position to me, to be honest.) Many Substack
writers are now outraged, claiming their purely human-written and
-polished prose is getting flagged by Pangram as AI-written.
This made me remember when, a few months ago, a facebook friend shared an essay on Tolkien by the Substack writer Genny Harrison, and reading it immediately tripped my internal AI detector:
"Not because the claim was obviously absurd, but because it was obviously compressed."
"That question is not radical. It is basic historiography."
"Recognizing that possibility is not an attack on Tolkien. It is what historical reading looks like."
"Tolkien is more than an author. He is a formative experience."
"it is no longer a myth. It has become a shrine."
"Middle-earth is not a crime scene. It is an inheritance."
If you've spent any time reading AI-generated prose, it couldn't be any more obvious. And indeed, if you copy a chunk of it into Pangram, it scores it as 100% AI-generated.
I was curious if she admitted to AI use or not, so I searched Harrison's facebook profile to see if she mentioned it, and found a post where she decried the integration of Substack's AI detector:
The classifier is also only as reliable as the moment its labels were collected. Models change constantly. Humanizer tools exist for the sole purpose of defeating detection, and many of them work well enough to make the entire exercise resemble airport security for adjectives.
Think about who that leaves. The person running an essay mill through a laundering tool walks through clean. The writer most likely to be flagged is the one with a consistent voice, careful structure, and a habit of revising until the sentences behave.
That is the familiar genius of automated enforcement. The sophisticated learn how to evade it. The honest remain available for inspection.
Again, if you've read a lot of LLM-generated stuff, this sounds exactly like it. And indeed, well,
I leave you to guess what the Pangram result is. One might note with a careful read of Harrison's piece that she
never actually says, "I do not use AI to write my posts." She says she puts a lot of work into them, she frets about a lot of harms that AI detection might do in the abstract. I feel like this non-denial is pretty telling.
(One also notes that Harrison claims of Pangram's training data, "
Someone
wrote all of it. Somebody's novels, somebody's blog posts, somebody's
dissertation, somebody's newsletter about mushrooms, all of it harvested
and processed into a classifier that now stands at the door of my
essay, deciding whether I am real. The exact grievance the anti-machine
crowd has been shouting for three years, that these systems were built
out of human labor nobody agreed to donate, applies with full force to
the referee they just installed. They did not defeat the thing they
hate. They gave it a badge and a percentage sign." This in fact totally untrue. Pangram states its datasets of human writing are "commercially licensed," which they actually paid for it, they didn't just rip it off the Internet as so many LLM creators did.)
But... can we trust them? I mean, Genny Harrison very strongly implies she doesn't use AI, enough that I would feel free to call her purposefully dishonest, and yet her essays score very highly. Anyone on Reddit can say, "oh, AI detectors are so unreliable," but what proves that they didn't use an LLM? Hence why the Pangram tests I covered in my previous post used pre-2022 material as a control when calibrating Pangram.
My favorite commenter on that reddit posts is the one who says,
I've got an open offer for anyone reading this. If you can find a piece of pre-2022 writing, with an accompanying web.archive.org link proving it's pre-AI writing, that is at least 100 words in length, and which Pangram flags as AI-generated, I will donate $100 to the charity of your choice. Everyone says they're getting tons of false positives, so this should be pretty easy, but for some reason every time I've posted this offer before, nobody has been able to do it.
As far as I could tell, no one ever took them up on this offer.
Case closed?
Well, maybe not. In the course of researching this post, I came across another Substack post by the blogger Freddie deBoer (I read his work occasionally maybe a decade ago): "
I Wouldn't Say Pangram is Broken, But I Would Say That It's Brittle." DeBoer says a reader accused him of using LLM to write part of one post, and linked to a Pangram result; deBoer actually paid for a Pangram account so he could mess around with it a bit and was able to reproduce this result. I am familiar enough with deBoer to believe that he is telling the truth when he says he did not use AI.
His post is worth reading in full, but points out the issues Pangram has with 1) very short texts, and 2) texts that switch back and forth between AI- and LLM-generated writing. One of the cofounders of Pangram actually pops up in the comments to concede Pangram needs to be better about handling these things. So though it seems that while Pangram might be very good... it also has its vulnerabilities. I do think, going by deBoer's post, that the frequency with which Pangram rates even longer texts as either 0% or 100% is a bit suspect, never in between, and the fact that it doesn't better foreground its confidence probability is also something that could use some improvement.
No comments:
Post a Comment