There's a percentage on the screen. Ninety four percent AI generated, it says, in confident red text. It looks like a fact. It looks like something you could take to a dean or a client and win an argument with. Most of the time, it's closer to a coin flip dressed up in a lab coat, and I think that gap between how these tools present themselves and how they actually perform is the real story here.
A 2023 study ran fourteen different detection tools through their paces, including the ones everyone's heard of, Turnitin, GPTZero, the works. Every single one came in under 80 percent accuracy. Only five managed to clear 70. Sit with that. If a teacher hands a stack of forty essays to one of these tools, somewhere around ten of the verdicts it returns are simply wrong, and there's no way to tell from the output which ten.
The same research found something I find almost more troubling than the raw accuracy number. These tools lean toward calling ambiguous text human, not AI, when they're unsure. Which sounds like it should be reassuring, a built in benefit of the doubt. But it means the flagged results, the ones that actually get acted on, carry less real confidence behind them than the score implies. And the moment text gets edited or paraphrased, even slightly, accuracy drops further still.
"If a teacher hands forty essays to an automated detector, around ten verdicts are wrong, and there is no indicator showing which ten."
False Positives Are Not Edge Cases
False positives aren't some edge case you can wave away either. A human writes something, entirely on their own, and a detector says a machine did. This has happened to professional, published writers. Janelle Shane, who's written an actual book, found chunks of her own earlier work flagged as AI generated by detection software, work she wrote years before tools like this even existed in their current form. If that can happen to someone with a publishing history and the means to push back, I don't want to think too hard about how many students and freelance writers have eaten a false accusation with nowhere to appeal.
Part of why this keeps happening comes down to what these tools are actually measuring. If your natural writing style is clean, a little formal, maybe repetitive in the way technical writing tends to be, you can land in the exact same statistical neighborhood as machine generated text without a language model ever entering the picture. Non native English speakers get hit by this constantly, and it's a genuinely unfair pattern, because a lot of them were taught to write in exactly the tidy, rule following style that now reads as suspicious to a machine built to flag tidiness.
The Moving Target Problem
There's also a moving target problem baked into all of it. Detectors get trained on datasets tied to specific models, and language models keep changing. A tool tuned to catch GPT-3.5's habits doesn't necessarily catch GPT-5's, while it keeps right on misjudging human writing that happens to resemble the old patterns it learned. The tools are always a little bit behind, and human writers pay the price for that lag.
None of this means detection is worthless or that nobody should ever think about it. It means the confidence people place in a single flagged score is wildly out of proportion to what these tools have actually earned. If you're a student, a writer, anyone whose work might end up run through one of these things, that's reason enough to care about how your writing actually reads on its own terms. Not to game a score. Just because a system with this kind of error rate shouldn't get the final word on whether you did your own work.
Discover how WeCatchAI leverages peer consensus and human judgment over automated algorithms.