In September 2025, the University of Waterloo published a short administrative notice turning off a feature. Not a product launch — a shutdown. Turnitin's AI-writing-detection overlay was being discontinued for every instructor on campus, with immediate effect. The reason given was blunt: internal testing found the tool's benefit "inconclusive," and more than once it flagged human-written text as 100% AI-generated.
Curtin University did the same thing on January 1, 2026, disabling the feature across every campus and study period. Curtin's public statement cites reliability concerns and the tool's tendency to penalize students who write in a more formal or predictable register — non-native English speakers most of all.
Neither of these is a blog post arguing detectors are flawed. They're university IT and academic-integrity offices turning a purchased feature off, on the record, because it did not do what it was sold to do.
This is the same failure we've been writing about
We've written before about why AI detectors can't reliably answer the question they claim to answer, and about the Stanford research showing detectors falsely flag a majority of non-native-English essays because they key on vocabulary predictability, not authorship. What's new isn't the argument. It's that the people who bought these tools are now confirming it with their own data, in public notices, at their own cost.
That distinction matters. It's one thing for a vendor with a competing product to say detection doesn't work. It's another for the institution running the tool to say the same thing after testing it on its own students' writing.
We keep a running list of the ones we can verify: eight institutions so far, including Waterloo, Curtin, Washington State, Vanderbilt, Johns Hopkins, Yale, Berkeley and UCLA. Every row on that page cites the institution's own published notice, never a news round-up, an aggregator or a tracker. That makes it a floor rather than a headcount, and deliberately so: if the only evidence for a school is somebody else's list, it doesn't go on the page. Across the ones that do qualify the pattern is consistent: false positives, a documented bias against non-native English writers, and no reliable way to tell an honest revision process from a lucky guess.
The alternative isn't a better detector
Turning the feature off doesn't remove the underlying problem — instructors still need a way to tell whether a student did the work. What Curtin's and Waterloo's decisions confirm is that scoring the finished text was never going to answer that, no matter how the model improves. The information was never in the final draft.
It's in the process: whether a document grew over sessions or arrived in one paste, whether the sources cited were actually read, whether there's a revision history that looks like thinking. That's what Folio's authorship record captures while a document is being written, and what Folio Classroom hands an instructor alongside every submission — not a probability score to argue with, but evidence to read.
We're not claiming credit for a shift we didn't cause. Universities turned this off because it failed their own students, independent of anything we've built. But it's worth saying plainly: the industry's own evidence now backs the bet we made — that the honest question was never "does this look AI," and the answer was always going to have to come from the process, not the page.