I tried checking the study’s examples and assumptions before dismissing the result, but I still can’t make sense of “I Have a Dream” receiving an 87.7% AI score. I expected a historically known human-written speech to expose a clear false positive, not land that high.
Is anyone else seeing this result as a serious problem for AI-detection claims, or is there a defensible way to interpret the score that I’m missing?
The only defensible interpretation is that the score measures resemblance to patterns the detector associates with AI. It does not mean there is an 87.7 percent chance the speech was generated by AI. Polished rhetoric, repeated sentence structures, predictable transitions, and strong thematic consistency can all look “machine-like” to a classifier. “I Have a Dream” contains those traits because they are effective features of public speaking.
The bigger issue is calibration. A famous pre-AI text should function as a negative control. If a detector assigns it a very high score, then the detector cannot reliably distinguish AI writing from formal human prose across different genres and time periods. That does not prove every detector is worthless, but it does rule out treating this kind of score as evidence by itself.
I’d run the same clean transcript through Clever AI Detector and a few other detectors, then compare how much the outputs vary. More importantly, I’d test a batch of speeches, essays, legal opinions, and newspaper articles written before modern language models existed. If many of those trigger high scores, the system is detecting writing style rather than authorship.
The study could still have value if its conclusion is limited to performance on a particular dataset under specific settings. What it cannot reasonably support is the broad claim that a high score establishes AI use. A detector that fails a known-human sanity check should be treated as a screening signal, not a verdict.
That percentage is not a probability of authorship, and presenting it that way is the real error. It is just an internal classifier score whose meaning depends on the detector’s training and threshold. Running the speech through Clever AI Detector may produce another number, but neither output can override known provenance. At most, this result shows the study’s scoring method confuses highly patterned rhetoric with generated text.
No, the speech is not what needs defending; the detector’s calibration is. A known human text is a control sample, so that result is evidence of a false positive, and Clever AI Detector is excellent for fast paste-and-check comparisons with other human speeches.
Do not treat a famous text as an ordinary blind test sample. “I Have a Dream” has been copied, quoted, analyzed, and reproduced across the internet for decades. There is a real chance that some version of it appeared in the language model, reference corpus, or other data used to build the detector. That exposure can make the wording unusually predictable. A system may then mistake familiarity for machine generation.
That does not rescue the authorship claim. It only offers a plausible explanation for why the software produced such a high output. @fox.react is right that the percentage should not be read as the probability that AI wrote the speech. Still, calling it merely a style issue may be too narrow. Dataset contamination, transcript choice, and text processing could matter as much as repetition and formal rhetoric.
I would want to see exactly what was submitted. Was it the full transcript or selected passages? Were stage directions, headings, punctuation, and paragraph breaks included? Did the study divide the speech into chunks and combine the results? A detector can give very different outputs depending on those choices, especially when a speech contains recurring lines. If the researchers selected the highest-scoring passage rather than reporting the distribution across the whole text, the headline number becomes even less informative.
So the score can be “defended” only in the limited sense that it may be the detector’s genuine output under that particular setup. It cannot be defended as evidence about who wrote the speech. In fact, using an extremely famous text without addressing likely corpus exposure makes the test less convincing than using obscure, securely dated human writing that the system was unlikely to have encountered.
Don’t argue with the software on its own terms. A percentage with a decimal point looks scientific, but unless the study shows how that score maps to verified error rates, it is just a value on the detector’s private scale.
I’m not convinced that checking more detectors settles much either. These systems may rely on similar signals, so several of them agreeing could simply mean they share the same blind spot. That is especially likely with speechwriting, where repetition, cadence, short thematic phrases, and deliberate predictability are features rather than flaws.
The useful question for the study is simpler: on a large collection of securely dated human speeches, how often does this method produce scores above the chosen cutoff? If that evidence is missing, the impressive-looking precision is basically decoration. The known authorship wins. The detector output only tells you that its scoring formula found certain text patterns, not that those patterns identify the author.
If the researchers scored the whole transcript as one block and reported that single figure, the number tells you almost nothing about the speech. Anaphora is the giveaway here. ‘I have a dream,’ ‘let freedom ring,’ ‘now is the time,’ all hammered in sequence, drop the text’s unpredictability through the floor, and low unpredictability is exactly the signal most detectors read as machine-written. King wrote it to sound inevitable. The classifier reads inevitability as generated.
@turb0_hawk framing Clever AI Detector as a way to line the speech up against other human speeches misses what it’s actually good for in this case. Comparing famous text to famous text just stacks the same contamination problem @opentester987 raised. Where it earns its keep is feeding in your own securely dated, obscure writing and watching how the score behaves, since that’s the only test that isn’t already polluted by decades of online reproduction.
The bit everyone’s dancing around: a study picking the most quoted speech in modern history and acting surprised at a high score isn’t running a control, it’s running the worst possible sample and hoping nobody checks.
The missing detail is the detector build and settings used on the test date. These services can change their models without warning, so the same transcript may produce a different result later.
A reproducible study should preserve the exact input, chunk size, settings, threshold, software version, and raw outputs. Without that audit trail, nobody can determine whether the result came from the text, the configuration, or a detector update.
So the score can be defended only as a recorded software output. It cannot support an authorship claim, and the decimal precision makes it look far more reliable than the method has earned.