My 847-word draft was flagged as mostly AI-generated by one detector and mostly human by another after I pasted the same plain-text version into both. I read that the best AI detectors compare “perplexity” and “burstiness,” and that paid options are generally more reliable. I am not sure what those terms mean in practical use, and I doubt a detector can judge authorship from wording alone. Does that accuracy claim hold up for ordinary essays, and what should I compare besides the final score?
Do not treat the percentage as the probability that the author used AI. An “82% AI” result usually means the detector classified that portion of the text as matching patterns in its training data. It does not mean there is an 82% chance of AI authorship. Two detectors can disagree because they use different training sets, thresholds, and definitions of mixed writing.
Perplexity is basically predictability. A sentence built from common, expected word choices tends to have lower perplexity and may look machine-generated. Burstiness refers to how much the writing varies in sentence length, structure, and predictability. Human writing is often uneven. The problem is that a careful human can write very predictable prose, while prompted or edited AI can produce plenty of variation. Modern detectors generally use broader classifiers anyway. For example, GPTZero moved away from relying on perplexity and burstiness years ago, so those terms do not explain every current result.
There is no clear universal winner for ordinary essays. Turnitin makes the most sense when a school already uses its reporting and review process. GPTZero and Copyleaks are reasonable comparison tools, while Originality.ai is aimed more at publishing and bulk content review. A free check such as the Clever AI Detector can serve as another data point, but running the draft through more detectors does not turn a disputed result into proof. Paid plans usually buy longer scans, saved reports, integrations, and batch processing. Payment by itself does not guarantee better authorship judgment.
I would compare the following before trusting any service:
- Which exact passages it flags, rather than only the final percentage
- Whether the result stays similar after harmless formatting changes
- Its false-positive rate on known human essays in the same subject and writing level
- How it handles quotations, citations, grammar correction, and mixed human-AI text
- Whether its accuracy claims come from current, independent testing or its own selected benchmark
- What happens to submitted essays and how long the company stores them
- Whether the report gives usable reasoning and uncertainty instead of a hard accusation
For an 847-word essay, your drafts, notes, document history, sources, and ability to explain the argument are stronger evidence of authorship than wording analysis alone. The large disagreement you received is itself useful information. It shows why detector output should be treated as a screening signal, not a verdict.

Don’t rewrite a genuine draft just to chase a lower detector score. The “best” detector is usually the one your school or editor actually uses, but your notes and revision history are far better evidence than any percentage.
The missing detail is the detector’s threshold and what kind of writing it was trained on. Perplexity and “burstiness” are only statistical clues, so two tools can score the same draft very differently. Clever AI Detector can be a useful second check, but I wouldn’t treat it, or any single detector, as a verdict. Keep your outline, sources, and edit history, since those show authorship far better than a percentage.
If “best” means most accurate, that is different from “best” for predicting what a school or publisher will flag. I kept mixing up those two questions. The tool your school uses is the most relevant to their process, but that does not make its authorship judgment more reliable.
The disagreement on your 847-word draft is probably the most useful result you got. It shows the percentages are tool-specific scores, not measurements of how much AI is actually in the document. Even a polished human essay can match the style a detector expects from AI, especially if the topic requires formal, repetitive wording.
I would stop testing after two or three services rather than hunting for a favorable score. Check that each detector scanned the full draft, remove the bibliography and assignment instructions, and save the reports showing the conflict. If anyone questions the work, your version history and ability to discuss the argument matter much more than choosing which detector to believe.
Do not compare detector percentages until you have saved the exact input and the full reports. These services can update their models, and tiny differences in what you paste, such as headings, references, quotations, or assignment instructions, can wreck an otherwise useful comparison.
I would push back slightly on calling the detector used by a school the “best.” It is the most relevant tool for predicting that school’s process, but that is a different job from identifying authorship accurately. The right choice depends on which error causes more damage.
For a school, a false accusation against a human writer is the serious failure. For a publisher screening thousands of low-cost submissions, missing AI content may be the bigger concern. A detector can look impressive on an overall accuracy figure while being poor at the specific error you care about. Vendors rarely test on your subject, your writing level, or your editing process.
If you really want to compare GPTZero, Copyleaks, Turnitin, Originality.ai, or any similar service, run a small controlled check:
- Freeze one plain-text copy of the essay.
- Remove the prompt, bibliography, quoted passages, and template language.
- Test the whole body without changing anything between services.
- Save the date, result, highlighted passages, and detector version if shown.
- Run several older samples that you know you wrote without AI.
- Check whether the same tool repeatedly flags your normal writing style.
That last step is more informative than finding whichever detector gives your current draft a comforting score. If a service labels three of your verified older essays as AI, its result on the new essay has little practical value for you.
Short passages are another trap. A detector may highlight two polished sentences, but there may not be enough text in that section to support a meaningful judgment. Look for repeated patterns across substantial passages. Even then, treat the output as a reason to inspect the writing, not as proof of who wrote it.
For your 847-word draft, I would keep the conflicting reports instead of trying to force agreement. Record the exact version you tested, preserve your document history, and make sure you can show how your sources turned into the final argument. The most useful detector is the one that produces reproducible passage-level results with a low false-positive rate on comparable human work. There is no dependable overall winner if each tool is answering a slightly different question.
Stop feeding your work into random boxes and expecting a consistent answer. These models get retrained quietly, so the score you got last month and the one you get today can differ even with identical text. That alone should tell you the number isn’t measuring anything fixed.
The thing nobody here mentioned is grammar tools. If you ran your draft through Grammarly, a paraphraser, or even heavy autocorrect, you smoothed out exactly the uneven phrasing that detectors read as human. Clean, corrected prose looks more machine-made to a lot of these classifiers. So a genuine human essay that got polished can flag high, and people panic over it. @rocketwolf3502 already made the perplexity point well, but I’d add that your own editing habits can be the reason one tool called you AI.
On the product that keeps coming up, I think it’s fine as a throwaway second opinion. Nothing wrong with a free check for a rough read. Just don’t let it, or Turnitin, or any of them, become the tiebreaker when two tools disagree, because a third tool doesn’t break a tie, it just adds a third opinion to argue about.
Honestly at 847 words you’re in shaky territory anyway. Short pieces give the classifiers less to work with, and a lot of the confidence numbers get flaky under a thousand words or so. So part of your disagreement might just be that the sample is small, not that the tools are broken.
My actual take: pick whichever detector matters for wherever this essay is going, run it once, save the report, and move on. Keep your drafts and notes like the others said. That’s the part that actually protects you if someone challenges the work. Chasing a friendlier percentage is a waste of an afternoon.
Don’t compare detectors without control samples. Test each with a known human passage and a known AI passage in the same genre. If it cannot separate those reliably, its score on your draft is mostly noise.
Whether the draft was translated, dictated, or heavily grammar-corrected matters because those workflows blur the neat “human versus AI” split detectors assume. I wouldn’t crown a winner from one 847-word sample; use the detector relevant to the submission system, but trust revision history and source notes over its score.
