Detector accuracy guide

How accurate are AI detectors?

An accuracy percentage is meaningful only for the test population, threshold and models used. It does not automatically tell you the probability that one person used AI.

In brief: Ask what was tested, how recent it was, how often human text was falsely flagged and whether the submitted text resembles the evaluation data. Never convert a detector score directly into proof.

“Accuracy” can hide the errors that matter

A single percentage can combine true positives and true negatives while concealing false accusations. Sensitivity describes how often tested AI text is detected; specificity describes how often tested human text is correctly left unflagged. The decision threshold changes both.

Base rates matter too. When most submissions are human-written, even a low false-positive rate can produce a substantial share of incorrect flags among all flagged work.

Check whether the test matches real use

Look for independent evaluation on unseen material, current generation models, relevant languages, realistic editing and the same kind of text you intend to assess. Results on long English essays may not transfer to short answers, translated prose, CVs or heavily edited work.

Models and writing practices change

Detectors learn patterns from particular datasets. New generators, prompts, paraphrasing and mixed human-AI workflows can shift performance. A published benchmark is a snapshot, not a permanent guarantee.

Read the output as a limited signal

Ask whether the number is a calibrated probability, a class score or simply a thresholded label. Check the tool’s minimum text length, supported language, version and stated exclusions. “80% AI” often does not mean there is an 80% probability AI wrote the passage.

Use a decision process that can be reviewed

For consequential decisions, combine the limited signal with factual and citation checks, drafts or version history, a conversation with the author and an appeal route. Record contrary evidence and do not treat multiple detectors trained on similar patterns as independent proof.

Research and further reading

Try the evidence-first approach

Check what the passage says and sources

Paste a short passage to check its factual support and public-source similarity separately from AI-associated language signals.

0 / 3,000No account required

The result does not prove authorship or misconduct.