Try one sample

Detector guide

AI detector scores: what they mean and what they cannot prove

An AI detector score is a model's estimate about patterns in a submitted passage. It is not a record of who wrote the text, and it should not be interpreted as one.

Unrobot EditorialPublished August 24, 20269 minute read
01

How an AI detector reaches a score

AI detectors analyze linguistic signals and compare them with patterns learned from human and machine-generated examples. Different products use different models, thresholds, and reporting formats.

A result such as '80% AI' may describe a confidence estimate, a proportion of flagged text, or a product-specific composite score. Read the tool's documentation before comparing numbers across services.

02

Why the same text receives different results

Two detectors can disagree because they were trained on different data or tuned for different tradeoffs. The same product can also change after a model update.

Text length, genre, formatting, and language background matter. A short formal paragraph offers fewer signals than a long, varied document. Technical and academic prose may repeat necessary terminology.

03

False positives and false negatives

A false positive occurs when human-written text is labeled as AI. A false negative occurs when AI-generated text is labeled as human. Every classification system has both.

The practical question is not whether errors exist, but how often they occur for the kind of text being evaluated and how the result will be used. High-stakes decisions require more evidence than a single automated score.

04

What detector scores are useful for

Detector feedback can identify unusually predictable passages. It can prompt a closer look at repetitive structures, generic phrasing, or abrupt changes in voice. It can also support controlled product testing when the same samples and detector versions are used consistently.

That makes detector scores valuable diagnostic data. It does not turn them into proof of misconduct or authorship.

05

How Unrobot uses detector feedback

Unrobot treats detector performance as one part of a larger evaluation. A candidate rewrite also has to preserve meaning, retain citations and numbers, sound natural to human readers, and avoid unnecessary loss of detail.

Any public benchmark should state the date, detector, sample set, pass criteria, and limitations. A claim without that context is difficult to verify and unlikely to remain true after tools change.

06

A responsible way to interpret a result

Review the full document, not just the score. Look for evidence in the writing process, source history, drafts, and the writer's ability to explain the work. If the score is being used in school or employment, provide a fair way to challenge it.

For writers, the durable response is to improve specificity, clarity, and voice while keeping accurate records of the writing process. Chasing one detector number can make the prose worse without resolving the underlying question.