Natural writing
Blind reviewers compare Unrobot with a strong general AI rewrite. They judge rhythm, sentence variety, specificity, clarity, and whether the prose feels assembled from familiar AI patterns.
Unrobot quality standard
We do not judge humanization with one detector score. Unrobot candidates are tested against a fixed private benchmark that combines blind human preference, semantic preservation, protected details, reliability, speed, and cost.
See the current statusA detector-friendly rewrite that sounds worse or changes the writer's point is a failed rewrite.
What we measure
No single metric stands in for human writing quality. Each candidate must earn trust across the complete scorecard.
Blind reviewers compare Unrobot with a strong general AI rewrite. They judge rhythm, sentence variety, specificity, clarity, and whether the prose feels assembled from familiar AI patterns.
When a genuine human reference is available, reviewers check whether the rewrite preserves the writer's visible formality, vocabulary, emphasis, punctuation, and rhetorical habits.
Every candidate is checked for changed claims, missing qualifications, altered causality, new facts, lost negation, and shifts in intent or tone.
Meaning Lock verifies names, numbers, dates, citations, quotations, links, email addresses, and technical identifiers before a result can count as a pass.
Completion rate, fallback use, response time, and estimated cost are measured alongside writing quality. A great result is not useful if it is slow, inconsistent, or uneconomical.
Detector classifications may be recorded as a dated secondary signal. They never override an awkward rewrite, a meaning failure, or a poor blind-review result.
Methodology
The same development, validation, and untouched holdout cases are used for every candidate. Changing the cases creates a new benchmark version.
Unrobot and a capable general AI baseline rewrite the same drafts. Candidate labels are hidden before review.
Reviewers choose the more natural and voice-faithful version without knowing which system produced it.
Automated checks and human review look for omissions, inventions, qualification changes, and protected-detail failures.
The scorecard adds completion, latency, fallback usage, and cost so quality is evaluated under realistic conditions.
A candidate can be considered for release only after every required quality and evidence gate passes. Promotion is never automatic.
Current benchmark status
The permanent benchmark protocol was established on August 25, 2026. The private corpus has strong academic coverage and untouched holdouts. We are adding genuine professional, content, and everyday writing before publishing the first complete performance scorecard.
We will not present a partial academic-heavy result as proof that Unrobot performs equally well for every kind of writer.
What we will publish
Published reports may include the benchmark version, date, writing contexts, case counts, system versions tested, aggregate quality rates, detector versions, limitations, and the final disposition.
They will not include private user writing, licensed passages, case-level rewrites, internal prompts, judge rationales, or implementation details that reveal the rewriting system.
Build a useful test case, compare outputs fairly, and apply a clear pass and fail rule.
Read the checklist →Check claims, exact details, citations, qualifications, causality, and intent line by line.
Protect the point →Understand probabilities, false positives, changing results, and responsible interpretation.
Read the detector guide →