Try one sample

Unrobot quality standard

A rewrite is better only when the writing, voice, and meaning all improve.

We do not judge humanization with one detector score. Unrobot candidates are tested against a fixed private benchmark that combines blind human preference, semantic preservation, protected details, reliability, speed, and cost.

See the current status
The rule

A detector-friendly rewrite that sounds worse or changes the writer's point is a failed rewrite.

What we measure

Six signals. One release decision.

No single metric stands in for human writing quality. Each candidate must earn trust across the complete scorecard.

01

Natural writing

Blind reviewers compare Unrobot with a strong general AI rewrite. They judge rhythm, sentence variety, specificity, clarity, and whether the prose feels assembled from familiar AI patterns.

02

Personal voice

When a genuine human reference is available, reviewers check whether the rewrite preserves the writer's visible formality, vocabulary, emphasis, punctuation, and rhetorical habits.

03

Meaning preserved

Every candidate is checked for changed claims, missing qualifications, altered causality, new facts, lost negation, and shifts in intent or tone.

04

Protected details

Meaning Lock verifies names, numbers, dates, citations, quotations, links, email addresses, and technical identifiers before a result can count as a pass.

05

Reliable delivery

Completion rate, fallback use, response time, and estimated cost are measured alongside writing quality. A great result is not useful if it is slow, inconsistent, or uneconomical.

06

Detector behavior

Detector classifications may be recorded as a dated secondary signal. They never override an awkward rewrite, a meaning failure, or a poor blind-review result.

Methodology

How a candidate earns promotion.

  1. 01

    Freeze the cases

    The same development, validation, and untouched holdout cases are used for every candidate. Changing the cases creates a new benchmark version.

  2. 02

    Create the comparison

    Unrobot and a capable general AI baseline rewrite the same drafts. Candidate labels are hidden before review.

  3. 03

    Review blind

    Reviewers choose the more natural and voice-faithful version without knowing which system produced it.

  4. 04

    Audit meaning

    Automated checks and human review look for omissions, inventions, qualification changes, and protected-detail failures.

  5. 05

    Measure operations

    The scorecard adds completion, latency, fallback usage, and cost so quality is evaluated under realistic conditions.

  6. 06

    Promote deliberately

    A candidate can be considered for release only after every required quality and evidence gate passes. Promotion is never automatic.

Current benchmark status

Method locked. Full scorecard in progress.

The permanent benchmark protocol was established on August 25, 2026. The private corpus has strong academic coverage and untouched holdouts. We are adding genuine professional, content, and everyday writing before publishing the first complete performance scorecard.

We will not present a partial academic-heavy result as proof that Unrobot performs equally well for every kind of writer.

Protocol
Locked
Corpus privacy
Private
Result status
In progress
Next publication
Balanced scorecard

What we will publish

Enough evidence to evaluate the claim. Nothing that compromises private writing.

Published reports may include the benchmark version, date, writing contexts, case counts, system versions tested, aggregate quality rates, detector versions, limitations, and the final disposition.

They will not include private user writing, licensed passages, case-level rewrites, internal prompts, judge rationales, or implementation details that reveal the rewriting system.

Use the method

A practical guide to judging any rewrite.

How to test an AI humanizer

Build a useful test case, compare outputs fairly, and apply a clear pass and fail rule.

Read the checklist →

How to preserve meaning

Check claims, exact details, citations, qualifications, causality, and intent line by line.

Protect the point →

What detector scores mean

Understand probabilities, false positives, changing results, and responsible interpretation.

Read the detector guide →