Evaluation methodology

AI Detector Evaluation Framework

A transparent framework for reporting performance by modality, dataset, transformation, model version, and error type.

AUROC

Score discrimination

FPR

False positive rate

FNR

False negative rate

p50/p95

Latency distribution

Methodology

The minimum information required for a useful detector benchmark

Datasets

  • Named public datasets
  • Representative held-out samples
  • Compression and editing variants
  • Separate test sets by modality

Evaluation

  • Model version recorded
  • Thresholds fixed before scoring
  • Repeatable test runs
  • Ambiguous samples retained

Reporting

  • AUROC, precision, and recall
  • False positives and negatives
  • Latency distribution
  • Known failure cases

Real-World Generalization

Compression Robustness

Tests should include recompression and conversion patterns common in social media and messaging apps.

Novel Generator Handling

New generator families and model versions should be added to held-out test sets before broad performance claims.

Adversarial Resilience

Evaluation should include common evasion, editing, paraphrasing, noise, and obfuscation transformations.