Datasets
- Named public datasets
- Representative held-out samples
- Compression and editing variants
- Separate test sets by modality
Evaluation methodology
A transparent framework for reporting performance by modality, dataset, transformation, model version, and error type.
AUROC
Score discrimination
FPR
False positive rate
FNR
False negative rate
p50/p95
Latency distribution
The minimum information required for a useful detector benchmark
Tests should include recompression and conversion patterns common in social media and messaging apps.
New generator families and model versions should be added to held-out test sets before broad performance claims.
Evaluation should include common evasion, editing, paraphrasing, noise, and obfuscation transformations.