Impronta reads the traces that human writers leave without knowing it — the clustering of uncertainty, the rhythm of thought, the marks of a mind at work. Not what AI sounds like. What cognition leaves behind.
Sign in to run an analysis — free for educators, API key not required.
Diagnostic indicators only · Not for use as sole evidence in any academic integrity proceeding · Your text is never stored
In the 1880s, Morelli discovered that art forgers faithfully reproduced the grand compositional elements — the Madonna's posture, the colour palette, the architecture. What they could not reproduce were the small, unconscious details: the fold of an earlobe, the curve of a knuckle. Those details emerged from the hand itself, not from intention. Forgers cannot fake what was never consciously made.
AI language models are skilled forgers. They reproduce the grand rhetorical structure of human essays. What they cannot reproduce are the cognitive process traces: hedges clustering around genuine uncertainty, sentence rhythm driven by cognitive load, the marks of a mind revising itself mid-sentence. These traces emerge from the writing process itself — and autoregressive token prediction has no writing process.
"The details the hand produces without knowing it
are the only details that cannot be forged."
"The features that reveal cognitive origin
correspond to the traces the hand leaves without knowing."
The core detection runs on Impronta's own server — no text is transmitted to external AI services until you request the rhetorical analysis layer.
Every human value in Fig. III is the mean from a corpus of genuine academic essays. Every AI figure is from the same texts rewritten by large language models with explicit humanization instructions — the hardest detection case.
The gap between the bars persists across every model family tested — Mistral, LLaMA, Qwen, Gemma, Yi-Large — because these are cognitive process absences, not distributional artifacts of any particular training corpus.
You cannot attack all twenty-six traces simultaneously with a single prompt.
When you submit a document to Impronta, what happens to it?
Text enters our server over HTTPS. It is processed in RAM only — never written to disk,
never logged, never passed to a database. Twenty-six features are extracted.
A score is computed and returned to you. The text ceases to exist in our systems
the moment the response is transmitted. Not eventually. Immediately.
There is no deletion schedule because there is nothing to delete.
"Impronta's zero-retention design is not a policy choice. It is an architectural consequence. You cannot delete data that was never stored. Submitted text is never logged, never backed up, never retained in any form. No Impronta employee can access it — because it is not there."
| Platform | Text stored? | Default retention | Trains on submissions? | Student IP? | Methodology published? |
|---|---|---|---|---|---|
| Turnitin | Yes | Indefinitely (unless inst. requests deletion) | Not publicly denied | Non-revocable license granted | No |
| GPTZero | Unverified claim | Session (unverified) | Not stated | Not addressed | No |
| Originality.ai | Yes — by default | Stored in account history | Not addressed | Not addressed | No |
| Copyleaks | Enterprise-configurable | Not published | Not addressed | Not addressed | No |
| Winston AI | Not published | Not published | Not addressed | Not addressed | No |
| Impronta | Never — architectural | Zero — nothing persisted | Never — architectural | No license acquired | Yes — peer reviewed |
Sources: Turnitin Privacy Policy (Feb 2026) · GPTZero privacy policy · Originality.ai terms · Competitor documentation reviewed May 2026. "Unverified" indicates a privacy claim was made but is not independently auditable.
Process archaeology targets the hardest case: AI text explicitly humanized to evade detection. On this task, no published system comes close.
Where text falls on the authorship continuum — from verified human to clean AI generation
The proportion of explicitly humanized AI texts correctly flagged when the detection threshold is set such that no more than 5% of genuine human texts are falsely flagged. This is the operationally correct metric for academic integrity. No other published detector reports this metric.
| System | HumAI@5%FPR | AUC (academic) | Cross-model (5 families) | Method published? | Notes |
|---|---|---|---|---|---|
| Impronta v5 | 79.5% | 0.9445 | AUC 0.925 mean | Yes — peer reviewed | +52.8pp over RADAR |
| RADAR (NeurIPS 2023) | 26.7% | 0.789 | AUC 0.774 mean | Yes (adversarial training) | Designed for adversarial — still fails |
| roberta-mixed-detector | 38.5% | 0.937 | In-distribution only | No | Excellent on clean AI; fails humanized |
| chatgpt-detector | 28.7% | 0.779 | Poor cross-model | No | Near-chance on non-ChatGPT models |
| Binoculars-Light | 24.1% | 0.605 | Requires 28GB GPU | Partial | CPU approximation only |
Source: corpus_v3 test set (n=784) · Primary benchmark · All systems evaluated on identical test set with identical threshold protocol. Cross-model test: gsingh benchmark, format-normalized, DeLong p<.001 on 4/5 comparisons. Limitations disclosed in full methodology.
Humanized AI recall at 5% false positive rate on adversarially humanized academic essays — surpassing RADAR by 52.8pp (z=8.56, p<.001)
Mean AUC advantage over RADAR across 5 completely independent model families (Mistral, LLaMA, Qwen, Gemma, Yi-Large) — zero training overlap, p<.001 on 4/5
Mean AUC on 5 independent model families tested on journalistic prose (domain shift confirmed) — process archaeology generalizes across models and domains
Detection maintained against targeted adversarial hedge-clustering prompt attack. The attack paradoxically increased detectability by over-regularizing hedge distribution.
False positive rate on Chinese L1 English exam essays — below the native English rate of 15.4% at the same threshold. ESL risk is lower than predicted by SLA literature for this population.
Impronta is the only AI detection system that has published false positive rates by population group. The documented 3.8% ESL FPR (Chinese L1) contradicts the 15–30% ESL false positive rates reported across competing tools. Important caveat: this result is based on high-school exam essays; university-level L1 validation is ongoing.
Requires ≥250 words. Paragraph-level features collapse on shorter submissions, inverting the decision boundary. Use CNI-Short mode (AUC 0.655 on RAID benchmark) for shorter texts.
Every score is a probability estimate with a documented ±0.08 confidence interval. Impronta is one signal among several in any human review process. Never sole evidence.
Impronta is designed for adversarially humanized academic prose. On journalism-domain clean AI text, distributional detectors (roberta-mixed AUC 0.991) outperform Impronta (0.925). Different tools for different tasks.
Calibrated on academic essays, tested on journalism. Legal text, technical documentation, and creative non-fiction have not been formally validated. Domain-appropriate confidence adjustments shown when genre is declared.
The primary benchmark was constructed using humanization strategies designed by CNI's authors. Mitigated by three independent external evaluations but disclosed honestly. See full methodology for details.
Each layer targets a different cognitive operation of the writing process — not what AI text looks like, but what human cognition produces.
Every distributional AI detector — systems trained to recognize what GPT-4 or Claude "sounds like" — must be retrained each time a new model is released. Their accuracy degrades as models improve and as humanization techniques evolve.
Process archaeology works differently. It measures the absence of cognitive production traces — traces that are present in human writing by definition and absent from AI generation regardless of the model, the scale, or the explicit humanization instruction. No language model generates text through the cognitive operations that produce these traces. That is not a training data limitation. It is a mechanistic difference.
The empirical confirmation: Impronta v5 achieves mean AUC 0.925 on five model families with zero training overlap, tested in a different domain (journalism) than its training domain (academic essays). The process archaeological features generalize because they measure something architecturally absent from all current language models.
University instructors, professors, and teaching assistants with an institutional email address receive full Impronta access at no cost for 12 months — a complete academic year cycle. Academic integrity requires tools that are transparent, evidence-based, and held to scientific standards. Annual reconfirmation keeps the relationship current.