PDF DOI: 10.5281/zenodo.22087724 Code & Data Data DOI Preregistration ResearchGate
We ask a narrow question that the self-recognition literature has not isolated: when two capability tiers of the same model family (Claude 4 — Opus 4.8, Sonnet 4.6, Haiku 4.5) each write open-ended prose, can a tier behaviorally recognize its own writing among a sibling's, beyond what a neutral third-party judge of equal or lower compute achieves on the identical texts? Across forced-choice experiments with two independent confound controls, framing variation, a disagreement-conditioned reanalysis, and a raw-API replication, we report a four-part result. (1) Tier style is externally legible, and the legibility is capability-graded: in the agent harness, neutral judges discriminate authorship at Opus 0.91 ≈ Sonnet 0.90 and Haiku 0.63 (chance 0.50); off-harness, distant pairs remain discriminable while the closest pair's legibility falls to chance in a small replication. Yet being the author confers no advantage — under both controls, the self-minus-neutral advantage is negative or brackets zero for every tier. (2) The models run inside an agent harness that amplifies tier-stylistic distinctiveness in how the models write, not in how they judge. (3) Attribution runs on stereotype: a stronger model's length-matched caricature of a weaker tier is judged more authentic than the genuine article, even by the author. (4) Stronger tiers better know when they are right: identity-judgment calibration tracks capability, with the weakest tier mildly anti-calibrated. Together these place a behavioral boundary on emergent-introspection claims: privileged access, even where it is real for this family, does not extend to recognizing one's own prose authorship among tiers. Our study is deliberately narrow: a single behavioral axis (no token logprobs) in a single model family and generation.
The study was preregistered on the Open Science Framework (osf.io/brdt8, osf.io/5azq8). All judgment records, generation and judging scripts, and analysis code are public at github.com/dbookstaber/claude-tier-self-recognition and archived at doi.org/10.5281/zenodo.20724958.