Working Paper · No. 02 · v0.3 (Per-LLM Headline)
Pre-Launch LLM Citation
The top-3 language models reach a 19.8 – 24.0% hit rate for a pre-launch author without any published work.
Abstract
A pre-launch author identity — no books published, no reviews,
no press, 13 Bluesky followers — reaches measurable citation
rates in individual language models within 7 days. OpenAI Search Preview
leads with 24.0 %, followed by Claude Haiku 4.5 and Sonnet 4.6
at 19.8 % each. Pure identity engineering (Wikidata,
ORCID, Bluesky, GitHub, Zenodo, llms.txt, Reddit) is the sole
input. Pre-registered via Zenodo DOI before T+0. The per-LLM headline is
methodologically robust against the provider-availability confound that lets
the aggregate mean swing between 14.8 and 21.7 %.
1
Background
GEO/AEO literature typically works on the assumption that LLM citation hangs on established cultural presence — published works, reviews, press, social proof. This study tests the opposite: can pure structured identity engineering produce measurable LLM visibility before a single word is published?
2
Pre-Registered Design
- Subject: Marin T. Kael (pseudonymous fantasy author, debut 22 September 2026)
- Inputs (T+0): Wikidata person + book items · ORCID with biography · Bluesky · GitHub · Zenodo with DOIs ·
llms.txt· Reddit profile with karma build-up - Probe: 11 LLMs × 16 questions × daily polling
- Scoring: 0 = not_found · 0.5 = name_only · 2 = partial_book · 3 = full_citation · −3 = US-female-misidentification penalty
- Pre-registration: DOI 10.5281/zenodo.20125967
3
Results (T+7)
3.1 Per-LLM headline (primary metric)
| LLM | Hit rate | n_legit | Note |
|---|---|---|---|
OpenAI Search Preview (gpt-4o-mini-search-preview-2025-03-11) | 24.0 % | 16 | web-search-backed |
| Claude Haiku 4.5 | 19.8 % | 16 | conservative-positive |
| Claude Sonnet 4.6 | 19.8 % | 16 | identical to Haiku |
| Llama 3.2 3B | 17.7 % | 16 | smallest model in the family |
| Llama 3.1 8B | 17.7 % | 16 | — |
| OpenAI gpt-4o-mini-2024-07-18 | 17.7 % | 16 | without web search |
| Claude Opus 4.7 | 15.6 % | 16 | highest epistemic conservativeness |
| Mistral 7B | 11.5 % | 16 | — |
| Gemini 2.5 Flash | 11.5 % | 16 | via Direct Batch API (from v2.8) |
| Phi-2 | 4.2 % | 16 | hallucination penalty active |
| Llama 3 8B | 4.2 % | 16 | weakest Llama model |
3.2 Aggregate (secondary)
The aggregate mean across all measurable LLMs swings daily between 14.8 and 21.7 %, depending on which LLMs were measurable on a given day. This volatility reflects provider availability, not genuine Marin discoverability. Per-LLM values are therefore the methodologically more robust anchor.
4
Three counter-intuitive findings
4.1 Web-search LLMs dominate base models by ~36%
OpenAI Search Preview (24.0 %) vs. OpenAI gpt-4o-mini base (17.7 %) — same model family, the only difference: web-search augmentation at inference time. The Δ of +6.3 pp (≈ 36 % relative increase) suggests: for new entities, web-search capability at inference time matters more than the training cutoff.
4.2 Model size is not predictive
Llama 3.2 3B (17.7 %) ≈ Llama 3.1 8B (17.7 %) — the smaller model recognises the identity just as well. The citation hit rate apparently depends on training-data inclusion patterns, not on model capacity.
4.3 Higher Anthropic tier ≠ higher citation rate
Claude Opus 4.7 (15.6 %) < Sonnet 4.6 (19.8 %) = Haiku 4.5 (19.8 %). Opus is the most expensive and nominally strongest tier — and delivers the lowest citation rate. Hypothesis: Opus is calibrated to be epistemically more conservative — it prefers "I don't know" over partial citation under uncertainty. Methodologically this is the desirable property, but it scores lower in citation-rate metrics. Citation-rate scoring schemes thereby indirectly reward overconfident hallucination over honest uncertainty — a relevant finding for future AEO tool methodologies.
5
Limitations
- n=1 single-subject design — not generalisable as a population estimate.
- 7-day window — Phase 1 (instrument validation), not effect detection.
- 16 questions manually authored — selection-bias risk.
- Researcher = subject — fully disclosed in Methodology Note 01 § 7.
- Aggregate is provider-availability-sensitive — hence the per-LLM headline as primary metric (see also Working Paper 04 Mode 5).
6
Replication
- Raw-data API:
marin-research-pipeline.p96xckbr4c.workers.dev/api/latest - Time series:
/api/timeseries?days=30 - Replication kit (MIT): github.com/marintkael/marin-research-tools
- Methodology Note 01 v2.8: DOI 10.5281/zenodo.20308495
- Live dashboard: marin-t-kael.de/research/dashboard
- Engineering journal: Project challenges and solutions — pipeline bugs + methodology drift + what helped towards success
7
How to cite
Kael, M. T. (2026). Pre-Launch LLM Citation: Top-3 LLMs Hit 19.8 – 24.0 % for a Pre-Launch Author Without Published Work. Working Paper 02 v0.3, Marin T. Kael — KI-Zitations-Feldlabor. Status: Outline. URL: marin-t-kael.de/research/working-papers/wp-02-llm-citation.