Marin T. Kael
DE / EN

Working Paper · No. 02 · v0.3 (Per-LLM Headline)

Pre-Launch LLM Citation

The top-3 language models reach a 19.8 – 24.0% hit rate for a pre-launch author without any published work.

Marin T. Kael · · 10 minute read

Abstract

A pre-launch author identity — no books published, no reviews, no press, 13 Bluesky followers — reaches measurable citation rates in individual language models within 7 days. OpenAI Search Preview leads with 24.0 %, followed by Claude Haiku 4.5 and Sonnet 4.6 at 19.8 % each. Pure identity engineering (Wikidata, ORCID, Bluesky, GitHub, Zenodo, llms.txt, Reddit) is the sole input. Pre-registered via Zenodo DOI before T+0. The per-LLM headline is methodologically robust against the provider-availability confound that lets the aggregate mean swing between 14.8 and 21.7 %.

1

Background

GEO/AEO literature typically works on the assumption that LLM citation hangs on established cultural presence — published works, reviews, press, social proof. This study tests the opposite: can pure structured identity engineering produce measurable LLM visibility before a single word is published?

2

Pre-Registered Design

  • Subject: Marin T. Kael (pseudonymous fantasy author, debut 22 September 2026)
  • Inputs (T+0): Wikidata person + book items · ORCID with biography · Bluesky · GitHub · Zenodo with DOIs · llms.txt · Reddit profile with karma build-up
  • Probe: 11 LLMs × 16 questions × daily polling
  • Scoring: 0 = not_found · 0.5 = name_only · 2 = partial_book · 3 = full_citation · −3 = US-female-misidentification penalty
  • Pre-registration: DOI 10.5281/zenodo.20125967

3

Results (T+7)

3.1 Per-LLM headline (primary metric)

LLMHit raten_legitNote
OpenAI Search Preview (gpt-4o-mini-search-preview-2025-03-11)24.0 %16web-search-backed
Claude Haiku 4.519.8 %16conservative-positive
Claude Sonnet 4.619.8 %16identical to Haiku
Llama 3.2 3B17.7 %16smallest model in the family
Llama 3.1 8B17.7 %16
OpenAI gpt-4o-mini-2024-07-1817.7 %16without web search
Claude Opus 4.715.6 %16highest epistemic conservativeness
Mistral 7B11.5 %16
Gemini 2.5 Flash11.5 %16via Direct Batch API (from v2.8)
Phi-24.2 %16hallucination penalty active
Llama 3 8B4.2 %16weakest Llama model

3.2 Aggregate (secondary)

The aggregate mean across all measurable LLMs swings daily between 14.8 and 21.7 %, depending on which LLMs were measurable on a given day. This volatility reflects provider availability, not genuine Marin discoverability. Per-LLM values are therefore the methodologically more robust anchor.

4

Three counter-intuitive findings

4.1 Web-search LLMs dominate base models by ~36%

OpenAI Search Preview (24.0 %) vs. OpenAI gpt-4o-mini base (17.7 %) — same model family, the only difference: web-search augmentation at inference time. The Δ of +6.3 pp (≈ 36 % relative increase) suggests: for new entities, web-search capability at inference time matters more than the training cutoff.

4.2 Model size is not predictive

Llama 3.2 3B (17.7 %) ≈ Llama 3.1 8B (17.7 %) — the smaller model recognises the identity just as well. The citation hit rate apparently depends on training-data inclusion patterns, not on model capacity.

4.3 Higher Anthropic tier ≠ higher citation rate

Claude Opus 4.7 (15.6 %) < Sonnet 4.6 (19.8 %) = Haiku 4.5 (19.8 %). Opus is the most expensive and nominally strongest tier — and delivers the lowest citation rate. Hypothesis: Opus is calibrated to be epistemically more conservative — it prefers "I don't know" over partial citation under uncertainty. Methodologically this is the desirable property, but it scores lower in citation-rate metrics. Citation-rate scoring schemes thereby indirectly reward overconfident hallucination over honest uncertainty — a relevant finding for future AEO tool methodologies.

5

Limitations

  • n=1 single-subject design — not generalisable as a population estimate.
  • 7-day window — Phase 1 (instrument validation), not effect detection.
  • 16 questions manually authored — selection-bias risk.
  • Researcher = subject — fully disclosed in Methodology Note 01 § 7.
  • Aggregate is provider-availability-sensitive — hence the per-LLM headline as primary metric (see also Working Paper 04 Mode 5).

6

Replication

7

How to cite

Kael, M. T. (2026). Pre-Launch LLM Citation: Top-3 LLMs Hit 19.8 – 24.0 % for a Pre-Launch Author Without Published Work. Working Paper 02 v0.3, Marin T. Kael — KI-Zitations-Feldlabor. Status: Outline. URL: marin-t-kael.de/research/working-papers/wp-02-llm-citation.