Everyone has an opinion on AI citations. We have 2,000 Hours of Research and one clear starting point. We cracked the code of ChatGPT.
For technical teams
Google and ChatGPT share about 12% of results. A brand ranked number one on Google sits in a completely different database from the one an AI assistant draws on when a buyer types a category question. The channels are not additive — they are largely separate.
The misrepresentation risk is quantified. Anthropic research demonstrated that roughly 250 documents are enough to change what a model believes about a fact, regardless of total dataset size. The research proved this by moving the Eiffel Tower from Paris to Lyon in a model’s knowledge base. The number was consistent whether the training corpus was 600 million or 13 billion documents. If the documents shaping what AI says about your brand are written by affiliates, review sites, and competitors, that is the answer buyers receive mid- evaluation.
| Metric | Result |
|---|---|
| AI mention rate (brand queries) | 89% |
| Own-domain citation rate | 11% |
| Own-domain citation rate (category queries) | 0% |
Wanna see proof?
The algorithm is called co-citation similarity.
AI citations analysed across ChatGPT and Gemini. The deepest study of its kind.
Overlap between what Google shows and what ChatGPT cites. Your Google ranking does not transfer to AI.
Documents is all it takes to shift what an AI model believes about a brand.
For technical teams
Co-citation similarity predates digitisation. It comes from print research-journal citation analysis: when many papers cite the same documents together for a topic, those documents are treated as authoritative for it. LLMs apply the same logic at scale. When two domains appear in the same AI answer, they share a co-citation edge. The more answers they co-appear in, the stronger the edge. WLDM ran this analysis on a real client’s blackjack category and found 127 domains recurring consistently across roughly 3,000 total AI citations. That recurring cluster is the authority bubble. Brands inside it get cited. Brands outside it get mentioned incidentally, or not at all.
Listicles dominate AI citation because they operationalise co-citation. A “best of” or “top 10” page explicitly groups the same sources together — the model reads that as a co-citation signal. Getting placed in a listicle the model already pulls from is co-citation hijacking. WLDM identifies the high-centrality clusters where competitors are co-cited and inserts the client into the same contexts.
| Metric | Result |
|---|---|
| AI mention rate (brand queries) | 89% |
| Own-domain citation rate | 11% |
| Own-domain citation rate (category queries) | 0% |
Harmonic Centrality (Marchiori & Latora, 2000) measures how close a domain is to every other domain in the co-citation graph by summing the reciprocals of shortest-path distances:
HC(v) = Σ 1 / d(v, u) for all u ≠ v
Higher HC means the domain is connected to more diverse citation sources — it is harder to displace from the authority cluster. HC is used because it handles disconnected graphs correctly. AI citation networks are not fully connected: not every domain is cited for every topic. PageRank and betweenness centrality fail on disconnected graphs. HC does not.
| Metric | What it measures | Verdict for AI citations |
|---|---|---|
| Mention count | How many times a brand appears | No. Gameable; ignores context |
| PageRank | Random-walk probability on directed graph | No. AI citations are undirected co-occurrences; fails Density Axiom |
| Betweenness | Fraction of shortest paths through a node | No. Unstable on disconnected graphs; fails monotonicity |
| Degree | Count of direct connections | No. Ignores topology |
| Harmonic Centrality | Sum of inverse shortest-path distances | Yes. Handles disconnected graphs; stable; interpretable |
Boldi & Vigna (2014, Axioms for centrality) prove HC satisfies the Size, Density, and Score-monotonicity axioms. PageRank fails Density; betweenness fails monotonicity.
Authority tiers (normalised HC):
| Tier | nHC band | Meaning |
|---|---|---|
| Dominant | ≥ 0.75 | Cited across most queries; central in the cluster; very hard to displace |
| Established | 0.50 – 0.74 | Cited often; a known node, not the default |
| Emerging | 0.25 – 0.49 | Cited sometimes; on the radar, not yet load-bearing |
| Peripheral | < 0.25 | Rarely or never cited; the cluster does not treat it as a source |
AI does not cite the most famous brand. It cites the most connected one.
In every category, a small recurring set of sources appears together in AI answers —across users, across prompts, across models. Brands inside that cluster get recommended.
Brands outside it are mentioned incidentally, or not at all.
That's super interesting though — how does it help me?
AI visibility tools don’t build citations. We do. At scale.

[PENDING]
For technical teams
The AI Citations Audit forces ChatGPT to run live, non-cached searches for your category across 20+ buyer-intent queries. Every cited domain is extracted, co-citation edges are mapped, and Harmonic Centrality is computed for your domain and every domain cited alongside it. The result is a tiered network map: where your brand sits (Dominant / Established / Emerging / Peripheral), which domains are inside the authority bubble for your topic clusters, and which competitors are Dominant while your brand sits Peripheral.
The Audit runs in “Confident 3×” mode: three parallel searches per query. A citation appearing in all three is load-bearing. One appearing in one of three is noise. WLDM acts only on the stable set.
WLDM maps topic clusters across four maximum groups (cluster sprawl dilutes nHC). For each cluster, the co-citation graph is built from the Audit’s cited-domain list. Domains are sorted by HC tier: Dominant first, then Established. Those are the domains whose citation of your brand meaningfully moves your network position. Domains at Emerging or Peripheral tier are logged for long-tail volume purposes only.
WLDM also flags domains where competitors are co-cited alongside high-centrality nodes but your brand is absent. Those are the co-citation hijacking targets.
Every candidate placement is scored on two dimensions before any outreach:
Harmonic Centrality (HC): where this domain sits in the topic’s co-citation graph
Cosine relevance: how closely this specific placement’s paragraph matches your target pages
These plot to a 2×2 decision matrix:
| Cosine HIGH (≥0.45) | Cosine LOW (<0.25) | |
|---|---|---|
| HC HIGH (Dominant / Established) | IDEAL — pursue aggressively | WASTED AUTHORITY — negotiate paragraph rewrite, or walk away |
| HC LOW (Emerging / Peripheral) | NICHE RELEVANCE — useful in volume | SKIP |
Most agencies buy by Domain Rating, a single number that ignores cosine. A DR 85 domain with a placement paragraph scoring 0.18 cosine is wasted spend. WLDM scores the paragraph, not just the domain.
The cosine model is three-signal:
score = 0.50 × cos(context paragraph, target page) + 0.30 × cos(referring page body, target page) + 0.20 × cos(anchor text, target page)
Generic anchors (“click here”, bare URLs) trigger redistribution: the 20% weight shifts 60/40 to the two context signals, preventing score inflation.
Because a backlink and an AI citation are the same page, acquisition is standard link outreach pointed at a different selection problem. The math changes; the engine is the same. WLDM contacts the target domain, negotiates paragraph placement or linked mention, and varies the wording across placements.
Wording variation is not stylistic preference — it is structural. AI retrieval searches the same question several ways and assembles the answer from chunks across varied queries. A single perfectly-worded placement is weaker than the same claim phrased differently across several sources. WLDM’s placement plan specifies wording variation by design.
Off-site entity cultivation:
Beyond placement acquisition, citation probability is raised by building the structured entity surfaces that LLMs treat as authoritative:
Wikidata QID mapping. Every entity the client references needs a Wikidata Q-number. QIDs are harvested or created for the organisation, key people, and distinctive products. Each entity is populated with complete, verified properties: instance of, industry, founding date, founder, headquarters, official site. The completed QID map feeds the on-site sameAs implementation — the seam between AI Citations (off-site) and GEO (on-site schema). See the GEO page for on-site implementation.
Author profiles, earned coverage, industry directories. The same entity consistently described across surfaces (name, location, identifiers) raises model confidence in its existence and claims. Deliberate wording variation across surfaces validates the entity.
Measurement cadence:
• Per-placement: cosine score before sign-off
• Monthly: brand-mention velocity, Bing AI Performance, LLM prompt visibility
• Quarterly: full Audit re-run, Δ nHC per cluster, Δ citation rate per cluster
• Annual: off-site audit (Wikidata, author profiles, directories, partners)
Target trajectory: Δ nHC +0.05 per quarter per cluster; Δ citation rate +5–10 points per quarter.
Proof
We ran the research. No other agency has done this at scale.
WLDM analysed more than 11 million AI citations across ChatGPT and Gemini in partnership with DataForSEO. Not what AI claims to prefer. What it actually does, at scale, measured.
The finding: a small set of sources, roughly 10 to 20 percent, appears together consistently across every answer in a category. The rest cycles in and out. Your competitors are in that stable set. Getting your brand into it is what this programme is for.
For technical teams
The study dataset was 1 million ChatGPT citations and 1 million Gemini citations scraped August to November 2024. Each citation URL was normalised and enriched with: authority metrics, page-level traffic, schema presence, content type, freshness, Harmonic Centrality (extracted from the Common Crawl graph), and three cosine similarity scores measured separately for keyword-to-title, keyword-to-content, and keyword-to-topic.
Statistical tests: Spearman correlation of citation rank against each feature; ordinal regression with rank as the dependent variable; cited-vs-not classification against SERP-controlled negatives. ChatGPT and Gemini were run separately for a two-model defensible comparison.
Cosine sweet spots from the study:
| Similarity measure | Sweet spot | What it means |
|---|---|---|
| Keyword ↔ page title | ~0.5 | Over-optimised titles (>0.8) rarely move citation; keyword stuffing overshoots |
| Keyword ↔ page content | 0.3 – 0.5 | Thematic relevance beats exact-match density |
| Keyword ↔ topic vector (BERT) | 0.7 – 0.8 | Most consistent single predictor of citation at position 1 |
The pattern holds across all models tested except Perplexity, which runs a different retrieval model.
A second study is currently running with Search Atlas and Surfer SEO. An exploratory 10,000-citation sample reproduces the patterns from the million-citation set, confirming the finding is not scale-dependent. The larger dataset is for publishability.
A brand with global recognition. Cited by its own domain in 11 of every 100 AI answers.
One of the most recognised brands in their category worldwide appeared in 89 out of every 100 AI responses about their space. Their own domain was the cited source in 11 of those responses. On category queries — the ones where buyers decide which brand to consider — it was zero.
Revenue was being influenced by sources the brand had no relationship with and no ability to influence. The shortlist was forming without them.
For technical teams
Stake.com is geo-blocked in the US and restricted in the UK and Australia. No GA, GSC, or CMS access was provided. Cloudflare blocks standard crawling. The entire baseline was built from public and third-party infrastructure: DataForSEO’s AI optimisation ChatGPT scraper (force_web_search=true), Ahrefs, and CrUX. Canada was used as the measurement proxy — an English, crypto-friendly market where Stake operates openly and SERP data is undistorted. The baseline reproduces next quarter from zero.
The prompt grid — 45 prompts across 8 buyer-intent groups:
| Prompt group | Focus |
|---|---|
| Category (casino) | “best crypto casino”, “top online casino for crypto” |
| Category (sportsbook) | “best crypto sportsbook”, “Stake alternatives” |
| Brand trust | “is Stake legit”, “Stake review” |
| Comparison | “Stake vs Roobet”, “Stake vs Cloudbet” |
| Crypto rails | “fastest crypto withdrawal casino” |
| Games / originals | best Plinko game online”, “Stake Plinko” |
| Responsible gambling | “does Stake have responsible gambling tools” |
| Sports entity | “which betting site sponsors UFC”, “Drake Stake sponsorship” |
Each prompt scored on six columns: mentioned (y/n), cited from own domain (y/n), prominence/position, sentiment, grounding sources (which third-party domains the AI used), competitors cited.
Result: 89% mention rate. 11% own-domain citation rate. 0% own-domain citations on category queries.
The P44 proof case: Asked “which betting site sponsors UFC and Drake”, ChatGPT cited stake.com/sponsorships/ufc — the single own-domain citation across the entire 45-prompt baseline. The reason is decisive: a dedicated, factual entity page exists for that claim. On every other prompt where a comparable page is absent, the answer is grounded on affiliates, or Stake is absent entirely. This single contrast is the business case for the entire programme: own-domain citation happens where a dedicated factual page exists.
The ghost-citation surface map: The domains AI repeatedly used to ground answers about Stake were identified as the prioritised placement target list: cryptoslate.com, cryptonews.com, cloudbet.com, coincarp.com, gamingamerica.com, cryptocasinohouse.com, tokenist.com, bitcoin.com, casino.org, 99bitcoins.com, and others. These domains are simultaneously proof that Stake’s own pages are not the citation source, and the off-site acquisition targets for the AI Citations programme.
AI vs Google competitor set: The operators ChatGPT cited alongside Stake on category queries (Rainbet, BC.Game, Cloudbet, Thunderpick, Roobet) barely overlap the regulated and offshore books Stake competes with on Google’s SERP. Two different competitive battlegrounds, which is why AI visibility must be measured directly.
Your brand is either inside the citation clusters or it isn't.
We will show you which.
We audit your current AI citation position across the queries that matter to your buyers. We map where your competitors are that you are not. We show you the sources grounding every answer — and score each one for how much it would move your position if you were in it.
You leave with a clear picture of the gap and a 90-day plan to close it.
Free. No obligation. Reviewed within 48 hours.
For technical teams
The audit produces a scored map of your AI citation position across four topic clusters. Deliverables:
Citation presence across target queries. Which queries your brand is cited on, which it appears in without citation, and which it is absent from entirely. Own-domain citation rate vs. third-party grounding rate per query group.
Authority bubble mapping by nHC tier. Every domain currently cited in your category is tiered Dominant, Established, Emerging, or Peripheral. You see which Dominant and Established domains are inside the clusters your competitors already occupy that your brand does not.
HC × cosine placement scoring. Existing placements you already hold — backlinks, directories, editorial — are scored against the 2×2 decision matrix. You see which are in the Ideal quadrant, which are wasted authority, and which need a paragraph rewrite to start contributing to your citation position.
90-day citation engineering roadmap. Prioritised placement targets by nHC tier and cosine fit, with acquisition approach, wording variation plan, and quarterly re-baseline schedule.
Target trajectory for a 12-month programme: Δ nHC +0.05 per quarter per cluster, Δ citation rate +5–10 points per quarter, rising median placement cosine quarter on quarter.
The baseline requires no access to your GA, GSC, or CMS. The method reproduces next quarter with the same instrumentation, so progress is measured against a locked benchmark, not recalibrated.
