Axy.digital

GEO Research

First-Party vs. Third-Party Citations: The Breakdown Per LLM (and What to Do About It)

First-Party vs. Third-Party Citations: The Breakdown Per LLM (and What to Do About It)

Understanding First-Party and Third-Party AI Citations, Per LLM (2026)

Most brands know whether they show up in ChatGPT, Gemini, Claude, or Perplexity. Far fewer know what's actually being cited when they do.

We pulled this apart with a full citation forensics study: 90 buyer-intent prompts about marketing software, put to ChatGPT, Claude, Gemini, and Perplexity via API, logging every source behind all 360 answers, 6,075 unique URLs across 2,396 domains. This sits on top of Axy.digital's broader LLM visibility database, which tracks citation patterns across the 1,200+ brands now running on the platform, so what follows isn't a one-off snapshot, it's a pattern that holds up at scale. Here's the real breakdown of first-party vs. third-party citations, per LLM, and what to do with each.

There Is No Single "AI Web"

Source-universe overlap between engines

The first thing the data makes clear: treating "AI search visibility" as one channel is a mistake. Across the 2,396 domains cited, only 18 domains (0.8%) were used by all four engines, the rest is effectively four separate internets. 87.6% of domains were cited by exactly one engine. Even the closest pair of engines, Claude and Perplexity, only overlaps by 16.2%; ChatGPT and Gemini overlap by as little as 3.6%.

Of the 18 domains every engine drew on, 15 are vendor-owned software properties (activecampaign.com, blaze.ai, braze.com, buffer.com, jasper.ai, monday.com, sproutsocial.com, zapier.com, and others). Only three non-vendor sources cleared all four engines: G2, TheCMO, and Venture Harbour. Not Wikipedia. Not Reddit. Not a single news outlet.

What this means: winning on Perplexity teaches you almost nothing about winning on ChatGPT. A citation strategy has to be built engine by engine, not as one generic "get cited by AI" motion.

First-Party Citations

First-party citations are URLs on your own website domain that LLMs are pulling into their answers directly.

The data makes a strong case for why this is the highest-leverage category, full stop: vendor and SaaS marketing content (overwhelmingly first-party pages) accounts for 75–93% of all citation volume across every engine, Claude 93.3%, Gemini 89.6%, ChatGPT 84.1%, Perplexity 75.4%. It's also the only source type that's near-saturated across all four: cited in 100% of ChatGPT's and Perplexity's answers, 94.4% of Gemini's, and 81.1% of Claude's.

And the bar isn't domain authority, it's relevance. Between 15% and 27% of cited domains sit on .ai or .io TLDs, and several sites with no meaningful backlink profile out-cited category giants. On Gemini, salesforce.com and hubspot.com were cited in 0.0% of the 90 answers, while two unknown .ai/.com startups were cited in 6.7% and 7.8%.

What to do: Pull the list of your own pages that are getting cited. Look for patterns in format, depth, and structure across them, then produce more content that matches those patterns. Don't assume you need incumbent-level authority to compete here, a page that directly answers the question can out-cite a market leader that doesn't have one.

Third-Party Citations, By Category and By Engine

Third-party citations by domain and by engine

Third-party citations don't split evenly across sources, and which third-party category matters depends entirely on which engine you're optimizing for.

Publishers (News & Trade Media)

News and trade media shows up in 83.3% of ChatGPT's answers, techradar.com alone appeared in 81.1% of them, but it's nearly absent everywhere else: 6.7% of Claude's answers, 4.4% of Perplexity's, 2.2% of Gemini's.

One TechRadar review of a single product was reused across 24 separate ChatGPT answers, the single most reusable asset found in the entire study.

What to do: if you're chasing trade-press coverage, prioritize ChatGPT-facing outlets like TechRadar over a scattershot press list. One strong, durable review compounds far more than broad coverage.

This is exactly the gap Intelligent PR is built to close. Instead of manually chasing individual writers and editors, Axy.digital programmatically drafts market reports and syndicates them across 200+ high-authority newswire endpoints, landing you directly on the publisher domains that are already being cited, without the pitch cycles.

Reddit

ChatGPT and Perplexity both lean on Reddit heavily, but in opposite ways. ChatGPT retrieves Reddit in 88.9% of its answers (roughly 12 Reddit URLs per answer) but only links or credits it in 1.1% of them, a ~1% conversion rate from "read" to "cited." Perplexity does the reverse: far fewer Reddit URLs overall (about one per answer), but those references are the visible citations, showing up in 80% of its answers. Gemini cites Reddit in 8.9% of answers; Claude essentially never (0%).

What to do: know which goal you're playing for. If your KPI is a visible citation, target Perplexity threads, that's where Reddit mentions convert. If your goal is shaping what ChatGPT's model believes about your category (even without a visible link), Reddit is still worth seeding, but don't expect credit for it.

Product Directories & Review Sites

Review sites and software directories appear in 31.1% of ChatGPT's answers, 14.4% of Claude's, 6.7% of Perplexity's, and 4.4% of Gemini's. G2.com is one of only 18 domains cited by all four engines, and on ChatGPT it converts from "retrieved" to "cited" at 45%.

What to do: treat G2 as the one directory worth prioritizing everywhere. Beyond G2, weight your directory efforts toward ChatGPT and Claude, where this category actually shows up.

Competitor Content (Vendor & SaaS Pages)

This is the category that matters most, by a wide margin, see "First-Party Citations" above. Vendor and SaaS content (yours and your competitors') makes up the majority of what every engine cites, and it's not gatekept by incumbency.

What to do: study the format and topic of what's working for competitors, and for the unknown, no-authority pages that are somehow out-citing them, then publish your own version on your own domain. Match the structure that's proven to get cited, then put your own data and positioning behind it.

Two cheap, buyable assets worth calling out specifically: a strong trade-press review (see Publishers, above) and a Wikipedia page for your company. Wikipedia pages for Predis.ai, Omnisend, and StructuredWeb, none of them household names, were each pulled into 11–12 separate answers. A Wikipedia page that survives notability review is disproportionately cheap leverage relative to what it returns.

Retrieved ≠ Cited: The ChatGPT Paradox

The ChatGPT retrieval-to-citation funnel

ChatGPT is the one engine that prints its links in the answer body, which means we can compare what it pulled against what it printed. Only 20.3% of everything ChatGPT retrieves survives into the visible answer, and the survival rate varies wildly by source type: Reddit converts at 1%, Wikipedia at 13%, AI-tool listicle farms at a flat 0%, crawled, then discarded. Vendor blogs, by contrast, convert at 76–88%: Sprout Social, HubSpot, and Zapier get read and credited almost every time they're pulled.

What this means: stop measuring GEO by "did the crawler see me." Measure conversion, retrieved-to-cited. Vendor-owned, direct-answer content is the only source type that reliably converts on ChatGPT.

One Fifth of Claude's Answers Have No Citations at All

Claude returned zero sources on 17 of 90 prompts (18.9%), issuing a median of just one search. Gemini went source-free on 4 prompts; ChatGPT and Perplexity never did. Those uncited Claude answers still recommended products, pulled from training data, with no retrieval involved at all.

What this means: on a meaningful share of Claude's answers, there's no citation to win. Being in the training corpus matters as much as being crawled.

The matrix below reads by row: each cell is the share of that engine's 90 answers containing at least one source of that type. Engines can score high on several rows at once, because a single answer usually mixes types.

Source-type reach, by engine

The Takeaway

Knowing you're "visible in AI search" isn't enough, and treating it as one channel is a mistake. The brands pulling ahead know exactly which of their own pages are getting cited, which engine each third-party category actually matters for, and where a citation is worth chasing at all.

Want your own breakdown? Axy.digital runs this analysis on your brand and your category, first-party and third-party, per LLM.

Get Your Citation Breakdown →

Methodology: 90 buyer-intent prompts spanning category discovery, evaluation and comparison, alternative-seeking, problem/solution, use-case, and thought-leadership intents were put to ChatGPT, Claude, Gemini, and Perplexity via API on July 3, 2026. Every source each engine exposed was logged and classified by domain type. Because "source" doesn't mean the same thing to every engine — ChatGPT exposes a wide retrieval set, Perplexity a tight citation list — all percentages are computed within each engine, as a share of its own 90 answers. This study is an aggregate view drawn from Axy.digital's LLM visibility database, which spans the 1,200+ brands currently tracked on the platform.