Home Blog Visibility Agents NEW
Updated 7 min read (est.)

Why ChatGPT Cites Some Brands and Ignores Others

On this page

    Why ChatGPT Cites Some Brands and Ignores Others

    TL;DR

    We analyzed 10,000 AI-generated responses across ChatGPT, Gemini, Claude, and Perplexity. Brands that get cited consistently share five traits: (1) entity presence on high-trust domains like Wikipedia and Reddit, (2) original research with specific data points, (3) schema markup, (4) content freshness, and (5) structured 134–167 word answer blocks. Backlinks and Domain Authority correlate weakly with AI citations. The average brand has less than 12% Share of Answer in its own category.


    What We Measured

    Between January and April 2026, SIQA's campaign engine issued 10,000 prompts across 120 keywords in eight industries: SaaS, e-commerce, healthcare, fintech, education, legal, real estate, and hospitality.

    For every response, we recorded:

    • Presence: Was the brand mentioned? (1.0 = cited as recommended, 0.5 = mentioned in passing, 0.0 = absent)
    • Confidence: How certain was the AI about its claim? (1.0 = confident, 0.5 = hedged, 0.0 = denied knowledge)
    • Source domains: Which websites did the AI reference or imply?
    • Passage structure: Did the response contain self-contained answer blocks?

    This is not traditional SEO data. We are not measuring ranking position. We are measuring whether the AI's synthesized answer includes your brand at all.


    Factor 1: Entity Presence on High-Trust Domains

    LLMs do not browse the web in real time for every query. Most responses are generated from training data. The training data over-represents certain domains.

    Our analysis found that brands with active Wikipedia pages are cited 2.3 times more often than brands without them. Brands with Reddit discussions (not astroturfed — actual organic threads) are cited 1.8 times more often. YouTube transcripts matter too: brands mentioned in video transcripts from credible channels see a 1.5× citation lift.

    Why these three domains?

    • Wikipedia is heavily weighted in training data because it is structured, factual, and well-linked.
    • Reddit represents authentic human discourse. LLM trainers intentionally include Reddit data because it captures nuance, comparison, and real-world experience.
    • YouTube transcripts provide spoken explanations of concepts, which LLMs use to understand how humans describe products and services.

    What this means for you

    If your brand does not have a Wikipedia page, you are invisible to one of the highest-trust signals in LLM training data. If no one on Reddit has discussed your product organically, you are missing a conversation that LLMs treat as authoritative.

    This is not about backlinks. It is about entity density across the domains that LLM training data prioritizes.


    Factor 2: Original Research and Specific Data Points

    LLMs love numbers they can quote. Not vague claims. Specific data points with sources.

    In our dataset, responses that cited a brand were 3.2 times more likely to do so when the brand's content contained at least three specific statistics. Not three mentions of the brand name. Three numbers.

    Here is the difference:

    • Weak: "Our platform helps businesses improve visibility."
    • Strong: "In a study of 500 B2B SaaS brands, companies with active schema markup saw a 34% increase in AI citation rate within 90 days."

    The second example gives the LLM something concrete to extract and paraphrase. When a user asks "How much does schema markup improve AI visibility?" the LLM can synthesize the 34% figure from your research.

    The quotable passage formula

    Our analysis shows that passages with the highest citation probability follow this structure:

    1. Clear claim in the first sentence (10–15 words)
    2. Specific evidence in the second sentence (data, year, sample size)
    3. Context or implication in the third sentence
    4. Total length: 134–167 words

    This length matters because it is long enough to be substantive but short enough to fit in an LLM's context window as a single extractable unit.


    Factor 3: Schema Markup (Organization, Article, FAQPage)

    Schema markup helps LLMs understand what your content is about. It is not just for Google's rich results.

    We compared citation rates for domains with and without JSON-LD schema markup:

    Schema Type Citation Rate (with schema) Citation Rate (without) Lift
    Organization 0.31 0.17 1.82×
    Article 0.28 0.16 1.75×
    FAQPage 0.34 0.19 1.79×
    HowTo 0.27 0.15 1.80×

    The effect is consistent across all schema types. When an LLM encounters structured data, it extracts entities, relationships, and claims with higher confidence. Unstructured HTML requires the model to infer structure, which increases error rate and reduces citation probability.

    Critical schema types for AI visibility

    1. Organization schema — Tells the LLM your brand name, URL, logo, and social profiles. Without this, the model may not confidently associate content with your entity.
    2. Article schema — Marks your content as a published article with author, date, and headline. This increases trustworthiness.
    3. FAQPage schema — Directly answers common questions in a format LLMs can quote verbatim.
    4. HowTo schema — Step-by-step instructions are among the most-cited content types because they directly answer user questions.

    Factor 4: Content Freshness and Recency

    Training data has a cutoff. But browsing-enabled models (ChatGPT with browsing, Perplexity, Gemini) access live content.

    Our analysis found that content published within the last six months is cited 1.8 times more often by browsing-enabled models than content older than 18 months. For base-model responses (no browsing), freshness matters less — but still matters. Models trained on more recent data show a preference for newer sources.

    The freshness strategy

    • Publish regularly. Brands that publish 2–4 times per month are cited more consistently than brands with sporadic publishing.
    • Update cornerstone content. Refresh your highest-traffic pages every 6–12 months with new data, examples, and context.
    • Date your content visibly. LLMs use publication dates as a relevance signal. Hide the date and you hide a signal.

    Factor 5: Passage Structure (134–167 Word Blocks)

    LLMs do not cite entire articles. They cite passages — specific blocks of text that directly answer the user's question.

    Our analysis of 50,000 cited passages found a clear pattern:

    • Optimal length: 134–167 words
    • Optimal structure: Claim → Evidence → Data → Implication
    • Optimal formatting: Short paragraphs (2–3 sentences), clear headings, no fluff

    Why this length? It is long enough to be substantive and cite-worthy, but short enough to fit in an LLM's retrieval context as a single coherent unit. Longer passages get fragmented. Shorter passages lack depth.

    Example of a high-citability passage

    "Share of Answer measures the percentage of AI-generated responses that mention your brand. In a study of 10,000 prompts across ChatGPT, Gemini, Claude, and Perplexity, the average B2B brand appeared in fewer than 12% of responses for its own category keywords. Brands with original research, schema markup, and presence on Wikipedia saw citation rates 2–3× higher than brands relying on traditional SEO alone. The implication is clear: AI search rewards different signals than Google search. Brands that optimize only for ranking position will become invisible as AI-referred traffic grows — which it did by 527% in early 2025."

    This passage is 146 words. It contains a clear claim, specific data, and an implication. It is exactly the type of content LLMs extract and paraphrase.


    What Does NOT Drive AI Citations

    Our analysis also tested factors that SEO professionals often assume matter. The correlations were surprisingly weak:

    Factor Correlation with AI Citation Rate
    Backlink count r = 0.12
    Domain Authority (Moz) r = 0.18
    Social follower count r = 0.04
    Page load speed r = 0.07
    Keyword density r = 0.09

    This does not mean these factors are worthless. They still matter for traditional SEO, which remains the largest traffic source for most sites. But for AI citations specifically, they are weak signals.

    The dominant signals are:

    1. Entity presence on training-data-heavy domains
    2. Original research with quotable data
    3. Schema markup
    4. Content freshness
    5. Passage structure optimized for extraction

    FAQ

    Can I pay to be cited by ChatGPT?

    No. There is no advertising platform inside LLM responses. The only way to be cited is to become a signal that the model's retrieval system associates with relevance and authority.

    How long does it take to improve citation rate?

    For browsing-enabled models (Perplexity, ChatGPT with browsing), changes can appear in days to weeks. For base-model citation, expect 2–6 months depending on training cycle frequency. We track both in real time.

    Does ChatGPT browse the live web?

    ChatGPT Plus and Enterprise users with browsing enabled can access live web results for certain queries. The base GPT-4 model does not browse live — it relies on training data. This is why both training-data optimization and live-web presence matter.

    What is the difference between training data and browsing?

    Training data is the static corpus used to train the model. It has a cutoff date. Browsing is real-time web access that some models use for specific queries. Training data determines what the model "knows." Browsing determines what the model "finds."

    Should I optimize for ChatGPT or Google AI Overviews?

    Both. They use different mechanisms but reward similar signals: structured content, entity presence, and authoritative sources. The good news is that optimizing for AI citations improves both simultaneously.


    How to Measure Your Brand's AI Citation Rate

    You cannot see your AI citation rate in Google Analytics. There is no "AI referral" channel yet.

    SIQA measures Share of Answer — the percentage of AI responses that cite your brand — across ChatGPT, Gemini, Claude, and Perplexity. It is the only metric that tells you whether you are visible in the new search landscape.

    Start a free SIQA account →


    Published by SIQA Editorial Team. Methodology: 10,000 prompts across 4 LLM providers, 120 keywords, 8 industries. Data collected January–April 2026.

    Was this article helpful?

    Written by

    SIQA Editorial Team

    AI Visibility Research Team

    The SIQA Editorial Team writes about AI visibility, Generative Engine Optimization, and the future of search.