How to Track Your Brand's Mentions in ChatGPT
If you want to know whether ChatGPT recommends your brand, asking it once tells you almost nothing. Answers vary between sessions, between accounts, and between phrasings of the same question. A single response is an anecdote.
This guide covers how to turn that anecdote into a measurement you can actually track.
TL;DR
- Ask many questions many times — a single response is noise, not data.
- Never ask "is [your brand] good?" Naming your brand in the prompt guarantees it appears in the answer, which tells you nothing about discovery.
- Test unaided questions ("best tools for X") separately from aided ones ("how does X compare to Y").
- Log who else appears — your competitive set in AI answers is often not the one you expect.
- Re-test on a fixed schedule. Trend beats snapshot every time.
Why one-off checks mislead you
Three things make manual spot-checks unreliable:
Responses are non-deterministic. Ask the same question twice and you may get two different vendor lists. Any single answer sits somewhere in a distribution you cannot see from one draw.
Your session is not neutral. Memory, custom instructions, and prior turns in the conversation all shape what comes back. If you have discussed your own company in that account, results are contaminated.
Phrasing dominates. "Best project management software" and "project management tool for small agencies" produce different lists. Testing one phrasing measures one narrow slice of reality.
Step 1 — Build a question set, not a keyword list
Keywords are the wrong unit. People type keywords into search boxes; they ask assistants full questions. Write out the questions your buyers actually ask, grouped by intent:
- Informational — "What is [category]?", "How does [category] work?"
- Comparative — "Best [category] tools", "[Category] software compared"
- Purchase-intent — "Cheapest [category] for [segment]", "[Category] with [specific feature]"
- Problem-first — "How do I fix [problem your product solves]?"
Aim for thirty to fifty questions. Fewer than that and one odd result skews everything.
Step 2 — Ask blind, not branded
This is the mistake that invalidates most DIY tracking.
If you ask "Is Acme a good CRM?", Acme appears in the answer. Of course it does — you put it there. That measures nothing about whether a buyer who has never heard of you would discover you.
Run two separate tests:
- Unaided (blind): your brand is never mentioned in the prompt. "What are the best CRMs for a 20-person sales team?" This measures genuine discovery, and it is the number that matters.
- Aided: your brand is named. "How does Acme compare to Salesforce?" This measures how you are characterised once someone already knows you — useful for reputation, useless for discovery.
Keep the two scores apart. Blending them produces a flattering number that hides the problem.
Step 3 — Use clean sessions
Before each run:
- Turn off memory and custom instructions, or use a fresh account.
- Start a new conversation for every question. Prior turns bias what follows.
- Do not "correct" the model mid-conversation and then re-ask.
Step 4 — Record more than yes or no
For each response, log:
| Field | Why it matters |
|---|---|
| Brand mentioned? | The headline number |
| Position in the list | First mention carries far more weight than fifth |
| Framing | Recommended, listed neutrally, or mentioned as a negative |
| Competitors named | Reveals your real AI competitive set |
| Sources cited | Shows which pages the model leaned on |
That last row is the most actionable thing on the list. When a model grounds an answer in live results, the sources it names tell you exactly which pages earn citations in your category — and they are frequently review sites, comparison pages, and community threads rather than vendor sites.
Step 5 — Repeat on a schedule
Run the full set weekly or fortnightly, at a consistent time, and chart the percentage of responses that mention you. The absolute figure is less important than the direction.
Expect noise of a few percentage points between runs. Treat a change as real only if it persists across two or three consecutive runs.
Doing this manually vs. automating it
A thirty-question set, run across three assistants, repeated three times for stability, is 270 prompts per cycle. Done by hand, that is a slow day's work every fortnight, and consistency degrades quickly — different people phrase things differently and log results unevenly.
Manual tracking is a reasonable way to find out whether you have a problem. It is a poor way to manage one over time. This is the point at which most teams move to a platform, whether that is SIQA or something else.
Frequently asked questions
How often should I check? Weekly or fortnightly for most categories. Daily is noise unless you are running an active campaign.
Does asking about my brand teach ChatGPT to recommend it? No. Ordinary conversations do not update the underlying model.
Why do I appear for a colleague but not for me? Personalisation, memory, and account history. This is exactly why clean sessions matter.
What is a good score? It depends entirely on category concentration. Judge yourself against your own trend and your named competitors, not an industry average.
Published by the SIQA Editorial Team.