How to Track Brand Mentions in AI Answers: A Step-by-Step Playbook
Vendor homepages in this category make confident-sounding claims, "tracks brand mentions across seven engines," "30,000-plus users," case studies citing thousands of tracked mentions, but rarely explain the actual method behind the number in enough detail for you to replicate it, or trust it, independently. This is the method we use ourselves, described in enough detail that you could run it by hand tomorrow with nothing more than a spreadsheet and access to a few chat interfaces. It's the same process behind every citation percentage published on this site.
Why This Needs Its Own Tracking Process, Not a Repurposed One
Tracking a brand mention in an AI-generated answer is a different problem than tracking a brand mention in a news article, a tweet, or a Google search result, for one structural reason: the content doesn't exist anywhere until the exact moment you run the prompt. There's no crawlable archive of "everything ChatGPT has ever said about your brand" the way there's a crawlable archive of news coverage or social posts. That means brand mention tracking for AI answers is inherently an active, repeated-sampling process, you generate the data yourself, on a schedule, rather than passively monitoring an existing feed.
Step 1: Define Your Real Competitive Set
List every brand a genuine buyer in your category would consider alongside you, not just your two or three closest rivals. Our own tracked set spans 17 named players, from category leaders to tools that showed up in a single response out of 240. A narrow competitive set understates how contested your category actually is; a set that's too broad (including brands no real buyer would ever compare you to) dilutes the signal with noise.
Step 2: Write Real Buying-Intent Prompts, Not Generic Keywords
The prompts you run matter more than almost any other variable in this process. "Best [category] tools" is a fine starting prompt, but it's the easiest one to answer well, and it inflates every brand's apparent visibility relative to more specific, comparison-heavy questions a real buyer closer to a decision actually asks. Build a mix: a couple of broad discovery prompts ("what are the best tools for X"), a couple of comparison prompts naming specific competitors ("how does A compare to B"), and a couple of decision-stage prompts ("is there a free way to test A before committing", "alternatives to [specific competitor]"). Ten to twenty prompts covering this mix is enough to produce a meaningful result without becoming an open-ended project.
Step 3: Choose Your Engines and Languages Deliberately
At minimum, cover the engine or engines with the highest research volume for your category, don't default to "just ChatGPT" without checking. Our own data shows citation share for the same brand can swing more than 25 percentage points between engines, so a single-engine check systematically misrepresents your real visibility. If a meaningful share of your buyers research in a language other than English, run the same prompt set translated into that language too, don't assume an English-language result generalizes, our per-language data shows swings of 15 points or more on identical brands.
Step 4: Run Every Prompt Through Every Engine, and Log the Raw Response
For each prompt, in each engine, in each language, record two things separately: whether your brand's domain was explicitly cited as a source (the harder signal), and whether your brand's name appeared anywhere in the generated text (the softer signal). Keep the full raw response text, not just a yes/no flag, you'll want it later to audit borderline cases and to understand how you were mentioned, favorably, neutrally, or as part of a "here are some other options" list, not just whether you were.
Step 5: Calculate the Percentage, Then Break It Down
The core formula: mentions divided by total responses in your measurement set, times 100. Calculate this once for your blended total, then again separately per engine, and again per language if applicable. The blended number is useful as a single headline figure; the broken-down numbers are where the actionable insight actually lives, exactly as covered in our companion piece on calculating AI Share of Voice.
Step 6: Distinguish the Two Detection Signals When You Report the Result
Don't present a combined mention count without disclosing which detection method produced it. A citation-only count is the more conservative, more defensible number. A count that includes name-matches will run higher, sometimes substantially, and carries a real limitation: a brand name that's also a common word (a hypothetical tool called, say, "Rank" or "Scale") can register a false positive on the name-match signal when the word appears in an unrelated sentence. If you're publishing this number anywhere a stakeholder might scrutinize it, disclose the method plainly rather than letting an inflated combined number stand unexplained.
Step 7: Re-Run on a Fixed Schedule and Track the Trend
A single run is a snapshot, not a stable score, AI model behavior updates more frequently than a traditional search ranking algorithm's update cycle, and outputs carry real variance even on an identical prompt run twice in a row. Monthly is a reasonable default cadence for most categories. What you're actually building, over three or four consecutive monthly runs, is a trend line, which is a meaningfully more useful thing than any single month's number viewed in isolation.
A Worked Example, Start to Finish
Here's the process applied, using our own July 2026 measurement as the concrete example. Competitive set: 17 named players in the AI-visibility-tools category. Prompts: 20 buying-intent questions, covering discovery, comparison, and decision-stage intents. Engines: ChatGPT, Claude, Gemini, Perplexity. Languages: English, Portuguese, Spanish. Total responses: 20 × 4 × 3 = 240. Detection: combined citation-domain and name-match signals. Result: Semrush led at 33.3% (80 of 240), followed by Profound (25.0%), Otterly.AI (23.3%), and a long tail down to single-response mentions.
A Sample Prompt Library You Can Adapt
To make Step 2 concrete, here's the mix of prompt types we actually use, generalized so you can swap in your own category:
- Broad discovery: "What are the best tools for [category]?" and "What should I use to [core job your product does]?"
- Problem-first: "How do I [specific problem your product solves]?", useful because it surfaces which brands get recommended when the buyer hasn't yet named the category themselves.
- Named comparison: "How does [Competitor A] compare to [Competitor B]?" and "Is [Competitor A] worth it compared to alternatives?"
- Decision-stage: "Is there a free way to try [category] tools before committing?" and "What's the cheapest good option for [category]?"
- Switching intent: "What are alternatives to [Competitor A]?" and "Why would someone switch away from [Competitor A]?"
- Recency-checking: "What are the newest [category] tools in 2026?", useful for catching whether an engine's knowledge is current enough to know about newer entrants, including you if you're one.
Six categories, two to four prompts each, gets you comfortably into the 10-to-20 range recommended above, with enough variety that you're not just measuring performance on the single easiest question type.
What Good Tracking Looks Like vs. What Doesn't
A tracking process worth trusting produces a result that's specific, dated, and disclosed: "cited in 6 of 20 prompts (30%) across ChatGPT and Perplexity, English only, measured July 2026, citation-domain detection only." A tracking process not worth trusting produces a vague, undated, method-free claim: "we're highly visible in AI search," with no percentage, no prompt count, no engine list, and no way for anyone reading it to check the work. The difference isn't sophistication, a spreadsheet-based manual process done with full disclosure is more trustworthy than an expensive tool's dashboard presented without any of that context. Build the habit of disclosing your method every time you report this number, internally or externally, it costs nothing and it's the single biggest driver of whether anyone should believe the result.
DIY vs. Tool-Assisted: Where the Tradeoff Actually Lives
Running this manually, prompt by prompt, engine by engine, logged in a spreadsheet, is entirely workable for a modest scope: one brand, 10 to 15 prompts, two or three engines, checked monthly. It becomes a meaningfully heavier lift once you're tracking multiple competitors' full response text, multiple languages, or a weekly rather than monthly cadence, at that point the labor scales roughly linearly with scope, and a purpose-built tracking tool (whether a GEO-native platform or an AI-visibility module inside a broader SEO suite) starts saving real time rather than just adding a subscription cost. The underlying calculation doesn't change either way, what a tool buys you is automation and scale, not a different or more accurate formula.
Turning a Single Run Into an Ongoing Program
The step-by-step process above describes one measurement run. Turning it into an ongoing program that actually informs decisions requires two additional habits. First, keep every run's raw data, not just the summary percentage, in a consistent format (same spreadsheet structure, same prompt wording, same competitor list) so that month-over-month comparisons are measuring the same thing rather than silently changing methodology run to run. Second, pair each run with a short written note on what changed since the last one, in your content, in the competitive landscape, or in nothing you're aware of, so that a future reader (including a future version of yourself) can tell the difference between a real trend and unexplained noise. A tracking process that produces a clean number but no institutional memory of what drove it tends to get abandoned after two or three cycles; one that builds a running narrative alongside the data tends to survive long enough to actually inform strategy.
Common Pitfalls That Undermine This Process
Only checking once and treating it as a stable score. The single most common mistake. Report every number with the measurement date attached, and never present a one-time check as an ongoing trend.
Using prompts that already name your brand. If every prompt contains your own company name, you're measuring "does the engine repeat back the name I fed it," a materially easier bar than "does the engine recommend this brand unprompted," and the resulting number will overstate your real visibility.
Skipping the raw-response review. A binary yes/no mention count misses whether you were recommended favorably, mentioned neutrally in a list of five options, or referenced only in a "you might also consider" afterthought. These are meaningfully different outcomes that a percentage alone can't distinguish.
Ignoring engine and language breakdowns to save time. Covered in detail above, but worth repeating here: a blended number can hide you being completely invisible on the one engine or in the one language your actual buyers use most.
Checklist: Your First Brand Mention Tracking Run
- [ ] Competitive set defined, 8 to 15 named brands minimum
- [ ] 10 to 20 buying-intent prompts written, covering discovery, comparison, and decision-stage intent
- [ ] Engines selected based on real buyer research behavior, not assumption
- [ ] Languages selected, beyond English, if a meaningful buyer share isn't English-first
- [ ] Detection method decided (citation domain, name-match, or both) and disclosed when you report results
- [ ] Raw response text logged, not just a yes/no flag, for later review of mention quality
- [ ] Blended percentage calculated, then broken down by engine and by language
- [ ] Re-measurement scheduled on a fixed monthly cadence
Related Reading
For the formula and worked calculation behind the percentage this process produces, see AI Share of Voice: the formula and a worked example. For building a more structured, ongoing version of this process, see how to set up prompt monitoring and how to run a full GEO audit. For the complete leaderboard this exact methodology produced in our own category, see which brands AI engines actually recommend.
This playbook describes the exact method behind original research by the GeoHero Research Team: 240 AI-engine responses across ChatGPT, Claude, Gemini, and Perplexity (20 buying-intent prompts, three markets, July 2026). Every percentage referenced reflects a single measurement run, not an average across repeated runs. We re-run this monthly. Want this process run for your own brand instead of building the spreadsheet yourself? Get the full report.
Frequently asked questions
What counts as a 'mention' when tracking brand visibility in AI answers?
Two distinct signals, worth tracking separately rather than combining into one number. A cited source domain (the AI engine explicitly links to your site) is the harder, more reliable signal. A name-match in the generated text (your brand name appears in the prose, with or without a link) is a softer signal that catches mentions a citation alone would miss, but carries a real risk of false positives if your brand name is also a common word.
How many prompts do I need to get a reliable brand mention count?
Ten to twenty real buying-intent prompts is a reasonable working range for most categories, enough to produce a meaningful percentage without turning the exercise into an open-ended research project. Fewer than five prompts makes any single citation swing the whole result; more than thirty rarely changes the headline finding enough to justify the added labor of running and logging them manually.
Can I track brand mentions in AI answers without paying for a tool?
Yes. Run your prompt list manually through each engine's chat interface, log which responses mention your brand and how, and calculate the percentage yourself. It's entirely workable for a monthly check across a handful of prompts and engines; it becomes labor-intensive once you're tracking many competitors, multiple languages, or a high-frequency schedule, which is when a purpose-built tracking tool starts paying for itself in time saved.
How is this different from a social listening tool?
Social listening tracks what people say about your brand across social platforms, reviews, and news. This tracks what AI engines say about your brand when someone asks them a direct question, a distinct, machine-generated source most social listening tools were never built to capture, since the content doesn't exist anywhere public until the moment someone runs the prompt.
How often should I re-run this tracking process?
Monthly is a reasonable default for most categories, frequent enough to catch a real trend without over-reacting to single-run noise. AI model behavior updates more often than a search engine's core ranking algorithm typically does, so treat each run as one data point in a trend line, not a permanent score.