Every score on this site, yours or ours, traces back to a declared set of prompts, a declared set of engines, and a declared number of runs, all logged. This page is the how, in full, not a marketing summary of it.
Prompts, Engines, Runs: The Three Variables We Declare Every Time#
A citation rate is only meaningful once you know exactly what produced it. We declare all three inputs before we show you a single number.
1
Prompts come from real demand, not our guesses
Where we can, prompt sets are derived from measured search-demand data for the category, in the buyer's own words, not questions we hand-picked because they made a good screenshot. Our own category baseline used 20 buying-intent prompts, pulled the same way.
2
Every engine is measured, and labeled by family
Chat engines and AI-generated search results are different surfaces with different mechanics. We measure both, but we never merge them into one blended score without saying which family each number came from.
3
Runs are counted, and shown as a range when there is more than one
One run is a snapshot; several runs are a range. We say which one you're looking at every time, from our own published research to the report you'd get as a customer.
Two families, never blended
Chat Engines and Search-AI Surfaces Are Not the Same Measurement#
A citation inside a ChatGPT conversation and a citation inside a Google AI Overview happen through different mechanics. Reporting them as one number would hide which one actually matters for a given buyer journey. We keep the two apart, always labeled.
Chat family
ChatGPT, Perplexity, Gemini, and Claude. You ask a direct question and get a conversational answer, sometimes with a cited source. Model sampling has real randomness here, so this family gets measured multiple times per cycle.
Search-AI family
Google AI Overviews, Google AI Mode, and the AI answer panel inside Bing. An AI-written summary embedded directly in a familiar results page, not a chat conversation. It behaves close to deterministically for a fixed query, so it needs fewer runs, but it's still its own number.
The 7 surfaces
Every Score Is Labeled By the Surface It Came From#
Two families, seven surfaces, one public method label attached to every number, so nobody has to guess what produced it.
ChatGPT
measured · native chat. Segmented by whether the answer used live web grounding.
More repetitions produce a tighter, more trustworthy number, and cost more to run. We never let a cheap tier borrow the label a more rigorous one earned. Every score says which tier produced it.
Free scan
One run per prompt, on every engine.
single-run snapshot
Compete report
Five runs per prompt on chat engines, one to two on Search-AI surfaces.
report-grade, with a wide interval shown
Deep Research report
Seven to eight runs per prompt on chat engines, two on Search-AI surfaces, matching the floor set by the published GEO reproducibility survey.
research-grade
Confidence, not false precision
Every Repeated Score Ships With a 95 Percent Confidence Interval#
We compute the interval by bootstrap resampling the runs behind a score, not by assuming a bell curve that AI-answer data does not actually follow. A number without an interval is theater: it looks precise and tells you nothing about how far it could have landed somewhere else. At a single run, we do not fake an interval, we show the single-run-snapshot label instead, and say so.
A citation rate of 40% with a 95% interval of 10% to 70% is a very different claim from 40% with an interval of 35% to 45%, even though the headline number is identical. We show the width every time, never just the midpoint.
Selection vs Absorption
Getting Cited and Shaping the Answer Are Two Different Events#
Published research on LLM citation behavior found that inclusion happens in two stages with different drivers. Treating them as one blended number hides which one you are actually improving.
Selection
Did the brand enter the pool of sources the model drew from at all? Driven mostly by classic authority signals: backlinks, brand recognition, third-party mentions. This is what most published citation-rate numbers actually measure.
Absorption
Once selected, how much did the brand actually shape the substance of the answer, versus sitting unused in a source list? Driven by extractable evidence density: definitions, statistics, comparisons, and structured content a model can lift directly. Ranking well solves Selection. It does not automatically solve Absorption.
Within-tool, never across
We Never Compare an Absolute Score Between Two Different Engines#
Citation behavior varies enormously from one AI engine to the next. Published research measured swings of up to 615 times between engines on the same category, and found only about 11% overlap between the sources ChatGPT cites and the sources Perplexity cites for comparable queries. A change inside one engine, tracked over time, is a real signal. A comparison of one engine's raw score against a different engine's raw score is not a finding, it is noise dressed up as one.
It means the number came from exactly one pass, not an average of repeated runs. LLM outputs are not fully deterministic: ask the same question twice and you can get two different citations. A single-run number is real data, but it's a photograph, not a trend line, and we label it that way every time, including our own 240-response category baseline.
How many runs go into a number that isn't a single-run snapshot?
Chat-family engines are sampled multiple times per measurement cycle; Search-AI surfaces need fewer runs because they're closer to deterministic for a fixed query. Where we run more than once, we show a range instead of one flattering figure.
Do you ever combine chat-engine and Search-AI results into one score?
No. They're different surfaces, measured differently, and reported separately every time, labeled by family. A blended "AI visibility score" would look cleaner and tell you less.
Why won't you promise a ranking or a guaranteed citation?
Because nobody controls what a model chooses to cite, not us, not you, not even the lab that trained it. Anyone promising a guaranteed AI citation is promising something outside their own control. What we can guarantee is that the number we hand you was produced the way this page describes, not made up.
In July 2026 we ran the exact methodology described above on our own category, before selling it to anyone. Here is what each module measured, pulled straight from the logged research runs.
Competitor keyword mapping
Every keyword 14 competitor domains already rank for: 1,721 unique keywords, the raw material for the mandatory-ranking queue.
Search results + AI Overview landscape
Who Google's AI Overview cites for the category today, sampled across five languages.
Citation baseline (the 240-response study)
The buying-intent prompt study behind every leaderboard number on this site: 240 scored answers across four engines.
Gap analysis
Cross-referencing all of the above into a validated content queue.
We run the exact methodology described on this page against our own category, and publish the results as a dated, running case study, updated in place as the numbers move, or don't. If a method can't survive being applied to its own maker, it isn't a method.
Published, dated, and updated in place, not a retrospective
Every number on this page comes from asking real AI engines real questions and counting what came back. Here is exactly how to run the same check yourself: no account, no tool, just the engines and a stopwatch.
1
Ask the same question
Take a prompt from a declared set, like the example below, and ask it to each engine you care about, exactly as written, in a fresh conversation with no prior context.
2
Count what gets cited
For each answer, note whether your brand or domain shows up, and which sources the engine names or links to. A single run is a snapshot, not a trend.
3
Compare within the same engine over time
Repeat the same prompt on the same engine on different days to see how citation behavior moves. Never compare one engine's raw count directly against a different engine's raw count.
Example prompt from our buying-intent setbest CRM software providers
See a real measured example
A real measured example is not published yet. It is queued from our own next self-scan and will replace this note once it lands.
Sampling tiers renamed for the one-time launch: Starter and Pro become the Compete and Deep Research reports, the monitoring-grade label becomes report-grade, and the Compete tier locks at five runs per prompt. Run counts, the confidence-interval method, and every honesty label are unchanged.
v1.0 July 18, 2026
Initial publication: sampling tiers with honest labels (single-run snapshot, monitoring-grade, research-grade), Selection and Absorption measured separately, within-tool comparison only, versioned prompt sets, and a method label attached to every surface.