GeoHero

Methodology

How We Measure AI Visibility, and Show Our Work

Every score on this site, yours or ours, traces back to a declared set of prompts, a declared set of engines, and a declared number of runs, all logged. This page is the how, in full, not a marketing summary of it.

Methodology v1.1

The measurement

Prompts, Engines, Runs: The Three Variables We Declare Every Time#

A citation rate is only meaningful once you know exactly what produced it. We declare all three inputs before we show you a single number.

Prompts come from real demand, not our guesses

Where we can, prompt sets are derived from measured search-demand data for the category, in the buyer's own words, not questions we hand-picked because they made a good screenshot. Our own category baseline used 20 buying-intent prompts, pulled the same way.

Every engine is measured, and labeled by family

Chat engines and AI-generated search results are different surfaces with different mechanics. We measure both, but we never merge them into one blended score without saying which family each number came from.

Runs are counted, and shown as a range when there is more than one

One run is a snapshot; several runs are a range. We say which one you're looking at every time, from our own published research to the report you'd get as a customer.

Two families, never blended

Chat Engines and Search-AI Surfaces Are Not the Same Measurement#

A citation inside a ChatGPT conversation and a citation inside a Google AI Overview happen through different mechanics. Reporting them as one number would hide which one actually matters for a given buyer journey. We keep the two apart, always labeled.

Chat family

ChatGPT, Perplexity, Gemini, and Claude. You ask a direct question and get a conversational answer, sometimes with a cited source. Model sampling has real randomness here, so this family gets measured multiple times per cycle.

Search-AI family

Google AI Overviews, Google AI Mode, and the AI answer panel inside Bing. An AI-written summary embedded directly in a familiar results page, not a chat conversation. It behaves close to deterministically for a fixed query, so it needs fewer runs, but it's still its own number.

The 7 surfaces

Every Score Is Labeled By the Surface It Came From#

Two families, seven surfaces, one public method label attached to every number, so nobody has to guess what produced it.

ChatGPT

measured · native chat. Segmented by whether the answer used live web grounding.

Perplexity

measured · native chat

Gemini

measured · native chat

Claude

measured · native chat

Google AI Overviews

measured · AI-answered search

Google AI Mode

measured · AI-answered search. English-language queries only.

Bing Copilot

proxy · AI answer in search. The AI answer panel inside search results, not the separate Copilot chat app.

The sampling tiers

How Many Times We Run Each Prompt, By Tier#

More repetitions produce a tighter, more trustworthy number, and cost more to run. We never let a cheap tier borrow the label a more rigorous one earned. Every score says which tier produced it.

Free scan

One run per prompt, on every engine.

single-run snapshot

Compete report

Five runs per prompt on chat engines, one to two on Search-AI surfaces.

report-grade, with a wide interval shown

Deep Research report

Seven to eight runs per prompt on chat engines, two on Search-AI surfaces, matching the floor set by the published GEO reproducibility survey.

research-grade

Confidence, not false precision

Every Repeated Score Ships With a 95 Percent Confidence Interval#

We compute the interval by bootstrap resampling the runs behind a score, not by assuming a bell curve that AI-answer data does not actually follow. A number without an interval is theater: it looks precise and tells you nothing about how far it could have landed somewhere else. At a single run, we do not fake an interval, we show the single-run-snapshot label instead, and say so.

A citation rate of 40% with a 95% interval of 10% to 70% is a very different claim from 40% with an interval of 35% to 45%, even though the headline number is identical. We show the width every time, never just the midpoint.

Selection vs Absorption

Getting Cited and Shaping the Answer Are Two Different Events#

Published research on LLM citation behavior found that inclusion happens in two stages with different drivers. Treating them as one blended number hides which one you are actually improving.

Selection

Did the brand enter the pool of sources the model drew from at all? Driven mostly by classic authority signals: backlinks, brand recognition, third-party mentions. This is what most published citation-rate numbers actually measure.

Absorption

Once selected, how much did the brand actually shape the substance of the answer, versus sitting unused in a source list? Driven by extractable evidence density: definitions, statistics, comparisons, and structured content a model can lift directly. Ranking well solves Selection. It does not automatically solve Absorption.

Within-tool, never across

We Never Compare an Absolute Score Between Two Different Engines#

Citation behavior varies enormously from one AI engine to the next. Published research measured swings of up to 615 times between engines on the same category, and found only about 11% overlap between the sources ChatGPT cites and the sources Perplexity cites for comparable queries. A change inside one engine, tracked over time, is a real signal. A comparison of one engine's raw score against a different engine's raw score is not a finding, it is noise dressed up as one.

The honesty rules

What "Measured" Actually Means Here#

What does a "single-run snapshot" actually mean?

It means the number came from exactly one pass, not an average of repeated runs. LLM outputs are not fully deterministic: ask the same question twice and you can get two different citations. A single-run number is real data, but it's a photograph, not a trend line, and we label it that way every time, including our own 240-response category baseline.

How many runs go into a number that isn't a single-run snapshot?

Chat-family engines are sampled multiple times per measurement cycle; Search-AI surfaces need fewer runs because they're closer to deterministic for a fixed query. Where we run more than once, we show a range instead of one flattering figure.

Do you ever combine chat-engine and Search-AI results into one score?

No. They're different surfaces, measured differently, and reported separately every time, labeled by family. A blended "AI visibility score" would look cleaner and tell you less.

Why won't you promise a ranking or a guaranteed citation?

Because nobody controls what a model chooses to cite, not us, not you, not even the lab that trained it. Anyone promising a guaranteed AI citation is promising something outside their own control. What we can guarantee is that the number we hand you was produced the way this page describes, not made up.

What we never do

Five Things You Will Never See From Us#

  • One blended "AI visibility" score averaged across different engines.
  • An absolute score from one engine compared directly against a different engine and presented as a ranking.
  • A number on this site without the method label that produced it: the engine, the tier, and the run count.
  • A guaranteed citation, ranking, or placement promise. Nobody controls what a model chooses to cite.
  • A missing data point silently treated as zero. Missing is labeled missing.

The research

Our Own Category Research, Line by Line#

In July 2026 we ran the exact methodology described above on our own category, before selling it to anyone. Here is what each module measured, pulled straight from the logged research runs.

Competitor keyword mapping

Every keyword 14 competitor domains already rank for: 1,721 unique keywords, the raw material for the mandatory-ranking queue.

Search results + AI Overview landscape

Who Google's AI Overview cites for the category today, sampled across five languages.

Citation baseline (the 240-response study)

The buying-intent prompt study behind every leaderboard number on this site: 240 scored answers across four engines.

Gap analysis

Cross-referencing all of the above into a validated content queue.

Same ruler

We Measure Ourselves With the Same Ruler#

We run the exact methodology described on this page against our own category, and publish the results as a dated, running case study, updated in place as the numbers move, or don't. If a method can't survive being applied to its own maker, it isn't a method.

Published, dated, and updated in place, not a retrospective
The state of the category: 1,721 keywords and 14 competitors, mappedThe citation leaderboard: who gets recommended todayOur own public case study, updated as the numbers move

Verify it yourself

You Do Not Have to Take Our Word For It#

Every number on this page comes from asking real AI engines real questions and counting what came back. Here is exactly how to run the same check yourself: no account, no tool, just the engines and a stopwatch.

Ask the same question

Take a prompt from a declared set, like the example below, and ask it to each engine you care about, exactly as written, in a fresh conversation with no prior context.

Count what gets cited

For each answer, note whether your brand or domain shows up, and which sources the engine names or links to. A single run is a snapshot, not a trend.

Compare within the same engine over time

Repeat the same prompt on the same engine on different days to see how citation behavior moves. Never compare one engine's raw count directly against a different engine's raw count.

Example prompt from our buying-intent setbest CRM software providers
See a real measured example

A real measured example is not published yet. It is queued from our own next self-scan and will replace this note once it lands.

Changelog

Every Measurement Change, Dated#

v1.1 July 19, 2026

Sampling tiers renamed for the one-time launch: Starter and Pro become the Compete and Deep Research reports, the monitoring-grade label becomes report-grade, and the Compete tier locks at five runs per prompt. Run counts, the confidence-interval method, and every honesty label are unchanged.

v1.0 July 18, 2026

Initial publication: sampling tiers with honest labels (single-run snapshot, monitoring-grade, research-grade), Selection and Absorption measured separately, within-tool comparison only, versioned prompt sets, and a method label attached to every surface.

References

The Research Behind These Rules#

Every rule on this page traces to a specific published source, not an internal opinion about what sounds rigorous.

Aggarwal et al., "GEO: Generative Engine Optimization," KDD 2024 (arXiv:2311.09735)

The original reproducibility baseline for this category: five random seeds, public code. Still not followed by most commercial tools.

Optimizing Visibility in Generative Engines: A Critical Survey, 2023 to 2026 (arXiv:2607.14035)

Sets the seven-to-eight-run floor behind our research-grade tier, and the separate estimands we sample for.

Zhang, He, and Yao, 2026 (arXiv:2604.25707), dataset "geo-citation-lab"

602 prompts and 21,143 citations behind the Selection versus Absorption split.

Seer Interactive, AI Overview click-through research (3,119 queries analyzed)

The organic click-through swing under AI Overviews that motivates measuring this category at all.

US Patent 12,158,907 B1, query fan-out for generative search

Why a single seed query is not enough: AI Overviews expand a prompt into sub-queries before answering it.

See where you stand, measured the same way.

Free scan, one run, labeled honestly. No credit card.

Run my free scan