GeoHero
Data & Benchmarks

Which Brands Do AI Engines Actually Recommend? We Asked 240 Times to Find Out

By GeoHero7 min read

When someone asks ChatGPT, Claude, Gemini, or Perplexity to recommend an AI-visibility tool, the answer isn't random and it isn't evenly distributed. We ran 20 buying-intent prompts, questions like "what are the best AI visibility tools" and "what tools track my brand's visibility in ChatGPT", across all four engines, in English, Portuguese, and Spanish, for 240 total answers, and counted exactly who got cited. Semrush leads the category at 33.3% of answers. The next four names, Profound, Otterly.AI, Peec AI, and Ahrefs, sit between 19.6% and 25.0%.

This is a baseline, not a claim of causation, and it's a single measurement run rather than an averaged score with a confidence interval, a point we come back to in the caveats section. But nobody else has published this specific number for this specific category, which is exactly why it's worth publishing.

Methodology, in Full

Twenty buying-intent prompts, covering questions like "what are the best AI visibility tools," "what tools track my brand's visibility in ChatGPT or Perplexity," "how do I measure my AI search visibility," and "best alternatives to Semrush AI Visibility Toolkit / Profound / Peec," were run through four engines (OpenAI's ChatGPT, Claude, Gemini, and Perplexity), in three languages: English, Portuguese, and Spanish. That's 20 prompts × 4 engines × 3 languages = 240 total answers, captured in July 2026. Every response was scanned for brand mentions using two signals together: an explicitly cited source domain (the harder, more reliable signal) and the brand's name appearing in the answer text via a name-matching pattern (a softer signal, with a known limitation covered below). This is the same scoreboard-of-record we track for our own site, we run this exact methodology on our own domain first, a loop described in our prompt monitoring guide.

The Leaderboard

Ranked by share of the 240 total answers that cited each player:

  • Semrush: 33.3% (80 of 240 answers)
  • Profound: 25.0% (60 of 240)
  • Otterly.AI: 23.3% (56 of 240)
  • Peec AI: 19.6% (47 of 240)
  • Ahrefs: 19.6% (47 of 240)
  • SE Ranking: 15.4% (37 of 240)
  • Scrunch AI: 7.5% (18 of 240)
  • Writesonic: 5.0% (12 of 240)
  • Geoptie: 5.0% (12 of 240)
  • AthenaHQ: 3.3% (8 of 240)
  • LLMrefs: 2.9% (7 of 240)
  • Rankscale: 2.9% (7 of 240)
  • Bluefish AI: 2.1% (5 of 240)
  • Conductor: 1.7% (4 of 240)
  • Knowatoa, Gauge, Gumshoe: 0.4% each (1 of 240)

Two things stand out immediately. First, Semrush's lead is real but not overwhelming. It's cited in one in three answers, which means two in three answers don't cite it, leaving real room for a specialist tool to be the recommendation instead. Second, the four names right behind Semrush, Profound, Otterly.AI, Peec AI, and Ahrefs are AI-visibility-native tools clustered close together, not a single runaway winner. The category has a leader, not yet a monopoly.

The Real Story Is Per-Engine, Not the Blended Average

The blended leaderboard above hides something the per-engine breakdown makes obvious: the four engines don't agree with each other, and the disagreement is large.

Perplexity is already deep into the specialist category. Semrush still leads at 43%, but Otterly.AI (42%) and Profound (40%) are essentially tied with it, and Peec AI sits at 28%. Perplexity, in other words, treats AI-visibility-native tools as legitimate default answers nearly as often as it does the generalist incumbent.

Gemini looks similar. Semrush 38%, Otterly.AI 32%, Profound 30%, Peec AI 28%. Specialist tools are close behind the category leader here too.

Claude sits in the middle. Semrush 35%, then Profound and SE Ranking tied at 23%, Ahrefs at 18%, Peec AI at 15%, a mix of generalist and specialist names, with more spread between them than Perplexity or Gemini show.

OpenAI's ChatGPT is a different world entirely. Semrush 17%, Ahrefs 13%, and every AI-visibility-native tool falls under 8%: Otterly.AI 8%, Profound 7%, Peec AI 7%. ChatGPT, at the time of this scan, still defaults overwhelmingly to generalist SEO brands it presumably has more training exposure to, and hasn't "caught up" to the specialist category the way the other three engines have.

The practical takeaway: if you're trying to be the AI-recommended tool in this category, the opportunity is not evenly distributed across engines. Perplexity and Gemini have already made room for specialist brands next to the incumbent. ChatGPT, by far the highest-traffic of the four, has not, which makes it simultaneously the hardest engine to break into today and the one with the most headroom for whoever establishes clear authority content first.

It Also Varies by Language: More Than You'd Expect

Restricting the same 240 answers by language instead of by engine shows another gap:

  • English (80 answers): Semrush 40%, Otterly.AI 33%, Profound 31%, Peec AI 28%, Ahrefs 21%
  • Portuguese (80 answers): Semrush 35%, Profound 23%, Otterly.AI 19%, Ahrefs 19%, Peec AI 16%
  • Spanish (80 answers): Semrush 25%, Profound 21%, SE Ranking 20%, Otterly.AI 19%, Ahrefs 19%

Semrush's own citation share swings from 25% in Spanish to 40% in English, a 15-point difference on the identical brand, identical category, different language only. That's not noise; it's a real signal that the models have different amounts of training exposure to this still-young category depending on language, and it means a citation strategy built only around English-language content is, by definition, leaving the Portuguese and Spanish-language answer space to whoever shows up there instead.

What This Means If You're Evaluating AI-Visibility Tools

If you asked ChatGPT which AI-visibility tool to buy, you'd get a meaningfully different-shaped answer than if you'd asked Perplexity. Not because one engine is "wrong," but because they've converged on different defaults from different training data at different points in time. Comparing your own experience across only one engine and generalizing to "what AI recommends" is a mistake this data makes obvious. It's also, separately, one of the exact use cases prompt monitoring exists to catch on an ongoing basis rather than as a one-time curiosity.

What This Means If You're Building GEO Strategy

Zero percent is not a failure state, it's a starting measurement, the same way a technical SEO audit's first crawl report is never "done," it's a baseline. Two things follow directly from the data above. First, ChatGPT is the engine where the category is least consolidated, which means it's the one where new, well-structured authority content has the clearest runway to become a default answer, the opposite of Perplexity and Gemini, where specialist brands are already established defaults you'd be competing directly against. Second, a citation strategy that only targets English content is measurably incomplete given how much Semrush's own share moves by language alone; if your buyers search in Portuguese or Spanish, that's a separate battle with a separate baseline, not an afterthought.

Caveats: Read Before You Cite This Number Anywhere

In the interest of the same honesty we're asking AI answer engines to apply to citations: this is a single measurement run, not an average across repeated runs with a confidence interval. AI model outputs vary from run to run even with an identical prompt, so treat every percentage above as a snapshot, not a permanently stable score. We plan to re-run this with multiple repetitions before treating any percentage as fully stable for a positioning decision, and we'll publish the update when we do. Brand detection combines an explicitly cited source domain (the harder signal) with name-matching in the answer text (softer), common-word brand names can, in principle, produce a false positive on the name-matching signal alone, though the domain-citation signal is unaffected by this. And citation in an AI answer is an observational signal, not proof of causation or of product quality. It reflects what these four models currently associate with the category, which is itself shaped by everything from publish date to training cutoff to how much content a company has published, not solely by product merit.


GeoHero applies this exact methodology to its own GEO work and tracks the result in the open, see how to run a GEO audit and how to set up your own prompt monitoring. For the underlying concept this entire category is built on, start with What Is Generative Engine Optimization (GEO)? Curious where your own brand currently lands on a leaderboard like this one? Run a free scan and find out in about two minutes.

Topics

  • how ai engines pick brands
  • ai search visibility tools
  • ai brand monitoring
  • which ai cites brands
  • generative engine optimization statistics