AI book for children AI book for teens AI book for families New series Have you seen our book series yet? AI for kids, teens and adults Discover the books

Measure AI Visibility: Method, Not a Screenshot

geo ai-visibility measurement

A customer asks an AI assistant for a provider of your service. Do you show up in the answer? And if so, how often, for which questions, on which platform? That question has an answer, but not one you reach with a screenshot and a hunch. It needs a method.

The reason is plain: a language model rolls the dice on every response. If you want to measure visibility seriously, you build that variance into the method instead of ignoring it. Here is what that looks like in practice.

Why a single screenshot proves nothing

Language models do not answer deterministically. Ask the same question five times and you often get five slightly different answers, sometimes with entirely different company names. A screenshot freezes one of those outcomes. It tells you that you were named once, or missing once, and nothing more.

That is too little to decide on. Appearing in three of ten answers is a very different situation from missing in nine of ten, and both can produce the same screenshot. The only defensible statement is a frequency across many runs. Anything below that is an anecdote.

Three levels: mention, citation, recommendation

Before you measure, decide what you are counting. "The AI knows us" can mean three things, each with a different business value.

  • Mention: your name appears in the answer, with no source. The model saw you during training.
  • Citation: the answer points to a verifiable source, usually your website or an article. This happens mostly with web search enabled.
  • Recommendation: the AI actively presents you as the fitting choice. This level drives the most revenue and is the hardest to earn.

Blend these three into one metric and you measure past the point. Being named but never recommended is a different problem from not appearing at all. We covered the diagnosis behind that in ChatGPT recommends a competitor, and the first self-test in Does ChatGPT mention my company. This piece goes one step further: from a spot check to a systematic measurement.

From a single question to a query battery

A good measurement does not rest on one question. It rests on a battery of questions that mirrors how customers actually search. A single phrasing only ever hits one slice.

Cover the intent types people use to approach a provider:

  • Plain search: "Who offers service X in region Y?"
  • Comparison: "Which providers of X exist and how do they differ?"
  • Problem solving: "I have problem Z, who can solve it?"
  • Follow-up: "Is company A a good choice for X?"

Phrase each intent several ways and in your customers' language. A German firm gets asked in German, an exporter also in English or Spanish. Answers differ noticeably by language, because the underlying sources differ.

The brand usually does not belong in the prompt

The most common measurement error is writing your own name into the question. "What does company A do?" gets a friendly answer from almost any model as soon as company A is documented somewhere. That only measures whether the model can produce something about a known name, and it usually can.

The commercially important question is different: are you named when nobody supplies your name? Only the brand-free question reflects the situation of a real prospect who does not know you yet. A small share of brand-bearing questions may sit in the battery to test the citation level, but the backbone stays neutral.

Web search on or off changes everything

Whether a model searches the live web or answers from training changes the result at the root. With web search on, the answer pulls in freshly retrieved pages, and what counts is who is well findable today. With web search off, the model answers from what it saw during training, a state months old.

That creates a practical trap. The chat interface of many providers has web search on by default, the API often does not. Measure conveniently through the API and you see a different picture than the customer asking in the app. Keep the two modes apart and record, per run, which was active. Otherwise you compare apples to oranges and puzzle over the contradiction.

Frequencies instead of snapshots

The output of a clean measurement is not a pile of screenshots. It is a table of shares. For each question and each platform you record in how many of N runs you were mentioned, cited or recommended.

Visibility then becomes a number you can compare: over time, across platforms, against the competitor. Improve a source, measure again, and you see whether the share rises. One good hit in April and one bad one in August tell you nothing. The trend of the frequency tells you plenty. More on this setup sits on our AI visibility page.

What the method cannot do

Honesty is part of the method. A measurement gives you a current state with variance, not a forecast and not a guarantee.

  • It cannot promise a placement. Models and sources change, and every number is a snapshot.
  • It captures only what you query. A question no customer asks measures nothing useful, so everything hinges on the choice of battery.
  • It samples, it does not survey every possible conversation. More runs shrink the uncertainty, they never remove it.

These limits are not a flaw. They are the condition under which the numbers mean anything. Hide them and you are selling a certainty that does not exist.

If you want to know how often the big AI systems already name you rather than the competitor, a pilot project is the fastest route to real numbers. How measurement pairs with local, data-sovereign processing we show under local AI. Through the contact page we will look at your case.

Frequently asked questions

Why is a screenshot of an AI answer not proof?

Because language models do not answer deterministically. The same question can name different companies minutes later. A single screenshot shows one of many possible outcomes, not the rule. Only many repetitions reveal how often you actually appear.

Why should my brand name stay out of the prompt?

If you put your own name in the question, you only measure whether the model can say something about a known name, which it almost always can. The useful question is the one without your brand: are you named when a customer asks neutrally about the service.

What does it mean that API answers differ from the web interface?

The chat interface of many providers enables live web search, the API often does not. With web search the answer comes from freshly retrieved sources, without it from training. Measuring only through the API shows a different picture than the customer sees in the app.

Can a measurement guarantee placements?

No. A measurement describes the current state as a frequency and makes progress visible. Models change, sources change, so every number is a snapshot with variance. Anyone promising guaranteed placements has not understood the mechanism.

Share this article