AI Visibility Monitoring
Across 12 AI models, did your brand move since the last period?
The same questions, asked once in Chinese and once in English, re-run on a fixed schedule. Every number comes with its margin of error, and we call something a change only when it clears that margin. We do not sell optimisation.
Three metrics, reported separately
Each metric shows how it is computed, right underneath. We never fold them into a composite score, and every number traces back to the raw answers of that period.
Mention rate
Of all valid answers, the share that name your brand.
Answers naming you ÷ valid answers, with a 95% Wilson interval
Average position
When you are named, where you sit in the model's list.
Sum of your positions ÷ answers naming you (first place = 1). Computed only over answers that name you; not reported below 20 mentions
Model coverage
How many of the 12 models name you consistently across repeated asks.
Models that name you in at least 50% of their own answers ÷ models that returned valid answers this period
The denominator is every valid answer, including answers that name no brand at all. Mention rate therefore answers: each time someone asks, how likely is your name to come up?
How many answers each period buys
Valid answers per period (maximum) = 5 scenario questions × 7 phrasings × 5 repeats × 12 models × 2 languages = 4,200
One scenario in one language is 35 asks per model — 420 answers on a 12-model panel. Near a 5% mention rate, the 95% interval is about ±2.1 percentage points (Wilson), and it widens as the rate approaches 50%. Failed calls are logged, left out of the denominator, and the report states the actual valid count.
Two modes, reported apart
The same questions can be asked two ways. Each is measured on its own, reported in its own table, never subtracted from the other and never merged into one score.
Model-only panel
- Measures
- How a model answers offline, from its training data alone — where your brand sits in its memory.
- Does not measure
- Whatever was published online this week. It barely moves between model releases, so it tracks long-run standing, not the short-term effect of a campaign.
- Used for
- Public rankings, certificates, scenario collections, and the baseline of every client project.
Web-enabled panel
- Measures
- How a model answers when it can search the web. Recent content moves it, which makes it the layer your GEO work acts on directly; every cited URL is stored.
- Does not measure
- What you see in a consumer app. Apps add account memory and location, so web-enabled results are not what any one person sees in a chat window either.
- Used for
- Client projects only. Never in public rankings, never on a certificate.
- What keeps it neutral
- The measurement parameters are frozen before your campaign starts, and their hash goes into the report. We take no money from any GEO provider. Every cited URL is listed in full in the report.
One-off web-enabled demoIn a one-off, small-sample AI-glasses demo, the same question with and without web search shared only five or six of the top-10 brands on average, and Chinese answers with search often copied a single ranking page almost verbatim. With search on, a model reads a few pages first and its list follows them; without it, answers reflect long-run impressions from training data. The two measure different things, which is why public rankings use the model-only panel and the web-enabled panel is used only in client projects.
The report states the two results side by side — for example, "model-only panel: no significant change; web-enabled panel: significant increase, with k answers citing URLs you submitted." We do not attribute the change to your campaign: competitors publish in the same window, and search indexes change on their own.
What each period delivers
- 01
A verifiable archive
Every number drills down to that period's raw model answers. Once issued, the archive does not change.
- 02
A three-way verdict
A change in mention rate between two periods gets one of three verdicts: significant increase, no significant change, significant decrease (two-proportion test at 95% confidence). If the phrasings, the panel structure or a methodology change affecting that cell differ, there is no verdict — only a note saying why.
- 03
Model-change markers
Trend charts carry a vertical line at every panel or methodology change. Numbers on either side of a line are not compared directly.
- 04
Exclusions and known biases
Which model failed on which day, how it was handled, which spellings were merged into one brand, and which kind of question the set leans towards — itemised at the end of every report.
A method note under every chart
The method sits directly below each chart, so a reader never has to go looking — and it travels with the chart when someone shares a screenshot.
Based on answers from {models} models to {scenarios} scenario questions, {phrasings} phrasings each, asked {repeats} times per phrasing, computed separately for Chinese and English. Collected {from} to {to}, panel snapshot {version}. {N} valid answers per brand; intervals are 95% Wilson intervals. Every raw answer is listed at {verify link}.
What models say about you
When a model talks about you, which aspects of your brand does it bring up — and does it mention drawbacks when it does? This is an add-on to monitoring. The approach:
- Two kinds of question, counted separately. One describes a use case and names no brand, to see which brands a model thinks of and which it picks first. The other asks about your brand by name, to see which aspects the model brings up unprompted. The two sets of numbers are never added together or mixed in one table.
- Every answer is coded sentence by sentence against a list of aspects. The list is frozen before the first period and its hash recorded; every later period uses the same list. Changing it means a new version and a new baseline, with no comparison across versions.
- Each aspect is reported as a mention rate (k/N) with a 95% interval, plus the share of those answers that mention a drawback. Aspects mentioned fewer than 20 times get the mention rate only. There is no overall sentiment score.
- About 20% of answers, drawn at random, are coded again independently by a second coder; agreement and Cohen's κ are reported per aspect, and any aspect with κ below 0.7 gets no numbers.
- We report what models said; what you publish is up to you.
Sample reports
Two complete samples, no email required. One is built from our published weekly data: each brand's rank over time, how many models named it, 95% intervals, panel-change markers and the exclusions box. The other is an aspect-analysis demo on AI glasses: which brands come up in a use case, and what models say when asked about a brand by name.
Questions
- No significant change between two periods — is that normal?
- It is the most common result. Ask the same questions of the same models again and the numbers move by a few percentage points on their own; we only call it a change when it moves further than that.
- Why do I get different answers in the app?
- We ask through each model's developer API: no personal account, no chat memory, no location. What we measure is the model's default knowledge. Apps add web search, account memory and your region to what the model says. When web answers are what you need, we run a separate web-enabled panel and report it apart.
- Will you help me improve my numbers?
- No. We only measure. The measurement parameters are frozen before your campaign starts, and afterwards we re-measure with the same ruler.
Independence and contact
This service runs on the same method and rules as the public rankings. Four pages you can check against each other:
- How we measure → Methodology
- What we have declined → Refusal log
- What we got wrong → Corrections
- Can every number be checked → Verify
The algorithm and panel configuration are published on the methodology page. You are welcome to replicate the measurement from it.
We do not sell optimisation and we do not promise to raise your recommendation rate. Rankings cannot be changed for money.