// Methodology · Answer Monitor

Most AI visibility scores are a single coin flip.

Ask an AI engine the same question five times and you'll get different sources each time. A dashboard that reports one number from one run isn't measuring your visibility — it's sampling noise and calling it a metric. Here's how we do it instead: like polling, with the error bars shown.

// The field is bifurcating

Rigorous, or convenient. Pick one.

The research is blunt about it: identical prompts return wildly different citations run to run, and the major engines swap out most of their cited sources every week. That means monthly rank-style tracking is worse than useless — it reports last month's noise as this month's trend. The measurement tools splitting the market into the ones that sample properly and the ones that ship a reassuring single number. We planted our flag on rigor.

// The standard

Six commitments we don't break.

Polling, not rank tracking

AI visibility is a distribution you sample, not a position you read. The same prompt returns different sources on different runs — so we measure it like an election poll, with repetition and error bars, not like a keyword rank.

Per engine, never blended

ChatGPT, Perplexity, Gemini and Google AI Overviews disagree constantly — most citations appear on only one of them. A single blended "visibility score" hides the only thing that matters: which engine, for which prompt. We never average them into one number.

Five-plus samples per prompt

Within-model variance on an identical prompt runs high, and only a small fraction of citations survive being asked the same thing three times. One run is a coin flip. We sample every prompt at least five times before we report anything.

Confidence intervals on every number

If a number has no error bar, it is a guess wearing a suit. Every figure we report carries the range around it, so you can tell a real move from sampling noise.

Persona-segmented prompts

A cautious first-time buyer and a technical evaluator ask different questions and get different answers. We segment prompt sets by persona instead of pretending your market is one voice.

Journeys, not one-shot prompts

Buyers ask follow-ups. We track five-stage buyer journeys as single conversations — Problem → Exploration → Comparison → Validation → Selection — because that is how the answer engines actually build a recommendation.

// How we prompt

We track the whole buyer journey, as a conversation.

One-shot prompts miss how recommendations actually form. We run each persona through five connected stages and watch where you enter the answer — and where you fall out.

01

Problem

"Why do my crypto transactions keep failing?" — the moment before your category is even on the table.

02

Exploration

"What are the best options for X?" — where the model builds its shortlist, and where category ownership is won or lost.

03

Comparison

"X vs Y" — head-to-head, where sentiment and framing decide the winner.

04

Validation

"Is X legit / audited / safe?" — where a wrong or stale fact quietly kills the deal.

05

Selection

"How do I get started with X?" — where the model either hands the customer to you, or to a competitor.

// What lands in your report

Not one score. The full picture, with its uncertainty.

Per engine, per persona, per journey stage: how often you appear, where you rank in the shortlist, whether you're recommended ahead of competitors, the sentiment and framing around you, and whether the facts the model states are actually true — each with the confidence range around it. Mention, citation, recommendation, and accuracy are four different things, and we report them as four different things.

Measurement you can defend to your board.

Answer Monitor applies this standard to crypto and fintech brands first. Request a monitored panel and we'll baseline your category with the error bars shown.

Request a monitored panel →

Our approach is informed by published research on AI-answer variance — notably Kevin Indig's Growth Memo and its data partners (AirOps, SISTRIX). The measurement standard and its application are our own.