← Signal

The Fixed Set Method: How to Measure AI Visibility Without Lying to Yourself

Ask any AI visibility tool the same question twice and you may get two different answers. Ask it next month, after someone quietly reworded the prompts, and the number will have moved for reasons that have nothing to do with your brand. Most AI visibility measurement is built on sand, and the vendors selling it either do not know or do not say.

The Fixed Set Method is our answer to that problem. It is the way Zion Labs measures generative engine optimization so the numbers survive contact with a second look. This post names the method, defines the terms it rests on, and argues each of its four planks from work we have actually run.

What generative engine optimization and answer engine optimization actually mean

Before you can measure AI visibility, you have to be precise about what you are measuring. Two terms carry the weight.

Generative engine optimization (GEO) is the practice of making your brand the source that AI answer engines cite, quote, and recommend when they answer your market’s questions. Where classic SEO earns you a ranking in a list of blue links, GEO earns you a place inside the answer itself: the paragraph ChatGPT writes, the sources Perplexity lists, the summary Google shows above the results. It is measured by citation share and answer accuracy, not by keyword rankings.

Answer engine optimization (AEO) is the practice of being the single direct answer to a question, whether that answer is a Google featured snippet or a reply from an AI assistant. AEO sits one layer below GEO. SEO earns a position, AEO earns the answer, and GEO earns a citation inside a synthesized answer that blends several sources at once. We break the three apart in full in AEO vs GEO vs SEO, and define GEO from the ground up in What Is Generative Engine Optimization.

The reason this matters for measurement is simple. All three disciplines share the same hard question: when the results page is a generated paragraph that reads differently every time, how do you get a number you can trust? That is the whole problem the Fixed Set Method exists to solve.

Plank one: a fixed prompt set

The first rule is the one the whole method is named for. The questions never change.

An AI visibility measurement is only as stable as the prompts behind it. If you baseline your brand against fifty questions this quarter and forty different questions next quarter, any movement in the score is meaningless. You cannot tell whether your visibility improved or the test simply got easier. The prompt set is the ruler. You do not get to swap the ruler between measurements and call the difference growth.

A number that moves because someone changed the questions is a story, not a measurement. This is why the prompt set is fixed at the start of an engagement, written down, and reused verbatim on every run after. When we add a question, it starts its own baseline from zero rather than being folded silently into an existing score. The set is versioned, because the alternative is a number that drifts for reasons nobody can name.

This is unglamorous, and it is exactly why most tools skip it. A live demo looks more impressive when the questions adapt to flatter the result. A measurement you can defend six months later looks like discipline instead.

Plank two: multi-run measurement

The second rule follows from a fact about the systems themselves. AI answer engines are non-deterministic. Ask the same fixed question three times and you will often get three different answers, citing different sources, sometimes contradicting each other.

That is not a bug you can engineer around. It is the ground you are standing on. Any tool that asks each prompt once is reporting a single sample from a distribution it never looked at, and calling that sample a finding.

We ran the experiment. In one scan of a client’s brand across ChatGPT, Perplexity and Gemini, running every prompt three times, we found 18 contradictions between what the engines said and the client’s own fact sheet. 7 of those 18 appeared in only one of the three runs. A single-pass tool would have sold all seven as findings, shipping noise at a 39% rate without knowing it. Running each prompt several times is what let us separate the contradictions that repeated from the ones that were a coin landing tails once.

The same discipline caught a second trap. Providers sometimes answer without searching at all, and a run with no retrieval is not a measured zero, it is a missing observation. In that scan, 15 of 108 runs returned nothing because the model answered from memory, and 12 of those were ChatGPT, a third of its sample. Counting those as absences would have understated the client’s visibility. You only find that out by running the set enough times to see it happen.

Multi-run is what turns an anecdote into a measurement. One run tells you what an engine said once. Several runs tell you what it tends to say, and how much it wobbles, which is the thing a brand actually needs to know.

Plank three: error tracing

Finding that an engine says something wrong about you is half a result. The half that matters is knowing why, because that is the half you can act on.

Nobody can edit a model. We control every input it reads. You cannot call OpenAI and ask them to correct their weights. What you can do is find the page the engine pulled from and fix that. So every wrong answer in a Fixed Set measurement is traced back to the source that caused it, using the citations, footnote numbers and inline links the engines attach to their own answers.

When we can trace it, the fix becomes concrete. In that same scan, one contradiction traced straight back to a page on the client’s own domain, a page they could edit that afternoon. That is the difference between a report that says “the AI is wrong about you” and one that says “the AI is wrong about you because of this paragraph on this URL, and here is the change.” One is a complaint. The other is a work order.

When we cannot trace it, the report says so in those words. An honest gap is worth more than a confident guess about which of five listed sources was really responsible. Blaming whichever source happened to be cited first is how you send a client to edit a page that was never the problem.

Plank four: the Retrieval Readiness Score

The first three planks measure what the engines say about you. The fourth measures whether your page is even built to be used.

An engine can only cite a page it can retrieve, read, and lift a clean answer from. Plenty of pages fail at that gate before citation is ever on the table: they block the AI crawlers, bury the answer under three paragraphs of preamble, or fragment into chunks that make no sense out of context. The Retrieval Readiness Score is our nine-check scan of a single page against exactly those failure modes: is it retrievable, is it machine-readable, does it answer the question first-party and early, and does it resolve as a clear entity.

It is already built and it is free. Run any URL through it from the tools page and you get the nine checks back with the specific reasons a page is hard to use, not a vague grade. Readiness does not, on its own, create demand for you, and we have published our own finding that it does not predict citation rate once brand size is held constant. What it does is remove the reasons an engine cannot use your page, which is a precondition for everything the other three planks measure.

How the four planks fit together

Read as a loop, the method is straightforward. A fixed prompt set gives you a ruler that does not move. Multi-run sampling gives you a reading you can trust instead of a single lucky or unlucky pull. Error tracing turns each wrong reading into a specific page to fix. The Retrieval Readiness Score checks that the page you are about to fix is built to be used at all. Then you run the same fixed set again and watch the number move for a real reason.

Every plank exists to close a specific way of fooling yourself. Change the questions and you fake progress. Run each once and you ship noise. Skip the trace and you have a complaint instead of a fix. Skip readiness and you optimize a page an engine cannot even read. The method is nothing more than refusing to do any of those four things.

Why this is the honest way to measure

The uncomfortable truth about AI visibility measurement is that it is easy to produce a number and hard to produce a number that means anything. A shifting prompt set, a single run, an untraceable finding, and a page nobody checked for readiness will together give you a dashboard that looks alive and tells you nothing. It will move. It just will not mean what you think it means.

The Fixed Set Method is slower and less flattering by design. It fixes the questions, runs them enough times to see the variance, traces every error to a page you can edit, and checks that the page is built to be retrieved before anyone touches it. That is the whole discipline, and it is why our numbers hold up when a client’s own analyst reruns them.

Where to start

If you want this run against your own brand, our AI Visibility Audit applies the full method: a fixed prompt set built for your market, multi-run measurement across the major engines, every contradiction traced to a source page, and a readiness pass on the pages that matter.

If you would rather start on your own, the free tools are open with no login. Run the Retrieval Readiness Score on your most important page first. It is the cheapest way to see whether your page is even in the game before you measure how often it gets named.