Agency Operator Economics · 12 min read

Reporting GEO Results When Nothing Moved

Once you put an honest error bar on an AI visibility number, most months are flat. Here are the four conversations that creates, written out word for word, and the report structure that makes each one survivable.

1.5%of the variance in a single LLM brand answer is attributable to the brand itselfZatuchin, 12,933 LLM responses across 20 brands, 8 languages, 3 models
The short version
  • A 20 percent visibility reading taken from 250 sampled answers carries a 95 percent confidence interval of roughly 15.5 to 25.4 percent, so a move to 24 percent is not a result.
  • Brand identity accounts for just 1.5 percent of the variance in a single LLM brand answer, while pure resampling accounts for 34.8 percent.
  • Build your report in three layers: a census layer with no error bar, a sample layer with one, and an outcome layer you must label as the weakest.
  • The four hard calls are the move inside the band, the drop, the rise you did not cause, and the guarantee request. Write the sentences before you need them.
  • Repeating one prompt past five runs is almost worthless. Widen the prompt set instead.

The number on your dashboard is mostly not about your client

The short answer

Most months in a GEO engagement are genuinely flat, and the correct move is to say so on the call. A visibility reading of 20 percent taken from 250 sampled answers carries a 95 percent confidence interval of roughly 15.5 to 25.4 percent. A move from 20 to 24 is not a result. It is the instrument breathing.

Start with the most uncomfortable finding in the measurement literature. A July 2026 variance decomposition of 12,933 LLM brand answers across 20 brands, 8 languages and 3 models partitioned the variance in a single response. Brand identity accounted for 1.5 percent of it. Pure within-prompt resampling accounted for 34.8 percent. Query language accounted for 26.5 percent.

Read that in reporting terms. In one sampled answer, the wording of the prompt and the roll of the dice explain more than twenty times as much of the score as which brand is being measured. That is not an argument against measuring. It is an argument that one reading, or a small one, tells you nothing, and that a report built on a small sample is a report built on noise with a logo on it.

1.5%
of variance in a single LLM brand answer traces to the brand itself
1 in 1,000
runs before two AI answers return the same brands in the same order
61.7%
of AI citations never name the brand anywhere in the answer text

SparkToro and Gumshoe.ai ran 12 prompts 2,961 times with 600 volunteers and found it takes roughly 1,000 runs before two answers return the same brands in the same order. Position tracking inside an AI answer is not a metric. It is a coin flip you paid a vendor to photograph.

AIs do not give consistent lists of brand or product recommendations. If you don't like an answer, or your brand doesn't show up where you want it to, just ask a few more times.

Rand FishkinCo-founder, SparkToro

Build the error bar before you build the dashboard

Visibility percentage is a binomial proportion. Your brand either appeared in a sampled answer or it did not. That means the interval is computable, and the method is not exotic: the Wilson and Agresti-Coull intervals are the standard treatment in the NIST engineering statistics handbook, section 7.2.4.1.

Run the arithmetic on the three sample sizes practitioners actually use.

Observations per monthReading95% Wilson intervalBand width
50 (10 prompts, 5 runs)20%11.2% to 33.0%21.8 points
250 (50 prompts, 5 runs)20%15.5% to 25.4%9.9 points
1,000 (200 prompts, 5 runs)20%17.6% to 22.6%5.0 points

Most tools ship a default closer to the first row than the third. On a 50-observation month, your brand could sit anywhere between one answer in nine and one answer in three, and you would not be able to tell the difference. Quadrupling the sample halves the band. That is the whole lever, and it is the reason sample size is the first thing to settle in a GEO engagement.

Here is where I will contradict my own table. Those intervals are a floor, not a true band. The Wilson interval assumes independent trials, and five runs of the same prompt are not independent. They are clustered inside one wording, which is exactly the facet the variance decomposition found to be dominant. Your real uncertainty is wider than the arithmetic above. I publish the floor because it is checkable and because even the floor kills most reported movement.

The same paper explains where extra budget should go. Repeats past the fifth cut relative error variance by 0.0003, the smallest return of any facet studied.

Reliability is bought by breadth across languages and models, not by depth of repetition.

Dmitrij ZatuchinEstonian Entrepreneurship University of Applied Sciences and Rankfor.AI

Translate that into a line item. Five runs, many prompts, more than one engine. Not fifty runs of ten prompts. If your prompt set is built wrong, no amount of resampling rescues it, and most AI visibility tools will happily sell you the resampling anyway.

The three-layer report, and why the industry template is backwards

The standard GEO report you will find published today leads with share of voice, then business impact, then an action plan. That order is exactly wrong. It opens on the noisiest number in the document and closes on the only work you actually controlled.

Invert it.

The three layers, in reporting order
Step 01

Layer one, the census

Counts, not samples. Crawler hits by user agent from your logs, pages returning 200 to each bot, assets published, third-party placements earned. These have no error bar because nothing was sampled. Report them first.

Step 02

Layer two, the sample

The visibility reading, always printed with its interval and its sample size. Never a single figure. Never a trend line drawn through three points.

Step 03

Layer three, the outcome

Referral sessions, assisted conversions, closed revenue. Label this the weakest layer out loud, because it is.

Layer one is where your work lives and where the numbers are honest. If a crawler was getting a challenge page in March and gets a 200 in April, that is a fact, not an estimate, and it is the reason crawler access is the highest certainty lever in this whole discipline. Server log analysis is a census of every request that hit your origin.

Layer three deserves the warning label. Google's Search Console generative AI reports, rolled out to a subset of sites in June 2026, give impressions, pages, countries, devices and dates, and as Barry Schwartz noted at Search Engine Land, they do not include click data. AI Overviews, AI Mode and Discover generative features are counted in one combined view. That is the best first-party instrument that exists, and it still cannot tell you which surface produced anything. Anyone promising deterministic AI search attribution is selling you a number the platform itself does not publish.

Timeline of a twelve month GEO engagement showing the confidence interval corridor and the four client conversations that occur inside it
The interval figures are real Wilson arithmetic on a 20 percent reading at each stated sample size. The month markers are a template, not one client's data.Sources: NIST e-Handbook 7.2.4.1, Zatuchin arXiv 2607.13304, SparkToro and Gumshoe.ai, Ahrefs.
Use this graphic on your site

Free to republish with a link back to this page. Copy the embed code:

<a href="https://josephtimpson.com/insights/geo-client-reporting"><img src="https://josephtimpson.com/assets/infographics/geo-client-reporting.svg" alt="Timeline of a twelve month GEO engagement showing the confidence interval corridor and the four client conversations that occur inside it" width="1200" style="max-width:100%;height:auto"></a><p>Graphic by <a href="https://josephtimpson.com/insights/geo-client-reporting">Joseph Timpson</a></p>

Script one: the number moved, but not past the band

This is the most common call and the one that destroys credibility fastest, because the temptation to narrate a four point rise as momentum is enormous.

The move that makes this work is committing to the band in month one, in writing, before you know which direction the readings go. A threshold you set after seeing the data is not a threshold. It is a story. Put the interval in the scope of work alongside the deliverables so the client agrees to the rule while it is still symmetrical.

Script two: the number fell

The discipline here is separating the sampled number from the counted number. The sampled number fell inside its own noise. The counted number, third-party placements lost, actually fell. Report the second one as the finding, because where your citations come from is auditable in a way that a share of voice percentage is not.

Get the elephant in the room on the table immediately.

Wil ReynoldsCEO and Vice President, Seer Interactive

That line was written in February 2010, about rankings, and it is still the best sentence anyone has published on this. Which is a small indictment of a category that has produced a hundred dashboards and almost no language.

## Script three: the number rose and you did not cause it The hardest script, because the client is happy and you are about to make them less happy.

Book a working session

This is not hypothetical. Ahrefs published in March 2026 that only 38 percent of AI Overview citations come from top ten pages, against roughly 76 percent in its own July 2025 study. Author Louise Linehan notes in the same piece that Ahrefs improved its parsing methodology, and the article does not disentangle the measurement change from the behavioural one. That is a serious research team, publishing its own reversal, and even they cannot cleanly split the two causes. If they cannot, neither can your monthly dashboard, and any post still citing the 76 percent figure is quoting a number the publisher has itself replaced. It is also why "rank number one and you win AI" has to be stated as table stakes rather than as a mechanism.

Script four: the client asks for a guarantee

Google's live guidance is unambiguous. The Search Central page on whether you need an SEO, last updated 5 June 2026, states plainly that no one can guarantee a number one ranking and warns about firms that are secretive about what they intend to do. Handing the client that link does more for your credibility than any slide you could build. It is also the cleanest way to differentiate against the myths this category still runs on.

What belongs on the report that has no error bar

Every item below is a count from a log, a file or a URL. None of it is sampled, so none of it needs an interval, and all of it is defensible under audit.

The census layer
  1. Requests per AI user agent, from raw server logs, month over month
  2. Status codes returned to each agent, with any non-200 named and dated
  3. Pages newly reachable to crawlers this month, listed by URL
  4. Assets published or materially updated, with dates
  5. Third-party placements earned, lost or rewritten, by URL
  6. Brand facts corrected in answers, with before and after text captured
  7. Prompt set version, sample size and any change to either

That last line matters more than it looks. Changing the prompt set mid-engagement inflates the reading without any underlying improvement, and it is the single easiest way to accidentally lie to a client. Version the set, and if you widen it, re-baseline and say so.

Note also that a citation is not a recommendation. Semrush and Growth Memo logged 3,981 domain appearances across 115 prompts in 14 countries and found 61.7 percent were ghost citations, where the page was used as a source but the brand was never named in the answer. Reporting raw citation counts as brand presence overstates the result by roughly two and a half times, which is why citations and recommendations need separate rows and why how accurately the engine describes the brand belongs on the report at all.

Why the honest version is a commercial asset

The obvious objection is that this loses clients. The evidence points the other way, and the mechanism is not sentimental.

AgencyAnalytics surveyed more than 220 agency leaders and found 70 percent say client reporting plays a critical role in retention, with 43 percent reporting average client lifespans of two to five years. Reporting is not the paperwork around the retention problem. It is the retention problem.

Meanwhile the budget environment punishes vagueness specifically. The CMO Survey, fielded with 308 US marketing leaders in January 2026, found that when profits fall short, 53.1 percent of executives cut expenses first, and marketing gets cut 45.4 percent of the time, more than any other category. A retainer that cannot explain its own numbers is a line item waiting to be deleted.

Rather than investing in deeper customer insights, most marketers focus on developing stronger performance tracking as the primary way to demonstrate value.

Christine MoormanProfessor, Duke University's Fuqua School of Business, and Director of The CMO Survey

Moorman is describing the failure mode this entire post exists to prevent. Dashboards substituting for understanding. The escape is not a better chart. It is a report whose top layer is auditable and whose middle layer is honest about its own precision.

With generative AI, economic uncertainty, and evolving platform rules reshaping how agencies work, it's more important than ever to stay grounded in what actually drives results.

Joe KindnessCEO, AgencyAnalytics

One last piece of arithmetic to hand a client who expects AI referral traffic to appear in analytics. Ahrefs measured, across 76,000 sites in its Web Analytics panel, that Google sends 190 times more traffic to websites than ChatGPT. Your layer three sessions line is going to look like nothing for a long time. Say that in month one, and the month three call stops being a crisis.

If you are building or repricing this offer, the reporting rule belongs upstream of everything: in what you scope, in what you charge, in how you launch the service, and especially in what you hand a white-label partner, who will otherwise strip the interval off your number before it reaches the end client. The rest of the Cited Method is built on the same premise, and every other post in this insights library assumes it.

Frequently asked questions

How do I know if an AI visibility change is real or just noise?

Compute the confidence interval on the reading before you interpret it. At 250 sampled answers, a 20 percent visibility reading carries a 95 percent Wilson interval of roughly 15.5 to 25.4 percent. Any month-over-month move inside that corridor is statistically the same number.

How many prompts and runs should a GEO client report be based on?

Five runs per prompt, across as many distinct prompts and engines as the budget allows. The variance research shows repeats past the fifth reduce relative error variance by only 0.0003, the weakest return of any factor tested. Breadth buys reliability. Depth of repetition does not, so stop paying for it.

What should I report when the visibility number is flat?

Report the census layer instead: crawler requests by user agent, status codes returned, pages newly reachable, assets shipped, third-party placements earned or lost. These are counts from logs and URLs, so they carry no sampling error and hold up under audit.

Can Google Search Console tell me how I perform in AI Overviews?

Partially. The generative AI performance reports launched in June 2026 give impressions, pages, countries, devices and dates. They do not include click data, and AI Overviews, AI Mode and Discover generative features are combined into one view you cannot separate.

Is it safe to change the prompt set mid-engagement?

Only if you re-baseline and disclose it. Widening the set inflates the reading without any underlying improvement, which is the easiest way to accidentally mislead a client. Version the prompt set and print the version number on every report.

Should I report AI citations as brand visibility?

Not on their own. Semrush and Growth Memo found 61.7 percent of citations were ghost citations, where the page was used as a source but the brand was never named in the answer. Report citations and named mentions as separate rows.

What do I say when a client asks for a citation guarantee?

Decline, and point to Google's own hiring guidance, which states no one can guarantee a number one ranking. Guarantee the instrument instead: fixed prompt set, stated sample size, published interval, verified crawler access, named assets on named dates.

How often should GEO reporting run?

Sample monthly, but interpret quarterly. Monthly readings at realistic sample sizes rarely clear the confidence band, so treat the monthly report as an activity and access record, and reserve any trend interpretation for a full quarter of accumulated observations. Agree that split with the client in month one.

Sources

  1. arXiv (Dmitrij Zatuchin, Estonian Entrepreneurship University of Applied Sciences and Rankfor.AI). Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (2026-07)
  2. National Institute of Standards and Technology. NIST/SEMATECH e-Handbook of Statistical Methods, 7.2.4.1 Confidence intervals (2013-10)
  3. SparkToro with Gumshoe.ai. New Research: AIs are highly inconsistent when recommending brands or products (2026-01)
  4. Semrush with Growth Memo. The Ghost Citations Study (2026-06)
  5. Ahrefs. Only 38% of AI Overview citations come from pages ranking in the top 10 (2026-03)
  6. Ahrefs. ChatGPT has 12% of Google's search volume, and Google sends 190x more traffic (2026-02)
  7. Google Search Central. Do you need an SEO? (2026-06)
  8. Search Engine Land (Barry Schwartz). Google Search Console AI performance reports and controls to block your content in AI responses (2026-06)
  9. AgencyAnalytics. 2025 Marketing Agency Benchmarks Report (2025-09)
  10. Duke University Fuqua School of Business, Deloitte and the American Marketing Association. The CMO Survey, Highlights and Insights Report, Spring 2026 (2026-04)
  11. Duke University Fuqua School of Business. CMOs Face Headwinds Even as Marketing Value and AI Impact Grow (2026-03)
  12. Seer Interactive (Wil Reynolds). Avoiding Client SEO Failures, our near huge mistake (2010-02)
Joseph Timpson
Written by
Joseph Timpson

Joseph Timpson has worked in search since 2010 and runs Timpson Marketing out of St. George, Utah. He built The Cited Method, a five stage framework for earning and proving real citations in AI answers, and publishes what does not work alongside what does.

Want me to run this on your site and show you the before and after?

One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.

Book a free consultation