What Is Actually Inside a GEO Retainer
Eighteen tasks, an hour figure on each one, and a written test for when each is finished. This is the document that decides whether month four happens.
- A deliverable list is not a scope of work. A scope of work is tasks, cadence, hours, the artifact each task produces, and one testable sentence per line that defines done.
- In GEO the acceptance criterion must attach to the artifact, never to the engine's behavior, because no operator controls whether ChatGPT names a brand on a given day.
- My month runs about 30 hours for a single site, single location client, rising to 33.5 in a quarter-end month. Roughly 40% of it goes to content and off-site work.
- Five published GEO service pages I read list deliverables. None lists hours. None writes an acceptance criterion. That gap is where the churn starts.
- llms.txt, schema-as-a-citation-lever and Markdown-for-crawlers are out of my scope entirely, and there is a controlled experiment or a vendor statement behind each removal.
The proposal says "ongoing optimization" because nobody wrote the work down
Two thirds of agencies are already taking this call. AgencyAnalytics surveyed 494 agency professionals between February and April 2026 and found 66% are fielding client requests for answer engine optimization, with 44% reporting that clients now expect faster turnaround than a year ago. Demand arrived before the delivery document did.
So I went looking for the delivery document. I read five published pages that claim to describe what is inside a GEO or SEO engagement: Superlines, Aruntastic, LLM Pulse, Humans With AI and First Page Sage. All five list deliverables. Not one lists hours. Not one writes a sentence that says when a line item is finished.
Deliverable lists are the easy half. "Visibility tracking" and "content refresh" name a category of activity, not a unit of work, and a category cannot be reviewed, disputed in good faith, or handed to a new account manager in month seven.
A GEO scope of work is a list of tasks. Each task carries a cadence, an hour estimate, the artifact it produces, and one testable sentence defining when it is done. Anything that fails that four part test is marketing copy, not scope. The acceptance criterion is the load bearing part, because it is the only clause a client can check without having been in the room.
There is a reason this category skips the acceptance criterion specifically, and it is not laziness. The obvious criterion, "the brand appears in the answer," is not something any honest operator can promise. That constraint is real. It is also solvable, and solving it is the whole subject of this document.
An acceptance criterion is the only enforceable sentence in the document
Danish Butt, Managing Director at Swiftwater and Company, names the failure mode in his guide to drafting professional services statements of work. His prescription is a clause type most marketing SOWs never contain: criteria for deliverable acceptance, defined as the specific conditions that must be met for each deliverable to be accepted by the client.
Scope creep is the most common source of consulting engagement disputes.
Product teams solved this a decade ago and marketing borrowed none of it. Agile Sherpas draws the line between a Definition of Done, which applies across everything a team ships, and acceptance criteria, which are unique to each individual work item. A GEO retainer needs both: one house standard for what "delivered" means, and a separate test per line.
Here is the move that makes this work in a category where the outcome is probabilistic. Acceptance criteria attach to the artifact, never to the engine's behavior. You control whether the crawler received a 200 and the full HTML body. You do not control whether Gemini names your client on Thursday. Write criteria about the first kind of thing and the document is enforceable by both sides. Write them about the second and you have quietly signed a guarantee you will lose.
That single rule is what the published templates are missing, and it is why they stop at deliverable names. Once you accept that the engine is not yours to promise, the acceptance criteria write themselves.
| What the proposal usually says | What someone can actually review | |
|---|---|---|
| Crawler access | We optimize for AI crawlers | Every named agent returns 200 with the full HTML body on a live fetch, or the block is documented and signed off |
| Measurement | Monthly AI visibility tracking | Every prompt run at the agreed repetition count per engine, with run counts and dates printed on the report |
| Content | Ongoing content optimization | The URL is live, indexed, snippet eligible, and every major claim carries a named source |
| Off site | Digital PR and brand mentions | Each pursued placement has a sent date, a state, and either a resolving URL or a documented decline |
| Reporting | Transparent monthly reporting | Every figure carries its sample size, and the report names one metric that moved the wrong way |
The month, task by task
Below is a full month of delivery, banded by the five stages I run: ACCESS, MEASURE, MAP, EARN and PROVE. Cadence is either monthly or quarterly. The hour figures are mine, for a single site, single language, single location client with one decision maker. They are an allocation, not a benchmark, and I have not found a published study of GEO delivery hours that I would trust to correct them. Treat them as a starting shape and re-time your own month against them.
| Task | Cadence | Hrs | Artifact | Done when |
|---|---|---|---|---|
| Crawler access verification | Monthly | 1.5 | Agent access matrix | Every named agent returns 200 with the full HTML body on a live fetch, or the block is documented and signed off |
| Index and snippet eligibility sweep | Monthly | 1.0 | Eligibility sheet | Zero priority URLs carry noindex, nosnippet, data-nosnippet or a max-snippet limit in raw source |
| Log pull and AI agent segmentation | Monthly | 2.0 | Verified agent log extract | Every counted hit is IP verified against the vendor's published range, and spoofed hits are reported separately |
| Raw HTML render check | Quarterly | 1.0 | Source diff | The answer bearing text appears in the raw response, not only after JavaScript runs |
| Prompt set maintenance | Monthly | 1.5 | Versioned prompt file | Every prompt maps to a buying stage and a real query, and every addition or retirement carries a written reason |
| Multi run sampling across engines | Monthly | 3.0 | Raw run log | Each prompt run at the agreed repetition count per engine and locale, with counts and dates recorded |
| Citation versus mention split | Monthly | 1.0 | Two column visibility table | Every appearance is classified as cited, mentioned or both, and no blended visibility number is reported |
| Answer surface gap map | Monthly | 2.0 | Gap sheet | Every tracked prompt names the URL that currently owns the answer and the reason we are not it |
| Fan out subtopic coverage audit | Monthly | 2.0 | Subquery matrix | Each priority topic lists its subqueries with an owning URL or an assigned action |
| Third party source inventory | Quarterly | 1.5 | Ranked source list | Each recurring cited source has a route in and a current status |
| Content build or rebuild | Monthly | 6.0 | Published URL | The page is live, indexed, snippet eligible, and every major claim carries a named source |
| Refresh pass on decaying pages | Monthly | 2.0 | Per URL diff log | Substantive change, visible updated date, and the diff logged. A date bump is not a refresh |
| Off site placement work | Monthly | 2.5 | Placement tracker | Each placement has a sent date, a state, and a resolving URL or a documented decline |
| Entity consistency pass | Quarterly | 1.0 | Entity sheet | Name, category and description string match across the site, the listed profiles and the client's own materials |
| First party surface pull | Monthly | 1.0 | Raw platform exports | Search Console and Bing exports attached unedited, with the date range printed |
| AI referral segmentation | Monthly | 1.0 | Analytics segment | Referrers isolated by named host, and anything inferred is labelled an estimate rather than a count |
| Report build and written read | Monthly | 2.5 | Client report | Every figure carries its sample size and the report names one thing that got worse |
| Working session | Monthly | 1.0 | Decision log | Every decision has an owner and a date. Unowned items are not decisions |
That totals 30.0 hours in a normal month and 33.5 in a quarter-end month. The distribution matters more than the total: ACCESS, MEASURE, MAP and PROVE each take 5.5 hours, and EARN takes 11.5. If your month does not look roughly like that, you are either under-instrumenting or under-producing, and both show up in month four.
ACCESS: the 5.5 hours that decide whether the other 28 matter
Everything downstream is contingent on this stage, which is exactly why it is the stage most often skipped. Google's documentation states the prerequisite in one line: a page must be indexed and eligible to be shown in Google Search with a snippet to appear in AI Overviews or AI Mode. A stray nosnippet on a template zeroes out AI eligibility for every page using it, and no amount of content work later in the month recovers it.
Crawler access is not one checkbox either. OpenAI documents four separate bots and states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Anthropic documents three, and says plainly that blocking Claude-SearchBot may reduce a site's visibility in user search results. Perplexity documents two and admits that Perplexity-User generally ignores robots.txt because a human requested the fetch. Each vendor now publishes IP ranges, which is why my acceptance criterion says IP verified rather than user agent matched. A user agent string is a claim. An IP range check is a test.
The log task is the one clients push back on and the one I refuse to cut. Cloudflare's 2025 network measurement put AI bots at an average 4.2% of HTML requests across the year, with user-action crawling growing more than fifteenfold. That is small enough that a monthly log pull is cheap and large enough that a WAF rule quietly eating a third of it is invisible in every other report you run. The quarterly render check exists because Vercel and MERJ found no major AI crawler executes JavaScript, and that study is from December 2024, so I treat it as the best available evidence rather than settled fact, and I re-run the check myself rather than assuming.
Write these four lines properly and you own the strongest mechanical lever in the category. A full walkthrough of the check itself is in the AI crawler access audit, and the log side is covered in AI crawler log analysis.
MEASURE: the acceptance criterion is a sample size, not a screenshot
This is where most GEO SOWs collapse, because a screenshot of a good answer is easy to produce and impossible to falsify. The research that killed screenshot reporting is SparkToro's run variance study with Gumshoe.ai: 600 volunteers, 12 prompts, 2,961 combined runs, and a finding of less than a one in a hundred chance that two runs of the same prompt return the same brand list.
if you really want to know an AI's set of recommendations, you need to ask over and over again; usually at least 60-100X
That sentence is not a research note. It is a scope line. It means the repetition count belongs in the SOW as a number both parties agreed to, and the acceptance criterion for the measurement task is that the count and the dates appear on the report. If your prompt set is 40 prompts and your agreed count is 60 runs per engine across three engines, that is 7,200 samples a month, which is a tooling decision with real hours behind it. Say the number out loud in the document or you will quietly reduce it in a busy month and nobody will notice. The sizing argument is in AI visibility sample size, and the construction of the prompt list itself in building an AI visibility prompt set.
The split line is the other thing nobody writes down. Semrush and Growth Memo logged 3,981 domain appearances across 115 prompts in 14 countries and four engines, and found roughly 62% of citations are ghost citations where the page is linked but the brand name never appears in the answer text. A single blended "visibility" figure hides that entirely, which is why my acceptance criterion forbids reporting one. The distinction is unpacked in AI citations versus recommendations.
MAP and EARN: where seventeen of the thirty hours actually go
Mapping is the stage that stops the retainer becoming a content treadmill. Two findings set its shape. First, Surfer analyzed 10,000 keywords and 33,000 fan-out queries and found pages ranking for both the head query and its fan-outs were 161% more likely to be cited than pages ranking only for the head term. That is the best evidenced positive on-page finding in the category and it is about breadth of subtopic coverage, not page length. It is why the fan out coverage audit is a standing line rather than a one-time audit deliverable, and why the query fan-out mechanics matter more than keyword volume.
Second, and more uncomfortable for anyone selling on-site work: Aleyda Solis measured 84% to 93% of AI citation weight sitting on third-party properties across 15 SaaS brands. If that pattern holds in your vertical, an SOW with no off-site line is structurally incapable of moving the number it promises to move. My off-site line is 2.5 hours, which is honestly light, and I say so to clients rather than pretending it is sufficient. Peec AI's analysis of nearly 200,000 AI responses found that ranking first in a frequently cited third-party listicle was associated with a 16.5 percentage point visibility lift in B2B SaaS, which tells you where those hours should point. The mechanism behind it is covered in brand mentions and AI visibility and the source side in how to analyze AI citation sources.
The refresh line has the cleanest evidence of any task on the sheet. Seer Interactive studied 7,683 cited pages carrying 47,097 citations across ChatGPT, Gemini and Perplexity and found 75% had been updated within the last year, with 72% looking fresh by update date against only 42% by publish date.
The freshness LLMs reward is being manufactured by updates, not by new publishing.
That is why the refresh acceptance criterion explicitly rules out a date bump. It is the single easiest line item in this whole document to fake, and faking it is a fireable offence in my shop. Note also that longer is not the goal: DejanSEO's analysis of 7,060 queries and 883,262 snippets found grounding coverage falling from 61% on sub-1,000-word pages to 13% on pages over 3,000 words, so a rebuild that doubles length can measurably reduce retrievability.
If you want this table pressure tested against your own engagement before you sign it, bring the draft SOW and we will mark up the lines that cannot be reviewed.
Book a working session→PROVE: the report is a deliverable, so it gets its own criteria
Reporting is not the write-up of the work, it is one of the eighteen tasks, and it carries 2.5 hours plus a 1.0 hour working session. AgencyAnalytics found 70% of agency leaders say client reporting plays a critical role in retention in a survey of over 220 leaders, and 43% report average client lifespans of two to five years. Reporting is a retention mechanism, not an administrative tax.
With generative AI, economic uncertainty, and evolving platform rules reshaping how agencies work, it's more important than ever to stay grounded in what actually drives results.
Two first party surfaces now exist and both are free, which is why the pull is one hour rather than a tooling line. Google announced generative AI performance reports in Search Console in June 2026, and Microsoft shipped an AI Performance report in Bing Webmaster Tools in February 2026 covering Copilot, Bing AI summaries and the grounding queries the model used. Any SOW that bills for AI visibility measurement without pulling both is charging for something it could have got from the platform. The tooling landscape around that is compared in AI visibility tracking tools, and the report structure itself in GEO client reporting.
The acceptance criterion I care about most is the last one: the report names one thing that got worse. It costs nothing, it is trivially checkable, and it is the single strongest signal to a client that the numbers in front of them are not curated. Attribution deserves the same honesty, and its limits are set out in AI search attribution.
What I removed from the scope, and the evidence that removed it
A scope of work is defined as much by its exclusions. Three line items that appear routinely in GEO proposals are not in mine.
llms.txt. Ahrefs checked 137,210 domains and found that of the roughly 38,000 with a valid file, 97% received no requests for it at all in May 2026. Google's own guidance on optimizing for generative AI features states you do not need to create new machine readable files, AI text files, markup or Markdown. SE Ranking's analysis of nearly 300,000 domains found no relationship to citation at all, and removing the feature improved their model.
If your goal is showing up in ChatGPT, Perplexity, or AI Overviews, an llms.txt file is largely decoration.
Schema billed as an AI citation lever. Ahrefs tracked 1,885 pages that added JSON-LD against 4,000 matched controls and measured minus 4.6% in AI Overviews, plus 2.4% in AI Mode and plus 2.2% in ChatGPT. Structured data still belongs in the engagement for rich results and entity hygiene. It does not belong on a GEO line with a citation promise attached. The full argument is in does schema help AI citations, and the llms.txt case in does llms.txt work.
Serving Markdown to AI crawlers. Profound ran a controlled 381 page experiment over 21 days, 189 control against 192 treatment, and found no statistically significant increase in AI bot traffic.
Cutting these three is not austerity, it is capacity. Those hours move to ACCESS and EARN, where the evidence is strongest. And publishing the cut list is a sales asset in its own right, because it is the fastest way to demonstrate that the rest of the document was built the same way.
Five clauses that make the document survive month four
The task table is necessary and not sufficient. These five clauses are what stop a good SOW being quietly renegotiated by attrition.
- Exclusions, written as explicitly as inclusions, including the three cut line items above and anything the client's dev team owns.
- Client dependencies with a named owner and a service level, because a content line with no approver is not a 6 hour task, it is an open loop.
- A change control route: what happens to the hour budget when an emergency lands mid month, and who signs off the trade.
- Data ownership on exit, covering the prompt set, the run logs, the access matrix and the raw exports, all of which the client paid to create.
- A review cadence for the scope itself, quarterly, where line items are re-timed against the actual hours logged rather than the estimate.
The fifth one is the clause I have seen protect the most engagements. Every estimate on that table is wrong by some margin in month one. A quarterly re-time turns that from a credibility problem into a governance ritual, and it gives the client a legitimate place to raise a concern that would otherwise arrive as a cancellation. If you are standing this service up from scratch, sequence it alongside adding GEO services to your agency, and if you deliver for other agencies, the same table is the spine of a defensible white label GEO arrangement.
What this scope of work cannot promise
What the document does do is make the work reviewable. That is a lower bar than most GEO proposals set for themselves and a much higher bar than any of them clear. Price sits deliberately outside this post: the cost model, the packaging and the margin question are handled separately in GEO retainer pricing, and this table is the input to that one. Build the table first. You cannot price a month you cannot describe.
The rest of the evidence behind these choices, including the audit that precedes month one, is collected across my insights archive and worked through end to end in an AI visibility audit example.
Frequently asked questions
What should a GEO retainer scope of work actually contain?
Tasks, not outcomes. Each line needs a cadence, an hour estimate, the artifact it produces, and one testable sentence defining done. A deliverable list like visibility tracking or content refresh names a category of activity, which cannot be reviewed, disputed fairly, or handed to a new account manager later.
How many hours a month does GEO delivery take?
My month runs about 30 hours for a single site, single language, single location client, rising to 33.5 in a quarter-end month. Content and off-site work take roughly 40% of it. That is my allocation, not an industry benchmark, and I have found no published study of GEO delivery hours.
What is an acceptance criterion in a GEO statement of work?
One testable sentence stating the conditions under which a deliverable is accepted. In GEO it must attach to the artifact, never to the engine's behavior. You control whether a crawler received a 200 with full HTML. You do not control whether ChatGPT names the client on a given day.
Can a GEO SOW promise that a brand will appear in ChatGPT?
No. SparkToro and Gumshoe.ai found less than a one in a hundred chance that two runs of the same prompt return the same brand list, across 2,961 runs. Any document promising presence in an answer is promising something the operator cannot control and should be treated as a liability.
Should llms.txt and schema markup be in a GEO scope of work?
llms.txt, no. Ahrefs found 97% of files got zero requests across 137,210 domains and Google says the file is unnecessary. Schema, yes, but for rich results and entity hygiene rather than as an AI citation lever, because Ahrefs measured no meaningful citation uplift in a controlled test.
How often should the prompt set be re-run each month?
Put the number in the document. Rand Fishkin's guidance is at least 60 to 100 repetitions before averaging. Whatever count both sides agree, the acceptance criterion should require the run count and the dates to be printed on the report, so it cannot be quietly reduced in a busy month.
How is a GEO SOW different from a classic SEO SOW?
The measurement stage changes most. Classic SEO reports positions from a stable index. GEO reports sampled distributions from volatile engines, so sample size becomes a contractual term. Crawler access also splits into per-vendor agents rather than one Googlebot decision, and off-site work carries far more weight.
Who owns the prompt set and the run data when the engagement ends?
The client should, and the SOW should say so in a data ownership clause. The prompt set, the raw run logs, the crawler access matrix and the platform exports were all paid for by the client. Silence on this point is the most common way a departing agency retains leverage.
Sources
- AgencyAnalytics. 2026 Marketing Agency Benchmarks Report (2026-04)
- AgencyAnalytics. 2025 Marketing Agency Benchmarks Report (2025-09)
- Swiftwater and Company. How to Draft a Statement of Work for Professional Services Contracts (2026-03)
- Agile Sherpas. Definition of Done vs Acceptance Criteria (2026)
- Google Search Central. AI Features and Your Website (2025-12)
- Google Search Central. Optimizing your website for generative AI features on Google Search (2026-07)
- Google Search Central. Introducing Search Generative AI performance reports in Search Console (2026-06)
- OpenAI. OpenAI bots documentation (2026-07)
- Anthropic. Does Anthropic crawl data from the web, and how can site owners block the crawler (2026-04)
- Perplexity AI. Perplexity Crawlers (2026-07)
- Microsoft Bing Webmaster Blog. Introducing AI Performance in Bing Webmaster Tools (2026-02)
- Cloudflare. Cloudflare Radar 2025 Year in Review (2025-12)
- Vercel with MERJ. The Rise of the AI Crawler (2024-12)
- SparkToro with Gumshoe.ai. New Research: AIs Are Highly Inconsistent When Recommending Brands or Products (2026-01)
- Semrush with Growth Memo. The Ghost Citations Study (2026-06)
- Seer Interactive. Study: Content Recency's Impact on AI Visibility in 2026 (2026-07)
- Surfer. Query Fan-Out Impact on AI Overview Citations (2025-12)
- Aleyda Solis. SaaS AI Search Optimization (2026-07)
- Peec AI. The Listicle Rank Effect (2026-07)
- DejanSEO. How Big Are Google's Grounding Chunks (2025-12)
- Ahrefs. 97% of llms.txt Files Get Zero Traffic (2026-06)
- Ahrefs. Does Schema Help AI Citations (2026-05)
- SE Ranking. LLMs.txt study across 300,000 domains (2025-11)
- Profound. Does Markdown Increase AI Bot Traffic (2026-02)
- Superlines. How should a marketing agency package and price GEO services (2026)
- Aruntastic. The GEO Service Package: What to Include (2026)
- LLM Pulse. GEO Agency Guide (2026)
- Humans With AI. What Deliverables Should GEO Services Include (2026)
- First Page Sage. SEO Agency Scope of Work: What To Expect (2026)
Want me to run this on your site and show you the before and after?
One call, no pitch deck. We look at what is actually blocking you and tell you the truth about whether we can help.
Book a free consultation →