Generative engine optimization agencies are selling a service whose primary metric cannot yet be measured officially. That does not make the category fraudulent — the underlying work is real and the shift in search is real — but it does mean the buyer carries an unusual amount of diligence.
The good news is that most of the diligence is testable. Below are seven checks you can run yourself, including one that takes about two minutes and reveals whether an agency's reported "AI visibility" is durable or borrowed.
First, what the market actually costs
Pricing is opaque in this category because scope genuinely varies and because measurement standards are unsettled. Published ranges cluster like this:
| Engagement type | Typical monthly cost |
|---|---|
| DIY tools / AI visibility software | $10 – $1,000 |
| Small business retainer | $1,500 – $5,000 |
| Mid-market retainer | $2,000 – $8,000 |
| Enterprise / competitive brands | $10,000 – $25,000+ |
Sources including WebFX and Digital Elevator put the overall agency band at roughly $1,500 to $50,000+ per month.
Use these as sanity bounds, not benchmarks. A $2,000 retainer that fixes your canonical tags, rendering and entity consistency will outperform an $8,000 retainer producing AI-written blog posts.
Test 1 — Toggle web search off
This is the highest-value two minutes you will spend, and almost no agency will volunteer it.
Ask an AI assistant a question where you are reportedly visible, with live web search enabled. Note whether you appear. Then turn live retrieval off and ask again.
An SEO practitioner documented this pattern precisely in r/SEO, using an accounts-payable software vendor as the example:
"Ask ChatGPT for the best AP software with web search enabled, and SuitiAP shows up. Why? Because it appears in AI-SEO listicles and comparison pages that surface in search. Turn live retrieval (search) off, and SuitiAP disappears from AI answers."
The same analysis noted the revealing split: the top platform recommended from training data alone launched in 2016, while the top platform recommended with web search enabled launched in 2024. Those are two different visibility products.
Watch out
Neither result is worthless — retrieval-based visibility drives real traffic today. But an agency reporting "we increased your AI visibility by X%" without separating retrieval from trained knowledge is measuring something far more fragile than the number implies. The same practitioner's blunt summary: prompt tracking "can fool you into thinking you're doing great at AEO or GEO, when really you've just temporarily hijacked a search result."
Ask your prospective agency directly: does your reporting distinguish retrieval-driven mentions from model-knowledge mentions? The answer tells you how sophisticated they are faster than any case study.
Test 2 — Ask which deliverables Google says are unnecessary
Google published official guidance on optimizing for generative AI features that includes a section explicitly labelled mythbusting. Google states you do not need:
- llms.txt or other special markup — Google Search does not use these files
- "Chunking" content into small pieces — there is no ideal page length
- Rewriting content specifically for AI — models understand synonyms and intent
- Inauthentic "mentions" — spam systems target manufactured brand-drops
- Structured data specifically for AI — useful for rich results, not required for generative features
If a proposal leads with llms.txt setup and content chunking as headline deliverables, the agency is either not current with Google's published position or is billing for effort that Google has said does not drive the outcome.
This is not a fringe objection. A widely upvoted r/SEO thread captured the practitioner mood bluntly:
"I've been SO sick of getting reels sent to me from folks asking, 'Is this something we should be doing? Are we GEO optimized? Do we have an llms.txt setup yet?' 90% of what the GEO/AEO influencers push is just garbage and a waste of time."
The discipline is real. A large share of what is marketed under its name is not.
Test 3 — Demand the source of any citation-share number
If an agency shows you a precise figure — "you hold 12% citation share in your category" — ask exactly where it comes from.
Here is why that matters. Google launched Generative AI performance reports in Search Console on June 3, 2026. That is the official measurement surface, and it has hard limits:
- It reports impressions only — no clicks, no CTR, no query data
- It rolled out to a subset of site owners, starting in the UK
So any precise citation-share number is third-party prompt sampling. That is a legitimate technique, and it is an estimate with real variance — results shift by model version, by whether retrieval is on, and by personalization and memory. An agency that presents an estimate as a metric is either careless or counting on you not knowing the difference.
Good answer: "It's prompt sampling across a fixed panel, run monthly, here's the panel and here's the variance." Bad answer: "It's from our proprietary platform."
Our audit states its own limits: a fixed prompt panel with the methodology written down, a technical foundation review, and the specific pages to fix. You keep the findings whether or not you work with us.
Test 4 — Require independently verifiable references
Self-published proof is not proof. A Canadian agency roundup made this point about a competitor that markets itself as "Canada's #1 SEO agency" — a self-awarded claim on its own site rather than an independently verified ranking.
Ask for one of the following instead of a testimonial hosted on the agency's own domain:
- Reviews on a third-party platform such as Clutch or Google Business Profile
- A named case study with specific, checkable numbers
- A live before-and-after report you can inspect, not a screenshot
Test 5 — Confirm the engagement goes past diagnosis
The single most common complaint about this category is that it stops at telling you there is a problem. From the same practitioner discussion:
"Or pushing AI visibility tools — that just tell you that you have a problem, which you already know…"
Another put the standard clearly:
"Knowing you're invisible is worthless on its own. The only version worth anything is the one that hands you the specific pages to write, the exact phrasing to seed, and where to seed it. Diagnosis is the cheap part."
Ask what you receive in month two, after the audit. If the answer is a dashboard, you have bought monitoring, not optimization.
Test 6 — Verify they fix the technical foundation first
Generative engines retrieve from search indexes. A site that cannot be crawled, rendered or ranked will not be cited regardless of how quotable its prose is. Any competent engagement starts here:
- One canonical host — www or non-www, redirect permanently, with canonical tags, sitemap and internal links in agreement
- Server-side rendering so content does not depend on JavaScript execution
- AI crawler access in robots.txt — GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, Google-Extended
- Zero-click page triage — any page with impressions and no clicks has a title and description problem
These are unglamorous and they gate everything else. We found a canonical split on our own site that was quietly splitting ranking signals across two hosts; the full walkthrough with the actual Search Console numbers is here.
If a proposal jumps straight to content production without auditing this layer, the content is being poured into a leaking container.
Test 7 — Ask them to describe an engagement that failed
An agency that has run enough programs has lost some. Ask for one, and listen for whether the explanation is mechanical or evasive.
Useful answers sound like: the client was in a category where the model leaned almost entirely on two industry publications we could not get into, or their site was on a platform we could not server-render without a rebuild. Those describe real constraints.
Evasive answers blame the client's budget or "not being ready." Every category has losing conditions, and knowing yours is a sign of competence.
The short version
Generative engine optimization is worth paying for when it is competent technical SEO plus deliberate entity building plus genuine third-party presence work. It is not worth paying for when it is llms.txt files, AI-written filler and a dashboard.
The seven tests, condensed:
- Toggle web search off — is the visibility durable or retrieved?
- Do deliverables include things Google has said are unnecessary?
- Where does any citation-share number actually come from?
- Are references independently verifiable?
- What happens after the audit?
- Is the technical foundation fixed first?
- Can they describe a failure honestly?
If you want to understand the mechanics before you talk to anyone, start with how GEO compares to SEO — knowing what the work actually consists of is the best protection against buying the version that is mostly vocabulary. Our generative engine optimization guide covers the underlying theory, while ranking on ChatGPT and getting cited by Perplexity show how differently each engine behaves — useful context for judging whether an agency actually knows the difference.
