🧪This module tackles the most common and costly GEO mistake: "I made
Why GEO needs the scientific method
It is worth understanding why GEO is especially treacherous when it comes to knowing if something works, because if you do not internalize this you will fall into the traps again and again. GEO has a lot of noise: bots do not visit at a steady rate, but in bursts; AI products change their algorithms and sources from month to month; and there is seasonality in searches and consumption. All that noise means your metrics go up and down on their own, without you doing anything. And that is where the danger lies: if you make a change and your metric rises right after, your brain will want to attribute the rise to the change, even when the real cause was noise fluctuation. Without experiments, you will systematically confuse coincidence with causality, and make decisions based on phantom correlations.
The good news is the solution does not require being a data scientist. It is a DIY method of four steps anyone can follow: change one thing only, log it with its exact date, measure in clear, well-defined windows, and decide with rules set in advance. That minimum discipline is what separates real learning from the illusion of learning. It is not about complicating things; it is about not fooling yourself.
Where BotPass fits: your "reading" signal
In the GEO world, measuring is hard because many things that matter —what a chatbot answers exactly, whom it cites— are opaque or hard to track. BotPass gives you the most practical input metric of all: reading. It tells you, with data, whether bots are reading the page you changed, which answers the first question of any experiment: has my change even reached AI's eyes? It lets you compare baseline consumption with post-change consumption, broken down by URL and bot, which is the core comparison of every GEO experiment. And it helps you detect unwanted side effects: if after a change you see bots returning to old URLs, that is a signal you probably need redirects or clearer canonicals, a diagnosis that connects with modules 3 and 6.
The reason reading is so valuable as a signal is that it sits high in the causal chain. Before AI can cite you, it has to read you; before a content change can influence an answer, the bot has to consume the changed page. By measuring reading, you get the earliest and least noisy signal that your work is having an effect, weeks before it shows up in business metrics.
The three layers of metrics
One of the most useful frameworks in the entire module is to think of your metrics in three layers, because each answers a different question and none alone tells the full story. The first layer is reading, measured by BotPass: what bots consume. It answers "are they reading me?". The second is response: citations, referrals, and correct mentions. It answers "are they using and citing me correctly?". And the third is business: leads, demos, licenses. It answers "does this generate real value?". The three form a funnel: reading enables response, and response enables business.
There is a principle that should govern this entire measurement architecture, and it is worth engraving: a metric without an associated decision is vanity. If you measure something but are not clear what you would do differently if it rises or falls, that metric serves only to make you feel good or bad. Before adding any metric to your dashboard, ask what decision would change based on its value; if there is no answer, drop it. This discipline keeps your measurement focused on what is actionable and saves you from drowning in pretty, useless numbers.
Building a baseline that works
No experiment is worth anything without an honest baseline to compare against, and building one well has its technique. You need three things. A window of seven to fourteen days, long enough to average bot burst noise but not so long that the world changes underneath. A fixed URL list, decided in advance and not modified during measurement, because changing what you measure midstream invalidates the comparison. And a mandatory change log, where you note every modification with its exact date. That last piece is the most neglected and the most critical: without a dated record of what you touched and when, it will be impossible to attribute a metric change to a concrete cause later, and you return to "I think it worked" territory.
How to design a GEO experiment
With the baseline running, designing an experiment is filling out a template that forces you to think before acting. The fields are deliberately strict: a clear hypothesis —what you think will happen and why—; the exact change you will apply, described precisely so it can be reproduced; the affected URLs; the measurement window, typically fourteen days; the primary metric that will decide the outcome; the success threshold set before looking at data, to avoid retrospective cheating; and the result you record at the end. Setting the success threshold in advance is the safeguard against the most common self-deception: looking at data first and deciding afterward what would have counted as success.
A well-designed experiment, illustrated
Let us see how all of this looks in a concrete case. The hypothesis is: "If we add a TL;DR and intent-based FAQs on /pricing, AI referrals will increase and support confusion will decrease." Notice the hypothesis does not only predict improvement, but explains the mechanism by which it expects it to occur. The change is precise and reproducible: add a five-bullet TL;DR, add ten intent-based FAQs, and add the "last updated" date along with currency and VAT conditions. The window is fourteen days. The primary metric is AI referrals to /pricing, and the secondary is bot consumption on /pricing measured with BotPass, which checks that the change is being read. And the success threshold is set in advance: a twenty percent increase in referrals, or a consistent improvement sustained across two windows.
What makes this a method and not a trick is that anyone can repeat exactly the same procedure on another URL, with another change, and get an equally interpretable result. A well-designed experiment does not only tell you whether this change worked; it teaches you a pattern you can apply systematically across your site, turning every improvement into an accumulable unit of learning.
🎓 Academy Level — Experimentation and ROI with the scientific method
The above is open level. From here: full methodology, templates, practical labs, module checklist, and the 🏁 Certification milestone — free when you register at botpass.io.
🔒 Academy Level content
Register free at botpass.io (or sign in) to unlock the full module, templates, and certification milestone.
Sign in Create free account🎓Academy Level. The above is open level. Here you unlock the experiment
The Entity Equity Score: your north-star metric
Here it is worth presenting in depth the metric that, across the entire program, acts as your north star: the Entity Equity Score. While the GEO Score you met in earlier modules measures the technical quality of a specific page —its structure, schema, citability— the Entity Equity Score operates at a higher level: it measures the accumulated strength of your entity as a recognized source. It combines in one number signals you have been working on module by module: entity recognition (Module 5), schema quality (Module 5), the solidity of your sameAs chain and NAP consistency (Module 5), the health of your semantic clusters (Module 3), and the volume and quality of your citations (Module 8). It is, in a sense, the metric that summarizes whether the whole system is working together.
That is why it is the ideal metric to track long term: an individual experiment moves one page's GEO Score, but it is the sustained trend of Entity Equity Score that tells you whether your authority as a source is truly growing. Think of GEO Score as a page's pulse and Entity Equity Score as the health of the whole organism. Monitoring its trend before and after each experiment —and over months— gives you the big-picture view no single-page metric can offer.
Attribution in GEO: the art of not fooling yourself
Attribution —knowing which cause produced which effect— is notoriously hard in GEO, and it is worth attention because it is where most people deceive themselves. The underlying problem is that many things move at once: you make changes, but at the same time AI products update their models, seasonality raises or lowers demand, and a competitor may publish something that shifts the landscape. Any of those external factors can mimic the look of "success" —or mask a real success—. The only robust defense is the combination of a meticulous change log and the discipline of changing one thing at a time, so when a metric moves you can look at your changelog and have a credible hypothesis for the cause. Perfect attribution is impossible in GEO; reasonable attribution is achievable, and it is enough to decide well.
📋 Templates
Template A — GEO experiment worksheet. Fill it out completely before touching anything, especially the success threshold. A worksheet completed afterward is not an experiment, it is rationalization.
| Field | Content |
|---|---|
| Hypothesis | |
| Change (exact) | |
| URL(s) | |
| Window | 14 days |
| Primary metric | |
| Success threshold | |
| Result |
Template B — Three-layer dashboard: Reading (BotPass: consumption by URL/bot) · Response (citations/referrals/mentions) · Business (leads/demos/licenses), with baseline vs post-change column. Keeping all three layers in view avoids celebrating a reading spike that never translates to business, or discarding a change whose business effect has not matured yet.
🧪 Practical labs
Lab 1 — Design and launch 1 experiment (40 min). Fill out the complete worksheet, apply a single change on a single URL, and record the exact date in the changelog. The temptation to "touch a couple more things while I'm at it" is very strong; resist it, because every extra variable you touch is a cause you will not be able to isolate later.
Lab 2 — Build the dashboard (40 min). Set up the baseline vs post-change comparison with GEO Score, Entity Equity Score, and analytics, and track Entity Equity Score trend before and after. Watching that number move over the weeks is one of the most motivating ways to confirm accumulated work pays off.
⚡ Hacks
Three reflexes of a rigorous experimenter. Change one thing only: if you touch five variables at once and the metric moves, you will not know which one did it, and you will have spent an experiment without learning anything. Two windows to confirm: improvement sustained across two windows is worth much more than an isolated spike, which is almost always noise disguised as signal. And avoid false correlations: remember that changes in AI products or seasonality can mimic "success," so distrust improvements that coincide with known external events.
💎 Additional high-value content
On the Entity Equity Score, you already have it above: understanding that it combines recognition, schema, sameAs, NAP, clusters, and citations in one number gives you the holistic metric no isolated signal offers. On vanity metrics vs decision metrics: internalize that any metric without an associated action is excess, and prune your dashboard without mercy. And on attribution in GEO: accept that you need a change log to avoid confusing coincidence with cause, and that this methodological humility is precisely what makes your conclusions reliable.
📚 References and recommended reading
- External sources on measurement, experimentation, and ROI. They complement Academy Level; they do not replace hands-on practice with BotPass.
- Swydo — The Agency Guide to Tracking AI Traffic in GA4 — step-by-step setup, regex patterns, and custom channels to isolate AI traffic. ↗
- Two Octobers — Tracking AI Traffic in GA4 — A Step-by-Step Guide — segments and explorations to separate assistant traffic. ↗
- Contentful — What is GEO and how does it differ from SEO? — why measuring GEO requires shifting from volume KPIs to mentions and pipeline impact. ↗
- GEO — Generative Engine Optimization (paper) — the original visibility metrics framework for designing your own experiments. ↗
🏁 Certification milestone (Module 9)
🏁BotPass milestone (verifiable): dashboard running (baseline vs post-
📣 Share your progress (optional):"Result of my first GEO experiment: [change] on [URL] → [metric] went from [X] to [Y] in 14 days. My Entity Equity Score: [trend]. Hypothesis, window, and success threshold 👇 #GEOAcademy"
When you have completed the practice and the milestone is recorded in your BotPass plugin, mark the module. After completing all 10, submission for review unlocks.