🗺In this module you build the "map" that humans and agents use to reach your
Your site is not a list of pages, it is a system of routes
There is a mindset shift that unlocks everything in this module, and it is worth pausing on it. We usually think of a website as a collection of pages: the home page, the blog, pricing, contact — a list of rooms in a building. That is a comfortable way to think for a human navigating with a menu in front of them. But an AI agent does not experience your site as a list of rooms; it experiences it as a system of routes, paths that lead from a question to an answer. And what decides whether you get cited is not how many rooms you have, but whether there is a short, clear path from the user's question to the truth you hold.
The ideal route, the one you want to exist for every important topic, looks like this: question → answer-first page → source of truth → (optional) comparison or policy. Someone asks something; there is a page that answers immediately and clearly; that page points to the canonical source of truth where the official data lives; and, if needed, there is a link to a comparison or the policy that qualifies it. When that route exists, the agent follows it effortlessly and cites you with confidence. When it does not, something dangerous happens: the agent improvises. And improvisation, which is tolerable on an opinion topic, is exactly what you do not want on prices, policies, or critical claims — because improvising on your prices means AI invents them or pulls them from a worse source.
Why architecture wins or loses citations
Stop to think about the effort asymmetry. For an agent, a clear route is cheap to follow and a broken route is expensive to reconstruct. If your truth about "how much it costs" is spread across an old landing, a two-year-old blog post, and the current pricing page — all three saying similar but not identical things — the agent faces three competing signals. In the best case it picks one at random; in the worst, it mixes all three and produces an answer that is not true on any of your pages. Dispersion does not just reduce clarity: it actively introduces error.
This connects to what you learned in Module 1 about how AI retrieves fragments. If the same idea lives in five different, slightly contradictory chunks scattered across your site, none of those chunks is as strong as a single canonical chunk would be — and they cannibalize each other semantically. Architecture, at bottom, is the art of concentrating signal instead of diluting it.
Where BotPass fits: do not draw a map nobody travels
Before you redesign your site on a whiteboard, an important warning so you do not fall into the architect-in-love-with-the-blueprint trap. It is tempting to design a perfect, elegant, symmetric map… that turns out bots do not care about. GEO architecture must be informed by real consumption, and that is where BotPass anchors you to reality.
Use BotPass first to identify which pages bots already treat as entry points. You will almost always be surprised: it may be the home page, yes, but pricing, documentation, or certain key guides you never imagined were so visited usually show up too. Those real entry points are the foundations of your map. From there, design hubs and internal linking so those entry points lead to your canonical sources of truth in one or two clicks, not five. And pay special attention to one signal: if BotPass shows bots consuming a lot of a page that is not canonical — that old landing, that forgotten duplicate — do not celebrate it as traffic, treat it as an alarm. It is a leak: AI is feeding on a version you do not control. The correct response is redirect, merge, or at minimum put crystal-clear links to the canonical page.
Hubs: the chapters of your book
Here is the first piece of architectural vocabulary. A hub is the pillar page for a main topic. And it is important to understand that a hub is not simply "a long post." A long post is content; a hub is an entry point with a structural function. A good hub does three things at once: it defines the topic clearly so anyone — human or machine — understands what it is about; it links in an orderly way to subpages that develop each facet of the topic; and it makes unmistakably clear which URL is canonical for each subtopic, so there is no doubt where the truth lives.
The metaphor that works best is the book. Hubs are the chapters of your site, and clusters — subpages grouped under each hub — are the sections within each chapter. A well-organized book has a few clear chapters, each with logical sections, and an index that takes you to any idea in seconds. A poorly organized book has the same information scattered without order, and even if the content is excellent, finding something specific is a nightmare. Your site, for AI, is exactly that: a book read by jumping around, and hubs are your index.
Sources of truth: one truth, one URL
Within each chapter there are pages with special status: sources of truth, or single source of truth. A source of truth is the page you want AI to cite when someone asks about a price, a policy, a definition, terms, or a methodology. These are the data points where accuracy is non-negotiable.
The rule here is brutally simple: if it is critical, it cannot be spread across five URLs with different wording. Every critical truth in your business should have a single canonical home, and everything else that talks about that topic should point to that home instead of repeating the data in its own words. This is not just content hygiene; it is a risk decision. Every duplicate of a price is a future contradiction waiting to happen the moment you change the price in one place and forget the other. And AI, consuming both, will propagate that contradiction multiplied. Consolidating your sources of truth is, at once, GEO work and risk management.
Internal linking for agents, not just SEO
Here is a subtle turn that distinguishes a GEO practitioner from a classic SEO professional. Traditional SEO thinks about internal linking in terms of ranking: which links distribute "authority" and help you rank. That logic is legitimate, but it is a different game. GEO thinks about internal linking in terms of understanding and citation: which links help an agent understand the structure of your knowledge and reach the correct truth.
That gives rise to a type of link that in SEO would be almost redundant but in GEO is gold: intent links. Instead of a generic menu, think in explicit signals like "if your question is about pricing, go to the canonical pricing page," "if your question is about terms, go to the canonical policy," "if you want to compare, go to the alternatives page." They do not have to be written that literally on the page surface — though sometimes it helps — but your linking structure should make that intent-to-destination mapping obvious. You are, in essence, leaving breadcrumbs for the agent so it does not have to guess.
Build your GEO Map v1: the step-by-step method
With the concepts in place, let us build your first map. It is a focused session of work, and the result — GEO Map v1 — is one of the most valuable documents you will produce in the entire Academy.
Start by defining your hubs. Take the three to five main topics you already identified in Module 1 and create one hub per topic. No more: if you end up with eight hubs, some are probably subtopics of others, and your map will lose focus. Fewer hubs, clearer, always wins.
Next, design the clusters for each hub. Each hub should support between three and eight subpages covering the different facets of the topic: definitions, use cases, practical guides, comparisons, and related documentation or policies. The table below is the tool to capture this structure; pay special attention to the "action" and "target canonical URL" columns, because they turn a static inventory into a work plan:
| Hub | Cluster / subpage | Intent | Current URL | Action | Target canonical URL |
|---|---|---|---|---|---|
| Define / Compare / Buy / Implement | Create / Improve / Merge / Redirect |
The third step is to mark one to three sources of truth per hub. Not every subpage is a source of truth; only those that hold critical data. Identify them explicitly, because they will receive priority attention in the content and schema modules:
| Topic | Source of truth | URL | What it should answer |
|---|---|---|---|
| Pricing / Policy / Doc / Definition |
And the fourth step is to design at least ten reading routes. A reading route is a small recipe that goes from a real question to the sequence of pages that answer it: the question, the answer-first page that handles it, the source of truth that grounds it, and optionally the comparison or policy that completes it. Writing ten of these routes forces you to verify, one by one, that the important paths on your site actually exist and are short. When you discover a route cannot be completed because a page is missing, you have just found a piece of your backlog.
An illustrated map, start to finish
So the method does not stay abstract, let us follow it with a concrete example. Suppose one of your main topics is "AI scraping." Your hub would be the pillar page /ai-scraping, which defines the topic and serves as the entry point. From it would hang several clusters: /ai-scraping/what-is with the canonical definition, /ai-scraping/how-to-detect as a practical guide, /ai-scraping/pricing-impact as a data benchmark, and /ai-scraping/alternatives as a comparison. Your sources of truth for this topic might be /pricing, where plans and terms live, and /policy/bot-access, where the bot policy lives.
Now observe how a concrete reading route materializes on this map. Someone asks: "How do I detect AI bots on my website?" The answer-first page is /ai-scraping/how-to-detect, which answers immediately; the source of truth that grounds it is /docs/bot-detection; and the policy that qualifies it is /policy/bot-access. The agent enters through the answer, confirms in the source of truth, and finishes at the policy — all in a clean sequence. That structural clarity is exactly what reduces ambiguity and increases the probability of a correct citation: the system "sees" a sharp route and follows it instead of improvising.
🎓 Academy Level — Design a map humans and agents can follow
Everything above is open level. From here: full methodology, templates, hands-on labs, module checklist, and the 🏁 Certification milestone — free when you register at botpass.io.
🔒 Academy Level content
Register free at botpass.io (or sign in) to unlock the full module, templates, and certification milestone.
Sign in Create free account🎓Academy Level. Everything above is open level. Here you unlock GEO map
Fan-out: why one question becomes many
There is a technical phenomenon that has completely changed content strategy and that you must understand deeply, because it is the reason much of this module exists. Modern AI engines, when they receive a question, rarely answer it as-is. Instead, they expand it into a fan of sub-questions — this is query fan-out — and search for answers to each one before synthesizing a final response. If someone asks "what is the best GEO plugin for WordPress?", the engine may internally break it into "what is a GEO plugin?", "what GEO plugins exist for WordPress?", "how much do they cost?", "what do users think?", "how do they compare to Yoast?" and several more.
The strategic consequence is enormous and changes how you must think about your content. In the old SEO world, you optimized one page for one keyword. In the fan-out world, you win when your architecture covers the full tree of sub-questions the engine generates from your topic. It is not enough to have one great page about "GEO plugins"; you need solid answers among your clusters to each branch of that tree. That is why hubs with their clusters are not organizational luxury: they are literally how you cover fan-out. Every cluster you add is a branch of the tree you can now win, and every uncovered branch is a sub-question someone else will answer in your place.
📋 Templates
Template A — GEO Map v1 (hub → clusters → canonical). This is the working version of the open-level table, meant to stay alive throughout the program. Every time you create, improve, merge, or redirect a page, update the action column: the map becomes your architectural progress log.
| Hub | Cluster / subpage | Intent | Current URL | Action | Target canonical URL |
|---|---|---|---|---|---|
| Define / Compare / Buy / Implement | Create / Improve / Merge / Redirect |
Template B — Reading route (intent → route recipe). Write at least ten. Every row you cannot complete because a page is missing is, literally, a task on your content backlog.
| User question | Answer-first page | Source of truth | Comparison/Policy (optional) |
|---|---|---|---|
🧪 Hands-on labs
Lab 1 — Build a complete hub (40 min). Choose your number-one main topic and build it whole in one sitting, because doing one well teaches you the pattern to replicate on the rest. Create the pillar hub, define between three and eight clusters covering definition, use cases, guide, comparison, and docs or policy, mark one to three sources of truth, and obsessively verify that each cluster reaches its canonical in one or two clicks. If any need three or more, your route is too long and must be shortened.
Lab 2 — Fan-out coverage (30 min). Run BotPass Fan-out Coverage on your three main topics. It returns the tree of sub-questions AI expects you to answer, and you mark which your site covers today and which it does not. Uncovered sub-questions are your cluster backlog, ordered by the AI engine itself. It is one of the most honest ways to discover your gaps, because it is not your opinion of what is missing — it is what the system actually searches for.
| Topic | Expected sub-question | Covered? | Page that should answer it |
|---|---|---|---|
| Yes/No |
⚡ Hacks
Three GEO architect reflexes. The one-or-two-click rule is your acid test: if an agent cannot get from the home page to your canonical in two clicks, the route is too long and internal linking must be reorganized. Conflicting canonicals are your silent alarm: if BotPass shows bots consuming a lot of a page that is not canonical, do not read it as success, read it as a leak and redirect or merge immediately. And links for agents are your unfair advantage: add explicit intent links — "for pricing, the canonical pricing page" — instead of settling for "pretty" links designed only for SEO.
💎 Additional value content
On the anatomy of a hub that wins citations: a good hub defines the topic, links to its clusters, signals which URL is canonical for each subtopic, and answers the parent question in the first scroll without forcing a scroll down. On semantic clusters: BotPass automatically groups your pages by topic and detects fine or missing subtopics, giving you an objective read on whether your architecture has the pillar+support shape engines reward. And on orphan pages: any content with no internal links pointing to it is, for practical purposes, invisible to an agent, no matter how good it is; finding those orphans and reconnecting them to the map is one of the highest-return, lowest-effort improvements that exists.
📚 References and recommended reading
- External sources on RAG and retrieval architecture. They complement Academy Level; they do not replace hands-on work with BotPass.
- AWS — What is RAG (Retrieval-Augmented Generation)? — clear explanation of how systems retrieve external sources before answering. ↗
- Google — AI in Search: going beyond information (query fan-out) — official explanation of how AI Mode breaks a question into multiple sub-queries. ↗
- Semrush — What Is Query Fan-Out & Why Does It Matter? — how fan-out changes content strategy from keyword to topical coverage. ↗
- Wikipedia — Retrieval-augmented generation — neutral conceptual base with references. ↗
🏁 Certification milestone (Module 3)
🏁BotPass milestone (verifiable): GEO Map v1 approved + Fan-out Coverage check
📣 Share your progress (optional): "AI engines expand every question into dozens of sub-questions. I just checked my Fan-out Coverage with BotPass: of (N] sub-questions about [topic], my site only answers (M). That is my content gap. #GEOAcademy"
When you have completed the practice and the milestone is recorded in your BotPass plugin, mark the module. After completing all 10, submission for review unlocks.