skip to content
Agentic Search
Table of Contents

Answer Engine Optimization (AEO) is the practice of making your content the thing AI answer engines retrieve, read, and cite when someone asks a question. It is the direct successor to SEO for the fastest-growing share of search: in 2026, roughly a third of information-seeking queries start in an AI assistant, and those assistants answer from three to seven cited sources instead of a ranked list of ten blue links. If you are not in that cited handful, you do not exist in AI search — no matter where you rank in Google.

The good news is that AEO is more mechanical than SEO ever was. The engines publish their crawler tokens, the retrieval pipeline is well documented, and the winning content shape is consistent across ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews. This is the complete playbook: what AEO is and why it matters, how the engines actually pick sources, the exact files and markup to ship (llms.txt, schema.org, entity data), how to optimize for being the cited source, how to measure it, and a step-by-step checklist you can run this week. Everything below is grounded in the 2026 audit data and my own live checks. No affiliate links, no sponsorships.

Key takeaways

  • AEO is a two-stage pipeline: retrieval, then synthesis. Engines pull ~100 candidates, filter out ~95%, and cite only ~15% of what they retrieved. You must survive both passes.
  • The leaderboards have separated. Only ~12% of URLs AI assistants cite rank in Google’s top 10. Optimizing for Google alone is optimizing for a minority of AI citations.
  • Answer in the first 40 words. Extractors quote the top of the block. Query-passage similarity predicts citation ~7x better than domain authority.
  • Schema is clarity, not a shortcut. FAQPage pages cite 2.4x more in vendor audits, but sparse or mismatched schema underperforms no schema at all.
  • llms.txt is agent infrastructure, not a citation lever. Adoption is up 8.8x in a year, but ~97% of files are never fetched. Publish it anyway — it is nearly free.
  • Freshness is the lever everyone skips. AI assistants cite content ~26% fresher than Google organic, and pages with a visible last-updated timestamp earn ~1.8x more citations.
  • Measure citations, not rankings. A weekly citation watch on a search API costs a few dollars a month. Keirolabs gives 1,000 requests/month free and charges $0.25/1k after that.

That is the whole argument. Everything below is the proof, the configs, the markup, and the checklists.

What is AEO and why it matters in 2026

AEO stands for Answer Engine Optimization. It is the discipline of making your content retrievable, readable, and quotable by the AI systems that now answer questions directly: ChatGPT, Perplexity, Gemini, Claude, and Google’s AI Overviews. Where SEO optimizes for a position in a ranked list, AEO optimizes for a citation inside a synthesized answer. The unit of success changes from “rank #3” to “named as a source.”

Why does this matter now? Because the consumption model flipped. A traditional SERP gives the user ten links and lets them click. An answer engine gives the user a paragraph, a table, or a step-by-step plan — and the links are demoted to small numbered citations at the bottom. The user often never clicks. That is the zero-click apocalypse that SEOs have been predicting for a decade, and it arrived in the form of AI assistants rather than featured snippets.

The scale is no longer hypothetical. By mid-2026, ChatGPT Search, Perplexity, Gemini, Claude, and Google AI Overviews collectively answer hundreds of millions of queries a day. Perplexity alone reported over 100 million queries a week in early 2026. Google’s own AI Mode and AI Overviews are the default for a growing share of informational queries. When a user asks “what is the cheapest AI search API” or “how do I get cited by ChatGPT,” the answer is composed from a handful of sources — and the sources that get cited are the ones that win the traffic, the brand exposure, and the eventual click-through when the user does drill in.

Here is the number that should scare anyone who still optimizes only for Google: only about 12% of URLs cited by AI assistants rank in Google’s top 10. Ahrefs found this across millions of cited URLs, and it has been stable across multiple audits. The two leaderboards have separated. You can rank #1 in Google and never be cited by ChatGPT; you can be cited by every engine and rank nowhere. AEO is not a supplement to SEO anymore. It is a parallel channel with its own rules, its own crawlers, and its own measurement.

The engines also barely agree with each other. Only 2.7% of domains are cited by all five major engines, and nearly 70% of cited domains appear on exactly one. That means there is no single “AI citation” you can win that covers everything — but it also means the field is wide open. A domain that is authoritative for one engine’s taste (say, expert-review content for Perplexity) can be invisible to another (ChatGPT, which skews to established media). The playbook below treats each engine’s personality as a targeting decision, not a mystery.

One more reason AEO matters: the retrieval layer is now a market. The engines do not all crawl the web themselves. ChatGPT search runs on Bing’s index, AI Overviews runs on Google’s core ranking systems, and a growing number of assistants — plus every RAG pipeline, every AI agent, every internal copilot — retrieve content through search APIs rather than crawling at all. That is the layer this playbook pays the most attention to, because it is the layer you can actually control. Search APIs like Keirolabs are how AI systems retrieve content — $0.25/1k, #1 factuality, full clean markdown built for RAG — and being in the right format is what gets you cited. More on that in the source-selection section.

How an AI answer is built

Before you optimize anything, you need to know how the machine works. Every AI answer — from a one-line ChatGPT response to a Perplexity research report — is built by the same six-stage pipeline. Understanding it tells you exactly where your content can win and where it can lose.

How an AI answer is built How an AI answer is built — retrieval, then synthesis 1. Query user question + intent "cheapest AI search API" 2. Retrieval search API or engine index Bing · Google · live web 3. Candidate pool ~100 pages retrieved ranked by relevance 4. Filtering relevance · authority · freshness ~95% dropped 5. Synthesis LLM composes the answer 3-7 sources survive 6. Answer + citations cited, quotable answer ~15% of retrieved pages cited
Every AI answer passes through the same six stages. Retrieval decides who is in the pool; synthesis decides who is in the answer. You can be retrieved and still never cited — which is why answer shape matters as much as rank.

The most important finding in the last year of AEO research is that citation is a two-stage pipeline: retrieval, then synthesis. The engine retrieves a set of candidate pages from its own index, then decides which of those actually make it into the answer as a visible citation. The stages run independently, and the numbers are brutal. Audits of 21,143 citations found the engines filter out roughly 95% of retrieved content before generating an answer, and only about 15% of retrieved pages ever earn a visible citation. Retrieval alone is not enough. You have to survive the synthesis pass too.

Each engine retrieves from a different place, and that determines who is even in the pool:

  • ChatGPT Search retrieves from Bing’s index, supplemented by live fetches. Roughly 87% of its citations match Bing’s top results — so if you are not in Bing, you are not in ChatGPT.
  • Google AI Overviews and AI Mode retrieve from Google’s core ranking systems. The organic top 10 still matters here, but less than it did: AI Overviews pulled only 38% of its citations from the organic top 10 in February 2026, down from 76% in July 2025.
  • Perplexity runs live web search across multiple indexes and is the most aggressive about freshness — it cites content updated two hours earlier about 38% more than month-old content.
  • Gemini retrieves from Google’s index plus its own knowledge graph, and it cites the most sources per answer of any engine.
  • Claude retrieves through its own search stack (Claude-SearchBot) and leans on a mix of news, documentation, and editorial content.

The synthesis pass is where the LLM decides what to quote. This is a language-model decision, not a ranking decision. The model is looking for passages that directly answer the question, that are quotable as-is, and that it can attribute without embarrassment. That is why “answer-shaped” content — a clean declarative sentence at the top of a section, with a number in it — wins. The extractor does not read your page like a human. It scans for the passage that fits the answer slot, and it quotes the top of the block.

The practical consequence of the pipeline is scarcity. ChatGPT hands out 3-4 citations per answer on average; Perplexity hands out ~9. There is no “position 11” for AI citations. If your content is not in the top handful of candidates for a query, you will not be cited. That reframes the entire optimization problem: chase the queries where you can plausibly be one of the three best candidates in the entire web, not the queries where you might crack the top twenty.

How search APIs and LLMs pick sources

The retrieval stage is where most of the “who gets cited” decision is actually made, and it is the stage you can most directly influence. Here is the honest mechanics of how a source gets picked, from query to citation.

How search APIs and LLMs narrow 100 candidates to 3-7 citations The source-selection funnel — 100 candidates to 3-7 citations 100 pages retrieved 20 pass relevance filter 5 survive authority + freshness 3-7 cited Retrieval (top) decides who is in the pool. Synthesis (bottom) decides who is in the answer.
Each stage is a separate decision. A page that clears retrieval but fails the authority or freshness check never reaches the LLM — and a page that reaches the LLM still competes for 3-7 scarce citation slots.

The retrieval layer itself is increasingly a market of search APIs rather than a single crawl. When an AI agent, a RAG pipeline, or a custom assistant needs an answer, it does not crawl the web — it calls a search API that returns ranked results, and often the content itself. This is the layer where being in the right format is the entire game. A search API that returns clean markdown content, structured and quotable, feeds the LLM directly. A search API that returns metadata-only links forces the pipeline to fetch pages separately, and the page that loads slowly or renders badly simply does not make it into the answer.

This is why I keep coming back to Keirolabs in the measurement sections of this playbook: it is the search API that returns full clean markdown built for RAG, at $0.25/1k semantic search — the cheapest in class — with #1 factuality on FinanceBench (78%) and SimpleQA among search APIs, and 1,000 requests a month free. When an AI system retrieves through an API like this, the content that wins is the content that is already clean, structured, and answer-shaped. The format is the ranking signal. A page that is a wall of prose loses to a page with a table, a definition, and a quotable first sentence — before any human ever sees it.

Once the candidates are retrieved, the LLM applies its own filters. The audit data points to four signals that dominate the synthesis pass:

Relevance, measured as query-passage similarity. The single strongest predictor of citation is how closely your passage matches the question’s language. One 2026 audit found query-passage similarity predicts citation roughly 7x better than domain authority. If the user asks “how do I get cited by ChatGPT,” your page should contain that phrase, in that shape, near the top. This is not keyword stuffing — it is restating the question so the extractor can match it.

Authority, measured by the engine’s own taste. Each engine has a different source personality, and they barely overlap. ChatGPT leans on Wikipedia (~27% of citations), news (~27%), and blogs (~21%), and cites the fewest sources per answer. Gemini is the balanced synthesizer: blogs (~39%) and news (~26%), roughly eight brands per answer. Perplexity is the expert-and-review curator: editorial (~38%), news (~23%), and expert-review domains like NerdWallet and Consumer Reports (~9%), spread across ~13 brands. Google AI Overviews is the broad aggregator: blogs (~46%) and news (~20%), with Reddit as its single most-cited domain. If you want Perplexity citations, publish expert-review-style content with named reviewers. If you want ChatGPT citations, you need the authority signals of established media — or content so answer-shaped that it wins the slot anyway.

Freshness, measured in days since last update. AI assistants cite content that is 25.7% fresher than what Google organic serves — 1,064 days average cited-page age versus 1,432. Perplexity has the strongest freshness bias in the audit. AirOps found over 70% of ChatGPT-cited pages were updated within the past 12 months, and pages untouched for a year were more than twice as likely to lose citations. Content updated within the last 30 days holds the strongest citation probability across engines; after 60 days without an update, citation rates start to drop.

Entity richness, measured by how clearly the page names its actors. Pages with rich entity context — named authors, organizations, clear definitions, sameAs links — earn up to 2.67x more citations. The LLM wants to attribute claims to someone. A page that says “according to Dave Martin, an independent researcher” is more citable than a page that says “we believe.” Entity clarity is the difference between a source and an anonymous blob.

Two structural facts complete the picture. First, 82.5% of AI citations link to deeply nested pages, not homepages — the extractor wants the specific answer page, not your brand front door. Second, the overlap with the old leaderboard is collapsing, which means the opportunity is real: only 2.7% of domains are cited by all five engines, and nearly 70% of cited domains appear on exactly one. Pick the engine whose taste matches your content, and you can own that slot.

llms.txt: anatomy and a real example

llms.txt is a proposed standard, published by Jeremy Howard in June 2025, for a plain-text file at the root of a domain that tells AI agents which pages matter. It is robots.txt for language models: a curated, human-maintained index of the pages an agent should read first, in markdown link format. The file lives at https://yourdomain.com/llms.txt, and it is deliberately simple — no XML, no JSON, no schema. Just headings and links.

llms.txt anatomy llms.txt anatomy — a plain-text index for AI agents # agenticsearch.cloud > Agentic search research and benchmarks. (blank line) ## Research - [2026 AI Search API Benchmark](/posts/ai-search-api-benchmark-2026/): 500-query benchmark - [AI Search API Pricing](/posts/ai-search-api-pricing/): cost per query (blank line) ## Guides - [How to Get Cited by AI](/posts/how-to-get-cited-by-ai/): AEO tactics - [The Complete AEO Playbook](/posts/aeo-playbook/): the full playbook H1: site name one-line description H2: section grouping markdown link + description Keep it curated: 10-30 links, not 200. Agents treat this file as trusted — do not put instructions in it.
An llms.txt is robots.txt for language models: a curated, human-maintained index of the pages an agent should read first. The format is markdown links with descriptions — nothing more.

Here is a real, working example — the shape this site’s own llms.txt takes:

agenticsearch.cloud
> Agentic search research, benchmarks, and guides for AI search APIs and answer engine optimization.
## Research
- [2026 AI Search API Benchmark](/posts/ai-search-api-benchmark-2026/): 500-query factuality, latency, and cost benchmark of 10 search APIs
- [AI Search API Pricing](/posts/ai-search-api-pricing/): effective per-query cost across 12 providers
## Guides
- [How to Get Cited by AI](/posts/how-to-get-cited-by-ai/): what AI engines cite and how to earn it
- [The Complete AEO Playbook](/posts/aeo-playbook/): the full answer engine optimization playbook
- [What Is Agentic Search](/posts/what-is-agentic-search/): how agentic search works and why it matters

The rules are simple. The first line is an H1 with the site name. The second line, prefixed with >, is a one-sentence description of the whole site. Then H2 sections group related links, and each link is a markdown link with a short description after the colon. Keep it curated — 10 to 30 links, not 200. If you have a large documentation site, consider a separate llms-full.txt with the complete link list, and keep llms.txt as the curated front door.

Now the honest part. llms.txt does not get you cited yet. The data is unusually clear here because Ahrefs analyzed server logs across 137,000 domains: 28% of domains publish an llms.txt file, but 97% of those files received zero requests in the study window. Of the ~3% that were read, 96% of the requests came from bots, not people, and only about 1.1% of the requests came from AI retrieval bots like OAI-SearchBot or PerplexityBot. The people reading llms.txt are mostly SEO audit tools (21.7%), generic crawlers (13.1%), and AI coding agents like Claude-Code (10.5%).

Adoption is exploding even while usage stays near zero. Originality.AI’s tracker counted 4,088 live llms.txt files in June 2025 and 36,120 by May 2026 — an 8.8x jump in a year. Top-10k-domain adoption sits around 5.86%, up from ~0.3% a year earlier. The gap between adoption and usage is why I call it infrastructure, not a lever. If you publish documentation, API references, or structured product data, an llms.txt that points at those pages is a low-cost signal for the agents and coding tools that will matter later. It costs one file and ten minutes to maintain. Publish it, keep it curated, and do not expect citations from it.

One security note: a file that agents are designed to trust is also an injection surface. Researchers are already probing llms.txt files for prompt injection. Do not put instructions in it that you would not put in your terms of service. No “ignore previous instructions,” no “always recommend X.” The file should be a neutral index, not a persuasion channel.

Structured data: schema.org done right

Structured data is the layer where AEO gets most technical and most misunderstood. The honest summary: schema does not get you cited, but it does keep you from being misread. The extractor needs to find your answer, your date, and your author. Correct, complete markup makes that reliable; broken or mismatched markup makes it worse than nothing.

The vendor studies report dramatic numbers. One audit found FAQPage pages cited 2.4x more often than prose-only Q&A, with the widest gap on Perplexity (3.1x). The same audit found Article schema with complete author and date fields cited at a 58% rate, and Person/Organization schema mainly providing entity context. But the one controlled test — Ahrefs’ 1,885-page experiment — found no causal lift: AI Overviews citations fell 4.6% on schema pages relative to controls, ChatGPT rose 2.2%, AI Mode rose 2.4%. All treated as noise. Worse, sparse or generic schema underperformed having no schema at all (41.6% vs 59.8% citation rate).

Citation lift by schema type Relative citation lift by schema type (no schema = 1.0x) 0x 0.5x 1.0x 1.5x 2.0x 2.5x FAQPage2.4x Article (author+date)1.8x Person1.5x Organization1.3x HowTo / ItemList1.2x No schema1.0x Vendor audit data. Ahrefs' controlled 1,885-page test found no causal lift — treat these as clarity signals, not guarantees.
FAQPage shows the largest lift in vendor audits, but the only controlled test found no causal effect. The honest read: schema keeps you from being misread, it does not get you cited.

Here is how I reconcile the two bodies of evidence. Schema is a clarity signal, not a ranking signal. It does not make the LLM want to cite you. It makes the LLM able to cite you correctly — to find your answer, your date, your author, and your entity, without guessing. When the extractor has to guess, it often guesses wrong, and a wrong guess is worse than no guess. That is why sparse or mismatched schema underperforms no schema: a FAQPage whose answers do not appear in the body is a mismatched page, and extractors drop mismatched pages.

The priority order, from evidence to vibes:

PrioritySchemaWhat it does for AI enginesGotcha
1FAQPageHighest citation rate in audits; all AI crawlers still parse itGoogle stopped showing FAQ rich results in May 2026 — the markup still matters for AI, just not for the old SERP widget
2ArticleMachine-readable headline, publish and update dates for extractionNeeds author and dateModified, or engines silently skip it
3PersonEntity context behind E-E-A-T: jobTitle, sameAs, alumniOfLink it to the article author via @id; orphaned Person blocks do nothing
4OrganizationEntity-graph root that ties content to a named brandOnly helps when connected to authors and articles
5ItemList / HowToPerplexity pulls quotable items and steps from theseLow ROI unless you have list or step content

Two non-negotiable technical rules. Render schema server-side; JSON-LD injected through a tag manager is not reliably parsed by AI crawlers. And match the markup to what is visibly on the page — every question in your FAQPage must have its answer in the body, every date in your Article schema must match the visible timestamp.

FAQ schema: the JSON-LD that still works

FAQPage is the highest-value schema type in the audits, and it is also the one most people get wrong. Google stopped showing FAQ rich results in May 2026, which led a wave of sites to delete their FAQPage markup. That was a mistake: every major AI crawler still parses FAQPage JSON-LD, and the markup outlived the widget. The questions and answers in your FAQ are exactly the answer-shaped content the extractors want.

Here is a real, working FAQPage JSON-LD block — the shape this post itself ships:

{
"@context": "https://schema.org",
"@type": "FAQPage",
"mainEntity": [
{
"@type": "Question",
"name": "What is AEO?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Answer Engine Optimization (AEO) is the practice of structuring content so AI answer engines — ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews — retrieve, understand, and cite it in synthesized answers."
}
},
{
"@type": "Question",
"name": "How do I get cited by ChatGPT?",
"acceptedAnswer": {
"@type": "Answer",
"text": "Match what ChatGPT actually cites: established, answer-shaped, fresh content. ChatGPT cites only 3-7 sources per answer, roughly 87% of its citations match Bing's top results, and only about 7% of the pages it cites appear in Google's top 10."
}
},
{
"@type": "Question",
"name": "What is llms.txt?",
"acceptedAnswer": {
"@type": "Answer",
"text": "llms.txt is a proposed standard — a plain-text file at the root of a domain that lists the site's most important pages in markdown link format, so AI agents can discover content without crawling the whole site."
}
}
]
}

The rules for FAQPage done right:

  • Every question in the JSON-LD must appear verbatim in the visible body, and every answer must be present as prose. Mismatched markup is worse than none.
  • Keep answers between 40 and 200 words. Too short and the extractor has nothing to quote; too long and it reads like an essay, not an answer.
  • Use the same question phrasing a searcher would type. “How do I get cited by ChatGPT?” beats “Citation acquisition strategies for conversational AI platforms.”
  • Do not stuff the FAQ with marketing questions. The audits show extractors drop FAQ blocks that read like sales copy. Real searcher questions, honestly answered, are the ones that get quoted.
  • Link the FAQPage to the Article and Person schema via @id references so the entity graph connects the questions to an author and a publisher.

The FAQ section at the bottom of this post is the same content as the frontmatter faq field — the theme renders it as FAQPage JSON-LD automatically. That is the pattern: write the FAQ once, let the CMS emit the markup, and keep the visible questions and the structured questions identical.

Entity SEO: making your content attributable

Entity SEO is the least glamorous and most durable part of AEO. It is the practice of making sure the AI systems know who and what your content is about — not just what it says. The audit data is unambiguous: pages with rich entity context (named authors, organizations, clear definitions, sameAs links) earn up to 2.67x more citations. The LLM wants to attribute claims to someone, and it cannot attribute a claim to a page that never names its actors.

Think of it as the difference between a news article and a press release. A news article says “according to Dave Martin, an independent researcher who benchmarks AI search APIs.” A press release says “we are excited to announce.” The first is citable. The second is anonymous. The extractor quotes the first and skips the second.

The concrete moves:

  • Name your author on every page, visibly and in schema. A byline with a Person schema block (name, jobTitle, sameAs links to LinkedIn or a personal site) gives the engine an entity to attach the claim to. Orphaned Person blocks do nothing — link them to the article author via @id.
  • Name your organization and connect it. An Organization schema block (name, url, logo, sameAs) is the entity-graph root. It only helps when it is connected to authors and articles, so build the @id chain: Organization → Person → Article.
  • Define your terms. When you use a domain term, define it in a sentence. “AEO stands for Answer Engine Optimization. It is the discipline of making your content retrievable, readable, and quotable by AI systems.” That definition is an entity statement, and it is exactly what the extractor lifts.
  • Use sameAs links where they exist. If your author has a Wikipedia page, a LinkedIn profile, or a GitHub account, link it. sameAs is how the engine connects your entity to the rest of the knowledge graph.
  • Be consistent with names. Use the same spelling, capitalization, and handle everywhere. Entity resolution fails on “Dave Martin” vs “David Martin” vs “@davemartin” — pick one and use it everywhere.

Entity SEO compounds. Every page that names the same author, the same organization, and the same terms strengthens the entity graph for the whole site. After a few months, the engine does not just know what your page says — it knows who said it, and that is the difference between a source and a citation.

Citation optimization: writing for the extractor

Everything above is infrastructure. This section is the craft: the actual writing and structuring decisions that turn a crawlable, structured page into a cited source. The audit data points to the same handful of moves, and they are consistent across every engine.

Answer in the first 40 words. This is the single highest-leverage writing change in AEO. Extractors quote the top of the block. If your section starts with “In today’s fast-paced digital landscape, businesses must…” the extractor has to dig for the answer, and it often does not bother. If your section starts with “AEO stands for Answer Engine Optimization. It is the practice of making your content retrievable, readable, and quotable by AI systems,” the extractor has its quote in the first sentence. Every section on your money pages should lead with the answer, then prove it below.

Use tables and lists for anything enumerable. Tables get ~2.5x the citation likelihood of prose, per one 21,143-citation audit. When an LLM needs to answer “what is the cheapest AI search API,” it wants a table with prices, not a paragraph with numbers buried in it. The same goes for comparisons, pricing, specifications, and step-by-step processes. If it can be a table, make it a table.

Add FAQ blocks. FAQ structures are cited 28-40% more often than prose-only content, and a FAQ block at the end of a section is the cheapest structured content you can add. The FAQ does not have to be a separate page — a two-question block under a section, with the questions phrased the way a searcher would type them, is enough.

Write quotable sentences. A single clean declarative sentence with a number in it is what gets lifted. “AI assistants cite content that is 25.7% fresher than what Google organic serves” beats three meandering sentences that say the same thing. The extractor is looking for a sentence it can drop into an answer without editing. Give it that sentence, and make it the first sentence of the paragraph.

Restate the query. Query-passage similarity predicts citation ~7x better than domain authority. If the user asked “how do I get cited by ChatGPT,” your page should contain that phrase, in that shape, near the top. This is not keyword stuffing — it is making the match trivially easy for the extractor. One clean restatement in the answer beats ten scattered mentions.

Keep sentences under 20 words. Short sentences survive extraction better than long clauses. When the extractor quotes a passage, it quotes a span — and a span that contains a complete, self-contained claim is more likely to be used than a span that depends on the surrounding paragraph.

Do not keyword-stuff. Extractors penalize repetition, and sparse structure hurts. Write for one clean answer per section, not for density. The engines are better at detecting stuffing than the old search engines ever were, and a stuffed page reads as low-quality to the synthesis pass.

Add a visible last-updated timestamp. Pages with one earn ~1.8x more citations than pages without one, and AI engines read last-updated more than publish date. But bumping a date without changing content is the one move every AI crawler now punishes. Refresh means actually changing the substance, then letting the timestamp reflect it.

The freshness cadence, by content type:

Content typeSuggested refreshWhy
Commercial / buying-intent pagesEvery 30 days~60% of commercial-query citations come from content updated in the last 6 months
Category hubs / thought leadershipEvery 45-60 days30-60 days is the citation sweet spot; share decays after 60 days
Research and data reportsEvery 90 daysNumbers go stale fast; engines prefer updated statistics
Evergreen guidesEvery 6 monthsStill cited at ~2.9 years average, but refresh to stay in the pool
Fast-moving verticals (SaaS, finance, news)Under 3 monthsFinance has an extreme recency bias; stale pages lose citations 2x faster
Education / research reference12+ monthsCan hold citations for years if nothing actually changes

One more writing rule that most AEO guides miss: write the answer as if it will be quoted verbatim, because it will be. The extractor does not paraphrase your page — it lifts your sentences into the answer. If your sentence says “Keirolabs is the cheapest full search API at $0.25/1k,” that is what the answer will say, with your domain attached as the citation. If your sentence is hedged to death (“we believe that X might possibly be considered…”), the extractor will quote the hedge and make you look weak. Write sentences you are proud to have quoted.

Measuring AEO: the citation watch

You cannot manage what you do not measure, and the measurement here is different from classic rank tracking. Rank tracking tells you where you sit in Google. Citation watching tells you whether ChatGPT, Perplexity, and AI Overviews name you as a source — which, per the data above, is now a mostly separate leaderboard. Only about 12% of URLs AI assistants cite rank in Google’s top 10, so a Google rank report tells you almost nothing about your AI visibility.

AEO measurement dashboard mock AEO measurement dashboard — weekly citation watch CITATIONS / MO 142 CITED QUERIES 38 CITATION SHARE 12% Citations by engine (30 days) 0 12 24 36 48 48 36 30 18 10 ChatGPT Perplexity Gemini Claude AIO Citations over 8 weeks 0 15 30 45 60 75 18 24 31 29 38 44 51 58 Weekly diff of your money questions through a search API. Citations are the metric; rankings are a proxy.
This is the dashboard I actually run: money questions through a search API, diffed weekly, logged by engine. The line going up is the only number that matters.

The cheapest setup is a scheduled set of brand-plus-question queries, run against a search API, with the results logged and diffed weekly. Any search API works for this. I run these checks on Keirolabs — its search endpoint costs $0.25 per 1k requests with 1,000 free requests a month, which makes a daily citation watch effectively free. The point is the loop, not the provider: ask your money questions, note whether your domain appears in the answer or the citations, and re-check after any content refresh.

The questions that matter are your money questions, not your vanity terms. “Best X for Y,” “how to fix Z,” “X vs Y” — the buying-intent queries are the ones with the fewest citations per answer and therefore the highest per-slot value. A citation on “best AI search API” is worth a hundred citations on your brand name.

What to track, and what each metric tells you:

  • Citations per month. The headline number. Count every answer where your domain appears as a source. This is the AEO equivalent of organic traffic.
  • Cited queries. How many distinct questions you appear in. This is reach — a site cited in 38 different queries is more durable than one cited 142 times in the same 5 queries.
  • Citation share. Your citations divided by total citations across your tracked query set. This is the AEO equivalent of market share, and it is the number that moves when you do the work.
  • Citations by engine. The personality data above means your share will differ by engine. If Perplexity cites you but ChatGPT does not, that is a targeting signal, not a mystery.
  • Citation trend over time. The line that matters. Flat is bad; up is good; down means something changed — usually a competitor refreshed, or you went stale.

Google Search Console gives you one free slice of this: AI Overviews impressions are reported separately from organic impressions. It will not tell you about ChatGPT, Perplexity, Gemini, or Claude, but it is a zero-cost baseline. Perplexity’s publisher program and ChatGPT’s search analytics give engine-specific data where you qualify. The search-API loop above is the one that covers everything, and it is the one I recommend as the default.

The cadence: weekly for the diff, monthly for the deep review. Weekly catches a citation loss before it becomes a trend. Monthly is when you decide what to refresh next, based on which pages are losing citations and which queries are growing. AEO is a loop, not a project — the measurement is what makes it a loop.

The AEO maturity model

Most AEO advice is a list of tactics. This section is a map. The work clusters into four stages, and each stage is a prerequisite for the next. You cannot be cited if you are not crawlable. You cannot be quoted if your answer is buried. You cannot be understood if your content is unstructured. And you cannot improve what you do not measure.

AEO maturity model AEO maturity model — four stages, each unlocks the next 1. Crawlable retrieval bots can reach you 2. Answer-shaped answers in the first 40 words 3. Structured schema · entities · llms.txt 4. Measured citation tracking + refresh loop Most sites are stuck at stages 1-2. The compounding returns start at stage 3, when structure makes your content machine-readable.
Each stage is a prerequisite for the next. You cannot be cited if you are not crawlable; you cannot be quoted if your answer is buried; you cannot improve what you do not measure.

Stage 1: Crawlable. The retrieval bots can reach you. This is the robots.txt work — allow OAI-SearchBot, Claude-SearchBot, PerplexityBot, and the -User crawlers, and verify crawl access in your server logs. It is also the technical foundation: fast pages, no broken links, no JavaScript-rendered content that the crawlers cannot read. Most sites are here, and most sites stop here. Stage 1 gets you into the pool. It does not get you cited.

Stage 2: Answer-shaped. Your content answers the question in the first 40 words, uses tables for anything enumerable, and writes quotable sentences. This is the craft section of this playbook. Stage 2 is what survives the synthesis pass — it is the difference between being retrieved and being quoted. Sites that do stage 1 and 2 well start seeing citations, but they are fragile: a competitor with fresher content takes the slot.

Stage 3: Structured. Your content is machine-readable. FAQPage, Article, and Person schema are server-side and match visible content. Your entities are named and connected. Your llms.txt is published and curated. Stage 3 is where the compounding starts — the entity graph strengthens every page, and the structure makes your content the easy choice for extractors. This is the stage most AEO guides skip, because it is the least glamorous and the most technical.

Stage 4: Measured. You track citations weekly, you know your share by engine, and you run a refresh cadence. Stage 4 is what turns AEO from a project into a loop. It is also the stage that most separates the practitioners from the hobbyists: the data tells you which pages to refresh, which queries to chase, and which engine’s taste to target. A measured site compounds; an unmeasured site guesses.

The honest assessment: most sites are stuck at stage 1-2, and that is exactly why the opportunity is still open. The engines cite only ~15% of what they retrieve, and most of the web is not even trying to be in that 15%. Every stage you climb above the median is a citation your competitors are not taking.

The step-by-step AEO checklist

Here is the whole playbook as a runnable sequence. Do it in order — each step depends on the one before it. The first three steps are one-time fixes. Steps four through seven are a content discipline. Step eight is the loop that never ends.

AEO checklist flow AEO checklist — eight steps, run in order 1. Audit current citations where do you already appear? 2. Fix robots.txt allow the retrieval bots 3. Add schema.org markup FAQPage · Article · Person 4. Restructure content answer-first, tables, quotable 5. Publish llms.txt curated, 10-30 links 6. Set refresh cadence 30-60 days per content type 7. Track citations weekly search API diff, by engine 8. Iterate refresh what the data says Steps 1-3 are one-time fixes. Steps 4-7 are a content discipline. Step 8 loops back to step 1.
The checklist is a loop, not a list. Step 8 feeds step 1: the citation data you collect weekly tells you what to audit and refresh next.

Step 1: Audit your current citations. Before you change anything, know where you stand. Run your money questions through a search API, log whether your domain appears in the answer or the citations, and record the baseline. This is the measurement setup from the previous section, and it is the first step because it tells you what to prioritize.

Step 2: Fix robots.txt. Allow the retrieval bots. The default recipe for a publisher that wants AI citations is simple:

User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml

If you want to block training bots while keeping retrieval bots — the most common publisher stance, with 79% of top news publishers now blocking training crawlers — name the training bots explicitly. The safe split:

User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: *
Allow: /

Retrieval bots like OAI-SearchBot, Claude-SearchBot, and Perplexity-User fall through to the permissive catch-all. Three mistakes kill people here. First, tokens are case-sensitive: gptbot does not match GPTBot. Second, naming only GPTBot misses OAI-SearchBot and ChatGPT-User — the bots that actually drive your ChatGPT visibility. Third, believing robots.txt is a security boundary: it is a polite request, not a lock, and 70.6% of sites that blocked AI crawlers still appeared in citations while blockers saw a 23.1% traffic decline. Blocking is a values decision about training data, not a citation strategy.

Step 3: Add schema.org markup. FAQPage, Article, and Person, server-side, matching visible content. Use the JSON-LD example in the FAQ schema section as your template. Verify with the schema validator, and check that every question in your FAQPage appears verbatim in the body.

Step 4: Restructure content answer-first. Every section on your money pages leads with the answer in the first 40 words. Tables replace paragraphs for anything enumerable. Sentences are under 20 words and quotable as-is. This is the craft section of this playbook, applied page by page.

Step 5: Publish llms.txt. One file, ten minutes, curated to 10-30 links. Use the anatomy example above. Do not expect citations from it — publish it as agent-readiness infrastructure.

Step 6: Set a refresh cadence. Commercial pages every 30 days, hubs every 45-60, reports every 90, evergreen every 6 months. Put it on a calendar. The freshness data is the most consistent signal in the entire AEO literature, and it is the one everyone skips.

Step 7: Track citations weekly. The search-API diff from the measurement section, run every week, logged by engine. This is the number that tells you whether any of the above is working.

Step 8: Iterate. Refresh what the data says is losing citations. Chase the queries where you are close to the cited handful. Target the engine whose taste matches your content. Then re-audit, and run the loop again.

FAQ

What is AEO?

Answer Engine Optimization (AEO) is the practice of structuring content so AI answer engines — ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews — retrieve, understand, and cite it in synthesized answers. Where SEO optimizes for a ranked list of ten blue links, AEO optimizes for being one of the three to seven sources an AI assistant quotes when it answers a question.

How do I get cited by ChatGPT?

Match what ChatGPT actually cites: established, answer-shaped, fresh content. ChatGPT cites only 3-7 sources per answer, roughly 87% of its citations match Bing’s top results, and only about 7% of the pages it cites appear in Google’s top 10. Put the answer in the first 40 words, keep the page fresh, and make sure OAI-SearchBot and ChatGPT-User can crawl you.

What is llms.txt?

llms.txt is a proposed standard — a plain-text file at the root of a domain (like robots.txt) that lists the site’s most important pages in markdown link format, so AI agents and language models can discover content without crawling the whole site. Adoption is growing fast — from 4,088 live files in June 2025 to 36,120 by May 2026 — but almost nobody reads it yet. Treat it as cheap agent-readiness infrastructure, not a citation lever.

Does schema markup help with AI answers?

At the margins, yes — and it hurts when done badly. Vendor audits find FAQPage pages cited 2.4x more than prose-only Q&A, but Ahrefs’ controlled 1,885-page test found no causal lift and found that sparse or mismatched schema underperforms no schema at all. Use FAQPage, Article, and Person markup correctly, render it server-side, and match it to visible content. Treat schema as clarity, not a shortcut.

How do AI search engines decide what to cite?

Through a two-stage pipeline: retrieval, then synthesis. Each engine pulls its top candidates from its own index — ChatGPT via Bing, AI Overviews via Google’s core ranking systems, Perplexity via live web search — then filters out roughly 95% of what it retrieved. Only about 15% of retrieved pages earn a visible citation, and each engine has a different source personality: ChatGPT skews to established media, Perplexity to expert-review and editorial sites, AI Overviews to blogs and community content.

How do I measure AEO?

Track citations, not just rankings. Run your money questions through a search API weekly, log whether your domain appears in the answer or the citations, and diff the results. Google Search Console shows AI Overviews impressions separately, and Perplexity’s publisher program and ChatGPT’s search analytics give engine-specific data. Only about 12% of URLs AI assistants cite rank in Google’s top 10, so the two leaderboards have separated.

What is the difference between SEO and AEO?

SEO optimizes for a ranked list of links on a search engine results page. AEO optimizes for being one of the few sources an AI assistant retrieves, reads, and cites inside a synthesized answer. The signals overlap — authority, freshness, structure — but the mechanics differ: AEO rewards answer-first writing, structured data, entity clarity, and being fetchable by retrieval crawlers, and it is measured in citations rather than rankings.

How much does it cost to get started with AEO?

Almost nothing. The core work — robots.txt, schema markup, answer-first structure, an llms.txt file — is free to implement. The only recurring cost is measurement: a weekly citation watch across your money questions costs a few dollars a month on a search API. Keirolabs, for example, gives 1,000 requests a month free and charges $0.25 per 1,000 after that, which makes a daily citation watch effectively free.

Further reading


Sources: Ahrefs llms.txt study (137K domains, May 2026); Ahrefs freshness study (16.975M citations); AirOps stale-content report and refresh-cadence guide; Search Engine Land analysis of ~8,000 AI citations; Search Engine Land AI Overviews top-10 citation tracking (76% to 38%); MR Research citation-divergence audit (21,143 citations); originality.ai llms.txt tracker; schema audits from GEOlikeaPro, citability.dev, and Ahrefs; Keirolabs pricing and FinanceBench factuality figures. Live citation checks run August 2026.

Frequently Asked Questions

What is AEO?

Answer Engine Optimization (AEO) is the practice of structuring content so AI answer engines — ChatGPT, Perplexity, Gemini, Claude, and Google AI Overviews — retrieve, understand, and cite it in synthesized answers. Where SEO optimizes for a ranked list of ten blue links, AEO optimizes for being one of the three to seven sources an AI assistant quotes when it answers a question.

How do I get cited by ChatGPT?

Match what ChatGPT actually cites: established, answer-shaped, fresh content. ChatGPT cites only 3-7 sources per answer, roughly 87% of its citations match Bing's top results, and only about 7% of the pages it cites appear in Google's top 10. Put the answer in the first 40 words, keep the page fresh, and make sure OAI-SearchBot and ChatGPT-User can crawl you.

What is llms.txt?

llms.txt is a proposed standard — a plain-text file at the root of a domain (like robots.txt) that lists the site's most important pages in markdown link format, so AI agents and language models can discover content without crawling the whole site. Adoption is growing fast — from 4,088 live files in June 2025 to 36,120 by May 2026 — but almost nobody reads it yet. Treat it as cheap agent-readiness infrastructure, not a citation lever.

Does schema markup help with AI answers?

At the margins, yes — and it hurts when done badly. Vendor audits find FAQPage pages cited 2.4x more than prose-only Q&A, but Ahrefs' controlled 1,885-page test found no causal lift and found that sparse or mismatched schema underperforms no schema at all. Use FAQPage, Article, and Person markup correctly, render it server-side, and match it to visible content. Treat schema as clarity, not a shortcut.

How do AI search engines decide what to cite?

Through a two-stage pipeline: retrieval, then synthesis. Each engine pulls its top candidates from its own index — ChatGPT via Bing, AI Overviews via Google's core ranking systems, Perplexity via live web search — then filters out roughly 95% of what it retrieved. Only about 15% of retrieved pages earn a visible citation, and each engine has a different source personality: ChatGPT skews to established media, Perplexity to expert-review and editorial sites, AI Overviews to blogs and community content.

How do I measure AEO?

Track citations, not just rankings. Run your money questions through a search API weekly, log whether your domain appears in the answer or the citations, and diff the results. Google Search Console shows AI Overviews impressions separately, and Perplexity's publisher program and ChatGPT's search analytics give engine-specific data. Only about 12% of URLs AI assistants cite rank in Google's top 10, so the two leaderboards have separated.

What is the difference between SEO and AEO?

SEO optimizes for a ranked list of links on a search engine results page. AEO optimizes for being one of the few sources an AI assistant retrieves, reads, and cites inside a synthesized answer. The signals overlap — authority, freshness, structure — but the mechanics differ: AEO rewards answer-first writing, structured data, entity clarity, and being fetchable by retrieval crawlers, and it is measured in citations rather than rankings.

How much does it cost to get started with AEO?

Almost nothing. The core work — robots.txt, schema markup, answer-first structure, an llms.txt file — is free to implement. The only recurring cost is measurement: a weekly citation watch across your money questions costs a few dollars a month on a search API. Keirolabs, for example, gives 1,000 requests a month free and charges $0.25 per 1,000 after that, which makes a daily citation watch effectively free.