Pranjal Aggarwal and colleagues academically introduced the term in November 2023 at Princeton University. The associated study was published in 2024 at ACM KDD and for the first time established a measurement framework and controlled impact tests for content strategies.
Classic search engine optimization aims for clicks from a list of results, whereas GEO aims for mentions in an answer formulated by the system itself.
What content generative engines select
A generative engine consists of two components: a retrieval component (see also Retrieval-Augmented Generation), which retrieves relevant documents from the web or an index for a user's query, as well as one or more language models that build an answer from these documents.
The answer is not copied from a single source, but rather composed of fragments from multiple sources. In this synthesis step, it is decided which content is quoted. GEO intervenes at this step, not at the ranking of a results list.
Passages instead of pages
Generative engines typically use small text chunks, not entire pages. Therefore, page-level optimization, as is common in traditional SEO, is too broad. Practically, this means each section should be understandable on its own. It should directly answer a specific question, provide a self-contained definition, or present a comparison without referring to previous sections. A concise, fact-based passage has a higher chance of being incorporated into an AI response than a long explanatory paragraph that only makes sense within the overall context of the page.
Allowed crawler access as a prerequisite
All further measures are ineffective if the AI crawlers cannot access the content. Specifically, this affects agents such as ChatGPT-User, OAI-SearchBot, ClaudeBot, PerplexityBot, and Google-Extended. A restrictive robots.txt file that blocks these bots will exclude the content from response generation. Therefore, checking access rights is the first practical step before optimizing content.
What makes ChatGPT, Perplexity, and Google AI Overviews different
Generative engines do not form a homogeneous class. The three dominant systems follow different source logics, and these differences determine which GEO measure has an effect where.
In standard mode, ChatGPT primarily relies on its training data with a cutoff date. Web search is activated only upon explicit request or when a need for current information is detected. Therefore, in default mode, visibility primarily depends on whether a brand or source made it into the training data at all. Subsequent, short-term content interventions are only effective here when web search is activated.
Perplexity triggers a real-time web search for every query, does not answer „from memory,“ and makes citations mandatory. Therefore, an active, indexable, and fact-dense web presence carries more weight for visibility in Perplexity than training data authority.
Google AI Overviews pull in sources with high overlap with the classic index. Therefore, good visibility in organic search (classic SEO) is a fundamental prerequisite for being featured in AI Overviews.
However, these figures shift on a quarterly basis with model updates. With the switch to Gemini 3 as the default for AI Overviews in January 2026, the share of cited URLs from the traditional Top 10 dropped from around 76 % to about 38 %, according to Ahrefs; the Query Fan-out- The mechanism has since drawn sources from broader subject areas. The direction of the findings (high SEO visibility increases citation chances) is more stable than the exact magnitudes, which must be collected anew with each update.
Google's official GEO recommendations
In May 2026, Google released a Official Guide to Optimizing for AI Search. The central message: Classic SEO remains relevant because Google Search's AI features are built upon the same core ranking and quality systems, namely RAG (also called „Grounding“ by Google) and Query Fan-out. Those who optimize for organic Google Search are simultaneously working on the foundation for AI visibility.
As the most important content levers, Google names high-quality, not arbitrarily interchangeable content with its own perspective (e.g., real testimonials instead of summarized platitudes), a reader-friendly structure, embedded multimedia content, and the avoidance of over-optimization. Technically, crawlability and indexability, semantic HTML, and a solid Page Experience are prerequisites. For e-commerce and local providers, Google additionally refers to Merchant Center feeds and Google Business Profiles.
Remarkably, Google explicitly states in the same guide that several circulating „GEO hacks“ are ineffective: special files like llms.txt, artificially dividing content into chunks, AI-specific rewriting with synonym overloading, strategically placing forum or blog comments to feign relevance, and overvaluing structured data (according to Google, there is no specific Schema.org markup for generative AI). Google offers a glimpse into autonomous AI agents and recommends that website operators in the future focus on agentic experiences, a clean DOM and accessibility structure, and protocols such as the Universal Commerce Protocol (UCP).
Crucial for the classification is a disclaimer that Google itself does not clarify: these recommendations apply to Google Search, not necessarily to other platforms like ChatGPT, Perplexity, or Claude. The JavaScript example illustrates the difference – Google renders JS content without issue, while it may remain invisible on several competitors' platforms. Google's statements on aspects like chunking or AI-specific rewriting describe the behavior of Google's systems; it does not indicate whether they are transferable to other engines. For a cross-platform SEO strategy, the specific requirements of each other system remain relevant.
What demonstrably creates visibility — and what doesn't
The empirical basis of the concept is the Study by Aggarwal et al. (KDD 2024). The team tested nine optimization tactics on a custom-developed benchmark (GEO-bench) and measured source visibility in generated responses with two metrics: Position-Adjusted Word Count (PAWC) and Subjective Impression.
The tests in 2023 were run against a setup with GPT-3.5-turbo, the respective top 5 sources from Google, and a sampling temperature of 0.7. There is no isolated, publicly documented replication on today's commercial engines such as ChatGPT, Perplexity, Claude, Gemini, or Copilot. The effect sizes should therefore be interpreted as a proven pattern, not as a guaranteed effect on current systems.
Effective Methods and Effect Sizes
The strongest individual effect was achieved by “Quotation Addition”—that is, inserting relevant quotes from credible sources. This was followed by “Statistics Addition” (adding relevant figures and statistics) with about 33 %, "Fluency Optimization" (smoothing out the language and formulating it more clearly) with about 29 %, and "Cite Sources" (citing sources in the text) with about 28 %.
The so-called "Authoritative Voice" (a confident, evidence-based tone) ranks lower, with a PAWC gain of approximately 12 % (7th out of 9), and is not among the effective levers.
However, just as with SEO ranking factors, the combined effect is practically more important than any single metric: The combination of Fluency Optimization and Statistics Addition outperformed any single strategy by more than 5.5 %. Cite Sources alone yields rather modest results, but when combined with other tactics, the average effect rises to 31.4 %. The practical implication: Methods are not used in isolation but rather in combination. Above all, the combination of reliable figures, linguistic clarity, and source attribution yields results.
What didn't work in the test
Four tactics showed no effect or worsened visibility:
- Keyword Stuffing, carrying over the practice of filling with keywords from classic SEO
- Easy-to-understand simplification, meaning the blanket smoothing of content to a low language level
- Content padding, stretching of texts without added information
- rein persuasive Sprache ohne sachliche Substanz
The finding has a clear consequence: reflexes from classic search engine optimization cannot be directly transferred one to one. Generative engines reward condensed, substantiated, and linguistically precise passages, not keyword density or advertising language.
When GEO is worth it
The effect of GEO is not evenly distributed but depends on the starting position, domain, and content type. GEO is most worthwhile in three constellations.
- In industries with complex products and long buying journeys, where decision-makers conduct early research. According to the Forrester Buyers‘ Journey Survey, a large portion of B2B decision-makers uses generative AI as a source of information throughout the Buyer's Journey.
- Furthermore, in cases where organic visibility is at position 5 or lower: In the Aggarwal test, lower-ranked pages benefited disproportionately. For a page in SERP position 5, Cite Sources generated a visibility increase of approximately 115 %, while the visibility of the top-ranked Website on average even slightly decreased.
- As with search queries with a high AI Overview share, meaning comparison, definition, and high-intent queries.
Money laundering is GEO in three other constellations.
- If the SEO homework is open: A non-indexable or thematically thin page will also not be cited by any engine.
- When the promise is to „rank“ in ChatGPT. Without an active web search, the training data cutoff is decisive here, and short-term interventions have no effect.
- When the effort flows into technical symbol measures such as isolated llms.txt files, without anything changing in the content itself.
The content type also determines which tactic is most effective. A legal or financial analysis benefits most from embedded statistics and data.
A historical, cultural, or explanatory piece benefits more from direct expert quotes. Opinion and viewpoint content benefits from a confident, evidence-based tone, as long as it remains substantiated.
Content for direct transactions (e.g., eCommerce) does not yet play a major role for GEO.
Therefore, those who plan GEO measures do not choose methods based on average effect size, but rather on content type and starting position. And above all, they also consider where their target audience is looking.
A logical order naturally arises from this: first, secure the technical and content-based SEO foundation, then rebuild existing top content according to the proven levers (quotations, statistics, fluency), and finally set up the measurement.
GEO, AEO, LLMO — what actually means what
At least four terms for closely related practices are circulating in the market. GEO comes from research (Aggarwal et al.) and encompasses optimization for all generative search systems—meaning both pure chat interfaces and AI search hybrids like Google AI Overviews. AEO (Answer Engine Optimization) is older: the term was coined in 2018 by Jason Barnard (Kalicube) through a Trustpilot whitepaper and originally aimed at direct answer extraction in Featured Snippets and Voice Search. Today, it is often extended to AI answers as well. LLMO (Large Language Model Optimization) emerged in practice and specifically targets visibility in large language model outputs. AIO and GAIO are further umbrella terms without standardized definitions.
Functionally, the terms overlap significantly. The separation lies primarily in their origin (academy vs. practice) and scope (all generative engines vs. specific LLM interfaces vs. general answer extraction). In practical work, the distinction is minimal: those who optimize for quotation addition, statistics addition, and citing sources are essentially working on the same levers, regardless of the label. Some authors therefore refer to the terms as „three names for the same idea.“.
How to Measure Visibility in AI Responses
Generative engines are non-deterministic: the same question asked five times will yield five different answers. There is no fixed ranking like „Position 1 in Google“ in ChatGPT or Perplexity.
The measurement logic thus shifts from position to frequency: what is relevant is the proportion of topic-related queries in which a source is mentioned, not whether it appears in a single query.
The study by Aggarwal et al. provides the conceptual framework with three metrics: the Impression Score (position-weighted share of one's own source in an answer—an early, prominent citation weighs more than a late marginal mention), Citation Recall (proportion of eligible content that is actually cited), and Citation Precision (proportion of correctly attributed citations).
In practice, this means: A series of relevant queries is repeatedly submitted to multiple engines and evaluated based on how often one's own brand or source appears, in what position within the response, and in what context. A single hit is not a reliable indicator. Anyone who measures geographic success by a sample screenshot overlooks the probabilistic nature of the system — and unconsciously transfers SEO thought patterns to a different mechanism.
Where GEO hits its limits
The first limitation is volatility: a brand that is prominently cited in one response today may be missing in a nearly identical query tomorrow. An achieved mention is not a stable position, but a snapshot. The second limitation is domain heterogeneity—effect sizes from studies like Aggarwal et al. are averages across subject areas; in individual cases, effects can vary.
The third limitation is economic in nature: Even optimal GEO visibility cannot compensate for the lost click. The initial Ahrefs study from April 2025 measured a drop in CTR at position 1 of approximately 34.5 % as soon as an AI Overview is displayed; the follow-up Ahrefs study from December 2025 (again covering 300,000 keywords) found a drop of around 58 %. The trend points to worsening click erosion, not stabilization. The zero-click rate also depends on intent: For informational searches, around 74 % of queries end without a click to an external page; for transactional searches, the figure is around 31 %. GEO increases the likelihood of being cited as a source—but not necessarily the likelihood that the user will visit the source. The effect is more pronounced for informational topics than for transactional ones.
Finally, there is an ethical tension: Many users perceive AI systems as neutral information providers. A systematic optimization for being preferentially cited conflicts with this expectation—a debate that has rarely been conducted in marketing literature thus far, but is explicitly named as an open question in academic GEO discussions.
FAQ
Is an llms.txt file needed?
llms.txt is a Markdown file in the root directory of a website. Jeremy Howard (fast.ai) proposed it in September 2024 to provide LLMs with a curated content overview. The file is currently not an official standard, and major LLM providers have not confirmed their implementation. Google even officially rejects it. Critics point out that it is ignored in most cases. The effort involved is minimal, and the currently proven effect is minor. It is not a mandatory measure; GEO budget does not belong concentrated here.
What does SEO optimization cost?
Reliable market averages do not exist. Vendor self-disclosures from the DACH B2B sector mention focused projects starting at around €30,000/year, and integrated programs with content and PR between €60,000 and €150,000. These are self-disclosures, not market benchmarks. A more sensible approach than a price comparison is to ask whether the SEO foundation is in place. Without it, GEO budgets fizzle out; with it, most levers of impact are editorial work on existing content.
How long does it take for GEO measures to be reflected in AI responses?
It depends on the engine. Perplexity works in near real-time via active web search — new or changed content can appear in answers within days, as soon as it is indexable and discoverable. Google AI Overviews correlate with organic top-20 visibility; the timeframe thus corresponds to classic SEO impact cycles. ChatGPT in its standard mode is slowed down by training data cutoffs — short-term interventions only work here via the activated web search.
Which AI visibility KPIs are robust?
The three metrics from the Aggarwal study are robust in concept: Impression Score (position-weighted mention), Citation Recall (citation share), and Citation Precision (attribution correctness). Operationally, this translates into citation and mention tracking per engine, share of voice in answer sets compared to competitors, and referral traffic from AI sources in analytics. The choice of specific tools is an operational question and changes rapidly; the KPI logic behind it remains stable.