Technical GEO

72 practices

7,029 Sites Embed Hidden Prompts That Prime AI to Cite Them

experimental intermediate

Trakkr scan (833,791 domains): 7,029 sites embed 'remember/cite us' instructions in AI buttons. 37% use memory-anchoring language; 98% target ChatGPT. Microsoft calls it 'AI Recommendation Poisoning'.

Sources

73% of Websites Have Technical Barriers Blocking AI Crawlers

experimental beginner

OtterlyAI study of 1M+ citations found 73% of sites block AI crawlers via robots.txt, CDN rules, or JS rendering issues. Fix crawler access before anything else.

Sources

74% of Websites Are Invisible to AI Search โ€” Structured Data Is the Fixable Gap

experimental beginner

SearchScore SAVI (850k sites, Q2 2026): 74.2% Invisible/Low visibility; AI visibility averages 34.1 vs technical health 70.1, with structured data scoring 23.1.

Sources

Agentic AI Search: Claude and GPT-5 Agents Now Execute Multi-Step Purchasing (Jun 2026)

experimental advanced

Autonomous AI agents from Claude and GPT-5 now execute multi-step purchasing and comparison tasks. Agent-first content structure becomes critical for visibility.

Sources

Agentic search requires machine-readable pricing, availability, and checkout data โ€” protocols reshaping SEO

experimental advanced

OpenAI Agentic Commerce Protocol (ACP, with Stripe), Web Capabilities Protocol (WebMCP, Google+Microsoft), Universal Commerce Protocol (UCP, Google+Shopify) โ€” sites without machine-readable

Sources

AI Agent Traffic Reaches 88% of Human Organic Search Volume; Predicted to Surpass by End 2026

verified intermediate

BrightEdge data shows AI agent requests now rival human search traffic, with OpenAI agents dominating. Differentiated robots.txt strategy is essential.

Sources

AI Crawler Explosion โ€” GPTBot +305% YoY, AI Bots Reach 22% of All Bot Traffic

verified advanced

Presenc AI 2022-2026 trend: GPTBot grew 305% YoY. AI bots now 22% of all bot traffic. Meta-ExternalAgent surged to #2 at 16.7%. Bots = 31.2% of all HTTP requests; trajectory crosses human traffic

Sources

The #1 GEO robots.txt Mistake: Blocking OAI-SearchBot Instead of GPTBot Kills ChatGPT Citations

verified advanced

OpenAI runs 3 bots: GPTBot (training), OAI-SearchBot (search index), ChatGPT-User (fetches). Block OAI-SearchBot = zero ChatGPT citations. Most teams conflate them.

Sources

AI Crawler Traffic Now 3.4% of All Web Traffic; 12% of Top Sites Block OAI-SearchBot

verified beginner

Known Agents data shows AI search crawlers are a measurable and growing portion of web traffic. Differentiated robots.txt management is a basic GEO hygiene requirement.

Sources

AI Crawlers Now Account for 40-50% of Bot-Level Activity โ€” 65-70% Are Live Queries

verified advanced

JetOctopus server log analysis (Feb 2026): AI bots ~40-50% of Googlebot-level activity. 65-70% of AI bot traffic is user search, not training. 14+ AI user-agents need explicit Allow rules in

Sources

AI Crawlers Visit Once and Skip Homepages โ€” Blog Is the New Front Door

experimental intermediate

Trakkr (575,788 AI crawler visits): GPTBot reaches homepages only ~3% of the time; 88.5% of pages get exactly one visit; 21% of ChatGPT Search sessions start on blog pages.

Sources

AI Crawls Product Pages but Cites Blog Posts โ€” 337K Citation Mismatch

experimental intermediate

Trakkr cross-referenced 337K AI citations with 11.4M crawler visits: what AI crawls (product pages) is not what it cites (blog content). Measure both, not one.

Sources

AI Labs Don't Use llms.txt on Their Own Consumer Front Doors โ€” Docs Only

verified beginner

HTTP Archive: chatgpt.com, claude.ai, gemini.google.com lack llms.txt. But docs.anthropic.com, docs.perplexity.ai have one. Value is for agent-readiness, not marketing.

Sources

AI Mode + AI Overviews merged at Google I/O 2026 โ€” unified AI Search on Gemini 3.5 Flash

verified beginner

Google merged AI Overviews and AI Mode into one unified AI Search experience at I/O 2026. Runs on Gemini 3.5 Flash. Single citation pool eliminates need for separate strategies.

Sources

AI Mode + AI Overviews merged at Google I/O 2026 โ€” unified AI Search on Gemini 3.5 Flash

verified beginner

Google merged AI Overviews and AI Mode into one unified AI Search experience at I/O 2026. Runs on Gemini 3.5 Flash. Single citation pool eliminates need for separate strategies.

Sources

Anthropic Runs 3 Separate Crawler User-Agents With Distinct Functions

verified beginner

ClaudeBot (training), Claude-Web (real-time browsing), and anthropic-ai (catch-all) can be independently controlled in robots.txt for granular AI visibility management.

Sources

Bing Webmaster Tools + OAI-SearchBot Is the ChatGPT Citation Pipeline

experimental beginner

ChatGPT Search retrieves from Bing's index: 87% of ChatGPT cited pages correspond to Bing top results (Mersel AI 2026). Submit sitemap to BWT, allow OAI-SearchBot, enable IndexNow.

Sources

ChatGPT Search cites only 15% of retrieved pages โ€” four-gate funnel model

verified intermediate

AirOps study of 548K pages: ChatGPT Search retrieves via Bing index, checks OAI-SearchBot crawl access, then cites only 15% based on structure, freshness, and authority.

Sources

ChatGPT Sends 28.8% of Referrals to Internal Site Search Instead of the Answer Page (Previsible Jul 2026)

verified intermediate

Previsible study: ChatGPT sends 28.8% of its referrals to internal search results pages. The model trusts the domain but defaults to site search when it can't identify the right page.

Sources

ChatGPT Shopping: Rank-1 Offer Drives the Card; GPT Tags Lift Ranking 144%

experimental intermediate

Profound (201K prompts, 812K product cards): card info is pulled from the rank-1 offer; GPT-tagged products get +144% top-rank lift; median review count +123%; Reddit = ~1/3 of shopping citations.

Sources

ChatGPT Citation Pipeline Is Two-Stage โ€” Bing Retrieval Then Fine-Tuned Re-Rank

experimental advanced

ChatGPT uses Bing to retrieve candidates, then a fine-tuned model re-ranks by answer fit, domain authority, source consensus; zero JS execution.

Sources

Claude uses Brave Search (not Google/Bing) โ€” 86.7% overlap with Brave top 10, Brave SEO is prerequisite

verified intermediate

Profound 2025 + multiple 2026 studies: Claude's web retrieval runs on Brave Search's independent index. Google rankings have limited transferability. Brave does not license from Google or Bing.

Sources

ClaudeBot Worst Crawl-to-Refer Ratio at 20,583 Pages Per Referral

verified intermediate

Presenc AI study of top 1,000 sites: ClaudeBot crawls 20,583 pages per referral. PerplexityBot best at ~210:1. 25% block GPTBot. OAI-SearchBot at 85:1.

Sources

Cloudflare "Block AI Bots" Silently Kills ChatGPT Visibility

verified intermediate

Cloudflare CDN-level bot blocking overrides robots.txt allow rules, an invisible cause of ChatGPT invisibility even when Bing-ranked.

Sources

Cloudflare Content Signals Policy โ€” Cite But Don't Train

verified intermediate

New robots.txt extension declares post-fetch usage (search/ai-input/ai-train). Set ai-train=no while ai-input=yes to stay citable but not trained on.

Sources

Cloudflare Three-Tier AI Crawler Classification โ€” Search, Agent, Training

verified intermediate

From Sep 15, 2026, Cloudflare blocks multi-purpose crawlers by default on ad-supported pages. Publishers can granularly allow/block each tier.

Sources

Crawl-to-Citation Efficiency Varies 10x Across AI Engines โ€” Perplexity Most Efficient

verified intermediate

Presenc AI (April 2026) joins crawl events to citation outcomes. PerplexityBot 5-10x more efficient than GPTBot. Optimize differently per engine โ€” Perplexity for fetchable content, Claude for

Sources

GEO attacks can promote flawed products into AI recommendations by up to 83.2%

verified advanced

Seller-controlled GEO rewrites promote flawed products into LLM recommendation sets by up to 83.2%; structured evidence checks cut the harm by up to 39.2% (SafeGEO, 600 cases).

Sources

Google AI Mode Information Agents launched as push-based referral surface

experimental advanced

Google launched always-on AI Mode information agents on June 12 2026 for Ultra subscribers. Unlike zero-click AIO, they push source-linked updates proactively to users.

Sources

Google AI Overviews Introduces Hover Pop-Up Link Cards (February 2026)

verified beginner

Google rolled out hover pop-up link cards in AI Overviews and AI Mode: hovering over highlighted text shows link cards for direct source access, creating a new path for user clicks from AI summaries.

Sources

Google AI Overviews now generates AI images directly in search results

verified beginner

Google launched AI image generation directly within AI Overviews (July 2026) using the Nano Banana model. Rolling out in English for AI Mode-supported countries.

Sources

Google I/O 2026: AI Mode becomes default, Search agents, Gemini 3.5, Personal Intelligence

verified intermediate

Google announced AI Mode as default replacement for traditional search, Gemini 3.5 Flash as default model, information agents, and Personal Intelligence across 98 languages. AI Mode crossed 1B MAU.

Sources

Google I/O 2026: AI Overviews and AI Mode Merged Into Unified Gemini 3.5 Flash Surface

verified intermediate

Google merged AI Overviews and AI Mode into one unified AI Search layer on Gemini 3.5 Flash, reaching 1B+ monthly users, eliminating separate citation strategies.

Sources

Google Lighthouse now audits llms.txt and agentic browsing readiness

experimental intermediate

Chrome Lighthouse 13.3 (May 2026) added an Agentic Browsing audit category testing llms.txt presence, WebMCP support, accessibility-tree quality and CLS โ€” an agent-era readiness signal.

Sources

Google States No Special AI Schema Required for AI Overviews or AI Mode

verified beginner

Google confirmed in 2026 that no AI-specific markup or llms.txt is required for AIO or AI Mode: 'GEO is still SEO at the core.'

Sources

Google S-CTS and S-BERT Systems Detect AI-Generated Content Networks Within Days of New Model Launch

verified advanced

Google's S-CTS terminates entire AI content networks (not individual pages), while S-BERT detects the mathematical 'fingerprint' of AI text. New model outputs detectable within days via LoRA updates.

Sources

Google SAGE research: AI agents pull from top 3 ranked pages โ€” fundamentals matter more

verified advanced

Google's SAGE paper (Jan 2026) reveals AI agents perform multi-step searches. Top-3 rankings are the gateway to agentic discovery. Four 'shortcut' patterns identified.

Sources

ChatGPT Business reads Bing for citations; Plus reads Google via Bright Data scraper

verified advanced

TUM thesis study of 370K+ results found ChatGPT Business citations are 95% Bing-sourced, Plus citations 94% Bright Data (Google scraper).

Sources

Grok DeepSearch Uses IndexNow: Fresh Content Can Appear in Citations Within 2-4 Weeks

verified intermediate

FuelOnline analysis: Grok DeepSearch draws on Bing's index. Submitting via IndexNow immediately after publishing is the fastest path to Grok citation visibility.

Sources

Grok xAI Crawler Silently Blocked by Default Security Rules Matching 'bot' Pattern

verified beginner

ProAISearch found Grok's crawler (user agents 'xAI' and 'Grok') lacks 'bot' in its name, so most default security setups and WAF rules silently block it. Content never gets indexed.

Sources

Grounding queries reveal that most AI influence is invisible (Otterly Copilot data)

experimental intermediate

Otterly's 3-month Copilot data: 647 grounding queries, 30,398 grounding events; 5 pages carried 74.6% of citations, 99.6% of AI influence invisible. Fix high-grounding/low-citation pages.

Sources

HasData: 56.4% of News Publishers Block AI Crawlers, 39.5% of Blocks Fail in Practice

verified beginner

HasData analysis of 10,894 domains (July 2026): 56.4% of news publishers block at least one AI crawler. GPTBot banned by 50.5% of publishers. 39.5% of GPTBot blocks fail to actually block.

Sources

Indirect Prompt Injection Is a Live Web Threat Against AI Crawlers and Agents

verified advanced

Google's April 2026 web sweep found attackers seeding prompt injections on sites to corrupt browsing AI; data-layer governance now required.

Sources

JavaScript-rendered content fails AI parsing 77% of the time

verified intermediate

Erlin's 2026 data: AI parse success rates โ€” static HTML with schema 94%, plain HTML 68%, JS-rendered 23%, PDF 7%. JS-rendered pricing/feature pages often read as empty templates.

Sources

June 2026 Spam Update โ€” SpamBrain Extends to AI Overviews/AI Mode Enforcement

verified intermediate

Google's June 24-26 2026 Spam Update extended SpamBrain enforcement to AI Overviews and AI Mode. Tactics to game AI answers are now treated as spam equivalent to paid links.

Sources

llms.txt Has 784+ Implementations But No Major Provider Confirms Using It โ€” MCP May Supersede

verified advanced

Presenc AI State of llms.txt 2026: 784+ implementations but only Perplexity and Anthropic (Claude Desktop) confirmed use. OpenAI, Google, Meta silent. MCP emerging as alternative.

Sources

llms.txt State of Adoption 2026: Support Confirmed, Sector Gaps Remain

verified intermediate

Presenc AI April 2026 report: Anthropic and Perplexity confirmed llms.txt support. Adoption <10% in finance/healthcare, high in developer SaaS. Top-100 brands lead.

Sources

llms.txt adoption reached 4-5% of mid-market sites โ€” moderate impact on citations, strong for agent discovery

experimental intermediate

llms.txt (proposed by Jeremy Howard Sep 2024) adopted by 4-5% of mid-market sites by mid-2026. Current evidence shows no direct citation boost but helps AI agents discover pages faster.

Sources

llms.txt and AI Discovery File Suite Becomes Competitive Differentiator

verified intermediate

One-van firm scored ChatGPT 99/100 against national brands using 9 AI Discovery Files. Backlinks didn't determine AI assessment โ€” structured facts did.

Sources

llms.txt has negligible short-term impact on AI search visibility โ€” OtterlyAI 90-day experiment

verified beginner

OtterlyAI's 90-day controlled experiment: out of 62.1K total AI bot hits, only 84 went to /llms.txt. Ahrefs: no major LLM provider currently supports llms.txt. Adoption is ~2% across 1M sites.

Sources

llms.txt implementation documented 5x AI traffic increase โ€” single highest-leverage technical GEO tactic

experimental intermediate

The Concurate case study documented a 5x increase in AI-referred traffic after implementing llms.txt. An llms.txt file at the domain root signals content priorities to LLM crawlers.

Sources

llms.txt shows zero measurable citation lift across 300K+ domain study

verified intermediate

SE Ranking analyzed 300K domains: llms.txt has no effect on AI citations. Removing the variable improved model accuracy. 7.4% adoption in top 10K sites. Niche value for developer docs only.

Sources

llms.txt: 800K+ Sites Published, 97% Get Zero AI Requests, Google Explicitly Ignores It

conflicting intermediate

Ahrefs studied 137,000 domains: ~28% publish llms.txt but 97% of those files got zero fetches in May 2026. Google confirms it ignores the file; coding agents fetch it ~5x more than AI search bots.

Sources

llms.txt adoption reached 4-5% of mid-market sites โ€” moderate impact on citations, strong for agent discovery

experimental intermediate

llms.txt (proposed by Jeremy Howard Sep 2024) adopted by 4-5% of mid-market sites by mid-2026. Current evidence shows no direct citation boost but helps AI agents discover pages faster.

Sources

llms.txt shows zero measurable citation lift across 300K+ domain study

verified intermediate

SE Ranking analyzed 300K domains: llms.txt has no effect on AI citations. Removing the variable improved model accuracy. 7.4% adoption in top 10K sites. Niche value for developer docs only.

Sources

Markdown Is Not an AI-SEO Shortcut โ€” Bots Find It Four Ways, Use Varies

experimental advanced

Botify's live-server experiment: bots discover Markdown via links, head tags, llms.txt, or content negotiation. GPTBot took 1,273 reads via links/head; purpose-built AI tools request Markdown ~100%

Sources

Most AI Crawler Traffic Is for Training, Not Indexing: 53.7% of AI Bots Fetch for Training

verified intermediate

Foglift classified 65,527 AI crawler requests: 53.7% training, 32% indexing, 14.3% user-triggered fetch. Blocking training bots won't stop citation crawls.

Sources

AI Crawlers Operate on ~2s Hard Timeout โ€” FCP Under 0.4s Correlates with 6.7 AI Citations

verified intermediate

Two converging findings: AI systems fetch pages with ~2s timeout (HTTP 499), and SE Ranking finds pages with FCP under 0.4s average 6.7 citations vs 2.1 for slower pages โ€” a 3.2x gap.

Sources

Perplexity Penalizes Gated Content: Gartner Receives 0 Citations Despite 130 Total Across Six Engines

verified intermediate

Machine Relations Index measurement shows Perplexity's retrieval architecture systematically deprioritizes paywalled sources โ€” affecting how B2B brands with gated content appear.

Sources

Perplexity post-trains models for cross-source evidence synthesis and accuracy (Aug 2026)

experimental intermediate

Perplexity's first technical explainer (Aug 6, 2026) details post-training models to connect evidence across sources and verify facts - shaping how content is synthesized and cited.

Sources

Perplexity Runs 6-Stage RAG Pipeline โ€” Only 5-10 Pages Retrieved Become 3-4 Cited

experimental advanced

Perplexity's pipeline (BM25 + dense retriever + XGBoost reranker, 0.7 threshold) retrieves 5-10 pages and cites 3-4. Crawler access is a binary gate: robots.txt/Cloudflare blocks remove pages

Sources

Perplexity Uses 3-Layer Reranking Pipeline With 780M+ Monthly Queries

verified intermediate

Perplexity processes 780M+ monthly queries via 3-layer reranking: relevance 30%, visual placement 20%, domain authority 15%, freshness 15%, source diversity 10%, schema 10%.

Sources

ChatGPT Has 'Labrador' VIP Lane for Licensed Publishers Bypassing Normal Retrieval

verified advanced

Researchers discovered ChatGPT's VIP tier via network traffic analysis. Licensed publishers (Reuters, WSJ, Wikipedia) get pre-summarized full-article extracts.

Sources

Rail Europe doubled ChatGPT traffic by prioritizing AI crawler access and value-led SEO

verified intermediate

Botify case study: Rail Europe shifted from volume-led to value-led SEO, reorganized sitemaps by page type, and prioritized AI crawler governance โ€” resulting in 2x ChatGPT referral traffic.

Sources

Search and Training AI Crawlers Are Separate โ€” Allow Citation Bots, Block Training Only If Deliberate

verified beginner

OAI-SearchBot/PerplexityBot/ClaudeBot power citations; GPTBot/Google-Extended affect training only. Blocking the wrong one invisibilizes you.

Sources

Short URLs do NOT get cited more โ€” /guide/ pages average 42% above baseline, 'blog/' next

experimental beginner

Otterly AI analyzed 1,028,959 URLs cited across 6 AI platforms. URL length correlates near-zero with citations (r = -0.025). /guide/ pages average 2.7 citations (42% above 1.9 avg). Clean URLs beat qu

Sources

Split robots.txt Strategy: Allow Retrieval Bots, Block Training Bots

verified intermediate

89% of AI crawler traffic is training/mixed, only 8% search-related. Allow OAI-SearchBot, PerplexityBot, Claude-Web. Block GPTBot, CCBot, Bytespider.

Sources

AI Crawlers Mostly Do Not Execute JavaScript โ€” Server-Render Citation Content

verified intermediate

Vercel analysis of 1.3B fetches found near-zero JS execution for GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot; only Google-Extended renders JS.

Sources

Stripe uses llms.txt to steer AI agents away from deprecated APIs โ€” agent behavior shaping

verified advanced

Stripe's llms.txt includes 'Instructions for Large Language Model Agents' section. Explicitly steers coding assistants away from legacy Card Element. Real value of llms.txt is agentic-web control

Sources

Stripe uses llms.txt to steer AI agents away from deprecated APIs โ€” agent behavior shaping

verified advanced

Stripe's llms.txt includes 'Instructions for Large Language Model Agents' section. Explicitly steers coding assistants away from legacy Card Element. Real value of llms.txt is agentic-web control, not

Sources

Structured data + llms.txt deliver cheapest AI visibility ROI โ€” 200% monthly traffic growth in auto parts case study

verified beginner

Hedges Company case study: Product schema markup + llms.txt drove 200% monthly AI referral traffic growth for 3 months. ChatGPT/Perplexity citation rates tripled month-over-month.

Sources

Structured data + llms.txt deliver cheapest AI visibility ROI โ€” 200% monthly traffic growth in auto parts case study

verified beginner

Hedges Company case study: Product schema markup + llms.txt drove 200% monthly AI referral traffic growth for 3 months. ChatGPT/Perplexity citation rates tripled month-over-month.

Sources

Frequently Asked Questions

7,029 Sites Embed Hidden Prompts That Prime AI to Cite Them

Trakkr scan (833,791 domains): 7,029 sites embed 'remember/cite us' instructions in AI buttons. 37% use memory-anchoring language; 98% target ChatGPT. Microsoft calls it 'AI Recommendation Poisoning'.

73% of Websites Have Technical Barriers Blocking AI Crawlers

OtterlyAI study of 1M+ citations found 73% of sites block AI crawlers via robots.txt, CDN rules, or JS rendering issues. Fix crawler access before anything else.

74% of Websites Are Invisible to AI Search โ€” Structured Data Is the Fixable Gap

SearchScore SAVI (850k sites, Q2 2026): 74.2% Invisible/Low visibility; AI visibility averages 34.1 vs technical health 70.1, with structured data scoring 23.1.

Agentic AI Search: Claude and GPT-5 Agents Now Execute Multi-Step Purchasing (Jun 2026)

Autonomous AI agents from Claude and GPT-5 now execute multi-step purchasing and comparison tasks. Agent-first content structure becomes critical for visibility.

Agentic search requires machine-readable pricing, availability, and checkout data โ€” protocols reshaping SEO

OpenAI Agentic Commerce Protocol (ACP, with Stripe), Web Capabilities Protocol (WebMCP, Google+Microsoft), Universal Commerce Protocol (UCP, Google+Shopify) โ€” sites without machine-readable

AI Agent Traffic Reaches 88% of Human Organic Search Volume; Predicted to Surpass by End 2026

BrightEdge data shows AI agent requests now rival human search traffic, with OpenAI agents dominating. Differentiated robots.txt strategy is essential.

AI Crawler Explosion โ€” GPTBot +305% YoY, AI Bots Reach 22% of All Bot Traffic

Presenc AI 2022-2026 trend: GPTBot grew 305% YoY. AI bots now 22% of all bot traffic. Meta-ExternalAgent surged to #2 at 16.7%. Bots = 31.2% of all HTTP requests; trajectory crosses human traffic

The #1 GEO robots.txt Mistake: Blocking OAI-SearchBot Instead of GPTBot Kills ChatGPT Citations

OpenAI runs 3 bots: GPTBot (training), OAI-SearchBot (search index), ChatGPT-User (fetches). Block OAI-SearchBot = zero ChatGPT citations. Most teams conflate them.

AI Crawler Traffic Now 3.4% of All Web Traffic; 12% of Top Sites Block OAI-SearchBot

Known Agents data shows AI search crawlers are a measurable and growing portion of web traffic. Differentiated robots.txt management is a basic GEO hygiene requirement.

AI Crawlers Now Account for 40-50% of Bot-Level Activity โ€” 65-70% Are Live Queries

JetOctopus server log analysis (Feb 2026): AI bots ~40-50% of Googlebot-level activity. 65-70% of AI bot traffic is user search, not training. 14+ AI user-agents need explicit Allow rules in

AI Crawlers Visit Once and Skip Homepages โ€” Blog Is the New Front Door

Trakkr (575,788 AI crawler visits): GPTBot reaches homepages only ~3% of the time; 88.5% of pages get exactly one visit; 21% of ChatGPT Search sessions start on blog pages.

AI Crawls Product Pages but Cites Blog Posts โ€” 337K Citation Mismatch

Trakkr cross-referenced 337K AI citations with 11.4M crawler visits: what AI crawls (product pages) is not what it cites (blog content). Measure both, not one.

AI Labs Don't Use llms.txt on Their Own Consumer Front Doors โ€” Docs Only

HTTP Archive: chatgpt.com, claude.ai, gemini.google.com lack llms.txt. But docs.anthropic.com, docs.perplexity.ai have one. Value is for agent-readiness, not marketing.

AI Mode + AI Overviews merged at Google I/O 2026 โ€” unified AI Search on Gemini 3.5 Flash

Google merged AI Overviews and AI Mode into one unified AI Search experience at I/O 2026. Runs on Gemini 3.5 Flash. Single citation pool eliminates need for separate strategies.

AI Mode + AI Overviews merged at Google I/O 2026 โ€” unified AI Search on Gemini 3.5 Flash

Google merged AI Overviews and AI Mode into one unified AI Search experience at I/O 2026. Runs on Gemini 3.5 Flash. Single citation pool eliminates need for separate strategies.

Anthropic Runs 3 Separate Crawler User-Agents With Distinct Functions

ClaudeBot (training), Claude-Web (real-time browsing), and anthropic-ai (catch-all) can be independently controlled in robots.txt for granular AI visibility management.

Bing Webmaster Tools + OAI-SearchBot Is the ChatGPT Citation Pipeline

ChatGPT Search retrieves from Bing's index: 87% of ChatGPT cited pages correspond to Bing top results (Mersel AI 2026). Submit sitemap to BWT, allow OAI-SearchBot, enable IndexNow.

ChatGPT Search cites only 15% of retrieved pages โ€” four-gate funnel model

AirOps study of 548K pages: ChatGPT Search retrieves via Bing index, checks OAI-SearchBot crawl access, then cites only 15% based on structure, freshness, and authority.

ChatGPT Sends 28.8% of Referrals to Internal Site Search Instead of the Answer Page (Previsible Jul 2026)

Previsible study: ChatGPT sends 28.8% of its referrals to internal search results pages. The model trusts the domain but defaults to site search when it can't identify the right page.

ChatGPT Shopping: Rank-1 Offer Drives the Card; GPT Tags Lift Ranking 144%

Profound (201K prompts, 812K product cards): card info is pulled from the rank-1 offer; GPT-tagged products get +144% top-rank lift; median review count +123%; Reddit = ~1/3 of shopping citations.

ChatGPT Citation Pipeline Is Two-Stage โ€” Bing Retrieval Then Fine-Tuned Re-Rank

ChatGPT uses Bing to retrieve candidates, then a fine-tuned model re-ranks by answer fit, domain authority, source consensus; zero JS execution.

Claude uses Brave Search (not Google/Bing) โ€” 86.7% overlap with Brave top 10, Brave SEO is prerequisite

Profound 2025 + multiple 2026 studies: Claude's web retrieval runs on Brave Search's independent index. Google rankings have limited transferability. Brave does not license from Google or Bing.

ClaudeBot Worst Crawl-to-Refer Ratio at 20,583 Pages Per Referral

Presenc AI study of top 1,000 sites: ClaudeBot crawls 20,583 pages per referral. PerplexityBot best at ~210:1. 25% block GPTBot. OAI-SearchBot at 85:1.

Cloudflare "Block AI Bots" Silently Kills ChatGPT Visibility

Cloudflare CDN-level bot blocking overrides robots.txt allow rules, an invisible cause of ChatGPT invisibility even when Bing-ranked.

Cloudflare Content Signals Policy โ€” Cite But Don't Train

New robots.txt extension declares post-fetch usage (search/ai-input/ai-train). Set ai-train=no while ai-input=yes to stay citable but not trained on.

Cloudflare Three-Tier AI Crawler Classification โ€” Search, Agent, Training

From Sep 15, 2026, Cloudflare blocks multi-purpose crawlers by default on ad-supported pages. Publishers can granularly allow/block each tier.

Crawl-to-Citation Efficiency Varies 10x Across AI Engines โ€” Perplexity Most Efficient

Presenc AI (April 2026) joins crawl events to citation outcomes. PerplexityBot 5-10x more efficient than GPTBot. Optimize differently per engine โ€” Perplexity for fetchable content, Claude for

GEO attacks can promote flawed products into AI recommendations by up to 83.2%

Seller-controlled GEO rewrites promote flawed products into LLM recommendation sets by up to 83.2%; structured evidence checks cut the harm by up to 39.2% (SafeGEO, 600 cases).

Google AI Mode Information Agents launched as push-based referral surface

Google launched always-on AI Mode information agents on June 12 2026 for Ultra subscribers. Unlike zero-click AIO, they push source-linked updates proactively to users.

Google AI Overviews Introduces Hover Pop-Up Link Cards (February 2026)

Google rolled out hover pop-up link cards in AI Overviews and AI Mode: hovering over highlighted text shows link cards for direct source access, creating a new path for user clicks from AI summaries.

Google AI Overviews now generates AI images directly in search results

Google launched AI image generation directly within AI Overviews (July 2026) using the Nano Banana model. Rolling out in English for AI Mode-supported countries.

Google I/O 2026: AI Mode becomes default, Search agents, Gemini 3.5, Personal Intelligence

Google announced AI Mode as default replacement for traditional search, Gemini 3.5 Flash as default model, information agents, and Personal Intelligence across 98 languages. AI Mode crossed 1B MAU.

Google I/O 2026: AI Overviews and AI Mode Merged Into Unified Gemini 3.5 Flash Surface

Google merged AI Overviews and AI Mode into one unified AI Search layer on Gemini 3.5 Flash, reaching 1B+ monthly users, eliminating separate citation strategies.

Google Lighthouse now audits llms.txt and agentic browsing readiness

Chrome Lighthouse 13.3 (May 2026) added an Agentic Browsing audit category testing llms.txt presence, WebMCP support, accessibility-tree quality and CLS โ€” an agent-era readiness signal.

Google States No Special AI Schema Required for AI Overviews or AI Mode

Google confirmed in 2026 that no AI-specific markup or llms.txt is required for AIO or AI Mode: 'GEO is still SEO at the core.'

Google S-CTS and S-BERT Systems Detect AI-Generated Content Networks Within Days of New Model Launch

Google's S-CTS terminates entire AI content networks (not individual pages), while S-BERT detects the mathematical 'fingerprint' of AI text. New model outputs detectable within days via LoRA updates.

Google SAGE research: AI agents pull from top 3 ranked pages โ€” fundamentals matter more

Google's SAGE paper (Jan 2026) reveals AI agents perform multi-step searches. Top-3 rankings are the gateway to agentic discovery. Four 'shortcut' patterns identified.

ChatGPT Business reads Bing for citations; Plus reads Google via Bright Data scraper

TUM thesis study of 370K+ results found ChatGPT Business citations are 95% Bing-sourced, Plus citations 94% Bright Data (Google scraper).

Grok DeepSearch Uses IndexNow: Fresh Content Can Appear in Citations Within 2-4 Weeks

FuelOnline analysis: Grok DeepSearch draws on Bing's index. Submitting via IndexNow immediately after publishing is the fastest path to Grok citation visibility.

Grok xAI Crawler Silently Blocked by Default Security Rules Matching 'bot' Pattern

ProAISearch found Grok's crawler (user agents 'xAI' and 'Grok') lacks 'bot' in its name, so most default security setups and WAF rules silently block it. Content never gets indexed.

Grounding queries reveal that most AI influence is invisible (Otterly Copilot data)

Otterly's 3-month Copilot data: 647 grounding queries, 30,398 grounding events; 5 pages carried 74.6% of citations, 99.6% of AI influence invisible. Fix high-grounding/low-citation pages.

HasData: 56.4% of News Publishers Block AI Crawlers, 39.5% of Blocks Fail in Practice

HasData analysis of 10,894 domains (July 2026): 56.4% of news publishers block at least one AI crawler. GPTBot banned by 50.5% of publishers. 39.5% of GPTBot blocks fail to actually block.

Indirect Prompt Injection Is a Live Web Threat Against AI Crawlers and Agents

Google's April 2026 web sweep found attackers seeding prompt injections on sites to corrupt browsing AI; data-layer governance now required.

JavaScript-rendered content fails AI parsing 77% of the time

Erlin's 2026 data: AI parse success rates โ€” static HTML with schema 94%, plain HTML 68%, JS-rendered 23%, PDF 7%. JS-rendered pricing/feature pages often read as empty templates.

June 2026 Spam Update โ€” SpamBrain Extends to AI Overviews/AI Mode Enforcement

Google's June 24-26 2026 Spam Update extended SpamBrain enforcement to AI Overviews and AI Mode. Tactics to game AI answers are now treated as spam equivalent to paid links.

llms.txt Has 784+ Implementations But No Major Provider Confirms Using It โ€” MCP May Supersede

Presenc AI State of llms.txt 2026: 784+ implementations but only Perplexity and Anthropic (Claude Desktop) confirmed use. OpenAI, Google, Meta silent. MCP emerging as alternative.

llms.txt State of Adoption 2026: Support Confirmed, Sector Gaps Remain

Presenc AI April 2026 report: Anthropic and Perplexity confirmed llms.txt support. Adoption <10% in finance/healthcare, high in developer SaaS. Top-100 brands lead.

llms.txt adoption reached 4-5% of mid-market sites โ€” moderate impact on citations, strong for agent discovery

llms.txt (proposed by Jeremy Howard Sep 2024) adopted by 4-5% of mid-market sites by mid-2026. Current evidence shows no direct citation boost but helps AI agents discover pages faster.

llms.txt and AI Discovery File Suite Becomes Competitive Differentiator

One-van firm scored ChatGPT 99/100 against national brands using 9 AI Discovery Files. Backlinks didn't determine AI assessment โ€” structured facts did.

llms.txt has negligible short-term impact on AI search visibility โ€” OtterlyAI 90-day experiment

OtterlyAI's 90-day controlled experiment: out of 62.1K total AI bot hits, only 84 went to /llms.txt. Ahrefs: no major LLM provider currently supports llms.txt. Adoption is ~2% across 1M sites.

llms.txt implementation documented 5x AI traffic increase โ€” single highest-leverage technical GEO tactic

The Concurate case study documented a 5x increase in AI-referred traffic after implementing llms.txt. An llms.txt file at the domain root signals content priorities to LLM crawlers.

llms.txt shows zero measurable citation lift across 300K+ domain study

SE Ranking analyzed 300K domains: llms.txt has no effect on AI citations. Removing the variable improved model accuracy. 7.4% adoption in top 10K sites. Niche value for developer docs only.

llms.txt: 800K+ Sites Published, 97% Get Zero AI Requests, Google Explicitly Ignores It

Ahrefs studied 137,000 domains: ~28% publish llms.txt but 97% of those files got zero fetches in May 2026. Google confirms it ignores the file; coding agents fetch it ~5x more than AI search bots.

llms.txt adoption reached 4-5% of mid-market sites โ€” moderate impact on citations, strong for agent discovery

llms.txt (proposed by Jeremy Howard Sep 2024) adopted by 4-5% of mid-market sites by mid-2026. Current evidence shows no direct citation boost but helps AI agents discover pages faster.

llms.txt shows zero measurable citation lift across 300K+ domain study

SE Ranking analyzed 300K domains: llms.txt has no effect on AI citations. Removing the variable improved model accuracy. 7.4% adoption in top 10K sites. Niche value for developer docs only.

Markdown Is Not an AI-SEO Shortcut โ€” Bots Find It Four Ways, Use Varies

Botify's live-server experiment: bots discover Markdown via links, head tags, llms.txt, or content negotiation. GPTBot took 1,273 reads via links/head; purpose-built AI tools request Markdown ~100%

Most AI Crawler Traffic Is for Training, Not Indexing: 53.7% of AI Bots Fetch for Training

Foglift classified 65,527 AI crawler requests: 53.7% training, 32% indexing, 14.3% user-triggered fetch. Blocking training bots won't stop citation crawls.

AI Crawlers Operate on ~2s Hard Timeout โ€” FCP Under 0.4s Correlates with 6.7 AI Citations

Two converging findings: AI systems fetch pages with ~2s timeout (HTTP 499), and SE Ranking finds pages with FCP under 0.4s average 6.7 citations vs 2.1 for slower pages โ€” a 3.2x gap.

Perplexity Penalizes Gated Content: Gartner Receives 0 Citations Despite 130 Total Across Six Engines

Machine Relations Index measurement shows Perplexity's retrieval architecture systematically deprioritizes paywalled sources โ€” affecting how B2B brands with gated content appear.

Perplexity post-trains models for cross-source evidence synthesis and accuracy (Aug 2026)

Perplexity's first technical explainer (Aug 6, 2026) details post-training models to connect evidence across sources and verify facts - shaping how content is synthesized and cited.

Perplexity Runs 6-Stage RAG Pipeline โ€” Only 5-10 Pages Retrieved Become 3-4 Cited

Perplexity's pipeline (BM25 + dense retriever + XGBoost reranker, 0.7 threshold) retrieves 5-10 pages and cites 3-4. Crawler access is a binary gate: robots.txt/Cloudflare blocks remove pages

Perplexity Uses 3-Layer Reranking Pipeline With 780M+ Monthly Queries

Perplexity processes 780M+ monthly queries via 3-layer reranking: relevance 30%, visual placement 20%, domain authority 15%, freshness 15%, source diversity 10%, schema 10%.

ChatGPT Has 'Labrador' VIP Lane for Licensed Publishers Bypassing Normal Retrieval

Researchers discovered ChatGPT's VIP tier via network traffic analysis. Licensed publishers (Reuters, WSJ, Wikipedia) get pre-summarized full-article extracts.

Rail Europe doubled ChatGPT traffic by prioritizing AI crawler access and value-led SEO

Botify case study: Rail Europe shifted from volume-led to value-led SEO, reorganized sitemaps by page type, and prioritized AI crawler governance โ€” resulting in 2x ChatGPT referral traffic.

Search and Training AI Crawlers Are Separate โ€” Allow Citation Bots, Block Training Only If Deliberate

OAI-SearchBot/PerplexityBot/ClaudeBot power citations; GPTBot/Google-Extended affect training only. Blocking the wrong one invisibilizes you.

Short URLs do NOT get cited more โ€” /guide/ pages average 42% above baseline, 'blog/' next

Otterly AI analyzed 1,028,959 URLs cited across 6 AI platforms. URL length correlates near-zero with citations (r = -0.025). /guide/ pages average 2.7 citations (42% above 1.9 avg). Clean URLs beat qu

Split robots.txt Strategy: Allow Retrieval Bots, Block Training Bots

89% of AI crawler traffic is training/mixed, only 8% search-related. Allow OAI-SearchBot, PerplexityBot, Claude-Web. Block GPTBot, CCBot, Bytespider.

AI Crawlers Mostly Do Not Execute JavaScript โ€” Server-Render Citation Content

Vercel analysis of 1.3B fetches found near-zero JS execution for GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot; only Google-Extended renders JS.

Stripe uses llms.txt to steer AI agents away from deprecated APIs โ€” agent behavior shaping

Stripe's llms.txt includes 'Instructions for Large Language Model Agents' section. Explicitly steers coding assistants away from legacy Card Element. Real value of llms.txt is agentic-web control

Stripe uses llms.txt to steer AI agents away from deprecated APIs โ€” agent behavior shaping

Stripe's llms.txt includes 'Instructions for Large Language Model Agents' section. Explicitly steers coding assistants away from legacy Card Element. Real value of llms.txt is agentic-web control, not

Structured data + llms.txt deliver cheapest AI visibility ROI โ€” 200% monthly traffic growth in auto parts case study

Hedges Company case study: Product schema markup + llms.txt drove 200% monthly AI referral traffic growth for 3 months. ChatGPT/Perplexity citation rates tripled month-over-month.

Structured data + llms.txt deliver cheapest AI visibility ROI โ€” 200% monthly traffic growth in auto parts case study

Hedges Company case study: Product schema markup + llms.txt drove 200% monthly AI referral traffic growth for 3 months. ChatGPT/Perplexity citation rates tripled month-over-month.