The Costly robots.txt Mistake Killing AI Visibility
Over the past year, many site owners panicked about AI training and added blanket disallow rules to their robots.txt:
User-agent: *
Disallow: /
While intended to protect intellectual property, this created a devastating commercial side effect: The website disappeared from ChatGPT Search, Perplexity citations, and Claude recommendations entirely.
The critical distinction is understanding the difference between Training Scrapers and Live Search/Citation Bots.
AI Crawler Breakdown: Who Does What?
| Bot User-Agent | Operator | Purpose | Impact if Blocked |
|---|---|---|---|
GPTBot |
OpenAI | Web Crawling & Search | Your brand is excluded from ChatGPT Search answers and citations. |
OAI-SearchBot |
OpenAI | Live Search Engine | ChatGPT Search cannot link directly to your pages. |
PerplexityBot |
Perplexity | Real-Time Citation | Perplexity will never cite your domain as a primary source. |
ClaudeBot |
Anthropic | Web Crawling | Claude will not reference or verify your product specifications. |
Google-Extended |
Gemini Training Data | Blocks Gemini model training without hurting Google Search indexing. | |
CCBot |
Common Crawl | General Web Archive | Excludes your content from foundational model pre-training. |
The Recommended 2026 Robots.txt Configuration
To capture maximum commercial referral traffic from AI search engines while managing general training scrapers:
# Allow commercial AI search engines (Drives Real Traffic & Customers)
User-agent: GPTBot
Allow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: ClaudeBot
Allow: /
# Point AI crawlers to your sitemap and llms.txt
Sitemap: https://yoursite.com/sitemap.xml
Verify your live bot configuration right now using GeoVisible’s free AI Robots.txt Tester.