BrandGEO
Tutorials · · 5 min read

gptbot robots.txt: AI Crawler Rules That Work

A practical guide to blocking or allowing GPTBot, ClaudeBot, Google-Extended and other AI crawlers.

AI crawlers are not all doing the same job. Here’s how to separate training bots from retrieval bots, decide what to block, and copy the right robots.txt rules.

There is a good chance your robots.txt file is quietly deciding whether buyers find you in AI answers, and that nobody on your team has looked at it since a panicked afternoon two years ago when someone decided to "block the AI bots." That one file now shapes whether your content trains models, whether AI answer engines can retrieve your pages, and whether your brand shows up accurately when a buyer asks ChatGPT, Claude, Gemini, or Perplexity about your category.

The reason this is dangerous is that it feels settled. You decided once, the file has not changed, and you assume it still reflects what you want. Meanwhile the landscape it governs has shifted completely, and the most common configurations are now working against the teams that set them.

The mistake that started it all

Almost every problem here traces back to one flawed assumption: that "AI crawlers" are a single category you either allow or block. They are not, and treating them as one is how well-meaning teams accidentally make themselves invisible.

There are really two kinds of bots reaching your site, and they do opposite things for your business. One kind collects content that may be used to train or improve AI models. When you block those, you are trying to keep your work out of future training datasets. The other kind fetches or indexes pages so an AI product can answer a live user query, cite a source, or recommend a brand. Those are directly connected to whether you appear in AI answers at all.

A blanket "block all AI" policy treats both the same. It might succeed at keeping your content out of training data, which may or may not matter to you, while also cutting off the exact crawlers that could put you in front of a buyer asking for recommendations. You win an argument you may not care about and lose the visibility you almost certainly do.

Why this is a strategy decision, not a technical one

The instinct is to hand robots.txt to whoever manages the site and consider it done. But the real question is not technical. It is a trade-off between protecting content and being discoverable, and that trade-off depends on your business model, not on a best-practice snippet.

A publisher or research firm sitting on proprietary reports, paid editorial, or licensed data has real reasons to restrict AI crawling broadly. The concern is not just that a blog post gets copied. It is losing control over how expertise and data are absorbed into systems that may never cite the source. For some businesses that outweighs the visibility.

If your site exists to generate leads, sell software, or win B2B deals, the calculation inverts. Buyers increasingly ask AI tools for the best alternatives to a competitor, the top providers for a service, or whether a product fits their use case. If retrieval systems cannot reach your pages, they answer from older indexes, third-party summaries, review sites, or competitor pages instead. You do not go silent. Someone else narrates your category for you.

Why "block everything" quietly backfires

Here is the trap that catches marketing-led companies. Blocking training crawlers feels protective and costs little. Blocking retrieval crawlers feels like the same gesture but does something entirely different: it removes your official pages from the very moment a buyer is deciding. And because the two kinds of bots are easy to lump together, teams routinely block both while believing they only did the first.

Two subtler realities make this worse. Some crawler tokens are not ordinary crawlers at all. They are preference signals used to opt out of AI training while allowing the core search crawler to keep working, and confusing the two can either fail to protect what you meant to protect or damage search visibility you meant to keep. And some AI products lean on standard search indexes rather than crawling under an obvious AI user agent, so blocking the visible "AI bots" does not fully determine whether you appear in AI answers anyway.

There is also a self-defeating move worth naming: blocking retrieval bots while complaining that AI answers about your brand are inaccurate. If the engines cannot reach your canonical pages, they fall back to whatever outdated or third-party version of you exists elsewhere. You cannot both hide your facts and expect the facts to be right.

What a deliberate policy actually balances

For most B2B, SaaS, ecommerce, and professional services brands, the sane middle ground is not "all AI" or "no AI." It is a set of judgment calls: let standard search crawlers work, let AI retrieval and search crawlers reach the public pages that define your brand, restrict training crawlers from content you genuinely do not want used for model improvement, and keep private and transactional areas behind authentication rather than trusting robots.txt to hide them. Robots.txt is public and advisory, so anything truly confidential belongs behind a login.

The question is never "AI or no AI." It is which content, for which use, by which crawler, and that answer differs section by section across your site. Product and comparison pages usually want maximum retrieval access. Paid reports and account areas want none. Map crawler access to business value, not one rule to everything.

The step almost everyone skips

Whatever policy you choose, there is a measurement step that makes or breaks it, and it is the one teams leave out. Before you change anything, you should know how AI engines currently describe and recommend your brand. After you change it, you should check whether those answers got more accurate and better sourced. Otherwise you are adjusting a file that governs your visibility with no read on whether the adjustment helped or hurt.

This is where monitoring belongs in the workflow. BrandGEO audits how AI engines describe your brand across trained-data and live web-search modes, showing which pages and facts need to be clearer before you touch robots.txt, and whether your changes moved the answers afterward.

Crawler names and provider policies keep changing, so this is not a set-and-forget file. Revisit it whenever you launch a major content section, new pricing, a docs site, or a community, and treat every change as a deliberate trade-off between protecting content and controlling how your brand shows up in the answers buyers trust.

See how ChatGPT, Claude, Gemini, Grok, and DeepSeek currently describe and recommend your brand. BrandGEO audits your AI visibility across trained-data and live web-search modes, then turns the findings into a practical GEO action plan.

See how AI describes your brand

BrandGEO runs structured prompts across ChatGPT, Claude, Gemini, Grok, and DeepSeek — and scores your brand across six dimensions. Two minutes, no credit card.

Keep reading

Related posts

BrandGEO
Tutorials Sep 24, 2026

llms.txt Explained: Should Your Site Have One?

llms.txt is a proposed way to help AI systems understand your site’s most useful content. Here’s what to publish, what not to expect, and a template you can use.