AI crawler access: what to allow and how to verify it

AI products fetch your pages with their own crawlers, and you control which ones may do so in robots.txt. Allow the crawlers you want, block the ones you do not, and then check that a firewall is not blocking the ones you allowed.

Synopsis
Where you control it
robots.txt at the root of your domain, one group of rules per user agent.
OpenAI
GPTBot (training data), OAI-SearchBot (search) and ChatGPT-User (fetches a page for a user).
Anthropic
ClaudeBot.
Perplexity
PerplexityBot.
Google
Google-Extended is a robots.txt token that controls use of your content for Google's AI models. It is not a separate crawler.
Apple
Applebot-Extended is the equivalent token for Apple's AI features.
Check the vendors
Names and behavior change. Confirm them in each vendor's documentation before you rely on them.

Which AI crawlers should you allow?

It depends on what you want. Blocking training crawlers and allowing search crawlers is a common middle path, but the decision is yours.

Allow if

  • You want to be found and cited in AI search
  • Your content is public and meant to be read

Block if

  • Part of your site is private or licensed
  • You do not want your content used to train models

Remember

A blocked crawler cannot read your page, so it cannot cite it from that crawl.

What does a robots.txt for AI crawlers look like?

One group per user agent. An empty Disallow, or Allow: /, permits everything. Disallow with a path blocks that path.

# Allow the crawlers you want, block the ones you do not
User-agent: GPTBot
Allow: /

User-agent: OAI-SearchBot
Allow: /

User-agent: ClaudeBot
Allow: /

User-agent: PerplexityBot
Allow: /

User-agent: Google-Extended
Disallow: /private/

Sitemap: https://example.com/sitemap.xml

How do you verify that crawlers can reach your pages?

Rules are only half of it. Firewalls, bot protection and rate limits can block a crawler that your robots.txt allows.

Steps

  1. Open /robots.txt on your domain and read the groups.
  2. Fetch a page with a crawler user agent and check that the status is 200.
  3. Look at your firewall or CDN settings for rules that block bots or unknown user agents.
  4. Check your server logs for requests from the crawlers you allowed.
  5. Run a scan to see the crawler access checks for your site.

Related terms: robots.txt, AI crawler and user agent.

Questions.

Straight answers about AI crawler access.

Will blocking AI crawlers hurt my search ranking?

The AI crawler rules do not change how a traditional search engine ranks you, because those use their own crawlers. They do decide whether AI products can read your pages.

Do all AI crawlers obey robots.txt?

Well-behaved ones do, and the major vendors document that they do. Check each vendor's documentation, and use your firewall for anything you must enforce.

How often should I check AI crawler access?

After any change to robots.txt, your firewall, or your CDN, and whenever crawler names change.

Check crawler access on your site

Run a scan and see whether AI crawlers can reach your pages.