OpenAI Says Robots.txt May Not Apply To ChatGPT’s Fetch Bot via @sejournal, @MattGSouthern
All Paid Media PPC News Social MediaAdvertising Video Advertising Columns Ask A PPC ExpertNEW PPC Pulse Webinar Google Local Services Ads Are Moving To PMax: What To Check First Prepare for Google's LSA move into Performance Max with a before-and-after audit checklist from CallRail. Digital Marketing All Digital Marketing Analytics & Data Ecommerce Lead Generation Content Marketing Social Platforms Google YouTube Reddit LinkedIn TikTokNEW WordPress Other CMS Platforms Webinar Google Local Services Ads Are Moving To PMax: What To Check First Prepare for Google's LSA move into Performance Max with a before-and-after audit checklist from CallRail. SEJ Live Webinars Rundowns PodcastNEW Ebooks All Resources SEJ Pro SEJ Pro: compare AI search notes in private Ask the SEJ experts you already read what AI, agents, and algorithm changes mean for your traffic. Private, human to human. $97/mo. New data shows ChatGPT's page-fetching bot is reaching sites that have disallowed it, and OpenAI's own documentation explains why the file may not stop it. SEJ STAFF Matt G. Southern 3 hours ago ⋅ 3 min read SEJ STAFF Matt G. Southern Senior News Writer at Search Engine Journal Bio Follow ChatGPT’s page-fetching bot is disallowed by more sites than any other AI bot of its kind. It also reached disallowed pages on more sites than any other bot. OpenAI says robots.txt rules may not apply to it because a person asked for the page. TollBit’s latest State of the Bots report has the numbers for the first half of 2026. Here’s what else the data shows about how these crawlers behave and what it means for your site. In the European sites discussed in the report, about 15% of identified AI page-fetchers reached URLs that the sites had marked as disallowed. This happens mostly with a few specific agents. For example, ChatGPT-User, Bytespider, and Youbot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them. Among these, ChatGPT-User reached the most sites. Many of the newer page-fetching agents are hardly blocked at all. Only 9% of European websites disallow Claude-User, compared to 26% in North America. Perplexity-User sits at 13% versus 26%. Most of the newest agents have disallow rates in the single digits across Europe, but ChatGPT-User stands out as an exception. OpenAI’s crawler documentation says ChatGPT-User visits a page when a ChatGPT user asks a question, and that because those actions are initiated by a user, robots.txt rules may not apply. Perplexity says Perplexity-User generally ignores the file for the same reason, but Anthropic has a different view and states that all three of its bots respect it, as we reported in February. TollBit treats any request to a disallowed URL as a bypass, regardless of what the operator claims. A disallow line for ChatGPT-User is a request that OpenAI’s documentation says may not apply. It’s important to look at a different aspect here. According to OpenAI’s documentation, the agent responsible for deciding if a site shows up in ChatGPT search results is called OAI-SearchBot, not ChatGPT-User. Sites that block both agents to prevent AI traffic have traded away the visibility half of that deal and kept a fetching control that carries a carve-out. Server logs or CDN records show what actually arrived. The file only shows what you asked for. Cloudflare is making some updates to how it manages its crawler controls, moving the decision to the network layer. When it comes to the bots it recognizes, compliance is no longer left up to the crawler itself. Starting from September 15, new domains added to Cloudflare will have their Training and Agent crawlers blocked by default on pages with ads, while Search crawlers remain allowed. Whether the user-initiated loophole survives is the open question. It rests on the argument that requesting a page differs from a crawler taking it, and now all major assistants fetch pages this way. Read Full Bio SEJ STAFF Matt G. Southern Senior News Writer at Search Engine Journal See short video versions of news stories on YouTube and TikTok. Matt G. Southern is the Senior News Writer at ... Learn how to connect search, AI, and PPC into one unstoppable strategy. The Modern Guide To Robots.txt: How To Use It Avoiding The Pitfalls ChatGPT Search May Have A Shot At Google 8 Common Robots.txt Issues And How To Fix Them Join 75,000+ Digital Leaders. Learn how to connect search, AI, and PPC into one unstoppable strategy. Learn how to connect search, AI, and PPC into one unstoppable strategy. In a world ruled by algorithms, SEJ brings timely, relevant information for SEOs, marketers, and entrepreneurs to optimize and grow their businesses -- and careers.
Source: Search Engine Journal
This article has been carefully curated and reformatted for educational and informational purposes. Full credit goes to the original publisher.
📚 Visit more helpful articles on Joab Peters Blog
No comments