- →GPTBot governs training access only. ChatGPT Search visibility runs through OAI-SearchBot, and blocking one does not block the other.
- →Cloudflare measured GPTBot requests up 305% between May 2024 and May 2025, with its share of verified crawler traffic going from 4.7% in July 2024 to 11.7% in July 2025. Request volume, not page coverage.
- →Cloudflare's Year in Review analysis, October to November 2025: Googlebot touched 11.6% of unique pages against 3.6% for GPTBot. High growth from a low base is still a low base.
- →Chartbeat data reported by Axios on 17 March 2026 put ChatGPT referrals up more than 200% year over year and still under 1% of publisher page views. Optimise for citation and brand presence, not for that click.
- →Google's 15 May 2026 generative search guidance explicitly warns against manufactured mentions. Editorial links on real media remain the defensible play, for classic ranking and for AI answers alike.
- →User-agent strings are trivially spoofed. Verify against OpenAI's published IP ranges before you draw any conclusion from a log line.
What GPTBot actually collects
GPTBot is the user agent OpenAI sends to fetch public web pages that may end up in the corpora used to train and improve its models. That is the whole job. It does not build the index ChatGPT queries when a user asks a question with browsing enabled, it does not decide whether your brand gets named in an answer, and it does not send you a single visitor. The only lever it gives you is upstream: whether OpenAI is allowed to read your content for model building.
This is why the decision to allow or disallow it is a licensing question dressed up as a technical one. A publisher with a syndication business and a legal team has a reason to close the door. A B2B SaaS with forty pages of documentation and a content marketing budget has almost none, and the ones that block anyway are usually copying a robots.txt snippet they found in a thread without checking what else that snippet disallows. We audit sites every month where the blanket AI blocklist was pasted in eighteen months ago and nobody has revisited it since, while the same client asks in the next meeting why competitors get named in ChatGPT and they do not.
The other thing worth stating plainly: robots.txt is a request, and the compliance model differs per agent. PPC Land reported on 9 December 2025 that OpenAI revised its crawler documentation and removed the explicit robots.txt compliance language attached to ChatGPT-User, the agent that fetches a page because a human asked for it in a conversation. GPTBot and OAI-SearchBot kept their documented controls. If your policy needs to be enforced rather than merely declared, it belongs at the WAF or CDN layer, not in a text file at the root.
The 2026 crawler split, and why robots.txt alone is thin
Three agents, three jobs. GPTBot collects for training. OAI-SearchBot supports the ChatGPT Search index, which is the one that determines whether your pages are candidates for citation. ChatGPT-User fetches a specific URL in response to a live user request. Treating them as one entity is the single most common technical error we see on this topic, and it produces two opposite failures: sites that block everything and vanish from answers, and sites that allow everything while their legal position on training says otherwise.
The configuration that expresses the most common intent, meaning keep me out of training and keep me visible in search, looks like this:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: / That is a strategic choice, not a default recommendation. It only makes sense if you have something to protect and enough brand pull that OpenAI has other reasons to know who you are. Most sites should allow all three and spend the saved deliberation time on something that moves a number. Either way, the mechanics of writing directives for AI crawlers deserve the same care as any other directive: a rule that never matches the agent you meant to target is a rule that does nothing at all.
Two verification points before you act on anything you see. First, user-agent strings are a header, and scrapers set them to whatever gets them through. OpenAI publishes IP ranges per crawler, and a log line that does not resolve to those ranges is not GPTBot regardless of what it calls itself. Second, blocking a crawler in robots.txt is not the same as removing a page from an index, exactly as it never was with Googlebot. If a URL is already known, a crawl block prevents refetching, it does not retract what was already learned.
Volume versus coverage: read your logs, not the headlines
The growth numbers are real and they are widely misread. Cloudflare's measurements for May 2024 to May 2025 showed overall crawler traffic up 18%, Googlebot requests up 96%, and GPTBot requests up 305%. Its share of verified crawler traffic moved from 4.7% in July 2024 to 11.7% in July 2025. Those are request counts, and a triple-digit growth rate off a small base is exactly what an expanding crawler looks like when it starts refetching aggressively.
Set that against coverage. Cloudflare's Year in Review analysis of HTML requests across October and November 2025 found Googlebot reached 11.6% of unique web pages, against 3.6% for GPTBot and 0.06% for PerplexityBot. Google still sees roughly three times the surface of the web that OpenAI's training crawler does. Anyone selling you a strategy built on the premise that AI crawlers have overtaken classic search discovery is selling a narrative the measurement does not support, at least not yet.
The practical consequence is that this analysis lives in server logs and nowhere else. GPTBot generates no sessions, no referrals, no entries in GA4. If you want to know whether OpenAI is reading you, whether it is reading the sections you care about, and at what cadence, you parse access logs and segment by verified user agent. Google's Search Console reports for generative AI features, launched on 3 June 2026 and initially rolled out to a subset of sites, cover Google's own AI Overviews and AI Mode surfaces with impressions, pages shown, countries and devices. Useful, and orthogonal: they tell you nothing about OpenAI, and an AI impression is not a visit.
What GPTBot access changes in a netlinking operation
Here is where the strategy actually sits. Chartbeat data reported by Axios on 17 March 2026 showed ChatGPT referrals up more than 200% year over year while remaining under 1% of publisher page views, over a period where Google Search referrals fell 34% from December 2024 to December 2025. Read those two figures together and the conclusion is uncomfortable but clear: the click is not coming back through the chatbot. What you are competing for is being named, described accurately, and associated with the right category inside an answer someone else's model composes.
Which brings the question back to a very old one. If a model's picture of your market is assembled from what the web says about you, then the surface that matters is third-party editorial coverage on sites the crawlers actually read at depth, not your own about page. That is the same asset base that produces citations inside generated answers and the same one that has always produced rankings. It is why we treat generative visibility as a downstream effect of publishing on real media rather than as a separate discipline, and why the work of getting a brand named inside generative answers starts with the same media selection as a classic campaign.
Google said as much itself. The guidance published by Google Search Central on 15 May 2026, on optimising for generative AI features in Search, states that generative visibility depends on crawlability, useful non-commodity content and the existing Search systems, and it cautions specifically against artificial or inauthentic mentions engineered to influence AI answers. There is no AI-only markup to add and no separate ranking system to game. Digital PR, original research and genuine editorial citations hold up under that guidance because they build discoverability, not because a file in your root promises anything.
Operationally, that argues for depth over spread. A brand mentioned once across forty unrelated sites reads as noise to a model assembling entity associations. The same brand covered repeatedly across a coherent set of thematic media reads as a fact about the category. We run our own French editorial media in-house at nautilinks precisely so that this kind of thematic concentration is something we can plan rather than hope for, and it is the logic behind how we phase an editorial campaign across several months instead of buying a batch. If you want to see which media are involved before committing to anything, the catalogue is open without an account, metrics and prices included.
What we see go wrong
The blanket blocklist is the classic. A team decides it does not want its content training a model, pastes a list of AI user agents into robots.txt, and takes OAI-SearchBot down with it. Six months later the brand is absent from ChatGPT answers in its own category and someone proposes a GEO budget to fix a problem the robots file created for free. Read every line of that file before you accept a diagnosis about AI visibility.
The mirror error is treating GPTBot hits as a performance metric. We have seen dashboards where AI crawler requests are charted next to sessions, as if a fetch were an outcome. A crawler request tells you the door is open. It says nothing about whether the model retained anything, and nothing at all about whether a user ever sees your name.
Third, the expectation that Search Console's generative reports will attribute ChatGPT. They will not. Those reports cover Google's surfaces. Cross-engine visibility measurement in 2026 still means sampled prompt testing plus log analysis, and anyone offering a clean attribution number for a chatbot citation is modelling, not measuring. Say so to the client before they build a target around it.
Last, the sequencing mistake. Teams block training access first and then wonder why the model's description of them is thin or two years stale. If a model has no fresh reading of your category and no third-party coverage to triangulate against, it falls back on whatever it absorbed earlier. Blocking is a defensible choice when you have something worth defending. It is not a strategy on its own, and it is never a substitute for being written about.
Nautilinks operates an owned network of editorial media. In-house written articles, transparency disclosures respected, anchor mix calibrated.
Frequently asked questions
Does blocking GPTBot remove my site from ChatGPT answers?
No, provided OAI-SearchBot stays allowed. GPTBot governs training access, OAI-SearchBot feeds the ChatGPT Search index that determines citation candidacy. The two are separately documented and separately controllable. What blocking GPTBot does affect is the model's baseline knowledge of you built during training, which matters for answers generated without a live retrieval step. Most sites that lost ChatGPT visibility after a robots.txt change had blocked both agents without noticing.
GPTBot requests grew 305% in Cloudflare's data. Should I be worried about crawl load?
Check your own logs before deciding. That figure covers May 2024 to May 2025 across Cloudflare's network, and it is request growth from a small base. The same source found GPTBot reached 3.6% of unique pages in October and November 2025 against 11.6% for Googlebot. For most sites the load is negligible. For large catalogues with heavy dynamic rendering it can be real, and rate limiting at the CDN is the answer rather than an outright block.
How do I tell real GPTBot traffic from spoofed requests?
Verify the source IP against OpenAI's published ranges for each crawler. User-agent headers are attacker-controlled and scrapers routinely impersonate GPTBot to inherit the reputation of an allowed agent. Anything that fails the IP check is not OpenAI regardless of what the string says. Same discipline you would apply to a Googlebot claim, and it is the only way log-based conclusions about AI crawling hold up.
Is there any markup or file that improves GPTBot pickup?
Nothing that OpenAI documents as such, and Google's 15 May 2026 guidance on generative search says the equivalent about its own features: no AI-specific markup is required and manufactured mentions are explicitly discouraged. What helps is unremarkable. Server-rendered HTML, clean internal linking, content that is not a commodity restatement of the same five sources, and third-party editorial coverage that gives a model something to triangulate against.
If ChatGPT referrals are under 1% of page views, why invest in this at all?
Because the click is the wrong unit. The Chartbeat data reported by Axios on 17 March 2026 showed ChatGPT referrals up more than 200% year over year and still under 1% of publisher page views, over the same window Google referrals fell 34%. The value is being named as a credible option inside an answer, at the moment a buyer is shortlisting. That is brand presence, measured through sampled prompt testing, not a traffic line.
Should a client with a licensing business handle this differently?
Yes, and it is the one case where blocking GPTBot is clearly right. If your content is the product, allowing free ingestion into a training corpus undercuts the deal you are trying to sell. Keep OAI-SearchBot open so search visibility survives, enforce the training block at the WAF rather than trusting robots.txt alone, and treat the block as an opening position in a commercial negotiation rather than a permanent technical setting.
Test your knowledge
Quiz: GPTBot
1/3Which OpenAI user agent determines whether your pages are candidates for citation in ChatGPT Search?