Skip to content
HotelSEO Lab
← The Lab
AI Visibility (AEO/GEO)

The Robots.txt Hall of Shame: Which Big Travel Sites Block the AI Crawlers (and What It Signals)

A directional look at why some big travel brands block GPTBot, PerplexityBot, ClaudeBot and Google-Extended in robots.txt, what the choice signals, and why an independent hotel should almost always do the opposite.

HotelSEO LabJuly 1, 2026 12 min read

Opinion and directional analysis. This piece is based on public information at the time of writing and may be out of date by the time you read it. Any site behavior described here can change without notice, so verify the current state yourself before acting. We are not affiliated with OpenAI, Perplexity, Anthropic, Google, or any brand named below. This is not business, legal, or financial advice.

Animated infographic: robots txt ai crawler hall of shame

Download this study as a one-page PDF

There is a quiet decision buried in a plain text file on every website, and it is starting to shape who gets recommended when a traveler asks an AI where to stay. The file is robots.txt. The decision is whether to let the AI crawlers in.

I have spent the last stretch pulling up that file on travel sites big and small, and the pattern is more interesting than “everyone blocks the bots” or “nobody does.” Some very large travel brands slam the door on GPTBot and its cousins. Others leave it wide open. And a lot of independent hotels have no idea the door exists, which is its own kind of answer.

This is an opinion piece about what those choices signal, why the big players make them, and why your hotel should almost certainly make the opposite one. Let me walk through it.

What robots.txt actually does (and does not do)

robots.txt is a file that lives at the root of a domain. Type any domain into your browser, add /robots.txt on the end, and you will see it. It is a set of polite requests: “Dear crawler named X, please do not read these parts of my site.” Well-behaved bots honor it. It is not a lock. It is a sign on the lawn.

The lines that matter look like this. A User-agent line names a specific crawler. A Disallow line names a path that crawler should skip. Disallow: / means “skip the whole site.” An empty Disallow: or Allow: / means “come on in.”

Here are the AI-era user agents worth knowing:

CrawlerWho runs itWhat it feeds
GPTBotOpenAITraining and index refresh for ChatGPT
OAI-SearchBotOpenAILive search results inside ChatGPT
PerplexityBotPerplexityPerplexity’s answer engine
ClaudeBotAnthropicClaude’s crawling
Google-ExtendedGoogleWhether Google may use your content for Gemini and AI features
Googlebot / BingbotGoogle / MicrosoftClassic search indexing

One crucial nuance: Google-Extended is not a crawler at all. It is a toggle. Blocking it tells Google “you may still rank me in classic search, but do not use my content to train or ground your AI answers.” That distinction trips up a lot of people, and it matters when you read a big brand’s file and try to reverse-engineer their strategy. For the bigger picture on why any of this shows up in AI answers, we wrote is your hotel invisible to ChatGPT.

Why some big travel sites block the AI crawlers

When a large travel platform blocks GPTBot with a flat Disallow: /, it is usually not because they are technophobes. It is a calculated move, and the logic is worth understanding even if you should not copy it.

They are protecting a data moat. A giant aggregator’s core asset is its inventory, its pricing, and its review corpus, accumulated over years. If an answer engine can freely ingest all of that and hand travelers a synthesized recommendation, the aggregator becomes a commodity backend that never gets the click, the account, or the loyalty signup. Blocking is a way of saying “if you want our data, come license it, do not scrape it for free.” Some of the loudest robots.txt blocks in travel are pure negotiating posture.

They fear disintermediation. This is the funny part. The online travel agencies built their empires by sitting between the traveler and the hotel. Now they are staring at a new intermediary, the AI assistant, that could sit between the traveler and them. The same disintermediation the OTAs did to hotels is the thing they now dread. I have written about the original version of this dynamic in how OTAs steal search and in the OTA destination landers mega guide. Watching them defend against it is a study in irony.

They have a licensing strategy. Several large content and travel companies have signed paid deals with AI firms. If you are getting paid for your data, you block the free crawler and route access through the contract. That is a rational enterprise play. It also requires a legal team, a business development team, and enough leverage to get an AI company to sign. You do not have that leverage. Neither do I. That is fine, because we do not want it.

They are managing crawl load and cost. At the scale of hundreds of millions of pages, an aggressive crawler can be a genuine infrastructure expense. Blocking or throttling is sometimes just operations, not strategy. A 40-room hotel does not have this problem.

So when you see a travel behemoth blocking the AI bots, read it as: a company with a data moat, licensing leverage, and fear of being disintermediated, acting in its own narrow interest. None of those three conditions describes an independent hotel.

Why an independent hotel should do the opposite

Here is the whole argument in one line: the big platforms block AI crawlers to protect the thing that makes them the middleman. You are not the middleman. You are the destination the traveler actually wants. Being readable is your advantage, not your risk.

Think about what happens when a traveler asks an assistant, “Where should I stay in [your city] near the old town, walkable, with parking and a real breakfast?” The assistant assembles an answer from what it can read and ground itself in. If your site is open, structured, and specific, you can be in that answer. If you have blocked the crawlers, you have opted out of the shortlist before it was even drawn. You did not protect anything. You just disappeared.

For an independent, AI visibility is the rare channel where you are not paying a 15 to 25 percent toll for the introduction. Compare that to the OTA math, which we lay out in the book-direct math on OTA commission cost. Every AI-sourced guest who lands on your own site and books direct is a guest who did not cost you a commission. Blocking the crawlers is, functionally, choosing the expensive channel over the free one.

The mental model that helps: aggregators want to be the answer, so they hoard the ingredients. You want to be in the answer, so you make yourself easy to read. Those are opposite goals, which is why copying a big brand’s robots.txt is a mistake.

There is a reputational angle too. When your content, your policies, your neighborhood knowledge, and your honest description of the rooms are the source an assistant leans on, the assistant describes you the way you describe yourself. When you are absent, it describes you using whatever third-party fragments it can find, which often means an OTA listing, an old review, or a competitor’s comparison page. Silence does not protect your story. It hands the narration to someone else. This is the core of what we do in AI visibility, AEO and GEO.

How to check any site’s robots.txt in two minutes

You do not need tools for this. You need a browser.

  1. Take the root domain, for example the hotel’s main domain.
  2. Add /robots.txt to the end and load it.
  3. Scan for User-agent blocks that name the AI crawlers above.
  4. Look at the Disallow line under each. Disallow: / is a full block. A blank disallow or Allow: / is open.
  5. Check for Google-Extended specifically, since that is the AI-usage toggle rather than a crawler.

Do this for your own site first. I am consistently surprised how many hotels find a Disallow: / they never intended, usually a leftover from a staging environment, a security plugin’s default, or a web developer who copied a template. That single line can quietly wall you off from both classic search and AI. If you find it and did not put it there, that is your afternoon sorted.

Then check three or four competitors and, for fun, a couple of the big OTAs. You will start to see the strategic split I described: platforms hoarding, independents mostly wide open by accident rather than intent. Turning that accident into a deliberate, welcoming, well-structured posture is most of the game.

A few practical cautions when you read these files:

What “welcoming the crawlers” should actually look like

Opening the door is necessary but not sufficient. An open door onto an empty, unreadable room does not help. The hotels that win in AI answers do three things beyond simply allowing the bots.

They are specific. Vague marketing prose (“an oasis of tranquility”) gives an assistant nothing to ground a recommendation in. Concrete facts do: walking minutes to the station, exact parking situation, whether the breakfast is cooked to order, the real pet policy, the floor the quiet rooms are on. Specificity is what gets you matched to a specific query.

They are structured. Clear headings, honest FAQs, a genuine local guide, and clean policy pages give machines and travelers the same thing: fast, unambiguous answers. Our content and reputation work is largely about turning a hotel’s real knowledge into that kind of readable substance, and our hotel SEO foundation makes sure the crawlers can reach it.

They are consistent everywhere. Your name, address, amenities, and policies should match across your site, your Google Business Profile, and the major directories. Contradictions make you a low-confidence source, and low-confidence sources get dropped from answers. If you run longer-stay inventory, the same logic applies to the details covered in aparthotel and extended-stay marketing.

The through-line: being crawlable is the price of entry. Being the clearest, most specific, most consistent source is what actually earns the mention. For more on the emerging data pipelines that feed these systems, the Flip.to and Spacetime sleeping-giant piece is worth a read.

The one-line takeaway

Big travel platforms block AI crawlers to defend a middleman position you do not hold and do not want. For an independent hotel, the robots.txt question is not “how do I keep the AI out,” it is “have I accidentally kept it out, and how do I make myself the easiest hotel in [your city] to recommend.” Go check your file. Then check your competitors’. The gap between those two is your opportunity.

If you want a second set of eyes on your robots.txt, your AI visibility, and the specific queries you should be winning, that is exactly the kind of audit we run. Book a call and we will pull up your file together, live, in the first ten minutes.

Reminder: this is opinion and directional analysis, based on public information at the time of writing, and may be out of date. Crawler behavior, robots.txt policies, and the AI products named here change frequently, so verify the current state before acting. Any figures referenced elsewhere on this site are third-party estimates subject to change and independent validation. We are not affiliated with the companies named. This is not business, legal, or financial advice.

How we got this: the patterns above come from manually reading publicly accessible robots.txt files and public statements about AI crawler policies at the time of writing. No private data was used, and no single site’s current behavior should be assumed from this general analysis.

FAQ

Quick answers

Does blocking AI crawlers in robots.txt keep my hotel out of ChatGPT?

Partly, and not in the way most people think. Blocking GPTBot mainly stops OpenAI from crawling your pages to freshen its own index and to fetch live context. But answers can still cite you from third-party sources, licensed data, or a search partner's index. Blocking is a lever, not a light switch, and for a small hotel it usually removes upside without removing risk.

Which crawlers should an independent hotel allow?

For nearly every independent hotel, allow the mainstream AI crawlers: GPTBot, OAI-SearchBot, PerplexityBot, ClaudeBot, and Google-Extended, plus classic search bots like Googlebot and Bingbot. You want to be readable by the systems travelers now use to shortlist places to stay. The exception is genuinely private or paywalled content, which you should protect regardless.

How do I check what a hotel or competitor is blocking?

Add /robots.txt to the root domain in your browser, for example type the domain then slash robots dot txt. Read the User-agent blocks and their Disallow lines. A block like User-agent GPTBot followed by Disallow slash means that site is asking OpenAI's crawler to stay out of everything.

Is blocking AI crawlers ever the right move for a hotel?

Rarely, and only for narrow reasons: protecting rate content behind a booking engine, throttling an abusive crawler that is hammering your server, or a deliberate licensing strategy at enterprise scale. None of those describe a typical independent property trying to win more direct bookings.

Keep reading

More from the Lab

Free intro call

Let's go find out why the OTAs are outranking you for your own name.

20 free minutes. We'll look at your hotel live, show you where you're invisible — on Google and in the AI answers — and tell you straight whether we can help.

No lock-in · No 12-month handcuffs · You talk to the strategist