Opinion and directional analysis. This piece is based on public information at the time of writing and may be out of date by the time you read it. Any site behavior described here can change without notice, so verify the current state yourself before acting. We are not affiliated with OpenAI, Perplexity, Anthropic, Google, or any brand named below. This is not business, legal, or financial advice.

Download this study as a one-page PDF
There is a quiet decision buried in a plain text file on every website, and it is starting to shape who gets recommended when a traveler asks an AI where to stay. The file is robots.txt. The decision is whether to let the AI crawlers in.
I have spent the last stretch pulling up that file on travel sites big and small, and the pattern is more interesting than “everyone blocks the bots” or “nobody does.” Some very large travel brands slam the door on GPTBot and its cousins. Others leave it wide open. And a lot of independent hotels have no idea the door exists, which is its own kind of answer.
This is an opinion piece about what those choices signal, why the big players make them, and why your hotel should almost certainly make the opposite one. Let me walk through it.
What robots.txt actually does (and does not do)
robots.txt is a file that lives at the root of a domain. Type any domain into your browser, add /robots.txt on the end, and you will see it. It is a set of polite requests: “Dear crawler named X, please do not read these parts of my site.” Well-behaved bots honor it. It is not a lock. It is a sign on the lawn.
The lines that matter look like this. A User-agent line names a specific crawler. A Disallow line names a path that crawler should skip. Disallow: / means “skip the whole site.” An empty Disallow: or Allow: / means “come on in.”
Here are the AI-era user agents worth knowing:
| Crawler | Who runs it | What it feeds |
|---|---|---|
| GPTBot | OpenAI | Training and index refresh for ChatGPT |
| OAI-SearchBot | OpenAI | Live search results inside ChatGPT |
| PerplexityBot | Perplexity | Perplexity’s answer engine |
| ClaudeBot | Anthropic | Claude’s crawling |
| Google-Extended | Whether Google may use your content for Gemini and AI features | |
| Googlebot / Bingbot | Google / Microsoft | Classic search indexing |
One crucial nuance: Google-Extended is not a crawler at all. It is a toggle. Blocking it tells Google “you may still rank me in classic search, but do not use my content to train or ground your AI answers.” That distinction trips up a lot of people, and it matters when you read a big brand’s file and try to reverse-engineer their strategy. For the bigger picture on why any of this shows up in AI answers, we wrote is your hotel invisible to ChatGPT.
Why some big travel sites block the AI crawlers
When a large travel platform blocks GPTBot with a flat Disallow: /, it is usually not because they are technophobes. It is a calculated move, and the logic is worth understanding even if you should not copy it.
They are protecting a data moat. A giant aggregator’s core asset is its inventory, its pricing, and its review corpus, accumulated over years. If an answer engine can freely ingest all of that and hand travelers a synthesized recommendation, the aggregator becomes a commodity backend that never gets the click, the account, or the loyalty signup. Blocking is a way of saying “if you want our data, come license it, do not scrape it for free.” Some of the loudest robots.txt blocks in travel are pure negotiating posture.
They fear disintermediation. This is the funny part. The online travel agencies built their empires by sitting between the traveler and the hotel. Now they are staring at a new intermediary, the AI assistant, that could sit between the traveler and them. The same disintermediation the OTAs did to hotels is the thing they now dread. I have written about the original version of this dynamic in how OTAs steal search and in the OTA destination landers mega guide. Watching them defend against it is a study in irony.
They have a licensing strategy. Several large content and travel companies have signed paid deals with AI firms. If you are getting paid for your data, you block the free crawler and route access through the contract. That is a rational enterprise play. It also requires a legal team, a business development team, and enough leverage to get an AI company to sign. You do not have that leverage. Neither do I. That is fine, because we do not want it.
They are managing crawl load and cost. At the scale of hundreds of millions of pages, an aggressive crawler can be a genuine infrastructure expense. Blocking or throttling is sometimes just operations, not strategy. A 40-room hotel does not have this problem.
So when you see a travel behemoth blocking the AI bots, read it as: a company with a data moat, licensing leverage, and fear of being disintermediated, acting in its own narrow interest. None of those three conditions describes an independent hotel.
Why an independent hotel should do the opposite
Here is the whole argument in one line: the big platforms block AI crawlers to protect the thing that makes them the middleman. You are not the middleman. You are the destination the traveler actually wants. Being readable is your advantage, not your risk.
Think about what happens when a traveler asks an assistant, “Where should I stay in [your city] near the old town, walkable, with parking and a real breakfast?” The assistant assembles an answer from what it can read and ground itself in. If your site is open, structured, and specific, you can be in that answer. If you have blocked the crawlers, you have opted out of the shortlist before it was even drawn. You did not protect anything. You just disappeared.
For an independent, AI visibility is the rare channel where you are not paying a 15 to 25 percent toll for the introduction. Compare that to the OTA math, which we lay out in the book-direct math on OTA commission cost. Every AI-sourced guest who lands on your own site and books direct is a guest who did not cost you a commission. Blocking the crawlers is, functionally, choosing the expensive channel over the free one.
The mental model that helps: aggregators want to be the answer, so they hoard the ingredients. You want to be in the answer, so you make yourself easy to read. Those are opposite goals, which is why copying a big brand’s robots.txt is a mistake.
There is a reputational angle too. When your content, your policies, your neighborhood knowledge, and your honest description of the rooms are the source an assistant leans on, the assistant describes you the way you describe yourself. When you are absent, it describes you using whatever third-party fragments it can find, which often means an OTA listing, an old review, or a competitor’s comparison page. Silence does not protect your story. It hands the narration to someone else. This is the core of what we do in AI visibility, AEO and GEO.
How to check any site’s robots.txt in two minutes
You do not need tools for this. You need a browser.
- Take the root domain, for example the hotel’s main domain.
- Add
/robots.txtto the end and load it. - Scan for
User-agentblocks that name the AI crawlers above. - Look at the
Disallowline under each.Disallow: /is a full block. A blank disallow orAllow: /is open. - Check for
Google-Extendedspecifically, since that is the AI-usage toggle rather than a crawler.
Do this for your own site first. I am consistently surprised how many hotels find a Disallow: / they never intended, usually a leftover from a staging environment, a security plugin’s default, or a web developer who copied a template. That single line can quietly wall you off from both classic search and AI. If you find it and did not put it there, that is your afternoon sorted.
Then check three or four competitors and, for fun, a couple of the big OTAs. You will start to see the strategic split I described: platforms hoarding, independents mostly wide open by accident rather than intent. Turning that accident into a deliberate, welcoming, well-structured posture is most of the game.
A few practical cautions when you read these files:
- Behavior changes. A brand that blocks today may open tomorrow, and vice versa. Treat any snapshot as a snapshot. Verify before you quote it to anyone.
- A block in robots.txt does not guarantee the content is absent from an AI answer. Assistants can surface you via search partners, licensed feeds, or third-party pages that mention you. Blocking reduces your control, it does not create a clean disappearance.
- Do not confuse robots.txt with a
noindextag or a login wall. They do different jobs. If content is genuinely private, protect it properly. robots.txt is a request, not a guard.
What “welcoming the crawlers” should actually look like
Opening the door is necessary but not sufficient. An open door onto an empty, unreadable room does not help. The hotels that win in AI answers do three things beyond simply allowing the bots.
They are specific. Vague marketing prose (“an oasis of tranquility”) gives an assistant nothing to ground a recommendation in. Concrete facts do: walking minutes to the station, exact parking situation, whether the breakfast is cooked to order, the real pet policy, the floor the quiet rooms are on. Specificity is what gets you matched to a specific query.
They are structured. Clear headings, honest FAQs, a genuine local guide, and clean policy pages give machines and travelers the same thing: fast, unambiguous answers. Our content and reputation work is largely about turning a hotel’s real knowledge into that kind of readable substance, and our hotel SEO foundation makes sure the crawlers can reach it.
They are consistent everywhere. Your name, address, amenities, and policies should match across your site, your Google Business Profile, and the major directories. Contradictions make you a low-confidence source, and low-confidence sources get dropped from answers. If you run longer-stay inventory, the same logic applies to the details covered in aparthotel and extended-stay marketing.
The through-line: being crawlable is the price of entry. Being the clearest, most specific, most consistent source is what actually earns the mention. For more on the emerging data pipelines that feed these systems, the Flip.to and Spacetime sleeping-giant piece is worth a read.
The one-line takeaway
Big travel platforms block AI crawlers to defend a middleman position you do not hold and do not want. For an independent hotel, the robots.txt question is not “how do I keep the AI out,” it is “have I accidentally kept it out, and how do I make myself the easiest hotel in [your city] to recommend.” Go check your file. Then check your competitors’. The gap between those two is your opportunity.
If you want a second set of eyes on your robots.txt, your AI visibility, and the specific queries you should be winning, that is exactly the kind of audit we run. Book a call and we will pull up your file together, live, in the first ten minutes.
Reminder: this is opinion and directional analysis, based on public information at the time of writing, and may be out of date. Crawler behavior, robots.txt policies, and the AI products named here change frequently, so verify the current state before acting. Any figures referenced elsewhere on this site are third-party estimates subject to change and independent validation. We are not affiliated with the companies named. This is not business, legal, or financial advice.
How we got this: the patterns above come from manually reading publicly accessible robots.txt files and public statements about AI crawler policies at the time of writing. No private data was used, and no single site’s current behavior should be assumed from this general analysis.