If you found HolidayHeroResearchBot in your server logs, this page explains what it was doing and how to stop it.

What it is

HolidayHero runs an industry study of Dutch hotel websites. We measure things guests and search engines run into: whether a link to the site shows a proper preview when shared on WhatsApp, whether the sitemap is valid, which booking engine and chat tools a site uses, and how accessible the homepage is.

We publish aggregate numbers only, for example "x% of 4-star hotels have a chat widget". Individual hotels are never named in published findings.

How it identifies itself

The crawler renders pages in a regular Chromium browser, so pages look the way they do for a visitor. It always adds its own name to the browser's user agent:

Mozilla/5.0 (...) Chrome/... Safari/537.36 HolidayHeroResearchBot/1.0 (+https://www.holidayhero.com/research-bot)

To check how link previews behave, it fetches the homepage and its preview image once more with the user agent that Facebook's and WhatsApp's link-preview crawler uses (Facebookexternalhit/1.1). That is the only request that does not carry our name, and it is only made when robots.txt allows our crawler.

What it reads

Per website, once per study run:

  • robots.txt, before anything else

  • the homepage, as raw HTML and rendered in the browser, with a screenshot

  • the booking page linked from the homepage, one click deep, if robots.txt allows it

  • /llms.txt, your sitemap(s), and a sample of at most 20 URLs listed in the sitemap, to check they still resolve

  • the og:image declared on the homepage

It does not fill in forms, make or hold bookings, log in, or follow links any further. Requests to one host are spaced at least one second apart.

How to opt out

The crawler follows robots.txt (RFC 9309). To block it from your whole site, add:

User-agent: HolidayHeroResearchBot
Disallow: /

A site that blocks the crawler is recorded only as "not crawled" and is left out of the measurements. If robots.txt is unreachable or returns a server error, we treat that as a full block and don't crawl.

Contact

Questions, or want your site removed from a study run that has already happened? Email [email protected] with your domain, and we will delete the collected data for it.