---
title: "Research Bot"
description: "Hospitality Research Bot"
url: "https://www.holidayhero.com/research-bot"
updated: "2026-10-04T10:23:49+00:00"
---

# Research Bot

<p>If you found <code>HolidayHeroResearchBot</code> in your server logs, this page explains what it was doing and how to stop it.</p><h2><strong>What it is</strong></h2><p>HolidayHero runs an industry study of Dutch hotel websites. We measure things guests and search engines run into: whether a link to the site shows a proper preview when shared on WhatsApp, whether the sitemap is valid, which booking engine and chat tools a site uses, and how accessible the homepage is.</p><p>We publish <strong>aggregate numbers only</strong>, for example &quot;x% of 4-star hotels have a chat widget&quot;. Individual hotels are never named in published findings.</p><h2><strong>How it identifies itself</strong></h2><p>The crawler renders pages in a regular Chromium browser, so pages look the way they do for a visitor. It always adds its own name to the browser&#039;s user agent:</p><pre><code>Mozilla/5.0 (...) Chrome/... Safari/537.36 HolidayHeroResearchBot/1.0 (+https://www.holidayhero.com/research-bot)</code></pre><p>To check how link previews behave, it fetches the homepage and its preview image once more with the user agent that Facebook&#039;s and WhatsApp&#039;s link-preview crawler uses (<code>Facebookexternalhit/1.1</code>). That is the only request that does not carry our name, and it is only made when robots.txt allows our crawler.</p><h2><strong>What it reads</strong></h2><p>Per website, once per study run:</p><ul><li><p><code>robots.txt</code>, before anything else</p></li><li><p>the homepage, as raw HTML and rendered in the browser, with a screenshot</p></li><li><p>the booking page linked from the homepage, one click deep, if robots.txt allows it</p></li><li><p><code>/llms.txt</code>, your sitemap(s), and a sample of at most 20 URLs listed in the sitemap, to check they still resolve</p></li><li><p>the <code>og:image</code> declared on the homepage</p></li></ul><p>It does not fill in forms, make or hold bookings, log in, or follow links any further. Requests to one host are spaced at least one second apart.</p><h2><strong>How to opt out</strong></h2><p>The crawler follows robots.txt (<a href="https://www.rfc-editor.org/rfc/rfc9309">RFC 9309</a>). To block it from your whole site, add:</p><pre><code>User-agent: HolidayHeroResearchBot
Disallow: /</code></pre><p>A site that blocks the crawler is recorded only as &quot;not crawled&quot; and is left out of the measurements. If robots.txt is unreachable or returns a server error, we treat that as a full block and don&#039;t crawl.</p><h2><strong>Contact</strong></h2><p>Questions, or want your site removed from a study run that has already happened? Email dev@holidayhero.com with your domain, and we will delete the collected data for it.</p>
