# Retro Delights — retrodelights.co.uk # # Our position, in short: you are welcome to READ this site and cite it in an # answer. You are not welcome to hoover it into a training corpus without # asking. That is the same line /legal/ draws in prose; this file is it in # machine-readable form. # # ── THE PRECEDENCE RULE (RFC 9309, and Google's robots spec) ────────────── # Named user-agent groups and the wildcard `*` group are NOT combined. A bot # obeys the MOST SPECIFIC group matching it and ignores every other group, # INCLUDING `*`. That is why `Disallow: /.netlify/` is repeated inside every # named group below. Delete a repetition and that bot starts crawling the # function endpoints. # # robots.txt is a request, not access control. Perplexity-User, # meta-externalfetcher and Google's user-triggered fetchers all document that # they ignore it when a human asked for the page. Anything that must not be # read by a machine needs a real auth gate, not a line in this file. # ── DEFAULT ─────────────────────────────────────────────────────────────── User-agent: * Disallow: /.netlify/ # ── ALLOWED: the crawlers that put us in AI answers ────────────────────── # These are RETRIEVAL crawlers. They build live search indexes and fetch # pages to answer a question someone is asking right now. Blocking one # removes Retro Delights from that assistant's answers. They are listed # explicitly so nobody tidying this file later blocks one by mistake, and # each repeats the Disallow above per the precedence rule. User-agent: OAI-SearchBot Disallow: /.netlify/ User-agent: ChatGPT-User Disallow: /.netlify/ User-agent: Claude-SearchBot Disallow: /.netlify/ User-agent: Claude-User Disallow: /.netlify/ User-agent: PerplexityBot Disallow: /.netlify/ User-agent: Perplexity-User Disallow: /.netlify/ User-agent: Googlebot Disallow: /.netlify/ User-agent: Googlebot-Image Disallow: /.netlify/ User-agent: bingbot Disallow: /.netlify/ User-agent: Applebot Disallow: /.netlify/ User-agent: DuckAssistBot Disallow: /.netlify/ User-agent: MistralAI-User Disallow: /.netlify/ User-agent: Amzn-SearchBot Disallow: /.netlify/ User-agent: YouBot Disallow: /.netlify/ User-agent: Bravebot Disallow: /.netlify/ User-agent: kagi-fetcher Disallow: /.netlify/ # ── BLOCKED: training-corpus and data-broker crawlers ──────────────────── # These fetch to build training sets or to resell scraped text. None feeds a # live answer surface, so blocking them costs nothing in citations — every # assistant listed above is unaffected. Cloudflare Radar puts Anthropic's # crawl-to-referral ratio near 4,580:1 and OpenAI's near 848:1, so this is # bandwidth we pay for and get almost nothing back from. # # GPTBot and ClaudeBot are the TRAINING bots. OAI-SearchBot and # Claude-SearchBot, allowed above, are the SEARCH ones. Do not confuse them — # blocking the wrong pair of these is how a site disappears from ChatGPT. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: CCBot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: FacebookBot Disallow: / User-agent: Bytespider Disallow: / User-agent: cohere-ai Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: AI2Bot Disallow: / User-agent: Ai2Bot-Dolma Disallow: / User-agent: Diffbot Disallow: / User-agent: Timpibot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / User-agent: Webzio-Extended Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: img2dataset Disallow: / User-agent: PanguBot Disallow: / User-agent: DeepSeekBot Disallow: / User-agent: Scrapy Disallow: / Sitemap: https://retrodelights.co.uk/sitemap.xml