# AI Crawler Guidance For Horizons Homecare ## Purpose This file explains intended AI crawler access in plain language. It complements /robots.txt but does not replace it. ## Allowed For Search And Citation These retrieve at query time and are the ones that produce citations. - Googlebot - Bingbot - OAI-SearchBot - ChatGPT-User - PerplexityBot - ClaudeBot ## Allowed For Training And Model Grounding - GPTBot - anthropic-ai - CCBot - Meta-ExternalAgent - Bytespider ## Permission Tokens (not crawlers, these make no requests) - Google-Extended (governs Gemini and Google AI Overviews) - Applebot-Extended (governs Apple Intelligence) Both tokens are opt-out by default, so they are named here deliberately rather than permitted by silence. ## Anything Not Named Above Also permitted. /robots.txt allows User-agent: * across the whole site, so a crawler absent from this list is still welcome. The site owner has chosen maximum visibility: Horizons Homecare wants to be found and cited everywhere, by search engines and AI assistants alike. One restriction applies to everyone, named or not: /api/ is disallowed for every group in /robots.txt. Those endpoints serve the site's own forms and already-public review data, and there is nothing there for a crawler to index. Crawler names change often, and a wrong user-agent string fails silently. This list was verified against current references on 28/07/2026. ## Canonical Files - Robots.txt: https://www.horizonshomecare.co.uk/robots.txt - AI permissions: https://www.horizonshomecare.co.uk/ai.txt - AI identity: https://www.horizonshomecare.co.uk/llms.txt - Structured identity: https://www.horizonshomecare.co.uk/identity.json