User-agent: * Allow: / Allow: /listings # --- AI / LLM discovery surface (Feb 2026, per Doug) --- # These paths MUST stay crawlable so agents / answer engines can discover # the Doogie tool schema and the citation / AI-plugin manifests. Placed # above the /api/ Disallow so longest-match Allow wins on Google / Bing. Allow: /api/doogie/tools.json Allow: /llms.txt Allow: /llms-full.txt Allow: /.well-known/ai.json Allow: /.well-known/ai-plugin.json # Explicit MLS listing lockdown (CREA DDF® §8) Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ Disallow: /my-journey/ Disallow: /admin/ Disallow: /market-report # Backend API surface is otherwise blocked from all crawlers. Provides # live CREA DDF® data under a limited board licence; scraping / # redistribution is prohibited by DDF® terms. Longest-match Allow rules # above keep /api/doogie/tools.json and the AI-discovery manifests # fetchable even though /api/ as a whole is blocked here. Disallow: /api/ # Task 10 (Feb 2026): Crawl-Delay directive intentionally REMOVED from # the wildcard group. Well-behaved AI crawlers (Googlebot, Bingbot, # GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended, # Applebot, DuckAssistBot) either ignore Crawl-Delay entirely or apply # their own adaptive throttle; setting a value here caused unnecessary # indexing lag on the educational corpus without helping anti-scraping # outcomes. Bulk-scraper defense-in-depth now relies on: # 1. Application-level per-IP rate limits on /api/listings and # /listing/{key} (still enforced server-side). # 2. 410-Gone responses for sold / expired records. # 3. Explicit Disallow: /listing/ + Disallow: /api/listings for every # AI user-agent (see individual groups below). # CREA DDF® Rule 3.1: raw per-listing pages (/listing/{key}) must not be # machine-scraped or AI-trained. The /listings search hub, however, does # NOT expose per-property MLS data and is safe (and beneficial) to index # so it can rank for "bc real estate search", "maple ridge homes for sale". # # /market-report was retired 2026-08-08 (Doug's decision). Disallow tells # Google/Bing to drop it from the index over the next few crawls. # AI answer engines welcome — for EDUCATIONAL content only # CREA DDF® terms prohibit scraping/AI-training against MLS® listing data, # so /listings and /listing/{id} are excluded for every bot. User-agent: GPTBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: ChatGPT-User Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: OAI-SearchBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: ClaudeBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: anthropic-ai Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: PerplexityBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: Perplexity-User Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: Applebot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: DuckAssistBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: cohere-ai Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ # --- Explicitly welcomed AI answer engines (added Feb 2026 per Doug's ask) --- # All the following bots return citations back to eztofind.ca, so they # get the same educational-content Allow policy as GPTBot/ClaudeBot above. User-agent: xAI-Bot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: Grokbot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: MistralAI-User Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: MistralAI-Bot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: YouBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: You.com Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: KagiBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: Neevabot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: ManusBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: BingBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: Bingbot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: msnbot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: AdsBot-Google Allow: / User-agent: AmazonQBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: Amazonbot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ Disallow: /admin/ User-agent: FacebookBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ Disallow: /admin/ User-agent: Meta-ExternalAgent Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ Disallow: /admin/ User-agent: GitHub-Copilot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: GitHubCopilotChat Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: DuckDuckBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ # LLM training corpora — mixed policy. # Google-Extended controls citation in Google AI Overviews / Gemini / Bard. # For AEO reach, we explicitly Allow it on educational content while keeping # MLS® listings disallowed (CREA DDF® licensing). CCBot (Common Crawl) is # similarly allowed for citation. Training-only bots that don't cite back # remain fully blocked below. User-agent: Google-Extended Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: CCBot Allow: / Allow: /api/doogie/tools.json Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ # Training-only AI crawlers below are fully blocked. # (FacebookBot / Amazonbot / Meta-ExternalAgent duplicates removed # Task 10 (Feb 2026) — /admin/ Disallow was folded into their primary # groups above so each bot now has ONE canonical rule set.) User-agent: Meta-ExternalFetcher Disallow: /admin/ User-agent: Applebot-Extended Disallow: / User-agent: cohere-training-data-crawler Disallow: / User-agent: ImagesiftBot Disallow: / User-agent: PetalBot Disallow: / User-agent: Timpibot Disallow: / User-agent: OmigiliBot Disallow: / # Block SEO scrapers and data brokers (not search / AI) User-agent: Bytespider Disallow: /admin/ User-agent: Diffbot Disallow: / User-agent: MJ12bot Disallow: / User-agent: DataForSeoBot Disallow: / User-agent: ZoominfoBot Disallow: / User-agent: SemrushBot Disallow: / # Ahrefs — the site owner runs Ahrefs Web Analytics + Ahrefs Webmaster Tools # on this property, so both AhrefsBot (backlink + keyword index) and # AhrefsSiteAudit (technical health crawler) are fully welcomed on the # educational corpus. Per-listing MLS® pages (/listings, /listing/) are # excluded to match the CREA DDF® bot policy applied to every other # crawler above — the data is licensed, hourly-changing, and non-canonical. User-agent: AhrefsBot Allow: / Allow: /api/doogie/tools.json Disallow: /admin/ Disallow: /my-journey/ Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ User-agent: AhrefsSiteAudit Allow: / Allow: /api/doogie/tools.json Disallow: /admin/ Disallow: /my-journey/ Disallow: /listings Disallow: /api/listings Disallow: /listing/ Disallow: /mockups/ Sitemap: https://eztofind.ca/sitemap.xml Sitemap: https://eztofind.ca/sitemap-ai.xml Sitemap: https://eztofind.ca/sitemap-snapshots.xml Sitemap: https://eztofind.ca/sitemap-index.xml # RSS 2.0 editorial feed — Feedly / Inoreader / NewsBlur crawl this # aggressively (often multiple times per hour), which is faster than the # regular sitemap crawl for new sites. Bing / DuckDuckGo also honour the # `pubDate` fields. Regenerated on server boot + hourly via cron. # Not a sitemap declaration (search engines only accept Sitemap: for XML # sitemaps), but listed here for feed-reader discovery. # Feed: https://eztofind.ca/feed.xml # (Task 10, Feb 2026 — the trailing `User-agent: *` API-lockdown block # that used to live here has been merged into the primary wildcard group # at the top of this file. `Disallow: /api/` with the AI-discovery # `Allow:` overrides now lives there.)