# robots.txt for iampro.io # AI-powered job matching platform # # Served statically from the host (/var/www/iampro/robots.txt, published by the # deploy job) so it always answers 200 — a 5xx here throttles crawling of the # WHOLE domain. See `location = /robots.txt` in nginx/iampro.conf. # # STRUCTURE, and it is not cosmetic: a crawler obeys exactly ONE group — the # most specific User-agent match — and ignores every other group, `*` included. # So a named group listing only `Allow:` lines grants that crawler the whole # site, Disallows and all. Before 2026-07-29 each AI crawler had its own # Allow-only group, which quietly re-opened /admin, /api/ and /dashboard to all # eleven of them. Any Disallow that matters must be repeated inside every group # that needs it — hence the shared multi-User-agent group below. User-agent: * Allow: / Allow: /login Allow: /register Allow: /extension Allow: /getting-started Allow: /tools Allow: /tools/ Allow: /careers Allow: /careers/ Allow: /static/ # Block private/authenticated areas. # /dashboard is slated for removal — logged-in users now land on the unified # landing (/#home, /#search, /#activity). No preview carve-out needed. Disallow: /dashboard Disallow: /admin Disallow: /api/ Disallow: /verify/ Disallow: /reset-password Disallow: /logout # Block utility endpoints Disallow: /health # --------------------------------------------------------------------------- # LES VARIANTES DE LANGUE, ET POURQUOI ELLES SE FERMENT AUSSI POUR GOOGLE. # # `?lang=xx` sert la MÊME page sous une chaîne de requête, et la page pointe sa # canonique vers l'URL nue. Le groupe des crawlers d'IA plus bas les bloquait # déjà ; Googlebot et bingbot en étaient exemptés, au motif que « la découverte # des hreflang vaut le budget ». # # Ce motif ne tenait pas. Mesuré le 2026-08-22 : # # · AUCUN hreflang sur ces pages. 24 entrées du sitemap en portent — les # pages statiques — et zéro des 4 728 pages /careers/. Il n'y avait donc # rien à découvrir. # · La traduction ne couvre que l'habillage : sur /careers/accountant/berlin, # 93,1 % du texte est IDENTIQUE entre la version par défaut et ?lang=fr. # Le titre passe de « Buchhalter/in » à « Comptable », les annonces ne # bougent pas. Ce ne sont pas des versions linguistiques, ce sont des # quasi-doublons. # · Google les explore et ne peut pas les indexer, puisqu'elles se déclarent # elles-mêmes doublons. Sur les 999 URL de l'export « Crawled – currently # not indexed », 496 — la moitié — portaient un ?lang=. Une même page y # figure jusqu'à quatre fois, une par langue. # · Coût mesuré dans les journaux nginx : 12,6 % des requêtes de Googlebot # sur deux jours portaient un ?lang=. # # On ferme donc la même porte pour tout le monde. La canonique reste en place # et sa cible est explorée et indexée séparément — c'est le traitement standard # d'un paramètre qui produit un doublon. # # ⚠️ À ROUVRIR si un jour les pages portent un vrai jeu hreflang ET un contenu # réellement traduit. Ce sont les deux conditions, pas une. Disallow: /*?lang= Disallow: /*&lang= # Sitemap location Sitemap: https://iampro.io/sitemap.xml # Crawl-delay for politeness (optional, respected by some bots) # Note: Removed global crawl-delay as it can slow Bing indexing # Crawl-delay: 1 # --------------------------------------------------------------------------- # MODEL-TRAINING crawlers — scoped to the pages that say what iampro.io IS. # # These collect a training corpus; by their operators' own documentation they do # not feed the live indexes that cite sources, so they cannot return a visitor. # Measured: over 14 days they took ~191 000 requests and ~11.8 GB, against # ~12 200 requests and ~0.6 GB for every crawler that CAN cite. Over Plausible's # full history (2026-05-27 → 07-29, 2 692 visitors) AI assistants returned 18 # visitors via chatgpt.com, 3 via Perplexity, 1 via Copilot — none ever from # claude.ai. Googlebot returned 1 056. # # The first instinct was `Disallow: /`. The logs argued for something better. # GPTBot + ClaudeBot pulled 113 412 DISTINCT urls; 48 890 were /o/* and 48 750 # /jobs/* — individual job postings that expire within days, many already # answering 410. Perishable listings are worth nothing to a training corpus and # every pipeline already knows it: near-duplicate removal collapses thousands of # templated career pages to a handful, and quality filters drop pages that are # mostly boilerplate. Those 179 259 requests bought roughly the corpus footprint # that the ~15 urls below would have bought — while causing all the contention. # # So the door stays open on the identity surface and closes on the churn. # Congestion drops to ~nothing and corpus presence gets BETTER: what survives # dedup is prose describing the product, not listings that are stale before the # model ships. Enforced in nginx/bot-throttle.conf for crawlers that ignore this. # # The cost was never the hosting bill — ~16 GB/month of GPTBot is under two # cents at Hetzner's overage rate and we sit far inside the included allowance. # It is contention: 3 vCPU, 3 uvicorn workers. On 2026-07-28 this traffic # pushed Googlebot to a 15 % error rate (176x 499 + 27x 504) mid-recovery from # the June collapse. # # Google-Extended and Applebot-Extended are NOT crawlers — they are policy # tokens opting content out of Gemini / Apple Intelligence training, so they # take no scope and are left out of this group entirely. # --------------------------------------------------------------------------- User-agent: GPTBot User-agent: ClaudeBot User-agent: Claude-Web User-agent: Anthropic-AI User-agent: CCBot User-agent: meta-externalagent User-agent: Bytespider Disallow: / Allow: /$ Allow: /llms.txt Allow: /about Allow: /reach Allow: /connector Allow: /tools Allow: /jobs$ Allow: /careers$ Allow: /places$ # --------------------------------------------------------------------------- # CITATION crawlers — welcomed, on a budget. # # These feed the live indexes that answer with a source link, so they are the # ones that can actually return a visitor. Kept open, but capped: an open # invitation with no ceiling is precisely what broke on 2026-07-28. # Crawl-delay is advisory and only some crawlers honour it — the enforced # ceiling lives in nginx/bot-throttle.conf. This states the intent; that file # makes it true. # # One group, many User-agent lines: identical rules, and no chance of a new # crawler being added with the Disallows forgotten. # # OAI-SearchBot powers ChatGPT Search's live index (distinct from GPTBot, which # is training-only and blocked above). Claude-SearchBot is its Anthropic # counterpart — it had visited exactly 14 times in 14 days, which is the whole # point: the crawlers that cite are tiny, the ones that train are enormous. # --------------------------------------------------------------------------- User-agent: OAI-SearchBot User-agent: Claude-SearchBot User-agent: PerplexityBot User-agent: Cohere-AI User-agent: Amazonbot User-agent: PetalBot Allow: / Allow: /llms.txt Disallow: /dashboard Disallow: /admin Disallow: /api/ Disallow: /verify/ Disallow: /reset-password Disallow: /logout Disallow: /health # Language variants are the same page under a query string, and the canonical # always points at the bare URL. Crawling them multiplied the surface by five # for nothing: 3 777 of ClaudeBot's 13 786 requests on 2026-07-29 carried a # lang= parameter. Left open for Googlebot and bingbot, which handle # canonicals correctly and where hreflang discovery is worth the budget. Disallow: /*?lang= Disallow: /*&lang= # 6 seconds = 10 requests/minute = EXACTLY the nginx ceiling in # bot-throttle.conf. Deliberately matched: Anthropic documents honouring this # directive (support.claude.com article 8896518), so a compliant crawler # self-limits to its budget and never sees a 429 at all. The nginx limit stays # as the backstop for crawlers that ignore the hint — which is most of them. Crawl-delay: 6 # --------------------------------------------------------------------------- # User-triggered assistant fetches — no crawl-delay, no query-string limits. # These fire when a person asks an assistant about iampro.io and it goes to read # the page: one human, waiting, right now. Slowing them down is slowing down a # prospect, and this is exactly the traffic the MCP connector exists to attract. # Kept out of the crawler group above on purpose, and out of the nginx rate # limit too (nginx/bot-throttle.conf). Only the private areas stay closed. # --------------------------------------------------------------------------- User-agent: ChatGPT-User User-agent: Claude-User User-agent: Perplexity-User Allow: / Allow: /llms.txt Disallow: /dashboard Disallow: /admin Disallow: /api/ Disallow: /verify/ Disallow: /reset-password Disallow: /logout Disallow: /health # --------------------------------------------------------------------------- # SEO-intelligence scrapers and audit scanners. No traffic ever comes back from # these; they resell our index to competitors. On 2026-07-29 SemrushBot (5 322) # and AhrefsBot (2 909) alone outweighed Googlebot 42-to-1. Disallowed here for # the ones that comply, dropped at the edge in nginx/bot-throttle.conf for the # ones that do not. # --------------------------------------------------------------------------- User-agent: SemrushBot User-agent: AhrefsBot User-agent: MJ12bot User-agent: DotBot User-agent: DataForSeoBot User-agent: BLEXBot User-agent: Barkrowler User-agent: SeekportBot User-agent: ImagesiftBot Disallow: /