Measure AI Crawlers Before You Block Them
TL;DR: AI crawlers are hitting sites at ratios as high as 70,900 requests per referred visitor, creating real infrastructure costs — but blocking them blindly risks competitive invisibility in LLM-driven discovery. Operators in high-CAC verticals need a structured audit process: identify which bots are visiting, calculate their cost, measure the visibility they generate, and assign each one a deliberate keep/restrict/block decision before touching robots.txt or WAF rules.
Why the Old Robots.txt Playbook No Longer Works
For years, SEOs controlled bot access with a simple robots.txt entry. That era is ending. OpenAI updated its documentation to clarify that ChatGPT-User, its user-triggered fetcher, no longer commits to honoring robots.txt. Perplexity-User behaves the same way. This means robots.txt still blocks compliant training crawlers and search-indexing bots, but it does nothing against the user-triggered fetches that represent genuine, in-funnel interest in your content.
For operators running paid acquisition at scale, this distinction matters immediately. User-triggered fetches happen when a prospect is already researching your brand or product — they are mid-funnel signals. Blocking them via robots.txt does nothing. Blocking them at the WAF layer (Cloudflare, AWS, or server-level user-agent rules) does. If you have not had a conversation with your infrastructure team about which AI bots your WAF currently allows, that conversation is overdue.
Before making any blocking decisions, understand the taxonomy. There are three distinct bot types: training crawlers that feed LLM knowledge bases without a direct referral loop, search-indexing bots that surface and cite pages in AI-generated answers, and user-triggered fetches that pull pages on demand when a user asks about your specific brand or content. Each carries a different risk/reward profile and warrants a different default response.
Pull the Log Files Before You Touch Anything
The biggest operational mistake operators make is adjusting bot access without knowing what is actually crawling their site. Log files are the authoritative source. A 30-day sample will show every bot that hit your server, the frequency of requests, and which pages were targeted. Traditional log analyzers can separate search engine bots from AI crawlers. Dedicated AI visibility tracking tools go further, tagging individual LLM agents and flagging crawl concentration on specific page types.
Pay attention to page-level patterns. A training crawler hammering your proprietary research pages is a different problem from a search-indexing bot reviewing your service landing pages. One is extracting competitive IP; the other is potentially building your citation footprint. Treating them identically is how operators make expensive mistakes in both directions.
If log access is restricted, referral traffic in Google Analytics gives a partial picture. GA’s new “AI Assistant” channel classification captures ChatGPT, Gemini, and Claude referrals via referrer header, but misses Perplexity entirely. Use it as a directional signal, not a complete inventory. The gaps in referral data are why a proper channel performance audit is the right starting point for operators who have never systematically reviewed their bot traffic.
Calculate the Actual Cost, Then the Actual Value
Cloudflare data from June 2025 put Anthropic’s Claude at a crawl-to-refer ratio of 70,900:1 — meaning Claude makes 70,900 page requests for every single visit it sends to a website. Google’s equivalent ratio is 9.4:1. That is not a rounding difference; it is a fundamentally different infrastructure burden. For large sites with aggressive AI crawler activity, this cost is frequently absorbed into general hosting fees until someone on the infrastructure side flags a bandwidth spike.
Get a number from your server team. What does it cost per month to serve the AI crawlers currently accessing your site? That figure becomes the denominator in your value calculation. On the revenue side, look at engagement metrics for LLM-referred visitors, not just session volume. AI-referred traffic is often smaller in raw numbers but highly qualified — users who arrived via an LLM answer have already passed through a filtering layer. Compare their on-site behavior against the behavior of converting users from other paid channels. If the conversion pattern aligns, the referral value is real even if the volume is modest today.
Citations and brand mentions are a second value dimension that does not show up in referral traffic at all. Track how often your site appears in AI-generated answers for your core commercial topics. For operators in regulated or high-competition verticals — iGaming acquisition, forex broker lead generation, or law firm client acquisition — brand mention frequency in LLM answers is becoming a leading indicator of organic pipeline health, especially as younger demographics shift their discovery behavior toward AI-assisted search.
What This Means for High-CAC Vertical Operators
In verticals where cost per acquisition runs $200 to $2,000 per converted lead, the stakes of AI crawler decisions are amplified. Every citation missed is a prospect who finds a competitor instead. Every improperly trained LLM answer that misrepresents your product is a brand problem that no retargeting budget can easily fix.
For crypto exchange and token operators, proprietary content — market analysis, yield data, trading logic — is often the primary competitive differentiator. Training crawlers ingesting that content without attribution or referral represent a direct IP cost. The decision to block training crawlers while allowing search-indexing bots is defensible and often correct in this vertical. Make that distinction explicit in your WAF configuration rather than applying a blanket rule.
For CDL fleet recruitment operators, the dynamics are different. Driver-facing content is rarely proprietary IP, and LLM citation in location-specific job search queries is an emerging acquisition channel worth cultivating. Blocking AI crawlers in recruitment could quietly remove your job listings from AI-assisted job discovery at exactly the moment when that channel is gaining traction.
Across all high-CAC verticals, the core principle holds: the competitive disadvantage of invisibility compounds over time. A crawler that sends minimal traffic today may belong to the dominant discovery platform in 18 months. Locking it out now to save server cost means re-entering a knowledge base that has already formed its citations around your competitors. The operators running structured paid and organic performance programs are the ones building measurement frameworks now, before the decision becomes urgent.
Build a Decision Matrix, Not a Policy
The practical output of this process is a bot-level decision matrix with three outcomes: keep, restrict, or block. For each AI crawler currently hitting your site, answer five questions. Does it generate converting revenue or measurable brand visibility? Is it accessing proprietary content or product data that constitutes your IP? Does the company behind it publish documentation on crawler behavior, data retention, and compliance with site directives? Is it creating infrastructure cost or server strain that affects real users? And what is the competitive cost of removing your site from its answers?
Bots that clear the value threshold, stay within trusted documentation standards, and do not hammer proprietary pages should default to keep. Bots where the business case is unclear — positive signals but no hard revenue attribution — should be restricted rather than blocked. Rate-limit their crawl access or restrict them to non-sensitive page categories while you continue gathering data. Bots that generate significant cost, access IP you cannot afford to externalize, or come from opaque operators with no published compliance documentation should be blocked at the WAF or server level.
Review the matrix quarterly. AI bot behavior changes. Platforms that respected robots.txt last quarter may update their documentation next quarter. New crawlers will appear. A crawler you blocked six months ago may have shipped a referral traffic integration that changes its value calculus entirely. Precision traffic analysis and regular log review are operational habits, not one-time projects. Build the cadence into your team’s workflow before a server invoice or a competitor citation forces the conversation.
Operators who want to audit their current AI crawler exposure and quantify what it is costing versus generating should start with a structured AI traffic and visibility review before making any permanent blocking decisions. A systematic data pull is faster and cheaper than reversing a bad block months later when you realize you have gone dark inside a platform your prospects now use daily.
Originally reported by Search Engine Journal, July 2026.
Get a playbook for your vertical
Forex lead gen
FTD acquisition, depositor funnels, regulated broker campaigns across Tier 1 & Tier 2 GEOs.
Explore → CryptoCrypto & Web3
Token launches, exchange user acquisition, DeFi protocol growth. Compliant campaigns only.
Explore → LegalLaw firm marketing
Mass tort, personal injury, immigration. High-intent lead gen for US law firms with $50K+/mo budgets.
Explore →