17 Σεπτεμβρίου 2026

    AI crawlers are not one thing: how to make better decisions about your content

    AI crawlers do more than one job. Learn how to distinguish search, training and agents, then use brand-representation data to make better access decisions.

    Control panel with a central mode selector and illuminated switches, illustrating separate access choices for AI crawlers
    photo by iSawRed | Unsplash
    Κοινοποίηση:LinkedInX (Twitter)FacebookWhatsAppΑνάλυση με AI:ChatGPTClaudePerplexity

    One line in robots.txt can now affect three very different things: model training, search discovery, and an agent visiting a page on a user’s behalf. Treating them as a single decision about “AI” creates two opposing risks. You may block a crawler to protect your content and, depending on the crawler and configuration, limit its discoverability at the same time. Or you may leave access open without knowing whether your own site is actually shaping the answers customers see, while models build their picture of your brand from third-party sources.

    Cloudflare’s latest changes show why the simple question of “allow or block?” is no longer enough. Site owners increasingly need to decide separately who may access their content, what they may use it for, and how to assess whether the result is an accurate and useful representation of the brand in AI-generated answers.

    One crawler can serve more than one purpose

    Cloudflare has introduced a Disallow AI Training setting, intended to let site owners refuse the use of their content for AI training without automatically blocking access for search. It addresses what Cloudflare calls mixed-use crawlers: bots that may support both search and other generative uses. Cloudflare explains the change here.

    The distinction matters. Cloudflare separates crawler activity into three broad categories:

    • search – crawling to build or maintain a search index

    • training – crawling to train or fine-tune a model

    • agent – a user-directed agent visiting a page on behalf of a person

    Cloudflare’s own Accountable designation applies to crawler operators that provide, or have committed to provide, greater transparency and more meaningful publisher controls. It is an important development, but it is not yet a universal standard for the whole web.

    The same need for distinction appears in the documentation of AI providers themselves. OpenAI separates OAI-SearchBot, used to surface sites in ChatGPT search features, from GPTBot, which may crawl content for use in developing foundation models. It also distinguishes ChatGPT-User, which visits pages in response to user-initiated actions. OpenAI’s crawler documentation explains these roles separately.

    Google provides another reminder that “AI” is not one system. Googlebot controls access for Google Search, including AI Overviews and AI Mode. Google-Extended is used by Gemini apps and the Vertex AI API for Gemini. Google’s crawler documentation and its guide to AI features in Search make that distinction explicit.

    This is not only a technical setting

    Allowing or blocking a crawler may look like an infrastructure decision. In practice, it is also a decision about content distribution, source control and the business model of a site.

    A publisher that depends on advertising may assess user-directed agents and answer summaries differently from a B2B brand that wants its expert content to remain available for search and research-oriented answers, while refusing its use for model training. An e-commerce business must additionally consider whether essential product information is being understood and presented accurately when AI helps a customer compare options or make a purchase decision.

    There is no single configuration that makes sense for every domain. There is, however, a strong case for making the decision with evidence rather than reacting only to a general distrust of AI or a fear of losing visibility.

    Three questions to ask before changing crawler policy

    It helps to separate three layers that are often treated as one.

    The first two questions concern access controls and stated preferences. The third concerns the outcome seen by customers: whether a brand appears in an answer, how it is described, whether it is recommended, and whether the system relies on its official sources.

    This is where Semantio can add a layer that technical logs alone cannot provide.

    How Semantio data can support crawler decisions

    Semantio does not replace server logs and does not identify traffic from individual crawlers. It cannot tell you whether a particular bot visited a particular page at a particular time.

    It can, however, provide evidence for assessing whether a content-access policy is aligned with the brand’s objectives.

    A useful starting point is a baseline study. Rather than asking a model a generic question such as “what do you know about our company?”, the research should test scenarios that reflect real customer decisions: selecting a supplier, comparing products, checking a specification, assessing risk, seeking a recommendation or asking a post-purchase question.

    Across such scenarios, Semantio can help identify:

    • whether the brand appears in answers relevant to its category and offer

    • whether it is merely mentioned or genuinely recommended

    • whether models describe its products, capabilities, reach and limitations accurately

    • whether answers rely on the official brand site or primarily on third-party sources

    • whether competitors dominate particular decision situations

    • how results differ across models, languages, markets and repeated measurements

    This matters when a company is considering restricting access across part of the AI ecosystem. If the official domain already appears as an important source for the accurate representation of the brand, that should inform the decision. If models largely ignore the company website and instead construct their answers from outdated directories, forums or third-party articles, the problem may not be crawler access itself. It may be the clarity, freshness and source authority of the brand’s own content.

    For a closer look at the distinction between mentions, recommendations, sources and the quality of representation, see How marketing teams can measure brand visibility and representation in AI. The question of when an AI answer draws on an official brand site rather than an unverified third-party source is explored in Understanding Sources of Truth.

    Measure before and after, but do not pretend certainty

    A change to robots.txt, WAF rules or a Cloudflare configuration should not be treated as an experiment with an immediate and unambiguous result.

    Generative answers vary. They depend on the model, date, location, phrasing of the question, available web sources and the internal systems used by each platform. In many cases, it is not possible to attribute a specific answer to one crawl or one URL.

    A more reliable approach is to:

    • define the scenarios, models, languages and markets that matter to the brand

    • establish a baseline before changing the configuration

    • retain full responses, cited sources, dates and relevant research metadata

    • make the technical change deliberately and document its scope

    • repeat the study with a comparable research design

    • interpret the results alongside crawl logs, CDN data and other changes to the site or wider brand environment

    This does not automatically establish causality. It can, however, reveal signals that warrant closer investigation: a decline in citations of the official site, a growing share of inaccurate claims, a change in the sources used by models, or weaker recommendations in commercially important scenarios.

    Platform data and brand-representation data answer different questions

    This distinction is becoming more practical as platforms begin to expose their own data. In February 2026, Bing Webmaster Tools introduced a public preview of AI Performance, covering Microsoft Copilot, AI-generated summaries in Bing and selected partner integrations. It reports citation activity, cited URLs and sampled grounding queries over time. Microsoft’s announcement is particularly useful because it shows what platform-level transparency may look like in practice.

    These data can help a publisher see which of its URLs Bing reports as being cited across supported AI experiences. They do not, however, tell the whole story. A citation count does not show whether the brand was presented accurately, whether it was recommended over a competitor, or whether the cited page influenced a meaningful customer decision.

    That is where the two layers can complement one another. Bing Webmaster Tools can provide operator-side evidence about citation activity within Microsoft’s ecosystem. Semantio can test how the brand is represented in defined decision scenarios across the models included in a study. Server and CDN logs can then provide the technical evidence about crawler access.

    These are not competing datasets. They describe different parts of the same problem.

    What this does not mean

    Robots.txt does not, by itself, solve the question of control over content. It is primarily a statement of preference, and its practical effect depends on whether and how a specific operator respects it.

    Blocking model training does not automatically mean that a brand will disappear from AI-generated answers. The brand may still be known from other sources, appear through search systems, or be retrieved from the live web for a particular response.

    Nor should a brand treat mere presence in an answer as success. It may be mentioned without being recommended, described inaccurately, or presented through an outdated third-party source.

    The most useful question is therefore not: “Should we allow AI to access our website?”

    It is: Which uses of our content are compatible with the brand’s business model, how can we control them technically, and what do the data tell us about the effect of those decisions in the answers customers actually see?

    Κοινοποίηση:LinkedInX (Twitter)FacebookWhatsAppΑνάλυση με AI:ChatGPTClaudePerplexity

    Michał Grzebyk
    Michał Grzebyk
    COO Brand Semantics

    Συνιδρυτής της Brand Semantics, με πολυετή εμπειρία στο μάρκετινγκ από το 2009. Διακεκριμένος εκπαιδευτής και στρατηγικός σύμβουλος, πρωτοπορεί στην εξερεύνηση νέων οριζόντων, μετατρέποντας τη διεπιστημονική γνώση σε καινοτόμες επιχειρηματικές λύσεις.