Your crawl log is your content calendar.
Every day, AI engines read your content and decide whether it belongs in their answers. Viz shows you that entire relationship, turns the gaps into topics, writes and reviews the drafts with your team as editor-in-chief, and tracks every published piece until the citations arrive. Track, publish, track again - the whole cycle.
What the machines' reading habits tell you.
What AI reads is a verdict
Retrieval bots fetch the pages they consider worth quoting. A crawl map of your content is an unsolicited, machine-written audit of what you are considered an authority on.
Citation is not scraping
An AI search bot fetching your comparison page to answer a live question can cite you and send a reader. A training crawler reading the same page feeds a model that never will. Same page, same moment, entirely different value exchange.
Your best work leaks first
The content AI engines pull hardest is exactly the content that took the most effort: original research, detailed guides, honest comparisons. That is an argument for measuring before blocking.
Four habits for content teams in the AI era.
Find your AI-favorite pages
Filter the request log to AI search bots and sort by URL. The pages they fetch most are the ones feeding answers right now - refresh those first, and write more where the crawls cluster.
Watch the payback channel
AI-referred visits appear next to crawl volume, and the sign-ups, purchases and leads those visitors produce are credited to the assistant that sent them, on your server. Referrals from assistants often land on deep pages, not your homepage - judge them by what they go on to do, not raw count.
Catch the silent republishing
A scraper walking your archive at machine cadence is tomorrow's duplicate content. The request log makes that pattern visible while it happens, not months later in a search result.
Gate with intention
If part of your library is your product - courses, research, premium guides - block AI training crawlers from it and verify the block holds, while leaving your marketing content open to be found.
Open, gated, or metered - decide per section, not per site.
The sites navigating this best treat AI access like distribution strategy, not an on-off switch. Marketing content stays open because answers are where discovery happens now. Premium content gets gated because it is the product. The middle - the guides and research that build your reputation - is where measurement earns its keep: if AI engines quote it and send readers, it is working; if it is only ever ingested by training crawlers, there is a licensing conversation to have, and AI Control gives you the enforcement to back it up.
Your crawl map writes the brief. The pipeline writes the draft.
Everything above tells you what to write: the questions retrieval bots keep asking, the gaps in your coverage, the queries you are starting to rise for. The AI Content Pipeline closes that loop - it proposes topics from that same evidence, writes and adversarially reviews the draft, and delivers it into Posts → Drafts with a hero image set. Your team stays editor-in-chief: you approve every topic, you edit every draft, and nothing publishes until you do. Then every published piece reports back - crawled, cited, visited, converted - and the researcher folds those results into its next briefs.
Being in the training set is not being citable.
A model that trained on your content may know things you taught it, but it will rarely say your name - training dissolves your work into weights, uncredited by construction. Citation runs on a different machine entirely: a retrieval bot fetches your live page at question time, and the engine links it under the answer. Training influence is diffuse and unattributable; citation is specific, creditable, and clickable. Only one of them can ever send you a reader.
Content that wins citations is current, specific, and structured to answer a question cleanly - the engine needs a reason to point at you rather than paraphrase you. And the two games have separate scoreboards on your own server: training crawlers reading your archive play the first, retrieval bots re-fetching your live pages play the second. Watching which sections attract which kind of reading tells you whether your work is becoming model knowledge, answer material, or both - and lets you set policy for each with AI Control instead of one switch for everything.
Content team FAQ.
How do I know which of my pages AI engines use most?
Filter server-side traffic to the AI search and AI training categories and look at URL frequency. Pages that retrieval bots fetch repeatedly are being used to answer live questions; pages only training crawlers touch are feeding models. Viz shows both per URL in the request log.
Are AI referrals actually worth anything compared to search traffic?
Often, per visitor, more than average: the assistant handled the shallow part of the question, so the human who clicks through tends to want depth. But volumes are smaller than classic search and some AI surfaces strip referrers, so measure conversions per assistant (Viz credits them on the server) and, with GA4, engagement per session, rather than celebrating or dismissing the raw count.
Should a content site block AI crawlers entirely?
Blanket blocking is rarely the right first move. Blocking retrieval bots removes you from AI answers - including the ones that cite and refer readers. A more surgical pattern: measure for a few weeks, keep marketing content open, gate the content people should pay for, and make training-crawler decisions separately from search visibility.
Can Viz tell me what AI engines say about my brand?
Paid plans include AI Visibility, which asks the major engines the questions your audience asks and records whether you are mentioned, cited, and how you compare to alternatives. Crawl data shows what machines read; AI Visibility shows what they say afterwards.
Can Viz write the content too?
On paid plans, yes - with your team holding the pen where it matters. The AI Content Pipeline proposes topics from your own AI-traffic evidence, and a draft is only written after you approve one. Adversarial reviewers score every draft before delivery, and a piece below the quality bar is refused rather than delivered (and never billed). Drafts land in Posts → Drafts at 1,200 to 1,800 words with a hero image set; you edit and publish like any post, and every published piece is tracked for AI crawls, citations, and referrals.
Does refreshing old posts matter for AI answers, or only new content?
It matters, and the crawl log shows you where. Retrieval bots re-fetch pages they consider live sources; when one of your older guides keeps drawing retrieval traffic, it is actively feeding answers today, and updating it improves what gets quoted almost immediately. A stale page still being fetched is the highest-payoff refresh on your list - the demand is proven and the fix is editing, not creating.