Skip to documentation
Browse 13 guides

User guide

AI crawler tracking

AI crawlers commonly request server-rendered pages without running browser JavaScript. Tailwin therefore records them at the edge or origin before a page response is sent.

Updated September 28, 2026

On this page
  1. Choose one installation
  2. Information collected
  3. Install the Cloudflare Worker
  4. Install on WordPress
  5. Install Node middleware
  6. Read the dashboard
  7. Troubleshooting
  8. Site content inventory

AI crawlers commonly request server-rendered pages without running browser JavaScript. Tailwin therefore records them at the edge or origin before a page response is sent.

Choose one installation

Site architectureRecommended installation
Cloudflare-proxied siteCloudflare Worker tracker
WordPressTailwin WordPress plugin in integrations/wp-plugin
Next.js or Expressintegrations/middleware-js server middleware
Another reverse proxy or serverAdapt the middleware contract or place the Cloudflare Worker in front

Do not install two trackers on the same request path unless you have deliberately implemented deduplication.

Information collected

The tracker sends a bounded event containing the site identity implied by its ingest key, matched bot identity, request time, path, and verification context. HTTP status and country are included only when the selected integration can observe them. The raw user agent is sent only for an unmatched, bot-shaped request and is capped at 256 characters. It does not send request bodies. Unknown bot-shaped user agents are aggregated for registry review rather than automatically trusted.

The ingest request uses a random twk_ bearer key. Tailwin stores only a peppered hash of that key. The endpoint validates batches of 1 to 200 hits, queues them when Redis is available, and processes inline when the queue is unavailable. Known hits and daily rollups are written by the same processing path.

Delivery is best effort. Integrations clear a batch before sending it and drop a failed delivery instead of retrying, because an ambiguous retry could duplicate a hit. A tracker placed at the origin also cannot observe responses served entirely by an upstream cache.

Install the Cloudflare Worker

  1. Open the site’s Tracker page in Tailwin.
  2. Create or copy the site ingest credential shown by the guided flow.
  3. Deploy the Worker to the customer’s zone with the twk_ site key stored as a Worker secret. Override the Tailwin API base only for a self-hosted deployment.
  4. Route the Worker over the public site hostname.
  5. Request a page with the verification method shown in the UI.
  6. Return to Tailwin and confirm that the live verification received a hit.

Cloudflare’s verified-bot signal strengthens classification when the customer’s plan and request context expose it. A registry match identifies the claimed crawler; it does not by itself prove that the caller owns the user agent.

Install on WordPress

  1. Package or copy integrations/wp-plugin into the WordPress plugins directory.
  2. Activate the plugin.
  3. Enter the twk_ site key in its settings. Change the Tailwin address only for a self-hosted deployment.
  4. Save and use its verification action.
  5. Confirm receipt on the Tailwin tracker page.

The plugin runs server-side. A cache or CDN in front of WordPress can prevent origin requests from reaching it; use the edge tracker instead when the origin does not see most traffic.

The optional WordPress verification setting uses documented, forward-confirmed reverse DNS for Googlebot, Bingbot, and Applebot. AI crawler matches remain unverified because this implementation does not currently validate their published IP ranges.

Install Node middleware

  1. Add the local package from integrations/middleware-js to the site project.
  2. Configure the twk_ site key and, only when self-hosting, the Tailwin API base in server-only environment variables.
  3. Register the Next.js or Express adapter before the final response handler.
  4. Deploy and make a controlled verification request.
  5. Confirm the event in Tailwin.

Never expose the bearer ingest key in a client bundle.

The Node middleware does not perform IP or reverse-DNS verification. Its verified value therefore remains false unless a future adapter supplies a trusted platform signal.

Read the dashboard

  • Fetcher/search/agent traffic can affect live answer retrieval.
  • Training traffic can be relevant to content-rights policy but is not automatically a live-answer visit.
  • Top pages show which URLs bots requested.
  • Unknown crawlers are untrusted leads until promoted by an administrator with evidence.
  • The robots audit compares the site’s policy with the visibility goal and separates live-answer access from training access.

Troubleshooting

  1. Confirm the event happens at the edge or server, not in browser JavaScript.
  2. Confirm the selected organization and site ID match.
  3. Confirm the endpoint is public and the twk_ bearer key has no copied whitespace.
  4. Check that a CDN is not bypassing origin middleware.
  5. Check Tailwin worker health; the HTTP endpoint can accept while asynchronous processing is delayed.
  6. Review unknown-crawler telemetry if the user agent is new.

Do not weaken verification just to make a test event appear as a known bot.

Site content inventory

On a site page, use Site content inventory to check your public pages. Confirm that you own the site or have permission, then choose Crawl site. The worker checks up to 25 pages, using the homepage, same-host sitemaps and page links. Status refreshes while the crawl is queued or running.

The crawler respects TailwinCorpus robots rules, waits at least one second between requests, and excludes query-string URLs, credentials, private network addresses and other hosts. It uses public HTTPS sites with IPv4 DNS. A redirects-only robots file, unavailable robots policy, oversized response or unsupported host can prevent a crawl. Repair that condition and retry. If a job stays queued for ten minutes, check worker availability and start another crawl.

Snapshots show HTTP status, title, content changes, technical findings and a deterministic citability assessment. Unavailable HTML has an unmeasured score. A first observation has no change baseline. A page limit or fetch failure produces a partial result; pages outside the crawl are not declared missing or deleted.

Keep up to ten snapshots per site and choose a snapshot to inspect older findings. Delete crawl history erases all stored snapshots and prevents an active job from saving them again. An already-started request may finish before cancellation is observed. The inventory stores hashes and findings, not complete page HTML. It measures page content, while crawler tracking measures bot visits and Rankings measures search evidence.