Log file analysis for SEO and AI crawlers

  • AI crawlers and Googlebot in detail
  • Telling genuine bots from fake ones
  • Finding errors that crawls and web analytics miss
Request a log file analysis
4.9
Based on 33 reviews
PremierBadge
Microsoft Partner Badge
SISTRIX certified agency

Making crawl behaviour visible

Which pages does Googlebot fetch frequently and which never? Which content do ChatGPT, Claude and Perplexity read? The log files show it for every URL.

Separating genuine and fake bots

Many requests claim to be Googlebot or an AI crawler but come from entirely different servers. We check every request against the providers’ official IP lists, complemented by reverse DNS checks.

Finding errors that otherwise go unnoticed

URLs invented by AI systems, crawlers stuck in endless loops, huge files nobody sees: problems like these rarely show up in crawls or web analytics.

Bot verification

Matching every request against the official IP lists of Google, Microsoft, OpenAI, Perplexity and other providers, complemented by reverse DNS checks.

AI crawlers and user-triggered agents

Separate analysis of training crawlers, search crawlers and agents that fetch pages live for a user’s question, including patterns across the day and the week.

Googlebot and crawl budget

Which pages and resources use up the crawl budget? How often are important pages fetched? Where does Googlebot end up in redirect chains?

404 errors and invented URLs

Which addresses do crawlers request that do not exist? AI systems invent URLs. We provide a redirect list that catches these visitors.

Data volume and outliers

Which files cause the most traffic? Crawlers in loops or oversized images become obvious immediately.

Comparison with analytics

We compare the requests from AI agents with the visits from ChatGPT and others in GA4. This gives a realistic picture of AI demand.

1

Clarifying format and scope

Which servers, which fields, which period? Seven to fourteen days are usually enough. We also clarify the data protection basis, as log files contain IP addresses.

2

Taking over and checking the data

We check the data for gaps and special features, such as missing time windows or load balancers that only pass on the real IP address in an additional field.

3

Verifying bots

Every request is assigned to the providers’ official IP ranges. We report requests with a false sender separately.

4

Analysing

Crawl behaviour by bot, directory, status code and file type, outliers, errors and data volume.

5

Results meeting

We present the findings with direct links to the pages and files concerned.

6

Implementing actions

Redirects, robots.txt, caching, image sizes and content. On request, we repeat the analysis to check the impact.



 What do AI systems read on your website? 

Your log files know the answer.

We show you which pages Googlebot, ChatGPT, Claude and others fetch and where things go wrong.


Request a log file analysis
 Insights from practice 

What a log file analysis brings to light

The following findings come from an analysis of around 12 million log lines over ten days (29 August to 7 September 2026) for a large, high-traffic website (anonymised).

  • ChatGPT is reading along: we counted almost 29,000 requests from ChatGPT-User, an average of almost 2,900 a day (as an upper limit, see methodology below), with a clear daily rhythm. All 100 pages with the most Google clicks were among them. For roughly every 35 requests there was one page view referred by ChatGPT, a ratio consistent with the figures in GA4.
  • Invented addresses: 15% of Claude’s requests ended with a 404 error because the model had invented URLs. For ChatGPT, the figure was 0.1%. We created a redirect list for the most frequent invented addresses.
  • Crawler in a loop: a crawler run by a large platform operator made around 370,000 requests, three quarters of them to just nine URLs. Among them was a 32 MB image hidden in an invisible element on the home page. Browsers never loaded it, but crawlers did so around 8,800 times. That added up to about 280 GB of data traffic in ten days.
  • Fake bots: 23 network ranges of a cloud provider appeared with at least three different AI bot names. All requests claiming to be Perplexity-User were fake.
  • llms.txt: the website had no llms.txt file, and none of the major AI providers even requested one during the period analysed.
 Methodology 

Why standard tools often get log files wrong

Log file tools quickly produce nice charts. The figures are only correct, however, if the data is read correctly. Two typical pitfalls:

  • Load balancers and CDNs: if there is a load balancer in front of the web server, the standard column often shows its IP address rather than the crawler’s. The real address is then in an additional field. If you overlook this, you cannot verify a single bot correctly.
  • User-triggered agents are not users: requests from ChatGPT-User and similar agents happen when someone asks a question, but also because of monitoring tools that query prompts automatically. We therefore report these figures as an upper limit.

We also check whether the data is complete. In the analysis described above, for example, individual hours were missing on six days because the export was limited. We point out gaps like these rather than keeping quiet about them.

 Log file analysis with Rheinwunder 

Let's take a look at your log files

Send us a short extract and we will tell you which questions it can answer.

Microsoft Partner Badge
SISTRIX certified agency
4.9
Based on 33 reviews
A man in a white shirt leans against a wall next to the Rheinwunder logo
Founder and Managing Director Ralph Grundmann 0228 243 313 53




Information on how we process your data can be found in our Privacy Policy.