Aethelforge
Read the Index

Research · The channels

Twenty-eight times baseline: the crawler surge that followed Muse

The day after Muse shipped, Meta's indexing crawler went from 725 hits a day across our monitored sites to a peak of 20,177. What the surge was, what it was not, and what it cost.

Published 24 Sep 20263 min readAethelforge editors

In one paragraph

What happened to web crawler traffic after Muse launched?

Across 25 monitored sites, Meta-ExternalAgent rose from an average of 725 hits a day to a peak of 20,177 on 16 September, nine days after the 8 September launch. The nine-day surge totalled 104,058 hits, 44 percent of all bot traffic in the window.

What does a product launch look like from the other side of the web? For the businesses whose sites we monitor, Muse looked like this: nothing on the day, and then, the morning after, a crawler that had been visiting a few hundred times a day began visiting thousands of times a day, and kept climbing for a week.

The numbers

Across 25 monitored sites, Meta-ExternalAgent averaged 725 hits a day between 24 August and 8 September, the day Muse launched. On 9 September it recorded 3,125. On 10 September, 6,355. On 11 September, 13,860. After a weekend dip it climbed again, to 18,234 on 15 September and a peak of 20,177 on 16 September, roughly twenty-eight times the pre-launch average. On 17 September it was falling, and on 18 September it was back to 437. Over the nine days from 9 to 17 September the crawler made 104,058 requests, which is 44 percent of every bot hit the fleet saw in the entire 31-day window of the first Index edition.

The chart on the edition page draws it day by day. The whole shape of the month's bot traffic is that one crawler's surge; the AI category as a whole moved with it, and everything else stayed flat.

What it was

Meta documents Meta-ExternalAgent as the crawler it uses to index web content and gather training data for its AI products. It reads pages. It does not call, chat, book or submit anything. In the vocabulary of the Index it is the reading layer, the same layer as a search engine's crawler or an assistant fetching a page because someone asked a question.

The distinction matters because the cost profile is different. A reading-layer surge is measured in bandwidth, cache misses and, for the smallest sites, a hosting bill. It never rings a phone. An acting-layer surge, an assistant calling five clinics for one patient, is measured in staff minutes. One is a hosting question and the other is a staffing question, and a count that mixes them tells a business nothing it can act on.

What it was not

Two things the surge was not are worth stating plainly. It was not Muse itself browsing. Meta publishes a separate user agent for on-demand fetches, the kind an assistant makes when a person asks it to look at a page right now, and that pattern recorded no hits on any monitored site in the window. This is consistent with Amazon's public complaint that Muse browses without identifying itself; whatever Muse fetched, it did not fetch as Meta's fetcher.

And it was not, on the evidence, caused by the launch. The timing is as tight as timing gets: baseline until the launch day, a sharp rise the day after, a peak eight days later, baseline again two days after that. Meta has not said what the crawler was doing. The edition records the correlation as strong and the cause as plausible, and it stops there, because a number that is published and signed should not carry a guess.

What the sites did

Nothing, which is the right answer for a reading-layer event. Every monitored site allows Meta's crawler in its robots file, along with the other AI crawlers and fetchers a business should expect, and none of them changed that during the surge. The crawler that read the site in September is the same one that will let an assistant describe the business accurately in October, and being described accurately is the beginning of every inquiry that turns into a customer.

Whether to keep allowing it is a decision each business gets to make, and the point of counting is that it can now make it on its own numbers rather than on a feeling. The channels page describes what the web instrument reads, and the layers page explains why a crawler and an agent are never added into the same total.

FAQ

Questions this note answers.

  1. 01What is Meta-ExternalAgent?

    The user agent Meta documents for its crawler that indexes and gathers training data for its AI products. It reads pages; it does not act on them. In the Index it belongs to the reading layer, alongside search crawlers and answer-engine fetchers.

  2. 02Was the surge caused by the Muse launch?

    The timing is tight: baseline until 8 September, a sharp rise from 9 September, a peak on 16 September, back to baseline on 18 September. Meta has not said why. The Index treats the correlation as strong and the cause as plausible but not proven.

  3. 03Did the surge cost the businesses anything?

    Bandwidth and cache misses, not people. Reading-layer traffic never rings a phone or fills a form, which is exactly why the Index keeps it separate from the acting layer. A surge in reading is a hosting question; a surge in acting is a staffing one.

  4. 04Should a business block Meta's crawler?

    Only with a reason. The crawler that read the site in September is the same one that lets an assistant describe the business accurately later. The monitored sites allow it in robots.txt and count it; blocking is a choice each business can make on its own numbers.