When an AI agent reaches your website, your phone line or your chat window, does it tell you what it is? Some do, in ways you can check against the operator's own documentation, and others do the same work without a word.
That split runs through the acting layer — the agents that fetch, call, type and submit on someone's behalf, as distinct from the crawlers that only read. The wider map sits on the layers overview and in the third kind of caller; this note goes one layer deep.
All operator documentation and legal texts cited below were checked on 5 October 2026. Where a source does not settle a point, we say so and stop.
What Makes An Agent Declared
An agent is declared when it presents a claim about what it is, and that claim can be checked against something its operator has published. The claim might be a user agent that names the operator, a cryptographic signature on a web request, or a plain-language disclosure at the start of a call or chat.
An agent is undeclared when it presents none of that. It may still be entirely legitimate — a browser automation running for one person, a voice assistant that never states its nature — but the only evidence of what it is comes from how it behaves.
Keep in mind that the distinction is about evidence, not intent. A declared agent can misbehave, and an undeclared agent can be doing exactly what its user asked.
The Web: What A Declared Agent Presents
The web is where declaration is most developed, because the major operators publish their agents' identities. A declared agent on the web can present up to three layers of evidence:
- A named user agent. OpenAI's bot documentation lists ChatGPT-User for “certain user actions in ChatGPT and Custom GPTs,” alongside GPTBot for training crawls and OAI-SearchBot for search. Perplexity's bot guide lists Perplexity-User for user-initiated fetches, and Google's user-triggered fetchers page lists a Google-Agent token “used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request.”
- A published IP list. OpenAI publishes a separate range file per agent, including chatgpt-user.json, Perplexity publishes perplexity-user.json, and Google points Google-Agent at user-triggered-agents.json. A request that names the operator and arrives from that operator's own published range is a far stronger claim than either signal alone.
- A signature. Google's fetcher page states that it is “experimenting with the Web Bot Auth protocol, using the https://agent.bot.goog identity.” The mechanics — which headers carry the signature, where the key directory lives, what a server verifies — are laid out in what a signed agent request looks like, and the wider shift is logged in agents have started signing their requests.
Here is how a declared user agent reads on the wire, exactly as OpenAI publishes it for ChatGPT-User:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot
Note that the string carries the operator's product name, a version number and a URL pointing back to the documentation. That is the whole point of a declaration — it hands the receiving server the means to check it.
That said, a user agent on its own is still only a claim. Google's same page warns that “the user agent string can be spoofed,” which is why the IP list and, increasingly, the signature carry the weight of verification.
The operators also document different crawl policies for these user-initiated agents. OpenAI says robots.txt rules “may not apply” to ChatGPT-User, and Perplexity says Perplexity-User “generally ignores robots.txt rules” because a user asked for the page.
Anthropic takes the other position. Its crawler help article says Claude-User honors robots.txt exclusion signals, as do ClaudeBot and Claude-SearchBot.
For an operator, these differences matter more than the user agent names themselves. A declared agent tells you not only who sent it but, through its operator's published policy, how it intends to behave once it arrives.
The Web: What An Undeclared Agent Leaves Behind
An undeclared agent on the web typically drives an ordinary browser, often headless or remotely hosted, and sends an ordinary browser user agent. What it leaves instead of a declaration is a pattern, and the patterns include but are not limited to:
- Origin. Traffic from cloud-provider address space rather than residential or mobile networks. Plenty of people browse through cloud-hosted VPNs, so this shifts the odds without settling them.
- Cadence. Page-to-page timing that is too regular, too fast, or free of the idle gaps a person leaves while reading. Agents can add jitter, and some people are simply quick.
- Browser environment. Automation markers, missing fonts, or a rendering stack that does not match the browser the user agent claims. Privacy extensions produce many of the same mismatches in ordinary human browsers.
- Path shape. Direct navigation to deep pages, no scrolling, and structured extraction of the same elements on every page. Screen readers and power users can produce similar paths.
Every one of these has an innocent human explanation. This is why we treat them as evidence toward a label, never as an identification.
The Phone: Disclosure Versus Attestation
On a phone line, the declared case is a spoken disclosure — the agent says, near the start of the call, that it is an automated assistant acting for a named person or business. Some operators build this in; where an operator's own documentation does not commit to it, we do not assume it.
Remember that STIR/SHAKEN does not fill this gap. It attests to the calling number and the originating carrier's knowledge of the caller, not to whether a person or a model is speaking — the full breakdown is in what an AI agent's phone call carries on the line.
An undeclared voice agent leaves only acoustic and conversational traces. Those include response latency that sits in a narrow band, turn-taking that never overlaps, unusually even prosody, and recoveries from interruption that restart a sentence rather than adjusting it mid-phrase.
Each new generation of speech synthesis narrows those gaps. As a result, the phone supports the weakest “likely” of the four channels, and STIR/SHAKEN offers no field in which a voice agent could prove what it is.
Chat: Where Disclosure Is Becoming A Rule
Chat is the channel where law most directly reaches the declaration itself. California's bot statute, operative 1 July 2019 under Business and Professions Code section 17943, makes it unlawful to use a bot “with the intent to mislead the other person about its artificial identity” to incentivize a sale or influence a vote, and it permits bot use where the bot is disclosed.
In the EU, Article 50(1) of the AI Act requires providers to ensure that AI systems “intended to interact directly with natural persons” are designed so those people “are informed that they are interacting with an AI system.” Article 50 applies from 2 August 2026 under Article 113.
Neither text, read on its own, prescribes a machine-readable format for the disclosure or a fixed position for it in a conversation. We stop there, because the sources do not go further.
A declared agent in chat says what it is in its first message, often naming the person or service it acts for. An undeclared one leaves textual and timing traces instead:
- Typing cadence. Whole messages arriving at once, or at speeds no keyboard produces, where the chat widget reports keystroke events at all.
- Register. Uniform formatting, complete sentences on every turn, and an absence of typos that holds across a long exchange.
- Session behavior. No page navigation before or during the chat, or a session that opens directly onto the chat endpoint.
Each of these also describes a careful person drafting in a text editor and pasting. Accordingly, these traces move a session toward “likely agent” and no further.
Forms: The Quietest Channel
Forms are the channel where agents say the least. A form submission is still a web request, so an agent that signs its requests can carry that signature into the submission — the same mechanics apply, and the standards hub covers them.
Beyond that, we have not found an operator that documents a form-level way for an agent to announce itself. An undeclared submission leaves the following instead:
- Fill timing. Every field completed in milliseconds, or the form submitted without the page's own scripts having run.
- Field behavior. Hidden fields populated, fill patterns that match no browser's autofill, or identical values repeated across many submissions.
- Session context. A submission with no preceding page view, or one arriving with the same origin and environment patterns described for the web above.
Keep in mind that password managers and browser autofill produce fast, perfect fills for ordinary people every day. Form signals are useful mainly in combination with the web signals from the same session.
Why “Likely” Is The Ceiling
Every undeclared signal above is probabilistic, and stacking them raises confidence without ever converting it into fact. Five weak signals pointing the same way make a strong inference; they do not make an identification.
The practical consequence is a vocabulary rule. A declared agent can be labeled with its operator's name, while an undeclared one can be labeled “likely agent,” with the supporting signals attached, and nothing stronger.
There is also a direction-of-error problem. A behavioral model tuned to catch more undeclared agents will mislabel more humans, and a model tuned to protect humans will miss more agents — no threshold removes both errors at once.
What's more, the signals decay. Each published detection heuristic gives agent builders a target, so a pattern that separated agents from people last quarter may not separate them this quarter.
Why The Index Counts Declared Agents On Their Own
The Aethelforge Index reports declared agents as their own count, separate from any behavioral estimate. The reasoning is set out on the method page, and it comes down to what each number can be checked against.
A declared count rests on claims an outside reader can verify — the operator's published user agent, its published IP list, its signing key. An undeclared estimate rests on a model's judgment, and blending the two would give a probability the appearance of a headcount.
Keeping them apart also makes the trend legible. As more operators move from user-agent claims to signed requests, traffic that was once only inferable becomes countable, and that movement shows up only if the declared figure stands alone.
For a sense of how quickly the mix of automated traffic can move, see the crawler surge that followed Muse. For the current figures themselves, go to the current Index edition rather than any number quoted secondhand.
Recognise, Route, Receipt
None of this is an argument for turning agents away. The point of separating declared from undeclared is to treat each one appropriately:
- Recognise. Verify declared agents against the operator's published identity, and label undeclared traffic “likely” with its supporting signals attached.
- Route. Send verified agents to paths built for them — structured endpoints, machine-readable confirmations — and keep the human path clear for people.
- Receipt. Record what each agent presented, what was verified and what was inferred, so the record distinguishes fact from estimate long after the session ends.
After all, an agent that declares itself is offering a better interaction than one that does not. The sensible response is to make declaring worth its while.
If you're mapping which of your channels can already verify a declared agent and which can only infer one, start with the method and the layers overview. Then set your own picture against the current Index, and read what a signed agent request looks like before you decide what to verify first.