TL;DR: TIME now serves some AI agents a lean markdown version of its site, including sponsored facts that human readers don't see. Our 19 August 2026 retest found that access already varies by bot. The new internet has one URL, several audiences and a new market for machine attention.
I opened one TIME article through several crawler identities on 19 August 2026. The URL didn't change. Almost everything else did.
Googlebot received about 305,000 bytes of HTML. ClaudeBot and PerplexityBot each received 13,409 bytes of clean markdown. ChatGPT-User, OAI-SearchBot and GPTBot received HTTP 406 and no article at all.
Then I tested a TIME collection page. ClaudeBot and PerplexityBot received a markdown page containing an Ally Bank sponsored FAQ, campaign tracking and FAQPage schema. That paid material wasn't on the page a person reads.
Vincent Schmalbach documented the split on 5 August. Two weeks later, the bot matrix had already changed. That's the point. We're no longer publishing one stable page to one general audience. We're negotiating with several machine audiences that arrive with different names, permissions and commercial value.
What did TIME serve to AI bots on 19 August 2026?
I sent one GET request per crawler using the current full User-Agent string published by its operator, with an Accept header that allows any media type and identity encoding. Anthropic publishes the ClaudeBot token rather than a browser-style string. That's worth being precise about. If you don't use the full identity, you can get a different answer. Bare and partial bot names aren't included below. The dated response receipt includes every User-Agent, status, header subset and body hash.
The six official crawler identities produced three distinct response groups in our live check.
| Request identity | HTTP result | Format | Bytes returned | What it means |
|---|---|---|---|---|
| Googlebot | 200 | HTML | About 305,000 | Google's search crawler received the full web page |
| ClaudeBot | 200 | Markdown | 13,409 | Anthropic's model-development crawler received a machine-ready page |
| PerplexityBot | 200 | Markdown | 13,409 | Perplexity's search crawler received the same machine-ready page |
| ChatGPT-User, OAI-SearchBot and GPTBot | 406 | None | 0 | These OpenAI crawler identities were refused in this test |
The markdown response was about 23 times smaller than the HTML response. It's also carrying a fresh Mobian impression identifier, a no-store cache rule, a markdown format marker and a count of 3,323 tokens in its headers.
The collection page went further. Its machine copy included a clearly labelled sponsored block from Ally Bank. The block answered questions such as who Ally Bank is and whether customers can deposit cash. It also carried structured FAQ data and a campaign label. There's a sponsored label for the machine, but a person reading the public page won't encounter the unit.
What surprised me wasn't the markdown. It was the business model sitting inside it.

Why does this look like a new internet rather than a website trick?
We've always had web servers that vary what they return. A site can change language, compression, device layout or logged-in content. HTTP calls each result a representation of the same resource. RFC 9110 says the Vary header should name request fields that may influence representation selection. TIME returned a Vary header naming Accept, yet our controlled requests show that User-Agent also changed the result. You can't see the whole routing rule in the public header.
TIME's experiment crosses a more important line. The representation changes because the reader is a machine, and the machine version has its own advertising inventory, measurement and facts.
| The familiar internet | The new machine internet |
|---|---|
| A person loads a page | A crawler or agent requests a representation |
| Attention is measured in pageviews | Attention can be measured in fetches, impressions and tokens |
| Ads sit beside the article | Sponsored facts can sit inside the machine copy |
| One public page is the reference | One URL can return different formats and different content |
| Search crawlers are treated as one technical class | Training, search and user-triggered agents get separate policies |
| A click carries value | A cited answer or recommendation can carry value without a click |
This is the same economic pressure behind Cloudflare's Pay Per Crawl. Cloudflare's private beta lets publishers allow a crawler, block it or return HTTP 402 with a price. TIME and Mobian show another route: let selected machines in and place sponsored material in the representation they consume.
Both models point to the same change. Machine access has become a product in its own right, beyond the old side effect of running a website.
How are AI crawlers already different from one another?
The bot names sound interchangeable. They aren't.
OpenAI documents three separate access paths. OAI-SearchBot feeds ChatGPT search. GPTBot may collect material for model training. ChatGPT-User fetches pages when a person asks ChatGPT or a Custom GPT to visit them. OpenAI lets site owners allow search while refusing training.
Anthropic also documents three roles. ClaudeBot supports model development. Claude-SearchBot supports search. Claude-User retrieves a page at a person's direction.
A robots.txt decision is therefore a distribution decision. Blocking a training bot doesn't have the same effect as blocking a search bot. Blocking a user-triggered fetch can stop an assistant from reading a page at the moment somebody asks about your product.
TIME's changing responses show why a static crawler checklist doesn't age well. Schmalbach reported markdown from OAI-SearchBot and PerplexityBot on 5 August. Our full-identity test on 19 August still received markdown from PerplexityBot, but OAI-SearchBot received HTTP 406. You can't assume the response matrix will hold for a month, let alone a year. Even a shortened User-Agent can produce a different answer.
Is serving AI bots different content cloaking?
Format, intent and factual parity decide whether bot-specific content becomes cloaking.
Google defines cloaking as presenting different content to users and search engines with the intent to manipulate rankings and mislead users. In our test, Googlebot received the full HTML page. The separate markdown went to selected assistant agents. That's not the classic pattern of stuffing keywords into a page only Googlebot can see.
Still, the trust question remains. A clean markdown mirror of the same facts is a format choice. A machine page with extra commercial claims is a content choice. If an AI assistant absorbs those claims and drops the sponsored label when it answers a user, we've broken the disclosure chain somewhere between publisher, crawler and model.
My rule is simple: one truth, several formats. You can't judge parity by format alone. Headings may be cleaner. Navigation may disappear. Scripts and visual furniture may go. Material facts, claims, dates, prices and disclosures should remain aligned with the public page.
Why would a publisher put ads in pages humans never see?
Because the assistant doesn't need to send a person to the publisher before it can influence them.
Traditional display advertising needs a human pageview. AI search doesn't always need a click. It can complete the answer in its own interface after the model retrieves a source and extracts a passage. The user won't always visit the page that shaped the recommendation.
Machine-only sponsored facts move the ad closer to that retrieval step. The Ally block uses the form an answer engine likes: direct questions, short answers, named features and FAQ schema. It's built for extraction.
That doesn't mean an ad automatically becomes a recommendation. Models combine sources, apply their own ranking and can ignore sponsored material. The important shift is that the paid unit now competes for machine attention before a human sees the result.
The new audit question for every AI answer is this: which parts came from editorial pages, which came from clearly sponsored machine pages, and did the final answer preserve that difference?
What changes for GEO when the machine page can carry different facts?
The difference between SEO, AEO and GEO starts with retrieval versus selection inside an answer. I've stopped treating GEO as a content formatting job where we make pages clear, answer questions directly and add structured data. TIME's split shows that distribution and governance now matter just as much.
First, crawler access belongs in the measurement. An engine can't retrieve the page through a request that receives HTTP 406. It's possible the engine knows an older copy or reaches the information another way. Check the real bot response, not only the browser page and robots.txt file. AI visibility data only becomes useful when it leads to citation work.
Second, a brand needs evidence parity. If a machine version contains a different product claim, price or endorsement, you've created two sources of truth. That may win one retrieval and lose trust everywhere else.
Third, machine-readable advertising will create citation provenance problems. A model can quote a fact, but it won't necessarily keep the sponsored context around it. Brands and publishers need durable labels that survive extraction, not labels that only make sense inside the source file.
Fourth, bot policy can change without a redesign or public announcement. A marketing team can lose access in one answer engine while the website still looks perfect to every human visitor.
Should your website serve markdown to AI crawlers in 2026?
Start with clean HTML. It already gives serious crawlers headings, lists, tables, links and structured data. A separate markdown version adds another surface that can drift, leak drafts or carry stale claims.
I'd use a machine representation when it solves a measured problem: expensive rendering, blocked extraction, heavy navigation, a paid access model or a clear publishing partnership. Keep it generated from the same source as the human page. You'll need to test factual parity on every release.
For most brands, I'd start somewhere less exotic. Put the direct answer near the top. Use plain headings. Keep key facts in server-rendered HTML. Add schema that matches the visible copy. Make citations and author details easy to verify. Then test what each crawler actually receives.
What are the 6 moves brands should make now?
I'd put the following six-part check into the monthly GEO routine.
- Map every important crawler by purpose. Separate training, search and user-triggered agents. You'll need to decide what you want to allow for each one.
- Fetch your priority pages as those crawlers. Record status code, content type, byte size, canonical URL and the first meaningful passage. You'll want to repeat this monthly because policies move.
- Keep one factual source. Generate HTML, feeds or markdown from the same approved content. Don't hand-edit a hidden machine page.
- Test claim and disclosure parity. Prices, benchmarks, endorsements, dates and sponsored labels must survive every representation. If they don't, you've built two sources of truth.
- Log machine attention separately. Distinguish search bots, training bots and user-triggered fetches. You won't see the distribution decision in one total bot number.
- Measure the answer, not only the crawl. Check whether the brand is named, which source is cited, whether sponsorship survives and whether the answer is accurate.
I've come to see GEO as an operating system rather than a copywriting trick. The team that owns AI search visibility has to work across content, edge rules, analytics, brand safety and answer-level measurement.
Where does the new internet go next?
We're moving beyond a web with one page and one public audience. The next version has people, search crawlers, training crawlers and agents arriving at the same URL with different permissions and value.
TIME didn't invent content negotiation. It showed what happens when that technical feature meets advertising and AI retrieval. A page can now earn a machine impression, carry a machine-only sponsor and influence a person who never visits it.
Don't build a secret website for bots. Know exactly what each audience receives and keep the truth intact across all of them.
One address can now reach many readers. The real risk is letting those versions drift until nobody knows which one said what.



