In the previous blog post, I looked at a question that started with Scrunch AXP: if AI agents are becoming another consumer of our digital experiences, should we keep transforming websites afterwards, or should we start designing our frontends with those consumers in mind from the beginning?
My conclusion was fairly simple. We should not create a second website for AI agents, and we should not create a second editorial process either. The CMS should remain the single source of truth. The content stays the same, but the representation will change depending on the consumer.
For a human visitor, that representation might contain navigation, carousels, tabs, personalisation, and all the other things we have spent years building into modern digital experiences. An AI agent doesn't need many of those interactions. It might benefit more from a stable, semantic and complete representation of the same information. I called the capability agent rendering.
The next question is: what architecture do we need to accomplish this?
We are not building two websites
This is the most important constraint. If every component suddenly gets a completely separate "human version" and "agent version", editors start writing different content for agents, and every new feature needs to be implemented twice. This will surely mean the architecture has already failed. The outcome should be much simpler. The CMS provides structured content, the frontend understands that content, and the rendering layer optimises it for a particular consumer.
Conceptually: CMS → component data → rendering layer → representation
The representation comes after the content. We are not creating another content model specifically for AI agents.
Separate detection from rendering
One of the first decisions is where we determine what kind of consumer is making the request. A CDN or edge layer could inspect it, middleware could do it, or the application itself could decide which representation to return.
For existing enterprise websites, doing this at the edge can make a lot of sense. That is part of what makes an approach such as Scrunch AXP interesting: you can introduce agent-friendly delivery without rebuilding the systems behind it. For a new headless implementation, I would argue that we might want to take matters into our own hands. But I would not tightly couple the rendering architecture to bot detection.
User agents change, new agents appear, and not every agent will identify itself in the same way. One part of the architecture should therefore determine which representation is requested, while another part knows how to render that representation. The rendering logic shouldn't care whether that decision came from a CDN, middleware, a request header or something else.
The component data stays the same
This is where headless architectures give us an advantage. By the time a frontend receives content from a CMS such as Sitecore or Storyblok, it already knows what that content represents. It knows that something is a hero, a product list, an accordion or a comparison component. That knowledge exists before the frontend turns the data into the visual experience we see in the browser.
Take an accordion, for example: the CMS might provide five items, each containing a heading and some content. For a human visitor, we render those items as an accordion because it makes the page easier to scan. But the underlying information isn't really an accordion. It is five sections of related content. An agent renderer could simply expose those five sections as headings followed by their content. Nothing changes in the CMS or editorial process. Only the representation changes. The same applies to tabs, carousels, selectors and other interactive components.
Only specialise where it is needed
That doesn't mean every component needs a separate agent renderer. A heading is still a heading. A paragraph is still a paragraph. An image with useful alternative text might already be perfectly understandable in its normal HTML representation.
Special handling becomes useful when human interaction makes information harder to access. An accordion is an obvious example. A carousel containing several products might be another. A comparison where information only appears after selecting different options is an even clearer case.
So instead of asking how we create an AI version of every component, I think the more useful question is: Which components currently hides information behind human interaction? That gives us a much more manageable problem.
What should the agent representation look like?
I don't think there is a final answer yet. My starting point would still be good semantic HTML. It already works across the web and allows us to preserve hierarchy, links, media references and structured information without inventing a completely new publishing format.
But the architecture shouldn't depend on semantic HTML alone. Markdown is interesting because it offers a simpler representation of textual content. Structured data can even provide extra context, and content negotiation could eventually allow a consumer to explicitly request another representation of the same resource.
The important part is not whether the first version returns HTML or Markdown. It is having a clear point in the frontend where a representation can be created for a specific consumer.
Keep it deterministic
For the first version, I would keep the agent representation deliberately boring. No personalisation. No A/B testing. No user-specific state. No generative rewriting during the request.
If the published CMS content hasn't changed, the output should ideally stay the same as well. That gives us a canonical representation which is easier to cache, test and monitor, while the approved CMS content remains the source of truth.
This also prevents agent rendering from quietly becoming another content channel. The moment we start generating alternative marketing copy specifically for AI agents, we introduce new questions around approval, compliance, translations and ownership. That might become interesting later, but I wouldn't start there.
Caching and testing become easier
A stable representation also makes caching much more attractive. Human experiences often contain personalisation, experiments and other state that make responses vary between visitors. A public, canonical agent representation can avoid most of that. Instead of letting every agent execute the full frontend and trigger all backend calls used by the human experience, we may be able to serve a much smaller cached response.
At the same time, we need to prove that simplifying the representation hasn't removed important information. Because both versions use the same source content, much of that testing can be automated. We can compare important text, links, media references, heading structure and structured data between both representations.
The goal isn't to create the smallest possible page. It is to create the simplest representation that still preserves the meaning.
The architecture I would propose today
When designing a new headless DXP today, I wouldn't make agent rendering a completely separate subsystem. I would make it a capability of the rendering layer. The CMS would continue managing structured content, workflows, translations, permissions and publishing. The frontend would continue receiving component data, but it would no longer assume there is only one way to turn that data into an experience.
Most components would have one representation (semantic HTML). Only the components where interaction hides information would need something different. Consumer detection could happen earlier at the CDN, edge or application level, while the rendering layer remains independent of how that decision was made. I am not suggesting this because we already know exactly what AI agents will need in the future. We don't. The value is that the architecture gives us a place to evolve without redesigning the frontend later.
From architecture to implementation
All of this still looks relatively clean on paper. The real test is what happens when we actually build it. How much additional code does an agent renderer require? Can we avoid maintaining two separate component trees? What happens to components that don't have a special agent representation? And how do we test that both versions still contain the same meaningful information?
That is what I want to explore in part 3. Instead of another architecture diagram, I want to build a small agent-rendering implementation using a modern headless frontend. We can take a few typical interactive components, render them normally for a human visitor and then create a deterministic representation for an AI agent. That should show whether agent rendering remains a useful architectural idea once we actually have to maintain the code.

