Building Agent rendering in Next.js and Sitecore

Agent rendering
We've now come to the third and final part of our agent rendering series. In the first blog post, I wrote about the modern DXP frontends and how it's not always easy for AI agents to consume our content. In the second part, we've taken a look at the architecture that we need for this. In this final part, I've built a Proof of Concept to show how it might work.

For the Proof of Concept, I used the Sitecore Next.js Skate park starter, this is a starter that is provided by Sitecore to help you get up-and-running for a headless frontend against SitecoreAI. In this starter, I implemented two separate capabilities: a consumer mode that allows components to render a flattened, non-interactive representation for AI agents, and HTTP content negotiation so the page can be requested as Markdown. You can find the implementation in my repo kit-nextjs-skate-park example on GitHub.

The implementation deliberately keeps the normal human experience untouched. The last thing you want to do, is ruin your current site and the outstanding ranking that it might have. Therefore, a regular browser will still get the same rendering of the components as before. An AI agent can get a different representation of those same components, with interaction removed and hidden content exposed. When that agent asks for text/markdown, the page will be rendered first through the agent-friendly path before being converted to Markdown.

One pipeline, two rendering outputs

When first thinking about a proper solution to fix this. I thought about creating a separate pipeline to support the Markdown rendering. That would have created a whole new, separate pipeline to maintain. The first problem is not HTML versus Markdown. The first problem is that some components are designed around human interaction. A navigation menu may require a click, a carousel may only show one slide at a time, another component may even fetch additional information after it has rendered. Simply converting the initial HTML to Markdown does not solve any of that. So I ended up with the following concerns: Consumer: default | agent, Representation: HTML | Markdown

The consumer mode decides what the component should render. Representation decides how the result will be delivered. When you look at the code, you will find that the consumer detection lives in consumer-mode.ts, and the AI-specific User-Agents are stored in ai-user-agents.ts.

This is how the flow is implemented:
Request / response flow

The distinction between consumer and representation turns out to be important. The agent version of a component is allowed to have a different DOM structure, but it must remain the content equivalent to the human version. A carousel can become a simple stack of slides, a hamburger navigation will become an expanded list. If a component loads information through client-side calls, its agent representation must retrieve that information server-side instead. Allowing the agent to have all the relevant information it needs to provide good possible context for constructing an answer. Below here, you can see the flow how it goes:

Component rendering
By using this flow, we keep the logical in the rendering component for agent or humans, allowing for an easier overhaul to output Markdown rather than HTML. By doing so, we only need to have one function to convert all components into markdown, rather than having each component create it's own markdown representation. Of course you could also do it like that, but I think you might end up in a maintenaince nightmare real soon.

Detecting the consumer

In order to create an implementation for detecting which type of consumer was accessing our site, I needed to understand how we could identify each one. For this, I've created a list of AI agents types and stored them in ai-user-agents.ts. The list is kept deliberately small and explicit, I don't want to change the flow for none AI user agents, that might cause an unwanted side effect.
export const AI_AGENT_USER_AGENTS = [
  'GPTBot',
  'ChatGPT-User',
  'ClaudeBot',
  'Claude-Web',
  'anthropic-ai',
  'PerplexityBot',
  'Google-Extended',
  'FacebookBot',
  'cohere-ai',
] as const;
A deliberate decision I took, was to keep the classic search-engine crawlers out of this list. Agent mode is not meant to become a generic "bot mode" for now. Search crawlers are becoming more able to work with Javascript actions as well, I don't want to risk harming your SEO ranking in the process. Let's move on, on how I implemented the user-agent decetion in consumer-mode.ts. This code turns the User-Agent detection into a typed value:
export type ConsumerMode = 'default' | 'agent';

export interface ConsumerContext {
  mode: ConsumerMode;
}

export const detectConsumerMode = (headers: Headers): ConsumerContext => {
  const userAgent = headers.get('user-agent') ?? '';

  return isAiAgentUserAgent(userAgent)
    ? { mode: 'agent' }
    : { mode: 'default' };
};
From this point on, we can start creating the implementation into the components.

Adding consumer mode in Sitecore components

To determine the type of consumer requesting the content, I wanted to avoid introducing a separate React Context provider purely for agent rendering. Every Sitecore component already has access to the current page through useSitecore(), so it made more sense to extend the existing page context with the consumer information. I therefore introduced a ConsumerContext type and extended the Sitecore page object with a consumer property.

There was, however, a TypeScript issue I encountered with this approach. The SDK's Page is a type alias rather than an interface, so declaration merging was not an option. The solution was to define a local page type that extends the SDK type and becomes the application's source of truth:
import { Page as SdkPage } from '@sitecore-content-sdk/nextjs';
import { ConsumerContext } from 'lib/consumer/consumer-mode';
You can see that in src/lib/component-props/index.ts.

The actual page route then detects the consumer and adds the result to the page before it is passed into the existing Sitecore providers:
const headers = await nextHeaders();

const pageWithConsumer: Page = {
  ...page,
  consumer: detectConsumerMode(headers),
};
The wiring lives in src/app/[site]/[locale]/[[...path]]/page.tsx.

From that point onwards, any component can make one simple decision:
if (getConsumerMode(page) === 'agent') {
  // render the agent representation
}

Flattening interactive components

For the PoC, I implemented the pattern in two components: Navigation and Carousel. For the navigation, the default experience exposes the interactive navigation behaviour, including a hamburger-style interaction for smaller screens. That is useful for a person, but not needed for an agent that doesn't need the UI experience. So when in agent mode, Navigation.tsx renders an expanded recursive list instead. The underlying content does not change; the links are all the same, we only removed the experience around it.

In the Carousel.tsx component, we did more or less the same thing. The component uses an interactive carousel with previous and next controls. For a human, that is fine, for an agent, hiding the two other slides behind a next button makes little sense. So when run in the agent mode, the code simply stacks every slide in document order:
const CarouselAgent = ({ slides }: { slides: CarouselSlide[] }): JSX.Element => (
  <div className="flex flex-col gap-8">
    {slides.map((slide) => (
      <CarouselSlideContent key={slide.id} {...slide} />
    ))}
  </div>
);
You might think, but is that really a huge improvement? Perhaps not really, I would have been better to use an example that also uses client-side loading, but you get the idea: don't create an alternate version of every component by default. Most components are probably already fine. Only introduce an agent representation where interaction, client-side loading or presentation hides meaningful information.
The most important rule is content parity; ensure that every content remains in both representations.

Rendering marking to the consumer

Now that we've flattened our components for agent mode, we can continue one step further and render the content out as Markdown. Making our content even better to read for agents. For this to work, I didn't want to expose a new URL structure; you should simply be able to request a page, where, based on the Accept header, the rendering shifts from HTML to Markdown. In order to do this, I've altered the next.config.ts to use a header-based rewrite.
{
  source: '/:path*',
  destination: '/api/agent-markdown/:path*',
  locale: false,
  has: [
    {
      type: 'header',
      key: 'accept',
      value: '(.*)text/markdown(.*)',
    },
  ],
}
The caller still requests:/Carousel, but can ask for another representation by: Accept: text/markdown.  That request is internally routed to src/app/api/agent-markdown/[[...path]]/route.ts, therefore ensuring that the public URL never changes, based on the Accept header.

Why render HTML first?

Like I said before, I wanted to create a simple implementation that would work independently of the components. One might argue that it might be simpler to create a separate rendering that would fetch the raw Sitecore layout data and convert fields directly to Markdown. I deliberately did not want that, because of component behaviour. A component can do more than render the fields Sitecore originally gave it. It can resolve references, call another service, enrich data server-side, or replace client-side fetching with a server-side equivalent in agent mode. If Markdown starts from raw layout data, all of that is bypassed and needs to be rebuilt again. Or become highly intertwined with the default mode, waiting for a nightmare to happen.

Instead, the Markdown route asks the application to render the same page again as HTML. The helper for that lives in render-agent-page.ts. It forwards only the request context that matters:
copyHeader(request, headers, 'user-agent');
copyHeader(request, headers, 'accept-language');
copyHeader(request, headers, 'cookie');
copyHeader(request, headers, 'authorization');

headers.set('accept', 'text/html,application/xhtml+xml');
headers.set('x-agent-markdown-render', '1');
The original User-Agent is particularly important. If the original request came from GPTBot, the internal HTML render also comes from GPTBot. That means the normal page pipeline detects agent mode again, so the Navigation and Carousel components automatically render their flattened versions, thus making our life simpler with one implementation to maintain here. The Markdown implementation therefore inherits the agent rendering for free.
GET /Carousel
Accept: text/markdown
User-Agent: GPTBot
        ↓
Markdown rewrite
        ↓
Internal HTML render
User-Agent: GPTBot
Accept: text/html
        ↓
Agent-aware components flatten themselves
        ↓
Flattened HTML
        ↓
Markdown conversion
The x-agent-markdown-render header is used as a recursion guard. If you've ever written recursive code, you might learned that hard way that it needs a stopping point... The internal request hits the same origin and pathname, so there needs to be a way to distinguish the internal HTML render from the original Markdown request. This approach does have a cost: one Markdown request results in an additional internal HTTP round trip.

For a production implementation, that deserves measurement and probably optimisation, for example through caching. For this proof of concept, I preferred cleanliness and simplicity to understand the concept.

Turning the rendered page into Markdown

Now that we've the flattened HTML from our site, the final step is straightforward. We've built code to do the conversion in html-to-markdown.ts. Here we use Turndown together with the GFM plugin:
const turndown = new TurndownService({
  headingStyle: 'atx',
  bulletListMarker: '-',
  codeBlockStyle: 'fenced',
  emDelimiter: '*',
  strongDelimiter: '**',
});

turndown.use(gfm);
In order to further strip down the HTML, we look for Scripts, styles, templates, SVG markup and other irrelevant elements, which are removed before the HTML is converted to: Content-Type: text/markdown; charset=utf-8

There was only one minor problem when using Turndown. Removing the head section wasn't reliable enough in Node. The Title and metadata content could still end up in the Markdown output. So to guarantee a proper output, we strip the head block before Turndown sees it and keep the Turndown removal rules as a second line of defence. That is not the kind of thing that appears in the architecture diagram, but it is exactly the kind of thing you discover once you actually build the feature.

Testing the solution

The tests focus less on what the markup looks like and more on whether the contract still holds. consumer-mode.test.ts checks that the known AI User-Agents activate agent mode, while normal browsers and classic search crawlers do not. Next to that, the tests in Navigation.agent-mode.test.tsx checks that agent mode exposes the same links as the default navigation. Ensuring we don't break our existing SEO while outputting a different representation.
Same content, different DOM shape.
The point is not to create a stripped-down page that happens to be easier for a crawler. The point is to expose the same meaningful content without requiring human interaction to reach it.

What to do before going to production?

This is a proof of concept, nothing more. You should not put this into production, before making changes to performance and such. Also, ensure that you have at least a baseline of your site. That might include performance SEO, ranking and more. Once yoy have that in order, there are a few things I would want to improve before releasing into production.

The first item on the list is caching. The Markdown response currently uses private, no-store. That is the safest starting point when request-time context, personalisation, or enrichment might affect the output, but it is not necessarily where I would leave it. A real implementation needs an explicit caching and revalidation strategy.

Another thing to keep an eye out is personalisation, or A/B testing. What version are you going to show to an AI bot? Or perhaps even leave out the whole thing. It's up to you to decide.

The next item is the internal HTTP round trip. It is a very convenient way to reuse the complete Next.js and Sitecore rendering pipeline, but it adds latency and work. If this became a heavily used capability, I would investigate extracting the agent representation and server-side enrichment into shared server functions so HTML and Markdown can serialise the same enriched model without requesting the application over HTTP.

I would also make the consumer detection more configurable. User-Agent detection is a useful starting point, but the architecture should not depend on it forever. A future version could allow explicit agent-mode negotiation through a request header as well. Also, move the configuration to the configuration, rather than hardcoding it into the solution.

Finally, this implementation currently lives only in the skate park starter. I deliberately kept the scope small, so I could prove the pattern before trying to generalise it across other starters.

Source code

The complete source code is available here: github.com/rvdplas/agent-rendering/tree/main/examples/kit-nextjs-skate-park. I mentioned earlier, this code is Proof of Concept. It is at best a reference implementation of the idea, including the trade-offs and some of the rough edges that came with building it.

Where this leaves the idea

When I started this series, the question was whether modern DXP frontends should start treating AI agents as another type of consumer. After building the proof of concept, I think the architecture is less exotic than it first sounds. The CMS does not need to change. The structured content does not need to be duplicated. The human experience does not need to be compromised.

What changes is the assumption that every consumer needs the same representation.

For a human, an interactive carousel makes sense. For an agent, a list of the same slides may be better. For a browser, HTML is the obvious representation. For an agent, Markdown may be easier to consume. The content can remain the same while the rendering layer becomes more aware of who, or what, is asking for it. That is the part of this experiment I find most useful.

Agent-ready rendering does not need to become a separate delivery platform. In a modern headless architecture, it can simply become another capability of the frontend.

With that said, it's a wrap on agent-ready rendering. Thank you for reading!