What is your developer-dependent CMS actually costing you? Calculate your costs and get a full report
What is your developer-dependent CMS actually costing you? Calculate your costs and get a full report
Going Live
Treat AI agents as an audience: server rendering, llms.txt, Markdown twins, JSON-LD, stable headings and chunk-friendly content, with Next.js and Fetch SDK examples based on this docs site.
More of your readers now reach your content through an assistant. An AI agent fetches a page, extracts the text, and answers a question with it, often without a human ever loading your site. That agent is an audience, and it has different needs from a browser.
This guide covers six mechanisms that make an Agility-powered site easier for agents to fetch and understand:
llms.txt indexThe worked example throughout is this documentation site, which runs on Agility CMS and Next.js and implements all six. Each section shows a simplified version you can adapt for your own Next.js App Router site using the Agility Fetch SDK (@agility/content-fetch) through the lib/cms/ wrappers from the Agility Next.js Starter, which tag each request so the starter's publish webhook can revalidate it. The starter already server-renders its pages (section 1). It doesn't include llms.txt, Markdown twins or JSON-LD, so the snippets show how to add them.
What this does and doesn't do. These mechanisms make your content easier to retrieve, parse and attribute. None of them guarantees that a search engine or AI assistant will rank, cite or quote your pages. Treat them as removing obstacles, not as a ranking strategy.
| Mechanism | What it does | Who reads it |
|---|---|---|
| Server rendering | Puts the content in the initial HTML response, so a fetcher that doesn't run JavaScript still gets it | Every crawler and agent fetch tool |
llms.txt | Gives a plain-text, Markdown-formatted index of your important pages with one-line descriptions | Agents and tools that look for it (it's a proposed convention, not a standard every agent follows) |
| Markdown twin | Serves the page body as Markdown with no navigation, scripts or styling | Agents following links from llms.txt or a rel="alternate" link |
| JSON-LD | Describes the page in schema.org terms: what it is, who published it, when it changed, where it sits in the site | Search engines and any parser that reads structured data |
| Stable headings | Gives each section a predictable anchor and a meaningful label | Anything that splits pages into sections or links to them |
| Chunk-friendly content | Lets a section make sense when it's read on its own | Retrieval systems that index and quote passages rather than whole pages |
Start here, because nothing else matters if the content isn't in the response.
When a page fetches its content in the browser (a useEffect that calls an API, for example), the HTML that arrives from the server is an empty shell. A browser fills it in. Not every crawler or agent fetch tool runs JavaScript, and you can't count on the ones that do. To them, the page has no content.
The Agility Next.js Starter is server-first: components are React Server Components that fetch on the server, and the catch-all route prerenders the pages in your sitemap. Keep it that way:
lib/cms/ wrappers."use client" only for interactivity, and pass content into client components as props so it's already in the HTML.Check it by fetching the raw HTML and looking for a sentence from the page body:
curl -s https://www.example.com/blog/my-post | grep -c "a sentence from the post"
A count of 0 means the content isn't in the server response.
llms.txt is a proposed convention: a Markdown file at the root of your site with an H1, a short summary in a blockquote, and H2 sections of links, each with a one-line description. It gives an agent a map of your content without crawling it.
This site serves one at /docs/llms.txt. It's a Next.js route handler (app/llms.txt/route.ts) built from the cached flat sitemap, so it stays current with the same publish webhook that refreshes the pages. It skips folders, redirects and pages hidden from the sitemap, lists curated flagship pages first, and links each article to its Markdown twin.
Here's a simplified version for your site:
// app/llms.txt/route.ts
import { getSitemapFlat } from "lib/cms/getSitemapFlat"
const SITE_URL = "https://www.example.com"
// Optional one-line descriptions for the pages that matter most.
const PURPOSES: Record<string, string> = {
"/pricing": "Plans and what each includes",
"/blog": "Product news and guides",
}
export async function GET() {
// Tagged agility-sitemap-flat-{locale}, which the starter's webhook
// revalidates on page publish.
const sitemap = await getSitemapFlat({
channelName: process.env.AGILITY_SITEMAP || "website",
languageCode: "en-us",
})
const lines = Object.entries(sitemap)
.filter(([, node]) => !node.isFolder && !node.redirect && node.visible?.sitemap !== false)
.map(([path, node]) => {
// Pages backed by a content item have a Markdown twin (see section 3).
const url = node.contentID ? `${SITE_URL}${path}.md` : `${SITE_URL}${path}`
const purpose = PURPOSES[path]
return `- [${node.title || node.menuText}](${url})${purpose ? `: ${purpose}` : ""}`
})
const body = [
"# Example Co",
"",
"> One or two sentences that say what the company does and what this site covers.",
"",
"## Pages",
"",
...lines,
"",
].join("\n")
return new Response(body, {
headers: { "content-type": "text/plain; charset=utf-8" },
})
}
Notes:
getSitemapFlat is the starter's tagged wrapper around the Fetch SDK's getSitemapFlat, and takes the same channelName and languageCode parameters. Its keys are paths, and its nodes carry title, menuText, isFolder, redirect, visible and, for dynamic pages, contentID./home. Map it to / if your site serves it at the root.An HTML page carries navigation, footers, scripts and styling that an agent has to strip out before it reaches the content. A Markdown twin skips that step: the same content, served as Markdown, at a predictable URL.
On this site, every article has one. Add .md to any article URL, for example /docs/developers/agility-cms-mcp-server.md. Three pieces make it work:
proxy.ts sends any path ending in .md to an internal route, /api/article-md/{path}. It runs before the static-file check, because a .md path contains a dot.getContentItem wrapper, and serializes it (lib/cms-content/articleMarkdown.ts). The output is the H1, a Source: line with the canonical URL, then the body. Markdown articles are served nearly verbatim; older articles stored as editor blocks are converted block by block. Pages that aren't articles return a 404 that points to llms.txt.rel="alternate" link in each article's <head> advertises the twin with type="text/markdown", so an agent that lands on the HTML page can find it. It's added only for articles, because advertising an alternate that 404s is worse than advertising none.Here's a simplified version for a site where dynamic pages are backed by content items with an HTML rich text field. It uses turndown to convert HTML to Markdown; if your body field is already Markdown, return it as is.
The route handler:
// app/api/md/[...slug]/route.ts
import TurndownService from "turndown"
import { getSitemapFlat } from "lib/cms/getSitemapFlat"
import { getContentItem } from "lib/cms/getContentItem"
const SITE_URL = "https://www.example.com"
const turndown = new TurndownService({ headingStyle: "atx", codeBlockStyle: "fenced" })
interface IPost {
title: string
content: string // HTML rich text field
}
export async function GET(
_request: Request,
{ params }: { params: Promise<{ slug: string[] }> }
) {
const { slug } = await params
const path = "/" + slug.join("/")
const sitemap = await getSitemapFlat({
channelName: process.env.AGILITY_SITEMAP || "website",
languageCode: "en-us",
})
const node = sitemap[path]
if (!node?.contentID) {
return new Response("Not found. See /llms.txt for the index.", { status: 404 })
}
const item = await getContentItem<IPost>({
contentID: node.contentID,
languageCode: "en-us",
})
if (!item?.fields) return new Response("Not found", { status: 404 })
const markdown = [
`# ${item.fields.title}`,
"",
`> Source: ${SITE_URL}${path}`,
"",
turndown.turndown(item.fields.content || ""),
"",
].join("\n")
return new Response(markdown, {
headers: { "content-type": "text/markdown; charset=utf-8" },
})
}
The rewrite (Next.js 16 calls this file proxy.ts; on Next.js 15 it's middleware.ts with a middleware export). If your project already has one, add this branch near the top:
// proxy.ts
import { NextRequest, NextResponse } from "next/server"
export function proxy(request: NextRequest) {
const { pathname } = request.nextUrl
// /blog/my-post.md -> /api/md/blog/my-post
if (pathname.endsWith(".md") && !pathname.startsWith("/api/")) {
const url = request.nextUrl.clone()
url.pathname = `/api/md${pathname.slice(0, -3)}`
return NextResponse.rewrite(url)
}
return NextResponse.next()
}
Check your matcher config. Many projects exclude every path containing a dot from the proxy, which would also exclude .md URLs.
The alternate link, in the page's generateMetadata:
// inside generateMetadata in app/[...slug]/page.tsx
const canonical = `https://www.example.com${path}`
const isContentPage = !!node?.contentID
return {
title: page.title,
alternates: {
canonical,
...(isContentPage ? { types: { "text/markdown": `${canonical}.md` } } : {}),
},
}
Because the route reads through the same tagged wrappers as the HTML page, the publish webhook that revalidates those tags refreshes both. There's no second copy of your content to keep in sync.
JSON-LD describes the page in schema.org vocabulary: this is an article, this organization published it, it was last modified on this date, it sits under this section. A parser gets those facts directly instead of inferring them from layout.
This site builds one JSON-LD @graph per page in lib/cms-content/getRichSnippet.ts and renders it as an in-body <script type="application/ld+json">. Every page gets an Organization, a WebSite and a WebPage, plus a BreadcrumbList when the page is nested. Articles add a TechArticle with headline, description, datePublished, dateModified and articleSection, and a VideoObject for each embedded video. A few design choices are worth copying:
@ids. The organization has one @id and every page references it with {"@id": ...}, so a parser sees one publisher across all pages instead of hundreds of anonymous copies.Here's a simplified version for a blog post:
// lib/cms-content/getJsonLd.ts
const SITE_URL = "https://www.example.com"
const ORG_ID = `${SITE_URL}/#organization`
export const getPostJsonLd = (post: {
title: string
description?: string
url: string
datePublished: string
dateModified: string
}) =>
JSON.stringify({
"@context": "https://schema.org",
"@graph": [
{ "@type": "Organization", "@id": ORG_ID, name: "Example Co", url: SITE_URL },
{
"@type": "BlogPosting",
"@id": `${post.url}#article`,
headline: post.title,
description: post.description,
url: post.url,
datePublished: post.datePublished,
dateModified: post.dateModified,
author: { "@id": ORG_ID },
publisher: { "@id": ORG_ID },
},
],
})
// in the page or component that renders the post
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: getPostJsonLd(data).replace(/</g, "\\u003c") }}
/>
The replace escapes < so a value containing </script> can't break out of the tag. For dateModified, a content item's properties.modified is a good source. For a fuller treatment, including blog posts, events and articles, see Implementing JSON-LD Structured Data with Next.js.
Only describe what's on the page. Structured data that disagrees with the visible content is a liability, not a signal.
Headings are how both people and machines find their way through a page. On this site, every heading gets an ID generated from its text, so ## Server URL becomes #server-url and can be linked directly.
In Agility, this mostly comes down to editorial practice in your rich text and Markdown fields, plus making sure your components render real <h2>/<h3> elements rather than styled <div>s.
Many AI systems don't read a page top to bottom. They split it into passages, index the passages, and retrieve the few that match a question. A passage that only makes sense in context gets retrieved without that context.
Your content model can help. Separate fields for a summary, a body and FAQs give your components (and your Markdown twin) clean, predictable structure instead of one large blob of HTML.
| Mechanism | Where it lives in this site's code |
|---|---|
| Server rendering | React Server Components on Cache Components, prerendered from the cached sitemap |
llms.txt | app/llms.txt/route.ts |
| Markdown twin | proxy.ts (the .md rewrite), app/api/article-md/[...slug]/route.ts, lib/cms-content/articleMarkdown.ts |
| Alternate link | lib/cms-content/resolveAgilityMetaData.ts |
| JSON-LD | lib/cms-content/getRichSnippet.ts, lib/seo/schema.ts |
| Heading IDs | The Markdown renderer adds an ID to every heading |
All of it reads through the same cached, tagged wrappers, so the Agility publish webhook that refreshes the HTML also refreshes llms.txt and the Markdown twins. For the caching model behind that, see Caching with Next.js and Agility.
# 1. Content is in the server-rendered HTML
curl -s https://www.example.com/blog/my-post | grep -c "a sentence from the post"
# 2. llms.txt returns 200 as text
curl -sI https://www.example.com/llms.txt
# 3. The Markdown twin returns 200 as text/markdown
curl -sI https://www.example.com/blog/my-post.md
# 4. The HTML advertises the twin
curl -s https://www.example.com/blog/my-post | grep -o '<link rel="alternate"[^>]*>'
# 5. JSON-LD is present
curl -s https://www.example.com/blog/my-post | grep -c 'application/ld+json'
Then validate your structured data with the Schema Markup Validator.