# The Comprehensive Technical Guide to Building AI-Ready Documentation Systems

In my previous article on [From 71 to 100: What It Took to Make the Rootstock Developer Portal More Agent-Ready](https://blog.theowanateamachree.com/from-71-to-100-rootstock-developer-portal-ai-ready ),I shared what happened and what I learned while working on the [Rootstock Developer Portal](https://dev.rootstock.io).

The short version is that the portal moved from 71/100 to <mark class="bg-yellow-200 dark:bg-yellow-500/30">90/100</mark> to <mark class="bg-yellow-200 dark:bg-yellow-500/30">100/100</mark> on Fern's Agent Score. But getting there involved much more than changing content. The work touched the documentation build, delivery, representation, structure, and validation layers. This article goes deeper into those implementation details.

If you are new to concepts such as llms.txt, Markdown content negotiation, or MDX serialization, the first article is the better starting point.

If you work in docs-as-code, documentation engineering, developer experience, or technical writing and want to understand what happened underneath the documentation site, this is the indepth technical guide into building AI-ready documentation systems.

## TLDR:

| Concern | Before / problem | What changed |
| --- | --- | --- |
| Discovery | The generated machine index did not cover every useful hub. | Supplementary hub coverage |
| Representation | Human-facing HTML was not the only representation needed. | Markdown counterparts |
| Delivery | Machine consumers needed a reliable way to request Markdown. | Content negotiation |
| Components | MDX components did not always map directly to Markdown. | Serialization |
| Structure | Large/generated documentation needed better machine-facing structure. | Index and RPC restructuring |
| Size | Large representations create truncation risk. | Page size and export controls |
| Validation | Generated output could regress silently. | CI verification |
| Production | Local output does not guarantee deployed output. | Production validation |

## **The first lesson: fix machine access before rewriting content**

A lesson from the [project](https://blog.theowanateamachree.com/from-71-to-100-rootstock-developer-portal-ai-ready ) was that agent readiness is not necessarily a content rewrite exercise.

The current [AFDocs](https://afdocs.dev/checks/) reference groups its checks into seven categories:

1.  Content Discoverability
    
2.  Markdown Availability
    
3.  Page Size and Truncation Risk
    
4.  Content Structure
    
5.  URL Stability and Redirects
    
6.  Observability and Content Health
    
7.  Authentication and Access
    

Some of the highest-leverage work was infrastructural, and the work therefore focused on nine connected areas:

*   llms.txt coverage
    
*   Markdown availability
    
*   Machine-readable content negotiation
    
*   Markdown export
    
*   Content structure
    
*   Generated indexes
    
*   Page size
    
*   Documentation structure
    
*   CI validation
    

## **The 9 Areas We Changed**

### **llms.txt coverage**

**What is llms.txt?**

[llms.txt](https://llmstxt.org/) is a proposed Markdown-based convention for giving AI systems a concise, structured entry point into a website. Instead of mandating an agent to discover an entire documentation site through navigation, search, or HTML, an llms.txt file can provide a short description and links to relevant resources, ideally pointing to machine-friendly Markdown versions of those resources.

A simplified example looks like this:

```plaintext
# Rootstock Developer Portal

> Documentation for developers building on Rootstock.

## Developers

- [Smart Contracts](https://example.com/smart-contracts.md): Build and deploy smart contracts.
- [RPC API](https://example.com/rpc-api.md): JSON-RPC methods and examples.
- [Integrate](https://example.com/integrate.md): Integrate Rootstock into an existing application.
```

The goal here is to improve discovery for agents. That is what I mean by **llms.txt coverage**. Coverage is not simply "Does the file exist?" It is about whether the documentation that should be discoverable is represented in that machine-facing index. AFDocs now explicitly treats llms.txt coverage as a separate observability check, measuring how much of a site's documentation is represented in the file.

#### **Why Rootstock needed additional coverage**

For Rootstock, this became more interesting because not every useful documentation route was automatically represented by the generated index. So we supplemented the generated output with entries such as:

```javascript
const SUPPLEMENTARY_HUB_ENTRIES = [
  ['RIF Suite', '/concepts/rif-suite/', 'Open-source tools for building on Bitcoin.'],
  ['Blockchain Essentials', '/developers/blockchain-essentials/', 'Read transactions and deploy your first contract.'],
  ['Integrate', '/developers/integrate/', 'Integrate dApps with Rootstock SDKs and protocols.'],
  ['Libraries', '/developers/libraries/', 'SDKs and libraries for Rootstock development.'],
  ['RPC API', '/developers/rpc-api/', 'JSON-RPC providers and API guides.'],
  ['Smart Contracts', '/developers/smart-contracts/', 'Hardhat, Foundry, verification, and tooling.'],
  ['Node Setup', '/node-operators/setup/', 'Install and configure a Rootstock node.'],
  ['Developer Cheatsheet', '/cheatsheet/', 'One-page Rootstock developer quick reference.'],
];
```

*Note: These are not random pages. They are* ***category pages****. For a human, a category page provides navigation. For a machine, it can also provide an important map of the documentation structure.*

*That is an important connection between information architecture and agent readiness. The IA that helps a developer understand where information belongs can also help a machine understand how the documentation is organized.*

### **2\. Markdown availability**

This refers to giving machines a cleaner representation. So naturally, the next problem was representation.

Most documentation websites are built to deliver HTML. That makes sense for a browser. But HTML can include navigation, styling, scripts, layout components, and other interface elements that help humans but aren't needed for an agent trying to understand the documentation itself.

Markdown provides a much cleaner representation of the same knowledge.

AFDocs defines Markdown availability as whether agents can obtain documentation as Markdown rather than HTML, either through .md represented URLs or HTTP content negotiation.

So instead of:

```plaintext
GET /developers/rpc-api/
→ HTML
```

A documentation system can also expose:

```plaintext
GET /developers/rpc-api.md
→ Markdown
```

Or it can support content negotiation.

### **3\. Machine-readable content negotiation**

Markdown availability and content negotiation are related, but they are not the same thing. Content negotiation is an HTTP mechanism for allowing a client and server to agree on the representation of a resource.

For documentation, an agent can:

```plaintext
GET /developers/rpc-api/
Accept: text/markdown
```

and a server that supports this representation can respond with:

```plaintext
HTTP/1.1 200 OK 
Content-Type: text/markdown
```

Along with the Markdown representation of the page. This is important because the URL can remain the same while the representation changes according to what the client can consume. That is what **machine-readable content negotiation** means in this context. It is not a new documentation format. It is a delivery mechanism for providing an appropriate representation of existing documentation.

The final [Rootstock Agent Score](https://buildwithfern.com/agent-score/company/rootstock) reports that all 10 sampled pages supported content negotiation and that all 10 sampled pages supported .md URLs.

### **4\. Markdown serialization and export**

Had to learn about this distinction as the work progressed. The Rootstock source documentation uses Docusaurus MDX. MDX is useful because it allows Markdown to coexist with components and other JSX-based behavior. That is also the problem. Not everything that works as an interactive documentation component maps cleanly to plain Markdown.

So this:

```plaintext
MDX source 
↓ 
Human-facing page
```

does not automatically mean:

```plaintext
MDX source 
↓ 
Clean Markdown
```

We needed a serialization layer.

**What is Markdown serialization?**

In this context, serialization means taking a richer content representation and producing another representation that preserves the information needed by a different consumer.

It included:

*   Markdown counterparts for documentation pages
    
*   MDX-to-Markdown serialization
    
*   Generated content handling
    
*   Preservation of useful structure
    
*   Removal of presentation concerns that were not useful to the machine consumer
    

This became especially important for pages containing components that do not naturally translate into plain Markdown.

The goal was not to create a second version of the documentation that writers had to maintain manually. The goal was to make the published documentation system capable of producing a machine-friendly representation from the same underlying content and delivering it appropriately to different consumers.

## **The part I got wrong initially**

This was one of the more useful mistakes in the project.

My first instinct was to make existing components more Markdown-friendly by rewriting some of the content itself.

For example, an accordion component might work very well for a human reader:

```plaintext
▶ Advanced configuration
```

But flattening that interaction into Markdown can produce something closer to

```plaintext
## Advanced configuration

Long section of content...

## Another configuration

Another long section...
```

The component may be better for human navigation, while the flattened representation can become substantially larger.

That created a second problem.

**Trying to make every component directly "Markdown-friendly" could produce unnecessarily large machine-facing pages.**

So the question changed from How do I rewrite the content so the component becomes Markdown? to **How do I make the documentation system produce an appropriate representation for another consumer?**

That led toward serialization and content-delivery changes, which was a much better systems-level solution than rewriting the underlying content every time.

### **5\. Content structure**

Making Markdown available is not enough. The Markdown still needs to be usable.

This is where **content structure and format** become even more important.

Things such as

*   Headings
    
*   Code blocks
    
*   Tabs
    
*   Lists
    
*   Section boundaries
    
*   Content start position
    

affect how a machine can interpret the resulting document.

AFDocs currently has three checks in its content-structure category, including tabbed-content serialization, section-header quality, and Markdown code-fence validity.

The current Rootstock result reports:

*   No problematic tabbed content in the sampled pages
    
*   All sampled code fences properly closed
    
*   Content beginning within the first 10% of sampled pages.
    

This is why I would distinguish **content structure** from **documentation structure**.

Content structure is about how information is represented **inside a page**. Documentation structure is about how pages relate to one another **across the site**. Both of these are important.

### **6\. Generated indexes and documentation hubs**

Generated documentation systems often have different types of pages:

*   Content pages
    
*   Category pages
    
*   Indexes
    
*   Generated references
    
*   API documentation
    
*   Navigation hubs
    

Some of these are obvious to a human browsing the site but can be missing or incomplete in a machine-facing representation. Another lesson was that the generated documentation structure could not always be assumed to be complete simply because the build succeeded. A build succeeding does not automatically mean every one of these has an appropriate machine-facing representation.

That is why generated-index coverage became part of the work. A category page, such as the structure below:

```plaintext
Developers
├── Smart Contracts
├── Integrate
├── Libraries
└── RPC API
```

For a human, this is navigation. For a machine, it can also describe relationships between pieces of knowledge.

***This is one of the places where traditional information architecture and agent readiness meet. The architecture is not suddenly irrelevant because machines are consuming the documentation. It becomes more important.***

### 7\. Page size and restructuring

When a machine consumes documentation, the size of the representation matters.

Very large pages can introduce truncation or make it harder for a consumer to process the relevant information efficiently. AFDocs therefore has a dedicated Page Size and Truncation Risk category with four checks.

This became particularly relevant for large documentation pages. Rather than treating a very large API reference as one indivisible document, we looked at:

*   How the information was structured
    
*   How large the generated Markdown became
    
*   Where the content could be split
    
*   Whether individual representations remained useful
    

That resulted in restructuring and splitting oversized documentation pages where necessary.

#### **The same principle applied to** `llms.txt`

The index itself also has a size constraint. Our implementation limited per-entry descriptions:

```js
/** Truncate per-entry descriptions to help keep llms.txt
 * under Fern's 50,000-character preferred agent-ingest budget.
 */
const LLMS_DESCRIPTION_MAX_LENGTH = 38;
```

This might look like a small piece of code, but it illustrates an important principle.

***Machine-readable documentation has constraints of its own.***

The content can be correct and still be poorly delivered if the representation becomes too large or noisy.

The current [Rootstock Agent Score](https://buildwithfern.com/agent-score/company/rootstock) reports the resulting llms.txt at **49,412 characters**, under the 50,000-character threshold used by the check. For more context, read the article on [How to Measure Agent Web Traffic](https://dacharycarey.com/2026/03/05/how-to-measure-agent-web-traffic/) by Dachary Carey.

### 8\. Documentation structure and delivery configuration

This is the part that is easy to miss when looking only at the generated output.

The agent-facing documentation did not happen through a single file. It depended on the documentation build configuration.

A simplified version of the relevant configuration looked like this:

```plaintext
'docusaurus-plugin-llms',
{
  title: 'Rootstock Developers Portal',
  description:
    'EVM-compatible Bitcoin sidechain docs for builders: RPC, contracts, bridges, and tooling.',
  docsDir: 'docs',
  generateLLMsTxt: true,
  generateLLMsFullTxt: true,
},
```

The configuration also included an additional processing layer for internationalised documentation:

```plaintext
[
  './plugins/llms-i18n.js',
  {
    title: 'Rootstock Developers Portal',
    description:
      'EVM-compatible Bitcoin sidechain docs for builders: RPC, contracts, bridges, and tooling.'
    excludeImports: true,
    removeDuplicateHeadings: true,
    pathTransformation: {
      ignorePaths: ['docs']
    }
  }
]
```

The documentation build becomes responsible for producing multiple machine-facing artifacts and representations rather than treating the browser page as the only output. That is a documentation-engineering decision.

### 9\. CI Validation

This was one of the parts of the work I found most important. It is one thing to fix a documentation system once. It is another thing to make sure the same problems do not quietly return and also survive future changes.

**CI, or continuous integration, means automated checks run as part of the software development workflow.**

For documentation, CI can turn assumptions into executable checks.

For example:

*   Does llms.txt exist?
    
*   Does it contain valid URLs?
    
*   Is it within the expected size?
    
*   Did the generated agent-facing artifacts actually get produced?
    

The verification script included checks such as

```js
const requiredFiles = ['llms.txt', 'llms-full.txt'];

/** Fern Agent Score preferred llms.txt character budget. */
const MAX_LLMS_TXT_CHARS = 50000;

let failed = false;

console.log('Verifying LLM and markdown build artifacts...\n');

if (file === 'llms.txt') {
  const chars = fs.readFileSync(filePath, 'utf8').length;
  const sizeOk = chars < MAX_LLMS_TXT_CHARS;

  if (!sizeOk) failed = true;

  console.log(
    `${sizeOk ? '✓' : '✗ OVER LIMIT'} ${file} size ${chars} chars`
    + ` (limit ${MAX_LLMS_TXT_CHARS})`
  );
}
```

There are two things I like about this approach.

First, the check is deterministic. If the file is missing, the build fails. If a URL check fails, the build fails.

Second, it happens before someone has to discover the problem manually. If a future change causes llms.txt to exceed the defined limit, the build can tell us immediately. That is a meaningful change in how documentation is maintained.

## **What This Means for Technical Writers and Documentarians**

The AI era does not make the technical writer less important. In many ways, it makes the role more demanding.

As documentation becomes increasingly consumed by both humans and machines, the writer is no longer responsible only for producing clear content. The writer also needs to understand the systems through which that content is discovered, delivered, represented, validated and maintained.

A page can be well written and still fail its users if an agent cannot discover it, cannot retrieve the relevant content, receives a truncated representation, or is served information that is no longer current.

That makes **content maintenance and freshness part of the documentation system, not an afterthought**.

This matters because an agent may become another path through which users experience your product. If the underlying documentation is outdated, inconsistent, or difficult to consume, the problem does not necessarily stop with the documentation itself. An agent can surface that outdated information directly to a developer who is trying to build with your product. That can create confusion, failed implementation attempts, loss of trust, or a poor first impression of the product.

From my perspective, this means the technical writer's role is shifting.

We still need writers who can research, explain complex systems, structure information, and communicate clearly. But increasingly, we also need writers who understand the documentation systems around their work: how content is published, how it is made available to machines, how changes are validated, how regressions are detected, and how content stays current over time.

The goal is to make it **maintainable, observable and trustworthy enough to remain useful as the product changes**.

That is where I think the role becomes even more valuable in an AI-assisted development environment. Technical writers are not simply producing content for another generation of readers. We are helping design and maintain the knowledge layer through which both humans and machines understand a product.

## **Conclusion**

The practical lesson I took from this project is simple **Agent-ready documentation is not one feature.** It is the result of several things working together:

**Discoverability + representation + structure + delivery + validation + maintenance**

*   `llms.txt` can provide a discovery path.
    
*   Markdown can provide a cleaner representation.
    
*   Content negotiation can provide the right representation through HTTP.
    
*   Serialization can translate richer documentation components into machine-readable output.
    
*   Information architecture can give both humans and machines a clearer map.
    
*   Page-size controls can reduce truncation risk.
    
*   CI can catch deterministic regressions before they reach production.
    

And benchmarking can give the team a measurable way to inspect the result.

None of these replaces good technical writing. They extend the system around it. That is the part of documentation engineering I became much more interested in through this work.

## **Resources**

Here are some useful resources

*   [Build with Fern](https://buildwithfern.com/agent-score/company/rootstock)
    
*   [AFDocs](https://afdocs.dev/)
    
*   [How to measure agent web traffic](https://dacharycarey.com/2026/03/05/how-to-measure-agent-web-traffic/)
    
*   [Docs Agent Readiness Toolkit](https://github.com/owans/docs-agent-readiness-toolkit)
    
*   [From 71 to 100%: What It Took to Make the Rootstock Developer Portal AI-readyore Agent-Ready Documentation System](https://blog.theowanateamachree.com/from-71-to-100-rootstock-developer-portal-ai-ready )
