Structure
Generated HTML and expected route families are compared across the corpus.
Method v6 Frozen build-system benchmark
Nift and Astro were measured against the same frozen Cloudflare Docs source: thousands of routes, strict output audits, and resource accounting at the cgroup boundary.
01 / Before speed
A quick build is irrelevant if it produces the wrong site. The frozen corpus is checked at several layers before timing data can qualify.
Generated HTML and expected route families are compared across the corpus.
Document-level metadata is audited rather than inferred from page count.
Automated checks form part of corpus qualification, alongside semantic review.
8,986 endpoints are audited for expected routing and unintended source leakage.
02 / Formal clean result
The completed clean series reports all raw runs and descriptive statistics. Nift incremental workflows and the final three Astro incremental invocations are reported separately below.
5.18seconds median
708,198,400 bytes (675.39 MiB) median aggregate peak
570.96seconds median
11,845,435,392 bytes (11.03 GiB) median aggregate peak
110.2239xwall-time median ratio
16.7262xaggregate peak median ratio
Timing scope: tool-native generated-page work. Nift static shell/assets were pre-staged; Astro copied public assets into dist during timing. This is not an equal empty-output complete-site assembly measurement.
Raw values and descriptive statistics
Population standard deviation; seconds and bytes are reported without unit conversion. All three runs are included.
| Measure | Statistic | Nift | Astro |
|---|---|---|---|
| Wall seconds | Raw, ABB order | 5.29 / 5.18 / 5.03 | 567.49 / 571.48 / 570.96 |
| Median | 5.18 | 570.96 | |
| Mean | 5.1666666667 | 569.9766666667 | |
| Min / max | 5.03 / 5.29 | 567.49 / 571.48 | |
| Population SD | 0.1065624491 | 1.7711076258 | |
| Aggregate peak bytes | Raw, ABB order | 713,728,000 / 708,198,400 / 700,067,840 | 11,845,435,392 / 12,016,066,560 / 11,753,394,176 |
| Median | 708,198,400 | 11,845,435,392 | |
| Mean | 707,331,413.333 | 11,871,632,042.667 | |
| Min / max | 700,067,840 / 713,728,000 | 11,753,394,176 / 12,016,066,560 | |
| Population SD | 5,610,332.267 | 108,823,691.381 |
n=3 valid pairs. Fixed ABB first-tool order. Wall-time CV was 2.06% for Nift and 0.31% for Astro, below methodology v6's 15% extension threshold. No valid runs were discarded.
Correctness status
Byte-content exact against the certified manifest for every formal run.
Normalized-equivalent across 21,610 entries / 12,447 files, with zero rejected or unclassified differences.
Common normalization digest b2261e6a694f26653cec8871a98571b93e6eda5452ba6fbf148a6d727e9b389d
The narrow normalization classes cover build-generated random IDs, nine sampled-prompt pages, one Shiki rule-order difference, one sitemap fallback timestamp, and content-identical alternate SVG names. Raw artifacts remain published for independent review; normalized correctness is not described as byte identity.
Qualification, not formal headline
These runs tested correctness and repeatability before the paired formal sequence. They are context, not substitutes for that sequence.
| Checkpoint | Elapsed | Aggregate cgroup peak | Note |
|---|---|---|---|
Nift -1 | 4.585 s median | 516,749,312 B median | 8 runs |
| Astro clean reference | 571.37 s | 11,359,162,368 B | First correctness-valid |
| Astro no-change | 626.52 s | 13,389,099,008 B | Reported 9,022 pages built |
03 / Workflows under test
Clean speed is only one operating condition. The formal suite is designed around the edits maintainers actually make.
Generate the production corpus from clean benchmark state.
nift build completed without touching output. Astro's three final incremental commands each rebuilt all 9,022 pages; two passed normalized correctness and one failed closed on changed Shiki styling.
A normal nift build rebuilt the edited page in 0.94 seconds median. This includes graph inspection; explicit targeting remains materially faster.
nift build workers/get-started/guide/ rebuilt exactly workers/get-started/guide/index.html. Astro has no equivalent explicit named-route production build command in this project/toolchain, so no counterpart or speed ratio is reported.
Normal graph discovery rebuilt five pages in 0.98 seconds median. Explicitly naming all five targets completed in 0.16 seconds median across 20 valid runs.
An inert metadata marker in the shared head template caused Nift's normal incremental path to rebuild all 8,803 pages in 4.89 seconds median.
Nift incremental workflows
Normal runs check the dependency graph. Targeted runs receive exact tracked names. These are distinct workflows, not interchangeable labels, and no Astro targeted counterpart is invented.
10,129,408 B median aggregate peak. Three runs rebuilt exactly one designated page.
7,280,640 B median aggregate peak. Thirty runs changed exactly one output.
10,563,584 B median aggregate peak. Three runs rebuilt the five designated pages.
8,681,472 B median aggregate peak. Twenty runs changed exactly five outputs.
444,751,872 B median aggregate peak. Three runs reported all 8,803 pages rebuilt.
| Workflow | Runs | Wall median | Wall range | Peak median | Validated outputs |
|---|---|---|---|---|---|
| One-page normal | 3 | 0.94 s | 0.89–0.95 s | 10,129,408 B | 1 each run |
| One-page targeted | 30 | 0.15 s | 0.14–0.25 s | 7,280,640 B | 1 each run |
| Five-page normal | 3 | 0.98 s | 0.94–0.99 s | 10,563,584 B | 5 each run |
| Five-page targeted | 20 | 0.16 s | 0.15–0.21 s | 8,681,472 B | 5 each run |
| Shared-template normal | 3 | 4.89 s | 4.73–4.96 s | 444,751,872 B | 5 representatives |
Validation scope: the lean normal-workflow runs checked marker presence and changed hashes only for the intended outputs, as declared before measurement. They did not rescan or serialize the full site. The shared-template series checked five frozen representative outputs and Nift's reported fan-out count.
04 / Resource envelope
Aggregate cgroup peak captures the workload boundary, not just one convenient process. That matters when child processes and toolchains share a build.
0.708 GB11.845 GBFormal medians, n=3; decimal GB shown for orientation. Ratio: 16.7262x.
05 / Dependency footprint
Dependency surface is part of operational cost. This inventory separates declared packages, resolved pnpm entries, generated caches, filesystem allocation, and output so unlike categories are not collapsed into one number.
Nift 4.5.0 is a native executable. The benchmark project contains no node_modules directory and requires no Node package graph for its build.
node_modulesAstro 7.3.2 runs through Node 24.21.0 and pnpm 12.4.2. Counts below distinguish direct declarations from resolved package identities.
node_modules.astro cacheDefinitions: byte values are logical lstat sizes with hard-linked objects counted once. Astro's node_modules total excludes node_modules/.astro, which is reported separately; it occupies 1,366,536,192 allocated bytes before that cache. Nift and Astro source-checkout figures exclude Git history, dependencies, generated metadata, and output. Generated output size is not dependency size.
06 / Reproducibility
The experiment repository contains the harness and stage branch. Methodology version 6 is the measurement contract for the numbers on this page.
Cloudflare Docs is frozen at commit bc2bdaee16098ec1b0bb782b80cf3a73f9557ddf.
Structural, metadata, accessibility, route, leakage, and sampled visual checks gate timing validity.
Elapsed time and aggregate cgroup memory are collected on the same 6 vCPU, no-swap machine.
Three infrastructure-valid pairs used fixed ABB first-tool order. No valid run was discarded and the 15% CV extension threshold was not triggered.
Nift is certified-manifest exact. Astro's published raw artifacts reduce to one common normalization digest with zero rejected or unclassified differences.
07 / Limitations
The formal clean result has three pairs. It reports dispersion, but n=3 remains a small sample and does not characterize every host.
Timing covers tool-native generated-page work, not equal empty-output assembly: Nift assets were pre-staged while Astro copied public assets.
Astro correctness is normalized-equivalent, not byte-exact. The narrow classes and raw artifacts remain available for scrutiny.
Machine-specific timing should not be generalized to every host, filesystem, container runtime, or CI provider.
This measures a static documentation workload, not framework quality across applications, islands, SSR, or integrations.
Nift normal edit runs use a deliberately narrow contract: intended outputs are marker-checked and hashed, while the rest of the site is not rescanned. Astro's final incremental campaign produced two formal-valid runs out of three, so no no-change speed ratio is claimed.
08 / Engineering interpretation
The formal clean gap is substantial. It is not permission to flatten two different tools into a universal ranking.
01
Cloudflare Docs is a large, mostly static documentation corpus. That shape rewards a builder that can map explicit source dependencies to independent outputs. Astro remains a rational choice when component integration, adapters, and ecosystem familiarity carry more value than a minimal build surface.
02
8,803 generated HTML pages describe scale, not mechanism. Scaling depends on how much work each page triggers: module evaluation, content loading, transforms, rendering, serialization, and process coordination. A large corpus amplifies those choices; it does not explain them alone.
03
Maintainers more often change one guide or one shared template than clone from zero. No-change, one-page, five-page, and fan-out scenarios reveal whether the system can limit work correctly. Those results may matter more to daily engineering than the clean headline.
04
An explicit nift build <name> target is useful because tracked relationships still define what must rebuild. Targeting without dependency knowledge can be dangerously fast; the completed formal series shows dependency-aware targeting changing exactly the intended output across all 30 runs.
05
Memory affects whether builds fit on modest CI runners, coexist with other jobs, or fail under pressure. Time affects iteration and deployment recovery. Neither metric proves output quality, but both are architecture inputs, not cosmetic optimizations.
06
Here, Nift needs zero Node packages and one 3.50 MB native executable. The Astro checkout declares 105 direct dev dependencies, resolves 1,120 pnpm store entries, and installs 1.19 GB of logical node_modules data before its 1.53 GB generated .astro cache. That ecosystem buys integrations and familiarity, but also expands cache weight, update churn, and supply-chain review.
07
Nift's core map is short: tracked pages, content, templates, and dependencies. That can reduce context switching for people and agents. Astro's conventions can provide a different kind of leverage to teams already fluent in its component model. Local expertise changes the equation.
08
The maintenance loop is edit, build, inspect, validate, decide, repeat. Nift evidence now spans a 0.15-second targeted page, a 0.16-second targeted five-page batch, normal graph-aware edits under one second, and a 0.90-second no-change check. Everyone still requires visual, accessibility, and behavioral acceptance tests. Cheap checks improve the economics of verifying continuously; they do not remove the need to check.
Invitation, not conclusion
The raw experiment is public precisely so that surprising results can be inspected rather than converted into marketing folklore.
AI-DX + human DX / qualitative
AI-DX is the developer experience of a coding agent. Agent-mediated human DX asks a related question: if an agent can own routine integration mechanics, how much complexity must the human personally carry? This is my assessment after working extensively across both implementations. It is opinion, not benchmark evidence.
Primary qualitative conclusion
For this mostly static workload, I would choose Nift for both agent DX and human DX when capable coding agents are part of the development workflow. Its explicit build behavior, targetable production commands, dependency-aware rebuilds, low resource footprint, ordinary tool boundaries, and cheap production verification make work unusually easy to inspect against the artifact that will ship.
With capable coding agents handling most integration plumbing, I would choose Nift for the human developer experience too. Humans can retain MDX, Vite, framework islands, browser HMR, and familiar frontend tooling without personally carrying most of the surrounding configuration and glue. In return, they share the same cheap production feedback loop: bounded edits can be built and checked without leaving the task.
Agent handles: MDX and package integration, Vite configuration, island wiring, bundle entries, asset references, integration tests, package synchronization, and routine configuration maintenance.
Human keeps: familiar content authoring, React, Vue, or other frontend tools where useful, browser and HMR workflows for islands, plus architecture, component boundaries, UX, accessibility, security, and semantic decisions.
Both gain: Nift's cheap, targetable production build and verification loop. Complexity does not vanish, but the human does not need to carry every mechanical seam personally.
A Nift project can already use a package or external MDX compiler, compose the rendered result with normal templates, emit mount points and checked assets, and let Vite plus React, Preact, Vue, Svelte, Solid, or another browser framework own selected islands.
The practical model is selective integration: humans keep the frontend capabilities they need, agents maintain much of the plumbing, and Nift remains the outer production composition layer.
Astro internalizes MDX, component rendering, Vite, hydration conventions, TypeScript relationships, image handling, content collections, one dev server, and framework-aware HMR. That is a smoother environment for intensive, tightly coupled frontend and component development.
Nift exposes more seams, even when an agent maintains them. Astro's conventions reduce integration risk and provide mature ecosystem behavior; Nift's countervailing advantage is that mostly static pages need not pay the production-generation cost of the complete integrated stack.
Cloudflare's Astro source is readable MDX with semantically named components. This Nift implementation is a compatibility reproduction containing numeric body fragments, generated HTML, importer rules, and a mixed Markdown/HTML pipeline.
If a human must work manually without agent assistance in these repositories exactly as they exist, Astro is the nicer authoring source. That is a limitation of the reproduction, not a requirement of Nift's architecture, and it does not reverse the broader agent or human-with-agent preference.
Nift's retained production medians range from 0.15 seconds for one explicit page to 4.89 seconds for an 8,803-page template fan-out. Astro's three final incremental commands took 680.27, 688.26, and 688.10 seconds, each reporting all 9,022 pages built; only two were formal-valid.
Optional MDX tooling, Vite, and framework islands can already compose with a Nift project. Astro's advantage is integrated convention, typing, dev-server behavior, and HMR, not exclusive access to those technologies.
A future tracked-source API could let Nift own processor identity, imported dependencies, stale selection, diagnostics, and opaque transformed content. It would improve an already viable workflow rather than unlock MDX or frontend composition.
Current composition + future integration
today: optional MDX toolingtoday: Nift production generationtoday: selective Vite islandstoday: agent-maintained seamsfuture: tracked-source API
Nift can already be the outer website layer while an MDX package renders authored content and Vite plus a browser framework own selected islands. A future tracked-source API would bring processor identity, imported dependencies, and incremental orchestration into Nift's graph more seamlessly. The current 0.15-second result does not include optional MDX compilation, a persistent worker, source maps, or Vite bundling, so no end-to-end package workflow latency is claimed here.
If a documentation build can know exactly what changed, how much of the stack really needs to wake up?