Method v6 Frozen build-system benchmark

What does a docs build cost when the corpus gets real?

Nift and Astro were measured against the same frozen Cloudflare Docs source: thousands of routes, strict output audits, and resource accounting at the cgroup boundary.

8,986audited endpoints
164visual comparisons
2browser engines
0swap bytes

01 / Before speed

Parity has to be earned.

A quick build is irrelevant if it produces the wrong site. The frozen corpus is checked at several layers before timing data can qualify.

01

Structure

Generated HTML and expected route families are compared across the corpus.

02

Metadata

Document-level metadata is audited rather than inferred from page count.

03

Accessibility

Automated checks form part of corpus qualification, alongside semantic review.

04

Routes & leakage

8,986 endpoints are audited for expected routing and unintended source leakage.

02 / Formal clean result

Three valid pairs. One bounded conclusion.

The completed clean series reports all raw runs and descriptive statistics. Nift incremental workflows and the final three Astro incremental invocations are reported separately below.

Formal clean series / final
n=3 valid pairs / ABB order
N

Nift

4.5.0

5.18seconds median

708,198,400 bytes (675.39 MiB) median aggregate peak

A

Astro

7.3.2

570.96seconds median

11,845,435,392 bytes (11.03 GiB) median aggregate peak

110.2239xwall-time median ratio

16.7262xaggregate peak median ratio

Timing scope: tool-native generated-page work. Nift static shell/assets were pre-staged; Astro copied public assets into dist during timing. This is not an equal empty-output complete-site assembly measurement.

Raw values and descriptive statistics

Population standard deviation; seconds and bytes are reported without unit conversion. All three runs are included.

Formal clean benchmark raw values and statistics, three valid runs
MeasureStatisticNiftAstro
Wall secondsRaw, ABB order5.29 / 5.18 / 5.03567.49 / 571.48 / 570.96
Median5.18570.96
Mean5.1666666667569.9766666667
Min / max5.03 / 5.29567.49 / 571.48
Population SD0.10656244911.7711076258
Aggregate peak bytesRaw, ABB order713,728,000 / 708,198,400 / 700,067,84011,845,435,392 / 12,016,066,560 / 11,753,394,176
Median708,198,40011,845,435,392
Mean707,331,413.33311,871,632,042.667
Min / max700,067,840 / 713,728,00011,753,394,176 / 12,016,066,560
Population SD5,610,332.267108,823,691.381

n=3 valid pairs. Fixed ABB first-tool order. Wall-time CV was 2.06% for Nift and 0.31% for Astro, below methodology v6's 15% extension threshold. No valid runs were discarded.

Correctness status

Exact and normalized evidence, named separately.

Nift

Byte-content exact against the certified manifest for every formal run.

Astro

Normalized-equivalent across 21,610 entries / 12,447 files, with zero rejected or unclassified differences.

Common normalization digest b2261e6a694f26653cec8871a98571b93e6eda5452ba6fbf148a6d727e9b389d

The narrow normalization classes cover build-generated random IDs, nine sampled-prompt pages, one Shiki rule-order difference, one sitemap fallback timestamp, and content-identical alternate SVG names. Raw artifacts remain published for independent review; normalized correctness is not described as byte identity.

Qualification, not formal headline

Harness checkpoints

These runs tested correctness and repeatability before the paired formal sequence. They are context, not substitutes for that sequence.

Benchmark qualification measurements
CheckpointElapsedAggregate cgroup peakNote
Nift -14.585 s median516,749,312 B median8 runs
Astro clean reference571.37 s11,359,162,368 BFirst correctness-valid
Astro no-change626.52 s13,389,099,008 BReported 9,022 pages built

03 / Workflows under test

One corpus, six different questions.

Clean speed is only one operating condition. The formal suite is designed around the edits maintainers actually make.

A

Clean build

Generate the production corpus from clean benchmark state.

Formal series completen=3 valid pairs
B

No-change build

nift build completed without touching output. Astro's three final incremental commands each rebuilt all 9,022 pages; two passed normalized correctness and one failed closed on changed Shiki styling.

Campaign completeAstro: 2 of 3 formal-valid
C

One-page, normal workflow

A normal nift build rebuilt the edited page in 0.94 seconds median. This includes graph inspection; explicit targeting remains materially faster.

Nift series completen=3 / 1 page each
D

Explicit page target

nift build workers/get-started/guide/ rebuilt exactly workers/get-started/guide/index.html. Astro has no equivalent explicit named-route production build command in this project/toolchain, so no counterpart or speed ratio is reported.

Formal series completen=30 valid Nift runs
E

Five-page batch

Normal graph discovery rebuilt five pages in 0.98 seconds median. Explicitly naming all five targets completed in 0.16 seconds median across 20 valid runs.

Nift series completen=3 normal / n=20 target
F

Shared-template fan-out

An inert metadata marker in the shared head template caused Nift's normal incremental path to rebuild all 8,803 pages in 4.89 seconds median.

Nift series completen=3 / 8,803 pages each

Nift incremental workflows

Graph discovery costs about a second. Explicit targets cost a fraction of it.

Normal runs check the dependency graph. Targeted runs receive exact tracked names. These are distinct workflows, not interchangeable labels, and no Astro targeted counterpart is invented.

Normal / one edit0.94 s

10,129,408 B median aggregate peak. Three runs rebuilt exactly one designated page.

Targeted / one edit0.15 s

7,280,640 B median aggregate peak. Thirty runs changed exactly one output.

Normal / five edits0.98 s

10,563,584 B median aggregate peak. Three runs rebuilt the five designated pages.

Targeted / five edits0.16 s

8,681,472 B median aggregate peak. Twenty runs changed exactly five outputs.

Normal / shared template4.89 s

444,751,872 B median aggregate peak. Three runs reported all 8,803 pages rebuilt.

Nift incremental workflow descriptive statistics
WorkflowRunsWall medianWall rangePeak medianValidated outputs
One-page normal30.94 s0.89–0.95 s10,129,408 B1 each run
One-page targeted300.15 s0.14–0.25 s7,280,640 B1 each run
Five-page normal30.98 s0.94–0.99 s10,563,584 B5 each run
Five-page targeted200.16 s0.15–0.21 s8,681,472 B5 each run
Shared-template normal34.89 s4.73–4.96 s444,751,872 B5 representatives

Validation scope: the lean normal-workflow runs checked marker presence and changed hashes only for the intended outputs, as declared before measurement. They did not rescan or serialize the full site. The shared-template series checked five frozen representative outputs and Nift's reported fan-out count.

04 / Resource envelope

Memory is a suitability constraint.

Aggregate cgroup peak captures the workload boundary, not just one convenient process. That matters when child processes and toolchains share a build.

05 / Dependency footprint

Count what must be trusted, installed, and kept current.

Dependency surface is part of operational cost. This inventory separates declared packages, resolved pnpm entries, generated caches, filesystem allocation, and output so unlike categories are not collapsed into one number.

N

Nift toolchain

Nift 4.5.0 is a native executable. The benchmark project contains no node_modules directory and requires no Node package graph for its build.

Node dependencies
0
node_modules
absent
Nift binary
3,497,360 B
Source checkout
255,100,730 B
Build metadata
49,979,498 B
Generated output
870,706,566 B
A

Astro toolchain

Astro 7.3.2 runs through Node 24.21.0 and pnpm 12.4.2. Counts below distinguish direct declarations from resolved package identities.

Direct dev dependencies
105
pnpm store entries
1,120
Lockfile snapshots
1,327
node_modules
1,194,890,090 B
Installed regular files
55,232
Generated .astro cache
1,528,202,322 B

Definitions: byte values are logical lstat sizes with hard-linked objects counted once. Astro's node_modules total excludes node_modules/.astro, which is reported separately; it occupies 1,366,536,192 allocated bytes before that cache. Nift and Astro source-checkout figures exclude Git history, dependencies, generated metadata, and output. Generated output size is not dependency size.

06 / Reproducibility

Freeze inputs. Expose uncertainty.

The experiment repository contains the harness and stage branch. Methodology version 6 is the measurement contract for the numbers on this page.

  1. 1

    Pin the source

    Cloudflare Docs is frozen at commit bc2bdaee16098ec1b0bb782b80cf3a73f9557ddf.

  2. 2

    Qualify correctness

    Structural, metadata, accessibility, route, leakage, and sampled visual checks gate timing validity.

  3. 3

    Measure the boundary

    Elapsed time and aggregate cgroup memory are collected on the same 6 vCPU, no-swap machine.

  4. 4

    Pair formal runs

    Three infrastructure-valid pairs used fixed ABB first-tool order. No valid run was discarded and the 15% CV extension threshold was not triggered.

  5. 5

    Audit artifacts

    Nift is certified-manifest exact. Astro's published raw artifacts reduce to one common normalization digest with zero rejected or unclassified differences.

07 / Limitations

What the evidence does not yet say.

01

The formal clean result has three pairs. It reports dispersion, but n=3 remains a small sample and does not characterize every host.

02

Timing covers tool-native generated-page work, not equal empty-output assembly: Nift assets were pre-staged while Astro copied public assets.

03

Astro correctness is normalized-equivalent, not byte-exact. The narrow classes and raw artifacts remain available for scrutiny.

04

Machine-specific timing should not be generalized to every host, filesystem, container runtime, or CI provider.

05

This measures a static documentation workload, not framework quality across applications, islands, SSR, or integrations.

06

Nift normal edit runs use a deliberately narrow contract: intended outputs are marker-checked and hashed, while the rest of the site is not rescanned. Astro's final incremental campaign produced two formal-valid runs out of three, so no no-change speed ratio is claimed.

08 / Engineering interpretation

The useful question is fit, not fandom.

The formal clean gap is substantial. It is not permission to flatten two different tools into a universal ranking.

02

Page count is not the explanation

8,803 generated HTML pages describe scale, not mechanism. Scaling depends on how much work each page triggers: module evaluation, content loading, transforms, rendering, serialization, and process coordination. A large corpus amplifies those choices; it does not explain them alone.

03

Rebuild model beats a clean-build trophy

Maintainers more often change one guide or one shared template than clone from zero. No-change, one-page, five-page, and fan-out scenarios reveal whether the system can limit work correctly. Those results may matter more to daily engineering than the clean headline.

04

Targeting should follow dependencies

An explicit nift build <name> target is useful because tracked relationships still define what must rebuild. Targeting without dependency knowledge can be dangerously fast; the completed formal series shows dependency-aware targeting changing exactly the intended output across all 30 runs.

05

Resources define suitability

Memory affects whether builds fit on modest CI runners, coexist with other jobs, or fail under pressure. Time affects iteration and deployment recovery. Neither metric proves output quality, but both are architecture inputs, not cosmetic optimizations.

06

Dependencies have carrying cost

Here, Nift needs zero Node packages and one 3.50 MB native executable. The Astro checkout declares 105 direct dev dependencies, resolves 1,120 pnpm store entries, and installs 1.19 GB of logical node_modules data before its 1.53 GB generated .astro cache. That ecosystem buys integrations and familiarity, but also expands cache weight, update churn, and supply-chain review.

07

A small mental model compounds

Nift's core map is short: tracked pages, content, templates, and dependencies. That can reduce context switching for people and agents. Astro's conventions can provide a different kind of leverage to teams already fluent in its component model. Local expertise changes the equation.

08

Cheap checks benefit people and agents

The maintenance loop is edit, build, inspect, validate, decide, repeat. Nift evidence now spans a 0.15-second targeted page, a 0.16-second targeted five-page batch, normal graph-aware edits under one second, and a 0.90-second no-change check. Everyone still requires visual, accessibility, and behavioral acceptance tests. Cheap checks improve the economics of verifying continuously; they do not remove the need to check.

Invitation, not conclusion

Audit the harness. Challenge the classifications. Re-run the pairs.

The raw experiment is public precisely so that surprising results can be inspected rather than converted into marketing folklore.

Inspect on GitHub

AI-DX + human DX / qualitative

Which architecture would I rather maintain?

AI-DX is the developer experience of a coding agent. Agent-mediated human DX asks a related question: if an agent can own routine integration mechanics, how much complexity must the human personally carry? This is my assessment after working extensively across both implementations. It is opinion, not benchmark evidence.

Primary qualitative conclusion

For this mostly static workload, I would choose Nift for both agent DX and human DX when capable coding agents are part of the development workflow. Its explicit build behavior, targetable production commands, dependency-aware rebuilds, low resource footprint, ordinary tool boundaries, and cheap production verification make work unusually easy to inspect against the artifact that will ship.

With capable coding agents handling most integration plumbing, I would choose Nift for the human developer experience too. Humans can retain MDX, Vite, framework islands, browser HMR, and familiar frontend tooling without personally carrying most of the surrounding configuration and glue. In return, they share the same cheap production feedback loop: bounded edits can be built and checked without leaving the task.

Agent DXNift preferred
Human DX with capable agentsNift preferred
Manual work in these exact repositoriesAstro preferred for authoring
Deeply integrated component workAstro retains conveniences
01 / Agent-mediated human DX

The human need not maintain every seam by hand.

Agent handles: MDX and package integration, Vite configuration, island wiring, bundle entries, asset references, integration tests, package synchronization, and routine configuration maintenance.

Human keeps: familiar content authoring, React, Vue, or other frontend tools where useful, browser and HMR workflows for islands, plus architecture, component boundaries, UX, accessibility, security, and semantic decisions.

Both gain: Nift's cheap, targetable production build and verification loop. Complexity does not vanish, but the human does not need to carry every mechanical seam personally.

02 / Composition today

MDX and islands do not need to live in Nift core.

A Nift project can already use a package or external MDX compiler, compose the rendered result with normal templates, emit mount points and checked assets, and let Vite plus React, Preact, Vue, Svelte, Solid, or another browser framework own selected islands.

The practical model is selective integration: humans keep the frontend capabilities they need, agents maintain much of the plumbing, and Nift remains the outer production composition layer.

03 / Astro's remaining advantage

Integrated conventions still have real value.

Astro internalizes MDX, component rendering, Vite, hydration conventions, TypeScript relationships, image handling, content collections, one dev server, and framework-aware HMR. That is a smoother environment for intensive, tightly coupled frontend and component development.

Nift exposes more seams, even when an agent maintains them. Astro's conventions reduce integration risk and provide mature ecosystem behavior; Nift's countervailing advantage is that mostly static pages need not pay the production-generation cost of the complete integrated stack.

04 / Exact-repository caveat

If manual authoring is restricted to these two trees, I pick Astro.

Cloudflare's Astro source is readable MDX with semantically named components. This Nift implementation is a compatibility reproduction containing numeric body fragments, generated HTML, importer rules, and a mixed Markdown/HTML pipeline.

If a human must work manually without agent assistance in these repositories exactly as they exist, Astro is the nicer authoring source. That is a limitation of the reproduction, not a requirement of Nift's architecture, and it does not reverse the broader agent or human-with-agent preference.

01 / Measured today

Production feedback

Nift's retained production medians range from 0.15 seconds for one explicit page to 4.89 seconds for an 8,803-page template fan-out. Astro's three final incremental commands took 680.27, 688.26, and 688.10 seconds, each reporting all 9,022 pages built; only two were formal-valid.

02 / Capability today

Composable tooling

Optional MDX tooling, Vite, and framework islands can already compose with a Nift project. Astro's advantage is integrated convention, typing, dev-server behavior, and HMR, not exclusive access to those technologies.

03 / Proposed improvement

First-class source tracking

A future tracked-source API could let Nift own processor identity, imported dependencies, stale selection, diagnostics, and opaque transformed content. It would improve an already viable workflow rather than unlock MDX or frontend composition.

Current composition + future integration

today: optional MDX toolingtoday: Nift production generationtoday: selective Vite islandstoday: agent-maintained seamsfuture: tracked-source API

Nift can already be the outer website layer while an MDX package renders authored content and Vite plus a browser framework own selected islands. A future tracked-source API would bring processor identity, imported dependencies, and incremental orchestration into Nift's graph more seamlessly. The current 0.15-second result does not include optional MDX compilation, a persistent worker, source maps, or Vite bundling, so no end-to-end package workflow latency is claimed here.

If a documentation build can know exactly what changed, how much of the stack really needs to wake up?