Benchmarks
Rennet measures itself. Every Repo Map build and every lens generation records a timing for each stage it actually ran, and each stage record names the harness and model that executed it. This page renders the exported data — no number on it was typed by hand.
What is measured
Section titled “What is measured”Three pipelines carry stage records:
-
The Repo Map build. Deterministic end to end:
scout,resolve,tree,workspace,conventions,symbols,build,verify,store, plustotal. No model runs in any of them, and the schema forbids one from claiming it.Which of those a run could record depends on who ran it, so every run says. The daemon scouts a repository before it maps it;
rennet maphas no scout pass at all, by design. Each group of runs below names its producer, because aresolverow with noscoutrow beside it otherwise reads as a lost measurement rather than as a stage that was never going to run.totalis the sum of that repository’s own stage durations, not the wall clock from its first stage to its last. A project scouts every repository in one pass and maps them in a later one, so a wall-clock span would charge each repository for every sibling scouted or mapped in between — a repo’stotalwould grow with the size of the project rather than with its own work. Within a pass no stage closes until the next opens, so the sum hides no waiting; the only thing it excludes is other repositories. -
Lens drafting. Per lens: the drafting turn (
lens-draft), the repair ladder (lens-repair), and the deterministic work between the ladder and the accepted write (lens-post-process). A lane with more than one seat records onelens-draftper seat, so a dual review names each provider rather than averaging them into one. -
The round report.
reportis the whole gate — building and measuring the evidence manifest, resolving the seat, the turn, deterministic verification — andreport-classificationis the provider turn inside it.
Beside those sit coverage and reveal, and two latency figures measured from
the reviewer’s own wait origin: first-core-board, the first core lane that
SETTLED, and first-element, the first element any board published on its
element stream. Boards draw themselves as they are written now, so the two answer
different questions and neither is redefined — the first element is what the
reviewer can see, the first settled board is what they can trust, and the gap
between them is the drafting they now watch instead of waiting through. Both name
the lane they belong to. A generation on which no lens ever wrote an element
records no first-element at all; a zero would claim a board drew itself the
instant the generation started.
How the modes are derived
Section titled “How the modes are derived”The Model Council routes per job, so one run can legitimately span providers. A run’s configuration label is therefore derived from its stage records, never read off a setting:
| Derived mode | What it means |
|---|---|
| Dual model (council) | Stages of this run named both Claude and Codex. |
| Claude only | Every attributed stage named Claude. |
| Codex only | Every attributed stage named Codex. |
| No provider stage | No stage named a provider — every Repo Map build, and a generation that failed before a seat resolved. |
No table below merges two modes. Averaging a Claude-only run with a council run would state a number describing no configuration that exists.
Failed and aborted runs are recorded and counted — a pipeline that archived only
its successes would report the fast half of its own latency — and they are never
mixed into a latency figure. Every stage row states the outcome of the run it
came from, and each group’s median run time is over its complete runs only. A
run that fell over halfway measured half a pipeline, and averaging it with one
that finished describes neither. A group where nothing completed states no median
at all rather than a 0 that would read as an instantaneous pipeline.
Medians are the middle value, or the midpoint of the two middle values on an even sample, truncated to whole milliseconds.
Measurements
Section titled “Measurements”Each section below is rendered from docs/data/benchmarks.json. A run kind with no
exported rows says so rather than showing an empty table — lens and report timings
need a dogfood run against a real coding harness, and only measurements that have
actually been recorded and committed appear here.
Measured on Apple M-series, macOS 15 (darwin arm64), 16 GB at revision 12a4bc56b38be76dd2975b66199a889a7bc45172, exported 2026-09-01.
Repo Map build
No provider stage
3 runs — 3 complete, 0 failed, 0 aborted. Median complete run: 15.5 s. Recorded by cli-map.
| Stage | Lens | Outcome | Samples | Median | Slowest |
|---|---|---|---|---|---|
resolve | — | complete | 3 | 50 ms | 52 ms |
tree | — | complete | 3 | 170 ms | 275 ms |
workspace | — | complete | 3 | 162 ms | 179 ms |
conventions | — | complete | 3 | 6 ms | 6 ms |
symbols | — | complete | 3 | 14.0 s | 14.1 s |
build | — | complete | 3 | 279 ms | 281 ms |
verify | — | complete | 3 | 158 ms | 159 ms |
store | — | complete | 3 | 688 ms | 763 ms |
total | — | complete | 3 | 15.5 s | 15.9 s |
Lens generation and round report
No run of this kind has been exported yet, so this page states nothing about it.
Where the numbers come from
Section titled “Where the numbers come from”Recording is on by default and controlled from Settings → Benchmarks, which
also renders the local history with the same per-stage breakdown. Records live in
~/.rennet/benchmarks.jsonl and never leave the machine.
The data on this page is written by the export command:
rennet benchmarks export --out docs/data/benchmarks.jsonThe export is deterministic, including its own timestamp: exportedAt is derived
from the end of the newest recorded run, not from the wall clock, so re-exporting an
unchanged archive is a genuinely empty diff. Pass --timestamp <iso> to state the
stamp explicitly. A fresh clock reading is the last fallback and the only part of
this path that is not reproducible.
The docs build fails if docs/data/benchmarks.json is missing or corrupt, because
a benchmarks page that rendered blanks would be indistinguishable from a very fast
pipeline. That check runs in an Astro integration, not only in the remark plugin that
renders the tables, so it runs on every build regardless of what any cache holds.
Validation cannot catch the opposite case, a valid edit to the data: there is nothing wrong to throw about. The page’s Markdown carries no number — the fence is empty and the plugin reads the file at render time — so its digest does not move when the measurements do, and Astro’s content layer reuses the render it already has. That is not hypothetical: a valid edit once completed a build whose page still carried the previous export’s provenance.
So the same integration does two more things. It drops Astro’s stored renders when the data’s hash changes, so the page is rebuilt; and after the pages are emitted it checks the built HTML, failing the build if any rendered page states an older provenance — or if no page rendered the data at all. The check is what makes the invalidation trustworthy: it reads what was actually written, so a stale page cannot ship whatever the cause.
Provenance is stated above the tables. Stale-but-labeled beats fresh-but-invented, and re-exporting is one command.