45 of 132, and what the score was worth
In September this fleet was audited against an enterprise rubric: every repository, every pipeline, the live state of both nodes, the firewalls, the storage box, the smart-home hub and every published site. It scored 45 points out of 132, with a median criterion score of 1 and five of nine critical rows failing. This note is about the method, the cost — seventy-two and a half million tokens measured across 358 agent runs, plus two passes whose accounting is missing entirely — and about the part worth more than the number: what the audit got wrong about itself. One in eight findings the first pass wrote down did not survive being re-read. Of 86 critical findings re-checked against live state, one was simply false. And the headline verdict was an artefact of the rubric, which asks for a second reviewer and an on-call rota that a one-person household fails by construction. As of 2026-09-08.
Everything else on this site is written from the inside. This one was written by something that was told to assume none of it was true.

What was graded. Everything in this note is about the machines in this photograph and the repositories that describe them.
What was read
Seven repositories, plus the plugin marketplace and the workstation’s own configuration. The git history of each. Every continuous-integration file. All sixty-five alert rules. Every published site.
Then the running estate, read-only: both Proxmox nodes, the firewall pair, the storage box, the smart-home hub, the metrics and log stores, and the forge. Where a claim in a file could be settled against a live API, it was settled against the live API — the audit’s standing instruction was that a document is a claim, not evidence.
How it ran
Each stage is a separate run of agents, with what it cost underneath. Two later passes are in the cost table and not drawn here.
Before any auditing, two inputs were built. The vendor’s own documentation per technology area was collected into 14 research checklists, so that “best practice” meant a fetched page rather than a memory. And 264 of this household’s own prompts — every time the operator had corrected something — were mined into 63 review lenses, then into 14 conduct rules and 57 standing rulings, and finally into a reviewer agent. The estate was graded partly against its own owner’s past objections.
The rubric itself scores 33 criteria from 0 to 4, so the ceiling is 132.
The audit proper produced 109 verified reports: 38 on components, 36 on documents, 17 through the lenses, 8 on git history, 4 on plan drift and 6 on the guard and lint rules. A further 11 domain consolidations sit on top of those. Underneath sits a register of 2,415 records, 2,154 of them canonical after duplicates were folded together.
What did not survive
Every one of the 109 reports was re-read by a different pass, and the result was tallied by one rule: inside each section, one heading or one table row is one item.
2,269 items went into verification. 275 came out as not real.
Then the 86 findings the register held at critical severity were re-checked individually against code and live state. Sixty-seven read confirmed, twelve confirmed but weaker than first written, and six belonged at a different severity altogether. One was false: the audit had recorded the vault restore as never having run, and it had succeeded on 2026-08-20 — a date already the backups note. The job number went into the record in place of the finding.
The audit also filed findings against its own inputs. A pattern search written with an unescaped alternation character, which had carried one recommendation up to high severity. A report graded against a repository head already two commits stale. Checks that were handed back as “run this” instead of being run. And a low-severity bucket appearing in at least 14 of the 36 documentation inputs, in defiance of the standing rule — mined from the operator’s own corrections, and adopted by the audit — that nothing is filed as minor.
The last check was aimed at the report’s own language. A search for the vocabulary of deferral — minor, nit, later, non-blocking, out of scope, consider, might — was run against the finished text. It returned 76 hits on 42 lines, and every one of them is accounted for in the report by location: a quoted line of source code that itself defers, a semantic-version term, a measured interval, or the record of the search itself. A softening check is only evidence if its result is read rather than reported.
And what could not be verified is a list rather than a silence. Console readings the read-only tooling cannot reach. Values the safety classifier refuses to hand over. Operator-only surfaces: who holds which key, which account has two-factor turned on, who performed a particular action on a particular day. And the mutations that would settle a claim outright but would change the estate in order to do it. Twenty-one questions in the critical set stay unanswerable from that vantage point, and each is recorded with the command that would settle it and whose job it is to run it.
The score, and why it was the least useful thing
The number on the left is real. It is also close to the only number that rubric could have produced here.
45 of 132. Median 1. Five of nine critical checklist rows failing. Not enterprise grade, and the report says so in its first line.
The rubric’s own passing line is that no criterion sits below 2, the median reaches 3, and no critical row fails. The first of those three is unreachable with one operator, and the report admits it in the same section that applies it. So the verdict was decided before the first file was read. Worse, it inflates the count: a bus factor of one, no second reviewer and no on-call rota are all scored as defects, and none of them is a defect. They are a description of a household.
One figure made that plain. The audit counted merge requests merged by the account that opened them and read the ratio as an absence of review. The count is accurate and useless: there is one account here, and it cannot distinguish the person from the agent working on their behalf. What it measures is how many people live in the flat.
A separate review followed, over the report and the live estate both, and named the bar that should have been used. Two questions. Can an outsider get in without a credential? If a piece dies, can it be rebuilt from git, the vault and an off-site copy? Measured that way the design is sound and the holes are countable — the review counts four, which is a different sentence from 45 of 132, and not a comfortable one either. The findings underneath the framing mostly held as well: token scopes, runner mounts, log contents and firewall rules were re-checked one by one and stood up in every case the reviewer went back to. What the audit got wrong was grading, ownership and framing. Not facts.
What it cost
Measured from the runtime’s own per-run accounting.
| Pass | Tokens | Tool calls | Agent runs |
|---|---|---|---|
| Component, lens, document, history and rules | 36,290,808 | 11,200 | 98 |
| Verification of all 109 reports | 27,460,793 | 8,106 | 218 |
| Findings register and inventory | 3,872,917 | 388 | 13 |
| Re-check of the 86 critical findings | 3,326,489 | 1,164 | 23 |
| Independent triple double-check | 1,369,950 | 471 | 5 |
| Editorial pass | 230,887 | 51 | 1 |
| Measured total | 72,551,844 | 21,380 | 358 |
Two further runs report no figure and are not in the total: the foundation pass that built the checklists, the lenses and the rubric, and one verification pass stopped part-way through to add a lighter mode. The total is a floor.
What is deliberately not here
Twelve findings are canonical criticals. Three have since been closed; nine have not. What they are, which machine they sit on and what closes them are not on this page, and will not be while they are open — the same goes for the ranked list of ways in. This paragraph exists so that the omission is on the record rather than hidden behind a cheerful summary.
What it changed, and what it did not
As of 2026-09-08, twenty-seven findings are closed. Sixty-one have a change written and waiting. 2,327 have nothing written at all.
The distance between those three numbers is the honest result. A merged fix is not an applied one, an applied fix is not a verified one, and none of the three is a closed finding. Clearing out the stale merge requests on 2026-09-08 returned 77 records to open in a single pass, because the change had been written and never landed.
Which is the same lesson the rest of this site keeps arriving at from the other direction. A thing is not done because it was decided, or written, or merged. It is done when it has been seen working.
Method, coverage, cost and the re-verification results read from the audit’s own generated figures and report on 2026-09-08; the closed and open counts from the findings register on the same date. The findings themselves are not published.