Escape log¶
A record of defects that reached production, and — the point of the file — what now stops each one from happening again.
Every entry is written by /postmortem. The command's deliverable is not the
code fix; it is the answer to a single question:
Which gate should have caught this, and why didn't it?
A defect that produced a fix commit and nothing else has taught the system
nothing, and will be paid for again. A defect that produced a new check, a new
fixture in RUNS, or a re-anchored assertion has been converted into
permanent capability.
How to read an entry¶
- Escape class determines the fix:
| Class | Meaning | Fix |
|---|---|---|
gate-not-run |
a catching gate existed but didn't run at that point | move the gate earlier |
path-never-exercised |
the gate ran but no fixture reached the code | build the portfolio, register it in RUNS |
test-shared-the-assumption |
a test covered it and passed, written from the same wrong sentence | re-anchor to a source of truth |
no-assertion-of-presence |
output was absent/null rather than wrong | assert presence |
wrong-premise |
the plan bullet was wrong and was faithfully implemented | strengthen Wave 0 |
no-gate-exists |
nothing could have caught it | create the gate |
ungateable |
not mechanically detectable | .claude/LESSONS.md entry, with reasoning |
caught-and-parked |
a gate fired, and the record of the finding became its resting place | ratchet the finding register; give every parked entry an owning bullet |
caught-and-parked was added on 2026-08-09. It is for the case where a gate did
catch it, said so, and the wrong number shipped anyway. It is a distinct class
because its fix targets the register of tolerated findings rather than the
output — the only one of the eight that does — and because the shape recurs
across at least four parallel registers here (KNOWN_DISAGREEMENTS,
classification_table.toml's [[known_disagreement]], known_broken_rules,
known_vacuous_rules) plus strict xfails and plan bullets. Note that
no-gate-exists would prescribe roughly the right fix, so narrative fidelity
alone is not the argument.
An entry may also carry no class, and say why. A defect in a gate is not a defect that escaped one, and the eight classes all presume the latter — so forcing a class on it would prescribe the wrong fix. Leave the field unclassed with the reasoning, and name the class you would coin if it recurs.
- Verified red records the command and the failure line, confirming the new
gate fails without the fix. A gate nobody has seen fail is not a gate. Along
with the escape class and the gate change, it is what closes the defect: an
entry missing any of the three means the defect is still open, whatever
landed in
src/.
Related¶
.claude/LESSONS.md— the working set of traps every agent reads before starting. Entries graduate out of it into executable checks.scripts/arch_check.py— the numbered architectural invariants. Each one is a lesson that graduated.tests/acceptance/reporting/test_supervisory_validations.py— the two-way-ratcheted register of published EBA/BoE rules; the estate's strongest oracle for reporting defects.
The first six entries were written together on 2026-08-09, from a review of the
validation estate rather than from six separate /postmortem runs. The first four
are the escapes this project had already established with evidence and never
recorded — the file having sat at zero entries while defects reached published
output is itself the first thing the review found. The last two came out of the
review itself, one of them a measured escape and one a defect in a gate that had
not yet shipped.
Their gate changes land in the same change-set as this file, not in an
earlier release, and each was still under adversarial review when its entry was
written. Every Gate change field therefore names the item that owns it; if a
review returns revise, that field and its Verified red are what must be
re-checked, and a red produced against a revised gate is not evidence for the
gate that shipped.
2026-08-09 — Estate coverage is measured to four decimal places and gated on nothing¶
- Defect: The four C 07.00 off-balance-sheet defects recorded in
.claude/LESSONS.mdB5 reached published output because every golden portfolio was 100% drawn loans. No data ever flowed through the off-balance-sheet columns, so the four published rules that tie them out (boe_b0471,v6364_m,v1659_m,v1661_m) were never evaluated, and the supervisory gate — which fails open — was green throughout. The metric that measures exactly this condition was already computed and consulted by nothing. Measured over the current 16-run matrix: 12.85% template-cell liveness, 55,553 dead cells, 785 never-evaluated rules, and only 257 CRR / 289 Basel 3.1 published rules binding. - Rule: Not a regulatory escape. The regulatory content at risk is whatever lives in the 55,553 dead cells, which is the point: an unlit cell has no direction and no magnitude until something lights it.
- Origin:
scripts/coverage_report.pyandscripts/coverage_baseline.json, merged with the independent validation system (2026-08-08).--checkwas written, documented as a ratchet, and never called — not by CI, not byscripts/arch_check.py, not by any test. - Escape class:
gate-not-run - What the fix does and does not cover — reachability, not correctness: every
ratcheted quantity here is value-insensitive by construction. Liveness
counts cells that are non-null; "binding" counts a rule that reaches
PASSorFAIL. So a defect that changes a number in a cell that stays populated cannot move any of these metrics at all — it moves them only if it happens to null a cell or make a rule unevaluable. The worked example is the dropped+ airb_sl_exclterm in C 02.00 row 0340: a full supervisory run does not detect it either (8 passed with the defect live). This gate makes blind spots visible; catching a wrong value in a reachable cell is the supervisory register's job, and the C 02.00 subtotal shows the register currently fails at that too. Two escapes, two separate fixes — a reader who merges them will conclude the estate is better defended than it is. - Why every gate missed it: the gate was not weak, absent or wrongly anchored — it was unwired. A ratchet that no runner invokes has exactly the same effect on a defect as no ratchet, while reading in the repository like coverage is under control.
Be precise about what wiring it would have bought, because the obvious claim is
false. coverage_report.py was added on 2026-08-08 (2a1e200c), and the
off-balance-sheet portfolio that surfaced the original B5 C 07.00 defects was
built on 2026-08-01 (00b13b83) — the ratchet did not exist when they escaped.
More fundamentally, a ratchet fails on movement: a cell that was already dead
moves nothing, so wiring this gate would not have caught either the original B5
defects or their recurrence. What it prevents is the next one — a live cell
going dead, or a binding rule un-binding — and what it makes visible is the
standing blind spot's size. That is worth having, and it is not the same claim as
"this would have caught B5".
And the same inertness rotted the baseline it ratchets against.
The figures banked until this batch (251 / 277 / 1298 / 52817) reproduce at
neither matrix: re-measured on today's tree against the exact RUNS tuple of
the commit that banked them (13046bee, recovered with git show), the same
code yields 253 / 279 / 1300 / 52803. Nothing had re-derived those numbers
since they were written, so they had stopped describing the estate before the
matrix moved at all — a stronger statement of the same escape than "the matrix
grew". Nothing noticed, because the baseline recorded no field saying which
matrix, or which tree, it was measured over. A stale baseline is the second
failure mode of an unwired instrument, and the one that survives the wiring.
Do not read the dead_cells rise (52,817 → 55,553) as lost coverage: live
cells rose 7,891 → 8,193, against 63,746 declared, so the ceiling moved because
the declared population grew. And do not reason from the metric families moving in
opposite directions — they are independent, so a real cell-coverage loss alongside
an unrelated rule-coverage gain would look identical. The live-cell count is the
decisive evidence; the direction of dead_cells alone is not evidence of
anything. That property is the subject of its own entry below — the two cell
metrics are not floors.
- Gate change: in this change-set, from task 0.2 —
tests/contracts/test_coverage_ratchet.py (three always-on structural tests,
including test_the_coverage_ratchet_is_invoked_by_ci, which asserts the CI job
still invokes the script so unwiring it again fails locally, plus one
@pytest.mark.slow test that shells out to the real ~46s measurement) and the
coverage-ratchet job in .github/workflows/ci.yml running
scripts/coverage_report.py --check. Deliberately not in
scripts/arch_check.py or arch_metrics.json: the measurement is ~46s warm and
arch_check runs on every commit via the pre-commit hook, so this is a
considered placement rather than an omission to file.
The staleness limb closed too, in the same change-set from task 0.2b: the
baseline is re-banked over the 16-run matrix at 257 / 289 / 1285 / 55553 / 785
and now carries a provenance block naming the runs it was measured over, so
--check reports a matrix change as INVALID rather than as a regression and
a baseline with no provenance is called out as predating the field. That is the
structural fix, not a re-measurement: the failure was that the numbers could not
say what they described. Reducing the blind spot itself remains task 1.4.
- Verified red: two — one attacking the wiring, one attacking the ratchet.
The one that matches this escape class — the coverage-ratchet job removed from
a scratch copy of ci.yml, which is precisely the state the estate was in:
.github/workflows/ci.yml has no `coverage-ratchet` job. The coverage ratchet is
implemented but unrun, which is how it spent its whole life before P5.21: a
change can kill a live cell or un-bind a published rule with every gate still
green (.claude/LESSONS.md B5). Restore the job.
And the ratchet itself rejecting a regression — a measurement moved as a defect that kills one column would move it, against the banked figures:
[REGRESSED] union_binding_rules_crr: 257 -> 256 (may not decrease)
[REGRESSED] cells_live: 8193 -> 8181 (may not decrease)
[REGRESSED] dead_cells: 55553 -> 55565 (may not increase)
[REGRESSED] never_evaluated_rules: 785 -> 786 (may not increase)
Both run without mutating the tree. Task 0.2 also drove the real test body
through six perturbations — including a typo'd metric name, which would otherwise
surface as a KeyError 46 seconds into CI — and ran the slow test for real to a
genuine 1 failed in 46.66s.
Live caveat, and it belongs in this field rather than a footnote: --check
does not run at all as of this commit. cells_live is in _RATCHET_MIN and is
not in the banked baseline, so _check_baseline raises KeyError: 'cells_live'
— which is the typo'd-metric-name failure mode arriving for real, from a metric
addition rather than a typo. The invocation guard passes throughout, because it
asserts that CI invokes the script, not that the script works. So the red
above was produced against the baseline with cells_live banked at 8,193, which
is the state task 0.2b is landing, not the state on disk. Until that lands the
gate is wired and broken, and the honest reading of this entry is that its
escape is closed and its replacement gate is not yet demonstrably running.
The invocation guard is weaker than its own red suggests, and saying so here
is the point of the field. Its verified red exercises only the form where the
--check invocation is deleted outright. A skeptic defeated it five other
ways, each leaving the guard green: run: commented out, if: false,
continue-on-error: true, the step deleted with the command left behind in a
comment, and the workflow's on: triggers removed. A hardening is landing in
this batch. Separately, the metric choice has a defect of its own — an absolute
dead_cells ceiling can reward coverage loss, since dropping a template
removes dead cells; ratcheting cells_live instead is filed. An escape log that
overstates a gate's strength commits the error it exists to record.
- Lesson: partially graduated — B5 stays as prose. The ledger's 2026-08-09
row records B5 as PARTIALLY GRADUATED … STILL OPEN, narrowed to three: the
cell-granular case (the C 08.01 r0253 shape), the row-granular case (C 08.04's
single column is live while six of nine movement rows never carry a figure), and
never_evaluated_rules, the supervisory-register half. The first of those is
exactly what the paragraph above concedes these metrics cannot see. Since the
ledger's convention is that graduated prose gets deleted, calling this "graduated"
would invite destroying the two-leg fixture pattern that is currently the only
form the cell-granular case has. Do not delete it.
2026-08-09 — A defect that empties a column leaves all five register ratchets green¶
- Defect: A supervisory rule whose operands are all null or zero evaluates
to
VACUOUS. The per-run summary countsVACUOUSseparately fromPASSandNOT_EVALUATED— and then nothing constrains it. So a change that empties a column flips its rulesPASS→VACUOUS, the register's five ratchet tests stay green, and the estate's strongest reporting oracle reports success for a column it stopped checking. - Rule: Not a regulatory escape. The exposure is every published EBA/BoE rule whose operands can be emptied — i.e. all of them.
- Origin:
tests/acceptance/reporting/test_supervisory_validations.py. The summary was deliberately built to keep the four statuses apart (test_the_summary_keeps_unevaluable_rules_apart_from_passesasserts they are all reported and sum to the enforced population) — a correct and useful design, one step short of a gate. - Escape class:
no-assertion-of-presence - Why every gate missed it: the register asserts that no enforced rule
breaks, and vacuity is not breakage — it is the absence of an evaluation. The
count was recorded and treated as informational, which means the number moved
and no test cared. Note this is the neighbour of
path-never-exercised, not an instance of it: the 2026-08-08 recurrence proved that class's prescribed fix (build the portfolio, register it inRUNS) is necessary and not sufficient — the portfolio was registered and the cell was dead. C 08.01 r0253 held0.00in all six goldens, so the mandatory Tier 2 gate was structurally incapable of seeing a change to that column. Closing it took a two-leg fixture (a live cell that survives the change plus one that moves) and activated five previously-VACUOUSrules toPASS, includingboe_b0752_27, the r0253 tie-out itself.
The same interlock is live right now on the FCSM path, which is what makes this
worth reading twice. The seven Art. 197 capital understatements in the last
entry are unreachable by the estate's only FCSM golden portfolio
(reporting_funded_protection_portfolio.py), because both of its pledges are
CQS 1 — a CQS 1 security carries the obligor's own weight, so the defect cannot
express itself there. And that portfolio is the one deliberately withheld from
RUNS, having been registered against a config that silenced the very feature
it exists to exercise (B5's third form). So the defect sits behind two
independent layers of unreachability: a portfolio outside the register, and a
fixture shape that would not show it even inside. Neither vacuity nor coverage
can see that; only the oracle did.
- Gate change: in this change-set, from task 0.3 — a two-way vacuity ratchet
in the same register, keyed on (regime, rule_id) and stored as
known_vacuous_rules in
tests/expected_outputs/reporting/validation_known_breaks.json:
test_no_rule_falls_to_vacuous_outside_the_baseline (leg f) fails a rule that
falls to vacuity outside the register, and
test_no_baseline_vacuous_rule_asserts_again_without_being_removed (leg g)
fails a register entry that starts asserting again, so the population can only
shrink deliberately. Both drive extracted predicates
(_rules_newly_vacuous / _rules_no_longer_vacuous) rather than inline
logic. Baselined at 218 rules — 57 CRR / 161 Basel 3.1, 143 Error and 75
Warning severity — each carrying a written reason.
It is not a per-run count. The key matches known_broken_rules because a
vacuous rule has no failing coordinate to key on, and the 85 rule ids shared
across the two published extracts would otherwise collide. Membership is the
union over the sixteen runs: a rule qualifies only if it reaches a verdict
somewhere and never reaches PASS or FAIL anywhere. The per-run VACUOUS
counts stay in the register's summary block, descriptive and unasserted — the
ratchet does not read them. The consequence is worth knowing before relying on
it: a defect that empties a column on one portfolio while another portfolio
still exercises the same rule does not move this population, and that case
belongs to the goldens. Leg (f) catches the rule that stops asserting anything
anywhere, which is the case no other gate saw.
- Verified red: both legs, driven through the real test functions with a
synthetic measured set against the real committed register — deliberately not
a faked pipeline run. Leg (f), inserting b31/boe_b0752_27 (the C 08.01 r0253
tie-out, which passes on irb-classes today and is therefore absent from the
register) as vacuity-only, with its real measured facts:
1 published rule(s) now hold ONLY VACUOUSLY, 1 of them Error-severity. Every
operand was null or exactly zero, so the rule asserts nothing about our figures
while still reporting a green outcome:
b31/boe_b0752_27 [ERROR] held vacuously on 3 coordinate(s) across 4 portfolio(s)
rule: {t: OF08.01.01.01, r: 0070, c: 0253} = sum({t: OF08.02.01.01, c: 0253})
This is how a defect that empties a column passes this gate (LESSONS B5,
recurrence 2026-08-08). Find what emptied the cells.
Leg (g), removing b31/boe_b0958 (the OF 07.00 defaulted-exposure footing)
from the measured set, reported the entry leaving the vacuity population and
demanded the distinction that matters — banked activation versus a cell or run
that went away, "the estate got WORSE — fix that instead of deleting the
entry". I independently exercised the same _rules_newly_vacuous predicate
against the committed 218-entry register while writing this entry and saw it
reject the same rule. Register regeneration is idempotent over the curated
reasons; the suite is green at 8 tests.
- Lesson: second production-class recurrence of .claude/LESSONS.md B5, and
the second time B5 has been fixed as prose. Its executable form is this ratchet
plus the coverage ratchet in the entry above; B5's prose should retain only the
two-leg fixture pattern, which neither ratchet can express. The recurrence case
is now load-bearing rather than illustrative: boe_b0752_27 passes on
irb-classes and remains vacuous on rich, crm-substitution and art199, so
that one registered run is the whole reason it sits outside the vacuity
population — re-empty r0253 and leg (f) fires. Its 26 siblings
(boe_b0752_*, boe_b0814_*, boe_b0757, all Error severity) are in the
register with the B5 discharge precedent written on each entry, so the family is
ratcheted rather than merely known.
2026-08-09 — The detection rate of the whole estate is unknown, and the instrument that measures it would have lied¶
- Defect: Two compounding things. (1)
scripts/defect_injection.py— 22 mutants, a data-driven gate ladder, reachability as a first-class verdict — has never been run as a campaign, so no scorecard exists and the estate's detection rate is unmeasured. The plan that commissioned it (docs/plans/independent-validation-system.md:455) says that before the harness existed nobody could say whether the rate was 40% or 90%; that sentence is still true, because building the instrument and reading it are different acts. (2) Every gate command in the ladder was hardcoded to spawn throughuv run. On a runner without a usableuv, every gate fails to spawn, each failure scores as a detection, and the harness publishes a fictitious detection rate near 100%. This is measured, not hypothetical: on this project's own sandbox the defaultuv runpath exits 2 withCould not acquire lock … Read-only file system, so--ladder legacyrun here before the fix would have reported ~100% detection and zero escapes. The one number the harness exists to produce was the number it was most likely to get wrong. - Rule: Not a regulatory escape.
- Origin:
scripts/defect_injection.py, merged 2026-08-08 with the independent validation system. - Escape class:
gate-not-run, for the unrun campaign. Limb (2) is a defect in a gate rather than one that escaped a gate, and the taxonomy has no class for that; it is recorded here rather than given a class it does not fit. The general shape is worth naming: a gate that can go red for a reason unrelated to the defect scores that red as success, so any instrument whose signal is "something failed" needs to distinguish failed from did not run. - Why every gate missed it: nothing consumes the scorecard, so its absence
is invisible — there is no baseline to regress against and no CI job to go
red. Limb (2) survived review because the ladder is declared in the form a
developer types, and on a developer's machine
uv runworks; the failure mode only appears on a runner nobody had tried. - Gate change: in this change-set, from the injection-harness runner
override —
DEFECT_INJECTION_PYTHON(INTERPRETER_ENV_VAR,scripts/defect_injection.py:127) retargets the ladder through a named interpreter via a singleresolve_commandchokepoint (:150) that every gate command and the baseline command pass through. A partially retargeted ladder is worse than an unretargeted one, so a command it cannot rewrite is a hard error rather than a silent pass-through.preflight()(:187) then importsrwa_calc.engine.pipeline— not barerwa_calc, whose lazy__init__imports in ~150µs without touching polars — and raisesInterpreterUnusable, exiting 2 frommain(), so a broken interpreter aborts the campaign instead of reddening every gate.
Owed, not done: no test guards any of this. Nothing under tests/ imports
defect_injection at all. The graduation target is a contract test asserting
that an unset env var leaves every LADDER command and baseline_cmd byte
identical, and that preflight() raises on a nonexistent interpreter. Until
that exists the guard is correct-by-inspection-and-one-manual-run, which is what
this file exists to stop people calling a gate.
- Verified red: the pre-flight aborting a real campaign invocation
(--ladder fast --mutants control-reachable-output-floor-schedule) with
DEFECT_INJECTION_PYTHON=/nonexistent/python, exit code 2, before the baseline
digest capture and before any mutant was applied:
SPAWN PRE-FLIGHT FAILED — CAMPAIGN ABORTED, NOTHING SCORED
command /nonexistent/python -c import rwa_calc.engine.pipeline
reason the executable does not exist ([Errno 2] No such file or directory: '/nonexistent/python')
Two further reds from the same guard: /usr/bin/python3 spawns but cannot
import (reason it exited 1, with the ModuleNotFoundError quoted), and the
default uv run path in this sandbox gives reason it exited 2 with
Could not acquire lock … Read-only file system — the escape this guard actually
closes. Separately, I exercised resolve_command in-process and saw it refuse
both shapes it cannot rewrite (a command not beginning uv run, and
uv run watchfire check) rather than passing them through, which is the
silent-partial-retarget failure mode.
The campaign itself is still unrun, and no scorecard exists. Only
--reachability-only probes have run, which execute no gates; their two outputs
were written under tmp/dij/ and deleted by another agent's rm -rf tmp, and a
third run died in out.write_text because main() never creates --out's
parent directory. The default output path is scripts/defect_scorecard.json,
which is gitignored. Anyone quoting a detection rate for this workstream today
is quoting a number that does not exist. Filed as task 0.1, with the nightly
campaign and a detection-rate ratchet as task S.3.
Closed for the runner override; the reachability probe is a separate instrument
and it is open. A 22-mutant probe run (1,164s) produced four mismatches out of
22, including the deliberate UNREACHABLE control moving output — so the probe
currently reports reachable for a mutant chosen to be unreachable, which would
corrupt the denominator of any detection rate it is used to compute
(UNREACHABLE mutants are excluded from numerator and denominator both). Task
0.1a. A third defect, task 0.1b, has the harness rewriting mutation targets with
CRLF line endings. Splitting the claim matters here: two of the three instrument
defects in this entry are still live, and only the spawn path is demonstrably
fixed.
- Lesson: this is the second of these four entries whose class is
gate-not-run for the same underlying reason — the estate's habit is to build
the measurement and stop before wiring it. That is a pattern rather than two
slips, and the coverage ratchet's test_the_coverage_ratchet_is_invoked_by_ci
is the shape of its fix: an instrument ships with a test that it is invoked.
2026-08-09 — Eleven wrong numbers found by the oracle and parked as accepted disagreements, eight of them understating capital¶
- Defect:
KNOWN_DISAGREEMENTSintests/oracle/test_oracle.pyholds 11 entries, allxfail(strict=True)rather than fixed, and eight of them understate capital. Seven of the eight were added inside this batch by the CRM oracle (7c454be1), which is the fact this entry is really about: the register grew 4 → 11 in a matter of hours with nothing constraining its size. ORC-280— the largest. Art. 197 collateral eligibility is never applied on the Art. 222 Financial Collateral Simple Method path. At full cover on a CQS 5 sovereign security the oracle gives 1,500,000 against the engine's 1,000,000 — an understatement of 33.3%, the whole exposure moving from the obligor's 150% to the security's own Art. 114(2) 100%.ORC-257,ORC-258,ORC-275,ORC-278,ORC-279,ORC-281— the same defect at 30% cover, each understating 10.0% (1,500,000 against 1,350,000), across Art. 197(1)(b) rated and unrated sovereigns, Art. 197(1)(d) rated and unrated corporates, Art. 197(1)(f) equity, and the Art. 218 credit-linked note on which the engine raisesCRM019and then recognises the pledge anyway. The family's own reason text is unambiguous: "DIRECTION IS UNIFORMLY ANTI-CONSERVATIVE OR NEUTRAL, never conservative." Mechanism:engine/crm/processor.pyrunscompute_fcsm_columnsat Step 3.8, beforeapply_haircutsat Step 4 — andapply_haircutsis the only place the engine overrides a firm-supplied eligibility attestation, so the Simple Method recognises collateral the Comprehensive Method rejects.ORC-282, the Comprehensive-Method control, passes, which localises it to the one method.ORC-109— CRR Art. 121(1) Table 5 not applied to the institution class: at CQS 6 the engine returned 100% against a required 150%, an understatement by a third, withORC-105(CQS 1) andORC-020(CQS 2) as the conservative limbs of the same unwired ladder. This family is being discharged as this entry is written — P1.316 has wiredcp_sovereign_cqsthrough Table 5 under task S.2, so all three leave the register. It is recorded here because it was parked for a day with a known capital shortfall in it, not because it is still open.ORC-142— PS1/26 Art. 154(4A)(b) limb (iii): the 10% IRB mortgage RWEA floor applied to residential property outside the UK (oracle 0.00, engine 373,345.27). Conservative in direction, and unrepresentable rather than mis-gated: no module underengine/irb/reads any obligor or property country column, so no input could switch it off. Rescoped under task #21 — the fix needs aproperty_country_codecarrier, not the obligor-country gate the original framing implied.
The count in this paragraph is a snapshot, and that is the point. It was 4
when the entry was drafted, 11 when it was corrected, and lower again by the time
P1.316 lands. A register whose size is recorded in prose is stale the moment the
register moves, which is exactly why the fix is a ratchet and not a sentence.
- Rule: CRR Art. 197(1)(b)/(d)/(f), Art. 198(1)(a), Art. 218, Art. 222,
Art. 114(2); CRR Art. 121(1) Table 5 and Art. 121(2); PS1/26 Art. 121(6),
Art. 154(4A)(b), Art. 163(1)(b)-(c).
- Origin: found 2026-08-08 by the independent oracle, on merge of the
validation estate. The engine defects themselves predate it.
- Escape class: caught-and-parked — the eighth class, added with this entry.
The case for a new class is not that the existing labels read wrong
narratively; this file's own discriminator is that the class determines the fix,
and no-gate-exists → "create the gate" would in fact produce the register
ratchet named below. A class added to fit one datum is fitted, not derived. It
earns its place on two other grounds. First, the shape recurs across at least
four parallel registers in this repository — KNOWN_DISAGREEMENTS,
classification_table.toml's [[known_disagreement]] D1-D7, known_broken_rules
and known_vacuous_rules — plus strict xfails and plan bullets, so it is a
standing structural feature rather than one incident. Second, its fix targets
the register rather than a detector, which none of the other seven prescribe:
every one of them ends in something that looks at the output, and this one ends
in something that looks at the list of things we have agreed to tolerate. The
4 → 11 growth inside hours of the class being coined is the class earning its keep.
- Why every gate missed it: no gate missed it. strict=True is real discipline
in one direction — it prevents a silent fix, because an entry that starts
agreeing becomes an XPASS and a hard failure — and none at all in the other.
KNOWN_DISAGREEMENTS has no size ratchet, no owning bullet per entry and no
expiry, so seven new capital understatements were added in one batch and every
gate stayed green. The register was built to make findings triageable and became
the place they are stored.
- Gate change: filed as task #28 while this entry was being corrected — a
two-way ratchet on the size of KNOWN_DISAGREEMENTS plus a requirement that each
entry names an owning plan bullet. The 4 → 11 growth is what moved it from a
nice-to-have to the urgent item: the entry described a mechanism, and the
mechanism then fired. Code fixes tracked separately: the Art. 121 family under
P1.316 (landing now, task S.2, which must delete all three entries in the same
change), the FCSM family needing the Art. 197 gate factored out of
apply_haircuts so it applies to the Simple Method input as well — explicitly
not a step reorder, since Step 3.8 must precede the Comprehensive computation
that IRB LGD still needs — and ORC-142 under task #21.
- Verified red: n/a for detection — the disagreements are red today, by design,
as strict xfails. NOT VERIFIED for the disposition ratchet, which does not
exist yet. By this file's closing rule the escape therefore remains open, which is
the correct state to record: what exists today is the detection, not the
correction.
- Lesson: candidate for .claude/LESSONS.md — a strict xfail is a decision
to ship the wrong number; it needs an owner and a date, not just a reason.
Filed with the team lead rather than added here, since this file does not own
that one.
2026-08-09 — The register does not notice a term dropped from a C 02.00 subtotal¶
- Defect:
reporting/corep/c02.pybuilds C 02.00 row 0340 (A-IRB corporate) asairb_corp + airb_sl_excl. With+ airb_sl_exclremoved — the A-IRB specialised-lending contribution silently leaving the row — a full run of the supervisory validation suite reported8 passed. The mutation was live in the tree while that run happened. Direction: the term is only ever added, so dropping it understates the reported A-IRB corporate figure, and its RWEA goes missing from the class breakdown while the approach total still counts it —.claude/LESSONS.mdB6's shape, arrived at through a dropped term rather than a re-key. - Rule: COREP C 02.00 row 0340 composition. Ten published rules name that
cell; the two that bear on it are
v0211_m(ERROR, footing identity{r0310} = {r0320} + … + {r0410}) andv4252_i(ERROR, cross-template identity{C 02.00, r0340, c0010} == {C 08.01.a, r0010, c0260, s0007}). - Origin: the mutation was transient, injected during task 0.3's work. The escape is the register's inability to see it, which is a standing property of the estate.
- Escape class:
gate-not-run. The catching gate is not missing — this repository ships it.v0211_mis a live ERROR-severity footing identity insrc/rwa_calc/reporting/validations/rules/crr-eba-v3.0-credit-risk.json, and it is never evaluated. That is the class's definition exactly, and it is whyno-gate-existswould be the wrong label: the fix is to make an existing rule run, not to invent a check. -
Why every gate missed it:
v0211_mis one of four live ERROR rules on the C 02.00 hierarchy that are never evaluated anywhere —v0204_m,v0207_m,v0210_m,v0211_m; the fifth rule in that family,v0205_m, is WARNING severity, andv0207_mdoes evaluate, so "none of them runs" is false and the split is the evidence. The mechanism is not that C 02.00 sits outside the machinery: the recorded reason is{'row_not_emitted': 8}, so C 02.00 is in the cellspec executor and the rows the rules name are not emitted.v0210_mneeds r0250-0300 andv0211_mneeds r0310-0410, which the repo does not emit;v0207_mneeds r0060-0211, which it does — hence one evaluates and the others do not. That mechanism is already written verbatim in plan item P1.318, uncited until now. Two consequences worth stating plainly: -
The estate ships the rule that detects its own headline own-funds defect and never runs it.
v0204_masserts{r0010} = {r0040} + {r0490} + {r0520} + {r0590} + {r0630} + {r0640} + {r0680} + {r0690}, which on a credit-only book forcesr0040 == r0010; ourr0040isr0010 / 12.5.v0210_mgivesr0250five children, so r0250 is a parent where the engine puts the institutions leaf — the row-axis shift of task #17, detected by a rule we already own. - The second candidate mechanism is real but secondary:
v4252_i, the only cell-level tie-out of r0340 (== {C 08.01.a, r0010, c0260, s0007}), carriesif_value_missing: do not run rule, so a missing sheet silently removes it. Fail-open by the publisher's own semantics, on top of a rule set that is not being evaluated anyway.
What is not the explanation: the path is exercised. The cell is populated and
C 02.00 is emitted on every portfolio. And the coverage ratchet cannot see the
value defect — its metrics are value-insensitive (first entry) — but it can see
precisely this: an ERROR rule that never runs is one of the 785
never_evaluated_rules that entry counts and that nothing gated. These five
are concrete instances of that aggregate, which is what an aggregate is for.
- Gate change: deferred and filed — task #16 for this data point (it feeds
step 0.1's scorecard), task #17 for the row-axis shift, task #19 for the four
unevaluated ERROR rules. Making v0204_m/v0210_m/v0211_m evaluate is the
fix that catches the row shift and the subtotal composition together, but it is
not cheap: emitting the rows those rules address is plan item P1.318,
Effort: L, single-stream, moving 10 golden frames plus the validation
baseline. I said "cheapest of the three" in an earlier draft and that was wrong.
Independent re-derivation of C 02.00's class rows in tests/conformance/ remains
the second layer.
- Verified red: inverted — the gate was observed not firing, which is the
strongest evidence in this file. A full supervisory run with the mutation live
reported 8 passed. That is a measured negative result rather than an inference
from reading the rules: whatever the register checks, it does not check this.
- Lesson: this is the estate's first ESCAPED verdict, and it arrived free as
a side effect of another item — before the injection campaign built to produce
such verdicts has run even once (tasks 0.1 / 0.1a). Logged in the same run and
deliberately not chased: C 02.00: row 0300 (14,625,069.66) exceeds its class
breakdown (21,574.13) — a headline own-funds row exceeding the sum of its own
class rows, a live B6 condition that the estate emits as a log line and nothing
fails on.
2026-08-09 — A ratchet that can be satisfied by deleting the coverage it measures¶
- Defect: two of the coverage ratchet's five metrics are not floors.
template_cell_liveness_bpis a ratio whose denominator shrinks with its numerator, anddead_cellsis an absolute count of the complement (declared − live). Analytically, dropping N declared cells of which K are live passes both ratchets wheneverK/N ≤ 0.1285— so deleting any region less live than the estate's own average improves both numbers. Measured: droppingb31/richloses 689 live cells whiletemplate_cell_liveness_bpimproves 1285 → 1374 anddead_cellsimproves 55,553 → 47,123. Across 16 leave-one-out runs the two cell metrics never caught anything on their own, and on 4 of 16 they registered an improvement while real liveness fell; every genuine red came from a binding-rule fall or anever_evaluatedrise. "Cell liveness may not FALL" is therefore not a coverage floor, and the CI comment and the script's docstrings say that it is. - Rule: Not a regulatory escape.
- Origin:
scripts/coverage_report.py,_RATCHET_MIN/_RATCHET_MAX— in this change-set. The gate had not shipped. - Escape class: none of the eight, and it should not be forced. Every
class presumes a defect that reached production; this one was caught by
adversarial review of a gate before it landed, and its subject is the gate
rather than the engine. It is recorded here because this file's question — which
gate should have caught this — has a real answer worth keeping (adversarial
review of a new gate's metric algebra, which is what did catch it), and because
anyone tracing the coverage ratchet's history needs to find it. If entries of
this shape recur,
gate-unfitis the name to give them; one instance is not a taxonomy. - Why nothing else would have caught it: the metric algebra is invisible to tests. Every structural test of the ratchet — including task 0.2's six perturbations — checks that a declared regression is rejected, which these metrics do correctly. None asks whether the quantity being ratcheted is the quantity that matters. Only leave-one-out measurement over the real matrix exposes it, and nothing in the estate does that automatically.
- Gate change: in this change-set, from task 0.2b, and half-landed as of this
commit —
cells_liveis in_RATCHET_MINin the code and is not in the banked baseline, so--checkcurrently raisesKeyError: 'cells_live'rather than gating. The floor value is 8,193; banking it is what completes this, and until then the gate this entry describes is broken rather than working. It is already computed aspayload["cells"]["live"], it fell in 15 of the 16 deletions and in both config-silencing variants, and it never rose on a loss. Filed separately:never_evaluated_error_severity_{crr,b31}(175 / 195) is computed and unratcheted, so swapping one ERROR-severity never-evaluated rule in for one INFO out is invisible to the flat total — while the script's own docstring calls an ERROR rule that never runs anywhere the worst case in the estate. - Verified red: the leave-one-out measurement is itself the red, and it is red
in the diagnostic direction — the two metrics passed while coverage fell, on
4 of 16 deletions.
cells_livewas then checked against the same 16 deletions before being adopted and fell in 15; the one deletion it did not catch is a residual the follow-up should name rather than leave implied. - Lesson: the executable form is the
cells_livefloor itself. The transferable rule — ratchet the quantity you care about, not a ratio of it and not its complement — is offered to the operator as a.claude/LESSONS.mdentry, since a ratio-shaped ratchet reads as a floor to every reviewer who does not do the algebra.
2026-08-11 — The release script's test run happens before the mutation it should catch¶
- Defect:
scripts/deploy.pybumped the package version and regenerated two of the four generated artifacts. Three targets embed the version in their own output —docs/data-model/regulatory-tables.md(generate_regulatory_tables.py:811),docs/development/confidence-matrix.mdandtests/contracts/data/confidence_snapshot.json(generate_confidence_matrix.py:525,747) — so the bump alone was sufficient to make all three stale, with no other change in the tree. v0.3.25 was committed and tagged in that state; CI on the release commit failedtest_regulatory_tables_page_is_freshandtest_confidence_matrix_is_fresh. A second, latent instance rode along:generate_citation_matrix.pywritestests/contracts/data/citation_snapshot.json, whichGIT_STAGE_FILESnever staged — the same defect, one release away from firing. - Rule: Not a regulatory escape.
- Origin:
scripts/deploy.py::build_release/GIT_STAGE_FILES, standing since the generated pages acquired their version stamps. Every prior release had the same hole; it only became visible when a freshness contract test covered the stamped targets. - Escape class:
gate-not-run, with a twist worth recording. A catching gate existed and was in fine health: both freshness contract tests ship, run in the default suite, and pass.deploy.pyruns the suite as step one and bumps the version as step two, so the gate measured a tree in which the defect did not yet exist. The class table prescribes "move the gate earlier"; here the correct move is the opposite, later — after the mutation. The class is about a gate that ran at the wrong point, and earlier is simply the common case, not the definition. If this shape recurs, the prescription column should read "move the gate to the other side of the mutation". - Why every gate missed it: ordering, and nothing else. Local
pytestpassed (measured pre-bump). The pre-commit gate passed (same reason). CI was the only gate positioned after the mutation, and CI is the last one — by the time it spoke, the version commit and the annotated tag existed, and the tag had been pushed. Note what this rules out: it is not that the freshness tests are weak or that a path went unexercised. They are strong and they ran. A gate's position in the sequence is part of its specification, and nothing in this estate had ever stated the position of these two. - Gate change:
scripts/deploy.py::build_releasenow runs both version-stamped generators after the bump, andGIT_STAGE_FILEScarries their three targets pluscitation_snapshot.json. That fixes the instance. The category is closed bytests/contracts/test_release_regeneration.py, which discovers version-stamped generators by inspection — anyscripts/generate_*.pyreadingpyproject.toml's version — and fails when one is not invoked bydeploy.py. Discovery is deliberately not a hand-maintained list, because a hand-maintained list is exactly what was wrong. Its companion test asserts the sweep is non-empty, so a drifted heuristic fails loudly instead of passing vacuously. - Verified red: run against the real pre-fix
deploy.pyat6f513697:
RED against pre-fix deploy.py (6f513697).
Not regenerated by the release script:
- generate_confidence_matrix.py
- generate_regulatory_tables.py
Green against the shipped deploy.py (2 passed). The red is produced from
the actual defective commit, not a reconstruction of it.
- Lesson: a gate that runs before the step it protects has not run. The
release script's ordering — test, then mutate, then commit — reads as
conscientious and is precisely backwards for anything the mutation itself can
break. Worth a .claude/LESSONS.md entry in the operator's judgement, because
the shape generalises past releases: any sequence that validates and then
transforms has this hole.
2026-08-11 — A wheel outgrew the pinned uploader, and no gate in the release flow could see it¶
- Defect: publishing v0.3.25 to PyPI failed with
Checking dist/rwa_calc-0.3.25-py3-none-any.whl: ERROR
InvalidDistribution: Invalid distribution metadata:
'2.5' is not a valid metadata version
uv build resolves the build backend fresh on every run, and the current
backend emits Metadata-Version: 2.5. .github/workflows/publish.yml pinned
pypa/gh-action-pypi-publish at cef22109… (v1.14.0), whose vendored Twine 6
predates 2.5 and refuses it. Nothing in this repository changed: v0.3.24
published on 5 August and v0.3.25 did not, because one of the two components
moved on its own.
- Rule: Not a regulatory escape.
- Origin: standing since the action was pinned. The pin is correct practice —
it is the reason the failure was a clean refusal rather than a supply-chain
surprise — but pinning one side of a two-sided compatibility relation converts
"we are current" into "we are frozen against a moving target".
- Escape class: no-gate-exists. Nothing anywhere — locally, in the
pre-commit gate, in CI, or in deploy.py — inspected a built distribution.
uv build was run for its exit code alone, and its exit code is 0 for a
perfectly well-formed wheel that this particular uploader happens to reject.
- Why every gate missed it: the whole estate tests the source tree, and
this defect does not exist in the source tree. It exists only in the artifact,
and only in relation to a version pinned in a YAML file that no test reads.
Note in particular that twine check alone would not have caught it: any
Twine new enough to install today accepts 2.5 happily, so a local twine check
is green on precisely the wheel that fails. The failure is not "malformed
distribution", it is "distribution newer than the pinned publisher" — a skew
between two versions, visible only when both are read together. A gate built on
the obvious reading of this incident would not have caught this incident.
- Gate change: scripts/check_distribution.py — reads Metadata-Version
from every built wheel and sdist, reads the pinned publisher version out of
publish.yml, and fails when the former outruns what the latter accepts
(PUBLISHER_METADATA_SUPPORT). Invoked from deploy.py::build_release and
from the CI build job, with uvx twine check dist/* alongside it in CI for
the malformed-distribution class it does cover. tests/contracts/test_distribution_gate.py
covers the checker and asserts both call sites still exist — this project has
shipped an inert ratchet before, and a script nothing calls reports success
forever. An empty dist/ is a failure, not a pass.
- Verified red: the shipped gate, via
check_distributions(dist_dir, workflow) — the gate exposes no path-typed CLI
argument, so a reproduction calls the function, exactly as the contract tests
do — against the real v0.3.25 artifacts with the pre-fix pin from 451e97db:
Built distributions declare core metadata newer than the pinned publisher accepts.
- rwa_calc-0.3.25-py3-none-any.whl declares Metadata-Version 2.5
- rwa_calc-0.3.25.tar.gz declares Metadata-Version 2.5
pypa/gh-action-pypi-publish is pinned at v1.14.0, which accepts up to Metadata-Version 2.4
Exit 1. Green (exit 0) against the shipped v1.14.2 pin. Both the artifacts and
the pin are the genuine article, so this reproduces the escape rather than
modelling it. The disarm case was checked too: replacing the deploy.py call
site fails test_gate_is_actually_invoked.
- Note — the gate's own first version failed the quality gate: it took
--dist-dir / --workflow as type=Path CLI arguments, which SonarCloud
flagged as pythonsecurity:S8707 (MAJOR), taking new_security_rating to C
against a required A. That is the third instance of this rule here, after
injection_ratchet.py and coverage_report.py's bank(). The remedy is
already settled and is not a containment guard: commit a5d34c0d records
two successive attempts at resolve-then-contain that left the finding in place.
Both path arguments were therefore removed rather than sanitised, and
test_gate_exposes_no_path_typed_cli_argument now asserts that no type=Path
argument returns.
Three instances of one rule, each fixed the same way, is a lesson that has
proven it cannot survive as prose, so it was graduated to arch_check.py
check 19: no type=Path argparse argument anywhere in scripts/. Verified
red by restoring this script's own pre-fix body from d4fdcee6 — the exact
code SonarCloud rejected — which the check names argument by argument:
scripts/check_distribution.py: add_argument(--dist-dir) uses type=Path...
scripts/check_distribution.py: add_argument(--workflow) uses type=Path...
arch_check exit=1
Exit 0 once restored. Running it for the first time found nine further
instances that no one had counted — coverage_report.py --out,
defect_injection.py --out, five in impact_report.py, two in
parity_gate.py. They ship as a shrink-only CLI_PATH_ARG_ALLOWLIST
rather than being fixed here, because draining them means touching four
scripts and their workflow call sites; filed as task #36. Two things are worth
recording about that number. It is more than double the instances anyone knew
about, which is the usual result of converting prose into a check. And
coverage_report.py is on the list despite commit 89bf0323 having already
fixed this rule in that same file — the earlier pass removed bank()'s
baseline_path and left --out untouched, which is precisely what
per-instance fixing looks like from the outside: a file that has been "fixed"
and still carries the defect.
- Lesson: pinning one side of a compatibility relation makes the other side
a moving target, and the skew is nobody's regression. Neither component was
wrong; both were doing their job. The general form — when you freeze one of two
things that must agree, something has to assert they still agree — is the part
worth carrying, and it applies to every pinned tool in this repo, not just the
uploader.
2026-08-12 — The input contract was written, unit-tested, and connected to nothing¶
- Defect: 10 of the 14 public validators in
src/rwa_calc/contracts/validation.py— 402 lines, every one of them carrying green unit tests — could not be reached from any production path undersrc/.validate_pd_range,validate_lgd_range,validate_ccf_modelled,validate_non_negative_amounts,validate_schema,validate_schema_to_errors,validate_required_columns,validate_raw_data_bundle,validate_resolved_hierarchy_bundle, andvalidate_aggregated_bundle— the last of which is the output-bounds guard (RW > 1250%, RW < 0, RWA < 0, null EAD) that would have run on every result the engine has ever produced. 48 unit tests pass over that code in 2.7 seconds, testing logic that never runs on customer data.
The consequence is measured, not inferred. On CRR, £1m senior corporate, F-IRB:
a risk feed sending PD = 1.5 to mean "1.5%" returns RWA £603.67 against a
correct £1,119,286.69 — an understatement of 99.9461%, a £1,118,683 capital
shortfall on every £1m of exposure, with no exception, no null and no
CalculationError. LGD = -0.2 returns £0.00. Maturity = -3 years returns
£776,750.85. CQS = 0, 7 or 99 on a corporate returns RW 100%, which moves
capital in both directions — 100% against a true 20% at CQS 1, and 100%
against 150% at CQS 6 on an institution. Every one of those is a number a
reviewer would accept.
- Rule: Not a single regulatory article — the exposure is the input domain of
every article the engine implements. The output-bounds limb is the closest to a
named rule: RW is capped at 1250% (CRR Art. 92 / PS1/26), and
validate_aggregated_bundle is the only thing in the estate that would have
said so.
- Origin: src/rwa_calc/contracts/validation.py. The escape is old and was
already documented: docs/plans/engine-defensiveness-boundary-hardening.md,
written 2026-05-29, states in its root-cause line that "the
contracts/validation.py bundle validators were never wired into
pipeline.py." That sentence was still true 2.5 months later, which is the part
of this entry worth reading twice — the finding did not escape detection, it
escaped conversion into a gate.
- Escape class: gate-not-run. The catching gate is not missing; this
repository ships it, with tests. It is never invoked on customer data. That is
the class's definition exactly, and it is why no-gate-exists would prescribe
the wrong fix — nothing needed inventing, only wiring. It is the fourth
entry in this file to carry that class, and the fifth measured instance of the
habit the log itself names: build the measurement and stop before wiring it.
- Why every gate missed it: no gate looks at reachability. Every existing
guard in the estate answers "is this output right?", and an unreachable
validator produces no output to be wrong. Worse, it produces the opposite of a
signal: 48 passing unit tests over 402 lines of validation logic read, to any
reviewer and to every coverage instrument, as an input contract that is
enforced. Coverage tooling counts those lines as covered, because they are —
by tests, not by the pipeline.
Note the asymmetry with the four earlier gate-not-run entries. Those were
instruments nobody had run yet; this one had been run, its finding recorded in
a plan document, and the plan document then sat. A written root-cause line is not
a gate. Prose has now failed at this five times.
- Gate change: scripts/arch_check.py check 20 —
check_guard_reachability, registered in main() so
python scripts/arch_check.py runs it on every commit via the pre-commit hook.
Every public function in contracts/validation.py (the input contract, which is
guard-shaped in whole) and every guard-named public function elsewhere under
contracts/ (validate_ / check_ / assert_ / require_ / ensure_) must
be transitively reachable from production code under src/. Scoped by shape
rather than by path so a future contracts/checks.py is covered on the day it
is written; deliberately not extended to the layer's create_empty_* factories
and *_error constructors, 10 of which are unreachable today and would have
given the check a ten-entry allowlist on arrival.
Three properties are load-bearing, each aimed at a way this file records gates going soft:
GUARD_REACHABILITY_ALLOWLISTis empty by design. Seeding it with the ten known-unreachable validators would have made the check vacuous on arrival — thecaught-and-parkedshape, where a gate fires, the finding is filed, and the wrong behaviour ships anyway. Stale entries are themselves violations, so the list can only be drained.CONTRACTS_GUARD_SURFACEpins the population, because reachability alone is satisfiable by removing the thing it measures: renamevalidate_pd_rangeto_validate_pd_rangeand it leaves the measured set; delete it and the violation leaves with it. This repo has shipped that shape before — four back-compat shells silently disarmed a regression guard that readmodule.__file__(check 18). Deleting a guard stays legitimate; it just has to be deliberate, so the pin is removed in the same change and a reviewer sees which guard went away. It fired for real during this batch (see below).tests/contracts/test_guard_reachability_gate.py::test_every_arch_check_check_is_registeredgeneralises the wiring assertion past check 20: nocheck_*function inarch_check.pymay exist withoutmain()invoking it. An unregistered check is this escape class committed inside the gate itself.
The analysis has exactly one implementation. scripts/validator_reachability.py
— the census this grew out of, and the reproduce command in the proposal — now
reads measure_guard_reachability from arch_check instead of keeping its own
copy, and test_the_reachability_analysis_has_one_implementation asserts both
that it does and that the two agree on the population. The dependency points
diagnostic → gate and never the other way, so the gate keeps working if the
diagnostic is renamed or deleted.
- Verified red: two, and the second was not planned.
(1) The escape itself, check 20 run against the committed pre-wiring tree
(git archive aa2182bd src, so the measurement is of the state that shipped
rather than of a scratch mutation):
[FAIL] Contracts guards are reachable from production (wire it, or delete it)
contracts/validation.py:1140: validate_aggregated_bundle -- guard unreachable from production
contracts/validation.py:403: validate_ccf_modelled -- guard unreachable from production
contracts/validation.py:373: validate_lgd_range -- guard unreachable from production
contracts/validation.py:317: validate_non_negative_amounts -- guard unreachable from production
contracts/validation.py:348: validate_pd_range -- guard unreachable from production
contracts/validation.py:236: validate_raw_data_bundle -- guard unreachable from production
contracts/validation.py:151: validate_required_columns -- guard unreachable from production
contracts/validation.py:277: validate_resolved_hierarchy_bundle -- guard unreachable from production
contracts/validation.py:64: validate_schema -- guard unreachable from production
contracts/validation.py:176: validate_schema_to_errors -- guard unreachable from production
Total: 10 violation(s)
Ten violations, zero allowlist entries, and the same ten the proposal's census named — an AST-based reference analysis independently reproducing the seed script's regex-based answer.
(2) The population pin, firing unprompted. While the wiring work was landing
in parallel, five of the ten validators were deleted rather than wired
(validate_schema, validate_schema_to_errors, validate_required_columns,
validate_raw_data_bundle, validate_resolved_hierarchy_bundle). The
reachability limb went green as they disappeared — exactly the defusal the pin
exists for — and the pin caught all five:
contracts: validation.py::validate_schema -- pinned in CONTRACTS_GUARD_SURFACE but no
longer a public module-level function there. Check 20 may not be satisfied by deleting,
privatising or relocating the guard it measures; if the removal is deliberate, delete
the pin in the same commit so a reviewer sees it
This is the strongest evidence in the entry, because it was not a constructed
probe: the check's anti-defusal limb fired on real concurrent work within an
hour of being written, on a removal nobody had announced. Two of those five
deletions are prescribed by the proposal; the other three are a wider call, and
the pin is what turned them into a decision somebody has to state. Before the
pins were dropped the removal was verified complete — definitions and their
private helpers gone, and no reference to any of the five surviving anywhere
under src/ or tests/ — so what the change-set records is a finished
deletion, not a half-migration. The consequence worth stating plainly: the
estate now ships no declared-vs-actual schema-drift check at all.
(3) Six adversarial perturbations, run against a throwaway copy of src/
so the real tree was never mutated (a killed injection run leaving a live
mutation in src/ has cost this project hours before). Control: 0 violations.
Privatising validate_pd_range to _validate_pd_range → 1 (the pin, not the
reachability limb, which goes quiet — that is the defusal). Deleting
validate_lgd_range outright → 1. A new contracts/checks.py::validate_nothing
that nothing calls → 1, which is the scope-by-shape decision earning its keep:
the check covers a module it has never seen, so hard-coding validation.py
would have let that through. Deleting contracts/validation.py entirely → 9.
And the negative control that matters, since a gate that fires on everything
is not a gate: a new non-guard public function in the same unreachable
position (create_empty_thing) → 0 violations, confirming the scope
boundary is the one documented rather than an accident.
- Lesson: this is the fifth instance of the estate's dominant meta-pattern and
the second time it has been "fixed" by writing it down. .claude/LESSONS.md
carries the graduation rule for exactly this case — a lesson that reaches
production twice cannot survive as prose — so the entry belongs in the
Graduation ledger rather than the working set. Check 20 is the same move check
14 made on Polars namespaces: it removes the category, not the instances.
What remains prose, because no check can express it: the proposal's other three
phases (declare the input domain as data, fuzz the pathology axis, make absence
loud) are what stop the next silent-plausible-number defect; check 20 only
guarantees that a guard we have already written is running.
2026-08-12 — The registers of tolerated findings had no size gate and named no owner, and the count everyone quoted was wrong in both directions¶
- Defect: This file's 2026-08-09 entry 4 recorded eleven wrong numbers parked as accepted disagreements and named its own gate change — a ratchet plus an owning bullet per entry — then closed its Verified red field with "NOT VERIFIED for the disposition ratchet, which does not exist yet. By this file's closing rule the escape therefore remains open." It stayed open for three days. This entry is that closure, and it is filed as its own entry rather than an edit because the interval produced two findings of its own.
First: the population was never counted again. Measured on
fix/input-domain-correctness, 2026-08-12 — KNOWN_DISAGREEMENTS holds 8
entries, 7 of them understating capital: six FCSM cases at 10.0% and
ORC-280 at 33.3%. The eighth, ORC-142, runs the other way (engine applies
a 373,345.27 mortgage-floor adjustment where the oracle applies 0.00) and its
limb is unrepresentable rather than mis-gated. So "11 entries, 8 understating"
— the figure in docs/plans/test-space-correctness-proposal.md, in
IMPLEMENTATION_PLAN.md P5.41, and in this file's own entry 4 heading — was
wrong on both numbers by the time anyone acted on it. Three of the eleven
(the Art. 121(1) Table 5 family) were discharged by P1.316 hours after entry 4
was written. Entry 4 predicted this in terms: "a register whose size is recorded
in prose is stale the moment the register moves, which is exactly why the fix is
a ratchet and not a sentence." Its own heading is now the worked example.
Second: not one of sixteen entries named an owner. Across both declared
registers — 8 in KNOWN_DISAGREEMENTS, 8 in tests/conformance/classification_table.toml
(D1–D7 including D1b, which is why the parallel register is 8 and not the 7
everything referring to it says) — 0/16 reason strings named the plan bullet
responsible for the fix, against .claude/LESSONS.md B7's explicit instruction
to put one there. All sixteen owning bullets already existed in
IMPLEMENTATION_PLAN.md; nothing linked an entry to one, so a register of
sixteen shipped-wrong numbers had zero accountable owners while every bullet
that would have fixed them sat filed and unreferenced.
- Rule: Not a regulatory escape in itself. The regulatory content at risk is
what the sixteen entries park: CRR Art. 197(1)(b)/(d)/(f), Art. 198(1)(a),
Art. 218, Art. 222, Art. 114(2); PS1/26 Art. 154(4A)(b); CRR Art. 154(4)(c) /
PS1/26 Art. 147(5A)(c); PS1/26 Art. 147(4C)(b)(ii); CRR Art. 147(2)/(3)/(7).
- Origin: tests/oracle/test_oracle.py::KNOWN_DISAGREEMENTS and
tests/conformance/classification_table.toml. Both were built correctly —
tests/oracle/README.md is emphatic that when the engine and the oracle
disagree you adjust neither — and both were built without a size gate.
- Escape class: caught-and-parked. Confirmed against the definition at the
top of this file rather than assumed: "a gate fired, and the record of the
finding became its resting place", fix "ratchet the finding register; give
every parked entry an owning bullet". Both limbs match exactly, and the
prescribed fix is the fix that landed. Worth noting that the class's own
justification — "the shape recurs across at least four parallel registers here"
— is what made the shared mechanism below the right answer rather than four
bespoke ratchets.
- Why every gate missed it: xfail(strict=True) is a real gate in exactly one
direction. It fails when an entry starts agreeing, so a silent fix is
impossible and requirement (b) has always been enforced at the only place that
can enforce it. Nothing anywhere constrained growth, and nothing read the
reason strings at all — they are prose attached to a marker, and prose is not
parsed. The population could therefore rise without limit while every gate
including the two mandatory Tier 2 ones stayed green, which is precisely what
4 → 11 in one batch looked like from the outside: a green suite.
- Gate change: in this change-set, from P5.41. ONE mechanism, two runners.
- scripts/tolerated_findings.py — the shared primitive: a direction-neutral
generic set-diff (diff) and the owner grammar (owner_of / unowned).
Extracted from the supervisory register rather than written fresh.
- scripts/check_parked_registers.py --check — the gate over the two
declared registers, against scripts/parked_registers_baseline.json.
- tests/contracts/test_parked_register_ratchet.py — 13 tests. Three run the
real gate over the real registers, including one that shells the CLI; ten
drive the mechanism synthetically so both directions are demonstrable in
milliseconds.
- tests/acceptance/reporting/test_supervisory_validations.py — all seven legs
now route their set arithmetic through the same diff. Semantics-preserving:
8 passed, unchanged.
Script and pytest, for a stated reason. The four registers split on what it
costs to compute current membership. The supervisory three are measured — a
union over sixteen pipeline runs, and only pytest owns the fixtures that produce
them — so that ratchet stays where the data is. The two declared ones are a dict
literal and a TOML array; reading them is a few milliseconds, so they get a
script, which buys the explicit --update-baseline verb that makes banking a
separate reviewable act from checking.
Deliberately in-suite rather than a CI job. This file's 2026-08-09 entry 1 records a CI-only invocation guard defeated six ways while staying green. A millisecond census belongs where nobody has to remember to run it.
Stronger than the bullet asked: additions are shrink-only, not two-way.
--update-baseline prunes departed ids and refreshes owners and refuses to
add, so banking a new parked finding means hand-editing the baseline — a diff
in a file whose whole purpose is to be reviewed. Every other ratchet in this repo
can be satisfied by re-banking a worse number; this one cannot, because what it
counts is figures an independent derivation has shown are wrong.
Removals stay free, and the docstrings say why: a gate that reddens on a fix
teaches people to stop fixing, and strict=True already forces the entry's
removal in the same change. The residual hole is stated in
scripts/tolerated_findings.py rather than hidden — while a baseline id is
stale it is slack in the addition gate — and closing it would mean gating
removals.
Requirement (c) landed as ownership, not as a grandfathered exception set.
All 16 entries carry an OWNER: P<n>.<n> token; every owning bullet already
existed and no new bullet was filed. ORC-142 → P1.337;
ORC-257/258/275/278/279/280/281 → P1.330; D1/D1b → P1.320; D2 →
P1.321; D3/D4/D6 → P1.303; D5/D7 → P1.322. The grammar is
an explicit token and not a bare P\d+\.\d+ scan, which was tried and rejected
on measurement: _ART_154_4A_B_SCOPE already reads "Since P1.319,
engine/irb/adjustments.py gates on the first two and not the third" — a
historical reference to the bullet that narrowed the gate, not an owner for
what remains. A bare regex would have passed the one entry whose ownership was
hardest to establish and failed the seven whose owner was obvious; it would have
measured prose style, not accountability. That case is pinned by
test_a_historical_bullet_reference_is_not_an_owner.
- Verified red: six, and the first two are the ones this entry's class
demands. All were produced against the real registers — four by writing the
perturbation to disk and restoring it in a finally, so nothing was faked
in-process.
(1) A ninth disagreement added to tests/oracle/test_oracle.py, driven through
the suite (2 failed, 11 passed):
NEW PARKED FINDING (tests/oracle/test_oracle.py::KNOWN_DISAGREEMENTS): 1 entry/entries outside the committed baseline:
ORC-999 (owner: P1.330)
This register may only SHRINK. Every entry is a number we have independent evidence is WRONG and are shipping anyway, so growing the population is a decision, not a side effect.
FIX THE DEFECT. If the finding is genuinely accepted, hand-edit parked_registers_baseline.json to add the id with its owning bullet — --update-baseline will not do it for you.
(2) The owner token stripped from _ART_197_FCSM_ELIGIBILITY — requirement (c),
firing on all seven consumers of the shared reason (3 failed, 10 passed):
NO OWNING BULLET (tests/oracle/test_oracle.py::KNOWN_DISAGREEMENTS): 7 entry/entries name no plan bullet:
ORC-257
ORC-258
ORC-275
ORC-278
ORC-279
ORC-280
ORC-281
An entry with no owner is the review finding, not the xfail (.claude/LESSONS.md B7). Add `OWNER: P<tier>.<n>` to the entry's reason text, naming the IMPLEMENTATION_PLAN.md bullet that owns the fix — and file that bullet in the same change if it does not exist yet.
(3) and (4) are the same two against the other register, proving the mechanism
is shared and not a single-register special case — a ninth
[[known_disagreement]] appended to classification_table.toml gives
NEW PARKED FINDING … D8-synthetic-ninth (owner: P1.303), and stripping D2's
token gives NO OWNING BULLET … D2-large-corporate-test-keys-on-one-entity-type-string.
Both exit 1.
(5) The whole gate before the fix, on the untouched tree, which is the state the
estate was actually in — two NO OWNING BULLET blocks, 8 entry/entries each,
naming all sixteen ids, exit 1.
(6) The shrink-only claim attacked directly, since a refusal that is only
documented is not a refusal. --update-baseline asked to bank the ninth entry:
baseline banked at 16 parked finding(s)
REFUSED to bank 1 NEW entry/entries:
oracle_known_disagreements: ORC-999
Additions are shrink-only. Fix the defect, or hand-edit parked_registers_baseline.json so the decision to ship a known-wrong number appears in the diff.
with the baseline file byte-identical afterwards — checked, not assumed.
What this gate does NOT do, and it is the thing a reader will assume. It
constrains the size of the register. It does not shrink it, and it says nothing
about whether any of the sixteen parked numbers is right. Seven capital
understatements are still shipping today; P1.330 and P1.337 own them. This is a
ratchet, so it fails on movement — the sixteen entries already there move it
not at all, and a reader who takes a green gate as evidence the estate is clean
has read it exactly backwards. Draining is Phase 4's actual objective; this only
stops the queue growing while that happens.
- Lesson: .claude/LESSONS.md B7 is now graduable, and the graduation is
filed rather than performed here — following entry 4's own precedent, since
this file does not own .claude/LESSONS.md and two other agents were editing
the tree concurrently. Both of B7's limbs are executable for both declared
registers: membership (--check on the baseline) and ownership (the OWNER:
token), so its Detect line — "a register entry naming no bullet is the thing to
catch in review" — is no longer a thing to catch in review. What should survive
the trim is the judgment B7 carries and no check can: that registering the
disagreement is the right first move, and that the failure is treating the
registration as the end of it. Keep B8 ("ratchet the accumulator, not a ratio")
as prose in full: this mechanism
applies it — the accumulator is the id set, so a register that grows by one and
shrinks by one is detected as movement — but B8's own worked example is about
choosing the right metric, which no check can express. What also does not
graduate, and belongs here rather than in a lesson: the count in the bullet was
wrong, in a bullet whose entire subject was that counts in prose go stale. The
ratchet fixes the register; nothing fixes a figure typed into a sentence.
2026-08-12 — Three input pathologies produce a plausible number and no signal, and the codes that would report two of them are declared with no producer¶
- Defect: three distinct silent-wrong-number paths, all found in the first
hours of the
tests/robustness/suite existing. Each returns a populated results row a reviewer would accept.
(1) An orphan or null counterparty_reference. Every counterparty-attribute
join in the hierarchy stage is how="left" — the obligor rating/CQS lift at
engine/stages/hierarchy/enrich.py:131-135, the entity-type gate lookup at
:253, and the graph joins at graph.py:388-420 — so a loan whose obligor
cannot be found survives with null obligor attributes, classifies to other,
and takes the 100% fallback risk weight. Measured on CRR, one GBP
1,000,000 senior loan: a CQS 6 corporate is CRR Art. 122 150% = GBP
1,500,000; pointed at a reference that does not exist it returns GBP
1,000,000, a 33.3% understatement on a one-character typo. The direction
reverses with the obligor's true class — a CQS 1 corporate is 20% = GBP 200,000
and orphaning it returns GBP 1,000,000, a 5x overstatement — so the defect
is not conservative in either direction. A null counterparty_reference
reaches the identical fallback, which matters because that is the more common
feed shape (an outer join upstream, or a column never populated).
ERROR_ORPHAN_REFERENCE (DQ005) is declared in contracts/errors.py:93 and
re-exported from contracts/__init__.py, and appears nowhere else under src/.
(2) An unreadable numeric is indistinguishable from an absent one.
FIXED while this entry was being written — see the closing section below;
the measurement here is the pre-fix state, and is what the new gate was verified
red against. EdgeContract.conform_lenient (contracts/edges.py:198-251) cast every
mismatched declared column with strict=False, so Polars turns a value it
cannot parse into a null, missing comes back empty, and no data-quality
error is emitted. Composed with the input contract's own (correct) rule that
null is never a domain violation, the result is a hole. Measured: the same GBP
1,000,000 loan whose drawn_amount arrives as the string "1,000,000.00" — a
plain CSV export with a thousands separator — reports ead_final = 0.00 and
rwa_final = 0.00 against a correct GBP 200,000, with no error. Seven ordinary
export artefacts do this (£1000000, 1 000 000, (1000000), n/a, the empty
string, a truncated 1.0e). CSVLoader reads with pl.scan_csv, which infers
such a column as String, so this was the shipped CSV path and not a contrived
bundle.
(3) A duplicated input row vanishes.
engine/stages/classify/permissions.py:261 de-duplicates on
exposure_reference after the model-permission join — correct for its own
purpose, which is to stop a fan-out when several permissions match — and in
doing so collapses genuine duplicate INPUT rows. Three input loan rows produce
two output rows and no error; a whole file delivered twice loses every duplicate
silently. ERROR_DUPLICATE_KEY (DQ004) exists but is emitted only for the
org-hierarchy multi-parent case (engine/stages/hierarchy/graph.py:474), never
for an exposure table.
- Rule: no single article. As with the 2026-08-12 input-contract entry above,
the exposure is the input domain of every article the engine implements. The
closest named rules are the ones the fallback silently substitutes for: CRR
Art. 122 (corporate SA risk weights by CQS) for (1), and CRR Art. 111 (exposure
value) for (2), where a GBP 1m exposure is reported at GBP 0.
- Origin: (1) engine/stages/hierarchy/enrich.py; (2)
contracts/edges.py::conform_lenient; (3)
engine/stages/classify/permissions.py.
- Escape class: no-gate-exists for all three, and the word exists is doing
precise work in (1) and (3). An error code is declared for each; no producer
is. That is a recognisable relative of the four gate-not-run entries above —
the estate's habit of building the instrument and stopping before wiring it —
but it is not the same class and would not take the same fix. A constant with no
producer is not a gate that failed to run; there is nothing to move earlier. It
has to be written. What it shares with gate-not-run is the appearance of
coverage: ERROR_ORPHAN_REFERENCE in an __all__ reads, to a reviewer and to
every grep, as referential integrity that is enforced.
- Why every gate missed it: the whole estate is organised on one axis — by
regulatory rule, does Art. 123 work? — and every generator in it starts from
a valid portfolio. tests/properties/strategies.py bounds PD to
[0.0003, 0.20] and amounts to at least GBP 10k with every field populated, and
its docstring says so plainly. The oracle, the conformance table, the goldens
and the supervisory register all consume well-formed fixtures. Nothing in 1,081
test files asked what happens when the data is wrong, so a wrong answer to
that question could not be observed.
Two structural specifics are worth recording beyond that general point. First,
tests/acceptance/stress/test_stress_pipeline.py::TestRowCountPreservation is
the only place in the estate that states a count identity between input and
output rows, and it counts a clean portfolio — so (3), which is exactly a
row-count defect, was outside its reach by construction. Second, (2) is
pinned at unit level: tests/contracts/test_edge_contracts.py:368-380
(test_dtype_mismatch_cast_not_raised) asserts that an uncastable value becomes
null with missing == []. That test is
correct about conform_lenient's contract and says nothing about the
end-to-end consequence, which is a GBP 1m exposure reporting zero capital. A
test can pin a mechanism faithfully and leave its consequence unexamined.
- Gate change: tests/robustness/ — a new suite driving the FULL pipeline
(not calculate_branch) over deliberately broken inputs, asserting a triage
invariant rather than a hand-derived number so it scales to as many generated
shapes as the runner will pay for. Six generators: unit-scale errors on every
ratio column, out-of-domain numerics read off the Phase 1 ColumnSpec.domain
declarations, one nulled optional field at a time, unknown enum strings with
case and whitespace variants, sign flips / duplicate keys / orphan foreign keys,
and structural extremes up to 1M rows. tests/robustness/harness.py owns the
invariant; .github/workflows/nightly-robustness.yml runs it nightly under CRR
and Basel 3.1 as separate matrix legs, and both pyproject.toml's addopts and
ci.yml's test job exclude the robustness marker so it never enters the dev
loop.
The invariant has four clauses, not the two the proposal drafted, and the
two additions are what make it usable rather than what soften it. Clause (c)
accepts a table/column-level aggregate error, because _collect_domain_violations
names at most sample_cap=5 rows per column and summarises the rest, and DQ001
/ DQ010 name no row at all — without it the suite reports false failures on
correct behaviour, which is how a suite gets switched off. Clause (d) — the row
produced no output row and no error mentions it — is the failure, and is what
catches (1) and (3). A fifth outcome, collapsed, counts input rows as well
as references, because a duplicate reference IS present in the output and is
therefore invisible to any per-reference identity.
The join back to the input row is on source_exposure_reference, never
exposure_reference: the RE splitter, guarantee substitution and the
facility-undrawn leg all make the latter non-unique per input row, and joining
on it would report every correct split as a vanished row.
Not gated per-PR, deliberately. This is a search, not a regression check; its output is a list of shapes to triage. It is also red on arrival on exactly the three defects above, which is the deliverable rather than a fault in the workflow — a green run would mean they were fixed, not that the search found nothing. - Verified red — each defect against its own new test, before any fix, on the quiet tree:
.venv/bin/python -m pytest tests/robustness/ -m robustness -o addopts= -q
E AssertionError: an exposure whose counterparty_reference 'CP_DOES_NOT_EXIST'
matches no counterparty produced GBP 1,000,000.00 against a correct GBP
1,500,000.00 — a 33.3% understatement — and the run raised NO error at all.
DQ005 ERROR_ORPHAN_REFERENCE is declared in contracts/errors.py and emitted
nowhere in src/.
E AssertionError: an exposure with no counterparty_reference at all produced
GBP 1,000,000.00 and no error names it
E AssertionError: 1 input row(s) unaccounted for (3 input rows -> 2 output rows)
injections: ['loans.loan_reference (duplicated row)']
accumulated error codes: ['<no errors at all>']
[collapsed] loans:LN000 — 2 input rows share this reference and collapsed
to 1 output row(s); no error says so
E AssertionError: drawn_amount='1,000,000.00' could not be cast, was silently
nulled, and produced a zero-capital exposure with no error
All four assertions state what ought to be true and none pins the wrong
number, so a fix turns each one green rather than requiring the test to be
rewritten (.claude/LESSONS.md C1) — which is not merely a claim about the
tests but a measured fact about one of them, since the fourth went green under
DQ014 without being touched. The measured control values are asserted
first in each test, so a failure is unambiguously about the missing signal and
not about the risk weights having moved underneath the test.
- Defect (2) is CLOSED, by the gate rather than around it. The eight
test_cast_failures.py tests were written red against the measurement above and
handed to the concurrent Phase 1 work as an acceptance check. DQ014
ERROR_UNREADABLE_INPUT_DTYPE now reports a column supplied in a dtype whose
cast is destructive: seal_lenient returns LossyCast findings alongside
missing, the loader turns them into one error per (table, column)
(engine/loader.py:223-234), and tests/fixtures/raw_bundle.py routes them into
the bundle's error list so an in-memory bundle carries the same load-boundary
errors a production load would. All eight tests flipped green with no edit to
any of them, which is the strongest available form of this evidence: the gate
was written first, observed red, and turned green by the fix rather than by being
rewritten. Deliberately NOT routed through DQ003 ERROR_TYPE_MISMATCH, which
Phase 0 retired as unfirable, nor through DQ001 — a value that is present and
unreadable needs the feed RE-SENT where a missing column needs it EXTENDED, and
test_an_uncastable_value_is_not_reported_as_a_missing_column pins the
distinction.
- Defects (1) and (3) are now CLOSED as well, and by the same route: the four
remaining tests were turned green by the fix, with no edit to any of them.
DQ005 has a producer, DQ004 has an exposure-table producer, and both are read
off DECLARATIONS in data/schemas.py rather than hand-written per column —
TABLE_FOREIGN_KEYS (a new ForeignKey declaration alongside NumericDomain
/ EnumDomain) and TABLE_UNIQUE_KEYS, both consumed by
contracts/validation.py::validate_referential_integrity and
::validate_duplicate_keys. Three decisions in that fix are worth recording
because each was a fork where the obvious move was wrong:
The join stays how="left". The recommendation above ("emit DQ005 from
the counterparty-enrichment join by counting left-join misses") was followed in
substance and not in location. The detection sits at the INPUT gate, which runs
on both pipeline entries, because the information is strictly richer there: at
the join the miss is a null obligor attribute, indistinguishable from an
obligor row that exists and has no rating, and the reference that was supplied
— the one thing an operator needs to repair the feed — has already been
consumed. Nothing in the fix drops a row: an exposure that has left the
portfolio is worse than one priced off a fallback, because its capital is gone
and no total says so.
Null and orphan carry DIFFERENT codes, against this entry's own
recommendation that they "should carry the same code". They reach the same
engine fallback, which is exactly why the distinction has to be drawn at the
gate or not at all — downstream both are a null attribute and the information
is gone. But they are repaired in different files: an orphan needs the PARENT
feed extended or corrected (DQ005), a null needs THIS row's column populated
(DQ001, absent_reference_error, category DATA_QUALITY to keep it apart from
the seal's missing-COLUMN DQ001 under SCHEMA_VALIDATION). One code would have
sent an operator looking in the wrong file.
DQ004 is uncapped here, breaking the module's own sample_cap contract,
and deliberately. The domain gates sample a property of a COLUMN — naming any 5
of 900 out-of-domain rows locates the repair — whereas a duplicate key is a
property of a ROW, and a sampled duplicate leaves every un-sampled row exactly
as unaccounted-for as it was before the gate existed. That is also precisely
what tests/robustness/harness.py encodes by refusing to let clause (c) excuse
a collapsed row. The population is bounded by the number of DISTINCT
duplicated keys: zero on well-formed input, equal to the corruption on broken
input.
Cost of the two new materialisations, measured on a synthetic portfolio where every reference resolves and every key is unique (so both checks pay their full scan and emit nothing): 1.06% of a full pipeline run at 100k loans / 10k counterparties, 0.63% at 1M loans. Both live in the input gate, which already collects per table; no collect was added to a lazy stage.
Two fixture repairs were needed and neither loosened an assertion.
tests/unit/test_loader.py::test_scrub_and_validate_returns_empty_for_valid_data
passed pl.LazyFrame() for loans and facilities — a zero-COLUMN frame is not a
zero-ROW one, and the seal's literals broadcast it to a single phantom
all-null row, so the "valid data" bundle contained an exposure with no
obligor. It now declares a key column and is genuinely empty.
tests/contracts/test_validation.py::test_clean_bundle_raises_nothing declared
a rating for a counterparty absent from the bundle and gave its loan and
facility no obligor at all; it now names an obligor that exists. Both were
asserting that a broken bundle is silent.
- Lesson: a declared error code with no producer is negative coverage. It
reads as enforcement to every reviewer, to every grep and to every import-graph
tool, while enforcing nothing — the same shape as the 402 lines of unreachable
validators in the entry above, one level further out. Check 20
(check_guard_reachability) closed the unreachable-validator case; the
unreachable-code case is its exact analogue and is mechanically checkable in
the same way: every ERROR_* constant in contracts/errors.py should either
have a producer under src/ or be explicitly listed as reserved, as Phase 0 did
for DQ003. Two of the three defects in this entry sat behind a declared code;
the third had no code at all until DQ014 was written for it, so the check would
have flagged exactly the two that were flaggable.
That is the graduation candidate this entry files; it is not performed here
because scripts/arch_check.py was being edited by another agent in the same
tree and check 20's own entry records the same reason for the same restraint.
2026-08-17 — Two fixes for one taint finding, neither of which touched the flow the tool reported¶
- Defect:
pythonsecurity:S2083(BLOCKER) onscripts/generate_regulatory_tables.pysurvived two fixes written to close it. Commit2b1be086introducedTARGET_PATHSand removedtargets.items()as the write-loop iterator; its successor went further and bound the loop variable directly off the constant at the sink, duplicating a comparison to do so. Neither moved the finding, because the flow SonarCloud actually reports is not about the path. It is two steps, identical on both versions (line numbers frommaster/ the PR head):
source line 821 / 848: path.read_text(...) in _splice → the file CONTENT
sink line 703 / 729: path.write_text(desired, ...) → "a malicious value
can be used as argument"
_splice reads the target to preserve the hand-written prose outside the
GENERATED markers; that content becomes targets[path], then desired,
then the write's data argument. The taint is the content. Writing
file-derived content to a compile-time-constant path is not path injection,
so there was never a structural fix to find — the flow is the feature.
- Rule: Not a regulatory escape. No RWA number is affected; the script is
developer/CI codegen whose only external input is a boolean --check flag.
- Origin: the diagnosis, not the code. sonar-project.properties asserted
that "scripts/ findings are always structural" and cited this very file as
the worked example, so both fixes started from a premise the tool's own
evidence contradicts.
- Escape class: wrong-premise. The premise ("the write path is taken from
a mapping whose values were read off disk, therefore S2083") was wrong and was
faithfully implemented twice, which is exactly what the class describes. Note
that test-shared-the-assumption also fits the symptom —
test_generator_write_paths_do_not_come_from_the_rendered_mapping passed
against an unfixed finding, because it encoded the same wrong sentence — but
it prescribes re-anchoring the test, and re-anchoring a test to a false
mechanism produces a better-engineered wrong gate. The first attempt at this
entry made that mistake: it classed the escape test-shared-the-assumption
and shipped an AST gate asserting the write path is bound from the constant,
a property that is real, cheap, and irrelevant to what was flagged. The class
has to name the wrong belief, not the instrument that agreed with it.
- Why every gate missed it: only SonarCloud can see this flow, and nothing
in the loop ever read what it said. The finding arrives as a rule name plus
one line number, and both fixes were designed from that — the rule's title
("I/O function calls should not be vulnerable to path injection") points at
the path, and the flow points at the content. The full source → sink was
available the whole time and never fetched: SonarCloud uploads SARIF to GitHub
code scanning, so codeFlows is two gh api calls away and needs no
SonarCloud credentials. Local gates cannot substitute — ruff and
arch_check have no taint model, and the freshness tests measure output bytes
and were green throughout, correctly, since both refactors were
behaviour-neutral.
- Gate change: the misdirecting premise is deleted at its source. The
S6549 / S2083 note in sonar-project.properties no longer claims
scripts/ findings are always structural; it records this flow verbatim, and
carries the gh api incantation that retrieves any flow from the SARIF
without SonarCloud access. .claude/LESSONS.md gains the trap in
Trap/Why/Detect form, Detect being that command. The finding itself is
resolved as Accepted in the SonarCloud platform (issue
AaAMt-IEKije7nS9AwhB), which is the only mechanism available: taint findings
cannot be suppressed from the properties file under Automatic Analysis, as the
same note already recorded.
This is prose-tier and that is a real weakness, so the stronger form is filed
rather than hand-waved: scripts/sonar_flow.py, a helper that takes an alert
number or rule key and prints the source → sink chain, would make the
retrieval a command nobody can skip rather than a paragraph they must
remember. It is not built here because the fix for a wrong premise is to
delete the premise, and shipping a new script alongside it would put a second
unreviewed thing in the same change-set.
- Verified red: the retrieval run against the pre-fix commit
(6456d808, master) returns the true flow, contradicting the premise that
shipped in both fixes:
RULE: pythonsecurity:S2083 in scripts/generate_regulatory_tables.py
step 1: line 821 Source: a user can craft an HTTP request with malicious content
step 2: line 703 Sink: this invocation is not safe; a malicious value can be
used as argument
Line 821 on master is for line in path.read_text(encoding="utf-8") inside
_splice — the content read, not a path expression. Two minutes of this
against the original finding would have prevented both fixes; that is the
sense in which the gate is observed red.
- Lesson: a taint finding is a flow, not a line. The rule name tells you
the sink family and nothing about the source, and a sink can be reported for a
tainted argument rather than a tainted path — which inverts the whole
remedy. The tell here was misread twice as evidence: read_text sitting
unflagged beside a flagged write_text looks like proof that provenance is
the variable, when it only means the read takes no tainted argument. Fetch the
codeFlows first; every fix designed from the rule title alone in this repo
has failed.
2026-08-30 — The two reconciliation surfaces disagreed about what an exposure is, and 135 tests could not express the disagreement¶
- Defect: This estate has two reconciliation surfaces, and they answered "is
this the same exposure?" differently. The exposure-grain component
reconciliation has always collapsed a split exposure's legs onto the base
reference —
ReconciliationRunnercallsaggregate_to_key_grain(analysis/reconciliation.py:178), which runs_coalesce_to_parent(analysis/_collapse.py:73-75). The cell-grain template waterfall inanalysis/return_recon.pydid not. The grain inconsistency is the defect and the wrong number below is its symptom, and that ordering is what makes the collapse the right answer rather than merely an effective one: "is this the same exposure" is a base-grain question, and the waterfall was asking it of legs.
The symptom. Our sealed ledger splits one exposure into several legs — a
guarantee into L1__G_BANK / L1__REM, a mixed property into M1_rre /
M1_cre, a facility into an _UNDRAWN row — each stamped with the pre-split
reference, while a projected legacy extract carries the whole loan under that
original reference and no base reference at all. Keyed on exposure_reference
the two sides shared no key, so a cell where both engines agree to the penny
reported the whole loan as missing from BOTH sides at once. The four-way
additivity contract could not catch it, because two equal and opposite
population terms sum to the same 0.00 as no terms at all:
CellDecomposition.reconciles stayed true throughout.
What the fix did, stated narrowly, because the tempting sentence is wrong.
It did not surface a finding that was hidden. The delta was always the cell's
delta and reconciles was always true. What was wrong is the attribution:
on the measured differing portfolio the signed terms explained a 5,400
difference as +6,918,900 ours-only against −6,913,500 theirs-only — two
large, equal and opposite scope terms netting to the right number for the
wrong reason, sending an analyst hunting 6.9m of missing scope when the finding
is 1,800 of expected-loss difference per cell. The claim that survives someone
checking it is the narrow one: the fix deleted a fabricated scope finding
that never existed, and re-attributed a real finding from scope to valuation.
It hid nothing. No cell's delta changes anywhere.
And it traded no detection away — measured, not argued. The obvious worry
is that collapsing the key blinds the waterfall to a genuinely dropped split
leg. Three cases were measured on C 08.03 corporate row 0080 col 0090 under
both frameworks, ours = the guaranteed leg (60k) only against theirs = the
whole loan (100k): (A) an agreeing split, where nothing is wrong; (B) a leg
moved to another row; (C) a leg genuinely dropped, which is a real scope
failure. After the fix B and C are byte-identical at the cell (measurement
−40,000, one key, reconciles true). Under the old keying B and C were also
byte-identical — and case A raised the same alarm, population_ours_only
100,000 against population_theirs_only −100,000. The old population term
fired on every split exposure in the book whether or not anything was lost, so
it had zero discriminating power for a dropped leg: it did not detect case C,
it drowned it. Discriminating power went from one signal across three cases to
two — silent on A, breaking on B and C. The closing argument: the only book in
which the old alarm carried information is one with no split exposures at all,
and in that book case C cannot arise, because there is no split leg to drop.
Case C is caught elsewhere and already shipping — post-collapse the base exists
on both sides, so it lands in the component reconciliation's break as a
40,000 value break on a named exposure. What separates "moved" from "dropped"
in the waterfall is sheet-wide non-conservation of the terms, and nothing
presents that to an analyst as such; that residual is pre-existing, is not
introduced here, and is filed rather than closed.
Not a corner case: 21 of 128 bases in the fixture estate hold more than one
leg, carrying 13.9% of RWA.
- Rule: Not a regulatory escape. No RWA figure and no published cell moves.
What was wrong is the explanation attached to a difference between two
returns — which is the evidence a migrating firm uses to decide whether its
parallel run ties out, and which of the two engines to go and look at.
- Origin: src/rwa_calc/analysis/return_recon.py, shipped in v0.3.28
(2026-08-29) with the four-way waterfall. The defect is as old as the feature.
- Escape class: path-never-exercised. The gate ran — 135 tests across
tests/unit/analysis/test_return_recon.py and
tests/unit/ui/test_views_return_recon.py, exercising this exact code — and no
fixture reached the defect, because every leg in both suites set
source_exposure_reference == exposure_reference. A split leg could not be
expressed in the fixture vocabulary, so no assertion about one could be
written, right or wrong. The class's prescribed fix is "build the portfolio,
register it", and that is what landed.
- Why every gate missed it: three independent reasons, and only the first is
the escape class.
The fixtures could not express the input. Every leg carried its own reference as its base reference, which is the one shape in which the buggy key and the correct key agree.
The output contract was satisfied by the wrong answer. reconciles scores the
four terms against the reported delta, and it is exact — but it is a
statement about a sum, and the defect's signature is two equal and opposite
terms. Every additivity assertion in the suite passed on the defective output,
correctly. An identity that a defect satisfies is not a weak identity; it is an
identity about the wrong quantity.
The one hazard adjacent to this fix announced itself at runtime for the
feature's whole life and nothing asserted on it: row_migration logs a
WARNING saying "the first is used and the matrix may be wrong". That is the
_group_legs collapse hazard, and it was carried in prose, and in a log line,
and in no test. A log line is not a gate.
- Gate change: four, all in this change-set.
- Split-leg portfolios, built and asserted —
tests/unit/analysis/test_return_recon.pyandtests/unit/ui/test_views_return_recon.py, 135 → 163 tests (c898d950). Split exposures on the IRB C 08.03 path and on the standardised C 07.00 path (test_a_standardised_re_split_pairs_against_the_legacy_whole_loanand three siblings — the real-estate and facility splits, which are the commonest split shapes in a migration book and which no earlier fixture reached, plus the first split-leg coverage ofsheet_placement); bothKEY_COLUMNSsettings (test_both_key_column_settings_name_the_same_grain); the plan-frame presence contract (test_every_recon_template_plan_frame_carries_the_base_reference); andtest_the_migration_matrix_conserves_a_split_exposures_money, a conservation detector for the_group_legshazard, which previously had none. - The mutation evidence made durable —
tests/mutations/(909e69cb): five isolating pytest plugins that each change exactly one thing in production code and restore it in afinally, plus a README carrying the measured red set of each. They lived in a job scratch directory and would have vanished with the session; this file's discipline closes a defect on evidence a gate was observed red, so that evidence has to outlive the run that produced it. - Fixture adequacy asserted inside the test — both assertions open
test_the_migration_matrix_conserves_a_split_exposures_money, and are greppable by their messages rather than by position (line numbers as atc898d950, wheregit showwill still find them once the file moves):bases.count(...) == 2attest_return_recon.py:1528— "the group holds no split exposure — collapsing the key would be a no-op and this test would prove nothing" — and_SPLIT_G_LEG.ead != _SPLIT_REM_LEG.eadat:1532— "equal legs make.first()undetectable". The pre-existing_combinedfixture fails the first of those, which is precisely why its sibling test was vacuous for the life of the feature. This is the batch's real graduation: a test that states the property its own fixture must have cannot silently become a no-op. - The drill-through's own suite —
tests/unit/ui/test_views_recon_placement.py(513937ef,d1538f0b), 348 → 355 tests, with six further isolating mutations each reddening exactly its own detector.
What these fixtures do not cover, said here rather than left to be
discovered. The decomposable-cell census taken over the projected legacy
mapping yields 1,820 decomposable cells on C 07.00 and none at all on
C 08.01 or C 08.03 — the mapping is too thin to make either template
decomposable on that run. So any sentence of the form "measured across all
three templates" is false of the census: it is a C 07.00 measurement. The new
fixtures do reach C 08.01 and C 08.03, by construction (5,400 of expected-loss
difference on C 08.01, 843,600 of row placement on C 08.03) — which is a
different and narrower kind of evidence, hand-built cells rather than a census.
- Verified red: the pre-fix keying, restored as an isolating mutation —
mutate_prefix_key swaps _comparison_key back to pl.col(key_column) and
restores it in a finally, so nothing on disk is modified while it runs. It
ships, and it reproduces against the finished 163-test baseline at c898d950:
PYTHONPATH=tests/mutations uv run pytest \
tests/unit/analysis/test_return_recon.py tests/unit/ui/test_views_return_recon.py \
-q -p mutate_prefix_key
22 failed, 141 passed
population_ours_only 100000.0 where 0.0 expected
measurement 0.0 where -30000.0 / -40000.0 expected
Those 22 are every split test plus both key_column guards, and they are the
measurement of the gate gap: on the pre-fix tree the same code produced no
failures at all, because no fixture could express a split leg. The four sibling
probes, from the same committed table — mutate_drop_base_ref 24 red,
mutate_no_third_rung 4, mutate_no_presence_filter 2,
mutate_collapse_group_legs 2.
An earlier red was taken during implementation, against a mid-tree at 145 tests that was never committed. It is deliberately not quoted here: nobody can ever check it, and a figure no one can check does not belong in this field.
The last two are the ones this entry turns on, because before their guards existed both mutations were completely silent.
mutate_no_presence_filterremoves the presence filter from_key_rungs. That filter is on the hot production path — the projected legacy plan frame has nosource_exposure_referencecolumn at all — so removing it raisesColumnNotFoundErroron every real reconciliation, and the suite was blind to it because every fixture pinned the column intoschema_overridesas a typed null. Null and absent are different code paths, and both occur on the same side at once.mutate_collapse_group_legsis the "tidy up the two inconsistent key expressions" change the module now forbids in prose._group_legsprices with.first(), so collapsing distinct split legs onto one base key keeps one leg's money and discards the other's: 40,000 ofrwa_finaland 400,000 ofead_finalagainst the conservation invariantrow_migration's own docstring states. Reproduced on two portfolios with different totals but identical losses to the penny, which is what shows the loss is one specific leg rather than something proportional to portfolio size.
Two caveats belong in this field rather than in a footnote. First,
c898d950's commit message lists six mutations and gives the pre-fix keying
as 20 red; the committed README lists five plugins and gives 22. Two of the
six — a reversed coalesce, and a break of the _key_money / _side_keys
lockstep — were not preserved as plugins, so they are not reproducible and
should not be quoted. Where the two disagree, the README is the durable artifact
and its figures are the ones to cite. Second, the discrepancy was not
re-measured for this entry: analysis/return_recon.py was under concurrent edit
in the same worktree while it was written, and running a mutation probe against
a mid-edit file is one of the four false-green mechanisms recorded below
(.claude/LESSONS.md C12).
What a later reader can reproduce, and what they can only read — because this
entry had to be made self-contained before it could close anything. Every
field that closes the defect is reproducible from the repository alone: the
escape class (check out 358e4ce2; the two pre-fix suites hold 135 tests in
which every leg sets source_exposure_reference == exposure_reference), the
four gate changes (all committed files, with the two adequacy assertions pinned
to c898d950), and the verified red (five plugins that ship, a baseline that is
a commit, the invocation printed above). Nothing in that set needs anything
outside git.
The supporting measurements — the three-case A/B/C refutation, the
+6,918,900 / −6,913,500 attribution terms and the 5,400 they concealed, the
843,600 of row placement, the 1,820-cell C 07.00 census, and the two-portfolio
reproduction of the 40,000 / 400,000 — are stated in full in this entry's own
prose above, and that is deliberate. Their working record was batch state
under .claude/state/, which .gitignore excludes wholesale, and the archive
path that convention names sits inside the ignored directory, so none of it
reaches the repository and all of it dies with the worktree. They are therefore
recorded findings rather than figures anyone can re-derive, and this entry is
their only durable home. Do not add a citation pointing at batch state to
"support" them; there will be nothing on the other end.
- The same escape, one level out, in the code this batch wrote. The
drill-through built on top of this fix (513937ef) shipped with five defects
its author found by reviewing its own new code, each reproduced before being
fixed, and all 340 existing tests passed against two of them — which is the
finding, not the fix. A composite reconciliation key was token-matched segment
by segment, so a segment coinciding with another exposure's reference merged
that exposure's rows in, fabricating a row the loan never reached and
destroying a real band move by adding a second provable leaf. A || inside a
legitimate reference under a single-column mapping resolved half a reference to
a whole loan, silently. A scheme-relative Referer (//evil.example/...)
skipped the same-origin test entirely — an empty scheme made the guard skip
itself — and a malformed one raised and 500'd the page. The tri-state
parent-row flag was flattened at the render boundary, so a row False on ours
and None on theirs claimed "provably contains and duplicates no other row"
for both sides. And the single-leaf rule was unasserted: C 08.01 rows 0020 and
0070 are different cuts of one book rather than a tree, so an exposure can
reach two provable leaves — under the mutation each side picks its leaf from
unguaranteed Polars row order, the two can differ by chance, and the panel
renders a fabricated band move on a byte-identical exposure. A sixth followed
in d1538f0b: a key that cannot be read was rendering as "this exposure
reached no instrumented row", which is a different and false claim. Four of the
six are one shape — two distinct facts collapsing into one blank — which is the
shape of the defect this whole batch exists to fix.
- Lesson: four entries added to .claude/LESSONS.md — B9 (a log line is
not a gate, and neither is an alarm that always fires), C10 (careless about
scope, not about facts), C11 (a fixture is a claim: assert its adequacy and
check its label against its cause) and C12 (four ways a green mutation probe
lies) — plus a fold into B1 for the null-versus-absent mirror image. C11 and
C12 are the two closest to graduating: C11 already has an executable form in a
single test, and C12's second, third and fourth mechanisms are all mechanically
checkable from a shared plugin base in tests/mutations/. Both are filed in the
ledger's candidates list.
2026-09-05 — Basel 3.1 C 07.00 deducts on-balance-sheet netting twice, and no registered portfolio had a netting agreement¶
- Defect: on the Basel 3.1 OF 07.00 template the netted deposit is deducted
once in column 0035 "(-) Adjustment due to on-balance sheet netting" (so column
0040 is already net of it) and again inside column 0130 "(-) Financial
collateral: adjusted value", because CRR Art. 219 turns the netted deposit into
synthetic cash collateral that the CRM stage merges into the ordinary
collateral frame, and
reporting/corep/c07.pysums the wholecollateral_adjusted_valuecarrier into 0130. Column 0150 (fully adjusted exposure value) is therefore understated by the netting amount while column 0200 (exposure value, fromead_final) is correct. Measured on the newnettingportfolio (five unrated corporates, all drawn, GBP): col 0040 19,000,000, col 0130 −3,000,000, col 0150 16,000,000, col 0200 19,000,000. Under CRR the template has no column 0035, so the chain nets once and ties. The defect is pre-existing and independent of PR #487's perimeter change: with the Feature disabled (only the same-counterparty agreement nets) the same sheet reads col 0150 20,000,000 against col 0200 21,000,000. - Rule:
boe_b0471(ERROR — "{c0200} = {c0150} − 0.9·{c0160} − 0.8·{c0170} − 0.6·{c0171} − 0.5·{c0180}", the end-exposure-value derivation) andboe_b0556(WARNING — "{c0200} ≤ {c0150}"), both live on OF 07.00.01.01;boe_b0555(an OF 07.00 collateral/exposure-value relation) left the vacuity population on the same run and now asserts. - Origin:
src/rwa_calc/reporting/corep/c07.py— col 0035 was added underif is_b31asSum("on_bs_netting_amount")while col 0130 kept summing the undifferentiatedcollateral_adjusted_value. The carrier that would separate them — the post-haircut, post-mismatch adjusted value of theNETTING_synthetic collateral per beneficiary — does not exist on the sealed ledger;on_bs_netting_amountis the pre-haircut pro-rata allocation, which coincides with it only for same-currency, maturity-matched deposits. - Escape class:
path-never-exercised.boe_b0471is a live ERROR rule and the register ran on every commit, but no portfolio inRUNScarried a netting agreement at all — the template-cell census had classifiedcorep/c07_00/0035asNO_FIXTURE("no fixture supplies a netting agreement") for the template's whole life, so the rule was evaluated over sheets on which 0035 was structurally null and the double deduction summed to zero. Only three fixture builders in the estate set anetting_agreement_reference(P1.238, P1.241,r1_negative_gross), none of them a reporting portfolio. - Why nothing else would have caught it: the unit tests for C 07.00 build
their frames from the sealed carrier names and never populate
on_bs_netting_amountandcollateral_adjusted_valueon the same row; the golden portfolios have no deposits; the P1.238/P1.241 acceptance twins assert RWA, not template cells; and the engine's own EAD path is correct (it nets once), so no engine-level test can see a reporting-layer double count. - Gate change:
tests/fixtures/reporting_netting_portfolio.pyregistered inRUNS(two runs,crr/b31×netting) and captured as goldens undertests/expected_outputs/reporting/netting_{crr,b31}/, pinned bytests/acceptance/reporting/test_reporting_netting_golden.py(per-leg netting, C 07.00 rows 0010/0070, the Basel 3.1-only col 0035, the CRR absence of it, the single CRM016 record, and a two-limbed Feature test). Landed in PR #487. The two breaks are banked invalidation_known_breaks.jsonwith this mechanism as their written reason, so the register shrinks by two the day the fix lands. Code fix deferred, design recorded here: seal a per-exposureon_bs_netting_adjusted_value(the adjusted value of theNETTING_collateral rows after Art. 223 haircuts and Art. 237-239 mismatch scaling) from the CRM stage throughcontracts/edges.pyand the re-split carriers; under Basel 3.1 report it in col 0035 and subtract it from col 0130, so 0150 = 0200 for an all-drawn book in every currency case; leave CRR untouched (no 0035).C 08.01col 0035 sits beside the LGD-collateral columns 0170-0210 and needs the same review. - Verified red:
uv run pytest tests/acceptance/reporting/test_supervisory_validations.py -n 4 -qon the unchanged reporting code with the two netting runs registered —2 failed, 6 passed:b31/boe_b0471 [ERROR] left = 19,000,000.0000 vs right = 16,000,000.0000 on 2 cell(s)atOF07.00.01.01[corporate][r0010]and[r0230], andb31/boe_b0556 [WARNING] left = 19,000,000.0000 vs right = 16,000,000.0000 on 3 cell(s). The golden test's own cell pins (col 0200 == 19,000,000) pass on the same run, which is the point: the exposure value is right and the intermediate column is wrong, and only the published identity between them sees it. - Lesson: LESSONS B5 in its plainest form — the column was dead, the census
said so, and the ERROR rule over it was green for the life of the template.
The transferable rule is already in the file: a
NO_FIXTUREclassification in the template-cell baseline is a list of gates that are not running, and the change that lights one should expect the register to move.
2026-09-05 — A window nested inside a window made the classifier eleven times slower, and the one instrument that measured it asserted nothing¶
- Defect: commit
8ec7d302(P1.320, "count each facility limit once in the QRRE obligor aggregate", authored 2026-08-15) rewroteengine/classify/subtypes.py::qrre_obligor_aggregate_limit_exprso that acum_sum().over([counterparty_reference, parent_facility_reference])ordinal and amax().over([...same keys])group limit sit inside the input ofdeduped_limit.sum().over("counterparty_reference"). Polars evaluates a window's input inside the outer group-by context, so both inner windows re-run once per obligor group: per-row cost stops being a function of row count and becomes a function of row count times group count. It first shipped in v0.3.27 on 2026-08-28 and was present through v0.3.32.
Measured 2026-09-04, same machine, back to back, CreditRiskCalc.calculate()
on the repository's 100k-counterparty synthetic book (373,568 result rows, CRR
standardised):
| v0.3.26 | v0.3.32 | |
|---|---|---|
| classifier stage | 495 ms | 5,640 ms |
| total run | 11.8 s | 18.6 s |
With only that one function reverted, v0.3.32 measured 9.4-11.3 s across
three runs against 9.5-11.8 s for v0.3.26, so the function accounts for the
whole regression and nothing else in the release range does. The expression
alone, in one with_columns at 800k legs: 46 ms in the v0.3.26 form
against 45,851 ms in the shipped form, growing roughly 4x per row
doubling. And the decisive control — each window alone costs ≤151 ms at
400k rows, so neither window is expensive and only the nesting is. A user
reported a real book going from under one minute to over three; the measured
curve predicts that at roughly 1.7M legs, which is a prediction about their
shape rather than a measurement on their data, and is recorded here as such.
- Rule: Not a regulatory escape. No RWA figure moves, in either regime,
on any portfolio. The P1.320 rewrite was regulatorily correct — CRR
Art. 154(4)(c) caps the aggregate nominal exposure to a single individual, so a
facility's limit must be counted once however many legs it carries — and the
fix below preserves it exactly. What escaped is cost, not capital.
- Origin: src/rwa_calc/engine/classify/subtypes.py, commit 8ec7d302,
2026-08-15. Released 2026-08-28 in v0.3.27 and live for five releases.
- Escape class: no-gate-exists. The instrument existed and the assertion
did not: the CI benchmarks job ran test_classifier_100k on every push from
the day the defect landed, recorded the 11x stage regression every time, and
uploaded it — so gate-not-run is the wrong label, because the job did run
and simply had no failure condition to reach.
- Why every gate missed it: three independent reasons, none of them a weak
assertion.
Nothing in the estate is big enough for the defect to be visible. Every fixture and golden portfolio is under 40k rows, where the penalty is about 60 ms — comfortably inside the noise of a test run. The defect's signature is a curve, and a single-point measurement cannot carry one.
The tests that cover this function pin behaviour, not cost. The P1.320 unit and acceptance tests assert the aggregate limit is right, and it is right; they pass identically on both forms of the expression, before and after. A correctness gate cannot fail on a performance defect however thorough it is, which is why the gate change below is a scaling assertion rather than more coverage of the same property.
The benchmark job produces a number and compares it to nothing.
.github/workflows/ci.yml's benchmarks job runs tests/benchmarks under
-m 'benchmark and not slow' and uploads benchmark-results.json as an
artifact. There is no baseline, no ratchet and no threshold, so the job cannot
fail. The 11x classifier regression was therefore recorded on every run from
2026-08-15 onwards and read by nobody. And the marker is deselected from both
places a developer would otherwise meet it — pyproject.toml's addopts for
the dev loop, and CI's own test job — so the only path to the number was
downloading an artifact from a green build. This is the estate's standing habit
in its purest form: build the measurement, ship it, wire it to nothing.
- Gate change: two, both in this change-set, and they close different halves
— one measures the cost, one forbids the shape.
tests/unit/classifier/test_p1_320_qrre_aggregate_scaling.py— drivesclassify_exposure_subtypesat 100k and 400k synthetic legs in the default dev loop (no marker, so it is not deselected anywhere) and asserts the 4x frame costs less than 8x the small one and under 3 s in absolute terms. The ratio is the assertion that matches the defect's shape; the absolute budget stops a uniformly slow machine passing a quadratic curve. ItsTestTheMeasuredPathIsLiveclass carries the adequacy assertions — the window keys are non-degenerate (multi-leg obligors, multi-facility obligors, and both null slices present), and the QRRE limb both fires and discriminates at each size, anchored toExposureClassrather than to a hand-written list — so the file cannot silently degrade into timing a dead path. Those assertions hold on the pre-change engine too, by design: they are the premise of the timing, not part of it.scripts/arch_check.pycheck 21 —check_no_nested_window_expressions— an AST scan ofengine/that fails any.over()whose input contains another.over(). It follows local-name bindings within one function body, because the shipped defect bound the nested expression to a local first and a scan of the call chain alone would not have seen it; each name resolves to the last binding that ends strictly before the outer window's own line, so a self-rebinding statement cannot make a window find itself. Only the input is scanned, not the partition arguments, since the input is what Polars re-evaluates per group. No allowlist. It is contract-tested bytests/contracts/test_nested_window_gate.py, which pins the check's registration in thearch_checkrun, the two flagged shapes (direct nesting and nesting through a local), the clean two-step remedy, and — the reason the source-order rule exists at all — the real false-positive shape inengine/supporting_factors.py, whose single legitimate window sits in a statement that rebinds the frame name and which a source-order-blind resolver flags as nested against itself.- Verified red: taken on this worktree at
cb59e53cwith the two new test files present and before the engine rewrite, so both gates are observed failing against the code that actually shipped.
The scaling test, 2 failed, 4 passed in 19.94s — the ratio limb and the
absolute limb, from the same pair of measurements:
test_classifier_cost_scales_linearly_in_row_count
assert 18.3 < 8.0
(100,000 rows: 0.281s -> 400,000 rows: 5.129s)
test_classifier_stays_within_the_absolute_budget_at_scale
assert 5.129 < 3.0
The contract test was 10 failed in 7.96s before check 21 existed, which is
the trivial red. The one that matters came next: with the check implemented and
the engine still on the shipped expression, it named exactly one violation in
the whole engine tree —
src\rwa_calc\engine\classify\subtypes.py:374: nested Polars window -- this .over()'s
input contains another .over(), which Polars re-evaluates once per outer group.
Compute the inner result as its own column in a preceding with_columns and read it
back with pl.col()
— with the contract suite at 1 failed, 9 passed, the single failure being
test_engine_has_no_nested_window_expressions. A check that finds one instance
and no others across the engine tree is the useful outcome: it is neither
vacuous nor a false-positive generator, and the one shape that could have made
it the latter is pinned as a test.
- The fix, stated narrowly: the same expression tree, split across two
with_columns. The per-leg deduplicated contribution is materialised as the
scratch column _qrre_deduped_limit, the obligor sum reads it back with
pl.col(), and the helper is dropped before the stage returns, so the classify
exit schema is unchanged. After the rewrite the same measurement gives
0.095 s at 100k rows and 0.382 s at 400k, a ratio of 4.02x against a 4x row
increase. The aggregate is bit-identical to the shipped form on a 400k-leg
frame — 0 mismatches, 20,581 QRRE rows on both sides — and the P1.320 unit and
acceptance tests pass unchanged. Both @cites("CRR Art. 154(4)") and
@cites("PS1/26, paragraph 147") are retained, and the function keeps its
name so tests/contracts/data/citation_snapshot.json does not move.
- Lesson: graduated on the first occurrence, so there is no prose entry.
The transferable rule — a nested Polars window is quadratic in group count,
and no correctness gate can see it — is mechanically checkable from the AST,
which is the strongest form available, so it lands as check 21 with a row in
the .claude/LESSONS.md Graduation ledger rather than as a Trap / Why /
Detect entry that would have to recur before earning a check. The residual
worth naming, because check 21 does not cover it: the benchmark job still
compares against nothing. Check 21 forbids one known-quadratic shape and the
scaling test guards one stage; neither would catch a different construct
slowing a different stage by the same order. Ratcheting benchmark-results.json
against a stored baseline is the general form, and it is not built here.
2026-09-05 — A matched short-dated interbank pair lost its whole netting benefit to a mismatch that did not exist¶
- Defect: an F-IRB loan to an institution, fully offset by a negative-balance
deposit under one on-balance-sheet netting agreement, in the same currency and
with the SAME maturity date, reported LGD 45% and full RWA.
on_bs_netting_amountwas populated (the pair pooled) andtotal_collateral_for_lgdwas 0 (the syntheticNETTING_cash row was zeroed). Reported by a user; confirmed by the operator on 2026-09-05: the shared maturity was under 91 days from the reporting date. Measured on an F-IRB institution obligor under CRR, 1m loan and 1m deposit: matched at 6 months and 5 years gives LGD 0.00 and RWA 0; matched at 91, 89, 30 and 7 days, and one day before the reporting date, gives LGD 0.45 and RWA 776,751. Basel 3.1 identical. Direct collateral with a matched short residual fails the same way (a 0.1-year cash deposit against a one-month loan). - Rule: CRR Art. 237(1) — "a maturity mismatch occurs when the residual maturity of the credit protection is less than that of the protected exposure. Where protection has a residual maturity of less than three months and the maturity of the protection is less than the maturity of the underlying exposure that protection does not qualify" (crr.pdf p.232); Art. 238(1) — "Subject to a maximum of five years, the effective maturity of the underlying shall be the longest possible remaining time" — a cap, no floor. PS1/26 Art. 219(3) routes netting through Art. 237-239 unchanged and Art. 219(2) prescribes the treatment for the currency case, so the regime carry-forward is word-for-word on the point that matters.
- Origin:
engine/crm/haircuts.py::apply_maturity_mismatchderived the exposure residual as.clip(lower_bound=0.25, upper_bound=5.0)and then askedcoll_maturity < _exposure_maturity_years. The collateral residual was not floored, so any matched pair under 0.25 years (91 days on the /365.25 basis) compared as protection-shorter-than-exposure, entered the gate chain, and hit the< 0.25three-month zero. A past-dated pair has a negative residual on both sides and fails the same comparison. The floor was there to keep the scaling denominatorT − 0.25positive, but the scaling branch is reached only witht >= 0.25andt < T, where the denominator is positive on its own. The guarantee twin (engine/crm/guarantees.py::_apply_maturity_mismatch_to_guarantees, P1.232) had already been rewritten to compare raw residuals and carries a comment recording that the collateral twin still floored — the divergence was known and never closed. - Escape class:
path-never-exercised. The estate's netting fixtures (P1.238, P1.241,reporting_netting_portfolio,r1_negative_gross) all carry deposits maturing months or years out; the P1.241 "matched" control is a six-year deposit against a five-year loan. No fixture anywhere had a matched pair inside three months, which is the ordinary tenor of interbank money. The one unit test that mentioned the floor,test_exposure_maturity_floored_at_0_25, covered only collateral LONGER than a one-month exposure (factor 1.0 either way), so it documented the clip without ever exercising the matched limb. - Why nothing else would have caught it: the treatment was silent —
ERROR_MATURITY_MISMATCH(CRM002) was declared incontracts/errors.pyand produced nowhere, so a zeroed deposit left no record; the supervisory validation register cannot see LGD* on an F-IRB row that never reaches a registered portfolio; and thelgdexport column carries the unsecured supervisory LGD on every F-IRB row, so the number the user read (0.45) was the same whether the collateral had been applied or not — onlylgd_flooredand the RWA told the two cases apart. - Gate change: (1)
tests/unit/crm/test_maturity_mismatch.py::TestMatchedAndShortDatedProtection— matched pairs at 7/30/60/89/91 days, a past-dated pair, a matched pair with a short original term, a genuine sub-three-month mismatch (still zeroed), a just-over-three-months mismatch (scales, denominator positive without a floor), and the Basel 3.1 twin; the pinning test retitled totest_collateral_outliving_a_one_month_exposure_keeps_full_value. (2)tests/acceptance/{crr,basel31}/test_art237_matched_short_dated_netting_firb.pyon the new in-memorytests/fixtures/matched_short_netting/builder — the user's shape exactly: F-IRB institution, same currency, same maturity, negative balance deposit — assertinglgd_floored == 0and RWA 0 at 7/60/89 days and past-dated, a six-month control, and a real 30-day-vs-2-year mismatch that must stay zeroed AND now raise CRM002. (3) P1.241 gainsmatched_shortandmatched_pastscenarios in both regimes with the hand-calc corrected (no floor on T). (4)engine/crm/haircuts.py::_record_maturity_mismatch_adjustmentsproduces CRM002 as one rolled-up count per run (zeroed and scaled rows), pinned bytests/unit/crm/test_crm002_maturity_zeroed_warning.py. Deferred to P1.367: a registeredRUNSinterbank netting portfolio so Tier 2 and the template-cell census see the treatment. - Verified red:
uv run pytest tests/unit/crm/test_maturity_mismatch.py -n 0 -qon the unchanged engine —8 failed, 13 passed, every new matched caseassert 0.0 == 1.0 ± 1.0e-06onmaturity_adjustment_factor; anduv run pytest tests/acceptance/crr/test_art237_matched_short_dated_netting_firb.py -n 0 -q—5 failed, 2 passed:matched_7d,matched_60d,matched_89dandmatched_pasteachassert 0.0 == 1000000.0ontotal_collateral_for_lgd, and the mismatch controlassert Falseon CRM002 presence. The six-month control and the no-warning test passed on the same run, which localises the defect to the sub-three-month window. - Lesson: the rule's own words are the assertion. "Less than" is a strict comparison on the residuals as they are; a floor applied to one side of it for an arithmetic convenience elsewhere changed the rule's meaning for an entire tenor bucket, and the twin that had been corrected said so in a comment nobody acted on. LESSONS C10's question — "which call site did I measure this on, and is there a second one?" — applies to a comment that names its sibling's divergence: a known divergence between twins is a plan bullet, not a comment. Second, the silence: a treatment that can remove 100% of a protection's value must say so; declaring an error code and producing it nowhere is the shape LESSONS B9 warns about, and the 2026-08-12 entry names two more codes in the same state.
2026-09-05 — A BLOCKER taint finding was recorded as accepted, was not accepted, and sat open on master for three weeks¶
- Defect:
pythonsecurity:S2083(BLOCKER) onscripts/generate_regulatory_tables.py:703— the single open code-scanning alert onmaster(GitHub alert 35), raised 2026-08-16 and still open on 2026-09-05. The 2026-08-17 entry above closed this finding by recording it as "resolved as Accepted in the SonarCloud platform (issueAaAMt-IEKije7nS9AwhB)". It was not.masterreports the issue under a different key entirely, and it has never been resolved:
AZ__l2AI-nu7-4qGDJNX BLOCKER OPEN scripts/generate_regulatory_tables.py:703
AZ8Ke_9KFFhzEQA0N41T BLOCKER ACCEPTED src/rwa_calc/ui/app/recon_signoff.py:244
The second row is the control: an accept in this project does land, and does
show as ACCEPTED/WONTFIX. The key named in the record is on neither list,
so the accept was applied to something else — most likely a branch-scoped
duplicate on the PR — while the issue on master was never touched.
- Rule: Not a regulatory escape. No RWA number is affected; the script is
developer/CI codegen whose only external input is a boolean --check flag.
- Origin: the closure, not the code. The generator is byte-for-byte the
version the 2026-08-17 entry analysed, and that analysis was correct — the
flow really does run from _splice's content read to the write's data
argument, re-fetched here from today's SARIF and unchanged. What failed was
the remedy: a platform action with no artefact in this repository was
recorded as done, in a commit message, a properties note and a test
docstring, and nothing anywhere could disagree with it.
- Escape class: caught-and-parked. The gate fired and kept firing; the
record of the finding became its resting place. It is the class whose fix
targets the register of tolerated findings rather than the output, and the
register here was three pieces of prose asserting an accept that no one had
checked. no-gate-exists is the near miss and is wrong: SonarCloud caught
this on the day it was introduced and has reported it on every analysis of
master since.
- Why every gate missed it: nothing in the repository can observe a platform
accept, so "accepted" was unfalsifiable prose. Local gates cannot substitute —
ruff and arch_check have no taint model, and the freshness tests measure
output bytes and were correctly green throughout. CI is also no help by
design: SonarCloud's finding surfaces as a GitHub code-scanning alert, not
as a failing job, so master CI has been green for three weeks with a BLOCKER
alert open on it. The alert was found by a human reading the security tab.
- Gate change: the finding is fixed structurally rather than accepted, and
the property is gated executably.
The 2026-08-17 entry concluded "there was never a structural fix to find —
the flow is the feature", and stopped one step short of its own lesson. The
flow is indeed the feature: _splice must read each target to preserve the
hand-written prose outside the GENERATED markers, and the script must write
the result back, so the content is file-derived and cannot stop being so. But
that entry had already named the mechanism — "a sink can be reported for a
tainted argument rather than a tainted path" — and the remedy follows
from it: the content does not have to be an argument to a call that also
takes a path. main() now opens the target and writes to the stream, so
the path-taking call carries only the constant-derived path and a literal
mode:
tests/contracts/test_docs_freshness.py::test_generator_keeps_spliced_content_out_of_a_path_taking_call
parses the generator and fails on any write_text / write_bytes call, so
the shape cannot come back. The S6549 / S2083 note in
sonar-project.properties records the remedy, drops the false claim, and now
carries the one-line command that makes any future accept falsifiable:
An accept remains legitimate where the path genuinely is the trust root
(recon_signoff.py above). What is no longer legitimate is recording one
without verifying it landed.
- Verified red: the new gate run against the pre-fix generator restored
from master (45f5b315) fails on the flagged line itself —
AssertionError: scripts/generate_regulatory_tables.py hands its rendered
content to a path-taking write call:
line 703: .write_text(...)
— and passes on the fix, with tests/contracts/test_docs_freshness.py
6 passed. The closure claim is independently red too: the anonymous
SonarCloud query above (no credentials needed) returns OPEN for the
generator's issue against a record that said Accepted, and
gh api .../code-scanning/alerts/35 --jq .state returns open with
dismissed_at: null.
- Lesson: a remedy that leaves no artefact in the repository is not closed
until something outside the repository has been asked. The closing rule
already says a defect is closed by its escape-log entry rather than its fix
commit — this entry is the case where the escape-log entry itself was the
resting place, because its named gate change was an action in another system
that nobody verified. Where a fix is a platform action, the entry must carry
the query that proves it, and the query must be run. Second, and narrower:
when an analysis correctly identifies a mechanism but concludes no fix
exists, re-read the mechanism for the remedy it implies. "The sink is
reported for a tainted argument" and "there is no structural fix" cannot both
be true — the first names exactly what to remove.
2026-09-08 — C 08.06's placement block was reported as a population block, blaming the one mapping that was correct¶
- Defect: a firm mapping a mixed legacy extract for return reconciliation — SA, FIRB and AIRB rows alongside its specialised-lending book — was told:
your mapping cannot produce this template — its approach labels
(advanced_irb, foundation_irb, standardised) fall outside the population it reports
Their extract did carry slotting, correctly mapped;
present_approaches contained slotting while that sentence was on screen.
The real cause was elsewhere entirely: the SL columns carried a placeholder
("N/A") on the non-slotting rows, which put sl_type and
slotting_category outside the engine vocabulary and made
_vocabulary_permits refuse C 08.06 on _SLOTTING_PLACEMENT_MAPPINGS. The
message named [components.approach], the one table already right, and said
nothing about the two [carriers.*] tables that were wrong.
And the block itself was wrong, which the first pass at this entry missed.
Told the placeholders were the cause, the firm cleared them and was refused
again. _label_facts measures the placement columns' vocabulary over the
WHOLE ledger, so a value on a row C 08.06 can never read — its population is
slotting-only — was refusing the template. The rows carrying the placeholder
were exactly the rows the template excludes. Naming the cause correctly is
worth nothing if the cause should not have been a cause.
- Rule: not a regulatory escape — no RWA number is affected. C 08.06 / OF
08.06 (Reg (EU) 2021/451 Annex II; PS1/26 Annex I/II) simply went unproduced
on the compare surface, with a reason that sent the analyst the wrong way.
- Origin: LedgerCoverage.blocking_labels is a set subtraction —
present_approaches - TEMPLATE_POPULATION_LABELS[template_id] — so on a mixed
extract it is non-empty whatever made the template unreachable. Its
docstring anticipated exactly this and guarded one case ("or when it is
blocked by a missing COLUMN instead — a different fix"), but the placement
block is a third cause and was added to _vocabulary_permits without a
matching accessor. ui/views/return_recon.py::_template_block then walked
columns → labels, found columns empty, and printed the subtraction.
_warn_unreachable had it right on the same coverage record and in the same
run — the log said "invalid or null slotting placement values in sl_type,
slotting_category" — so the two surfaces disagreed about the cause, and only
the one nobody reads was correct.
- Escape class: no-assertion-of-presence, in its diagnostic form. Every
gate asserted the block was present and none asserted the reason named the
cause. tests/unit/analysis/test_legacy_ledger.py covered all four placement
blocks and asserted blocking_columns("c08_06") == () — pinning that this is
not a column problem, one step from the defect, without ever asking what the
surface says instead. no-gate-exists is the near miss and is wrong: the
gates existed and ran green; they were pointed at reachability rather than at
the sentence.
- Why every gate missed it: the fixture had the right shape all along —
_LEGACY_ROWS's Approach column is SA, SA, AIRB, AIRB, FIRB, SA, SA,
SLOT, SLOT, SLOT, precisely the mix that makes the subtraction misfire — so
this was never path-never-exercised. Both existing placement tests ran over
it and passed. The UI test file tests _template_block's siblings
(test_an_unreachable_template_is_still_offered_with_the_blocking_columns)
but only that a blocked_reason is non-empty, which the wrong sentence
satisfies. Nothing anywhere compared the two surfaces' accounts of one
coverage record, which is what would have caught it: _warn_unreachable and
_template_block are the same decision written twice.
- Gate change: the cause becomes a named thing the analysis layer owns, so a
surface cannot invent its own answer.
LedgerCoverage.blocking_placement() (src/rwa_calc/analysis/legacy_ledger.py)
returns the placement mappings the refusal keyed on; blocking_labels()
returns () when placement is the cause, so it can no longer misattribute
through either of its two callers — nor through a third written later —
rather than only through the one surface that was reported;
and _warn_unreachable now reads the accessor instead of recomputing the
intersection inline, so the log and the UI cannot drift apart again. The UI
gains the matching branch, ordered columns → placement → labels.
Regression gates: test_a_placement_block_does_not_blame_the_approach_labels
(all three placement carriers; asserts slotting is in
present_approaches while the block stands) and
test_a_slotting_placement_block_names_the_carriers_not_the_approaches.
Second change, for the block itself: the refusal is now scoped to the
slotting book, matching the scoping the null check already had. The placement
scan counts rows of the slotting population whose discriminator is null or
out of vocabulary, and LedgerCoverage.invalid_placements carries that set;
_vocabulary_permits keys on it instead of on the whole-ledger
unmapped_labels. The whole-ledger measurement is still REPORTED — C 07.00's
rows 0021-0023 read sl_type off SA rows, so an unmapped value there is worth
naming — so nothing diagnostic is lost and only the refusal narrows. An
unmapped blank now renders <blank> instead of a hole in the remedy line,
because a blank is precisely what an analyst produces when told to clear a
placeholder. Gate:
test_a_placeholder_on_the_non_slotting_rows_does_not_block_c08_06, over five
placeholder shapes including the empty and whitespace strings.
Note what did NOT change: every pre-existing block test targets a slotting row
(index 7 or 9 of _LEGACY_ROWS) and all still fire. That is the evidence the
scoping is narrow rather than merely permissive — had the fix been "stop
blocking", those tests would have gone green-by-deletion.
- Verified red: both gates were run against the unfixed tree first. The UI
gate reproduces the reported sentence verbatim —
AssertionError: assert 'sl_type' in 'your mapping cannot produce this template —
its approach labels (advanced_irb, foundation_irb, standardised) fall outside
the population it reports'
— and the analysis gate fails AttributeError: 'LedgerCoverage' object has no
attribute 'blocking_placement' on all three parameters. The over-refusal gate
was run red too, on all seven of its parameters:
After both fixes, tests/unit/analysis/, tests/unit/ui/, tests/unit/api/
are 1037 passed and tests/contracts tests/integration 1510 passed.
- Lesson: a derived explanation must be refused when it is not the reason,
not merely computed when it is. blocking_labels answers "which approaches
are not in the population" — a question with an answer on almost every mixed
extract — and the caller treated a non-empty answer as evidence it was the
cause. Where several causes can block one outcome, each needs its own
accessor, and the accessors need to be mutually exclusive at the source
rather than by call order in each consumer; otherwise every new consumer
re-derives the precedence and one of them gets it wrong. The tell was
available for free: two surfaces printed different causes for the same
coverage record in the same run.
Second, and the one that cost the user a second round trip: a comment
admitting a check over-reports is a defect report, not a caveat.
_label_facts carried, in capitals, the sentence "it can over-report on an
extract that reuses one source column across approaches" — an accurate
description of this exact escape, written before it happened, directly above
the code that caused it. It read as a documented trade-off ("conservative in
the right direction") because the cost was invisible: an over-report on a
reachability check is not a noisy warning, it is a template silently
withheld. Conservatism is only free where the failure mode is a false alarm;
where it is a refusal, "conservative" and "wrong" are the same thing, and the
scoping that made it correct was three lines and available all along. When a
comment says a check may fire on rows the consumer cannot read, ask what
firing costs before accepting it.
2026-09-08 — The C 07.00 exposure-class axis was our own vocabulary, and the register's scope resolution summed two of our sheets into the class total it was checking¶
- Defect: a user reading the COREP output asked why
retail_mortgageswas appearing as an exposure class on C 07.00, since it is not one. They were right. C 07.00 and OF 07.00 keyed their sheet (z) axis on this repo's internalExposureClassvalues instead of on the Art. 112(1) classes the templates index.reporting/kernel/bases.py::sheet_axisbuilds the axis asdata[class_col].drop_nulls().unique(), so whatever string the sealed class carrier happens to hold becomes a sheet, with no membership test anywhere between the classifier and the submission.
So corporate_sme opened a sheet beside corporate, and retail_mortgage /
residential_mortgage / commercial_mortgage opened three sheets for the
single class (i). The Art. 112(1)(g) and (i) class totals were therefore
reported nowhere, and row 0020 "of which: SME" was null on the very sheet
that should have carried it. Measured on the b31 estate before the fix: sheet
corporate r0010 = 8,000,000 with r0020 NULL, beside a separate
corporate_sme sheet of 500,000; retail_mortgage 400,000 beside
commercial_mortgage 10,000,000 with no class-(i) total anywhere. Two further
non-conforming keys were reachable in production and hit by no fixture:
retail_qrre and residential_mortgage. Fixed in PR #497 (eb546935,
merged as f42cf9c4); after it corporate reports 8,500,000 with r0020 =
500,000, and real_estate 10,400,000 split across r0330 / r0340.
- Rule: CRR Art. 112(1)(a)-(q) — the class list the C 07.00 z-axis indexes.
COREP Annex II ¶47 and PS1/26 Annex II ¶47 both require the total and each
exposure class in a separate dimension, and ¶56 assigns them in the
Art. 112(2) Table A2 order (checked against
ps1-26-annex-ii-reporting-instructions.pdf p.77 by the review that filed
P5.65; not re-extracted for this entry). The published rule the axis broke is
EBA v4240_i — ERROR, live —
{C 02.00, r0130, c0010} == {C 07.00.a, r0010, c0220, s0008}.
- Origin: two places, and the second is why it survived.
reporting/corep/c07.py keyed the sheet off the sealed class directly, and
reporting/validations/scope.py then declared the same wrong axis to the
validator: _C07_SHEETS z0008 held ("corporate", "corporate_sme"), z0009
("retail_other", "retail_qrre") and z0010 the three mortgage classes, with
the same three shapes on the _OF07_SHEETS twin. Standing since C 07.00
acquired a sheet axis. Found by a user reading the output.
- Escape class: test-shared-the-assumption.
tests/unit/reporting/corep/test_c07.py covered the sheet axis and passed
while indexing three keys — secured_by_re_residential,
secured_by_re_commercial, secured_by_re_property — that are not
ExposureClass members at all. Production and its tests were written from
one invented vocabulary, so both sides agreed and both were wrong. The class's
prescribed fix, re-anchor to a source of truth, is exactly what shipped:
tests/contracts/test_c07_art112_sheet_axis.py anchors every assertion on
domain.enums.ExposureClass, on the sibling template's C02_00_SA_CLASS_MAP,
and on the published z-axis in SHEET_INDEX_MAPS — three things that cannot
drift with C07_00_SA_SHEET_MAP.
The register limb carries no class, and forcing one would prescribe the
wrong fix. The supervisory register did not fail to catch this; it converted
it into a pass (below). That is a defect in a gate, which this file's own
preamble says to leave unclassed — and it is the second instance of the
shape the 2026-08-09 ratchet entry conditionally named gate-unfit, a gate
measuring something that is not the quantity that matters. That entry said one
instance is not a taxonomy. Two may be, and the naming is the operator's call;
the difference worth recording is that the first was caught by adversarial
review before its gate shipped, while this one shipped, ran green, and
masked a live defect for as long as C 07.00 has had a sheet axis.
- Why every gate missed it: the strongest gate in the estate was
structurally incapable of binding, and its green was manufactured by the
defect itself. reporting/validations/scope.py::resolve_sheet_codes returns
every bundle_key a published z-code names, and the evaluator then
aggregates the rule's reference across all of them
(reporting/validations/evaluate.py::_sum_cells, whose own docstring names
this exact case: "or from a publisher sheet code that maps onto more than one
of our sheets"). With z0008 holding two keys, the live ERROR rule v4240_i
compared C 02.00 row 0130 against the sum of two of our sheets — a figure
appearing in no submitted workbook. A faithful pre/post comparison of the
whole register is identical in both states: 27 broken / 4 uncovered / 188
vacuous either way, not one rule outcome moved. Corroborating that from the
other side, eb546935 does not touch
tests/expected_outputs/reporting/validation_known_breaks.json at all. The
multi-key mapping did not merely fail to catch the defect; it masked it.
Nothing else was positioned to see it either.
- **The goldens pinned it.** The pre-fix estate carried
`corep__c07_00__corporate_sme.ndjson` and
`corep__c07_00__retail_mortgage.ndjson` as golden files — one per wrong
sheet — so the golden gate was green *because* it had recorded the wrong
axis. Goldens ratchet against change, not against conformance, and a wrong
axis presents to them as a set of filenames.
- **The coverage ratchets are blind to it by construction.** The cells were
populated and the rules bound; an axis error redistributes figures across
sheets without nulling a cell or un-binding a rule, so no liveness or
binding-rule metric moves (first entry in this file).
- **The two unfixtured keys fail open.** A sheet we emit that no z-code
addresses resolves to `sheet_not_emitted`, and every rule scoped to it is
skipped rather than failed — so `residential_mortgage` and `retail_qrre`
would have cost coverage silently rather than reddening anything.
- Gate change: the invariant is graduated executably in two halves that measure different things, and the wrong axis is additionally caught in positive form.
scripts/arch_check.py check 22 — check_sheet_code_single_bundle_key,
registered in main() so it runs on every commit through the pre-commit hook.
An AST scan of reporting/validations/**/*.py: every SheetCode(...) — and
any call passing bundle_keys= at all, which is what catches
replace(entry, bundle_keys=...) — may carry at most one bundle key. It
is unconditional and has no allowlist, which the estate can afford because
no map violates it after #497; measured population 67 SheetCode entries
across all four maps, zero violations. It walks whole modules, so a map
built inside a builder function or a class body is still seen, and it fails
loudly rather than silently if it finds zero SheetCode literals or if
reporting/validations/ moves out from under it — an unmeasured invariant
reads exactly like a satisfied one.
Two cardinalities stay legal, deliberately. An empty tuple is a skip, not
a zero (z0012 and s0015 use it; filed as P2.55 and P2.54). And the reverse
shape — several z-codes onto one of our sheets — is the DPM's Art. 147(2)(d)
IRB axis being genuinely finer than ours (z0013 SME and z0014 non-SME both
address our single retail_mortgage sheet), made safe by the existing
sheet_scope_not_closed refusal. A "simplification" to a 1:1 axis would break
the IRB templates. _C08_SHEETS / _OF08_SHEETS were clean before #497 and
after it, which is the evidence the invariant is aimed at the one direction.
tests/contracts/test_sheet_index_bundle_key_cardinality.py carries the
same invariant over the map the evaluator actually resolves, rather than over
source literals, and pins the reverse shape and its closure guard beside it so
neither can be removed without the other being read. Its last limb asserts the
arch-gate half exists and is dispatched by main(), locating the check by
what it reads rather than by its name — a guard that is built and never wired
is this estate's dominant meta-pattern, and it is why arch_check grew check
20.
tests/contracts/test_c07_art112_sheet_axis.py is the positive form: the
C 07.00 sheet axis must be in 1:1 correspondence with the C 02.00 SA class
rows, since COREP Annex II §1.3.1 makes C 02.00 rows 0070-0211 identities
against the C 07.00 sheet for the same Art. 112(1) letter. So the class is now
caught two ways — at pytest time by the axis correspondence, at arch-gate time
by the cardinality invariant — and the two are anchored on different sources.
The axis review that produced this entry also filed P5.65 (this
graduation) plus P1.371, P1.372, P2.53, P2.54, P2.55 and
P5.66: further Art. 112(1) axis gaps that the fix exposes rather than
causes, none of them closed here.
- Verified red: the shipped check run against the pre-fix source produces
6 violations — z0008, z0009 and z0010 in each of _C07_SHEETS and
_OF07_SHEETS — and 0 on the current tree:
scope.py:143: _C07_SHEETS sheet code '0008' maps to 2 bundle keys
(corporate, corporate_sme). resolve_sheet_codes returns EVERY key for a scoped
code, so the rule evaluator aggregates the publisher's reference across all of
them and asserts against a figure that appears in no submitted workbook ...
The sample is tests/contracts/data/scope_pre_art112_class_axis_fix.py.txt,
verbatim eb546935^:src/rwa_calc/reporting/validations/scope.py, so the red is
reproducible from the repository rather than only from history; it was also
reproduced independently straight from git show while this entry was
written. The measured population is 67 in both states, which matters — the
red is a change in cardinality, not in how much the check can see.
The contract file's limb 3 runs the identical predicate over those pre-fix
entries transcribed verbatim, carrying z0007 unchanged as a single-key
control, so a predicate that flagged every entry or none of them fails there.
And its companion holds the requested z-code and the emitted sheet set fixed
while varying only the map, showing the pre-fix entry handing the evaluator
both corporate and corporate_sme where the live one hands it one — the
masking measured rather than argued.
- Lesson: a gate that normalises its input before asserting on it will
report the normalisation as a pass. Summing a multi-cell reference is a
reasonable reading — _sum_cells documents it as "the only additive reading"
— and it quietly reassembled the class total the workbook had lost, so the
register asserted against a correct figure the submission did not contain.
Where a validator canonicalises its input, the canonicalisation is part of the
assertion and needs a gate of its own; here that gate is a cardinality
invariant on the mapping, not a check on any number.
Narrower, and the reason this was a user's finding rather than a test's: a
template's axis is a published vocabulary, not a free string. unique()
over a sealed column produces an axis-shaped object that is not an axis, and
once one exists the tests written against it inherit its vocabulary — which is
how secured_by_re_property, a string belonging to no enum in the repository,
came to sit in a passing test asserting on a regulatory return.