source verification/studies/calibration-v3/steelman-2026-09-02.md
sha256 835a8169dfa5ca9c12e5dba928c882bdaa3a566f22b5a38a77b760efe221d0fc
bytes 14,081
RESEARCH · NOT THE CORE STANDARD

Research extension. This page is a study, not a measurement. The core standard is CEBE, Senior Claims %, CEBE mNAV and Claims Grade, each published with its formula and date. Studies here test priors against outcomes and can be retired; retired studies stay on the site with their correction notes. Nothing on this page enters a tracker figure.

External cold review of the CEBE calibration program — steelman and surviving objections

Role: commissioned external cold reviewer (Claude, Anthropic). Date of review: 2026-09-02, adjudicating the seven-document response package against a cold review of cebetracker.io conducted 2026-09-01. Independence: the review was conducted cold. The site and the seven files were read without briefing, and no context outside the published pages, the public feed, the sheet tabs, and the seven documents was used.

Reproduced as delivered, in original order and wording. No part of the review was revised after pushback; nothing below is a final form of an earlier draft.


Part I — Calibration findings from the cold review of 2026-09-01

The seat-3 and calibration-program sections of the original review, reproduced because Part II adjudicates them by reference.

Seat 3 — Hostile quant

Objection 1: The launched figure is not the validated one, and the gap runs one way. /claims/adjusted/weight/ §7: "It is not the quantity Calibration 3 scored… At this mark the divergence is 16.627 points of face… The blend carries the convert stack lighter than the model." $4.807B published vs $5.923B validated. The site's own §1 says section 8 held the weight until calibration showed marks "price within a tolerance" — the object that cleared is not the object served.

Objection 2: The scored volatility input was ruled ineligible by the program's own pre-registration. /claims/calibration-2/ R-VOL §6.6: "A value captured before the rules governing its capture were written cannot be retrofitted into a pre-registered result"; Revision 2: "The twelve implied-volatility 30-day values remain unscored groundwork." /claims/calibration-3/ §2: "Volatility is the sourced 30-day implied volatility at each mark," evidence file R-VOL_table_MC-only_2026-08-13.tsv — the same values. The tenor scaling uses "the at-the-money term-structure ratio measured on the 2026-08-14 option chain" applied to 2024 marks: look-ahead. R-VOL §5 says interpolation "is the only arithmetic applied to a scored value"; the ratio is a second one.

Objection 3: The mixed-vintage headline. /: "Stock prices live; CEBE from latest company filing." Five of twelve balance sheets are 2026-06-30 at BTC ≈ $58.5k against $77.2k live; NAKA reads 55.78% or 42.30% claims depending on basis. /methodology/: "the same world" on both sides. The methodology's own price rule ("the following day's print at 00:00 UTC") is overridden in five source lines by "convention #A-45," which is defined nowhere on the site.

Page that loses trust fastest: /claims/calibration-3/. §4: "This structure was chosen with every one of those numbers already in hand, including the failure." §5 on the degeneracy floors: "constructed during the exploratory phase… not an independent prior test." §3, sealed slice: 5.077 against 5.0, then Revision 1 sets the retirement bound at −5.0 "safely wide of the known starting state… negative 3.516."

Cannot recompute: All of Calibration 3 — "docs/specs/cal-3-registration-2026-08-19.md," "study/cal3/holdout/," blob 0b5e95f2 — the repository is not public (GitHub search returns nothing). Calibration 2's result — /claims/spec/ says "Calibration 2 also failed"; /claims/calibration-2/ says "No result exists"; no page carries a number, and /claims/calibration-3/ describes a "ruled coupling… in both N slots" that the Cal 2 pre-registration never specifies. Calibration 1's build_calibration1.py reads /mnt/user-data/uploads/per_mark_exhibit_v2.csv and the cited TSV returns 404. The /framework/ calculator computes CEBE on "FD Shares Outstanding 315.0M," violating Principle three on the page that teaches the formula.

The calibration program

Fitted:

Cherry-picked:

Demand pre-registered, and isn't:


Part II — Steelman review of the response package, 2026-09-02

Files reviewed: 00-README-v2.md, 01-cal3-limitation-note.md, 02-cal3-retirement-and-scorecard.md, 03-cal-num-4-registration.md, 04-cal-num-5-6-ladder.md, 05-cal-den-2-skeleton.md, 06-program-governance.md. Registrations not yet sealed at time of review.

Steelman, per document

00 README. It converts the red team into a coverage table, so a reader can check that every finding has a named answer rather than a mood. Routing the package to a cold reviewer before sealing makes adversarial review a precondition of registration instead of a post-mortem.

01 Cal-3 limitation note. It puts the two sourcing defects into the record in the reviewer's own words, beside the result, without editing the result, which is the only honest way to amend a pre-registration. It also correctly refuses to let a study with no result (Cal-4) retroactively re-score Cal-3.

02 Retirement rule. Moving "this bound was set after the slice was visible" into the registered text, with the arithmetic (−3.5 known vs −5.0 bound), lets a stranger verify the rule is not born tripped. Binding quiet quarters and shipping per-instrument deltas as a file makes the scorecard a record rather than a press release.

03 Cal-4. Forward-only scoring is the right consequence of R-VOL 6.6, and inheriting the incumbent's floors unchanged is the correct fix for floors-set-with-candidate-in-hand. The engagement rule (§5) pre-commits the "indistinguishable from null" outcome, so an unengaged cap cannot be reported as a pass.

04 Ladder. Sealing the generalization test under the same hash as the champion selection commits the definition of success before a champion exists. "Ineligible rather than improvised" is the right default for issuers with no vol source.

05 Cal-den-2. It names the program's central untested claim and registers the test alongside, not after, Cal-4. Lifetime horizon on resolved instruments only is the correct lesson from D1.

06 Governance. G-4 draws the fence the standard needed: four core figures, one definition per label, research quarantined so the core inherits none of its fragility. G-1 ends self-attested timestamps for everything forward.

What survives

One person. Every threshold, disposition and ruling is still "Bobby's to set" (04 last line), "Bobby's decision, flagged" (06 G-2, G-3), "Ratification: Bobby" (03 §7). G-1 externalizes timestamps, not judgment. An OpenTimestamps hash proves when a file existed, not that anyone independent agreed with it.

The launched figure. 06 G-3 offers two dispositions: carry the divergence on the face, or move to research. The weight page already carries both numbers on its face, and G-4 already moves the weight to research. G-3 is therefore satisfied by the status quo and changes nothing served. The 16.6-point gap stays live.

The ineligible inputs are still serving. 01 concludes "no served figure moves," but the six published weights were computed on the 2026-08-07 IV30 and the 2026-08-14 chain ratio the same note declares inadmissible. The note downgrades the admission's status and leaves the figures it produced untouched, for up to four quarters (03 §2).

Cal-3's protocol selection. 01 covers sourcing and tenor only. "This structure was chosen with every one of those numbers already in hand," the demoted repurchase print, the demoted leave-two-dates-out, and a gating statistic that is nine-tenths in-sample are not mentioned anywhere in the seven files. The admission's worst problem was never sourcing.

Reproducibility. 06 G-2 binds "studies concluding after this date." Cal-1 through Cal-3, the only studies with results, stay unverifiable.

Priors vs conversions. 05 tests delivery rates by issuance class. The published four-band moneyness schedule on /claims-grade/ is still the "standing stated prior" and still isn't the object any study scores.

New weaknesses the old program did not have

Cal-4 has no absolute tolerance. 03 §4 is purely relative: a challenger needs RMSE 1.0 better than the incumbent, ties to the incumbent. Cal-1 through Cal-3 all carried 5/12/6. A champion can now be crowned at RMSE 9 and be described as the program's clean result.

Optional stopping. 03 §2: "concludes at the earlier of four scored quarters or the point where the selection rule's margin is decidable." "Decidable" is undefined, so the stopping time is chosen after seeing partial results. That is the textbook way to manufacture a margin.

The panel is undefined. 03 §2 scores "weekly capture-protocol marks," but the deviation statistic (§4, "as in Cal-num-3") needs an observed fair value, which only exists in the quarterly filing. Weekly captures are inputs, not marks. The scored panel is therefore six marks per quarter from one issuer, and a 1.0-point margin on twelve or twenty-four marks is noise.

The incumbent carries the contamination forward. 03 §3 candidate 1 inherits Cal-3's convention "unchanged." If A, B and SPREAD_ADD travel with it, and nothing in 03 says they don't, every candidate is scored on parameters fitted to the inputs 01 rules ineligible. The governor changes; the fit underneath is the retrofit.

The issuer's ruler enters the model. 03 §3 candidate 6 and 04 amendment 2 both take the issuer's own fair-value-note volatility as an input, then score the model against the issuer's fair value. That candidate can only agree with the ruler it was handed.

The scorecard keeps the condemned arithmetic. 02 scores "the same construction as the 2026-08-22 confirmation slice," which is the 2026-08-14 tenor ratio 01 calls a forbidden second arithmetic. One file bans it; the next file schedules it quarterly.

Cal-2's page was false at registration. 06's state sentence reveals a candidate "withdrawn as unspecified at registration" on a page that said "LOCKED. No open items," and re-homes a running cohort under a different study's rules. That is data collected under one registration and scored under another.

Governance by instances. 03 §7: registration by copilot, scoring by "a fresh study desk booted on the sealed file," review by "a cold strong model" (00). Every non-human role is a model; every human role is one person. The independence problem did not shrink, it changed costume. And the README names "Dylan" while 03 §3 says attribution is withheld pending permission; the leak is in the governance file.

The sentence I'd bet on

03, lineage line: "Cal-num-4 is a champion selection and the program's first fully clean-sourced result by construction."

By March 2027 there will be two filed quarters, roughly twelve marks. Under §4 that is not decidable at a 1.0 margin, so either the study stops early on an undefined rule, or the incumbent, whose parameters were fitted on the ineligible inputs, is confirmed as champion and the word "clean" attaches to a retrofit, or a cap candidate wins on noise and inherits the ladder. All three outcomes falsify the sentence as written, and 01 has already staked the disposition of the Cal-3 admission on it.

Runner-up: 04 amendment 2, "where absent, the issuer is ineligible." Of the eleven non-MSTR converts on the roster, count how many issuers publish a fair-value note with a volatility assumption. The tiers will be empty and the ladder will conclude with one rung.


End of external review. The production seat's dispositions publish separately.

More studies, and the construction they test, are listed on the research page. The core standard's construction is documented on the methodology page.