External cold review of the CEBE calibration program — steelman and surviving objections
Role: commissioned external cold reviewer (Claude, Anthropic). Date of review: 2026-09-02, adjudicating the seven-document response package against a cold review of cebetracker.io conducted 2026-09-01. Independence: the review was conducted cold. The site and the seven files were read without briefing, and no context outside the published pages, the public feed, the sheet tabs, and the seven documents was used.
Reproduced as delivered, in original order and wording. No part of the review was revised after pushback; nothing below is a final form of an earlier draft.
Part I — Calibration findings from the cold review of 2026-09-01
The seat-3 and calibration-program sections of the original review, reproduced because Part II adjudicates them by reference.
Seat 3 — Hostile quant
Objection 1: The launched figure is not the validated one, and the gap runs one way. /claims/adjusted/weight/ §7: "It is not the quantity Calibration 3 scored… At this mark the divergence is 16.627 points of face… The blend carries the convert stack lighter than the model." $4.807B published vs $5.923B validated. The site's own §1 says section 8 held the weight until calibration showed marks "price within a tolerance" — the object that cleared is not the object served.
Objection 2: The scored volatility input was ruled ineligible by the program's own pre-registration. /claims/calibration-2/ R-VOL §6.6: "A value captured before the rules governing its capture were written cannot be retrofitted into a pre-registered result"; Revision 2: "The twelve implied-volatility 30-day values remain unscored groundwork." /claims/calibration-3/ §2: "Volatility is the sourced 30-day implied volatility at each mark," evidence file R-VOL_table_MC-only_2026-08-13.tsv — the same values. The tenor scaling uses "the at-the-money term-structure ratio measured on the 2026-08-14 option chain" applied to 2024 marks: look-ahead. R-VOL §5 says interpolation "is the only arithmetic applied to a scored value"; the ratio is a second one.
Objection 3: The mixed-vintage headline. /: "Stock prices live; CEBE from latest company filing." Five of twelve balance sheets are 2026-06-30 at BTC ≈ $58.5k against $77.2k live; NAKA reads 55.78% or 42.30% claims depending on basis. /methodology/: "the same world" on both sides. The methodology's own price rule ("the following day's print at 00:00 UTC") is overridden in five source lines by "convention #A-45," which is defined nowhere on the site.
Page that loses trust fastest: /claims/calibration-3/. §4: "This structure was chosen with every one of those numbers already in hand, including the failure." §5 on the degeneracy floors: "constructed during the exploratory phase… not an independent prior test." §3, sealed slice: 5.077 against 5.0, then Revision 1 sets the retirement bound at −5.0 "safely wide of the known starting state… negative 3.516."
Cannot recompute: All of Calibration 3 — "docs/specs/cal-3-registration-2026-08-19.md," "study/cal3/holdout/," blob 0b5e95f2 — the repository is not public (GitHub search returns nothing). Calibration 2's result — /claims/spec/ says "Calibration 2 also failed"; /claims/calibration-2/ says "No result exists"; no page carries a number, and /claims/calibration-3/ describes a "ruled coupling… in both N slots" that the Cal 2 pre-registration never specifies. Calibration 1's build_calibration1.py reads /mnt/user-data/uploads/per_mark_exhibit_v2.csv and the cited TSV returns 404. The /framework/ calculator computes CEBE on "FD Shares Outstanding 315.0M," violating Principle three on the page that teaches the formula.
The calibration program
Fitted:
- SPREAD_ADD 8.803 is a free credit spread wearing a "widening" label;
/claims/calibration-3/§2: "the specification is exactly insensitive to the B band's level." Three parameters, one issuer, 41 marks, 9 dates. - The gating protocol. LODO refits per fold; the "frozen" parameters are not what the gating 4.747 measures. The only frozen-parameter, unseen-data cell (5.077) fails.
- The degeneracy floors: σ√T ceiling 2.00 against the candidate's 1.882; N(d2) floor 0.05 against 0.210 — set with the candidate in hand (§5).
- The retirement thresholds (Revision 1), set after the confirmation slice was visible.
Cherry-picked:
- Worst-of-four admission dropped after the leave-two-dates-out fail at 14.94 (§4, §7).
- The traded repurchase print, the third clause of a tolerance a candidate "must clear all three" of (
/claims/calibration-1/§1), demoted to non-gating in Cal 3 and then missed by 13.01. - Put leg over maturity leg, chosen at 17-of-42 vs 2-of-42 coverage; the page's own words: "it was made with that difference in view" (
/claims/calibration-2/Rev 2 §2). - MSTR alone, then a scope rule (R-23) that makes the population "the tested one" by definition.
- "Full panel" quietly re-denominated from 9 held-out dates (Cal 1) to 10 dates of which 9 are the fitting set (Cal 3), while asserting "the three numbers do not move."
Demand pre-registered, and isn't:
- Third-party timestamping of every tolerance and registration.
/claims/calibration-1/ruling-log/: "This lane cannot independently verify the timestamp… The ruling authority has stated that the governing thread's transcript carries the earlier timestamp." The 14.07 floor "was disclosed to the ruling authority before the tolerance was transmitted." Ruling authority, model author, and launch ratifier are one person on every page. - The admission protocol itself. Cal 3 registered the construction but chose which protocol gates after seeing all six.
- The pass/fail of the sealed slice as its own cell, before opening it. It was "not a registered test" precisely because it would have failed.
- The volatility source and tenor-scaling rule for the scored run, since R-VOL as registered forbids what Cal 3 did.
- A calibration against observed conversions, which is what the priors were published for (
/methodology/: "calibrated against observed conversions"). No study tests that; Cal 1 §3.6 concedes the schedule is "a different object" from what was scored, and the four-band table still stands on/claims-grade/as the "standing stated prior" while feeding no served figure. - Publication of Calibration 2's construction and numbers before Calibration 3 was allowed to cite its failure.
Part II — Steelman review of the response package, 2026-09-02
Files reviewed: 00-README-v2.md, 01-cal3-limitation-note.md, 02-cal3-retirement-and-scorecard.md, 03-cal-num-4-registration.md, 04-cal-num-5-6-ladder.md, 05-cal-den-2-skeleton.md, 06-program-governance.md. Registrations not yet sealed at time of review.
Steelman, per document
00 README. It converts the red team into a coverage table, so a reader can check that every finding has a named answer rather than a mood. Routing the package to a cold reviewer before sealing makes adversarial review a precondition of registration instead of a post-mortem.
01 Cal-3 limitation note. It puts the two sourcing defects into the record in the reviewer's own words, beside the result, without editing the result, which is the only honest way to amend a pre-registration. It also correctly refuses to let a study with no result (Cal-4) retroactively re-score Cal-3.
02 Retirement rule. Moving "this bound was set after the slice was visible" into the registered text, with the arithmetic (−3.5 known vs −5.0 bound), lets a stranger verify the rule is not born tripped. Binding quiet quarters and shipping per-instrument deltas as a file makes the scorecard a record rather than a press release.
03 Cal-4. Forward-only scoring is the right consequence of R-VOL 6.6, and inheriting the incumbent's floors unchanged is the correct fix for floors-set-with-candidate-in-hand. The engagement rule (§5) pre-commits the "indistinguishable from null" outcome, so an unengaged cap cannot be reported as a pass.
04 Ladder. Sealing the generalization test under the same hash as the champion selection commits the definition of success before a champion exists. "Ineligible rather than improvised" is the right default for issuers with no vol source.
05 Cal-den-2. It names the program's central untested claim and registers the test alongside, not after, Cal-4. Lifetime horizon on resolved instruments only is the correct lesson from D1.
06 Governance. G-4 draws the fence the standard needed: four core figures, one definition per label, research quarantined so the core inherits none of its fragility. G-1 ends self-attested timestamps for everything forward.
What survives
One person. Every threshold, disposition and ruling is still "Bobby's to set" (04 last line), "Bobby's decision, flagged" (06 G-2, G-3), "Ratification: Bobby" (03 §7). G-1 externalizes timestamps, not judgment. An OpenTimestamps hash proves when a file existed, not that anyone independent agreed with it.
The launched figure. 06 G-3 offers two dispositions: carry the divergence on the face, or move to research. The weight page already carries both numbers on its face, and G-4 already moves the weight to research. G-3 is therefore satisfied by the status quo and changes nothing served. The 16.6-point gap stays live.
The ineligible inputs are still serving. 01 concludes "no served figure moves," but the six published weights were computed on the 2026-08-07 IV30 and the 2026-08-14 chain ratio the same note declares inadmissible. The note downgrades the admission's status and leaves the figures it produced untouched, for up to four quarters (03 §2).
Cal-3's protocol selection. 01 covers sourcing and tenor only. "This structure was chosen with every one of those numbers already in hand," the demoted repurchase print, the demoted leave-two-dates-out, and a gating statistic that is nine-tenths in-sample are not mentioned anywhere in the seven files. The admission's worst problem was never sourcing.
Reproducibility. 06 G-2 binds "studies concluding after this date." Cal-1 through Cal-3, the only studies with results, stay unverifiable.
Priors vs conversions. 05 tests delivery rates by issuance class. The published four-band moneyness schedule on /claims-grade/ is still the "standing stated prior" and still isn't the object any study scores.
New weaknesses the old program did not have
Cal-4 has no absolute tolerance. 03 §4 is purely relative: a challenger needs RMSE 1.0 better than the incumbent, ties to the incumbent. Cal-1 through Cal-3 all carried 5/12/6. A champion can now be crowned at RMSE 9 and be described as the program's clean result.
Optional stopping. 03 §2: "concludes at the earlier of four scored quarters or the point where the selection rule's margin is decidable." "Decidable" is undefined, so the stopping time is chosen after seeing partial results. That is the textbook way to manufacture a margin.
The panel is undefined. 03 §2 scores "weekly capture-protocol marks," but the deviation statistic (§4, "as in Cal-num-3") needs an observed fair value, which only exists in the quarterly filing. Weekly captures are inputs, not marks. The scored panel is therefore six marks per quarter from one issuer, and a 1.0-point margin on twelve or twenty-four marks is noise.
The incumbent carries the contamination forward. 03 §3 candidate 1 inherits Cal-3's convention "unchanged." If A, B and SPREAD_ADD travel with it, and nothing in 03 says they don't, every candidate is scored on parameters fitted to the inputs 01 rules ineligible. The governor changes; the fit underneath is the retrofit.
The issuer's ruler enters the model. 03 §3 candidate 6 and 04 amendment 2 both take the issuer's own fair-value-note volatility as an input, then score the model against the issuer's fair value. That candidate can only agree with the ruler it was handed.
The scorecard keeps the condemned arithmetic. 02 scores "the same construction as the 2026-08-22 confirmation slice," which is the 2026-08-14 tenor ratio 01 calls a forbidden second arithmetic. One file bans it; the next file schedules it quarterly.
Cal-2's page was false at registration. 06's state sentence reveals a candidate "withdrawn as unspecified at registration" on a page that said "LOCKED. No open items," and re-homes a running cohort under a different study's rules. That is data collected under one registration and scored under another.
Governance by instances. 03 §7: registration by copilot, scoring by "a fresh study desk booted on the sealed file," review by "a cold strong model" (00). Every non-human role is a model; every human role is one person. The independence problem did not shrink, it changed costume. And the README names "Dylan" while 03 §3 says attribution is withheld pending permission; the leak is in the governance file.
The sentence I'd bet on
03, lineage line: "Cal-num-4 is a champion selection and the program's first fully clean-sourced result by construction."
By March 2027 there will be two filed quarters, roughly twelve marks. Under §4 that is not decidable at a 1.0 margin, so either the study stops early on an undefined rule, or the incumbent, whose parameters were fitted on the ineligible inputs, is confirmed as champion and the word "clean" attaches to a retrofit, or a cap candidate wins on noise and inherits the ladder. All three outcomes falsify the sentence as written, and 01 has already staked the disposition of the Cal-3 admission on it.
Runner-up: 04 amendment 2, "where absent, the issuer is ineligible." Of the eleven non-MSTR converts on the roster, count how many issuers publish a fair-value note with a volatility assumption. The tiers will be empty and the ladder will conclude with one rung.
End of external review. The production seat's dispositions publish separately.