Closed evidence: The pre-registered Touch-and-Turn video rule was refuted.
No edge. The pre-registered Touch-and-Turn rule loses money on every one of the nineteen symbols it was run on, and it loses money before costs as well as after. Pooled over the nine pre-registered tier-A primaries: 2,654 trades on 495 distinct session dates, a 29.9% win rate (date-clustered 95%: 28.0%–31.9%), −15.40 bps of net expectancy per trade (date-clustered 95%: −17.90 to −12.84), a 0.553 profit factor. Pooled intervals are quoted date-clustered throughout; see "A note on pooled precision". The seal is b5a705ef7206a755cc32107ae3291120c8db2752cf8b0de6b0146cf644cee33d.
- Evidence state:
refuted, both tiers. Gates 5 (positive expectancy) and 6 (Holm) fail on every symbol. No cell in this run is promotable, eligible, or worth tuning. - Costs explain most of it (78%), not all of it — the gross residual is negative with an interval excluding zero. On the tier-A primaries the 12 bp round trip is 12.00 of the −15.40, i.e. 77.9%. The remaining −3.40 bps of gross expectancy carries its own bootstrap 95% of [−5.54, −1.26], so charging zero cost would still leave a losing rule. The gross figure is no longer the one headline number without an interval.
- The video's headline claims are refuted, not merely unconfirmed. ">70% win rate" — the pooled interval tops out at 31.7%. "Losses under 30%" — 70.0% of trades lost money (95%: 68.2%–71.7%).
- Tier C answers the owner's question about smaller names, and answers it the same way. VRT and APP, on a thin-tape corpus that admits a hole inside the opening 15: 665 trades, 28.4% win rate, −21.46 bps. Reported as an exploratory section, never pooled, never charged as a test.
- One configuration was run. No grid, no sweep, no symbol selection after the fact. The exploratory grid remains unrun.
What was run
- The pre-registered primary genome, unchanged; the evaluator's execution rules as amended — see "Deviations from pre-registration" below. 15-minute opening range, qualify at
range × 100 ≥ 25 × ATR14(i−1), fade the opening candle, both sides, limit resting 09:45–11:00, target0.382 × range, stop0.191 × range(exactly 2:1), 1-tick penetration on both resting limits, exit flat atcloseMs − 5 min.genomeHash 179c37c6…. - Sizing:
fixedNotionalMinor = 100,000,000,000— 10% of the 1,000,000,000,000 initial capital, passed explicitly, not defaulted. Whole shares, entry cost reserved outside the notional. 0 unfundable orders across all 19 symbols. - Costs, both models, side by side: the intraday parity model (6 bp adverse on each of entry and exit, 0 fee) is charged; the daily lane's
ceil(5 bp) + ceil(1 bp)per fill is reported as an uncharged comparison. The two models agree to the milli-basis-point on every cell — on the tier-A primaries they differ by 1,438 minor units out of 318,139,000,000 charged, the integer-rounding residue the amendments predicted. No result here depends on the choice of cost model. - Corpora: 19 sealed Alpaca SIP M1 corpora,
barInterval PT1M,usage replay_only(tier A) /replay_only_sparse_minutes(tier B) /replay_only_thin_tape(tier C, exploratory), read from~/.local/share/tradegg/corpusand re-verified at load. Panel bindings:roster-a 6331c87a…,nflx-post-split 0d584d8e…,roster-b-sparse ef8e9b2c…,roster-c-thin 4ed8a27b…. - Byte-exactness gate, binding and without opt-out: before a single statistic is computed the runner re-hashes the five golden Discovery-D corpora against their pinned sha256 and byte counts — 302,915,331 bytes verified, receipt
d4f5376a1a99…— and that receipt hash is a REQUIRED input to the seal. TheTT_ALLOW_MISSING_CORPUS=1escape used by CI cannot reach this path; the only way past the gate is to have the right bytes. The other twelve roster members are bound by their pinnedrevisionId, which the adapter re-derives from each file's contents. - Determinism: the whole pipeline was re-run in a fresh child process and produced the same seal hash. Re-sealing twice inside one process would have proved nothing.
Tier A — the nine pre-registered primaries
w* is the pre-registered median break-even win rate at a 12 bp round trip, computed from entry-side geometry before any outcome existed.
| Symbol | n | Win rate | Wilson 95% | Required w* | Net bps/trade | Bootstrap 95% | Profit factor | Loss share |
|---|---|---|---|---|---|---|---|---|
| NVDA | 312 | 31.4% | 26.5–36.8% | 48.1% | −14.25 | −19.25 … −9.26 | 0.523 | 68.6% |
| GOOGL | 328 | 27.7% | 23.2–32.8% | 51.2% | −14.65 | −18.17 … −11.14 | 0.435 | 72.3% |
| MU | 303 | 28.1% | 23.3–33.4% | 42.6% | −16.04 | −23.46 … −8.33 | 0.615 | 71.3% |
| TSLA | 333 | 29.1% | 24.5–34.2% | 44.3% | −17.81 | −23.07 … −12.32 | 0.510 | 70.9% |
| MRVL | 316 | 31.6% | 26.8–37.0% | 42.7% | −13.01 | −21.56 … −4.16 | 0.683 | 68.4% |
| AMD | 325 | 31.7% | 26.9–36.9% | 44.0% | −15.56 | −23.14 … −7.79 | 0.582 | 68.3% |
| AVGO | 306 | 28.8% | 24.0–34.1% | 44.8% | −16.89 | −22.73 … −11.02 | 0.522 | 71.2% |
| UBER | 320 | 29.4% | 24.7–34.6% | 47.2% | −16.54 | −21.23 … −11.66 | 0.468 | 70.3% |
| NFLX | 111 | 34.2% | 26.1–43.5% | 47.2% | −10.93 | −17.91 … −3.79 | 0.619 | 65.8% |
| Pooled (9) | 2,654 | 29.9% | 28.0–31.9% (clustered) | — | −15.40 | −17.90 … −12.84 (clustered) | 0.553 | 70.0% |
- Every upper bound is below every requirement. The best observed interval tops out at 43.5% (NFLX, n = 111) against a 47.2% requirement. This is not an underpowered null; it is a consistent miss in one direction.
- Gross expectancy is already negative: −3.40 bps per trade, 95% [−5.54, −1.26] across the 2,654 tier-A primary fills, before a single basis point of cost is charged. Two caveats a reader should carry: the same figure for the controls is −0.67 bps with an interval of [−2.11, +0.78] that includes zero, and for tier B it is −4.00 with [−8.00, +0.05] that also includes zero. The gross claim is established on tier A and tier C, not everywhere. Gross break-even at 2:1 is a 33.3% win rate regardless of symbol; the pooled interval tops out at 31.7%, below it. Per symbol the picture is weaker — only GOOGL's upper bound (32.8%) clears below 33.3% and the other eight straddle it (33.4%–43.5%) — but every symbol sits far below its own cost-adjusted requirement, which is the number that decides whether it can be traded.
- The short leg is not the problem, and it is not the rescue. Long trades: n = 1,347, 29.9% win, −15.54 bps. Short trades: n = 1,307, 29.9% win, −15.26 bps. The free-borrow assumption flatters the short leg and the short leg still loses.
- Restricting the pool to one panel changes nothing. Dropping NFLX (which rides its own 181-session post-split panel) leaves 2,543 trades on the single roster-a panel at 29.7% (clustered 27.8–31.7%) and −15.60 bps (clustered −18.25 … −12.96).
Activity and fills
| Symbol | Sessions | Decidable | Sessions with signal | Fill rate | Time-exit share | Max drawdown |
|---|---|---|---|---|---|---|
| NVDA | 512 | 497 | 88.3% | 71.1% | 0.0% | 4.7% |
| GOOGL | 512 | 497 | 92.8% | 71.1% | 0.0% | 4.9% |
| MU | 512 | 497 | 88.5% | 68.9% | 0.7% | 4.9% |
| TSLA | 512 | 497 | 90.1% | 74.3% | 0.0% | 5.9% |
| MRVL | 512 | 497 | 87.9% | 72.3% | 0.0% | 4.3% |
| AMD | 512 | 497 | 90.7% | 72.1% | 0.0% | 5.1% |
| AVGO | 512 | 497 | 88.1% | 69.9% | 0.0% | 5.2% |
| UBER | 512 | 497 | 93.2% | 69.1% | 0.6% | 5.3% |
| NFLX | 181 | 166 | 94.0% | 71.2% | 0.0% | 1.3% |
- The pre-registered activity forecasts land on the nose. SPY 40.0%, MU 88.5%, NVDA 88.3%, GOOGL 92.8% were all predicted from the same corpora before the run and are reproduced exactly. TSLA came in at 90.1% against a predicted 90.3%, and the two are different quantities: the prediction is the range-qualified share (449 of 497) and the reported figure subtracts the sessions that qualified but submitted no order — for TSLA, its single doji (448 of 497). Six is the tier-A pooled doji count, not TSLA's.
- Almost nothing reaches the time exit: 4 trades in 2,654 (0.2%). The 2:1 geometry is the operative mechanism here; the rule resolves at one of its two levels essentially always. It resolves at the wrong one.
- Pooled drawdown is 34.2%, but read it carefully. Per-symbol drawdowns run 1.3%–5.9%. The pooled figure chains nine symbols' return paths end to end (sequential capital, the
score.ts:100-113discipline applied across symbols rather than folds), so it measures the compounded decline of running all nine in series, not a portfolio drawdown.
Disposition ledger — tier A primaries (pooled)
| Disposition | Count |
|---|---|
atr_warmup | 135 |
incomplete_tape | 0 |
unqualified | 403 |
doji_no_trade | 6 |
side_filtered | 0 |
expired_unfilled | 1,079 |
unfundable | 0 |
sub_share_notional | 0 |
stop | 1,856 |
stop_gapped_at_fill | 0 |
target | 794 |
time_exit | 4 |
| Total | 4,277 = 4,277 sessions |
- The ledger sums exactly, per symbol and pooled. A ledger that does not sum is a void run, so this is checked before any rate is reported.
stop_gapped_at_fillnever fired. The SPEC-AMENDMENTS bucket for an entry that fills already at or beyond its stop is empty across all 5,637 trades in this run, so no result depends on its clamping rule.- The round-3 evaluator repairs moved five trades and not one statistic. Re-running against the repaired evaluator (open-through-target resolved before a same-bar stop; sparse-tail fallback bounded by the target; short quantity capped at the actual fill price) changed 5 of 4,972 trades (before tier C existed), all shorts — MRVL 2026-02-09, NVDA 2025-10-22, COIN 2026-06-10, HOOD 2024-08-15 and 2024-08-21 — whose gap above the sell limit had produced an entry notional slightly above the fixed 100,000,000,000. The cap brings every notional to at or under it. No disposition changed, no trade changed side of zero, and every win rate, expectancy, profit factor and interval in this report is identical at the precision reported.
unfundableandsub_share_notionalare zero on every symbol in both tiers, and the sealer refuses to produce an artifact where a pre-registered primary carries either. That is what makes the 10% sizing decision inert: it changed no trade's existence, only its size. A primary whose signals had been refused by capital would have been a test of the balance sheet, not the rule.
The video's claims, tested
Each claim was pre-registered as a two-sided test against a Wilson 95% interval, so "inconclusive" was available as an answer. It was not needed.
| Claim | Pre-registered test | Observed (tier-A pooled) | Verdict |
|---|---|---|---|
| Win rate above 70% | confirms iff lower bound > 0.70; refutes iff upper < 0.70 | 29.9%, upper bound 31.7% | refutes |
| Losses under 30% | confirms iff upper < 0.30; refutes iff lower > 0.30 | 70.0% lost money, lower bound 68.2% | refutes |
| Stopped out under 30% | same test, on the stop disposition | 69.9% | refutes |
| "One loser in a month" (≈95% win rate) | is 0.95 inside the interval? | interval is 28.2%–31.7% | no — it is not close |
| "Works better on some stocks than others" | cost-geometry pre-screen | see controls below | supported, and it is a cost effect |
- Every per-symbol test agrees with the pooled one. All nineteen symbols — nine tier-A primaries, four controls, four tier-B primaries and the two exploratory tier-C names — refute ">70%" and all nineteen refute "losses under 30%". There is no symbol where the rule works and no symbol where the answer is ambiguous.
- The one claim that survives is the caveat, not the promise. "Works better on some stocks" is real — and it is entirely a cost-geometry effect, quantified below. It does not produce a profitable symbol anywhere.
Controls — the four symbols the cost screen rejected up front
Declared before the run as cost-infeasible or demoted by a 55% break-even cut. Reported here, never counted toward the primary claim.
| Symbol | n | Win rate | Wilson 95% | Required w* | Net bps/trade | Profit factor | Sessions with signal |
|---|---|---|---|---|---|---|---|
| SPY | 164 | 32.3% | 25.6–39.8% | 85.8% | −11.73 | 0.147 | 40.0% |
| QQQ | 246 | 31.3% | 25.8–37.3% | 71.2% | −12.42 | 0.224 | 59.0% |
| IWM | 288 | 32.3% | 27.2–37.9% | 67.5% | −12.33 | 0.266 | 73.2% |
| MSFT | 328 | 29.6% | 24.9–34.7% | 55.9% | −13.64 | 0.382 | 92.4% |
| Pooled (4) | 1,026 | 31.2% | 28.0–34.5% (clustered) | — | −12.67 | 0.288 | 66.1% |
- The pre-screen was right, and it is measurable in the outcomes. 14 of the 53 SPY trades that reached their target (26.4%) still lost money after costs — the target was inside the round trip. QQQ: 2 of 77 (2.6%). The pre-registered forecasts, computed from entry-side geometry alone, were 26.6% and 4.1% of setups. Across all thirteen primaries in both tiers, that count is 0.
- The controls' profit factors are the worst in the run (0.147–0.382) even though their win rates are the highest, because their winners are tiny: SPY's average trade that reached its target nets +5.88 bps, against −20.14 bps on an average stop. On MU the same two numbers are +90.92 and −58.53. The index ETFs' opening ranges are too small a fraction of their ATR for a 38.2% retracement to clear the spread.
- Controls confirm rather than surprise. They were declared unusable before anyone looked, and they are unusable. That is the point of declaring them.
Tier B — the sparse-minute thin-name roster
Reported separately and never pooled into tier A. These four corpora exist only because the importer admits a no-print minute under an explicit rule (≤ 2% missing minutes per session, opening 15 complete), sealed under usage: replay_only_sparse_minutes.
| Symbol | n | Win rate | Wilson 95% | Required w* | Net bps/trade | Bootstrap 95% | Profit factor | Sparse sessions |
|---|---|---|---|---|---|---|---|---|
| HOOD | 334 | 24.9% | 20.5–29.8% | 41.6% | −23.84 | −32.22 … −15.17 | 0.516 | 1 |
| COIN | 318 | 33.3% | 28.4–38.7% | 41.7% | −12.40 | −21.81 … −2.54 | 0.713 | 12 |
| PLTR | 317 | 29.3% | 24.6–34.6% | 42.7% | −20.45 | −27.85 … −12.74 | 0.536 | 1 |
| META | 323 | 36.8% | 31.8–42.2% | 50.6% | −7.08 | −11.71 … −2.38 | 0.697 | 1 |
| Pooled (4) | 1,292 | 31.0% | 28.2–33.9% (clustered) | — | −16.00 | −20.27 … −11.66 (clustered) | 0.600 | 15 |
- Same answer, on different symbols — but not an independent replication. Tier B refutes the same claims on a different sealed panel and a tape with holes in it. Its four names were written down at 23:20 UTC, after the 22:45 probe had reported tier A's direction. No tier-B outcome was visible then (its corpora were sealed at 23:14) and the selection rule was entry-side cost geometry — but the direction was known, so this is a second test, not a blind one. See "Disclosure".
- Sparsity touched almost nothing: 2 trades of 1,292 had a missing minute inside their resting-order window, and 0 had one inside a realised holding window. The sparse-admission machinery was necessary to test these names at all and irrelevant to the result.
- META was the best symbol in the entire run (−7.08 bps, 36.8% win rate) and it still loses, with a bootstrap upper bound of −2.38 bps. There is no cell anywhere in this artifact whose interval touches zero.
incomplete_tapenever fires in tier B, by construction: its rule guarantees the opening 15 minutes are present. It fires three times in tier C, whose rule does not — see the tier-C section below.
Tier C — the thin-tape roster (EXPLORATORY, not a test)
Not evidence, and not charged as any. Tier C carries no evidence state, no Holm family and no assessMultiplicity call. It exists because the owner asked about smaller, recently-added names and the honest answer needed a corpus rule that admits a hole inside the opening 15 — which means these sessions cannot promise the opening range every Touch-and-Turn setup is priced off. Same pre-registered parameters; different question; never pooled with tier A or tier B.
| Symbol | n | Win rate | Wilson 95% | Net bps/trade | Bootstrap 95% | Profit factor | Loss share | incomplete_tape |
|---|---|---|---|---|---|---|---|---|
| VRT | 311 | 30.9% | 26.0–36.2% | −17.58 | −26.77 … −8.73 | 0.613 | 69.1% | 2 |
| APP | 354 | 26.3% | 22.0–31.1% | −24.86 | −33.25 … −16.41 | 0.528 | 73.2% | 1 |
| Pooled (2) | 665 | 28.4% | 25.0–31.9% (clustered) | −21.46 | −27.69 … −15.30 (clustered) | 0.564 | 71.3% | 3 |
- Same answer, worse. Tier C's pooled expectancy of −21.46 bps is the worst of the three tiers, and its gross expectancy is −9.46 bps per trade, 95% [−15.76, −3.12] before a basis point of cost. Both symbols refute ">70% win rate" and "losses under 30%" on their own.
- Contamination, stated beside the numbers rather than under them. Of the 919 sessions that submitted a resting order, 93 (10.1%) carry a hole somewhere in a window the rule cares about — 92 in the holding window, 3 in the order window, one session in both. That is the wide exposure measure. The narrow one: only 3 of 665 trades (0.45%) had a hole inside the window they were actually held through, because most exit long before the liquidation bell.
- 3 sessions produced no opening range at all and are recorded as
incomplete_tape(VRT 2, APP 1) rather than priced off whatever printed. They enter no rate's denominator. - The tape is genuinely thinner. APP is short 213 minutes across 90 sessions, VRT 33 across 23. Both stay inside the pre-declared ≤5% budget (19 of 390 in a regular session).
- Zero unfundable, zero
sub_share_notional, zerostop_gapped_at_fill, and zero targets that lost money after costs. These are volatile names with wide opening ranges — VRT's average trade reaching its target nets +90.13 bps and APP's +104.88, against −65.67 and −72.10 on a stop. The cost screen is not what kills tier C; the win rate is.
RDDT was refused, and no threshold moved
- RDDT missed the tier-C budget by one minute. Its 2024-07-31 session has 370 of 390 bars — 20 missing against a budget of 19 (≤5% of a regular session). It was refused at import and has no corpus on disk, which is why it appears in no table here.
- The ≤5% budget was written down before RDDT was tried, and it was not touched afterwards. Relaxing a tape rule by one minute to admit the symbol that just failed it is the cheapest possible way to manufacture a result, and the whole point of declaring the tier's rule in advance was to make that visible if anyone did it. Nobody did.
- Subsequent completeness follow-up: the owner later pre-registered a distinct ≥360/390 tier D, disclosed its ad hoc origin, and ran RDDT once; see
docs/research/touch-and-turn-tier-d-rddt-2026-08-26.md. - COHR was not admitted either and has no corpus on disk. Under the tier-B rule it had been recorded as 55 sessions over budget; whatever the tier-C arithmetic came to, it did not qualify, and no threshold was moved for it.
A note on pooled precision
Pooled intervals are date-clustered. The earlier ones were not, and were about 15% too tight.
- A pooled cell's series is nine symbols' trades laid end to end in symbol order. A stationary block bootstrap over that vector captures dependence only between adjacent entries, so two trades on the same calendar day in different symbols sit hundreds of positions apart and were resampled as independent — when the whole roster fades the same opening move on the same day.
- The resampling unit is now the session date: draw whole dates with replacement, carry every symbol's trades on that date together, take the trade-level mean of the pooled draw.
- Observed design effects (clustered half-width ÷ iid half-width), tier-A primary pooled: 1.193 on expectancy, 1.119 on the win rate. The iid expectancy interval was [−17.51, −13.26]; clustered it is [−17.90, −12.84]. The Wilson win-rate interval was 28.21%–31.69%; clustered it is 27.99%–31.88%.
- An independent analytic probe using a normal clustered standard error put the same two effects at 1.34 and 1.25. Both methods agree on the direction and rough size; the percentile bootstrap reported here is the more conservative choice about distributional shape and the less conservative about the width, so the honest range for the correction is at least the ~12–19% shown.
- Per-symbol cells are unclustered, by construction, not by omission. One symbol has at most one trade per date, so its clusters are singletons and the design effect is exactly 1. Those cells seal
nullrather than a fabricated design effect of 1. - None of this touches the verdict. Both intervals are nowhere near 0.50 or 0.70, and both expectancy intervals sit entirely below zero. What was overstated was the precision of a refutation, in a report that quotes those bounds to two decimals.
Deviations from pre-registration
The genome is unchanged. Five execution rules are not. Each is recorded in SPEC-AMENDMENTS.md against the §1 rule it supersedes; all five move against the strategy, and one has a quantified effect.
| Amended | §1 rule superseded | What changed | Direction | Measured effect |
|---|---|---|---|---|
| §1.2-e | "the exit orders … may fill on any subsequent bar, but never on the entry bar itself" | On the entry bar the STOP applies (the target stays deferred). A fill already at or beyond its stop becomes stop_gapped_at_fill. | Conservative — it can only add losses | ≈0.6 pp of win rate (the disclosure's own figure): roughly 30.5% unamended against the 29.9% reported |
| R5 / R6 | "truncating toward the entry" | Both distances round away from the entry with ceiling division | Conservative — the target never moves closer, the stop never tightens | Sub-tick; below reporting precision |
| Resting-limit penetration | §1.2-b applied the 1-tick rule to the entry only | The same penetration applies to the take-profit limit; a bar opening clean through the target still fills at its open | Conservative — a bare touch of the target is no longer a win | Not separately measured |
| Fixed-notional funding | §1.2-i named fixed notional without a granularity rule | Zero-whole-share orders get their own terminal disposition sub_share_notional; short quantity is capped at the actual fill price | Conservative — it refuses trades rather than inventing them | 0 sessions in this run |
| Direction control naming | §4.4 called mode 1 "breakout" | Mode 1 is a passive-limit continuation control; a breakout needs a stop-entry order type this family does not implement | Neither — a naming correction | Makes gate 8 untestable as written |
- The earlier report said "unchanged" and disclosed the entry-bar amendment only inside the quoted disclosure block. That was true of the genome and false of the evaluated rule set, and a reader would not have separated the two.
- The direction of every amendment matters more than its size. A conservative correctness fix cannot turn a working rule into a losing one, so none of them can have manufactured this refutation — but the ≈0.6 pp figure means the unamended rule would have reported ~30.5%, still nowhere near 70%.
What the owner asked for, and what happened to each name
The request was for NFLX and NVDA specifically, then smaller / recently-added S&P 500 names — RDDT, HOOD, COIN, APP, VRT, COHR. Every one was checked; not every one could be tested, and the reasons are rules, not preferences.
| Name | Outcome | Rule that decided it |
|---|---|---|
| NVDA | Tier A primary | Cost screen: median break-even win rate 48.1% < 55% |
| NFLX | Tier A primary, on its own 181-session post-split panel | Cost screen 47.2%; the July 2025 split forced a separate panel |
| HOOD, COIN, PLTR, META | Tier B primaries | Refused by the importer's exactly-390-minutes rule; admitted under the ≤2% sparse-minute rule with the opening 15 intact |
| VRT, APP | Tier C, exploratory only | Refused at ≤2%; admitted under the ≤5% thin-tape rule, which permits a hole inside the opening 15 |
| RDDT | Refused. Not on any panel. | 370 of 390 bars on 2024-07-31 = 20 missing, against a ≤5% budget of 19 |
| COHR | Refused. Not on any panel. | Over the tape budget; 55 sessions over at ≤2% |
- The admission rule was never relaxed to admit a name that had just failed it. RDDT missed by one minute on one day. Widening the budget from 19 to 20 would have admitted it, and that is exactly the edit a pre-declared rule exists to make visible. Instead a third tier was declared, with its own rule, before anything was run under it.
- The ≤2% budget was an a-priori guess, not evidence — the decision record says so — which is why the response to it blocking four names was a new declared tier rather than a quiet adjustment to the old one.
- PLTR and META were not on the owner's list. They entered tier B on the same cost screen as everything else; META in particular falsified the assumption that a mega-cap must be cost-infeasible.
Gates
Ordered as pre-registered; each is necessary.
| # | Gate | Tier A | Tier B |
|---|---|---|---|
| 0 | Corpus binding — byte-exact where a golden pin exists | pass (5 golden symbols) | n/a — bound by revisionId re-derivation |
| 1 | Disposition ledger sums, per symbol | pass | pass |
| 2 | Determinism — two fresh processes, one hash | pass | pass |
| 3 | Evidence floor: ≥126 decidable sessions and ≥60 execution fills per scored cell | pass — min 166 sessions / min 111 trades | pass — min 497 sessions / min 317 trades |
| 4 | Win-rate interval excludes 0.50 upward | FAIL — excludes_below | FAIL — excludes_below |
| 5 | Expectancy net of costs > 0, bootstrap lower-95 > 0 | FAIL | FAIL |
| 6 | Survives Holm across the primary symbols | FAIL | FAIL |
| 7 | Family multiplicity charge (3 attempts) | FAIL (0 survivors) | FAIL (0 survivors) |
| 8 | The direction control does not also pass | NOT EVALUATED | NOT EVALUATED |
Tier C is absent from this table on purpose. It is not a pre-registered test, so it has no gates to pass or fail. Giving it a row would invite reading its numbers as evidence of something. It is still charged as an attempt in gate 7.
- Gate 0 does not cover tiers B and C. The byte-exactness gate re-hashes the five golden Discovery-D corpora and nothing else; HOOD, COIN, PLTR, META, VRT and APP are bound by the adapter re-deriving each revision hash from the file's own contents. The earlier table read "pass" for tier B, which described a check that did not happen on it.
- Gate 3 now names both quantities, and the second one is new. Every claim in this report is a per-trade statistic, but the floor was being read on decidable sessions. NFLX clears 166 sessions and is scored on 111 trades; a cell with five trades would have passed the old gate unchanged. Both floors are now imported from
lib/paper/experiments(MIN_COMPLETE_SESSIONS = 126,MIN_EXECUTION_FILLS = 60) rather than re-typed here, so raising the repo's floor raises this one. - Gate 4 now FAILS rather than being "satisfied downward". The old gate asked only whether the interval excludes a coin flip, which an interval sitting entirely below 0.50 satisfies — and the artifact sealed that as an unqualified
true. The sealed value is now directional (excludes_below/excludes_above/straddles) and passing requiresexcludes_above. Nothing about the run changed; the gate now says what the prose always did. - Gate 6 is not close. Every one of the nine tier-A bootstrap p-values for H0: mean net expectancy ≤ 0 is ≥ 0.9972, against a first-step Holm threshold of 0.00556. The smallest p-value in the entire run is COIN's 0.9920.
- Gate 8 cannot be read from this artifact. It executed exactly one configuration, and running the control is a second test that belongs to the exploratory lane. Note also that
directionMode: 1in this family is a passive-limit continuation control, not the breakout control the protocol imagined — a breakout needs a stop-entry order type this family does not implement, so gate 8 as originally written is not testable here at all.
Multiplicity
- One charge for the whole family: 3 attempts, one per declared tier (A, B, C), 0 survivors,
distinguishableFromNoise: false. Charged through the repo's ownassessMultiplicityat its ownONE_SIDED_FALSE_POSITIVE_RATE = 0.025, so 0.075 is expected by chance. - It used to be one 1-attempt charge per tier, which is the laundering the pre-registration forbids — "41 configurations do not buy 41 independent 2.5% budgets". Two tiers charged separately are two independent budgets. They are now one.
- With zero survivors this verdict carries no information beyond gates 4-6. It is arithmetically implied by them and is reported as a charge, not as independent corroboration. The sealed artifact says so in a field:
carriesInformationBeyondGates4To6: false. It starts to bite the moment more than one tier survives. - Tier C is an attempt even though it can never be a survivor. It has no gates, so it can only ever enlarge the denominator — which is the correct direction and the reason it is not quietly excluded.
- Holm–Bonferroni across each tier's primary symbols is reported alongside as gate 6. It rejects nothing at either tier.
- The exploratory grid has not run and is not charged here. When it does, its 41 configurations × the roster join the global attempt denominator; they do not buy independent budgets. Adding this family can only make promotion harder for every other family, which is the intended direction.
- Nothing in this run is a "best cell". A single pre-registered configuration has no siblings to hide.
Disclosure
Byte-identical to the decision record, markdown emphasis and ASCII apostrophes included. Two digests, because the two answer different questions:
shasum -a 256over the quoted paragraph gives1c63f463d30e9ce051b645d4c136133ef497dd75eb1b48d14923c45295e4a78a.- What the seal binds is its
intradayHash— sha256 over the canonical JSON encoding,sha256(JSON.stringify(text))—9c2ae4cc3031666cf38a7b087654c16d2b003f769fccba9d1b65ab61c3dc4fa1. - The earlier report published only the second under the label "sha256". A reader who checked it the obvious way got a different number and would have concluded the artifact was wrong. The sealed text also carried two U+2019 apostrophes and a stripped
**, so "verbatim" was false. Both are fixed and a test now asserts the bytes rather than four apostrophe-free substrings.
Disclosure (22:45 UTC): the P-TT1 adversarial panel's verifier probes ran the committed evaluator on the sealed corpora to QUANTIFY a bias (entry-bar stop deferral): pooled over 9 corpora it reported ~2,498 filled trades, a pooled win rate near 31% and negative mean gross expectancy on most names, with the bias worth ~0.6 pp of win rate. These numbers were observed by a verifier, not by the sealed run, and AFTER the primaries/controls/parameters above were written down. No parameter, roster label or rule was changed in response to them; the repairs that followed (entry-bar stop applies; gap-through-stop disposition; symmetric penetration; round-away-from-entry) are correctness fixes whose direction is known to be conservative. P-TT3's sealed report must carry this disclosure verbatim.
- The timeline, precisely, because the earlier report got it wrong. The primaries, the controls and the cost pre-screen are dated 22:36 — before the 22:45 probe. The roster tiers — tier B's four names and the tier-C rule — are dated 23:20, which is after it. The earlier draft cited 23:20 as if it supported "written down before the probe"; it does the opposite, and the decision record's own "before any outcome" label on that line is wrong too.
- What the probe did and did not see. It pooled the nine tier-A corpora (NVDA, GOOGL, MU, TSLA, MRVL, AMD, AVGO, UBER, NFLX). No tier-B corpus existed yet — they were sealed at 23:14 — so no tier-B outcome was observed, and tier B's four names were selected on entry-side cost geometry (median break-even win rate under 55%), not on outcomes.
- What that costs tier B as evidence. The tier-A direction was known when tier B was written down. Tier B is therefore a second test on a roster chosen with that knowledge — not an independent replication, and the phrase "same answer, independently reached" above should be read with that qualifier. The same applies to META being the best symbol in the run: META entered the roster at 23:20.
- The sizing decision is undated.
fixedNotionalMinor= 10% of initial capital sits in the decision record after the 22:45 entry with no timestamp of its own, so it cannot be claimed as pre-probe. It is inert here in any case: zero orders were unfundable and zero were sub-share, so sizing changed no trade's existence. - The run is therefore a confirmation of a known direction, not a blind discovery. That is a weaker claim than a fully blind test and it is the claim being made.
- The repairs made after the probe all move against the strategy, so they cannot have manufactured this refutation. A conservative correctness fix cannot turn a working rule into a losing one.
What would be needed to say anything beyond discovery_only
Nothing in this run gets there, and the ceiling was stated before it ran.
- A holdout. Discovery-D is selection-only. Every session here is in-sample by construction; there is no held-out period to confirm anything against.
- A validation panel. All three (v18-20, v20-24, v22-24) are PENDING and do not exist as sealed artifacts.
- The global promotion gate. It is arithmetically shut until the P14 decision lands, independently of anything this family measures.
- Gate 8, in the exploratory lane. The direction control has to run — and be reported with all 40 of its siblings, not alone.
None of that is worth doing here. The result is not "inconclusive pending more evidence"; it is a refutation with a 2,654-trade pooled interval that never approaches break-even on any symbol, in either tier, under either cost model. The honest next step is to record it and stop, not to tune it. No parameter search was run precisely so that this sentence can be written.
Provenance
- Lineage. Four seals, one conclusion.
da01a8b5…(round-2 evaluator, 17 symbols) →9f22a84a…(round-3 evaluator and the binding byte-exactness gate; 5 short trades re-priced, no statistic moved) →3e512130…(tier C added as an exploratory section; tiers A and B byte-identical) →b5a705ef…(this one: date-clustered pooled intervals, gross expectancy with its own interval, a directional gate 4, an execution-fill floor, one family multiplicity charge, and the disclosure restored byte-identical). Superseded seals are recorded, not deleted; every one of them refuted the same claims. - Artifact:
~/.local/share/tradegg/touch-and-turn/primary/b5a705ef7206a755cc32107ae3291120c8db2752cf8b0de6b0146cf644cee33d.json - Re-derive the digest: the file carries
canonicalPreimage— the exact bytes that were hashed.sha256(canonicalPreimage) == sealHash. - Re-derive the statistics: the file carries all 5,637 per-trade rows, bound into the seal through each cell's
tradeLedgerHash. Every reported rate and expectancy — gross, charged, daily-model, iid and date-clustered — is recomputable from those rows alone. - Re-run it:
node --max-old-space-size=6144 scripts/run-touch-and-turn-primary.mjs(~2 minutes including the byte-exactness gate and the fresh-process determinism check; ~1.9 GB peak RSS, one corpus live at a time). - Byte-exactness receipt:
d4f5376a1a9910214da3685b4a9b6df96088f438418ce11f7aefc67493535bd7over GOOGL, MU, NVDA, SPY and TSLA; also insideartifact.corpusByteExactness, where the seal depends on it. It covers those five and no others. - Read the verdict in the app:
/lab-docs/strategies. The/labcatalog row links there rather than printing a repo path as plain text. - Code:
lib/paper/families/touch-and-turn/v1/{report,run-primary,run-primary-disk}.ts; rules as amended in that directory'sSPEC-AMENDMENTS.md. - Principal weakness, unchanged and stated on the /lab row: short-side borrow and locate are assumed available and free; no live locate check is modelled anywhere in the repo.