FRAKTAL·ICIOUS
Metric research
Metric research · updated 2026-09-16

CFTC Commitments of Traders — leveraged fund positioning

We scored hedge-fund positioning in CME Bitcoin futures for months before anyone checked the distribution. Its bearish half had never fired once in 2,608 days, so we set it to zero.

Status Tested and rejected Weight in our model 0 Also called CFTC COT, Commitments of Traders, leveraged fund net positioning, CME Bitcoin futures positioning, smart money positioning

What it measures

The CFTC publishes a weekly Commitments of Traders report breaking open interest in CME Bitcoin futures down by trader category. We read one number out of it: the share of the leveraged-fund category's positions that are long. The intent was a contrarian read — a crowded long side is a crowded exit, and an extreme short side is capitulation that tends to mark lows.

Our data: The CFTC's own weekly report, banked as a daily series with each row carrying the publication date the figure actually became public, so nothing can be scored before it was knowable. Measured over the model's scoring era, 2019-07-01 onward, restricted to days with at least 3,000 contracts of open interest — 2,608 days. The way those weekly rows were mapped onto days turned out to be a second, separate defect; it is the subject of the second test below.

What happened when we put it in the model

A/B run 2026-08-21 · tested at weight 0.03 · verdict revert
MeasureWithout itWith it
Walk-forward geomean edge1.2111.214
Navigator GPA3.553.55
Cycle-phase GPA2.912.91

Read the direction of this one carefully, because it runs opposite to the other pages here. This signal was already LIVE at weight 0.03; the test removed it. So the column on the right is the model as it had been shipping, and the column on the left is the model after we took the signal out — which is where it sits today. Measured against the model as it stood on 2026-08-21; that baseline is several re-baselines old and is NOT the current published record. The register is the only authority on today's figures. Recorded at src/services/cycleProjection.ts:3347.

A/B run 2026-08-21 · tested at weight 0.03 · verdict revert
MeasureBroken wiring (shipped)Corrected wiring
Walk-forward geomean edge1.2041.220
2019-22 cycle fold2.122.30
Navigator GPA3.553.48
October 2025 top, letter gradeCD
Current-cycle dollar edge1.921.58

A different test entirely: the signal is live at 0.03 in both columns, and what changes is whether the weekly report is read on the days it was actually in force. The correction moves 2,296 of 5,045 days, 45.5 percent of history, and both arms reproduce byte-identically on re-run. Note that the shipped walk-forward figure reads 1.204 here and 1.214 in the ablation above — the two runs were recorded days apart against different dated baselines, and neither one is today's number. Recorded at brain/IMPROVEMENT_PROPOSALS.md, entry dated 2026-08-20.

What we concluded

The finding that matters is not that this signal was weak. It is that it was not the signal we said it was. A two-sided contrarian indicator whose bearish side is unreachable is a one-sided indicator, and we had been describing it the other way in our own source code while it voted on the composite every week.

Nothing exotic hid it. The threshold was written in the code and the distribution was in the data; nobody had put the two next to each other. It surfaced only when a per-signal forensic trace was built in August 2026 to answer an unrelated question, and the trace printed what the signal was actually contributing rather than what it was meant to.

Taking it out cost us. Walk-forward geomean edge fell from 1.214 to 1.211 — the signal was, very slightly, helping. We removed it anyway. Three thousandths sits inside the noise line this stack had already set when it rejected the dollar index at five thousandths, and the alternative was keeping a number that implies we read positioning data when what we actually read was a constant. Paying 0.003 to stop implying evidence we do not have is a trade we will make every time.

There is a second defect underneath, and it is the more uncomfortable one. The COT report is weekly, carried as daily rows; the loader collapsed every row in a week onto the single date the report was published, so the signal scored on 14 percent of days and was simply absent on the other 86 — 963 daily rows carrying only 138 distinct publication dates. Because the composite is normalised by the weights that are active on a given day, a signal popping into the vote one day a week and vanishing for six was not just adding its own reading, it was rescaling everything else.

Fixing that is one line, and it is strictly look-ahead-free — it reads the most recent report published on or before each day. It also makes the model worse exactly where we care most right now. The historical walk-forward improves, 1.204 to 1.220, and the 2019-22 fold improves with it. But the navigator GPA falls from 3.55 to 3.48, the October 2025 top drops from a C to a D, and the current cycle's dollar edge falls from 1.92 to 1.58.

A correctness fix that improves the distant past and degrades the recent past is not a tuning result you can shrug at. It says the model's calibration had come to lean on the defect. We are publishing that because it is the least flattering reading available and because it is ours — the alternative, quietly re-weighting the signal until the gate goes green, is fitting to the thing the walk-forward exists to prevent.

The signal is off, not deleted. Both arms of both tests re-run from a single environment variable, which is the only reason a page like this can be written at all.

Where our own notes disagree

This one is unresolved on purpose, and it is an open decision rather than a closed finding. The wiring correction above is a genuine bug fix that our own shipping gate refuses, and three responses are on the record: leave it off and keep the defect documented, which is where it sits today; ship it and accept that the October 2025 top grades a D, which is defensible if correct data semantics outrank a single event grade; or redesign the scorer against the distribution the data actually has before turning it back on. We have not picked one. Anyone reading this should know that the version of our model serving the site today is the one that still contains the wiring defect, with the signal weighted at zero so that the defect changes nothing — and that the case for the other two options has not been refuted, only deferred.

Other candidates we put through the same gate

Each of these was wired into the same model, measured against the same walk-forward gate, and written up with the numbers it produced — including the cells where it helped. A candidate can improve several measures and still be reverted, because the gate is a conjunction and buying one measure by selling another fails it.

Tested and rejected · 1 A/B run · 1 of 4 measured cells improved · 0 shipped

Reserve Risk

Reserve Risk prices long-term-holder conviction against the market. We tested it inside our model, measured it, and it made the model worse — so it carries zero weight.

Tested and rejected · 3 A/B runs · 7 of 15 measured cells improved · 0 shipped

The Trade-Weighted Dollar

A falling dollar really does precede Bitcoin strength, in every cycle we can measure. We swept it into our model at three weights and every one of them failed the out-of-sample test — so it carries zero weight.

Tested and rejected · 1 A/B run · 1 of 3 measured cells improved · 0 shipped

Short-Term-Holder SOPR

When recent buyers start selling at a loss, a bottom is usually near. That is true, and our model already knew it — adding this signal made the model worse, so it carries zero weight.

Tested and rejected · 2 A/B runs · 4 of 9 measured cells improved · 0 shipped

Net Unrealised Profit/Loss

The heaviest valuation signal this model ever carried, removed at a measured cost to its sharpest recent call — not because it scored badly, but because it is algebraically the same number as a signal we were already scoring beside it.

Every signal we add has to survive the same gate, and most do not. Our published record is measured over 18 walk-forward folds (the year and cycle folds tile the same span, so they are not independent of each other). The figures in the table above are dated results against the model as it stood on the day of the test — the register is the only authority on where the model stands today.

Read the current call in today's briefing, or the timestamped record of every call in the call ledger.

Fraktalicious Research · Metric index · Briefing archive · Call ledger · Privacy
Nothing on this site is financial advice. This page is a record of a test we ran on our own model, published with the numbers it produced.