FinChiefs Research
Under the hood

An autonomous research desk that shows its work.

This page is the machine looking at itself: the data it watches, the desks and scouts that run every day, the adversarial review every idea has to survive, the ideas it killed, and the changes it made to its own models. Everything below renders live from the system’s own records — nothing is written by hand.

In developmentAutomated research desk — in development. Notes are AI-generated, published unedited, and may contain errors. Not investment advice.
1,250
data series tracked
markets, economic releases, currencies, futures, options, search interest
120
ideas taken to a verdict
51 more open on the desks right now
93%
of tested ideas were rejected
most ideas die — that is the point
24
models live
each scored against a simple benchmark
0
research notes published
0 in the last 7 days
20
changes it shipped to itself
written, checked and merged by the system

The machinery

Every part of the system, as registered in its own database — titles, descriptions and activity come straight from the live registry. When a new agent joins the desk, it appears here the morning after, with no edits to this page.

The desks

7

Seven specialist desks, each covering one corner of the markets. Every desk publishes a dated note every day — no skipping quiet days, no editing after the fact.

Deskdailyactive today

Macro desk

Watches the global economic calendar and flags the data releases that came in meaningfully above or below what recent history predicted. Publishes a daily note on which surprises matter and why.

informsHouse viewDaily briefingWeekly roundup+1 more
Deskdailyactive today

Rates desk

Covers government bond markets: the US yield curve, its slope, and real yields. Publishes a daily note on what moved yields and what the curve implies about growth and policy expectations.

informsHouse viewDaily briefingWeekly roundup+1 more
Deskdailyactive today

Currencies desk

Covers the major world currencies. Publishes a daily note on what drove exchange-rate moves and a weekly deep dive into the frameworks behind them.

informsHouse viewDaily briefingWeekly roundup+1 more
Deskdailyactive today

Equities desk

Covers the major stock indices in the United States and Europe, along with the market's own gauge of expected turbulence. Publishes a daily note on what moved and how the two regions compare.

informsHouse viewDaily briefingWeekly roundup+1 more
Deskdailyactive today

Energy desk

Covers crude oil and natural gas: prices, the shape of the futures curve, and how traders are positioned. Publishes a daily note on what is driving the energy complex.

informsHouse viewDaily briefingWeekly roundup+1 more
Deskdailyactive today

Precious metals desk

Covers gold, silver, platinum and palladium, including the gold-to-silver ratio and how speculators are positioned. Publishes a daily note on what is moving the metals.

informsHouse viewDaily briefingWeekly roundup+1 more
Deskdailyactive today

Technical analysis desk

Reads price behaviour across every market the house covers: trends, momentum, and turning points. Publishes a daily cross-market technical read the other desks can lean on.

informsHouse viewDaily briefingWeekly roundup+1 more

The scouts

10

Standing watchers that brief every research session: what just surprised, what quietly broke, what new data arrived, and what the rest of the world published.

Scouteach research run

Data scout

Notices when newly collected data series arrive and puts them in front of the right desk, so fresh data never sits unused. Also confirms when a requested series has started flowing, closing the loop on data requests.

informsResearch analyst
Scouteach research run

Surprise scout

Flags which economic releases and market moves were genuinely surprising relative to recent history, so every desk weighs the week's big events rather than only its own corner.

informsResearch analyst
Scouteach research run

Structural radar

Watches long-standing relationships between markets for signs they are breaking down or reconnecting, and raises each shift as a question for the desk to investigate.

informsResearch analyst
Scouteach research run

Central-bank scout

Tracks upcoming central-bank decisions, what markets expect them to deliver, and how recent decisions compared with those expectations.

informsResearch analyst
Scouteach research run

Geopolitics scout

Filters the outside research feed for geopolitical developments worth weighing — a curated short list, not a headline firehose.

informsResearch analyst
Scouteach research run

Internet scout

Searches the live web for what is too fresh to be in the research feed yet: breaking events, upcoming calendar items, and named methods worth replicating on our own data.

informsResearch analyst
Scouteach research run

Prior-work reviewer

Re-reads the desk's own published research alongside outside commentary, and flags gaps, contradictions and stale claims worth a second look.

informsResearch analyst
Scouteach research run

Forecast catalog briefing

Reminds each desk which in-house estimates already exist and how accurate they have proven, so analysts reach for the best available number instead of a stale official one.

informsResearch analyst
Scouteach research run

Coverage checker

Keeps a list of collected data the desk has never analysed and steadily surfaces the oldest gaps until each one gets a verdict.

informsResearch analyst
Scoutweeklyactive this week

Idea distiller

Reads the week's outside research and boils it down to a ranked list of specific, testable ideas for the desks.

informsResearch analyst

The review panel

5

No finding counts until it survives review. A skeptic tries to knock it down, a statistics reviewer audits how it was produced, and a closer forces open questions to a verdict.

Review paneltwice weekly

Research analyst

The panel's working analyst: forms a hypothesis, tests it against the data, and writes up what held and what did not. One runs per desk, with every claim recorded for review.

informsSkepticMethodologistResearch ledger+1 more
Review paneltwice weekly

Skeptic

An adversarial reviewer who tries to knock each new finding down — hunting for lucky streaks, cherry-picked windows and simpler explanations — before it can influence anything.

informsResearch ledgerTrusted publisher
Review paneltwice weekly

Methodologist

A statistics reviewer who audits how each result was produced: sample sizes, test choices, and whether the evidence actually supports the stated conclusion.

informsResearch ledgerTrusted publisher
Review paneltwice weekly

Backlog closer

Works through the desks' open questions and settles them one by one — confirmed, refuted, or genuinely waiting on data — so the question pile shrinks instead of growing. A refuted answer counts as a win.

informsResearch ledgerFollow-through wiringTrusted publisher
Review panelmonthly

Retrospective reviewer

Periodically reviews a desk's own track record — what paid off, what was noise — and proposes concrete changes to how the desk works.

informsTrusted publisher

The models

11

The standing forecasting and interpretation models — each one scored against a simple benchmark on live data, and benched when it stops earning its place.

Modelweeklyactive this week

Fed policy path model

Estimates where US interest-rate policy is headed from the economic data, and compares that estimate with the path markets have priced in.

informsRates deskForecast digest
Modelweeklyactive today

Central-bank tone reader

Reads each central-bank statement and scores how hawkish or dovish it sounds, then compares that tone with what markets expect — flagging when words and prices disagree.

informsRates deskCurrencies desk
Modeldailyactive today

Search-interest monitor

Tracks what people are searching for online — job worries, mortgage costs, layoffs — and turns shifts in public attention into an early read on the economy that every desk can use.

informsMacro deskRates deskCurrencies desk+4 more
Modelweeklyactive this week

Crowding gauge

Measures how crowded the bets are across markets by combining futures positioning with price momentum — a contrarian warning light when everyone leans the same way.

informsRates deskCurrencies deskEquities desk+2 more
Modelweeklyactive this week

Growth tracker

Combines the monthly activity data as it arrives into a running estimate of the current quarter's US economic growth, well before the official number is published.

informsGrowth-forecast poolForecast digest
Modelweeklyactive this week

Financial-conditions model

Reads market-based stress gauges — borrowing costs, credit spreads, volatility — for an early signal on jobs, output and growth.

informsGrowth-forecast poolForecast digest
Modelweeklyactive this week

Goods-inflation model

Tracks the pipeline pressures behind US goods prices — the dollar, producer prices, global supply-chain strain — to anticipate where goods inflation is heading.

informsForecast digest
Modelweeklyactive this week

Inflation-expectations gap

Compares household and market inflation expectations with a model-based view, and uses the gap between them to anticipate inflation surprises.

informsForecast digest
Modelweeklyactive this week

Commodity pass-through model

Uses moves in oil, gas and copper futures to anticipate their knock-on effects on consumer energy prices and industrial activity.

informsForecast digest
Modelweeklyactive this week

Growth-forecast pool

Blends the house's independent growth models into one estimate, weighting each by its proven accuracy.

informsForecast digest
Modelweeklyactive today

ECB policy rule model

Estimates where a textbook policy rule says euro-area interest rates should be, given inflation and unemployment, and compares that with what markets have priced in — a stance read, not a forecast, shown alongside the US equivalent.

informsRates desk

Self-correction & delivery

19

The machinery that keeps the desk honest: synthesis and delivery of the research, scoring against professional research houses, and the checks that catch the system drifting.

Self-correctionweeklyactive today

Tone reader's critic

Audits the tone reader's interpretations against independent commentary and adds any angles it missed to the reading checklist — a model that improves its own rubric.

informsCentral-bank tone reader
Self-correctiondaily

House view

The morning-meeting synthesis: reads all seven desk notes, surfaces where they agree and disagree, and states the house's single top-of-book read for the day.

informsDaily briefing
Self-correctiondaily

Daily briefing

One email each morning carrying the house view and every desk's daily note — the single consolidated read of the day.

Self-correctionweeklyactive this week

Forecast digest

A weekly consolidated view of every in-house forecast and how each has been scoring against reality.

informsWeekly roundup
Self-correctionweeklyactive this week

Weekly roundup

A weekly newsletter pulling together the week's forecasts, refreshed research papers and model health checks.

Self-correctionweekly

Street benchmark

Scores the desks' published notes against what professional research houses published: did we explain the moves they explained, and where did we see something they missed?

informsRetrospective reviewer
Self-correctionmonthlylast active Jul 16

Model portfolio review

A monthly review of every model on the books that forces a decision on any that have drifted: recalibrate, retire, or keep.

Self-correctiondaily

Daily question drain

Every day, picks the highest-priority open questions and pushes them through to an answer, so the research backlog keeps moving between full review cycles.

informsResearch ledgerTrusted publisher
Self-correctiondaily

Change-queue drain

Retries proposed changes that had already passed review but were held up by a temporary failure, and reports how many remain waiting.

Self-correctiondaily

Cycle health check

Checks that the research cycle itself is working — sessions completed, approved changes landed — and raises an alarm, or fixes the known causes, when the loop silently stalls.

Self-correctionweekly

Forecast demotion gate

Automatically benches any forecasting model that keeps losing to a simple benchmark on live data — with a clear bar, a record of why, and an easy way back.

informsModel portfolio review
Shared machinery

Research ledger

The shared logbook of every hypothesis, finding and verdict. Questions keep their history, negative results are preserved, and nothing is asked twice.

informsBacklog closerRetrospective reviewerDaily question drain
Shared machinery

Question quality gate

Checks every new research question at the door: duplicates are merged, questions that cannot say what would prove them wrong are sent back for rework, and the rest are priority-scored.

informsResearch ledger
Shared machinery

Follow-through wiring

Makes sure a confirmed finding actually lands somewhere — a model update, a code change, or a wiring proposal — instead of being celebrated and forgotten.

informsResearch ledgerTrusted publisher
Shared machinery

Trusted publisher

Takes the work the research agents commit, runs the full battery of checks, and merges it only when everything passes and the review panel found no fatal flaw; everything else is held for human review.

Shared machinery

Market-driver book

A desk's running book of the stories believed to be moving its market, verified against how prices actually behave rather than how often the story is repeated.

informsEnergy desk
Shared machinery

Precedent engine

Finds comparable historical episodes — say, every inflation surprise of this size — and reports what actually happened next, always shown alongside the everyday base rate, so the desks can cite measured history instead of impressions.

informsMacro deskRates deskCurrencies desk+4 more
Shared machinery

Forecast harness

The shared machinery every forecasting model runs on: one honest scoring standard against a simple benchmark, one storage format, and one publishing path.

informsForecast digestForecast demotion gate
Shared machineryweeklyactive this week

Framework health check

A weekly check on the shared analytics every desk relies on — the market-relationship radar and regime gauges — so a break in common machinery is caught centrally, not desk by desk.

What the desk rejected

Most research ideas do not survive testing — and a desk you can trust is one that shows you the bodies. These are the most recent ideas the desks killed, with the evidence that killed them, quoted from the internal record.

rejectedEquities deskraised July 24, 2026

Well-powered null on 20y (n=1021, 2006+): S&P 500 leveraged-fund COT index rank-IC vs 4-week fwd SPY return +0.04 (ns, wrong sign for contrarian), quarter +0.10 (ns); at extremes crowded-LONG -> HIGHER fwd returns (spread -0.26%/-1.21%), opposite of the contrarian thesis. Tested via the new read_history loader (market-research#285). COT stays crowding CONTEXT, not a forward-return model.

rejectedRates deskraised July 22, 2026

Prior session (2026-07-22, fi_cot_tenor_2026_07_22.py): Spearman IC of belly-avg COT index (2y+5y average) minus 30y COT index vs forward 4w slope change = 0.021 (p=0.505). OLS slope = null. Conditional mean at extreme belly divergence not distinguishable from unconditional. Matches 'already refuted' entry.

rejectedRates deskraised July 22, 2026

Prior session (2026-07-22, PR #224): TIPS 30y-10y real yield slope as carry signal for 30y bonds, 2010-2026 (≈16 years). NW-t=0.424 — far below 1.65 bar. This matches the 'already refuted' entry. Consistent with general carry refutation findings.

rejectedRates deskraised July 22, 2026

1990-2010 (20.9y): NW-t daily=1.178. 1995-2010 (15.9y): NW-t daily=0.469. Pre-2022 ex-inversion (44.7y): NW-t monthly=0.916. All below 1.65 bar. The 5y OOS (2021-2026) NW-t daily=1.973 is above bar but includes the entire 2022-2023 inversion episode — confirming the inversion era is necessary for the signal to appear significant.

rejectedRates deskraised July 22, 2026

Prior session (2026-07-22, PR #224, scripts/fi_cot_tenor_2026_07_22.py): IC (Spearman) between COT index and forward 4w/8w yield change is near zero for all four Treasury tenors (2y/5y/10y/30y). Conditional means at extremes (≥85 or ≤15) not significantly different from unconditional. All p-values > 0.20. This matches the 'already tested & rejected' list entry.

rejectedRates deskraised July 22, 2026

Swept 2m–24m (42d–504d), full FRED daily 10y yield history (1962-2026, n≥15,616d per lookback). Sharpe at 252d=0.567; global peak at 483d (23m)=0.668; drop from peak to 252d=15.1% (below 20% cherry-pick threshold). 9-15m range (189-315d) shows 33.9% Sharpe variation (max 0.567 at 252d, min 0.375 at 315d) — above 20% stability threshold, so no stable plateau. Signal exists broadly: Sharpe 0.375–0.6

rejectedRates deskraised July 22, 2026

1990-2010 holdout (n=5255d, 20.9y): 2s10s_z carry Sharpe=0.249, NW-t daily=1.178, NW-t monthly=1.329 — both below 1.65 bar. 1995-2010 (15.9y): Sharpe=0.113, NW-t daily=0.469. Broader pre-2022 sample (44.7y ex-inversion): NW-t monthly=0.916. Inversion era 2022-2023: Sharpe 1.263 (2y window) — this episode carries the full-sample result. scripts/fi_executor_2026_07_22.py (2026-07-22).

rejectedRates deskraised July 22, 2026

The validation_ref in spec.py previously stated 'The signal bar has never been cleared here. Establishing or refuting it is a live research lead.' This was superseded by the bond-momentum-in-repo-validated finding (NW-t=4.077, Sharpe=0.568, 1962-2026). Updated to reflect: (1) validated NW-t=4.077 status; (2) 2y tenor dominance finding (Sharpe=2.09, NW-t=3.79); (3) lookback sweep result (no stable

rejectedRates deskraised July 22, 2026

n=1016 weekly obs (2006-2024) per tenor. 2Y: IC(COT vs dy4w)=-0.003 (p=0.917), mean dy4w at extreme-long (COT>=85, n=207) = -0.014% vs unconditional -0.002% — slightly NEGATIVE (yields fell when 2Y was at extreme-long, opposite of long-squeeze narrative). 5Y: IC=-0.019 (p=0.538), null. 10Y: IC=+0.047 (p=0.136), mean dy4w extreme-long=+0.030% vs 0.000% unconditional, NW-t=1.368 — closest to signal

rejectedRates deskraised July 22, 2026

n=1016 weekly obs (2006-2024). Spearman IC of mean(COT_2y, COT_5y) - COT_30y vs forward 4-week 5s30s slope change: 0.021 (p=0.505). 8-week horizon: IC=0.032 (p=0.308). OLS r=0.032 (p=0.303) at 4w, r=0.040 (p=0.202) at 8w. All null. Current divergence = 71.8 (93.5th pctile — historically extreme belly-long vs long-end). Conditional mean slope change at extreme belly readings (>85th pctile, n=153):

rejectedRates deskraised July 22, 2026

Full-sample (2010-08-17 to 2026-06-17, n=192 monthly): NW-t=0.424, Sharpe=0.427. Well below the 1.65 bar. Current real slope = 56bp, z-score = -1.75 (currently compressed vs rolling average). The signal is null — the real yield slope does not carry as a timing signal for long/short 30y vs 10y TIPS returns.

rejectedEquities deskraised July 22, 2026

32 independent regime-exit events identified (SPY 2006-2026). 20d after exit: mean=+0.41%, t=0.72, p=0.48 (clearly null). 40d after exit: mean=+1.04%, median=+1.75%, t=1.64, p=0.11 (borderline but not significant at α=0.05). Historical episode stats: 33 negative-corr episodes, median=10d, mean=33d, q90=79d. Current episode: 83 days at 90.9th pctile. With n=32 exit events, the test is underpowered

rejectedEquities deskraised July 22, 2026

4w: top-quintile XLE mean_fwd_SPY=+0.16%, t_vs_zero=1.06 (raw), adj≈0.23. OLS t=−4.65 raw → adj≈−1.01. Bottom quintile: +0.75%; mid: +0.94% — XLE leaders actually have lowest forward returns. 8w: top quintile +1.23% vs bottom +1.10% (essentially identical). Current XLE relative at 93rd pctile (+8.22% vs SPY 1m). No significant headwind signal at either horizon.

rejectedEquities deskraised July 22, 2026

OLS slope=−0.111 (t=−7.85, r²=1.2%, n=4 980 obs, 2006–2026 EWJ/SPY daily). Overlap-adjusted t≈−1.7 (21-day forward returns sampled daily, ÷√21). Quintile table perfectly monotone: Q1=−0.09%, Q2=−0.35%, Q3=−0.62%, Q4=−0.65%, Q5=−0.82% fwd relative return. Japan outperformers >3% (n=651): mean_fwd_rel=−0.81%, raw t=−5.4, adj t≈−1.2. Japan underperformers >3% (n=1067): −0.11%, t=−1.13 (not significan

rejectedEquities deskraised July 22, 2026

Genuine broadening (RSP>SPY AND SPY>0): mean_fwd_4w=0.79%, t=8.71 (raw), adj≈1.90. Defensive rotation (RSP>SPY AND SPY≤0): 0.21%, t=0.90 (raw), adj≈0.20. t_A_vs_B=2.84 raw → adj≈0.62 after ÷√21 overlap correction — not significant. The commentary.py _breadth_words() function ALREADY implements this distinction (SPY direction conditioning added previously). Current reading: Bucket A genuine broaden

rejectedEquities deskraised July 22, 2026

2026 recheck using index-futures data (n=818 days): negative-corr regime Sharpe=1.35 vs positive-corr Sharpe=−0.10. Sharpe diff=1.46. Daily-label permutation p=0.204 (5000 perms). Regime diff t=1.21, p=0.225. Consistent with prior results: block-bootstrap p=0.297, label-shuffle p=0.187. Adding ~4 years of futures data has not changed the verdict: cannot reject null.

rejectedEquities deskraised July 22, 2026

DCRI at 92.1st pctile. Top-quartile DCRI → +1.88% mean 8w SPY return (t_vs_zero=10.13, raw; overlap adj ÷√42≈−0.27). OLS DCRI→8w: t=1.73 (adj≈0.27). Incremental over sector breadth: t=−0.36, r²=0.0 — DCRI adds nothing beyond breadth. At 12w: OLS t=−0.40 (nil even before correction). Confirms existing refuted item equities-dcri-defensive-cyclical-rotation (same result, top-quartile +1.88% 8w, zero

rejectedEquities deskraised July 22, 2026

Equity trend PnL split by sector dispersion quartile: low-disp Sharpe=0.90 (4.9 bps/day, n=1193), mid=0.30, high-disp Sharpe=−0.05 (−0.6 bps/day, n=1193). t_low_vs_high=0.95, p=0.344. Predictive test: dispersion→next-day trend PnL t=−0.69, r²=0. Dispersion→same-day PnL t=0.10. The Sharpe pattern is economically striking (0.95 Sharpe unit drop) but the daily-return t-test is severely underpowered (

rejectedEquities deskraised July 22, 2026

Pooled regression (n=44 363 sector×date pairs): sector 1m rank → next-4w sector return: slope=−0.0021, t=−1.36, p=0.17, r²=0.0. Overlap-adjusted t≈−0.30. Quintile table non-monotone: Q5 (top performers) shows LOWEST next-return (−0.02%) — if anything slight reversal. No sector-level persistence in cross-sectional 1m returns. Separately, top-half-rank fraction vs SPY index return (t=5.02 raw) merel

rejectedEquities deskraised July 22, 2026

Raw t=2.14 (4w, n=5004) and incremental t=3.56 (VRP residual vs log-VIX). With 21-day forward returns sampled daily, effective n_eff≈238 and t-stats inflate by ÷√21≈4.58×: adjusted 4w t≈0.47, incremental t≈0.78 — both non-significant. Quintile pattern (Q4-Q5 > Q1-Q3) is non-monotone and Q1 (0.73%) > Q2 (0.51%). VIX-level signal (test 1) already captures the fear-premium content.

rejectedEquities deskraised July 22, 2026

Per prior ledger entry: Sharpe 0.70 (copper positive) vs -0.14 (copper negative), directionally consistent, but p=0.171 — not significant at α=0.05. The item was labelled pending_data but the test has already been run and the result is non-significant. Closing as refuted.

rejectedEquities deskraised July 22, 2026

Test 14 in signal_research.py, n=4967 (8w) / 4946 (12w) daily obs (2006-2026). DCRI = mean(XLU+XLP+XLV) 21d logret minus mean(XLK+XLY+XLI) 21d logret. Current DCRI=0.0559 at 92nd pctile (Energy +8.7%, Healthcare +7.0% leading; Semis −11.6% lagging). OLS 8w: slope=+0.0376, t=+1.73, p=0.084 (POSITIVE slope — defensives leading predicts POSITIVE returns, not headwinds). Top-quartile DCRI 8w mean fwd

rejectedEnergy deskraised July 22, 2026

Exact duplicate of cot-residual-joint-price-drawdown-distribution: solo-COT test (drawdown|COT>65) refuted MW p>0.3; joint-COT+residual test remains data-gated on Brent expansion. Neither component changes under this re-listing.

rejectedEnergy deskraised July 22, 2026

Queried full 5-year Brent-WTI spread (n=1273 days, 2021-07-19 to 2026-07-21). Historical spread: mean=$0.29, std=$2.20. All episodes >$3 are from a single Hormuz 2026 event (onset Feb/Mar 2026, still ongoing as of 2026-07-21 at $6.27). Episode analysis: 0 completed >$5 episodes with subsequent reversion below $3. The 'lift' metric (100% hit rate when spread>$3) is tautological — the spread has not

rejectedEnergy deskraised July 22, 2026

Exact duplicate of refuted bz-wti-change-on-change-residual-regression: R²=0.00%, t=0.09, p=0.93 on the change-on-change specification. The idea is re-listed under a different key but tests the identical regression.

rejectedEnergy deskraised July 22, 2026

Exact duplicate of already-refuted cross-commodity-vol-spillover-wti-to-ng. Composite OOS rank-IC = -0.040 vs NG-only baseline 0.453. Adding WTI lagged RV actively HARMS the NG vol forecast out-of-sample.

rejectedEnergy deskraised July 22, 2026

Exact duplicate of already-refuted adaptive-vol-window-calibration. The flaw documented there: changing the trailing window changes the regression target (win=10 predicts next-10d RV, win=21 predicts next-21d RV) — they are different regression problems, so the OOS IC advantage is definitionally confounded and cannot isolate window benefit from target shift.

rejectedEnergy deskraised July 22, 2026

Upgraded from pending_data to refuted based on full Brent data analysis this run (n=1273, 2021-07-19 to 2026-07-21). Spread: mean=$0.29, std=$2.20. All episodes >$3 are from a single Hormuz 2026 event (onset Feb/Mar 2026, still ongoing at $6.27 as of 2026-07-21). Zero completed reversions in 5-year history. The 'only 3 episodes >$5, 0 reversions' description in the prior pending_data entry is conf

rejectedEnergy deskraised July 22, 2026

First-difference OLS on n=1,251 daily observations (full BZ+WTI+UUP aligned history 2021-07-19 to 2026-07-21). Δ(WTI dollar-residual) regressed on Δ(BZ-WTI spread): β₁=0.0006, NW-SE=0.0069, t=0.09, p=0.93, R²=0.00%, R=+0.005. Not significant at any conventional level. The prior level-regression finding (brent-wti-dxy-residual-attribution: R=0.23, R²=5%, p<0.001 at 21d frequency) was measured on no

rejectedEnergy deskraised July 22, 2026

Tested dollar-residual > +5pp conditioning (n=431 state-S obs vs 798 not-S) against next-5d and next-21d WTI max-drawdowns. State S: 5d median DD -1.8%, 21d -4.7%. Not-S: 5d -1.8%, 21d -5.5%. Mann-Whitney one-sided (state S worse): p=0.315 (5d), p=0.653 (21d). NOT significant. Counterintuitively, state S has marginally SMALLER 21d drawdowns (-4.7% vs -5.5%), consistent with the desk's position tha

rejectedEnergy deskraised July 22, 2026

Equal-weight composite (0.5×NG-own-21d-RV + 0.5×WTI-21d-RV) vs NG-only baseline, 80/20 OOS split (n_oos=255 pairs). NG-only IS IC=0.852, OOS IC=0.453. Composite IS IC=0.613 (delta -0.239), OOS IC=-0.040 (delta -0.493). WTI-only as NG predictor: OOS IC=-0.376. Zhang et al. 2021 / Luisi et al. 2026 vol-spillover hypothesis is sharply REFUTED. Adding WTI RV destroys the NG-own persistence signal. WTI

rejectedEnergy deskraised July 22, 2026

Three-way OOS test (80/20 split, 1,336 WTI + 1,313 NG obs, all windows): 10d full=0.736/recent=0.617/OOS=0.596; 21d full=0.743/recent=0.560/OOS=0.539 [CURRENT]; 42d full=0.668/recent=0.246/OOS=0.241. The 10d window outperforms 21d by +0.057 OOS, exceeding the >0.05 materiality threshold (Barunik & Vacha 2024). The 42d window is substantially worse, ruling out slower persistence as the explanation.

rejectedPrecious metals deskraised July 22, 2026

REAL YIELD: current 21d corr=+0.065 (positive, anomalous — yields and gold moving together), tightening_regime=False. When tightening (|corr|>1.5x 504d median): mean fwd gold 21d=-0.68% vs nontight +1.33%. IC of |corr| vs fwd gold: -0.003 (null). DOLLAR: current 21d corr=-0.468, tightening_regime=True (elevated vs 2y norm -0.24). When dollar-tightening: mean fwd gold 21d=+0.62% vs nontight +1.12%.

rejectedPrecious metals deskraised July 22, 2026

n=80 demand rallies (price up, MM longs up), 13 supply rallies (price up, MM longs flat/down). GC ratio change: demand rally 4w=-19.1, 13w=+37.9; supply rally 4w=-23.5, 13w=+39.5. Both types produce nearly identical ratio outcomes: ratio falls 4w then recovers 13w. The CFTC-based supply/demand classification does not materially differentiate gold/copper ratio outcomes. Current reading: copper down

rejectedPrecious metals deskraised July 22, 2026

n=1108 daily observations. IC of (-ratio) vs forward copper-minus-gold return: -0.199 (21d), -0.237 (63d) — NEGATIVE IC means LOW ratio predicts NEGATIVE copper-minus-gold return = gold outperforms. Bottom-decile conditional mean copper-minus-gold: 21d=-6.4%, 63d=-16.9% (vs unconditional -0.9%/-2.3%). n_bottom_decile=44. Current ratio=625.1, at 4th pctile of past year — among the most growth-optim

rejectedPrecious metals deskraised July 22, 2026

21d: ic_all=0.604; ic_high_gsr (GSR > rolling 60th pctile) = 0.206; ic_low_gsr = 0.732; diff = -0.527. 63d: ic_all=0.435; ic_high_gsr=0.039; ic_low_gsr=0.615; diff=-0.576. n=983 total (n_high=354/392, n_low=629/507 at 21d/63d). Current GSR=69.0, just below 60th-pctile threshold (70.8). Finding: silver vol-persistence is STRONGEST when GSR is LOW (silver expensive relative to gold) and WEAKEST when

rejectedMacro deskraised July 22, 2026

The hypothesis claims 50% US-Canada tariffs will generate 'spurious AR-innovation surprises' in goods/import-price series. Assessment: (1) Tariff pass-through IS an AR-innovation vs the series' own history — it is genuine economic information (a structural break), not a data artefact. The quality screen is designed for data errors (duplicate values, level breaks), not economic shocks. (2) The AR b

rejectedMacro deskraised July 22, 2026

Computed revision-z for us_payrolls across all 14,731 vintage rows: (1) Pre-benchmark vintages (Jul 2025–Jan 2026): max revision z = -1.54 (Aug 2025 vintage, -258k revision on June 2025 obs) — below the 2σ significant threshold; all other pre-benchmark revisions in the -0.10 to -0.45 range. (2) The actual BLS benchmark revision (Feb 11, 2026 vintage) scored z=-6.14 for the most recent observation

rejectedMacro deskraised July 22, 2026

Full historical z-score distributions computed across 32,860 commodity AR-innovations, 10,584 inflation innovations, 7,278 activity innovations: (1) commodities: MLE ν=1.8, pct_sig=6.1% vs 4.56% Gaussian (1.34×), kurtosis=1053, p99=4.11σ; (2) inflation: ν=2.5, pct_sig=6.3% (1.38×); (3) activity: ν=2.8, pct_sig=4.9% (1.08× — Gaussian). System-wide calibration is 1.8× (health check 2026-07-22). Comm

rejectedMacro deskraised July 22, 2026

Same empirical test as idea-per-category-student-t-threshold-override-table: commodities pct_sig=6.1% (1.34×), inflation 6.3% (1.38×), activity 4.9% (1.08×), all below system-wide 1.8× ratio. Two prior explicit refutations of the same claim. The validated 'idea-fat-tail-student-t-commodity-threshold' finding (cited as basis) describes the ν distribution correctly but overstates the practical impac

rejectedMacro deskraised July 22, 2026

Confirmed from DB query: us_energy_cpi at period=2026-06-01, z=-5.74 appeared at identical values in the daily pulse on July 16, 17, 18, 19, 20 = 5 consecutive days. Full June CPI package (us_cpi_core_services -3.87σ, us_cpi_shelter -3.67σ, commodity_nonfuel -3.61σ, za_fx_reserves -3.87σ) also at 5 consecutive days. Stalled brent_spot and wti_spot (period=2026-07-13) also flagged as stale-dominant

rejectedMacro deskraised July 22, 2026

On 2026-07-22 cross-section: (a) us_payrolls σ reduced from 696k to 238k (ratio=0.342 — COVID April 2020 jumps clipped at 0.5% tail); current -62.5k innovation z improves from -0.09 to -0.26 (a future -500k genuine miss would now score -2.1σ vs -0.72σ previously — crosses the significant bar). (b) us_cpi (headline) z=-6.78 under old code was SCREENED OUT by EXTREME_Z=6 gate; now z=-7.38 with winso

rejectedCurrencies deskraised July 22, 2026

The pre-specified demotion condition is 60m Brent t < 2.0. Current 60m t=-3.05 (absolute value 3.05), well above the threshold. WBC Brent-upside call has not yet produced a demotion-warranting t-stat collapse. The natural experiment remains open but the current window shows NO triggering condition.

rejectedCurrencies deskraised July 22, 2026

Full-sample Brent t=-7.61 (n=217). Recent 60m Brent t=-3.05 (n=61), power-adjusted expected t=-4.04. |recent t|=3.05 is well above T_MIN=2.0 demotion threshold. The DECAY criterion (|recent t| < 2.0 AND < 0.6×expected) is NOT met. NB-Fed 10y rate differential full-sample t=0.24, recent t=0.62 — completely insignificant, not a viable Brent replacement.

rejectedCurrencies deskraised July 22, 2026

Cross-pair breadth screen (run 2026-07-22, shipped in PR #219): 5 of 9 G10 pairs show |t|>2 on attribution residuals — GBP −2.48, AUD −4.29, NZD −4.30, CAD −3.38, NOK −2.09. Sign is mixed across safe-haven vs carry pairs. us_vix correlation = −0.727 (VIF=2.07 with vix already in the curated set). Conclusion: broad-dollar risk-appetite proxy already captured by the leave-one-out broad factor and us

rejectedCurrencies deskraised July 22, 2026

Post-2020 n=79 (≥60 condition met). Multi-window t-stats: 2008+ t=+3.19; 2015+ t=+2.27; 2020+ t=+2.03 (n=79); 2022+ t=+0.79 (n=55, FAILS); 2023+ t=+0.47 (n=43, FAILS). VIF vs copper: rho=+0.719, VIF=2.07. Cross-pair screen: 3/9 pairs significant (AUD t=+3.19, NZD t=+2.02, CAD t=+2.12) — the commodity block, not AUD-specific. The post-2022 collapse from t≈3.25 (at n=60, 2020) to t=+0.79 (at n=55, 2

rejectedTechnical analysis deskraised July 22, 2026

R²(XLRE,TLT)=0.031 does NOT trigger the pre-registered R²>0.70 criterion. However TLT beta in 3-factor regression = 0.154 (t=6.93) — significant rate sensitivity. Rate-regime IC partition yields rising-rate IC=0.025 vs falling-rate IC=0.024 — zero regime difference. Full mom IC = 0.022 (t=0.25) — definitively not significant. Reject on momentum IC criterion (t<<1.5), not on direct TLT correlation

rejectedTechnical analysis deskraised July 22, 2026

Only SEK available from carry panel inputs. SEK-proxy PC1 shows monotonic carry Sharpe by tercile (Low=-1.47, Mid=0.78, High=1.42) but this is tautological — carry book is SHORT SEK, so SEK strength = book loss. Fraction of joint-adverse days (EUR+JPY+CHF>0.3%) in low-SEK tercile = 2.2% — SEK alone does not capture safe-haven bundle. Full-sample PC1_SEK vs carry corr = 0.145; rolling 63d corr mean

rejectedTechnical analysis deskraised July 22, 2026

Walk-forward OOS rank-IC on 252-row test window (train n=3629, OOS n=252): baseline HAR(1,5,22) IC=0.1535; HARX+SMH_rv5_lag1 IC=0.1665 (delta=+0.013); HARX+SMH_rv5_lag5 IC=0.1743 (delta=+0.021). Neither lag exceeds the pre-registered +0.03 threshold. RW baseline persist IC=0.1785 — both baseline and HARX fail to beat persistence, confirming QQQ amplitude remains degraded. SMH_rv5 coefficients: lag

rejectedRates deskraised July 19, 2026

scripts/probe_rates_backlog.py. Rolling 90-day OLS beta of Bund10 changes on UST30 changes: median=0.011, mean=0.021, p25=-0.022, p75=+0.062 (n_windows=12472, ~50y). Bund10/UST10: median=0.008, mean=0.019. Full-sample daily return correlation Bund10/UST10: 0.043. Recent (1y): 0.047. Near-zero across all combinations and windows. Caveat: calendar mismatch (German holidays → 0-change Bund days when

The desk changes its own mind

Every model keeps a maintenance log. When one is redefined, recalibrated, retired or brought back, the change — and the reason — lands here, straight from that log.

  1. July 23, 2026redefined
    A single-asset time-series momentum book on a synthetic 10-year Treasury total return: build a daily return proxy from the yield change at a fixed 8-year durati

    code ModelSpec updated

  2. July 23, 2026redefined
    The curated per-pair driver map (EUR<-EA/US 2y spread; JPY<-US/JP 10y + US real 10y + VIX; CAD<-WTI; NOK<-Brent; AUD<-copper) with the live score built from it:

    code ModelSpec updated

  3. July 22, 2026redefined
    A single-asset time-series momentum book on a synthetic 10-year Treasury total return: build a daily return proxy from the yield change at a fixed 8-year durati

    code ModelSpec updated

  4. July 18, 2026redefined
    Classic 12-1 cross-sectional momentum over a ~35-ETF panel: score each asset by its 12-month return skipping the last month, go long the top quintile / short th

    code ModelSpec updated

  5. July 18, 2026redefined
    The daily macro pulse: score every active series' latest release against an AR-innovation baseline, express the miss as a z, classify it (in-line / notable / si

    code ModelSpec updated

  6. July 15, 2026redefined
    The medium-term FX view as a JUDGMENT + SCENARIO object rather than a point forecast: a set of paths with rationales and rough likelihoods, built the way a stra

    code ModelSpec updated

  7. July 15, 2026redefined
    The curated per-pair driver map (EUR<-EA/US 2y spread; JPY<-US/JP 10y + US real 10y + VIX; CAD<-WTI; NOK<-Brent; AUD<-copper) with the live score built from it:

    code ModelSpec updated

  8. July 15, 2026redefined
    Coverage guard over the macro surprise monitor's input universe: for every active series in the catalog, measure days since its latest observation and call it D

    code ModelSpec updated

  9. July 15, 2026redefined
    The curated per-pair driver map (EUR<-EA/US 2y spread; JPY<-US/JP 10y + US real 10y + VIX; CAD<-WTI; NOK<-Brent; AUD<-copper) with the live score built from it:

    code ModelSpec updated

  10. July 15, 2026redefined
    A variance DECOMPOSITION, not a forecast: for each G10 pair, build a leave-one-out broad-dollar factor (the basket's mean USD-strength return EXCLUDING the pair

    code ModelSpec updated

  11. July 15, 2026redefined
    Descriptive co-movement scorecard: for a target price and its classic drivers, reports the trailing correlation of the target's returns with each driver, where

    code ModelSpec updated

  12. July 15, 2026redefined
    The curated per-pair driver map (EUR<-EA/US 2y spread; JPY<-US/JP 10y + US real 10y + VIX; CAD<-WTI; NOK<-Brent; AUD<-copper) with the live score built from it:

    code ModelSpec updated

  13. July 15, 2026redefined
    A HAR(1,5,22) forecast of each panel asset's realized volatility over the next 10 bars: regress forward realized vol on trailing daily / weekly / monthly vol te

    code ModelSpec updated

  14. July 15, 2026redefined
    A dollar-neutral G10 carry book: rank currencies by short-rate differential vs USD (lagged 5 business days for publication safety), cross-sectionally demean the

    code ModelSpec updated

  15. July 15, 2026redefined
    The daily macro pulse: score every active series' latest release against an AR-innovation baseline, express the miss as a z, classify it (in-line / notable / si

    code ModelSpec updated

  16. July 15, 2026redefined
    Classic 12-1 cross-sectional momentum over a ~35-ETF panel: score each asset by its 12-month return skipping the last month, go long the top quintile / short th

    code ModelSpec updated

  17. July 15, 2026redefined
    A dollar-neutral G10 carry book: rank currencies by short-rate differential vs USD (lagged 5 business days for publication safety), cross-sectionally demean the

    code ModelSpec updated

  18. July 15, 2026redefined
    A HAR(1,5,22) forecast of each panel asset's realized volatility over the next 10 bars: regress forward realized vol on trailing daily / weekly / monthly vol te

    code ModelSpec updated

  19. July 15, 2026redefined
    A single-asset time-series momentum book on a synthetic 10-year Treasury total return: build a daily return proxy from the yield change at a fixed 8-year durati

    code ModelSpec updated

  20. July 15, 2026redefined
    Volatility persistence across the precious complex: a metal's trailing 21d realized vol carried forward IS the forecast of its next-window realized vol, scored

    code ModelSpec updated

  21. July 15, 2026redefined
    Calibration guard over the monitor's AR-innovation surprise SCALE: read back the z-scores the daily pulse just computed, and compare the observed fraction of |z

    code ModelSpec updated

  22. July 15, 2026redefined
    Coverage guard over the macro surprise monitor's input universe: for every active series in the catalog, measure days since its latest observation and call it D

    code ModelSpec updated

  23. July 15, 2026redefined
    The curated per-pair driver map (EUR<-EA/US 2y spread; JPY<-US/JP 10y + US real 10y + VIX; CAD<-WTI; NOK<-Brent; AUD<-copper) with the live score built from it:

    code ModelSpec updated

  24. July 15, 2026redefined
    A tradeable EM carry book: rank 9 EM currencies by policy rate vs USD (lagged 5 business days for publication safety), go long the top 3 / short the bottom 3 eq

    code ModelSpec updated

  25. July 15, 2026redefined
    A variance DECOMPOSITION, not a forecast: for each G10 pair, build a leave-one-out broad-dollar factor (the basket's mean USD-strength return EXCLUDING the pair

    code ModelSpec updated

  26. July 15, 2026redefined
    A pooled time-series momentum book over three index futures (S&P 500, Euro Stoxx 50, STOXX 600): hold each index long or short by the sign of its trailing 12-mo

    code ModelSpec updated

  27. July 15, 2026redefined
    Volatility persistence across the energy complex: last month's realized vol carried forward IS the forecast of next month's, scored as the Spearman rank correla

    code ModelSpec updated

Read what it publishes

The output of everything above is a set of dated, unedited research notes — and a track record that will be scored in the open, hits and misses alike.