Macro desk
Watches the global economic calendar and flags the data releases that came in meaningfully above or below what recent history predicted. Publishes a daily note on which surprises matter and why.
This page is the machine looking at itself: the data it watches, the desks and scouts that run every day, the adversarial review every idea has to survive, the ideas it killed, and the changes it made to its own models. Everything below renders live from the system’s own records — nothing is written by hand.
Every part of the system, as registered in its own database — titles, descriptions and activity come straight from the live registry. When a new agent joins the desk, it appears here the morning after, with no edits to this page.
Seven specialist desks, each covering one corner of the markets. Every desk publishes a dated note every day — no skipping quiet days, no editing after the fact.
Watches the global economic calendar and flags the data releases that came in meaningfully above or below what recent history predicted. Publishes a daily note on which surprises matter and why.
Covers government bond markets: the US yield curve, its slope, and real yields. Publishes a daily note on what moved yields and what the curve implies about growth and policy expectations.
Covers the major world currencies. Publishes a daily note on what drove exchange-rate moves and a weekly deep dive into the frameworks behind them.
Covers the major stock indices in the United States and Europe, along with the market's own gauge of expected turbulence. Publishes a daily note on what moved and how the two regions compare.
Covers crude oil and natural gas: prices, the shape of the futures curve, and how traders are positioned. Publishes a daily note on what is driving the energy complex.
Covers gold, silver, platinum and palladium, including the gold-to-silver ratio and how speculators are positioned. Publishes a daily note on what is moving the metals.
Reads price behaviour across every market the house covers: trends, momentum, and turning points. Publishes a daily cross-market technical read the other desks can lean on.
Standing watchers that brief every research session: what just surprised, what quietly broke, what new data arrived, and what the rest of the world published.
Notices when newly collected data series arrive and puts them in front of the right desk, so fresh data never sits unused. Also confirms when a requested series has started flowing, closing the loop on data requests.
Flags which economic releases and market moves were genuinely surprising relative to recent history, so every desk weighs the week's big events rather than only its own corner.
Watches long-standing relationships between markets for signs they are breaking down or reconnecting, and raises each shift as a question for the desk to investigate.
Tracks upcoming central-bank decisions, what markets expect them to deliver, and how recent decisions compared with those expectations.
Filters the outside research feed for geopolitical developments worth weighing — a curated short list, not a headline firehose.
Searches the live web for what is too fresh to be in the research feed yet: breaking events, upcoming calendar items, and named methods worth replicating on our own data.
Re-reads the desk's own published research alongside outside commentary, and flags gaps, contradictions and stale claims worth a second look.
Reminds each desk which in-house estimates already exist and how accurate they have proven, so analysts reach for the best available number instead of a stale official one.
Keeps a list of collected data the desk has never analysed and steadily surfaces the oldest gaps until each one gets a verdict.
Reads the week's outside research and boils it down to a ranked list of specific, testable ideas for the desks.
No finding counts until it survives review. A skeptic tries to knock it down, a statistics reviewer audits how it was produced, and a closer forces open questions to a verdict.
The panel's working analyst: forms a hypothesis, tests it against the data, and writes up what held and what did not. One runs per desk, with every claim recorded for review.
An adversarial reviewer who tries to knock each new finding down — hunting for lucky streaks, cherry-picked windows and simpler explanations — before it can influence anything.
A statistics reviewer who audits how each result was produced: sample sizes, test choices, and whether the evidence actually supports the stated conclusion.
Works through the desks' open questions and settles them one by one — confirmed, refuted, or genuinely waiting on data — so the question pile shrinks instead of growing. A refuted answer counts as a win.
Periodically reviews a desk's own track record — what paid off, what was noise — and proposes concrete changes to how the desk works.
The standing forecasting and interpretation models — each one scored against a simple benchmark on live data, and benched when it stops earning its place.
Estimates where US interest-rate policy is headed from the economic data, and compares that estimate with the path markets have priced in.
Reads each central-bank statement and scores how hawkish or dovish it sounds, then compares that tone with what markets expect — flagging when words and prices disagree.
Tracks what people are searching for online — job worries, mortgage costs, layoffs — and turns shifts in public attention into an early read on the economy that every desk can use.
Measures how crowded the bets are across markets by combining futures positioning with price momentum — a contrarian warning light when everyone leans the same way.
Combines the monthly activity data as it arrives into a running estimate of the current quarter's US economic growth, well before the official number is published.
Reads market-based stress gauges — borrowing costs, credit spreads, volatility — for an early signal on jobs, output and growth.
Tracks the pipeline pressures behind US goods prices — the dollar, producer prices, global supply-chain strain — to anticipate where goods inflation is heading.
Compares household and market inflation expectations with a model-based view, and uses the gap between them to anticipate inflation surprises.
Uses moves in oil, gas and copper futures to anticipate their knock-on effects on consumer energy prices and industrial activity.
Blends the house's independent growth models into one estimate, weighting each by its proven accuracy.
Estimates where a textbook policy rule says euro-area interest rates should be, given inflation and unemployment, and compares that with what markets have priced in — a stance read, not a forecast, shown alongside the US equivalent.
The machinery that keeps the desk honest: synthesis and delivery of the research, scoring against professional research houses, and the checks that catch the system drifting.
Audits the tone reader's interpretations against independent commentary and adds any angles it missed to the reading checklist — a model that improves its own rubric.
The morning-meeting synthesis: reads all seven desk notes, surfaces where they agree and disagree, and states the house's single top-of-book read for the day.
One email each morning carrying the house view and every desk's daily note — the single consolidated read of the day.
A weekly consolidated view of every in-house forecast and how each has been scoring against reality.
A weekly newsletter pulling together the week's forecasts, refreshed research papers and model health checks.
Scores the desks' published notes against what professional research houses published: did we explain the moves they explained, and where did we see something they missed?
A monthly review of every model on the books that forces a decision on any that have drifted: recalibrate, retire, or keep.
Every day, picks the highest-priority open questions and pushes them through to an answer, so the research backlog keeps moving between full review cycles.
Retries proposed changes that had already passed review but were held up by a temporary failure, and reports how many remain waiting.
Checks that the research cycle itself is working — sessions completed, approved changes landed — and raises an alarm, or fixes the known causes, when the loop silently stalls.
Automatically benches any forecasting model that keeps losing to a simple benchmark on live data — with a clear bar, a record of why, and an easy way back.
The shared logbook of every hypothesis, finding and verdict. Questions keep their history, negative results are preserved, and nothing is asked twice.
Checks every new research question at the door: duplicates are merged, questions that cannot say what would prove them wrong are sent back for rework, and the rest are priority-scored.
Makes sure a confirmed finding actually lands somewhere — a model update, a code change, or a wiring proposal — instead of being celebrated and forgotten.
Takes the work the research agents commit, runs the full battery of checks, and merges it only when everything passes and the review panel found no fatal flaw; everything else is held for human review.
A desk's running book of the stories believed to be moving its market, verified against how prices actually behave rather than how often the story is repeated.
Finds comparable historical episodes — say, every inflation surprise of this size — and reports what actually happened next, always shown alongside the everyday base rate, so the desks can cite measured history instead of impressions.
The shared machinery every forecasting model runs on: one honest scoring standard against a simple benchmark, one storage format, and one publishing path.
A weekly check on the shared analytics every desk relies on — the market-relationship radar and regime gauges — so a break in common machinery is caught centrally, not desk by desk.
Most research ideas do not survive testing — and a desk you can trust is one that shows you the bodies. These are the most recent ideas the desks killed, with the evidence that killed them, quoted from the internal record.
Well-powered null on 20y (n=1021, 2006+): S&P 500 leveraged-fund COT index rank-IC vs 4-week fwd SPY return +0.04 (ns, wrong sign for contrarian), quarter +0.10 (ns); at extremes crowded-LONG -> HIGHER fwd returns (spread -0.26%/-1.21%), opposite of the contrarian thesis. Tested via the new read_history loader (market-research#285). COT stays crowding CONTEXT, not a forward-return model.
Prior session (2026-07-22, fi_cot_tenor_2026_07_22.py): Spearman IC of belly-avg COT index (2y+5y average) minus 30y COT index vs forward 4w slope change = 0.021 (p=0.505). OLS slope = null. Conditional mean at extreme belly divergence not distinguishable from unconditional. Matches 'already refuted' entry.
Prior session (2026-07-22, PR #224): TIPS 30y-10y real yield slope as carry signal for 30y bonds, 2010-2026 (≈16 years). NW-t=0.424 — far below 1.65 bar. This matches the 'already refuted' entry. Consistent with general carry refutation findings.
1990-2010 (20.9y): NW-t daily=1.178. 1995-2010 (15.9y): NW-t daily=0.469. Pre-2022 ex-inversion (44.7y): NW-t monthly=0.916. All below 1.65 bar. The 5y OOS (2021-2026) NW-t daily=1.973 is above bar but includes the entire 2022-2023 inversion episode — confirming the inversion era is necessary for the signal to appear significant.
Prior session (2026-07-22, PR #224, scripts/fi_cot_tenor_2026_07_22.py): IC (Spearman) between COT index and forward 4w/8w yield change is near zero for all four Treasury tenors (2y/5y/10y/30y). Conditional means at extremes (≥85 or ≤15) not significantly different from unconditional. All p-values > 0.20. This matches the 'already tested & rejected' list entry.
Swept 2m–24m (42d–504d), full FRED daily 10y yield history (1962-2026, n≥15,616d per lookback). Sharpe at 252d=0.567; global peak at 483d (23m)=0.668; drop from peak to 252d=15.1% (below 20% cherry-pick threshold). 9-15m range (189-315d) shows 33.9% Sharpe variation (max 0.567 at 252d, min 0.375 at 315d) — above 20% stability threshold, so no stable plateau. Signal exists broadly: Sharpe 0.375–0.6
1990-2010 holdout (n=5255d, 20.9y): 2s10s_z carry Sharpe=0.249, NW-t daily=1.178, NW-t monthly=1.329 — both below 1.65 bar. 1995-2010 (15.9y): Sharpe=0.113, NW-t daily=0.469. Broader pre-2022 sample (44.7y ex-inversion): NW-t monthly=0.916. Inversion era 2022-2023: Sharpe 1.263 (2y window) — this episode carries the full-sample result. scripts/fi_executor_2026_07_22.py (2026-07-22).
The validation_ref in spec.py previously stated 'The signal bar has never been cleared here. Establishing or refuting it is a live research lead.' This was superseded by the bond-momentum-in-repo-validated finding (NW-t=4.077, Sharpe=0.568, 1962-2026). Updated to reflect: (1) validated NW-t=4.077 status; (2) 2y tenor dominance finding (Sharpe=2.09, NW-t=3.79); (3) lookback sweep result (no stable
n=1016 weekly obs (2006-2024) per tenor. 2Y: IC(COT vs dy4w)=-0.003 (p=0.917), mean dy4w at extreme-long (COT>=85, n=207) = -0.014% vs unconditional -0.002% — slightly NEGATIVE (yields fell when 2Y was at extreme-long, opposite of long-squeeze narrative). 5Y: IC=-0.019 (p=0.538), null. 10Y: IC=+0.047 (p=0.136), mean dy4w extreme-long=+0.030% vs 0.000% unconditional, NW-t=1.368 — closest to signal
n=1016 weekly obs (2006-2024). Spearman IC of mean(COT_2y, COT_5y) - COT_30y vs forward 4-week 5s30s slope change: 0.021 (p=0.505). 8-week horizon: IC=0.032 (p=0.308). OLS r=0.032 (p=0.303) at 4w, r=0.040 (p=0.202) at 8w. All null. Current divergence = 71.8 (93.5th pctile — historically extreme belly-long vs long-end). Conditional mean slope change at extreme belly readings (>85th pctile, n=153):
Full-sample (2010-08-17 to 2026-06-17, n=192 monthly): NW-t=0.424, Sharpe=0.427. Well below the 1.65 bar. Current real slope = 56bp, z-score = -1.75 (currently compressed vs rolling average). The signal is null — the real yield slope does not carry as a timing signal for long/short 30y vs 10y TIPS returns.
32 independent regime-exit events identified (SPY 2006-2026). 20d after exit: mean=+0.41%, t=0.72, p=0.48 (clearly null). 40d after exit: mean=+1.04%, median=+1.75%, t=1.64, p=0.11 (borderline but not significant at α=0.05). Historical episode stats: 33 negative-corr episodes, median=10d, mean=33d, q90=79d. Current episode: 83 days at 90.9th pctile. With n=32 exit events, the test is underpowered
4w: top-quintile XLE mean_fwd_SPY=+0.16%, t_vs_zero=1.06 (raw), adj≈0.23. OLS t=−4.65 raw → adj≈−1.01. Bottom quintile: +0.75%; mid: +0.94% — XLE leaders actually have lowest forward returns. 8w: top quintile +1.23% vs bottom +1.10% (essentially identical). Current XLE relative at 93rd pctile (+8.22% vs SPY 1m). No significant headwind signal at either horizon.
OLS slope=−0.111 (t=−7.85, r²=1.2%, n=4 980 obs, 2006–2026 EWJ/SPY daily). Overlap-adjusted t≈−1.7 (21-day forward returns sampled daily, ÷√21). Quintile table perfectly monotone: Q1=−0.09%, Q2=−0.35%, Q3=−0.62%, Q4=−0.65%, Q5=−0.82% fwd relative return. Japan outperformers >3% (n=651): mean_fwd_rel=−0.81%, raw t=−5.4, adj t≈−1.2. Japan underperformers >3% (n=1067): −0.11%, t=−1.13 (not significan
Genuine broadening (RSP>SPY AND SPY>0): mean_fwd_4w=0.79%, t=8.71 (raw), adj≈1.90. Defensive rotation (RSP>SPY AND SPY≤0): 0.21%, t=0.90 (raw), adj≈0.20. t_A_vs_B=2.84 raw → adj≈0.62 after ÷√21 overlap correction — not significant. The commentary.py _breadth_words() function ALREADY implements this distinction (SPY direction conditioning added previously). Current reading: Bucket A genuine broaden
2026 recheck using index-futures data (n=818 days): negative-corr regime Sharpe=1.35 vs positive-corr Sharpe=−0.10. Sharpe diff=1.46. Daily-label permutation p=0.204 (5000 perms). Regime diff t=1.21, p=0.225. Consistent with prior results: block-bootstrap p=0.297, label-shuffle p=0.187. Adding ~4 years of futures data has not changed the verdict: cannot reject null.
DCRI at 92.1st pctile. Top-quartile DCRI → +1.88% mean 8w SPY return (t_vs_zero=10.13, raw; overlap adj ÷√42≈−0.27). OLS DCRI→8w: t=1.73 (adj≈0.27). Incremental over sector breadth: t=−0.36, r²=0.0 — DCRI adds nothing beyond breadth. At 12w: OLS t=−0.40 (nil even before correction). Confirms existing refuted item equities-dcri-defensive-cyclical-rotation (same result, top-quartile +1.88% 8w, zero
Equity trend PnL split by sector dispersion quartile: low-disp Sharpe=0.90 (4.9 bps/day, n=1193), mid=0.30, high-disp Sharpe=−0.05 (−0.6 bps/day, n=1193). t_low_vs_high=0.95, p=0.344. Predictive test: dispersion→next-day trend PnL t=−0.69, r²=0. Dispersion→same-day PnL t=0.10. The Sharpe pattern is economically striking (0.95 Sharpe unit drop) but the daily-return t-test is severely underpowered (
Pooled regression (n=44 363 sector×date pairs): sector 1m rank → next-4w sector return: slope=−0.0021, t=−1.36, p=0.17, r²=0.0. Overlap-adjusted t≈−0.30. Quintile table non-monotone: Q5 (top performers) shows LOWEST next-return (−0.02%) — if anything slight reversal. No sector-level persistence in cross-sectional 1m returns. Separately, top-half-rank fraction vs SPY index return (t=5.02 raw) merel
Raw t=2.14 (4w, n=5004) and incremental t=3.56 (VRP residual vs log-VIX). With 21-day forward returns sampled daily, effective n_eff≈238 and t-stats inflate by ÷√21≈4.58×: adjusted 4w t≈0.47, incremental t≈0.78 — both non-significant. Quintile pattern (Q4-Q5 > Q1-Q3) is non-monotone and Q1 (0.73%) > Q2 (0.51%). VIX-level signal (test 1) already captures the fear-premium content.
Per prior ledger entry: Sharpe 0.70 (copper positive) vs -0.14 (copper negative), directionally consistent, but p=0.171 — not significant at α=0.05. The item was labelled pending_data but the test has already been run and the result is non-significant. Closing as refuted.
Test 14 in signal_research.py, n=4967 (8w) / 4946 (12w) daily obs (2006-2026). DCRI = mean(XLU+XLP+XLV) 21d logret minus mean(XLK+XLY+XLI) 21d logret. Current DCRI=0.0559 at 92nd pctile (Energy +8.7%, Healthcare +7.0% leading; Semis −11.6% lagging). OLS 8w: slope=+0.0376, t=+1.73, p=0.084 (POSITIVE slope — defensives leading predicts POSITIVE returns, not headwinds). Top-quartile DCRI 8w mean fwd
Exact duplicate of cot-residual-joint-price-drawdown-distribution: solo-COT test (drawdown|COT>65) refuted MW p>0.3; joint-COT+residual test remains data-gated on Brent expansion. Neither component changes under this re-listing.
Queried full 5-year Brent-WTI spread (n=1273 days, 2021-07-19 to 2026-07-21). Historical spread: mean=$0.29, std=$2.20. All episodes >$3 are from a single Hormuz 2026 event (onset Feb/Mar 2026, still ongoing as of 2026-07-21 at $6.27). Episode analysis: 0 completed >$5 episodes with subsequent reversion below $3. The 'lift' metric (100% hit rate when spread>$3) is tautological — the spread has not
Exact duplicate of refuted bz-wti-change-on-change-residual-regression: R²=0.00%, t=0.09, p=0.93 on the change-on-change specification. The idea is re-listed under a different key but tests the identical regression.
Exact duplicate of already-refuted cross-commodity-vol-spillover-wti-to-ng. Composite OOS rank-IC = -0.040 vs NG-only baseline 0.453. Adding WTI lagged RV actively HARMS the NG vol forecast out-of-sample.
Exact duplicate of already-refuted adaptive-vol-window-calibration. The flaw documented there: changing the trailing window changes the regression target (win=10 predicts next-10d RV, win=21 predicts next-21d RV) — they are different regression problems, so the OOS IC advantage is definitionally confounded and cannot isolate window benefit from target shift.
Upgraded from pending_data to refuted based on full Brent data analysis this run (n=1273, 2021-07-19 to 2026-07-21). Spread: mean=$0.29, std=$2.20. All episodes >$3 are from a single Hormuz 2026 event (onset Feb/Mar 2026, still ongoing at $6.27 as of 2026-07-21). Zero completed reversions in 5-year history. The 'only 3 episodes >$5, 0 reversions' description in the prior pending_data entry is conf
First-difference OLS on n=1,251 daily observations (full BZ+WTI+UUP aligned history 2021-07-19 to 2026-07-21). Δ(WTI dollar-residual) regressed on Δ(BZ-WTI spread): β₁=0.0006, NW-SE=0.0069, t=0.09, p=0.93, R²=0.00%, R=+0.005. Not significant at any conventional level. The prior level-regression finding (brent-wti-dxy-residual-attribution: R=0.23, R²=5%, p<0.001 at 21d frequency) was measured on no
Tested dollar-residual > +5pp conditioning (n=431 state-S obs vs 798 not-S) against next-5d and next-21d WTI max-drawdowns. State S: 5d median DD -1.8%, 21d -4.7%. Not-S: 5d -1.8%, 21d -5.5%. Mann-Whitney one-sided (state S worse): p=0.315 (5d), p=0.653 (21d). NOT significant. Counterintuitively, state S has marginally SMALLER 21d drawdowns (-4.7% vs -5.5%), consistent with the desk's position tha
Equal-weight composite (0.5×NG-own-21d-RV + 0.5×WTI-21d-RV) vs NG-only baseline, 80/20 OOS split (n_oos=255 pairs). NG-only IS IC=0.852, OOS IC=0.453. Composite IS IC=0.613 (delta -0.239), OOS IC=-0.040 (delta -0.493). WTI-only as NG predictor: OOS IC=-0.376. Zhang et al. 2021 / Luisi et al. 2026 vol-spillover hypothesis is sharply REFUTED. Adding WTI RV destroys the NG-own persistence signal. WTI
Three-way OOS test (80/20 split, 1,336 WTI + 1,313 NG obs, all windows): 10d full=0.736/recent=0.617/OOS=0.596; 21d full=0.743/recent=0.560/OOS=0.539 [CURRENT]; 42d full=0.668/recent=0.246/OOS=0.241. The 10d window outperforms 21d by +0.057 OOS, exceeding the >0.05 materiality threshold (Barunik & Vacha 2024). The 42d window is substantially worse, ruling out slower persistence as the explanation.
REAL YIELD: current 21d corr=+0.065 (positive, anomalous — yields and gold moving together), tightening_regime=False. When tightening (|corr|>1.5x 504d median): mean fwd gold 21d=-0.68% vs nontight +1.33%. IC of |corr| vs fwd gold: -0.003 (null). DOLLAR: current 21d corr=-0.468, tightening_regime=True (elevated vs 2y norm -0.24). When dollar-tightening: mean fwd gold 21d=+0.62% vs nontight +1.12%.
n=80 demand rallies (price up, MM longs up), 13 supply rallies (price up, MM longs flat/down). GC ratio change: demand rally 4w=-19.1, 13w=+37.9; supply rally 4w=-23.5, 13w=+39.5. Both types produce nearly identical ratio outcomes: ratio falls 4w then recovers 13w. The CFTC-based supply/demand classification does not materially differentiate gold/copper ratio outcomes. Current reading: copper down
n=1108 daily observations. IC of (-ratio) vs forward copper-minus-gold return: -0.199 (21d), -0.237 (63d) — NEGATIVE IC means LOW ratio predicts NEGATIVE copper-minus-gold return = gold outperforms. Bottom-decile conditional mean copper-minus-gold: 21d=-6.4%, 63d=-16.9% (vs unconditional -0.9%/-2.3%). n_bottom_decile=44. Current ratio=625.1, at 4th pctile of past year — among the most growth-optim
21d: ic_all=0.604; ic_high_gsr (GSR > rolling 60th pctile) = 0.206; ic_low_gsr = 0.732; diff = -0.527. 63d: ic_all=0.435; ic_high_gsr=0.039; ic_low_gsr=0.615; diff=-0.576. n=983 total (n_high=354/392, n_low=629/507 at 21d/63d). Current GSR=69.0, just below 60th-pctile threshold (70.8). Finding: silver vol-persistence is STRONGEST when GSR is LOW (silver expensive relative to gold) and WEAKEST when
The hypothesis claims 50% US-Canada tariffs will generate 'spurious AR-innovation surprises' in goods/import-price series. Assessment: (1) Tariff pass-through IS an AR-innovation vs the series' own history — it is genuine economic information (a structural break), not a data artefact. The quality screen is designed for data errors (duplicate values, level breaks), not economic shocks. (2) The AR b
Computed revision-z for us_payrolls across all 14,731 vintage rows: (1) Pre-benchmark vintages (Jul 2025–Jan 2026): max revision z = -1.54 (Aug 2025 vintage, -258k revision on June 2025 obs) — below the 2σ significant threshold; all other pre-benchmark revisions in the -0.10 to -0.45 range. (2) The actual BLS benchmark revision (Feb 11, 2026 vintage) scored z=-6.14 for the most recent observation
Full historical z-score distributions computed across 32,860 commodity AR-innovations, 10,584 inflation innovations, 7,278 activity innovations: (1) commodities: MLE ν=1.8, pct_sig=6.1% vs 4.56% Gaussian (1.34×), kurtosis=1053, p99=4.11σ; (2) inflation: ν=2.5, pct_sig=6.3% (1.38×); (3) activity: ν=2.8, pct_sig=4.9% (1.08× — Gaussian). System-wide calibration is 1.8× (health check 2026-07-22). Comm
Same empirical test as idea-per-category-student-t-threshold-override-table: commodities pct_sig=6.1% (1.34×), inflation 6.3% (1.38×), activity 4.9% (1.08×), all below system-wide 1.8× ratio. Two prior explicit refutations of the same claim. The validated 'idea-fat-tail-student-t-commodity-threshold' finding (cited as basis) describes the ν distribution correctly but overstates the practical impac
Confirmed from DB query: us_energy_cpi at period=2026-06-01, z=-5.74 appeared at identical values in the daily pulse on July 16, 17, 18, 19, 20 = 5 consecutive days. Full June CPI package (us_cpi_core_services -3.87σ, us_cpi_shelter -3.67σ, commodity_nonfuel -3.61σ, za_fx_reserves -3.87σ) also at 5 consecutive days. Stalled brent_spot and wti_spot (period=2026-07-13) also flagged as stale-dominant
On 2026-07-22 cross-section: (a) us_payrolls σ reduced from 696k to 238k (ratio=0.342 — COVID April 2020 jumps clipped at 0.5% tail); current -62.5k innovation z improves from -0.09 to -0.26 (a future -500k genuine miss would now score -2.1σ vs -0.72σ previously — crosses the significant bar). (b) us_cpi (headline) z=-6.78 under old code was SCREENED OUT by EXTREME_Z=6 gate; now z=-7.38 with winso
The pre-specified demotion condition is 60m Brent t < 2.0. Current 60m t=-3.05 (absolute value 3.05), well above the threshold. WBC Brent-upside call has not yet produced a demotion-warranting t-stat collapse. The natural experiment remains open but the current window shows NO triggering condition.
Full-sample Brent t=-7.61 (n=217). Recent 60m Brent t=-3.05 (n=61), power-adjusted expected t=-4.04. |recent t|=3.05 is well above T_MIN=2.0 demotion threshold. The DECAY criterion (|recent t| < 2.0 AND < 0.6×expected) is NOT met. NB-Fed 10y rate differential full-sample t=0.24, recent t=0.62 — completely insignificant, not a viable Brent replacement.
Cross-pair breadth screen (run 2026-07-22, shipped in PR #219): 5 of 9 G10 pairs show |t|>2 on attribution residuals — GBP −2.48, AUD −4.29, NZD −4.30, CAD −3.38, NOK −2.09. Sign is mixed across safe-haven vs carry pairs. us_vix correlation = −0.727 (VIF=2.07 with vix already in the curated set). Conclusion: broad-dollar risk-appetite proxy already captured by the leave-one-out broad factor and us
Post-2020 n=79 (≥60 condition met). Multi-window t-stats: 2008+ t=+3.19; 2015+ t=+2.27; 2020+ t=+2.03 (n=79); 2022+ t=+0.79 (n=55, FAILS); 2023+ t=+0.47 (n=43, FAILS). VIF vs copper: rho=+0.719, VIF=2.07. Cross-pair screen: 3/9 pairs significant (AUD t=+3.19, NZD t=+2.02, CAD t=+2.12) — the commodity block, not AUD-specific. The post-2022 collapse from t≈3.25 (at n=60, 2020) to t=+0.79 (at n=55, 2
R²(XLRE,TLT)=0.031 does NOT trigger the pre-registered R²>0.70 criterion. However TLT beta in 3-factor regression = 0.154 (t=6.93) — significant rate sensitivity. Rate-regime IC partition yields rising-rate IC=0.025 vs falling-rate IC=0.024 — zero regime difference. Full mom IC = 0.022 (t=0.25) — definitively not significant. Reject on momentum IC criterion (t<<1.5), not on direct TLT correlation
Only SEK available from carry panel inputs. SEK-proxy PC1 shows monotonic carry Sharpe by tercile (Low=-1.47, Mid=0.78, High=1.42) but this is tautological — carry book is SHORT SEK, so SEK strength = book loss. Fraction of joint-adverse days (EUR+JPY+CHF>0.3%) in low-SEK tercile = 2.2% — SEK alone does not capture safe-haven bundle. Full-sample PC1_SEK vs carry corr = 0.145; rolling 63d corr mean
Walk-forward OOS rank-IC on 252-row test window (train n=3629, OOS n=252): baseline HAR(1,5,22) IC=0.1535; HARX+SMH_rv5_lag1 IC=0.1665 (delta=+0.013); HARX+SMH_rv5_lag5 IC=0.1743 (delta=+0.021). Neither lag exceeds the pre-registered +0.03 threshold. RW baseline persist IC=0.1785 — both baseline and HARX fail to beat persistence, confirming QQQ amplitude remains degraded. SMH_rv5 coefficients: lag
scripts/probe_rates_backlog.py. Rolling 90-day OLS beta of Bund10 changes on UST30 changes: median=0.011, mean=0.021, p25=-0.022, p75=+0.062 (n_windows=12472, ~50y). Bund10/UST10: median=0.008, mean=0.019. Full-sample daily return correlation Bund10/UST10: 0.043. Recent (1y): 0.047. Near-zero across all combinations and windows. Caveat: calendar mismatch (German holidays → 0-change Bund days when
Every model keeps a maintenance log. When one is redefined, recalibrated, retired or brought back, the change — and the reason — lands here, straight from that log.
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
code ModelSpec updated
The output of everything above is a set of dated, unedited research notes — and a track record that will be scored in the open, hits and misses alike.