Tests for robustness

Evidence away from the path that built the model.

A full-period result can hide an uncomfortable fact: the model may be leaning on the same historical episodes that shaped its design. RAD-AA was developed while evaluating a walk-forward U.S. backtest from 2002 through 2026 because several regimes are rare. These tests ask whether the five-regime framework still behaves coherently in settings that were not the original design target.

Bottom line

The evidence is constructive, not conclusive. Germany and the UK show that the architecture can be replayed in other developed markets using local calendars and local-currency returns. The U.S. rolling audit gives stronger support to drawdown control than to raw return dominance.

§01Germany replication

A local-currency replay on the German market.

The Germany test uses EUR-listed substitutes, DAX returns in euros, euro-area macro/stress proxies, and the XETR trading calendar. It is a robustness test of process portability, not a finished German product.

Why and how

The Germany replication asks whether the same five-regime classifier and allocation machinery can be replayed outside the U.S. using local-currency returns, the XETR trading calendar, and Germany/euro-area macro inputs. The full-sample classifier is shown first as a plausibility check: before presenting any allocation result, the regime path should identify recognizable stress episodes.

The two forward cuts answer different questions. The 2023-forward cut is a transition stress test that begins immediately after the 2022 inflation shock, when the classifier is most likely to carry inflation-stress information into normalization. The 2024-forward cut is the cleaner live-forward view after the model has seen the inflation shock and the first leg of normalization. If only one Germany allocation result is ultimately shown, the 2024-forward cut is the more presentation-ready candidate.

Regime identification diagnostic

The diagnostic identifies the GFC as Crash and the 2021-2025 energy/inflation episode as Inflation Shock. During that Inflation Shock run, the saved German bond proxies also fell: the long and intermediate euro government-bond substitutes were both roughly -12%, while commodities and gold were positive.

200420082012201620202024

Limitations

The asset map is defensible, but not identical. The German test cannot perfectly reproduce the U.S. investable universe. The local sleeve uses DAX exposure, euro fixed income, Europe real estate, euro cash, commodities, and gold; some U.S. style, small-cap, and sector roles require broad UCITS placeholders.

That makes this a robustness test of the framework, not evidence that the exact U.S. portfolio map transfers one-for-one.

Findings

The post-normalization path is constructive. In the 2024-forward view, Germany RAD-AA produced +21.96% CAGR with a 1.87 Sharpe.

Across 5 seed/lambda variants, Germany RAD-AA had a CAGR range of 12.5% to 14.3%, a Sharpe range of 1.16 to 1.42, and a worst max drawdown of -13.4%.

Primary forward view · 2024

This is the cleaner presentation cut: it starts after the inflation shock and first normalization year are in the calibration set.

Cumulative return

Drawdown and regime path

2024
StrategyCAGRVolSharpeMax DDTotal
Germany RAD-AA+21.96%11.71%1.87-10.41%+62.4%
DAX only+14.94%16.05%0.93-15.96%+40.5%
EUR 60/40+9.43%11.01%0.86-10.43%+24.6%

Growth

100.0%

616 days

Calm

0.0%

0 days

Macro Stress

0.0%

0 days

Inflation Shock

0.0%

0 days

Crash

0.0%

0 days

Limitations

The start date is intentionally less clean. The 2023-forward cut begins closer to the post-shock normalization boundary, so it is more exposed to whether the classifier carries Inflation Shock forward too long.

Findings

The transition view remains informative. Inflation Shock accounted for 73.4% of evaluated days, and Germany RAD-AA produced +14.28% CAGR with a 1.42 Sharpe.

Transition stress view · 2023

This cut is retained to show the sensitivity around the disinflation boundary, not because it is the preferred headline result.

Cumulative return

Drawdown and regime path

2024
StrategyCAGRVolSharpeMax DDTotal
Germany RAD-AA+14.28%10.05%1.42-11.63%+58.6%
DAX only+16.00%15.24%1.05-15.96%+67.0%
EUR 60/40+10.92%9.94%1.10-9.12%+43.1%

Growth

0.0%

0 days

Calm

26.6%

232 days

Macro Stress

0.0%

0 days

Inflation Shock

73.4%

639 days

Crash

0.0%

0 days

§02UK replication

A second local-market replay with a less clean classifier diagnostic.

The UK test uses GBP/LSE-listed substitutes, FTSE 100 returns in pounds, UK inflation and stress proxies, and the LSE trading calendar. The allocator replay is useful; the full-sample regime-identification diagnostic is less clean than Germany.

Why and how

The UK replication uses GBP/LSE-listed substitutes, FTSE 100 returns in pounds, UK inflation and rate data, a sterling corporate-bond-versus-gilt credit proxy, and the LSE trading calendar. The goal is the same as Germany: test whether the research architecture can be transported into another developed market without reusing U.S. returns.

The 2024-forward cut is the primary view because it starts after the 2021-2023 inflation/gilt shock has entered the calibration sample. The 2023-forward cut is retained as a transition stress view because it begins closer to the normalization boundary and exposes how sensitive the UK classifier is to that handoff.

Regime identification diagnostic

The UK GMM/CSJM identifies the GFC as Crash and the 2021-2023 gilt/inflation episode as Inflation Shock. It is less clean in the middle of the sample, where some post-crisis and Brexit-era observations are labelled Macro Stress even when the economic interpretation is more ambiguous.

20082012201620202024

Limitations

The UK proxy set is useful, but less stable. The largest limitation is the credit feature. The U.S. setup uses a clearer corporate-credit spread, while the UK replication uses a public sterling corporate-bond-versus-gilt relative-performance proxy.

The investable universe also cannot perfectly reproduce the U.S. sleeve map. The allocation results should be interpreted as a foreign-market robustness replay, not as a finished UK strategy.

Findings

The allocation replay is useful, with lower label confidence. In the 2024-forward primary view, UK RAD-AA produced +15.76% CAGR with a 1.89 Sharpe.

Across 5 seed/lambda variants, UK RAD-AA had a Sharpe range of 1.41 to 1.69. The Inflation Shock share ranged from 40% to 100%, which is useful evidence but also a reminder that the UK replication is the noisier test.

Primary forward view · 2024

This cut starts after the inflation/gilt shock and first normalization year are available to the classifier.

Cumulative return

Drawdown and regime path

2024
StrategyCAGRVolSharpeMax DDTotal
UK RAD-AA+15.76%8.34%1.89-9.05%+42.3%
FTSE 100 only+12.88%11.66%1.10-13.23%+33.9%
UK 60/40+8.61%8.04%1.07-8.50%+22.0%

Growth

100.0%

607 days

Calm

0.0%

0 days

Macro Stress

0.0%

0 days

Inflation Shock

0.0%

0 days

Crash

0.0%

0 days

Limitations

The transition boundary is the hardest part of the UK replay. The 2023-forward cut begins while the inflation/gilt shock is still fresh in the calibration history, so it is more sensitive to semantic label instability than the 2024-forward cut.

Findings

The transition path is still useful context. The 2023-forward transition view produced +15.68% CAGR with a 1.41 Sharpe.

Transition stress view · 2023

This cut is retained to show how the classifier behaves closer to the post-shock handoff.

Cumulative return

Drawdown and regime path

2024
StrategyCAGRVolSharpeMax DDTotal
UK RAD-AA+15.68%11.16%1.41-11.97%+64.2%
FTSE 100 only+9.95%11.57%0.86-13.23%+38.1%
UK 60/40+7.02%7.80%0.90-8.49%+26.0%

Growth

0.0%

0 days

Calm

59.8%

513 days

Macro Stress

0.0%

0 days

Inflation Shock

40.2%

345 days

Crash

0.0%

0 days

§03U.S. rolling-window audit

A moving-decade test after enough rare regimes are visible.

The first scored decade begins on 1991-04-15. The model receives only pre-window data, then a three-year causal warmup, before the 10-year scored period begins.

Why and how

The U.S. rolling-window audit asks whether the five-regime framework still works when each scored decade is treated as unknown future data. The model receives only pre-window observations, then a three-year causal warmup, before the ten-year scored period begins.

The first scored decade begins on 1991-04-15 because the framework needs enough prior history to identify all five model-relative regimes before it is judged. The initial training set already includes the 1970s inflation episode, Volcker tightening, the 1973-1974 stress period, the 1987 crash, and several monetary-cycle environments.

26 annual windows · April 15 anchor · 10 scored years each · first score 1991-04-15 · last score 2016-04-15 · minimum initial model-relative regime count 122 days

Limitations

The windows are useful, but not independent proof. The audit is still constrained by regime rarity. Even when every model-relative regime clears the minimum observation threshold, expanding training data can change which observations define a fitted component. The windows also overlap heavily, so the test should not be read as 26 independent live track records.

The value of the audit is narrower and more honest: it checks whether the loss-control result survives when the scored decade moves through time after enough rare macro environments have been observed.

Findings · win rate

The five-regime framework had a smaller maximum drawdown than 60/40 in 25 of 26 windows. Sharpe beat 60/40 in 16 windows, while CAGR beat 60/40 in 10. This is a risk-control finding first.

Findings · start-date map

Positive values mean the five-regime framework beat 60/40 for that scored decade. The map runs from the 1991 score window through the 2016 score window and shows where the advantage is path-dependent.

Findings · Sharpe persistence

Sharpe advantage is shown on its own ratio scale. Positive values mean the five-regime framework produced a higher Sharpe ratio than 60/40 for that scored decade.