A diversified portfolio rests on a reassuring idea: different investments will not all behave the same way at the same time.

Stocks may fall while high-quality bonds rise. Energy producers may gain from the same inflation that hurts businesses with fixed costs. Combine enough of these and weakness in one place should be cushioned by strength in another. Then the market changes, and the assets meant to balance one another start falling together.

The usual explanation is a new "regime" — a word in every hedge-fund letter because it names something investors keep living through. Relationships that looked dependable in one environment turn unreliable in the next. Knowing that regimes exist is easy. Recognizing one in time is the hard part.

Diversification is a relationship, not a shopping list

Owning many investments does not make a portfolio diversified. Ten technology stocks are ten securities and may be one economic bet. A stock index and a corporate-bond fund look unrelated on a statement, yet both suffer when investors start doubting whether the same companies can grow and refinance.

What matters is the forces underneath the returns. Quantitative investors call them factors: cheap against expensive, recent winners against recent losers, small against large, and a few more built on profitability, interest rates and credit. They are not laws of nature. They are portfolios whose rewards come and go. Value can struggle for years. Momentum can reverse violently. Assets that normally defend a portfolio turn fragile when the dominant fear is inflation rather than weak growth.

A correlation is a description, not a contract.

It tells you how two investments moved together during one particular mix of economic conditions. Change the mix and the relationship is free to change with it.

What counts as a regime?

A regime is a stretch in which the market runs on a different set of relationships — calm or violent prices, rising or falling rates, easy or stressed credit. You can name one with rules, and pick thresholds nobody can defend. Or you can let a model sort the periods for you, and receive confident labels nobody announced. Either way, a method that names yesterday's regime precisely can still be late to today's turn.

Then there is the version that only works backwards. Read the whole record, mark every turning point with hindsight, and show what a strategy would have earned if it had known those labels at the time.

That is not forecasting. It is annotating history.

The test: what did the portfolio know then?

The experiment compares three approaches on public factor returns and market indicators. The first divides its risk equally and rarely changes. The second adjusts how much risk it takes as volatility moves, but never tries to name a regime. The third judges whether the market is calm, inflationary, stressed or recovering, and shifts its factor weights to match.

Every decision uses only what was knowable at the time, and trading costs come out whenever the portfolio changes. That second point carries more weight than it looks. Dynamic strategies look cleverest exactly when they trade most. A model can catch a subtle shift every month and still leave its owner poorer once turnover and false alarms are paid for.

So the test is not which portfolio ends up richest. A regime strategy earns its keep if it shortens the falls, avoids leaning on one lucky episode, and survives more than one definition of the market environment. If moving a single threshold destroys the result, the intelligence on display probably came from tuning.

Did adaptation help?

The measured result

The best regime-aware portfolio returned 9.75% a year with 5.36% volatility and a −12.48% maximum drawdown — a Sharpe ratio of 1.05.

The static equal-risk portfolio returned 7.76% at 4.01% volatility, a Sharpe of 0.94. The volatility-targeted portfolio — which never tries to name a regime, and only changes how much risk it takes — returned 11.12% at 6.84% volatility, a Sharpe of 1.02.

So the regime model won, by three hundredths of a Sharpe point, by earning less money.

That sentence is the finding, and it is worth sitting with. The regime strategy did not win by seeing a crisis coming and getting out of the way. It won by running 1.48 percentage points less volatility and taking a drawdown 3.21 points shallower. Put both at the same volatility and the baseline returns 9.59% against the strategy's 9.75% — an advantage of 0.16 percentage points a year.

In the months the baseline lost money, the strategy took 73 per cent of the loss. In the months it made money, 83 per cent of the gain. That is a defensive portfolio, not a better-timed one.

The strategies that did what the article promised, lost

The word "regime" hides a distinction that decides everything. One kind of strategy uses the regime label only to set how much risk to take. The other uses it to pick which factors to own — the version with the story attached, the one that says value works here and momentum works there.

The second kind is what the phrase "regime-aware investing" is usually selling. It lost to both baselines.

strategyreturnvolatilitySharpemax drawdownturnover
Static equal-risk7.76%4.01%0.94−10.00%0.09×
Volatility target, no regime model11.12%6.84%1.02−15.69%0.56×
Regime-aware, rule, risk sizing only9.75%5.36%1.05−12.48%2.44×
Regime-aware, model, risk sizing only10.78%6.58%1.01−15.69%1.10×
Regime-aware, rule, factor tilts9.95%7.75%0.77−20.93%6.21×
Regime-aware, model, factor tilts10.38%8.25%0.77−22.37%2.25×

557 months from February 1980. Costs charged at 10 basis points one-way on gross notional traded.

Look at the last two rows. Both factor-tilting strategies fell nearly twice as far as the winner, for a Sharpe ratio about a quarter lower. The rule-based one traded its whole book 6.21 times a year to get there — a cost drag of 0.68% a year against the winner's 0.27%.

The map from regimes to factors was written down in advance, as an economic hypothesis, before anyone scored the data. Its failure is evidence against that map, not proof that no map could work. But it is a specific and expensive failure of the exact thing the industry sells when it sells regime analysis.

The signal is real. The edge is not reliable.

Two tests point in opposite directions, and both belong here.

The first asks whether the timing is real. Slide the regime signal to the wrong decade, so it keeps its shape but loses its alignment with what markets did. Do that 200 times and the strategy beats every displaced copy of itself, at p = 0.005. Break the alignment and the advantage dies — exactly what should happen if the signal carries information.

The second asks whether the edge is worth anything. Change one assumption at a time and rerun the whole study. Across 33 such changes, a regime strategy beat the volatility target in 20, with a median advantage of 0.031 Sharpe points in a range from −0.043 to +0.228.

The signal knows something. It is not worth enough to survive the researcher moving a slider.

Try it

Move one assumption and watch the conclusion change sign

Every variation below is a real rerun of the whole study with one thing changed. Nothing about the market moves between them. Only a choice the researcher makes before seeing the answer.

This demonstration needs JavaScript. All 33 specifications are in the repository's reports/robustness.csv.

Two of those variations deserve naming. Charge 50 basis points to trade and the advantage falls to −0.043, so the regime strategy loses. That is steep for institutional factor trading, but not absurd for a book that turns over 2.44 times a year. Start the evaluation in 2000 instead of 1980 and +0.033 becomes −0.027. Neither is a hostile choice. Both are choices a careful person could have made first, and never known they had decided the answer.

Is it one crisis in a trench coat?

The other way a backtest earns an edge it cannot repeat is by owning one spectacular month. So each named episode gets deleted from the sample and everything runs again.

Try it

Delete a crisis from history and rerun the study

If a strategy's advantage is really one lucky episode, removing that episode takes the advantage with it. Here it does not — and in three cases the strategy looks better without the episode than with it.

This demonstration needs JavaScript. The episode-removal table is in the repository's reports/results.json.

The worst case is the 1987 crash: cut four months and 75 per cent of the advantage is still there. Nothing collapses. So this criticism does not land. The strategy is not one crisis wearing a portfolio — it is a small, defensive, unstable edge, which is a duller thing to be and a more honest one.

The deeper lesson sits underneath the ranking. Adjusting risk when volatility rises is not the same as predicting an economic regime. The volatility target holds no theory of the world at all. It reacts to what just happened, trades a fifth as much as the regime rule, and finished a rounding error behind it. A sophisticated model may offer a much richer story while reacting more slowly and trading too often. Complexity earns its place only when it improves a decision. Here it improved one by 0.16 percentage points a year, in one sample, on one country's factor data, in 20 specifications out of 33.

The portfolio is not broken — its assumptions are visible

When diversification fails, investors tend to conclude the idea stopped working. The better reading is that it always depended on conditions, and the conditions just became visible. Bonds protect stocks when falling growth pulls rates down, and may not when inflation pushes rates up. Momentum diversifies value right up until a reversal punishes recent winners. International assets spread you across economies and still fall together in a global rush for cash.

None of that makes diversification useless. It makes the assumptions behind it worth knowing. A portfolio should not depend on one historical correlation holding forever. It should stay tolerable across several plausible environments without asking its owner to call every turning point correctly.

The perfect portfolio is not the one that performs best in the regime we can name after it ends. It is the one that stays tolerable while the next regime is still unnamed.

What this study cannot support

The largest caveat is that this is not a point-in-time study. The factor files are rebuilt from a database that gets revised, so their history can change after the fact. The macro inputs were deliberately chosen because they cannot be revised. Even so, nothing here can be described as what an investor would definitely have achieved live.

It is one country, one asset class, and one set of six factors. Trading costs are a flat rate on everything traded, while real costs move with size, liquidity and urgency. Long-short factor portfolios can cost far more to trade than a flat rate suggests.

The Sharpe ratios carry two caveats rather than the usual one. The usual one: a ratio measured over a window by someone who has already seen that window overstates what a new investor should expect. The extra one: these six strategies share five of their six ingredients. A tenth of a Sharpe point between them sits well inside the range that can arise by chance. The winner beat the runner-up by 0.03.

Companion repository when-portfolios-change-regime, published with the verified results.

Repository specification

Build a publication-quality, fully reproducible GitHub repository named when-portfolios-change-regime. Test whether a regime-aware factor allocation improves on transparent static and volatility-responsive baselines after turnover and trading costs. Do not label regimes with hindsight or optimize the entire history before reporting an “out-of-sample” result.

Use Python 3.12 and monthly data. Download U.S. market, size, value, profitability, investment and momentum factor returns from the Kenneth French Data Library. Add a small predeclared set of public state variables from stable official sources: market volatility or realized volatility, Treasury-curve slope, a credit-stress measure and inflation or inflation expectations. Record exact definitions, publication timing, transformations, revisions, URLs, retrieval dates and file hashes in data/data_manifest.csv. When a macro series is revised, either use a point-in-time vintage or state clearly that the test is not point-in-time and exclude the revised value from any claim of live tradability.

Compare at least three portfolios: static equal-risk or equal-weight factor allocation; volatility-targeted allocation with no regime model; and regime-aware allocation. Implement both an interpretable rule-based regime model and one statistical alternative such as a hidden Markov model. For the statistical model, use filtered probabilities only and fit exclusively on the expanding historical window. Never use full-sample smoothed states to make past allocation decisions.

Predeclare the number of regimes, minimum training period, rebalance frequency, weight limits, leverage limit, volatility target, transaction-cost assumption and handling of negative long-short factor weights in configuration. Include a cash or Treasury-bill return. Constrain the regime strategy so that any gain cannot come from unlimited leverage. Report annualized return, volatility, Sharpe ratio with caveats, maximum drawdown, downside deviation, turnover, cost drag, worst rolling year and performance by decade. Attribute performance to factor exposures, risk level and a small number of major episodes.

Run robustness checks across regime counts, rule thresholds, training windows, rebalance frequencies, costs and evaluation start dates. Include a placebo or shuffled-state test to show how easily regime stories can arise by chance. Compare the regime model with the simpler volatility-targeted baseline, not only with a naïve equal-weight portfolio.

Create at least five publication-ready SVG and PNG figures: factor behavior across identified regimes; timeline of real-time regime probabilities; cumulative wealth and drawdown; turnover-versus-benefit decomposition; and robustness across specifications. Write exact article replacement fields to reports/article_values.md, including a table showing whether the conclusion survives each robustness check.

Include README.md, LICENSE, pyproject.toml, Makefile, src/, tests/, configs/, data/, reports/figures/, notebooks/, and GitHub Actions. Test temporal alignment, filtered-versus-smoothed states, portfolio weight constraints, transaction costs and P&L identities. make reproduce must regenerate the analysis. The README should explain factors and regimes without assuming prior knowledge and prominently distinguish an explanatory historical model from a dependable live forecasting system.