Every investing claim is a claim about the future.
"Buy when the 200-day average turns up." "Sell when RSI is above 70." "Bitcoin does not top until the ISM peaks." Each one says, "This pattern held before, so it will hold again."
That deserves a test. Not a chart with three arrows on it. A test over many years, across many markets, with the answer written down before you look.
This piece explains how we test, what we have learnt from doing it, and why most of the work is about losing less, not winning more.
1. Markets change their behaviour, but they do not announce it
A trading rule is a bet that the market keeps behaving the way it did when the rule was found. Markets do not cooperate. Rates rise. Inflation returns after a decade away. A pandemic closes the economy. Momentum leads for three years, then value leads for two.
We call these periods regimes. You cannot tell which one you are in until it is well underway. You cannot tell when it ends either. A rule tuned to the last regime works beautifully right up to the change.
Then it stops. Without warning.
The chart below shades the stretches when SPY was most volatile, with its daily swings running above 20% a year. Each stressed period came from a different direction. None looked like the one before.

Source: daily SPY adjusted closes.
There are two ways out.
One is to spot the new regime the moment it starts. In practice, that is close to impossible.
The other is to build methods that survive every regime: not the best in any one of them, but acceptable in all of them. That is the path we take. It gives up some upside in exchange for not being wrong-footed at the turn.
2. Most claims fail on data before they fail on logic
Here is a real one. Since 2010, SPY has fallen 15% from its high six times. Three months after each crossing, it was higher. Six for six.

Source: daily SPY adjusted closes.
Even a win can hide pain. In 2022, the three-month return was +3.8%. The fall then carried on to −24.5%.
Two things are wrong with treating this as a rule. First, six events are not enough. Since 2010, a three-month hold in SPY has been positive about three-quarters of the time, whatever the start date. From those odds, six wins in a row happen by chance about once in six.
Second, the pattern was found on the same sixteen years it is tested on. It has never faced a period it was not fitted to. The one time it fails is the one that matters, because that is when conviction is highest and the position likely largest.
Well-known “expert” indicators are not exempt. The famous Sahm rule links a rise in the unemployment rate to the start of a recession. It had matched every US recession since 1970. It reached its trigger line again in July 2024. But no recession followed.

Source: FRED (SAHMCURRENT, USREC).
The inverted yield curve, where two-year Treasuries yield more than ten-year ones, had a similar record. It inverted in mid-2022 and stayed inverted for longer than at any time in the series. The recession did not arrive on schedule.

Source: FRED (T10Y2Y, USREC).
Neither indicator had missed a recession before. Each was right too few times to know why it was right. Eight recessions is a sample of just eight.
So the first question we ask of any input is how much history it has. We want ten years or more, covering at least one bear market, one crash, one sideways grind and one strong rally. New names now need at least seven years of testable history. We pass on anything shorter.
3. How we run a backtest so that it can fail
A backtest applies a rule to past prices as if you had traded it. It is easy to make one look good. The discipline is in making it able to fail.
Split the data before you look. Settings are chosen on one slice of history, then judged on a later slice they never saw. The window then rolls forward, and the process repeats. This is called a walk-forward. The record it builds consists only of decisions made without knowing the answer.

Source: Yimin Xu.
Freeze the prediction first. Before a study runs, we write down what counts as success and what counts as failure. The goalposts cannot move after the numbers arrive.
Judge consistency, not the total. One large win can carry a ten-year result. It can also be luck. So we look at each period on its own. The chart below does it year by year. A rule that beats buy-and-hold by a little in most periods is worth more than one big year.

Source: Yimin Xu backtest; SPY adjusted prices.
Clear a margin, not a line. Beating buy-and-hold on raw return is not the test. A rule sits in cash some of the time, so it often makes less in a strong rally. The test is return for the risk taken. It must beat simply holding the asset by a clear margin, either for the volatility it carries or for the drawdown it suffers. Drawdown is the fall from a peak to the next trough. Its worst fall must also stay within a limit. Matching buy-and-hold with more effort is not worth running.
Test it against nothing. Every model component is also tested against a scrambled version of itself. If the scrambled version does about as well, the component carried no information. It is removed, however good its backtest looked.
Test across assets. Each model is tested across every name we cover: equities, bonds, commodities, currencies and crypto. A model that works on only one stock has probably fitted that stock's history. The choice for each name is then made within that tested set.
Expect to throw most of it away. In our latest review, we passed on nearly two-thirds of the tickers we tested. Fewer than half of everything we have tested is in the live set today. Most failed on history or on margin.
Then test it live before trusting it. Any change to a model or a portfolio first runs in the shadow of the live system. It trades on paper beside what we publish for several weeks. Every promotion so far has surfaced something the backtest missed.
4. A rule is a probability, not a prediction
A tested rule does not say "this trade will work". It says that across many trades, the gains have outweighed the losses, by a modest margin, over many years. On any single trade, the outcome is mostly noise.
So you only collect the edge by taking every signal, across many names, for a long time. Ten trades are mostly luck. A few hundred start to show the average. A casino does not know the next spin. It knows the next ten thousand. The chart below shows the idea with a simple rule that wins 55% of the time.

Source: Yimin Xu.
It follows that one trade tells you nothing about the rule, win or lose. We judge a rule on its record across all our names and across years.
The hardest part is following the rule after three losses in a row. Skipping the next signal is a new rule, untested, chosen at the worst moment. The moment you pick and choose, you are no longer running the thing that was tested.
5. What the models look at, and the four proprietary models we run
Every name in our Multi-model Signals is scored by four models each day. Each casts a vote: long or flat.
The models read only public market data. That covers the name's own price history, trend and volatility. It covers how the name moves against the wider market. It also covers context from other assets, such as bonds, credit, currencies and commodities. No alternative data. No sentiment feeds. The edge, such as it is, comes from testing which inputs carry information for each name. The rest are thrown away.
Machine learning. This model, built from gradient-boosted trees, weighs many inputs at once in ways a human rule cannot. Its weakness is flexibility. With enough settings, it can fit the past perfectly and learn nothing.
Neural networks. These are forecasting models built for time series. They read the recent sequence of prices rather than a single snapshot, so they see patterns the tree model cannot. They are also the most flexible thing we run, which makes them the most able to fool a backtest. So they face the same tests as everything else.
Trend. The oldest idea in the set. It compares fast- and slow-moving averages of the name's own price. It is slow to turn and gives back gains at tops. On its own, it is a weak way to choose what to hold. Its value is timing: it gets a name out of a long decline.
Market regime. This is a hidden Markov model, a statistical model that sorts each name's history into rising, sideways and falling states. The state is never observed directly, only inferred from the price. It votes flat only in the falling state. Like any regime detector, it recognises a change after it has begun, not before.
Why several, and why they vote. Each model fails in a different way. Trees overfit snapshots. Networks overfit sequences. Trend is late. The regime model is late in a different way. When two of them agree, the odds that both are wrong together are lower. Most names trade a majority vote. The rest trade the combination that tested best for that name. Once a position changes, it is held for a minimum number of days, so a single day's disagreement does not become a trade.

Source: Yimin Xu.
That holding rule costs a little in fast reversals. It removes most of the churn, which is the thing a subscriber cannot live with.
6. Why the work is about drawdowns, not alpha
A private investor cannot. A 50% loss needs a 100% gain to get back to even. The maths is bad. The behaviour around it is worse: selling near the bottom and missing the recovery turns a temporary loss into a permanent one.

Source: Yimin Xu.
So almost none of our testing hunts for the best stock or the best month. It looks for a path smooth enough to stay on.
In the ten-year backtest of the Macro & Megacaps Systematic Portfolio, the deepest fall was 15.8%. SPY's was 32.0%. The falls came at the same times, but the portfolio's were shallower and shorter.

Source: Yimin Xu backtest; SPY adjusted prices.
The portfolio outperformed SPY in 59% of months. That is a modest edge. The number that matters is the other one: in the 35 months when SPY fell, the portfolio's mean loss was 1.6%, against the index's 4.0%. Same direction, less than half the damage.

Source: Yimin Xu backtest; SPY adjusted prices.
These are backtest figures, before trading costs, from June 2016 to July 2026. The portfolio's settings were chosen on 2016 to 2022, then checked on 2023 to 2026. Live results are published separately.
That is the trade we make on purpose. Give up some of the best months. Avoid most of the worst ones. Stay invested long enough for the average to work.
7. Learn more about our Systematic work
This essay is the method behind everything we publish at YX Insights. We apply it in two products.
The Systematic Portfolio turns the models into ready-made portfolios for Macro & Megacaps and for Commodities. We publish the holdings and every change.
The Multi-model Signals give a daily long-or-flat call on every name in both groups. Each call shows how strongly the four models agree.
Alongside them, we publish research: company deep dives, macro commentary and essays like this one.
If this way of testing makes sense to you, you can see all of it on the website.
DISCLAIMER: This newsletter is strictly educational. Any information or analysis in this note is not an offer to sell or the solicitation of an offer to buy any securities. Nothing in this note is intended to be investment advice and nor should it be relied upon to make investment decisions. Any opinions, analyses, or probabilities expressed in this note are those of the author as of the note's date of publication and are subject to change without notice.