Scope and learning objectives
This article examines attributed historical cases and observations from the supplied course archive. The figures are self-reported or participant-reported, use differing strategies and methods, and are not presented as comparable results or general trading guidance.
- Distinguish a strategy’s documented development history from evidence produced through backtesting or live tracking.
- Interpret reported win rates alongside their labels, strategy context, and stated risk-reward observations.
- Identify how time period, overlap treatment, and calculation method affect the meaning of reported performance.
- Recognize when favorable historical results coexist with an explicit methodological weakness.
01
Case 1: Development Without Backtesting
In the documented development history, the speaker says the M3, M3.4.4, and bearish-butterfly strategies were created without backtesting. The M3 account adds historical context: use began in late 2006 or early 2007, when the speaker describes backtesting software as unavailable to them and the accessible data as limited and spotty. [1][2]
02
Case 2: Reported Win Rates in Context
The records contain several non-comparable win-rate observations. One speaker reports that backtests and live tracking since 2010 produced outcomes that diverged sharply from two probability labels. Elsewhere, a speaker characterizes a 70% futures win rate as extremely high and says their own hit rates are usually nearer 50%, while recalling a profitable day trader with roughly a 30% win rate and favorable risk-reward ratios. [3][4][7]
- A setup labeled at about 90% probability reportedly won about 75% of the time, whereas another labeled around 40% reportedly won about 90% of the time in the speaker’s backtests and live tracking since 2010. [3]
- In a separate futures observation, the speaker says 70% would be an extremely high win rate and places their usual hit rates closer to 50%. [4]
- The recalled day-trader case associates profitability at about a 30% win rate with favorable risk-reward ratios, but the causal attribution is unverified. [7]
03
Case 3: Performance Changed Across Periods
Two historical reports illustrate time-dependent outcomes without establishing a common pattern. A participant says the RUT bull strategy performed well from 2011 through 2017 but then consistently underperformed the S&P 500 from the start of 2018 through the backtest date. Separately, a speaker says referenced strategies traded at consistent size and according to their guidelines over 10–15 years produced annual results ranging from gains near 80% to gains near 10%, with losing years as well. [5][6]
- The RUT bull report changes from favorable performance in 2011–2017 to benchmark underperformance beginning in 2018 and continuing through the test date. [5]
- The longer-horizon strategy report includes large gains, modest gains, and losses despite the stated use of consistent size and strategy guidelines. [6]
04
Case 4: Calculation Choices and an Explicit Warning
The final cases show that a reported return can change with its calculation frame and that strong historical figures can coexist with a stated weakness. For an unspecified strategy, the speaker reports about 25–30% annually when overlap was considered and roughly 40% trade for trade when overlap was excluded. In another case, a manually managed methodology reportedly produced a 50% ROI in both 2020 and 2021, yet the same speaker describes its stops as insufficiently rigorous and expects some future condition to make it fail. [8][9]
- The overlap-aware annual figure and the overlap-excluding trade-for-trade figure are different measurements and are reported as approximately 25–30% and 40%, respectively. [8]
- The manual methodology’s reported 50% ROI in 2020 and 2021 does not remove the speaker’s explicit concern about inadequate stops. [9]
- The speaker’s expectation that the manual methodology would eventually encounter a condition that made it fail is part of the same historical account, not a documented subsequent outcome. [9]
Review
Key takeaways
- In these cases, strategy development history and validation evidence are distinct: several named strategies were reportedly developed without backtesting, and the M3 account attributes that process to limited tools and data at the time. [1][2]
- The reported probability labels, observed win rates, and low-win-rate day-trader recollection cannot be interpreted as interchangeable measures or universal performance rules. [3][7]
- The historical records include both regime-like changes across periods and substantial year-to-year variability within longer reported histories. [5][6]
- Return figures in these cases retain meaning only with their measurement frame and stated limitations, including overlap treatment and the explicit warning about insufficiently rigorous stops. [8][9]
Self-check
Review questions
What does the M3 development case establish, and what does it leave unresolved?
It records the speaker’s claim that the M3 was developed without backtesting software in late 2006 or early 2007 amid limited, spotty data; it does not establish later validation quality or performance. [2]
Why should the probability labels in the tracked-setup case not be treated as observed win rates?
The speaker reports materially different observed outcomes: the setup labeled about 90% won about 75%, while the one labeled around 40% won about 90% in their backtests and live tracking since 2010. [3]
What prevents the RUT bull result and the 10–15-year annual-results account from being directly compared?
They cover different and partly unspecified strategies, periods, and evaluation frames: one is a participant-reported backtest against the S&P 500, while the other is a speaker-reported range of annual strategy results. [5][6]
How should the two return figures in the overlap case be interpreted?
They belong to different calculation frames: about 25–30% annually with overlap considered and roughly 40% trade for trade with overlap excluded; the calculation method and strategy are unspecified. [8]
What tension is explicit in the manually managed methodology case?
The speaker pairs reported 50% ROI in both 2020 and 2021 with a warning that the stops were insufficiently rigorous and an expectation that some future condition would make the methodology fail. [9]
Traceability
Evidence index
Canonical source claims used in this guide. Open a session link to verify the underlying passage at its original timestamp.