Painterly dusk scene of people climbing a hillside path toward a monumental cantilevered modern structure.

Monte Carlo vs Historical Backtesting in Retirement Plans

History asks whether a plan survived the actual past. Simulation asks how often it survives in a world the model believes in. Killion runs both.

A historical backtest replays an allocation and a spending rule through real market start years. FI Calc is the best-known free version of that idea. A Monte Carlo draws (or, in our case, walks a regime model through) many synthetic lifetimes. They are not two implementations of one test. They answer different questions, fail in different ways, and become dangerous when a website treats the output of one as if it were the output of the other.

The rest of this page is the distinction, with a count of how many independent retirements the U.S. tape actually contains. Once you have seen that count, the habit of quoting a historical success rate as a probability gets harder to defend — and the habit of ignoring history because “we have simulation now” gets harder too.

Monte Carlo vs historical backtesting retirement: two questions

Monte Carlo vs historical backtesting retirement is not a branding choice. A historical backtest retirement study asks: if this household had retired in 1966, or 1973, or 2000, with this allocation and this rule, would the money have lasted? The inputs are a real return series and a start-year index. The output is a count of surviving windows. A Monte Carlo asks: in a world whose returns are generated the way this model generates them, how often does the same household last? The inputs are a process and a seed. The output is a proportion with a standard error.

Those sentences look similar because both end in a percentage. They are not similar. History cannot invent a 40% crash that almost happened in 1988 and did not. Simulation cannot put you in 1929 unless the model has a 1929-shaped state. When a marketing page says “96% historical success” and a planning page says “87% simulated success,” they are not disagreeing about the same coin. They are reporting two experiments.

What a historical backtest retirement can say

History includes wars, inflations, and recoveries that actually happened. That is its entire advantage. A sequence that includes 1929–32, 1973–74, 2000–02 and 2008 is not a draw from a nice distribution. It is a specific, ugly, true path. If a rule fails those windows, you should not soothe yourself with a simulated 95%. The tape already voted.

History does not include the crashes that almost happened. It does not include a 1970s that lasted twelve years instead of ten, or a 2009 that never recovered. There are only so many independent 40-year windows since 1926, and they overlap. The U.S. annual record most withdrawal papers start from is about 100 years long (1926–2025). That is not a small sample for a stock-bond premium. It is a tiny sample for 40-year retirements.

0204060801520253035404550planning horizon (years)overlapping start yearsindependent windows (reference)
History is a short tape, counted twice. From 1926 through 2025 there are 71 overlapping 30-year retirements and only three that do not share years. Stretch the plan to 50 years and the independent count is two. A 96% on overlapping windows is not a 96% probability.

The overlapping-window problem

Count the windows. On a 100-year tape, a 30-year retirement can start in 71 different years. A 40-year retirement can start in 61. A 50-year retirement can start in 51. Those are overlapping counts: 1965–94 and 1966–95 share 29 years. They are not 71 independent lives. The dashed line on the chart is the non-overlapping count — floor(years / horizon) — and it is brutal. Thirty-year retirements: 3 independent windows. Forty-year: 2. Fifty-year: 2.

A 96% historical success rate on 71 overlapping 30-year tapes is not a 96% probability. It is “this rule failed in a handful of neighboring start years.” Failures cluster because the windows share the same crash. That is why historical success rates jump when you add or drop a single ugly decade, and why two researchers can report 95% and 80% on “the same 4% rule” after changing the end date. The sample is small and dependent. Treating it as a binomial with n = 71 is already too kind; treating it as n = 1,000 is fiction.

A concrete pair makes the dependence obvious. The 1965–94 window and the 1966–95 window share twenty-nine years. If 1973–74 wrecks the first, it wrecks the second. Counting both as independent “failures” double-weights one bear market. Counting both as independent “successes” double-weights one recovery. Either way the headline percentage is a vote by a small committee that keeps recasting the same members. That is useful information about the decades we actually got. It is a poor denominator for a probability.

Researchers sometimes report an “effective sample size” that tries to discount the overlap. The dashed line on the figure is the blunt version of that discount: how many retirements you could line up end to end on the same tape. Three 30-year lives. Two 40-year lives. Two 50-year lives. Anything you do to sound more precise than that — a HAC standard error, a block bootstrap — is arguing about the fraction of a handful. It will not turn 61 overlapping 40-year windows into a 1,000-path Monte Carlo.

What a Bengen simulation is (and is not)

A Bengen simulation, in the 1994 paper, is a rolling historical window: pick a start year, withdraw a fixed real amount, step through the actual subsequent returns, see if the portfolio survives 30 years. It is a backtest with a researcher's discipline, not a Monte Carlo. Later writers sometimes say “we ran a Bengen simulation” when they mean “we ran a Monte Carlo with Bengen-like assumptions.” Those are different machines. One consumes a tape. The other consumes a process.

The original paper’s power is that the tape is real. Its limit is that the tape is short, U.S.-heavy, and already lived. You cannot add a 50-year early-retirement case without either overlapping the same crises even more or leaving the historical method for a model. That is why our withdrawal-rates-by-age table is a simulation, and why this page exists: so nobody mistakes that table for another Bengen chart.

Regime switching vs history

Regime switching vs history is the comparison the product is built to show. A five-regime engine tries to put 1929-shaped clusters back into the draws: crashes that persist, correlations that break, inflation that has a memory. It is still a model. It can invent ugly decades the 20th century did not bother to run, and it can miss a texture the 20th century actually had. The methodology page states the regimes. The June Monte Carlo article shows what happens when you replace those regimes with independent yearly draws: the success rate rises, because risk has been told to arrive politely.

Simulation can give you 1,000 roughly independent lifetimes, a seed, and a standard error — see how many paths are enough. It cannot give you 1929. History can give you 1929. It cannot give you a standard error that means what a statistics textbook says it means, because the windows are not independent. Use the historical number as a veto (“this failed the tapes we have”) and the simulated number as a range (“in the world this model believes in, here is how often it works”). Do not average them. Do not pick the prettier one.

When the two disagree

If history says a rule never failed and the Monte Carlo says it fails 8% of the time, the plan is more fragile than the tape suggests — or the model is harsher than the 20th century. Both readings are possible. Look at which simulated failures: clustered inflation, a long equity drought, a spending rule that cannot cut. If those states look like history’s bad decades, believe the simulation’s warning. If they look like a model that cannot stop raining, believe the tape.

If both say the rule fails often, believe them. If history fails and the Monte Carlo prints 98%, the engine is too gentle; that is the subject of the overstatement article. The product backtest starts in 1928; the projection is 1,000 seeded paths. The Guyton-Klinger calculator now has a historical-window mode and a 1,000-path mode for the same reason. Two percentages, two questions, one household. That is the whole method.

Tools that only replay history are still worth using. They will show you 1966 and 2000 in a way no regime model can, because those years happened. Tools that only simulate are still worth using. They will show you a drought the 20th century did not happen to run, and they will give you a standard error you can write down. The failure mode is a page that offers one percentage and lets you decide, silently, that it is the other kind. Monte Carlo vs historical backtesting retirement only works as a pair.

If you want a working order: run the historical window first as a veto. If the rule dies in 1966 or 2000 on your actual spending, stop and change the rule. Then run the 1,000-path projection and read the range, not the single percent. If the simulated failures look like the historical ones, you have a consistent story. If they do not, you have a modeling problem — and that is a better problem to have before you retire than after. The demo and the calculators are built so both steps fit on one screen.


Notes on the figures. Window counts are closed-form on a 1926–2025 annual tape (100 years, the Ibbotson / SBBI start year used in Bengen 1994). Overlapping windows of length h equal 100 − h + 1. Independent windows equal floor(100 / h). No market-return series is fetched at render time; the chart is arithmetic. Product backtests start in 1928; that is one year later than the 1926 convention and does not change the shape of the figure. This article is educational analysis, not investment advice, and does not recommend any security or strategy.

References

  1. William P. Bengen, “Determining Withdrawal Rates Using Historical Data,” Journal of Financial Planning (1994): rolling 30-year historical windows, not Monte Carlo.
  2. Philip L. Cooley, Carl M. Hubbard, and Daniel T. Walz, “Retirement Savings: Choosing a Withdrawal Rate That Is Sustainable,” AAII Journal (1998).
  3. On overlapping windows as a dependent sample, see any discussion of rolling-return statistics; the independent-window line on the figure is the blunt version of that critique.
  4. Killion methodology and the regime engine: /methodology/.