More gold meant more Sharpe — until it didn’t
- Core
- 60% CRSP equities / 40% 10y Treasury
- Funding
- Pro-rata from both sleeves
- Rebalance
- Annual, December · gross of cost
- Sweep
- 0–50% in 1pt steps · 9 shown
| Gold weight | Sharpe | Vol % | MaxDD % | Real % | Ulcer % |
|---|
Full sweep runs 0–50% in one-point steps; the chart shows 0, 1, 5, 10, 15, 20, 25, 30 and 50%. Sharpe uses T-bill excess returns. Select a row for the full metric strip.
Against the cheaper alternative, gold still wins — mostly
- Gold leg
- 10% gold, pro-rata funded
- Matched leg
- ~6.5% equity → bonds, same volatility
- Metrics
- Sharpe · real return · maxDD · ulcer
- Also tested
- Drawdown-matched · same-weight cash
| Sample | ΔSharpe | ΔReal pp | ΔMaxDD pp | Wins / 4 | Verdict |
|---|
Δ is gold minus the volatility-matched de-risk. Against all three matched variants jointly, gold sweeps all four metrics in 3 of 7 samples (1990, 2000, 2005). The drawdown-matched portfolio reaches gold’s drawdown at lower volatility in the two longest samples — at a real-return give-up of ~0.98pp.
Seven eras that don’t overlap — and don’t agree
- Comparison
- 10% gold vs 60/40 within segment
- Segments
- 7 disjoint eras, 36–258 months
- Robustness
- 6-of-7 drawdown survives 9 re-cuts
- Honest frame
- Drawdown, not Sharpe, is the claim
| Segment | Months | Gold % | ΔSharpe | ΔMaxDD pp | ΔReal pp |
|---|
ΔSharpe reads +0.69 (1972–74), +0.23 (1975–79), −0.10 (1980s), −0.08 (1990s), +0.07, +0.08. The 1968–71 pegged segment is shown for completeness; this page’s base case starts in 1972.
Eight of nine crises: the sleeve softened the fall
- Episodes
- 9 worst equity drawdowns, 1968–2026
- Measure
- Peak-to-trough total return
- Legs
- 60/40 vs 10% gold, pro-rata
- Exception
- Nov 1980 – Jul 1982 only
| Episode | Mo | Equities % | Bonds % | Gold % | 60/40 % | 10% gold % |
|---|
Equity figures are CRSP total returns — the 1980–82 gap vs an S&P 500 price fall near −24% is +7.5 points of dividends at 5%+ yields. All figures month-end; the COVID daily trough (−34%) is deeper than any monthly print.
The data cannot pick a weight
- Edge
- ΔSharpe vs vol-matched, 7 samples
- Noise
- Weight-specific bootstrap SE
- Worst case
- 1980-start at every weight
- Span
- 2.5–25% inside noise
| Weight | Median Δ | Worst Δ | Boot SE | Worst / SE | Samples + |
|---|
Median edge is positive at every weight; the worst sample (1980-start) is negative at every weight. “Samples +” counts samples with a positive ΔSharpe: 6 of 7 at every weight except 25% (5 of 7). The SE is computed separately per weight — it scales 0.011 at 2.5% to 0.106 at 25%.
Executive summary — what the evidence supports
- No identifiable optimum. The Sharpe-optimal weight ranges 2–33% across the 7 samples, and every weight from 2.5% to 25% sits inside the estimation band against the matched benchmark — median and worst case alike.
- The risk case is real; the return case is not claimed. At 10% gold, 1972–2026 Sharpe rises 0.54 → 0.59, volatility falls 10.2 → 9.5%, max drawdown improves −29.1 → −21.9%. Real return also rose in this sample — but it falls at 7.5% weight in 3 of 7 samples, so no return uplift is found.
- The comparison that matters is the matched one. Gold’s edge over plain 60/40 conflates gold with de-risking. Against a volatility-matched de-risk, gold wins all four metrics in 5 of 7 samples — and including a drawdown-matched variant, in 3 of 7.
- The failure mode is a twenty-year drought. Real drawdown reached 82.9%, troughing March 2001; the January-1980 real peak took 44.8 years to regain on a month-end basis. The 1980–2001 episode cost −1.10pp of CAGR at 7.5% weight and −3.72pp at 25%.
- Nothing is statistically significant — and that is disclosed, not hidden. Dependence-robust p for the headline test is 0.19 (iid Memmel: 0.078); 0 of 21 tests survive Holm correction. Only the 2000-start sample clears 5% on its own.
- No substitute beat gold. Across 40 alternatives — TIPS at five maturities, commodities, managed futures, miners, silver, platinum, long Treasuries, bitcoin — none beat gold on Sharpe over its own overlap window.
- Low hurdle, modest cost sensitivity. At a 10% weight the compound break-even is ~1.9% nominal (−0.5% real) under base assumptions — cleared in every sample. A 25bp vehicle leaves the Sharpe gain intact (0.054 vs 0.056); 75bp trims it to 0.048.
Sources & methodology
- Gold: LBMA Gold Price AM daily fixings, Jan 1968 – Aug 2026 — a three-regime splice (telephone fixing pre-20 Mar 2015, ICE electronic auction after). London was closed 15–31 Mar 1968; that month’s 0.0% is a closed-market artefact. PM fixings start Apr 1968; the choice changes nothing.
- Equities: CRSP value-weighted total return (Ken French library), Jul 1926 – Jun 2026. In the 2000s this runs 0.63pp/yr above the S&P 500 — small/mid-cap breadth, not an error.
- Bonds: synthetic 10y constant-maturity total return repriced monthly from FRED DGS10 (par bond, semi-annual coupons; validated against IEF at 0.990 correlation).
- Inflation & cash: CPI-U NSA (Oct 2025 interpolated across the shutdown) and 1-month T-bills. Common monthly sample: 701 observations, Feb 1968 – Jun 2026.
- Verification: an adversarial audit (57 checks: independent re-implementation, alignment, external validation) reports 48 PASS, 9 FLAG, 0 FAIL — the flags mark known limitations, not errors.