Gold, at Weight

Is there an ideal allocation to gold? Fifty-four years of portfolio evidence, measured against the benchmark it deserves — a volatility-matched 60/40, not an un-de-risked one.

Coverage
Feb 1968 – Jun 2026
Sample
701 months
Base case
1972–2026 · 60/40 core
Source
LBMA AM · CRSP · FRED
Metric
Sample
Theme
Gold sleeve (pro-rata funded) 60/40, no gold Volatility-matched de-risk Adverse / indistinguishable from zero
2–33%
Sharpe-optimal gold weight across 7 samples
0.54 → 0.59
Sharpe, 1972–2026, 0% → 10% gold
−29.1 → −21.9%
Max drawdown, 1972–2026, 0% → 10% gold
8 of 9
Equity crises where a 10% sleeve helped
Exhibit 01 · Weight sweep

More gold meant more Sharpe — until it didn’t

Risk-adjusted return rises with weight to about 20%, then falls; at 50% the 1972 Sharpe is back near zero-gold. Switch sample to see the one that breaks the pattern: 1980.

Core
60% CRSP equities / 40% 10y Treasury
Funding
Pro-rata from both sleeves
Rebalance
Annual, December · gross of cost
Sweep
0–50% in 1pt steps · 9 shown
Gold weightSharpeVol %MaxDD %Real %Ulcer %

Full sweep runs 0–50% in one-point steps; the chart shows 0, 1, 5, 10, 15, 20, 25, 30 and 50%. Sharpe uses T-bill excess returns. Select a row for the full metric strip.

Exhibit 02 · Matched benchmark

Against the cheaper alternative, gold still wins — mostly

Pro-rata funding cuts equity 60→54, so the honest comparator is a volatility-matched de-risk, not plain 60/40. Gold wins all four metrics in 5 of 7 samples — and loses all four in the 1980 sample.

Gold leg
10% gold, pro-rata funded
Matched leg
~6.5% equity → bonds, same volatility
Metrics
Sharpe · real return · maxDD · ulcer
Also tested
Drawdown-matched · same-weight cash
SampleΔSharpeΔReal ppΔMaxDD ppWins / 4Verdict

Δ is gold minus the volatility-matched de-risk. Against all three matched variants jointly, gold sweeps all four metrics in 3 of 7 samples (1990, 2000, 2005). The drawdown-matched portfolio reaches gold’s drawdown at lower volatility in the two longest samples — at a real-return give-up of ~0.98pp.

Exhibit 03 · Disjoint segments

Seven eras that don’t overlap — and don’t agree

The headline samples are perfectly nested, so “positive in six of seven” is one observation restated. Cut into non-overlapping eras, the Sharpe gain concentrates in the 1970s; drawdown improvement holds in 6 of 7.

Comparison
10% gold vs 60/40 within segment
Segments
7 disjoint eras, 36–258 months
Robustness
6-of-7 drawdown survives 9 re-cuts
Honest frame
Drawdown, not Sharpe, is the claim
SegmentMonthsGold %ΔSharpeΔMaxDD ppΔReal pp

ΔSharpe reads +0.69 (1972–74), +0.23 (1975–79), −0.10 (1980s), −0.08 (1990s), +0.07, +0.08. The 1968–71 pegged segment is shown for completeness; this page’s base case starts in 1972.

Exhibit 04 · Equity crises

Eight of nine crises: the sleeve softened the fall

Total return from equity peak to trough. Gold itself fell in two episodes (1980–82, 1998); the 10% sleeve hurt the portfolio in only one — the 1980–82 disinflation that followed gold’s own peak.

Episodes
9 worst equity drawdowns, 1968–2026
Measure
Peak-to-trough total return
Legs
60/40 vs 10% gold, pro-rata
Exception
Nov 1980 – Jul 1982 only
EpisodeMoEquities %Bonds %Gold %60/40 %10% gold %

Equity figures are CRSP total returns — the 1980–82 gap vs an S&P 500 price fall near −24% is +7.5 points of dividends at 5%+ yields. All figures month-end; the COVID daily trough (−34%) is deeper than any monthly print.

Exhibit 05 · Weight choice

The data cannot pick a weight

Median edge over the matched benchmark rises with weight, but so does the worst-sample shortfall — and every weight from 2.5% to 25% sits inside estimation noise. The median edge flattens past 20% while the worst case keeps deteriorating.

Edge
ΔSharpe vs vol-matched, 7 samples
Noise
Weight-specific bootstrap SE
Worst case
1980-start at every weight
Span
2.5–25% inside noise
WeightMedian ΔWorst ΔBoot SEWorst / SESamples +

Median edge is positive at every weight; the worst sample (1980-start) is negative at every weight. “Samples +” counts samples with a positive ΔSharpe: 6 of 7 at every weight except 25% (5 of 7). The SE is computed separately per weight — it scales 0.011 at 2.5% to 0.106 at 25%.

Findings

Executive summary — what the evidence supports

  • No identifiable optimum. The Sharpe-optimal weight ranges 2–33% across the 7 samples, and every weight from 2.5% to 25% sits inside the estimation band against the matched benchmark — median and worst case alike.
  • The risk case is real; the return case is not claimed. At 10% gold, 1972–2026 Sharpe rises 0.54 → 0.59, volatility falls 10.2 → 9.5%, max drawdown improves −29.1 → −21.9%. Real return also rose in this sample — but it falls at 7.5% weight in 3 of 7 samples, so no return uplift is found.
  • The comparison that matters is the matched one. Gold’s edge over plain 60/40 conflates gold with de-risking. Against a volatility-matched de-risk, gold wins all four metrics in 5 of 7 samples — and including a drawdown-matched variant, in 3 of 7.
  • The failure mode is a twenty-year drought. Real drawdown reached 82.9%, troughing March 2001; the January-1980 real peak took 44.8 years to regain on a month-end basis. The 1980–2001 episode cost −1.10pp of CAGR at 7.5% weight and −3.72pp at 25%.
  • Nothing is statistically significant — and that is disclosed, not hidden. Dependence-robust p for the headline test is 0.19 (iid Memmel: 0.078); 0 of 21 tests survive Holm correction. Only the 2000-start sample clears 5% on its own.
  • No substitute beat gold. Across 40 alternatives — TIPS at five maturities, commodities, managed futures, miners, silver, platinum, long Treasuries, bitcoin — none beat gold on Sharpe over its own overlap window.
  • Low hurdle, modest cost sensitivity. At a 10% weight the compound break-even is ~1.9% nominal (−0.5% real) under base assumptions — cleared in every sample. A 25bp vehicle leaves the Sharpe gain intact (0.054 vs 0.056); 75bp trims it to 0.048.
Valuation warning. Gold’s real price sits at the 98.9th percentile of its own history, ~40% above the 1980 real peak. A four-quarter phase-in cuts first-year dispersion 37% but costs 4.4pp of average first-year return.

Sources & methodology

  • Gold: LBMA Gold Price AM daily fixings, Jan 1968 – Aug 2026 — a three-regime splice (telephone fixing pre-20 Mar 2015, ICE electronic auction after). London was closed 15–31 Mar 1968; that month’s 0.0% is a closed-market artefact. PM fixings start Apr 1968; the choice changes nothing.
  • Equities: CRSP value-weighted total return (Ken French library), Jul 1926 – Jun 2026. In the 2000s this runs 0.63pp/yr above the S&P 500 — small/mid-cap breadth, not an error.
  • Bonds: synthetic 10y constant-maturity total return repriced monthly from FRED DGS10 (par bond, semi-annual coupons; validated against IEF at 0.990 correlation).
  • Inflation & cash: CPI-U NSA (Oct 2025 interpolated across the shutdown) and 1-month T-bills. Common monthly sample: 701 observations, Feb 1968 – Jun 2026.
  • Verification: an adversarial audit (57 checks: independent re-implementation, alignment, external validation) reports 48 PASS, 9 FLAG, 0 FAIL — the flags mark known limitations, not errors.