The source files can be found in examples/.
Train/test splitting
The cheapest honest question you can ask of a portfolio is: does it survive data it has never seen? Cross-validation answers that many times over, on many folds, at many times the cost. Before any of that, there is the holdout split — train on the first 80 % of history, score on the last 20 %, once.
It is the protocol everyone reaches for first, and the one easiest to get quietly wrong. Not because the split itself is hard, but because everything upstream of it has to respect it. An imputation fill value computed from the whole price history has seen the test window. So has a missing-data filter that chose the asset universe by looking at all five years. The returns matrix you split was already contaminated before you split it.
TrainTestSplit (alias TTS) closes that hole by making the split a step of the pipeline — the first step, before any other step has touched the data. Everything fitted downstream sees the training window and nothing else, by construction rather than by discipline.
This example covers the free function and its sizing rules, the pipeline step and the leakage argument that pins it to first position, the payoff (fit_predict scoring on the held-out window), the embargo, and why a pipeline carrying a holdout is refused by cross-validation.
See docs/adr/0031-holdout-split-as-a-pipeline-step.md for the design rationale.
using PortfolioOptimisers, PrettyTables
resfmt = (v, i, j) -> begin
if j == 1
return v
else
return isa(v, Number) ? "$(round(v*100, digits=3)) %" : v
end
end;1. Setting up
Four years of daily prices for ten S&P 500 names. As in the pipelines example we start from prices, not returns, because the whole point is that the split happens before the data is cleaned and converted.
using CSV, TimeSeries, DataFrames, Clarabel, Statistics, Dates
X = TimeArray(CSV.File(joinpath(@__DIR__, "..", "SP500.csv.gz")); timestamp = :Date)[(end - 252 * 4):end]
X = X[Symbol.(["AAPL", "MSFT", "JNJ", "JPM", "XOM", "PG", "KO", "WMT", "PEP", "MRK"])...]
pr = PricesResult(; X = X)
slv = Solver(; name = :clarabel, solver = Clarabel.Optimizer,
settings = Dict("verbose" => false),
check_sol = (; allow_local = true, allow_almost = true));
size(values(X))(1009, 10)2. The free function
train_test_split cuts price- or returns-level data into two windows: the training window is the head of the observations, the test window is the tail. Time-ordered, so the test window is always the most recent data — a random shuffle would let the model train on tomorrow to predict yesterday.
The windows are views, so nothing is copied.
train, test = train_test_split(pr; test_size = 0.2)
DataFrame(; window = ["train", "test"],
rows = [size(values(train.X), 1), size(values(test.X), 1)],
from = [first(timestamp(train.X)), first(timestamp(test.X))],
to = [last(timestamp(train.X)), last(timestamp(test.X))])| Row | window | rows | from | to |
|---|---|---|---|---|
| String | Int64 | Date | Date | |
| 1 | train | 808 | 2018-12-27 | 2022-03-11 |
| 2 | test | 201 | 2022-03-14 | 2022-12-28 |
2.1 Sizing: counts, fractions, and complements
A size is either an Integer row count or an AbstractFloat fraction of the observations in (0, 1). test_size = 250 means the last 250 rows; test_size = 0.25 means the last quarter.
Give one side and the other is its complement — the two windows partition the data. Give neither and the split is 75/25. The four spellings below all describe the same idea.
N = size(values(X), 1)
spellings = ["train_test_split(pr)" => (nothing, nothing),
"train_test_split(pr; test_size = 0.2)" => (nothing, 0.2),
"train_test_split(pr; train_size = 0.8)" => (0.8, nothing),
"train_test_split(pr; test_size = 202)" => (nothing, 202)]
sizes = DataFrame(; spelling = String[], train = Int[], test = Int[])
for (label, (tr_s, te_s)) in spellings
a, b = train_test_split(pr; train_size = tr_s, test_size = te_s)
push!(sizes, (label, size(values(a.X), 1), size(values(b.X), 1)))
end
sizes| Row | spelling | train | test |
|---|---|---|---|
| String | Int64 | Int64 | |
| 1 | train_test_split(pr) | 756 | 253 |
| 2 | train_test_split(pr; test_size = 0.2) | 808 | 201 |
| 3 | train_test_split(pr; train_size = 0.8) | 807 | 202 |
| 4 | train_test_split(pr; test_size = 202) | 807 | 202 |
2.2 The embargo
Give both sizes and they need not cover the data. The head supplies the training rows, the tail supplies the test rows, and anything in between is embargoed — it belongs to neither window.
That gap is not waste, it is insurance. Financial features are autocorrelated: a return computed over a 20-day window straddling the boundary leaks test information into the last training row. Dropping the rows around the seam severs that link. (This is the same idea as KFold's purged_size/embargo_size, expressed for a single split.)
Windows that would overlap are rejected outright, as is any split that would leave a window empty.
tr_e, te_e = train_test_split(pr; train_size = 0.6, test_size = 0.2)
DataFrame(; window = ["train", "embargoed", "test"],
rows = [size(values(tr_e.X), 1),
N - size(values(tr_e.X), 1) - size(values(te_e.X), 1),
size(values(te_e.X), 1)])
# Overlapping windows are a mistake, not a preference:
try
train_test_split(pr; train_size = 0.9, test_size = 0.2)
catch e
println(e.msg)
endthe training and test windows of a train/test split must not overlap, but 908 training and 201 test observations exceed the 1009 available; rows between the two windows are embargoed, so their sizes may sum to less than 1009 but never to more3. Why the manual composition is a trap
Nothing stops you from writing this:
res = fit(pipe, pr) ## fitted on EVERYTHING
pred = predict(res, pr, test_window)It runs. It produces a number. And that number is not out-of-sample — the pipeline's missing-data filter chose its universe using the test rows, and its imputer computed fill values from them. The call shape is identical to the honest one; only the data the fit saw differs. Nothing warns you.
The fix is to make the split part of the workflow, so the fit cannot see the test rows.
4. The split as the first pipeline step
TrainTestSplit narrows whichever data slot the pipeline input filled — prices here — and hands the training window to every step downstream. It also stashes the held-out window in its fitted result, which is what makes the evaluation a one-liner later.
Note the auto-generated step name: "split".
pipe = Pipeline(;
steps = (TrainTestSplit(; test_size = 0.2), MissingDataFilter(), Imputer(),
PricesToReturns(), EmpiricalPrior(),
MeanRisk(; r = Variance(),
opt = JuMPOptimiser(; slv = slv, pe = EmpiricalPrior()))))
pipe.names("split", "prices_1", "prices_2", "returns", "prior", "opt")Fitting runs the steps left to right. Everything after the split saw 806 training prices, not the 1009 in the sample.
res = fit(pipe, pr)
split_res = res["split"]
DataFrame(;
quantity = ["prices in the sample", "prices the workflow was fitted on",
"prices held out", "returns reaching the optimiser"],
rows = [N, size(values(split_res.train.X), 1), size(values(split_res.test.X), 1),
size(res.ctx.returns.X, 1)])| Row | quantity | rows |
|---|---|---|
| String | Int64 | |
| 1 | prices in the sample | 1009 |
| 2 | prices the workflow was fitted on | 808 |
| 3 | prices held out | 201 |
| 4 | returns reaching the optimiser | 807 |
4.1 The position rule is enforced
A TrainTestSplit must be the first step. This is not tidiness — it is the entire safety argument. A MissingDataFilter fitted before the split would choose the asset universe using the held-out rows; an Imputer fitted before it would compute its fill values from them. Their fitted state would carry test information into the training workflow, which is exactly the leak the split exists to prevent.
So the constructor refuses, rather than letting you find out from a suspiciously good backtest.
try
Pipeline(;
steps = (MissingDataFilter(), Imputer(), TrainTestSplit(; test_size = 0.2),
PricesToReturns(), EqualWeighted()))
catch e
println(e.msg)
end
# The same applies to a split hidden inside a nested pipeline:
try
inner = Pipeline(; steps = (TrainTestSplit(; test_size = 0.2), PricesToReturns()))
Pipeline(; steps = (inner, EqualWeighted()))
catch e
println(e.msg)
enda TrainTestSplit step must be the first step of a Pipeline, but one appears at step 3; a stateful step fitted before the split would have seen the held-out test rows, leaking them into the fitted workflow
a TrainTestSplit step is nested inside a Pipeline step of a Pipeline; the holdout must be the first step of the outermost Pipeline, where no step has yet touched the data5. The payoff: fit_predict scores on the held-out window
For a pipeline carrying a split, fit_predict fits on the training rows and predicts on the window the split reserved. That is the whole holdout protocol, in one call — and it is out-of-sample by construction, because the split is what defined "sample".
For a pipeline without a split, fit_predict keeps its old meaning and stays in-sample. The same call means different things only because the workflows genuinely differ.
pred_test = fit_predict(pipe, pr)
# The in-sample counterpart: replay the fitted workflow on its own training window.
pred_train = predict(res, split_res.train)PredictionResult
res ┼ MeanRiskResult
│ jr ┼ JuMPOptimisationResult
│ │ pa ┼ ProcessedJuMPOptimiserAttributes
│ │ │ pr ┼ LowOrderPrior
│ │ │ │ X ┼ 807×10 Matrix{Float64}
│ │ │ │ mu ┼ 10-element Vector{Float64}
│ │ │ │ sigma ┼ 10×10 Matrix{Float64}
│ │ │ │ chol ┼ nothing
│ │ │ │ w ┼ nothing
│ │ │ │ ens ┼ nothing
│ │ │ │ kld ┼ nothing
│ │ │ │ ow ┼ nothing
│ │ │ │ rr ┼ nothing
│ │ │ │ f_mu ┼ nothing
│ │ │ │ f_sigma ┼ nothing
│ │ │ │ f_w ┴ nothing
│ │ │ wb ┼ WeightBounds
│ │ │ │ lb ┼ 10-element StepRangeLen{Float64, Base.TwicePrecision{Float64}, Base.TwicePrecision{Float64}, Int64}
│ │ │ │ ub ┴ 10-element StepRangeLen{Float64, Base.TwicePrecision{Float64}, Base.TwicePrecision{Float64}, Int64}
│ │ │ lt ┼ nothing
│ │ │ st ┼ nothing
│ │ │ lcsr ┼ nothing
│ │ │ ctr ┼ nothing
│ │ │ gcardr ┼ nothing
│ │ │ sgcardr ┼ nothing
│ │ │ smtx ┼ nothing
│ │ │ sgmtx ┼ nothing
│ │ │ slt ┼ nothing
│ │ │ sst ┼ nothing
│ │ │ sglt ┼ nothing
│ │ │ sgst ┼ nothing
│ │ │ tn ┼ nothing
│ │ │ fees ┼ nothing
│ │ │ plr ┼ nothing
│ │ │ ret ┼ ArithmeticReturn
│ │ │ │ ucs ┼ nothing
│ │ │ │ lb ┼ nothing
│ │ │ │ mu ┴ 10-element Vector{Float64}
│ │ retcode ┼ OptimisationSuccess
│ │ │ res ┴ Dict{Any, Any}: Dict{Any, Any}()
│ │ sol ┼ JuMPOptimisationSolution
│ │ │ w ┴ 10-element Vector{Float64}
│ │ model ┼ A JuMP Model
│ │ │ ├ solver: Clarabel
│ │ │ ├ objective_sense: MIN_SENSE
│ │ │ │ └ objective_function_type: QuadExpr
│ │ │ ├ num_variables: 11
│ │ │ ├ num_constraints: 4
│ │ │ │ ├ AffExpr in MOI.EqualTo{Float64}: 1
│ │ │ │ ├ Vector{AffExpr} in MOI.Nonnegatives: 1
│ │ │ │ ├ Vector{AffExpr} in MOI.Nonpositives: 1
│ │ │ │ └ Vector{AffExpr} in MOI.SecondOrderCone: 1
│ │ │ └ Names registered in the model
│ │ │ └ :G, :bgt, :dev_1, :dev_1_soc, :k, :lw, :obj_expr, :ret, :risk, :risk_vec, :sc, :so, :variance_flag, :variance_risk_1, :w, :w_lb, :w_ub
│ fb ┴ nothing
rd ┼ PredictionReturnsResult
│ nx ┼ 10-element Vector{String}
│ X ┼ 807-element Vector{Float64}
│ nf ┼ nothing
│ F ┼ nothing
│ nb ┼ nothing
│ B ┼ nothing
│ ts ┼ 807-element Vector{Date}
│ iv ┼ nothing
│ ivpa ┴ nothingBoth are scored on the realised portfolio return series, so the measures must be ones that can read a bare return series: SCM() is the scoreable spelling of standard deviation (Variance and StandardDeviation consume portfolio weights instead — see the asset pre-selection example, which hits the same wall).
DataFrame(; window = ["train (in-sample)", "test (held out)"],
observations = [length(pred_train.rd.X), length(pred_test.rd.X)],
std = [expected_risk(SCM(), pred_train), expected_risk(SCM(), pred_test)],
cvar = [expected_risk(ConditionalValueatRisk(), pred_train),
expected_risk(ConditionalValueatRisk(), pred_test)])| Row | window | observations | std | cvar |
|---|---|---|---|---|
| String | Int64 | Float64 | Float64 | |
| 1 | train (in-sample) | 807 | 0.000121968 | 0.0253435 |
| 2 | test (held out) | 200 | 0.000113586 | 0.0238769 |
The gap between the two rows is the number the whole exercise exists to produce. In-sample risk is an estimate the optimiser minimised; held-out risk is an estimate it merely inherited. The first is optimistic by construction. Only the second is evidence.
# The weights are indexed by the training universe, and the test prediction is aligned to them.
DataFrame(; asset = res.ctx.returns.nx, weight = res.w)| Row | asset | weight |
|---|---|---|
| String | Float64 | |
| 1 | AAPL | 6.55457e-8 |
| 2 | MSFT | 6.53644e-8 |
| 3 | JNJ | 0.227308 |
| 4 | JPM | 5.61597e-8 |
| 5 | XOM | 0.0371455 |
| 6 | PG | 0.0488489 |
| 7 | KO | 0.163248 |
| 8 | WMT | 0.357139 |
| 9 | PEP | 1.14119e-7 |
| 10 | MRK | 0.166309 |
6. Replaying a fitted split on new data is a pass-through
A fitted holdout's rows are a fact about the fitting window. Replaying them on some unseen window would be meaningless, so apply_preprocessing on a TrainTestSplitResult passes the data straight through.
That is what keeps the ordinary workflow — fit on history, predict on genuinely new observations — working. Every other fitted step still replays properly: the training asset universe, the training fill values, the returns conversion. Only the split stands aside.
# Pretend the last 40 rows are data that arrived after the model was built.
future = PricesResult(; X = X[(end - 40):end])
pred_future = predict(res, future)
DataFrame(; source = ["held-out window (fit_predict)", "fresh data (predict)"],
observations = [length(pred_test.rd.X), length(pred_future.rd.X)],
std = [expected_risk(SCM(), pred_test), expected_risk(SCM(), pred_future)])| Row | source | observations | std |
|---|---|---|---|
| String | Int64 | Float64 | |
| 1 | held-out window (fit_predict) | 200 | 0.000113586 |
| 2 | fresh data (predict) | 40 | 7.75352e-5 |
7. One evaluation protocol per call
A holdout and a cross-validator are two answers to the same question, and they do not compose. Cross-validation already defines the train and test window of every fold; a split left in the pipeline would shave a second, redundant holdout off each fold's training data and stash a test window nobody ever reads.
That is silent loss of training data, so the library refuses it rather than quietly doing it.
gscv = GridSearchCrossValidation(Dict("returns" =>
[PricesToReturns(; ret_method = :simple),
PricesToReturns(; ret_method = :log)]);
cv = KFold(; n = 3), r = SCM())
try
search_cross_validation(pipe, gscv, pr)
catch e
println(e.msg)
endthis Pipeline contains a TrainTestSplit step, so it cannot also be cross-validated: cross-validation already defines the train and test window of every fold, and the split would shave a second holdout off each fold's training data. Remove the TrainTestSplit step, or evaluate the pipeline with fit_predict instead of cross-validating it.Drop the split and the same pipeline tunes happily — cross-validation supplies the discipline the split was providing.
pipe_cv = Pipeline(;
steps = (MissingDataFilter(), Imputer(), PricesToReturns(),
EmpiricalPrior(),
MeanRisk(; r = Variance(),
opt = JuMPOptimiser(; slv = slv,
pe = EmpiricalPrior()))))
scv_res = search_cross_validation(pipe_cv, gscv, pr)
scv_res.idx28. Choosing between them
Holdout (TrainTestSplit) | Cross-validation | |
|---|---|---|
| Fits | one | one per fold |
| Answers | does this survive unseen data? | how does this behave across regimes? |
| Test rows | the most recent tail, once | every row, in turn |
| Use it for | a final, honest score on a locked model | selecting hyperparameters, estimating variance of the estimate |
The holdout is a verdict, not a search tool. Tune with cross-validation, and keep the holdout untouched until the model is frozen — the moment you tune against the held-out score, it stops being held out, and you are back to reporting an in-sample number with extra steps.
Summary
train_test_splitcuts data into a head (train) and a tail (test); one size given makes the other its complement, both given embargoes the rows between them.TrainTestSplitmakes that the first step of aPipeline, so every fitted step downstream — universe selection, imputation, prior, optimiser — sees the training window alone. The position rule is enforced at construction.fit_predicton a split-bearing pipeline scores on the held-out window; the split's fitted result carries both windows (res["split"].train,res["split"].test).Replaying a fitted split is a pass-through, so predicting on genuinely new data still works.
A pipeline carrying a holdout cannot also be cross-validated: one evaluation protocol per call.
This page was generated using Literate.jl.