Skip to content

Make the question precise.
Make the test repeatable.

Eighteen years of historical data supports our research and system training. The usefulness of an experiment depends on how that history is prepared, divided and evaluated.

Six steps that belong to the same experiment.

01

Define the target

State the economic outcome or strategy question, the forecast horizon, units and decision time. A first-release estimate and an eventual revised value are separate targets.

02

Freeze the eligible data

Identify the input sources and the information available at the decision cutoff. Keep the dataset snapshot and its coverage limitations with the experiment.

03

Prepare and fit

Fit transformations and model parameters on the development sample. Record preprocessing, missing-value handling and selection decisions.

04

Evaluate forward

Test the model or rules on later observations. Keep the primary evaluation separate from the data used to select the version being assessed.

05

Investigate sensitivity

Inspect behaviour across periods and reasonable alternative assumptions. For strategies, include costs, exposure and execution assumptions.

06

Review and version

Retain the experiment, comparisons and findings. A model change should lead to a new version that can be evaluated against an explicit baseline.

Avoid letting the future into the sample.

01

Publication timing

Eligibility is determined by release availability, not simply the observation period. Later publications belong to a later decision.

02

Revision history

Preserve which data vintage an experiment used. Later revisions are separate evidence and should not overwrite the earlier research record.

03

Preprocessing

Normalisation, imputation and feature selection should be fitted using eligible development data, then applied to later evaluation inputs.

04

Coverage gaps

Missing first-release history and inconsistent definitions should be documented. A series can have less usable history than the wider research dataset.

Train earlier. Evaluate later.

Window design depends on the experiment. Explore two common ways of moving forward without mixing the illustrated evaluation sample into fitting.

Move forward through time.

Period 1Period 2Period 3Period 4Period 5
Test 1TrainEvaluate———
Test 2TrainTrainEvaluate——
Test 3TrainTrainTrainEvaluate—

The training history grows at each step. Evaluation always uses a later period, with preprocessing fitted only on the eligible training sample.

A workflow illustration, not a report of actual model results. Production experiments need source-specific timing rules and suitable gaps between samples.

ECONOMIC FORECASTS

Evaluate the defined release target.

Keep forecast origin, horizon, issued estimate and outcome together. Our preferred primary contract uses the first official release for training labels and scoring, with later revisions studied separately.

TRADING SYSTEMS

Evaluate the full rule set.

Specify the instrument, decision timing, position rules, exits and modelled execution costs. Inspect sample size, behaviour across periods and sensitivity alongside the aggregate result.

Keep enough detail to reconstruct a result.

Question

Target, horizon, units and decision cutoff

Dataset

Sources, snapshot, coverage and vintage assumptions

Model

Version, parameters and fitted transformations

Evaluation

Sample dates, baseline and scoring definition

Execution

Costs, fills, exposure and operational assumptions

Review

Comparisons, findings, limitations and next steps

These principles describe our research approach and development objectives. They do not imply that every model, data source or execution feature has completed validation.