# Trading Statistics and Uncertainty

Interpret performance estimates with appropriate uncertainty.

Use this workbook alongside the course. Write your answers before opening the solutions. Practical work is self-reviewed; scored knowledge checks are in the Academy.

## 1. Sample properties

### Returns versus trades

A trade sample must define its unit: completed trade, daily return or another observation. Mixing units changes the meaning of averages and uncertainty. State inclusion rules and the treatment of overlapping positions before calculation.

### Skew and tails

Returns can be skewed and heavy-tailed relative to simple symmetric models. A few large observations may dominate the mean. Inspect the distribution and extreme cases rather than relying only on average and standard deviation.

### Serial dependence

Serial dependence means nearby observations can contain related information. Ten overlapping trades are not necessarily ten independent experiments. An uncertainty method that assumes independence may be inappropriate without examining the sample structure.

### Selection bias

Stationarity assumptions concern whether the process is sufficiently stable for the intended inference. Market behavior can change. A historical estimate is conditional evidence, not a permanent property of a strategy.

### Worked example

A strategy produces ten entries during one event and all move together. Treating them as ten independent confirmations overstates how much distinct evidence the event provides.

### Independent exercise

Choose a sampling unit for that case and explain how you would retain the underlying trades without claiming independence.

My inputs and assumptions:

My calculation or decision:

Evidence that would change my conclusion:


## 2. Estimates

### Expectancy intervals

A point estimate summarizes a sample, such as mean net return. It is not the true future value known with certainty. Pair it with the sample definition and an appropriate uncertainty measure.

### Win-rate intervals

A confidence interval is produced by a procedure under assumptions. Its frequentist coverage describes repeated applications of that procedure, not a guarantee that a particular trading outcome will fall inside it. State the model and dependence assumptions.

### Bootstrap limitations

Bootstrap resampling approximates variability from the observed data under a resampling scheme. Resampling individual observations assumes a structure that may be unsuitable for dependent time series; blocks can preserve some dependence but introduce choices of their own.

### Drawdown distribution

Quantiles describe distribution thresholds and can be unstable in small samples, especially in the tails. A historical worst loss is only the worst observed loss, not a proven bound on future loss.

### Worked example

An estimated mean is positive but its uncertainty interval spans negative values under a defensible method. The sample does not establish a positive mean with the chosen level of confidence merely because the point estimate is above zero.

### Independent exercise

Write a conclusion for that result without converting uncertainty into either a guaranteed edge or proof that the strategy can never work.

My inputs and assumptions:

My calculation or decision:

Evidence that would change my conclusion:


## 3. Decisions

### Practical significance

A decision threshold should relate to the practical question, such as whether net benefit survives costs and uncertainty. Statistical significance alone does not establish economic usefulness or executable scale.

### Multiple testing

Testing many variants increases the chance of selecting an apparently strong result by chance. Preserve the search history and use a validation design that accounts for selection. A single final p-value cannot erase an undisclosed search.

### Sensitivity

Effect size measures the magnitude of the difference, while uncertainty describes precision. A tiny estimated advantage may be irrelevant after implementation costs even if a large sample makes it statistically distinguishable from zero.

### Evidence thresholds

A no-go decision is a legitimate research result. Define what evidence would reject the hypothesis before inspection, and record negative findings. This reduces pressure to keep changing the analysis until it produces an attractive answer.

### Worked example

A rule shows a tiny gross advantage of 0.02 per trade while realistic costs are 0.10. Even a precisely estimated gross advantage would not establish positive net economics.

### Independent exercise

Write a decision rule that considers both net effect and uncertainty. State how repeated variant searches should be reported.

My inputs and assumptions:

My calculation or decision:

Evidence that would change my conclusion:


## 4. Statistical reporting workshop

### Show confidence intervals with point estimates

Show estimates and intervals together with the sample size, unit and method. Avoid excessive decimal precision that suggests more certainty than the observations support. Make transformations and exclusions reproducible.

### Check serial dependence

Inspect serial dependence and changes over time before choosing a simple uncertainty formula. Diagnostics do not prove an ideal model, but they can reveal when its assumptions are visibly questionable.

### Document all tested variants

Document all tested versions, including abandoned ones. A final report should distinguish exploratory results from a frozen evaluation. Otherwise the reader cannot assess how much the reported result benefited from selection.

### Explain what the sample cannot establish

State what the data cannot establish: future stability, unobserved execution quality or causation may remain unresolved. A limitation is part of the result, not an optional disclaimer after the conclusion.

### Worked example

A report shows a mean to six decimals from twelve highly variable trades but no interval or search history. Its numerical precision does not imply evidential precision.

### Independent exercise

Rewrite the report's minimum results table and include a field that exposes whether the study was exploratory or confirmatory.

My inputs and assumptions:

My calculation or decision:

Evidence that would change my conclusion:


## Course project

Write a performance report with confidence intervals and dependence caveats.

### Self-review rubric

- Concepts and reasoning: 25%
- Calculations, data and evidence: 30%
- Process and risk controls: 25%
- Limitations and communication: 20%

Record one correction and one next practice task. This rubric is not automatically graded.

## Worked solutions

### Exercise 1

An event-level or time-block analysis may be useful depending on the question. Preserve individual trades for accounting, identify their shared event and use an uncertainty method that respects clustering rather than simply inflating the observation count.

### Exercise 2

Report the estimate, interval, method and sample limitations. Explain that the data does not separate the mean from nonpositive values under those assumptions; additional evidence or a different justified design may change the assessment.

### Exercise 3

Require the predefined net measure to meet the stated practical criterion under an appropriate uncertainty analysis, and disclose all material variants and selection stages. The exact criterion belongs to the research protocol, not a retrospective justification.

### Exercise 4

Include sample definition and count, net estimate, dispersion or interval with method, data period, costs, number of variants and evaluation status. Explain any dependence or instability that limits interpretation.

## Further reading

- https://www.itl.nist.gov/div898/handbook/eda/section3/eda35c.htm
- https://www.itl.nist.gov/div898/handbook/eda/section4/eda4232.htm
