Take repeated bootstrap samples from training set . Train many copies of a high-variance base learner (like decision trees) on resampled datasets and average their predictions to reduce variance.
Bootstrap sampling
Given set containing training examples, create by drawing examples at random with replacement from .
i.e. Averaging predictors trained on independent datasets of size reduces variance by a factor of . We don’t have independent datasets, so we cheat and resample with replacement from the one we have. Not truly independent, but works well in practice.
Example: from :
Each bootstrap sample contains of the unique original points. The remaining “out-of-bag” points give you a free held-out set per tree: you can evaluate tree on the points it never saw.
Since the samples aren’t independent, the variance doesn’t literally drop by , but empirically it still drops substantially.
Algorithm
-
Bootstrap-sample datasets of size
-
Train a classifier on each
-
Aggregate:
- Regression:
- Classification: