Boosted Trees
Boosted Trees
These are popular ensemble methods that build a sequence of tree models.
Each tree uses the results of the previous tree to better predict samples, especially those that have been poorly predicted.
Each tree in the ensemble is saved, and new samples are predicted using a weighted average of its votes.
We’ll focus on the popular xgboost implementation.
Boosted Tree Tuning Parameters
Some possible parameters:
trees: The number of trees (\([1, \infty]\), but usually up to thousands).
min_n: The number of samples needed to further split (\([1, n]\)).
learn_rate: The rate that each tree adapts from previous iterations (\((0, \infty]\), usual maximum is about 0.1).
stop_iter: The number of iterations of boosting where no improvement was shown before stopping (\([1, trees]\)).
We’ll focus on the learning rate for now.
Your turn
![]()
Open the help page for boost_tree() (via ?boost_tree).
Look at the engine documentation for xgboost models.
Grid search
A small grid of points trying to minimize the error via learning rate:
Grid search
In reality, we would probably sample the space more densely:
Iterative Search
We could start with a few points and search the space:
Specifying the model
# xgboost is the default engine
bst_spec <- boost_tree(trees = 500, learn_rate = tune(), mode = "classification")
bst_wflow <- workflow(class ~ ., bst_spec)
bst_wflow
#> ══ Workflow ══════════════════════════════════════════════════════════
#> Preprocessor: Formula
#> Model: boost_tree()
#>
#> ── Preprocessor ──────────────────────────────────────────────────────
#> class ~ .
#>
#> ── Model ─────────────────────────────────────────────────────────────
#> Boosted Tree Model Specification (classification)
#>
#> Main Arguments:
#> trees = 500
#> learn_rate = tune()
#>
#> Computational engine: xgboost
Working with Parameters
tidymodels provides pre-defined information on tuning parameters (such as their type, range, transformations, etc.).
We can:
- Extract the parameter information
- Update their characteristics
- Give their information to
tune_grid(), etc.
mtry_param <-
boost_tree(mtry = tune()) |>
extract_parameter_set_dials()
mtry_param
#> Collection of 1 parameters for tuning
#>
#> identifier type object
#> mtry mtry nparam[?]
#>
#> Model parameters needing finalization:
#> # Randomly Selected Predictors ('mtry')
#>
#> See `?dials::finalize()` or `?dials::update.parameters()` for more
#> information.
mtry_param |>
update(mtry = mtry(c(1, 10)))
#> Collection of 1 parameters for tuning
#>
#> identifier type object
#> mtry mtry nparam[+]
#>
Different types of grids ![]()
Space-filling designs (SFD) attempt to cover the parameter space without redundant candidates. We recommend these the most, and they are the default.
Try out multiple values
tune_grid() works similar to fit_resamples() but covers multiple parameter values:
set.seed(22)
bst_res <- tune_grid(
bst_wflow,
cls_folds,
grid = 15
)
Compare results
Inspecting results and selecting the best-performing hyperparameter(s):
Saving predictions
Let’s repeat this process but make sure that we keep the out-of-sample predictions
set.seed(22)
bst_res <- tune_grid(
bst_wflow,
cls_folds,
grid = 15,
control = control_grid(save_pred = TRUE)
)
Checking (Approximate) Calibration
![]()
bst_res |>
cal_plot_windowed(
# truth = class,
# estimate = .pred_event,
window_size = 0.2,
step_size = 0.025,
)
Compare results
Which is numerically best?
show_best(bst_res, metric = "brier_class")
#> # A tibble: 5 × 7
#> learn_rate .metric .estimator mean n std_err .config
#> <dbl> <chr> <chr> <dbl> <int> <dbl> <chr>
#> 1 0.00568 brier_class binary 0.0565 10 0.00409 pre0_mod05_post0
#> 2 0.00876 brier_class binary 0.0576 10 0.00450 pre0_mod06_post0
#> 3 0.0110 brier_class binary 0.0584 10 0.00454 pre0_mod07_post0
#> 4 0.00367 brier_class binary 0.0592 10 0.00401 pre0_mod04_post0
#> 5 0.0131 brier_class binary 0.0595 10 0.00474 pre0_mod08_post0
best_parameter <- select_best(bst_res, metric = "brier_class")
best_parameter
#> # A tibble: 1 × 2
#> learn_rate .config
#> <dbl> <chr>
#> 1 0.00568 pre0_mod05_post0
collect_metrics() is also available.
Running in parallel
Grid search, combined with resampling, requires fitting a lot of models!
These models don’t depend on one another and can be run in parallel.
We can use the future or mirai packages to do this:
cores <- parallelly::availableCores(logical = FALSE)
library(future)
plan(multisession, workers = cores)
# Now call `tune_grid()`!
library(mirai)
daemons(cores)
# Now call `tune_grid()`!
We’ll use mirai as our parallel backend for our notes.
Distributing tasks
When only tuning the model:
Running in parallel
Speed-ups are fairly linear up to the number of physical cores (10 here).
Your turn
![]()
Modify your model workflow to tune additional parameters.
Use grid search to find the best parameter(s).
The whole game - status update
The final fit ![]()
Suppose that we are happy with our boosted tree model.
Let’s fit the model on the training set and verify our performance using the test set.
We’ve shown you fit() and predict() (+ augment()) but there is a shortcut:
bst_final_wflow <- finalize_workflow(bst_wflow, best_parameter)
# cls_split has train + test info
set.seed(690)
final_res <- last_fit(bst_final_wflow, cls_split)
final_res
#> # Resampling results
#> # Manual resampling
#> # A tibble: 1 × 6
#> splits id .metrics .notes .predictions .workflow
#> <list> <chr> <list> <list> <list> <list>
#> 1 <split [800/200]> train/test split <tibble> <tibble> <tibble> <workflow>
What is in final_res? ![]()
collect_metrics(final_res)
#> # A tibble: 3 × 4
#> .metric .estimator .estimate .config
#> <chr> <chr> <dbl> <chr>
#> 1 accuracy binary 0.905 pre0_mod0_post0
#> 2 roc_auc binary 0.957 pre0_mod0_post0
#> 3 brier_class binary 0.0697 pre0_mod0_post0
These are metrics computed with the test set
What is in final_res? ![]()
collect_predictions(final_res)
#> # A tibble: 200 × 7
#> .pred_class .pred_class_1 .pred_class_2 id class .row .config
#> <fct> <dbl> <dbl> <chr> <fct> <int> <chr>
#> 1 class_1 0.977 0.0233 train/test split class… 1 pre0_m…
#> 2 class_2 0.164 0.836 train/test split class… 3 pre0_m…
#> 3 class_1 0.977 0.0233 train/test split class… 7 pre0_m…
#> 4 class_1 0.977 0.0233 train/test split class… 9 pre0_m…
#> 5 class_2 0.0460 0.954 train/test split class… 12 pre0_m…
#> 6 class_1 0.924 0.0760 train/test split class… 22 pre0_m…
#> 7 class_2 0.0460 0.954 train/test split class… 25 pre0_m…
#> 8 class_1 0.530 0.470 train/test split class… 27 pre0_m…
#> 9 class_1 0.977 0.0233 train/test split class… 28 pre0_m…
#> 10 class_1 0.977 0.0233 train/test split class… 32 pre0_m…
#> # ℹ 190 more rows
What is in final_res?
![]()
final_res |>
cal_plot_windowed(window_size = 0.2, step_size = 0.025) +
geom_rug()
What is in final_res? ![]()
extract_workflow(final_res)
#> ══ Workflow [trained] ════════════════════════════════════════════════
#> Preprocessor: Formula
#> Model: boost_tree()
#>
#> ── Preprocessor ──────────────────────────────────────────────────────
#> class ~ .
#>
#> ── Model ─────────────────────────────────────────────────────────────
#> ##### xgb.Booster
#> call:
#> xgboost::xgb.train(params = list(eta = 0.0056848552336798, max_depth = 6,
#> gamma = 0, colsample_bytree = 1, colsample_bynode = 1, min_child_weight = 1,
#> subsample = 1, nthread = 1, objective = "binary:logistic"),
#> data = x$data, nrounds = 500, evals = x$watchlist, verbose = 0)
#> # of features: 2
#> # of rounds: 500
#> callbacks:
#> evaluation_log
#> evaluation_log:
#> iter training_logloss
#> <num> <num>
#> 1 0.6313728
#> 2 0.6268702
#> --- ---
#> 499 0.1109659
#> 500 0.1107976
Use this for prediction on new data, like for deploying
Class Boundary
The whole game