Training# Run recipes, change them, run them on more GPUs and clusters, and monitor your runs. Train and pretrain Pick a recipe Check it on a few events Keep a variant as a child config Start the run From a notebook Configs and overrides Experiment configs Overrides on the command line Precedence Common experiment fields Check before running Scale up More GPUs on one machine Run on Slurm Interactive allocations and chains Launch recipes exex Checkpoints and resume What’s saved Warm start and resume Resume a run Changing the number of GPUs model_best.pth Errors Evaluate During training On a finished run Monitor and debug Weights & Biases TensorBoard Diagnostic hooks Structured traces Troubleshooting Installation Data Launcher and Slurm Training Checkpoints and evaluation Reporting a problem