Restart / Resume System#
Fault-tolerant training that resumes exactly where it left off. Set
save_restart_every: 500 in your YAML (or TrainingConfig) and an interrupted run
automatically resumes from the last snapshot the next time it’s launched — no code
changes needed.
How resumption is gated
Solvers check the done flag in meta.json. If done: false (the run was
interrupted), the snapshot is restored automatically — regardless of config
changes. Use python -m underPINN resume config.yaml to verify config integrity
before resuming a completed run. To force a fresh start, delete <out_dir>/restart/
manually.
How it works#
Every save_restart_every epochs, RestartManager writes params.msgpack,
opt_state.msgpack, hists.npz, and meta.json to <out_dir>/restart/.
done and resumesOn re-run, RestartManager reads meta.json. If done: false, params, optimizer
state, and loss histories are restored — training continues from the saved epoch.
Plots stay continuous across restarts.
resumeSolvers do not hash-check configs by default — an interrupted run resumes even if
you changed lr, epochs, etc. Run python -m underPINN resume config.yaml to
detect config drift before resuming.
After training finishes — normally or via early stopping — done() writes
"done": true to meta.json. The next run with the same config starts fresh
instead of re-resuming a completed run.
Snapshot directory contents#
File |
Contents |
|---|---|
|
Flax-serialised model parameters at the snapshot epoch |
|
Flax-serialised optimizer state (Adam moments, step count) |
|
All loss history arrays accumulated so far ( |
|
|
Configuration#
training:
save_restart_every: 500 # 0 to disable
config = TrainingConfig(
epochs = 10000,
out_dir = "outputs/burgers",
save_restart_every = 500,
)
solver.train(*data, config=config)
# If killed at epoch 3700, the next run resumes
# from epoch 3500 (last snapshot) automatically.
Verifying config changes before resuming#
python -m underPINN resume examples/burgers/config.yaml
resume computes the MD5 of the current YAML, compares it against the hash stored in
meta.json, and warns you if any field changed since the last snapshot (learning rate,
epoch count, network layers, physics parameters, …). If everything is consistent, it
resets done to false so the next run continues training.
Tip
To force a fresh start without the resume command, simply delete
<out_dir>/restart/, or set save_restart_every: 0 temporarily.
See also
Model Checkpointing & Inference covers the separate, longer-lived params.msgpack artifact that
every runner writes on successful completion — distinct from the in-progress
restart snapshot described here.