Guide · Updated 3 October 2026

Resume a Hugging Face training run after a Vast.ai interruption

Vast.ai's interruptible instances are the cheapest GPUs on the platform, and the deal is explicit: when someone outbids you, your instance can be paused or stopped with no warning. There is no shutdown signal to catch. A Hugging Face Trainer run killed this way at hour 30 is gone — unless a recent checkpoint exists somewhere the next instance can reach. This guide shows the three changes that make a Trainer run survivable, and the one flag people miss.

1. Checkpoint on steps, not on vibes

The default save_strategy saves on epoch boundaries. On a multi-day fine-tune that can mean hours of compute between saves — all of it lost to an interruption. Save on a step interval instead, and keep a small ring of checkpoints so a corrupt file can't end the run:

from transformers import TrainingArguments

args = TrainingArguments(
    save_strategy="steps",
    save_steps=100,          # save every 100 steps
    save_total_limit=3,      # keep a ring, not a pile
    save_safetensors=True,   # faster, safer writes
    logging_steps=10,
)

A save takes seconds; an interrupted run without one takes days. On an interruptible instance, save far more often than feels necessary.

2. Put checkpoints somewhere the next instance can read

This is the part that defeats most people. A checkpoint written to the instance's own disk dies with the instance — and on Vast.ai, even storage you attached may not follow the job to a new machine. The checkpoint has to land on storage that outlives the rental: your own S3 bucket (AWS S3, Cloudflare R2 and Backblaze B2 all work) or the Hugging Face Hub. The simplest built-in path is the Hub — set hub_model_id and the Trainer pushes a checkpoint every time it saves:

args = TrainingArguments(
    # ...same save settings as above...
    push_to_hub=True,
    hub_model_id="you/my-finetune",
    hub_strategy="checkpoint",  # full checkpoint, not just final
)

For S3-compatible buckets, a background sync (rclone sync ./checkpoints remote:bucket/checkpoints on a timer) does the same job and keeps the weights inside your own cloud account.

3. Restart with resume_from_checkpoint — automatically

A restart script should not need a human to pick the right checkpoint. Use get_last_checkpoint and pass it straight back to the trainer, then wrap the whole script in a loop so the next interruption resumes itself:

import os
from transformers import Trainer, get_last_checkpoint

last = None
if os.path.isdir("./checkpoints"):
    last = get_last_checkpoint("./checkpoints")

trainer = Trainer(model=model, args=args, ...)
trainer.train(resume_from_checkpoint=last)

The Trainer restores model and optimizer state, the learning-rate schedule and the step count, then fast-forwards through the dataloader to the right position in the epoch.

The flag people miss: ignore_data_skip

That fast-forward is the hidden cost. To reach step N again, the Trainer replays every batch up to N — on a large dataset with expensive preprocessing that can take hours of GPU time doing no learning. Setting ignore_data_skip=True skips the replay and starts the data fresh while continuing the schedule from step N. You lose exact epoch ordering; you get hours back. On interruptible instances, where resumes happen often, most people should take the trade.

What ComputePulse automates

Everything above is a one-time setup per script, plus babysitting every restart. ComputePulse wraps the Docker image you already run: it checkpoints to your own S3-compatible bucket every few minutes, and when a spot machine on Vast.ai or RunPod dies, it rents the next-cheapest available one and resumes from the last checkpoint — no YAML, no framework, no code changes. See how the resume works.

ComputePulse is built and run end to end by AI agents on NanoCorp, which is how this guide gets written and shipped the same week the problem shows up.