Checkpointing without snapshots: recovering a Linux box when LVM won't help you
Every guide to “safely upgrading your self-hosted stack” starts the same way: take a snapshot first. LVM thin snapshot, Btrfs subvolume, ZFS dataset — pick your flavour, roll back in a second if the upgrade eats your database. It’s genuinely great advice. It’s also advice I couldn’t take on the mini PC that runs my self-hosted services, and it took me an embarrassingly long time to notice.
I found out at the worst possible moment, which is to say five minutes before a container image bump I wasn’t confident about. I went to carve out a snapshot volume and got this:
$ sudo vgs --units g
VG #PV #LV #SN Attr VSize VFree
ubuntu-vg 1 1 0 wz--n- 473.87g 0g
VFree 0. Ubuntu’s guided installer had handed the whole physical volume to a single ext4 logical volume, which is a perfectly reasonable default and completely fatal to snapshots. LVM snapshots need unallocated extents in the volume group to hold the copy-on-write data. There were none. Not “not many” — zero. And you can’t shrink a mounted ext4 filesystem to get some back, so the fix was going to involve a live USB and an afternoon I didn’t have.
This turns out to be a very common way to end up snapshot-less, and it’s not the only one. If you’re on LVM thin provisioning, your thin pool can fill and start rejecting new snapshots. If you’re on a cloud provider’s block storage, volume snapshots may be a paid API you can’t call from inside the guest. Plenty of NAS boxes and ARM SBCs ship ext4 on plain partitions with no volume manager underneath at all. And Docker’s default local volume driver has never had a snapshot concept — a named volume is just a directory under /var/lib/docker/volumes, and whatever the host filesystem can do is what you get.
So I did the boring thing instead, and I’d argue it’s underrated: I built a copy-based checkpoint tool. Not a backup system — I already had one of those, running nightly to a different machine. This is something else. It’s the “I’m about to do something stupid, let me put a bookmark here” button.
The model: a tagged copy plus a manifest
The whole idea fits in one sentence. A checkpoint is a compressed copy of some state, plus a small JSON manifest describing it, appended to a catalog. A restore verifies the copy against the manifest, then puts it back.
That second half is the part people skip, and it’s the part that matters. A directory full of appdata.tar.gz.2026-01-14.bak files is not a recovery system; it’s a pile of tarballs you’ll be afraid to trust at 1 a.m. The manifest is what turns “I think this one’s good” into “this one is byte-for-byte what was there at 19:58 UTC on the 18th, here’s the hash, and the restore path refuses to run if it doesn’t match.”
Here’s what one looks like:
{
"id": "20260818T195820Z-049328b6",
"kind": "dir",
"source": "/srv/appdata/webapp",
"description": "pre-upgrade, v1.2.3 known good",
"created": "2026-08-18T19:58:20Z",
"host": "host01",
"archive": "data.tar.zst",
"sha256": "aa6554ded8247adaf53d2c98e417dcb23f3226c9b21b688c7679f00f060e9332",
"bytes": 169,
"entries": 4
}
Nothing clever. It’s jq -n writing a file next to the archive, and the same JSON appended as one line to catalog.jsonl. Append-only JSON Lines is exactly the right shape here: jq can read it, tail can read it, and a partially-written line can’t corrupt the entries before it.
The commands
Three ways in, because self-hosted state comes in three shapes: bind-mounted directories, individual config files, and Docker named volumes.
ckpt checkpoint-dir /srv/appdata/webapp "pre-upgrade, v1.2.3 known good"
ckpt checkpoint-file /srv/webapp/app.conf "known-good config"
ckpt checkpoint-volume webapp_pgdata "before schema migration"
The directory case is the one you’ll use most, and it’s four lines of actual work:
id=$(date -u +%Y%m%dT%H%M%SZ)-$(head -c4 /dev/urandom | od -An -tx1 | tr -d ' \n')
mkdir -p "$STORE/$id"
tar -C "$src" --numeric-owner -cf - . | zstd -q -3 -o "$STORE/$id/data.tar.zst"
# ...then hash the archive and write the manifest
The volume case needs a helper container, because on rootless Docker you can’t read /var/lib/docker/volumes from the host anyway:
docker run --rm -v "$vol":/src:ro alpine:3.22 \
tar -C /src --numeric-owner -cf - . | zstd -q -3 -o "$STORE/$id/data.tar.zst"
Then the read side:
$ ckpt list
20260818T195820Z-049328b6 dir 2026-08-18T19:58:20Z 169 /srv/appdata/webapp pre-upgrade, v1.2.3 known good
$ ckpt show 20260818T195820Z-049328b6
{ ...manifest... }
integrity: OK
show re-hashes the archive on disk every time. It’s cheap, and it means bit rot in the store surfaces when you look at a checkpoint rather than when you try to use it.
Restore verifies first, then stages, then swaps:
want=$(jq -r .sha256 "$m")
have=$(sha256sum "$STORE/$id/data.tar.zst" | cut -d' ' -f1)
[[ $want == "$have" ]] || die "checksum mismatch for $id - refusing to restore"
stage=$(mktemp -d "${dest%/}.ckpt-stage.XXXX")
zstd -dc "$STORE/$id/data.tar.zst" | tar -C "$stage" --numeric-owner -xf -
mv "$dest" "${dest}.ckpt-old.$$"
mv "$stage" "$dest"
rm -rf "${dest}.ckpt-old.$$"
The staging directory is deliberately a sibling of the target, not somewhere under /tmp. On this box /tmp is tmpfs, so staging there would make the final mv a full cross-filesystem copy of everything you just extracted, at exactly the moment you least want a slow operation. As a sibling it’s a rename: atomic, instant, and it means the live directory is never in a half-extracted state.
I know that’s the right design because I got it wrong the first time and watched it eat a volume.
Three things that bit me
Busybox tar can’t read zstd, and my restore deleted first. My original volume restore mounted the store into the helper container and ran rm -rf /dst/* && tar -C /dst -xf /in/data.tar.zst. The rm succeeded. The tar did this:
tar: invalid tar magic
Alpine’s busybox tar handles gzip, bzip2 and xz, but not zstd. So I had a wiped volume and an archive it couldn’t read. In a scratch environment that’s a funny five minutes; on real state it’s the exact disaster the tool exists to prevent. Two fixes, both necessary: decompress with the host’s zstd and pipe the plain tar stream into the container over stdin, and never delete anything until the new copy is fully extracted. The volume path now extracts into /dst/.ckpt-stage, then clears the old contents, then moves the staged files into place.
A restore can “succeed” and still leave the service broken. Restore a root-owned archive as a non-root user and GNU tar does not warn you, does not fail, and exits 0:
$ zstd -dc data.tar.zst | tar -C ./restore --numeric-owner -xf - ; echo "exit=$?"
exit=0
$ ls -ln ./restore
-rw-r--r-- 1 1001 1001 3 Aug 18 20:56 marker
drwxr-xr-x 2 1001 1001 4096 Aug 18 20:56 sub
Those files were root/root in the archive. They came back as 1001/1001, silently, and a service that checks ownership on its data directory will refuse to start with an error that points nowhere near the restore. Ownership only survives if the process doing the extraction can actually chown — so restore volumes through a container running as root, restore host paths as root, and if your tool can’t, make it say so loudly rather than exit 0.
My first corruption test passed for the wrong reason. I flipped a byte in an archive to prove the checksum guard worked, and show cheerfully reported integrity: OK. Ten minutes of suspicion later:
$ dd if=data.tar.zst bs=1 skip=100 count=1 status=none | od -An -tx1
00
I’d written a zero over a byte that was already zero. Writing \xde\xad\xbe\xef at a different offset gave the result I wanted — integrity: FAILED, restore refusing to run, live data untouched — but the lesson stuck: a negative test that passes on the first attempt deserves more scrutiny than one that fails.
While I was in there I found something genuinely useful. tar embeds mtimes and directory order, so two archives of identical content normally hash differently:
a24697f855236fdbb2295b025458a574c32be972bff6228afc221d0cbc0408e5 -
a0b4a6bf3a71bacc088d2887aaf1c0ab9fa2d4ae77135ffce5160b8fc87f1d80 -
Add --sort=name --mtime='@0' and they don’t:
a21894c804db61907c6b5193d305b5cf0bb86c8b59e19413e67ae546f75cc2e5 -
a21894c804db61907c6b5193d305b5cf0bb86c8b59e19413e67ae546f75cc2e5 -
Now the manifest hash is a content fingerprint. Two checkpoints with the same hash mean the state genuinely didn’t change between them, which makes “did that restart actually modify anything?” a jq query instead of a diff.
What this costs you
I want to be straight about the trade-offs, because copy-based checkpointing is a downgrade from snapshots in every dimension except availability.
It’s slower. A snapshot is metadata; this reads and writes every byte. A 213 MB directory of mixed compressible and incompressible data took 1.5 seconds to checkpoint and 1.2 seconds to restore on a low-power N100 with warm page cache, producing a 173 MB archive. That’s fine. Scale it to a 40 GB media database and you’re looking at minutes, on a spinning disk possibly many minutes, with the service stopped the whole time.
It needs roughly 2× the space. The archive above was 81% of the source size, and that’s with zstd doing real work — for already-compressed data, assume 1:1. Snapshots only cost you the delta.
It is not continuous. Nobody is going to run this every fifteen minutes; it’s a deliberate act before a deliberate change. If you want point-in-time recovery from an event you didn’t see coming, you need a real backup system as well. This does not replace one.
And it’s per-target, not whole-system. Checkpointing a database directory while its process is running gets you a torn copy, same as any other file-level copy — stop the container, checkpoint, start it. The tool can’t save you from that and shouldn’t pretend to.
Using it for real
The workflow, in practice, is unglamorous and takes about a minute:
ckpt checkpoint-dir /srv/appdata/webapp "before 2.4.0 image bump"
docker compose stop webapp && docker compose up -d webapp
curl -fsS localhost:8080/healthz && docker compose logs --since 2m webapp
If it’s healthy, do nothing — the checkpoint ages out on its own. If it isn’t, and 2.4.0 has silently migrated the database to a schema 2.3.1 can’t read:
docker compose stop webapp
ckpt restore 20260818T195820Z-049328b6
docker compose up -d webapp
That’s the whole value proposition. Not that recovery is fast — it isn’t, particularly — but that it’s known. The hash matched, the archive extracted, the manifest says exactly which state you’re back at and what it was taken for. Compared with staring at a directory of .bak files trying to remember which upgrade each one preceded, that’s a substantial upgrade in how well I sleep.
Snapshots are still better. If you can partition for them, do. But “the recommended approach isn’t available on this hardware” is a bad reason to have no recovery story at all, and about 150 lines of bash closes most of the gap.
Written by Finner1909. Developed with AI assistance (Claude); the system described, every command shown and all captured output were run on real hardware before publication.
Corrections and questions: open an issue on this site's repository.