Watchdogs That Actually Restart Things: Supervising Long-Running Processes on Linux
flock, pidfiles with /proc verification, and pgrep patterns that don't match your own shell — patterns that keep 24/7 data pipelines alive on a Linux VM.
A watchdog that crashes, starts duplicates, or kills the wrong process is worse than no watchdog at all. I keep a few 24/7 data pipelines running on a small Linux virtual machine — chart renderers and FFmpeg encoders feeding live streams — and the supervision around them has taught me more than the pipelines themselves. These are the patterns that survived contact with reality.
1. One instance at a time: flock
A watchdog that runs every few minutes must never overlap with itself: a run that overruns its interval would otherwise start a second copy, then a third. The fix is flock -n on a lockfile — the second instance exits silently if the lock is held. One line, and an entire class of duplicate-start bugs disappears.
# one instance at a time — a late second run exits silently
flock -n /var/lock/stream-watchdog.lock -c /opt/pipeline/watchdog.sh2. Restart only what is missing — never kill first
The watchdog's job is to check each component (tunnel, supervisors, renderers) and start the ones that are absent. It must never kill anything as part of a restart: when two streams share one tunnel process, a supervisor that kills "its" tunnel on restart takes down the other stream too — which then restarts, kills the first, and the two supervisors spend the night assassinating each other. Start what's missing, leave the rest alone.
3. Verify the PID is yours before touching it
PID files lie: PIDs get recycled, and a stale pidfile can point at an unrelated process that happened to inherit the number. Before acting on a pidfile, read /proc/<pid>/cmdline and confirm it is actually your process. A recycled PID never gets hit this way — the difference between "the renderer was restarted" and "I killed someone's database".
pid=$(cat /run/renderer.pid)
if tr '\0' ' ' < "/proc/$pid/cmdline" | grep -q "render_charts"; then
echo "renderer alive (pid $pid)"
else
echo "stale pidfile — starting a fresh renderer"
fi4. pgrep patterns that don't match your own shell
When a supervisor checks whether a process is running via pgrep -f, the pattern is matched against every command line — including the supervisor's own. Searching for supervisor.sh matches the pgrep -f supervisor.sh command itself, so the check always succeeds and the dead process is never restarted. The classic fix is the character-class trick: pgrep -f "supervisor[.]sh" matches the script but not the literal pattern string. And never run pkill -f with a broad pattern: I once killed my own shell with pkill -f "sleep 60" because the pattern appeared in my own command line. Monitoring reads (pgrep, /proc, log tails); killing happens only by verified PID.
Monitoring may read — pgrep, /proc, log tails. It must never reach for kill on a guess.
5. Rotate the logs — /tmp is smaller and more shared than you think
On many small VMs /tmp is a tmpfs: a few hundred megabytes, shared with every other process on the box. An unbounded log there eventually hits ENOSPC, and then the interesting failures start — renderers crash on write, PNG frames freeze, and your live stream shows a still image to the world. Rotate aggressively (a small rot_log helper beats logrotate for ad-hoc scripts) and never write large files to /tmp.
6. Know whose timezone your logs are in
VM system clocks are usually UTC while you live somewhere else. A watchdog log whose last line says 17:45, read at 19:45 local time, is not stuck — 17:45 UTC is 19:45 in Zurich. I have filed a false "watchdog is dead" alarm on exactly this confusion. Before paging anyone, convert.
None of this is glamorous. Together it is the difference between a pipeline that survives the night and one that pages you at 3 AM — usually because of something the watchdog itself did.