A 25-hour hole in every schedule
At 06:41 UTC on 31 August 2026, the OOM killer took out cron.service on a Hetzner box. It stayed down until someone restarted it by hand at 07:48 UTC the next day, and the team wrote it up in a public GitHub issue on 1 September.
For 25 hours nothing scheduled ran. The hourly archiver, the hourly auto-pull deploy from main, health snapshots, batch collection: all of it stopped. The memory came from eight parallel opj_decompress processes at roughly 1 GB each on a 15 GB machine.
That root cause is ordinary. The interesting part is how one greedy job stopped the scheduler, and why a dead scheduler produced zero errors for a day.
How one job can stop the scheduler
The job lives inside the cron unit
On Debian and Ubuntu, the processes cron starts stay in the control group of cron.service. Run systemctl status cron during a job and its shell and children show up under the unit's CGroup, right next to the daemon.
systemctl status cron
# ● cron.service - Regular background program processing daemon
# CGroup: /system.slice/cron.service
# ├─ 812 /usr/sbin/cron -f
# ├─ 90411 /bin/sh -c /usr/local/bin/convert-scans.sh
# └─ 90415 opj_decompress -i scan-0042.jp2 -o scan-0042.tif
To systemd, your batch job and the daemon are one unit. Some distributions move jobs into session scopes through PAM, so check your own hosts.
OOMPolicy decides what happens next
The systemd.service man page gives OOMPolicy= three values. With continue, an OOM kill inside the unit is logged and the unit keeps running. With stop, "the unit's processes are terminated cleanly by the service manager". With kill, the kernel kills everything left in the unit.
The default comes from DefaultOOMPolicy= in system.conf. systemd ships it as stop. Only units with Delegate= turned on default to continue.
Combine the two. The kernel kills one runaway image decoder, systemd sees an OOM kill inside cron.service, applies stop, and terminates the whole unit, cron daemon included. Result: oom-kill.
Whether cron comes back
After an OOM stop, the man page says Restart= "may apply". Its restart table counts termination due to OOM as a trigger for always, on-failure and on-abnormal. The other settings ignore it.
Packaged units differ here. Debian's cron.service sets Restart=on-failure and no RestartSec=, so it restarts after the 100 ms default. cronie's upstream crond.service, used on the Red Hat family, sets Restart=on-failure and RestartSec=30s. Both restart after an OOM stop, if the unit on your host is really the packaged one.
Then there is the start limit. It counts restarts too, and the shipped defaults allow five starts within ten seconds. systemd.unit(5) says a unit with Restart= that hits the limit is "not attempted to be restarted anymore". If memory is still exhausted when cron comes back, a 100 ms restart delay burns through five attempts fast. After that it stays down until a human steps in.
The issue asks for Restart=always and an OOMPolicy on the unit, so on that box cron did not recover by itself. Check yours:
systemctl show cron -p Restart -p RestartUSec -p OOMPolicy -p Result
# Restart=on-failure
# RestartUSec=100ms
# OOMPolicy=stop
# Result=success
Hardening the cron unit
A drop-in override changes this without editing the packaged file.
sudo systemctl edit cron
[Service]
OOMPolicy=continue
Restart=always
RestartSec=10s
OOMPolicy=continue matters most. A scheduler unit holds arbitrary child processes, and one of them hitting the memory wall should not take the scheduler down. The killed job still fails. Every other schedule keeps going.
RestartSec=10s spaces restarts out so a box under sustained pressure cannot hit the start limit in half a second.
Do not make cron unkillable
The tempting next step is OOMScoreAdjust=-1000. systemd.exec describes that value as disabling OOM killing "of processes of this unit", and on Debian the jobs are processes of the unit. The image decoder that caused the trouble would become immune, and the kernel would go after your database instead. Keep any adjustment on the daemon moderate.
Give heavy jobs their own cgroup
Better still, get memory-hungry jobs out of the cron unit. systemd-run --scope starts a command in a transient scope unit, and -p sets resource properties on it:
# root crontab: the job gets its own scope and a hard memory ceiling
0 * * * * systemd-run --scope -p MemoryMax=4G /usr/local/bin/convert-scans.sh
An overrun now hits the job's own ceiling, and cron.service never hears about it. A systemd timer with MemoryMax= gives the same isolation plus a unit state you can query.
Why nobody noticed for 25 hours
The issue lists three blind spots.
An empty log looks like no work
The archive log was 0 bytes, which looked exactly like "nothing to process".
Nobody asked systemd
Nobody checked systemctl list-units --state=failed. The unit sat there failed the entire time.
The heartbeat ran on cron
The heartbeat check was itself a cron job. It died with cron. Fix this one first: a monitor scheduled by the thing it watches cannot report that thing's death, so something outside the box has to notice the missing check-in.
# one canary line per host: while cron is alive, this pings every five minutes
*/5 * * * * curl -fsS --retry 3 https://cronguard.app/api/ping/host-canary-id > /dev/null
The canary proves the daemon is alive, per-job check-ins prove each job finished. With both, a dead cron shows up within minutes.
Frequently asked questions about cron and the OOM killer
Can one cron job's memory use stop the cron daemon? Yes. On Debian and Ubuntu, job processes stay in the cron.service control group, and systemd's default OOMPolicy is stop. One OOM-killed job process makes systemd terminate the whole unit, daemon included.
Does systemd restart cron after an OOM kill? Only if the unit's Restart setting is always, on-failure or on-abnormal. Debian and cronie both ship on-failure, but a unit that hits the start limit of five starts in ten seconds is not restarted again until someone intervenes.
Should I set OOMScoreAdjust to -1000 on cron? No. That value disables OOM killing for every process of the unit, including the jobs cron starts. The job causing the pressure would become immune and the kernel would kill something else, such as your database.
What is the safest OOMPolicy for the cron unit? Set OOMPolicy=continue in a drop-in override, with Restart=always and a restart delay of a few seconds. The job that ran out of memory still dies, but the scheduler and every other schedule on the host keep running.
How do I find out that cron itself has stopped? Not with a check that cron schedules, because it dies with the daemon. Add a per-host canary job that pings an external monitor every few minutes, so the missing check-in raises the alert.
Further reading
- Monitoring systemd Timers: The Cron Successor Fails Just as Quietly
- Everything Runs at Midnight: Spreading Out Your Scheduled Jobs
- Dead Man's Switch Monitoring: The Only Reliable Way to Watch Cron Jobs
Conclusion: On most Debian-style hosts your jobs run inside cron's own systemd unit, and the default OOM policy treats a job's memory problem as the unit's problem. Check three things today: what OOMPolicy and Restart say for your cron unit, whether your heaviest jobs have a memory ceiling of their own, and whether anything outside the box would notice if cron stopped at 06:41 tomorrow.
Sources: cron.service OOM incident (GitHub issue #4528), systemd.service(5), systemd.unit(5), systemd.exec(5), Debian cron.service.