Back to blog

systemd 262 Queues Your Timer Jobs, and One Hang Blocks Them All

systemd 262 adds ActivatingConcurrencyMax= to cap how many timer jobs start at once. Queued jobs get no error, and one hung oneshot can stall the whole slice.

CronGuard TeamCron Job Monitoring Experts
6 min read
The front of a dark network rack with rows of patch panels, tangled blue and black Ethernet cables, and a single red cable looping down across the ports

A new knob for the 03:00 pile-up

systemd 262 shipped on 22 September 2026. One line in its release notes matters if your scheduled work runs as timer units: slice units gained ActivatingConcurrencyMax=. It caps how many units in a slice hierarchy can be activating at the same time. Anything over the cap waits in a queue and starts when a slot frees up.

That is a job pool inside PID 1. Put the backup and the search reindex in one slice with a limit of one and they stop fighting over the disk at 03:00. It handles the problem from spreading out scheduled jobs better than random delays do. It also gives a job a new way to not run without anyone getting an error.

What ActivatingConcurrencyMax actually limits

Activating, not active

systemd 258 already gave slices ConcurrencySoftMax= and ConcurrencyHardMax=. The systemd.slice man page says those count units that are active. Over the soft limit, new starts queue. Over the hard limit, they fail with an error.

The new option counts units "while they are starting up", in the man page's words. When a unit leaves the activating state, for whatever state comes next, its slot goes to the next unit in line.

Why oneshot jobs live in activating

For a daemon, activating lasts a few seconds. Timer jobs are almost always Type=oneshot, and systemd.service says a oneshot without RemainAfterExit= never reaches active. It goes from activating straight to deactivating or dead.

So the unit is activating for its entire run. For timer jobs the new limit caps how many run at once, and a 40-minute backup holds its slot for 40 minutes.

Putting timer jobs in a bounded slice

Two small files. The slice:

# /etc/systemd/system/batch.slice
[Unit]
Description=Nightly batch jobs

[Slice]
ActivatingConcurrencyMax=2

Then each job's service points at it. The timers stay as they are:

# /etc/systemd/system/report-export.service
[Service]
Type=oneshot
Slice=batch.slice
ExecStart=/usr/local/bin/report-export.sh
systemctl --version | head -1
systemctl daemon-reload

Check that version line first. A systemd older than 262 does not know the setting, and your jobs keep starting together.

Where the queue goes wrong

One hung job holds a slot forever

The man page says it plainly: a unit stuck in activating, on a hung process for example, keeps its slot. It tells you to set TimeoutStartSec=, and for timer jobs that matters. systemd.service disables TimeoutStartSec= by default for Type=oneshot, and RuntimeMaxSec= has no effect on oneshot services at all. A backup hanging on a dead NFS mount sits in its slot until someone kills it. With a limit of two, two hangs freeze the slice.

Queued runs report nothing

At the limit, the man page says, "No error is returned to the caller." The timer fires, systemd queues a start job, and that is it. Nothing failed, so OnFailure= stays quiet and the unit looks normal for as long as the job waits.

A fired timer can collapse into the queue

systemd merges a new start request into a start job already queued for the same unit. An hourly job that waits three hours runs once when the slot opens. Two runs are gone and no log line says so.

Zero means frozen

ActivatingConcurrencyMax=0 blocks every activation in the slice tree until someone raises the limit. Handy during maintenance. Less handy when a config management run sets it and nobody sets it back.

Seeing the queue from the host

list-jobs shows what is waiting

systemctl --failed shows nothing here, because a queued job has not failed. Look at the job queue:

systemctl list-jobs
JOB  UNIT                   TYPE  STATE
4127 search-reindex.service start waiting
4126 report-export.service  start waiting
4119 invoice-sync.service   start running
4118 db-backup.service      start running

4 jobs listed.

Two running and two waiting at 03:10 is the limit doing its job. The same list at 09:00 is an incident.

Guarding the slot with a start timeout

Give every oneshot in a bounded slice a start timeout above its worst normal run:

[Service]
Type=oneshot
Slice=batch.slice
TimeoutStartSec=45min
ExecStart=/usr/local/bin/report-export.sh

A hang now turns into a failed unit after 45 minutes. The slot frees up, OnFailure= gets its chance, and the queue moves again.

What a heartbeat sees that the queue hides

The case that hurts is the job that should have finished by 04:00 and is still in line. From the host it looks fine. The unit has not failed and list-timers shows a recent trigger.

A check-in from the job itself catches it. Send it on success only, from ExecStartPost=, which runs after ExecStart= exits cleanly:

[Service]
ExecStartPost=/usr/bin/curl -fsS --retry 3 https://cronguard.app/api/ping/your-monitor-id

Size the monitor's grace period by the longest queue delay you will put up with. The job's runtime alone is too tight a window. Whether the job sat behind a hung neighbour or the slice was frozen at zero, the symptom is the same: no ping by the deadline. That is all a dead man's switch for a timer has to see.

Frequently asked questions about systemd slice concurrency limits

What does ActivatingConcurrencyMax do in systemd 262? It limits how many units in a slice hierarchy can be in the activating state at the same time. Once the limit is reached, further start requests are queued and dispatched automatically when a running activation completes. It arrived in systemd 262, released on 22 September 2026.

Why does it limit oneshot timer jobs for their whole runtime? A Type=oneshot service without RemainAfterExit never enters the active state, so it stays in activating until its process exits. For timer jobs that means the limit caps how many jobs run at once, and a long job holds its slot until it finishes.

Can one hung job block every other job in the slice? Yes. A unit stuck in activating keeps its slot, TimeoutStartSec is disabled by default for oneshot services, and RuntimeMaxSec has no effect on them. Set TimeoutStartSec on every oneshot in a bounded slice so a hang turns into a failure and frees the slot.

Will systemd tell me when a timer job is stuck in the queue? No. The man page says no error is returned to the caller when the limit is reached, so the unit never fails and OnFailure never fires. systemctl list-jobs shows waiting start jobs, but only if someone runs it at the right moment.

How do I monitor timer jobs that run in a concurrency-limited slice? Have each job send a check-in on success from ExecStartPost, and set the monitor's grace period to the longest queue delay you can accept. A job that is queued or hung then shows up as a missing check-in, whatever the cause on the host.

Further reading


Conclusion: ActivatingConcurrencyMax= gives systemd a real job pool, and for a crowded 03:00 it beats hoping randomized delays spread the load. The price is a waiting room with no alarm. A oneshot holds its slot for its whole run and, without a start timeout, for as long as it hangs. Set TimeoutStartSec= on every job in the slice and alert on the check-in that never arrived.

Sources: systemd v262 release notes, systemd.slice(5), systemd.service(5).

Share

Related posts

A close-up of a screen showing a green Ubuntu Linux shell prompt reading ubuntu@ubuntu:~$ with the word sudo being typed and a bright blinking cursor
MonitoringMonitoring systemd Timers: The Cron Successor Fails Just as Quietly
Rows of patch panels in a server rack with looped black network cables, lit in deep green and half hidden behind the dark edge of a neighbouring cabinet
ReliabilityGang-Scheduled CronJobs in Kubernetes 1.37: The Run That Never Starts
Three server cabinets with mesh doors in a dark room, packed with equipment, orange and blue cables looping down the front and rows of small green status lights
ReliabilityWhen the OOM Killer Takes Down Cron Itself, Not Just Your Job

Set up your first monitor.
It'll take 30 seconds.