On schedule, and still nothing ran
The CronJob fired at 02:00. LAST SCHEDULE reads 5h, ACTIVE reads 1, and four pods sit in Pending. None of them has failed. None of them will. The nightly run started on time and did no work.
That is all-or-nothing scheduling on a cluster one GPU short. Kubernetes v1.37, released on 26 August 2026, moved gang scheduling to beta and let a Job request it. CronJobs get it through jobTemplate: the API server checks the same feature gate when it validates that template.
What changed in Kubernetes 1.37
Gang scheduling reached beta
The GenericWorkload feature gate went to beta in v1.37 and now covers gang scheduling and workload-aware preemption too. The Workload and PodGroup APIs moved to scheduling.k8s.io/v1beta1. Beta or not, the gate is off by default. You switch it on for kube-apiserver, kube-controller-manager and kube-scheduler.
Jobs can now ask for it
The Job API gained .spec.scheduling. Its schedulingPolicy is either basic, pod by pod as before, or gang. Leave out gang.minCount and it defaults to the Job's parallelism. The field needs the WorkloadWithJob gate, alpha since v1.36 and also off by default, and only minCount can change after creation.
Nobody gets this by accident. The teams that switch it on run multi-pod batch work on scarce hardware, often nightly:
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-embeddings
spec:
schedule: "0 2 * * *"
timeZone: "Europe/Amsterdam"
concurrencyPolicy: Forbid
jobTemplate:
spec:
parallelism: 4
completions: 4
completionMode: Indexed
scheduling:
schedulingPolicy:
gang: {} # minCount defaults to parallelism: 4
template:
spec:
restartPolicy: Never
containers:
- name: worker
image: registry.example.com/embeddings:2026.09
resources:
limits:
nvidia.com/gpu: 1
How an all-or-nothing run stalls
The scheduler waits for a quorum
Per the docs, the scheduler holds a gang's pods in PreEnqueue until the PodGroup exists and at least minCount pods have been created. Before that the group never enters the active queue. So anything that stops the Job controller from creating every pod also stalls the gang. A namespace pod quota below minCount will do it.
Then it waits for capacity
With the quorum in place, the scheduler tries to place the whole group in one cycle. If it finds room for fewer than minCount pods, it binds none. The docs say the pods "are moved to the unschedulable queue to wait for cluster resources to free up". No timeout is mentioned.
That is on purpose. Three workers sitting on GPUs while the fourth waits is the deadlock gang scheduling was built to prevent. A nightly job pays for it by waiting.
Nothing in the Job counts as a failure
backoffLimit counts failed pods, and a pod that never got a node never failed. The Job carries no Failed condition. The CronJob looks fine too: it did its part and created a Job on time.
kubectl get cronjob nightly-embeddings
# NAME SCHEDULE TIMEZONE SUSPEND ACTIVE LAST SCHEDULE AGE
# nightly-embeddings 0 2 * * * Europe/Amsterdam False 1 5h12m 41d
kubectl get pods -l batch.kubernetes.io/job-name=nightly-embeddings-29321640
# NAME READY STATUS RESTARTS AGE
# nightly-embeddings-29321640-0-x7k2p 0/1 Pending 0 5h12m
# nightly-embeddings-29321640-1-q9d4m 0/1 Pending 0 5h12m
# ...indexes 2 and 3 look the same
How concurrencyPolicy changes the damage
The CronJob default is Allow, so every tick adds another gang Job fighting over the same missing capacity. With Forbid, the controller skips new runs while the stalled one is active. One stuck night blocks every night after it. Replace cancels the stalled Job and creates a fresh one, which waits the same way.
Whichever you pick, kubectl get cronjob looks healthy.
Putting a ceiling on the wait
Set activeDeadlineSeconds
activeDeadlineSeconds counts from the Job's status.startTime, which the API defines as the time "when the job controller started processing a job". Pending time counts. When the deadline passes, the Job becomes type: Failed with reason: DeadlineExceeded.
jobTemplate:
spec:
activeDeadlineSeconds: 14400 # 4h: fail well before the next 02:00 run
parallelism: 4
completions: 4
Keep it shorter than the schedule interval. Your alerting can see a failed Job. It cannot see a pending one.
Ask for the gang size you actually need
minCount defaults to parallelism. If three of four workers can still do useful work, set minCount: 3. At the default, one missing node blocks the whole run.
Reading the PodGroup condition
With the gate on, the Job controller creates a Workload and a PodGroup for each Job, basic ones included. After a scheduling cycle the scheduler sets PodGroupInitiallyScheduled on the PodGroup. A gang it could not place reads False with reason Unschedulable.
kubectl get podgroup
kubectl get podgroup <name> -o jsonpath='{.status.conditions}'
Good enough for a dashboard, with a caveat from the docs: once the condition is True it never changes, and the scheduler stops updating PodGroup status after the first placement. It says nothing about whether the run finished.
Monitor the end of the run
Every run in this post started on time, so any check on Job creation passes. Only completion separates a stalled night from a good one. Send a check-in when the work is done, and alert when it does not arrive.
In an Indexed gang Job, let one pod send it. Each pod gets its index in JOB_COMPLETION_INDEX:
#!/bin/sh
set -eu
/app/build-embeddings --shard "$JOB_COMPLETION_INDEX"
if [ "$JOB_COMPLETION_INDEX" = "0" ]; then
curl -fsS --retry 3 https://cronguard.app/api/ping/nightly-embeddings-id > /dev/null
fi
Set the grace period to the expected runtime plus a margin. A gang that waits all night then misses its check-in, and you find out over breakfast.
Frequently asked questions about gang-scheduled CronJobs
Can a Kubernetes CronJob use gang scheduling? Yes. From Kubernetes 1.37 a Job can request it through its scheduling field, and a CronJob passes its job template on to the Jobs it creates. It needs the alpha WorkloadWithJob gate and the beta GenericWorkload gate, and both are off by default.
What happens when the cluster cannot fit the whole gang? The scheduler binds none of the pods and moves them to the unschedulable queue until resources free up. The Kubernetes docs mention no timeout for that wait, so the pods stay Pending.
Why does a stalled gang Job not trip backoffLimit? The backoff limit counts failed pods, and a pod that was never bound to a node has not failed. The Job stays active with no Failed condition, and the CronJob looks healthy because it created the Job on time.
Does activeDeadlineSeconds stop a gang Job that never got scheduled? Yes. The deadline counts from the Job's startTime, which is set when the Job controller starts processing it, so time spent Pending counts. When it expires the Job is marked Failed with reason DeadlineExceeded.
How do I find out that a gang-scheduled run never finished? Monitor completion. Have one pod send a check-in to an external monitor when the work is done, and alert when that check-in does not arrive within the expected runtime.
Further reading
- Kubernetes Job Suspension and Why It Creates a Monitoring Blind Spot
- The Kubernetes CronJob Deadlock That Stops Your Schedule For Good
- When a Scheduled Job Runs Long: Overlap, Skip, or Replace
Conclusion: Gang scheduling solves a real problem for multi-pod batch work. The cost is that a short cluster no longer produces an error, just a wait with no end. If you enable WorkloadWithJob for CronJobs, put an activeDeadlineSeconds below the interval on every gang Job and size minCount to what the work needs. Then watch the end of each run. The start will always look fine.
Sources: Kubernetes v1.37: Garhwal, Kubernetes v1.37: Advancing Workload-Aware Scheduling, Gang Scheduling, Jobs, PodGroup scheduling.