Back to blog

Gang-Scheduled CronJobs in Kubernetes 1.37: The Run That Never Starts

Kubernetes 1.37 lets a Job ask for all-or-nothing gang scheduling. On a CronJob, a short cluster means pods that wait forever while nothing fails. How to bound it.

CronGuard TeamCron Job Monitoring Experts
6 min read
Rows of patch panels in a server rack with looped black network cables, lit in deep green and half hidden behind the dark edge of a neighbouring cabinet

On schedule, and still nothing ran

The CronJob fired at 02:00. LAST SCHEDULE reads 5h, ACTIVE reads 1, and four pods sit in Pending. None of them has failed. None of them will. The nightly run started on time and did no work.

That is all-or-nothing scheduling on a cluster one GPU short. Kubernetes v1.37, released on 26 August 2026, moved gang scheduling to beta and let a Job request it. CronJobs get it through jobTemplate: the API server checks the same feature gate when it validates that template.

What changed in Kubernetes 1.37

Gang scheduling reached beta

The GenericWorkload feature gate went to beta in v1.37 and now covers gang scheduling and workload-aware preemption too. The Workload and PodGroup APIs moved to scheduling.k8s.io/v1beta1. Beta or not, the gate is off by default. You switch it on for kube-apiserver, kube-controller-manager and kube-scheduler.

Jobs can now ask for it

The Job API gained .spec.scheduling. Its schedulingPolicy is either basic, pod by pod as before, or gang. Leave out gang.minCount and it defaults to the Job's parallelism. The field needs the WorkloadWithJob gate, alpha since v1.36 and also off by default, and only minCount can change after creation.

Nobody gets this by accident. The teams that switch it on run multi-pod batch work on scarce hardware, often nightly:

apiVersion: batch/v1
kind: CronJob
metadata:
  name: nightly-embeddings
spec:
  schedule: "0 2 * * *"
  timeZone: "Europe/Amsterdam"
  concurrencyPolicy: Forbid
  jobTemplate:
    spec:
      parallelism: 4
      completions: 4
      completionMode: Indexed
      scheduling:
        schedulingPolicy:
          gang: {}   # minCount defaults to parallelism: 4
      template:
        spec:
          restartPolicy: Never
          containers:
          - name: worker
            image: registry.example.com/embeddings:2026.09
            resources:
              limits:
                nvidia.com/gpu: 1

How an all-or-nothing run stalls

The scheduler waits for a quorum

Per the docs, the scheduler holds a gang's pods in PreEnqueue until the PodGroup exists and at least minCount pods have been created. Before that the group never enters the active queue. So anything that stops the Job controller from creating every pod also stalls the gang. A namespace pod quota below minCount will do it.

Then it waits for capacity

With the quorum in place, the scheduler tries to place the whole group in one cycle. If it finds room for fewer than minCount pods, it binds none. The docs say the pods "are moved to the unschedulable queue to wait for cluster resources to free up". No timeout is mentioned.

That is on purpose. Three workers sitting on GPUs while the fourth waits is the deadlock gang scheduling was built to prevent. A nightly job pays for it by waiting.

Nothing in the Job counts as a failure

backoffLimit counts failed pods, and a pod that never got a node never failed. The Job carries no Failed condition. The CronJob looks fine too: it did its part and created a Job on time.

kubectl get cronjob nightly-embeddings
# NAME                 SCHEDULE    TIMEZONE           SUSPEND   ACTIVE   LAST SCHEDULE   AGE
# nightly-embeddings   0 2 * * *   Europe/Amsterdam   False     1        5h12m           41d

kubectl get pods -l batch.kubernetes.io/job-name=nightly-embeddings-29321640
# NAME                                  READY   STATUS    RESTARTS   AGE
# nightly-embeddings-29321640-0-x7k2p   0/1     Pending   0          5h12m
# nightly-embeddings-29321640-1-q9d4m   0/1     Pending   0          5h12m
# ...indexes 2 and 3 look the same

How concurrencyPolicy changes the damage

The CronJob default is Allow, so every tick adds another gang Job fighting over the same missing capacity. With Forbid, the controller skips new runs while the stalled one is active. One stuck night blocks every night after it. Replace cancels the stalled Job and creates a fresh one, which waits the same way.

Whichever you pick, kubectl get cronjob looks healthy.

Putting a ceiling on the wait

Set activeDeadlineSeconds

activeDeadlineSeconds counts from the Job's status.startTime, which the API defines as the time "when the job controller started processing a job". Pending time counts. When the deadline passes, the Job becomes type: Failed with reason: DeadlineExceeded.

  jobTemplate:
    spec:
      activeDeadlineSeconds: 14400   # 4h: fail well before the next 02:00 run
      parallelism: 4
      completions: 4

Keep it shorter than the schedule interval. Your alerting can see a failed Job. It cannot see a pending one.

Ask for the gang size you actually need

minCount defaults to parallelism. If three of four workers can still do useful work, set minCount: 3. At the default, one missing node blocks the whole run.

Reading the PodGroup condition

With the gate on, the Job controller creates a Workload and a PodGroup for each Job, basic ones included. After a scheduling cycle the scheduler sets PodGroupInitiallyScheduled on the PodGroup. A gang it could not place reads False with reason Unschedulable.

kubectl get podgroup
kubectl get podgroup <name> -o jsonpath='{.status.conditions}'

Good enough for a dashboard, with a caveat from the docs: once the condition is True it never changes, and the scheduler stops updating PodGroup status after the first placement. It says nothing about whether the run finished.

Monitor the end of the run

Every run in this post started on time, so any check on Job creation passes. Only completion separates a stalled night from a good one. Send a check-in when the work is done, and alert when it does not arrive.

In an Indexed gang Job, let one pod send it. Each pod gets its index in JOB_COMPLETION_INDEX:

#!/bin/sh
set -eu
/app/build-embeddings --shard "$JOB_COMPLETION_INDEX"
if [ "$JOB_COMPLETION_INDEX" = "0" ]; then
  curl -fsS --retry 3 https://cronguard.app/api/ping/nightly-embeddings-id > /dev/null
fi

Set the grace period to the expected runtime plus a margin. A gang that waits all night then misses its check-in, and you find out over breakfast.

Frequently asked questions about gang-scheduled CronJobs

Can a Kubernetes CronJob use gang scheduling? Yes. From Kubernetes 1.37 a Job can request it through its scheduling field, and a CronJob passes its job template on to the Jobs it creates. It needs the alpha WorkloadWithJob gate and the beta GenericWorkload gate, and both are off by default.

What happens when the cluster cannot fit the whole gang? The scheduler binds none of the pods and moves them to the unschedulable queue until resources free up. The Kubernetes docs mention no timeout for that wait, so the pods stay Pending.

Why does a stalled gang Job not trip backoffLimit? The backoff limit counts failed pods, and a pod that was never bound to a node has not failed. The Job stays active with no Failed condition, and the CronJob looks healthy because it created the Job on time.

Does activeDeadlineSeconds stop a gang Job that never got scheduled? Yes. The deadline counts from the Job's startTime, which is set when the Job controller starts processing it, so time spent Pending counts. When it expires the Job is marked Failed with reason DeadlineExceeded.

How do I find out that a gang-scheduled run never finished? Monitor completion. Have one pod send a check-in to an external monitor when the work is done, and alert when that check-in does not arrive within the expected runtime.

Further reading


Conclusion: Gang scheduling solves a real problem for multi-pod batch work. The cost is that a short cluster no longer produces an error, just a wait with no end. If you enable WorkloadWithJob for CronJobs, put an activeDeadlineSeconds below the interval on every gang Job and size minCount to what the work needs. Then watch the end of each run. The start will always look fine.

Sources: Kubernetes v1.37: Garhwal, Kubernetes v1.37: Advancing Workload-Aware Scheduling, Gang Scheduling, Jobs, PodGroup scheduling.

Share

Related posts

A fiber-optic patch panel in a rack, green connectors and yellow cables looping across rows of numbered ports
ReliabilityWhen a Scheduled Job Runs Long: Overlap, Skip, or Replace
A person at a desk in a dark room, lit by two monitors filled with lines of code, looking toward the camera, in black and white
ReliabilityThe Kubernetes CronJob Deadlock That Stops Your Schedule For Good
Colorful syntax-highlighted source code on a black screen viewed at a slight angle, with a database query and functions like parseFloat and forEach visible in orange, blue and pink
ReliabilityThe Python Cron Parser Behind Airflow Is Not the Cron You Know

Set up your first monitor.
It'll take 30 seconds.