Ai2 swaps its priority-based GPU scheduler for budgets, and debug jobs now start in seconds
The Allen Institute for AI says replacing priority levels with GPU-time budgets, fair-share scheduling and a time-slicing contract delivered teams 98% of the GPU hours they were owed, while cluster occupancy held at 98%.

Photo: Florian Hirzinger / Wikimedia Commons, CC BY-SA 3.0
The Allen Institute for AI's infrastructure team described on the Hugging Face blog on October 9 how it rebuilt scheduling for clusters holding thousands of NVIDIA H100, B200 and B300 GPUs, used by about 150 researchers. Demand outstrips supply by two to three times, the team says.
Under the old priority system, the team says, researchers parked idle jobs to hold GPUs, nearly every job ended up marked high priority, and on-call engineers spent much of their time negotiating shutdowns for maintenance. Ai2 replaced it with GPU-time budgets set by managers, a hierarchical fair-share scheduler over a seven-day window, and a contract in which each job declares a minimum runtime, capped at eight hours, during which it cannot be preempted.
Over a 30-day test, Ai2 reports, teams received 98% of the GPU hours they were owed and occupancy stayed at 98%. The 90th-percentile wait for small debug jobs fell from 2 hours to 30 seconds, and repairs needing a human fell by 74%.
The team says not everything improved: researchers relying on long interactive sessions were hurt by preemption, so Ai2 is building a CPU-only cluster and restorable sessions, and it is investigating longer waits for the largest jobs.