Ai2’s AI Infrastructure team, which manages thousands of NVIDIA GPUs across large clusters, has replaced a legacy priority-based scheduler with a system that ties GPU access to budgeted time, uses hierarchical fair-share allocation, and introduces a time-slicing contract. The move targets the institute’s large, distributed training workloads and aims to align resource allocation with program strategy rather than case-by-case prioritization.
WHAT CHANGED: BUDGETS, FAIR-SHARE, AND TIME SLICING
The organization migrated from a priority-driven approach to a governance model where managers allocate a share of GPU time to projects and researchers. GPU time is now funded by budgets tied to programs, and allocations feed into a hierarchical fair-share scheduler that tracks occupancy over a lookback window. This approach aims to ensure that demand does not outpace supply and that time is spent on the most valuable workloads rather than on ad hoc scheduling decisions. TechStaged has also covered Asana slashes browser-model costs and speeds up using GPT-6.1 Sol in Codex tests.
HOW THE FOUR-METRIC PYRAMID GUIDES SCHEDULING
Ai2 describes a four-tier metric framework that builds from availability up to utilization. Availability measures hardware health and readiness, occupancy tracks the fraction of time assigned to workloads, impact reflects how often high-value workloads are chosen for resources, and utilization represents how much GPU capacity is used over the lifetime of a workload. The scheduling system targets improving impact while maintaining full utilization.
MECHANICS OF ALLOCATION AND PROTECTION
The new scheduler introduces two kinds of occupancy: allocated occupancy, charged to a workload’s budget and protected from preemption for a minimum runtime, and unallocated occupancy, which is uncharged, unprotected, and can be preempted. A workload must be funded by a budget to be protected, and the system allows rebalancing once minimum progress is banked. If a workload sets its minimum runtime to zero, it is unallocated and subject to preemption but not charged to any budget.
THE SCHEDULING CONTRACT AND ITS EFFECTS ON MAINTENANCE
To address long-running training jobs, the system requires workloads to declare a minimum runtime. This creates a protected window for progress while enabling automatic re-queuing and rebalancing. The contract helps retired or unhealthy hosts drain workloads so repairs can be automated. In practice, this reduced human-in-the-loop maintenance, lowering on-call toil by a substantial margin.
EARLY RESULTS AND ONGOING CHALLENGES
Ai2 notes that demand for GPU time often exceeds supply, with 2-3 times as many GPUs requested as are available at any moment. By cushioning allocations with budgets and a transparent governance process, teams can advocate for time without resorting to problematic behaviors observed under the old system, such as squatting or priority inflation. A cited benefit from practitioners is the impression of more available compute when capacity is underutilized, effectively creating more flexible compute through the new scheduling model.
WHY THIS MATTERS FOR LARGE-SCALE AI RESEARCH
The shift is framed as a response to the tragedy of the commons in shared GPU resources: private incentives had previously undermined global throughput. By converting scheduling from a case-by-case negotiation to a budget-driven, program-wide decision, Ai2 aligns resource use with strategic goals and reduces the need for manual intervention in daily scheduling.
WHAT HAPPENS NEXT
The post outlines ongoing iteration on budget levels and the governance structure to balance competing research needs with limited GPU capacity. The approach emphasizes proportional ownership of GPU time and a mechanism to rebalance as workloads evolve, aiming to sustain high occupancy while protecting important workloads from preemption.
RELATED COVERAGE
- Asana slashes browser-model costs and speeds up using GPT-6.1 Sol in Codex tests
- Sophos cuts threat investigation time by 96% with OpenAI Daybreak
- NVIDIA Demonstrates Frontier AI Agents Turning Simulation Ideas Into Working Omniverse Apps
- Rally Up: Gears of War: E-Day Lands on GeForce NOW with RTX Cloud Power
- AI & Automation articles
SOURCES
- Hugging Face - Blog: Impactful scheduling for GPU clusters Published · Primary source





