Trending:

Ai2 overhauls GPU scheduling with budgets, hierarchical fair-share, and time-slicing contract

Data center racks with GPU clusters and scheduling dashboards
TechStaged-owned

Summary

  • Ai2 replaced a priority-based GPU scheduler with a system including GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract.
  • The scheduling framework uses four metrics: availability, occupancy, impact, and utilization.
  • Ai2’s GPU infrastructure manages thousands of NVIDIA H100, B200, and B300 GPUs across clusters sized 88 to 1024 GPUs and serves roughly 150 internal researchers.

Ai2’s AI Infrastructure team, which manages thousands of NVIDIA GPUs across large clusters, has replaced a legacy priority-based scheduler with a system that ties GPU access to budgeted time, uses hierarchical fair-share allocation, and introduces a time-slicing contract. The move targets the institute’s large, distributed training workloads and aims to align resource allocation with program strategy rather than case-by-case prioritization.

WHAT CHANGED: BUDGETS, FAIR-SHARE, AND TIME SLICING

The organization migrated from a priority-driven approach to a governance model where managers allocate a share of GPU time to projects and researchers. GPU time is now funded by budgets tied to programs, and allocations feed into a hierarchical fair-share scheduler that tracks occupancy over a lookback window. This approach aims to ensure that demand does not outpace supply and that time is spent on the most valuable workloads rather than on ad hoc scheduling decisions. TechStaged has also covered Asana slashes browser-model costs and speeds up using GPT-6.1 Sol in Codex tests.

HOW THE FOUR-METRIC PYRAMID GUIDES SCHEDULING

Ai2 describes a four-tier metric framework that builds from availability up to utilization. Availability measures hardware health and readiness, occupancy tracks the fraction of time assigned to workloads, impact reflects how often high-value workloads are chosen for resources, and utilization represents how much GPU capacity is used over the lifetime of a workload. The scheduling system targets improving impact while maintaining full utilization.

MECHANICS OF ALLOCATION AND PROTECTION

The new scheduler introduces two kinds of occupancy: allocated occupancy, charged to a workload’s budget and protected from preemption for a minimum runtime, and unallocated occupancy, which is uncharged, unprotected, and can be preempted. A workload must be funded by a budget to be protected, and the system allows rebalancing once minimum progress is banked. If a workload sets its minimum runtime to zero, it is unallocated and subject to preemption but not charged to any budget.

THE SCHEDULING CONTRACT AND ITS EFFECTS ON MAINTENANCE

To address long-running training jobs, the system requires workloads to declare a minimum runtime. This creates a protected window for progress while enabling automatic re-queuing and rebalancing. The contract helps retired or unhealthy hosts drain workloads so repairs can be automated. In practice, this reduced human-in-the-loop maintenance, lowering on-call toil by a substantial margin.

EARLY RESULTS AND ONGOING CHALLENGES

Ai2 notes that demand for GPU time often exceeds supply, with 2-3 times as many GPUs requested as are available at any moment. By cushioning allocations with budgets and a transparent governance process, teams can advocate for time without resorting to problematic behaviors observed under the old system, such as squatting or priority inflation. A cited benefit from practitioners is the impression of more available compute when capacity is underutilized, effectively creating more flexible compute through the new scheduling model.

WHY THIS MATTERS FOR LARGE-SCALE AI RESEARCH

The shift is framed as a response to the tragedy of the commons in shared GPU resources: private incentives had previously undermined global throughput. By converting scheduling from a case-by-case negotiation to a budget-driven, program-wide decision, Ai2 aligns resource use with strategic goals and reduces the need for manual intervention in daily scheduling.

WHAT HAPPENS NEXT

The post outlines ongoing iteration on budget levels and the governance structure to balance competing research needs with limited GPU capacity. The approach emphasizes proportional ownership of GPU time and a mechanism to rebalance as workloads evolve, aiming to sustain high occupancy while protecting important workloads from preemption.

Reporting by Nora Ellington; editing by TechStaged editors

Editorial disclosure: This article was prepared with AI assistance from a source-limited research package and passed TechStaged's automated factual, originality, licensing, and publication checks.

Our Standards: The TechStaged Editorial Principles.

f in

Nora Ellington

Nora Ellington

AI & Automation Reporter

Nora reports on AI tools, automation workflows, and the product updates shaping modern business operations.