AI Workload Scheduler With Hierarchical Resource Ticket Balancing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for managing AI workloads in cloud infrastructure face challenges in scalability, efficiency, and fairness, as general-purpose cloud environments struggle to handle the unique requirements of AI workloads, leading to inefficiencies and resource fragmentation.

Innovation Solution

A global scheduler distributes AI workloads across nodes based on resource ticket values, with local schedulers managing workload execution on infrastructure resources, enabling fair and efficient scheduling through multiplexing, topology-aware scheduling, and tier-based priority management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If general-purpose cloud infrastructure is used to execute AI workloads, then existing infrastructure can be leveraged, but resource utilization efficiency deteriorates due to fundamental differences between AI workloads and general-purpose cloud environments

Engineering Contradiction:
Improveinfrastructure reusabilityVSAvoidresource utilization efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent implements purpose-built AI infrastructure with specialized components optimized for AI workloads rather than general-purpose resources. The AI infrastructure includes dedicated AI workload managers, specialized scheduling mechanisms, and optimized resource pools that provide local quality enhancements for AI-specific requirements, thereby improving resource utilization efficiency while maintaining adaptability through targeted optimizations.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If AI workloads are scheduled on infrastructure, then workload execution is enabled, but scheduling fairness and efficiency deteriorate due to substantial challenges in managing resource distribution

Engineering Contradiction:
Improveworkload execution capabilityVSAvoidscheduling efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent introduces an AI workload manager as an intermediary layer between AI workloads and infrastructure resources. This manager implements specialized scheduling algorithms that balance fairness and efficiency by mediating resource allocation, prioritizing workloads appropriately, and managing resource distribution without requiring direct infrastructure modifications, thereby maintaining ease of operation while improving scheduling efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If AI workloads are distributed across cloud nodes, then scalability is improved, but resource ticket value balancing becomes complex requiring sophisticated distribution mechanisms

Engineering Contradiction:
Improvesystem scalabilityVSAvoidscheduling mechanism complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the scheduling system into hierarchical levels with an AI workload manager handling high-level resource allocation and node-level schedulers handling local distribution. This segmentation allows scalability to be improved by adding nodes while maintaining manageable complexity through divided responsibilities, where each segment operates with appropriate autonomy and follows standardized protocols for resource ticket value balancing.

Inventive Principle:
Principle #1Segmentation

4Reliability

If resource ticket values are used to balance workload distribution, then fair resource sharing is achieved, but scheduling overhead increases due to continuous tracking and balancing requirements

Engineering Contradiction:
Improveresource distribution fairnessVSAvoidscheduling overhead
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent implements periodic resource ticket value balancing where the AI workload manager performs resource distribution optimization at scheduled intervals rather than continuously. This periodic action maintains fair resource sharing by regularly adjusting allocations while reducing scheduling overhead by avoiding constant tracking and rebalancing operations, thereby achieving reliability without excessive time loss.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS12353908B2Scheduler for planet-scale computing system
Publication Date: 2025.07.08 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12353908B2 patent drawing
  • US12353908B2 patent drawing
  • US12353908B2 patent drawing

AI summary

The disclosure herein describes scheduling execution of artificial intelligence (AI) workloads in a cloud infrastructure platform. A global scheduler receives AI workloads associated with resource ticket values. The scheduler distributes the AI workloads to nodes based on balancing resource ticket values. Local schedulers of the nodes schedule AI workloads on resources based on the resource ticket values of the AI workloads. Based on scheduling the AI workloads, coordinator services of the local schedulers execute the distributed AI workloads on the infrastructure resources of the nodes. The disclosure further describes scheduling AI workloads based on priority tiers. A scheduler receives AI workloads, and each AI workload is associated with a priority tier indicative of a preemption priority while being executed. The AI workloads are scheduled for execution on a distributed set of nodes based on the priority tiers and then execute based on the scheduling.