Accelerator Workload Scheduling With Fine-Grained Resource Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current workload orchestration platforms lack the flexibility to support fine-grained accelerator requirements for scheduling accelerator-enabled workloads, leading to inefficient resource allocation and potential performance issues.

Innovation Solution

A framework that includes an accelerator discovery component, workload specification component, context-aware scheduler, and resource recommendation engine to manage and automate the scheduling of workloads based on detailed accelerator requirements, such as type, architecture, resource allocation, and service level agreements, while also adjusting these requirements dynamically.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If users can only specify the number of GPUs for workload deployment, then the scheduling process is simple, but the flexibility and granularity of accelerator requirements are insufficient

Engineering Contradiction:
Improveflexibility of accelerator requirementsVSAvoidcomplexity of scheduling system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments accelerator requirements into multiple independent dimensions including accelerator type, architecture, model, resource allocation, service level agreements, and prioritized lists. This segmentation allows users to specify fine-grained requirements for each dimension separately, transforming a single complex specification into multiple manageable parameters that can be independently configured and evaluated.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scheduling system dynamically evaluates and adjusts workload placement based on real-time accelerator availability, performance metrics, and changing workload requirements. The system can adaptively select from multiple scheduling strategies and automatically adjust resource allocation decisions, enabling flexible response to dynamic cluster conditions while maintaining operational simplicity through automated decision-making.

Inventive Principle:
Principle #15Dynamics

2Productivity

If basic GPU count specification is used, then the deployment process is straightforward, but resource allocation efficiency and workload performance are suboptimal

Engineering Contradiction:
Improveresource allocation efficiencyVSAvoidsimplicity of workload specification
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The scheduling system performs self-service by automatically evaluating accelerator requirements, matching workloads to appropriate accelerators, and optimizing resource allocation without requiring manual intervention. The system autonomously processes complex multi-dimensional requirements, evaluates accelerator suitability based on detailed specifications, and makes intelligent scheduling decisions, thereby improving resource allocation efficiency while maintaining ease of operation through automation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the single parameter of GPU count into multiple independent parameters including accelerator type, architecture, model, resource allocation, service level agreements, and prioritized lists. This parameter expansion enables precise matching of workload requirements with accelerator capabilities, significantly improving resource allocation efficiency and workload performance while the system handles the increased complexity through automated parameter evaluation and optimization.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If fine-grained accelerator requirements are supported, then workload scheduling precision is improved, but the complexity of the scheduling framework increases

Engineering Contradiction:
Improveprecision of accelerator matchingVSAvoidcomplexity of scheduling framework
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The scheduling framework segments fine-grained accelerator requirements into distinct, independently evaluable dimensions including accelerator type, architecture, model, resource allocation, service level agreements, and prioritized lists. This segmentation enables precise matching by evaluating each dimension separately while maintaining overall scheduling coherence, achieving high measurement precision without proportionally increasing framework complexity through modular evaluation architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The scheduling framework implements a universal evaluation mechanism that handles multiple accelerator requirements dimensions through a single integrated system. This multi-functional approach allows the same scheduling infrastructure to process diverse requirement types (type, architecture, model, resources, SLAs, prioritized lists) uniformly, achieving precise accelerator matching while avoiding complexity multiplication through consistent evaluation patterns across all requirement dimensions.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12517765B2Framework for scheduling accelerator-enabled workloads
Publication Date: 2026.01.06 VMWARE INC
  • US12517765B2 patent drawing
  • US12517765B2 patent drawing
  • US12517765B2 patent drawing

AI summary

A framework that may be implemented by a workload orchestration platform for scheduling accelerator-enabled workloads on the accelerators in a cluster is provided. In one set of embodiments, the framework enables the platform to schedule accelerator-enabled workloads based on a multitude of user-provided, fine-grained accelerator requirements. In another set of embodiments, the framework enables the platform to automatically recommend an initial set of accelerator resource requirements for an accelerator-enabled workload and automatically right-size such requirements based on telemetry data collected during the workload's runtime.