QoS Partition Configuration for Compute Resource Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional scheduling techniques lack insight into resource utilization, leading to inefficient use and increased power consumption, especially in real-time execution scenarios where multiple workload requests cause service degradation and resource overallocation.
Innovation Solution
A scheduler exposes a quality-of-service (QoS) API that allows applications to specify parameters like latency and throughput, enabling the configuration of partitions within a hardware compute unit to optimize resource allocation and minimize power consumption based on workload statistics and operation data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional scheduling techniques are used to maximize available resources for each workload request, then resource utilization is improved, but power consumption increases and service degradation occurs in real-time execution scenarios
Solution Approach 1:
The patent changes the scheduling parameters from binary priority levels to continuous QoS parameters (latency, throughput, jitter). This allows the scheduler to allocate resources based on actual performance requirements rather than fixed priorities, reducing overallocation and associated power consumption while maintaining service quality guarantees for real-time workloads.
Solution Approach 2:
The patent introduces feedback mechanisms where applications specify QoS parameters and the scheduler monitors actual resource utilization and service quality. This feedback loop enables dynamic adjustment of resource allocation to match actual needs, preventing both overallocation (and wasted power) and underallocation (and service degradation).
2Reliability
If conventional scheduling techniques allocate resources to meet real-time workload requests, then service quality is maintained, but resource overallocation and inefficiency occur
Solution Approach 1:
The patent replaces fixed priority scheduling with QoS-based scheduling that uses continuous parameters (latency, throughput, jitter) to define service quality requirements. This allows precise matching of resource allocation to actual service quality needs, eliminating resource overallocation while maintaining reliable service for real-time workloads.
Solution Approach 2:
The patent makes resource allocation dynamic by allowing applications to specify varying QoS parameters and the scheduler to adjust resource allocation in real-time based on actual workload characteristics and system state. This dynamic approach maintains service quality while optimizing resource utilization efficiency.
3Speed
If applications default to real-time priority to ensure timely processing, then latency requirements are met, but multiple workload requests cause service degradation
Solution Approach 1:
The patent changes the scheduling approach from binary real-time vs. non-real-time priorities to continuous QoS parameters including latency, throughput, and jitter. This allows the system to differentiate between workloads based on actual performance requirements, enabling multiple real-time workloads to coexist without service degradation by allocating resources based on their specific QoS needs rather than a uniform real-time priority.
Data Source
AI summary
A scheduler of an apparatus exposes an application programming interface (API) usable to specify quality-of-service (QoS) parameters, e.g., latency, throughput, and so forth. An application, for instance, specifies the QoS parameters for a workload to be processed using a hardware compute unit. The QoS parameters are employed by the scheduler as a basis to configure a partition within a hardware compute unit. The partition is configured such that processing resources that are available via the partition to process the workload comply with the specified quality-of-service.


