CPU Performance Hints for QoS-Aware Inference Workloads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Inference workloads in devices with power-saving policies often experience insufficient bandwidth, leading to performance degradation and inefficient resource allocation due to default real-time workload prioritization and lack of insight into actual resource utilization.
Innovation Solution
A power manager exposes an API to specify priority and QoS parameters, configuring clock speeds and operating voltages to optimize resource allocation and reduce power consumption, while considering workload statistics and operational data to ensure compliance with specified parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If power-saving policies are applied to CPUs, then power consumption is reduced, but inference workload performance deteriorates due to insufficient bandwidth
Solution Approach 1:
The system dynamically adjusts CPU frequency and power settings based on workload type. Inference workloads receive higher frequency allocations compared to traditional workload types, allowing the system to optimize power consumption while maintaining performance for specific workload categories through dynamic parameter adjustment.
Solution Approach 2:
The patent changes the parameter of CPU frequency allocation for different workload types. By assigning higher frequency ranges to inference workloads specifically, the system resolves the contradiction between power-saving mode and inference performance without affecting other workload types.
2Reliability
If default real-time workload prioritization is used, then real-time applications receive sufficient resources, but inference workloads experience resource starvation and performance degradation
Solution Approach 1:
The system applies different resource allocation qualities to different workload types. Inference workloads receive a specific quality of service with higher frequency allocations, while other workload types maintain their traditional allocations, allowing localized optimization without compromising overall system reliability.
Solution Approach 2:
The patent segments the CPU resource allocation by workload type, creating distinct frequency allocation zones for inference workloads versus other workload types. This segmentation allows inference workloads to receive dedicated high-frequency resources while real-time applications maintain their priority status.
3Speed
If CPU frequency is increased to improve inference performance, then processing speed increases, but power consumption increases
Solution Approach 1:
The system dynamically adjusts frequency based on workload type, providing high frequency only when inference workloads are detected. This dynamic adjustment allows the system to achieve high processing speeds for inference tasks while maintaining lower power consumption during other operational states.
Solution Approach 2:
The patent applies partial action by allocating high frequency resources specifically to inference workloads rather than maintaining high frequency for all workloads continuously. This selective application of high performance resources reduces overall power consumption while achieving the necessary speed for inference tasks.
Data Source
AI summary
A power manager of an apparatus receives priority and quality-of-service (QOS) parameters (e.g., latency, throughput) for an inference workload. An application, for instance, specifies the priority and QoS parameters for an inference workload to be processed using a central processing unit. The priority and QoS parameters are employed by the power manager as a basis to configure the power setting of the central processing unit. In particular, resource prioritization for central processing units is extended to both real-time and best-effort workloads to satisfy specified QoS parameters for inference workloads.


