CPU Performance Hints for QoS-Aware Inference Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Inference workloads in devices with power-saving policies often experience insufficient bandwidth, leading to performance degradation and inefficient resource allocation due to default real-time workload prioritization and lack of insight into actual resource utilization.

Innovation Solution

A power manager exposes an API to specify priority and QoS parameters, configuring clock speeds and operating voltages to optimize resource allocation and reduce power consumption, while considering workload statistics and operational data to ensure compliance with specified parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If power-saving policies are applied to CPUs, then power consumption is reduced, but inference workload performance deteriorates due to insufficient bandwidth

Engineering Contradiction:
Improvepower consumptionVSAvoidinference workload performance
Core Design Contradiction:
Use of energy by moving objectVSProductivity

Solution Approach 1:

The system dynamically adjusts CPU frequency and power settings based on workload type. Inference workloads receive higher frequency allocations compared to traditional workload types, allowing the system to optimize power consumption while maintaining performance for specific workload categories through dynamic parameter adjustment.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of CPU frequency allocation for different workload types. By assigning higher frequency ranges to inference workloads specifically, the system resolves the contradiction between power-saving mode and inference performance without affecting other workload types.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If default real-time workload prioritization is used, then real-time applications receive sufficient resources, but inference workloads experience resource starvation and performance degradation

Engineering Contradiction:
Improvereal-time application responsivenessVSAvoidinference workload execution efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system applies different resource allocation qualities to different workload types. Inference workloads receive a specific quality of service with higher frequency allocations, while other workload types maintain their traditional allocations, allowing localized optimization without compromising overall system reliability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the CPU resource allocation by workload type, creating distinct frequency allocation zones for inference workloads versus other workload types. This segmentation allows inference workloads to receive dedicated high-frequency resources while real-time applications maintain their priority status.

Inventive Principle:
Principle #1Segmentation

3Speed

If CPU frequency is increased to improve inference performance, then processing speed increases, but power consumption increases

Engineering Contradiction:
Improveinference processing speedVSAvoidCPU power consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts frequency based on workload type, providing high frequency only when inference workloads are detected. This dynamic adjustment allows the system to achieve high processing speeds for inference tasks while maintaining lower power consumption during other operational states.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies partial action by allocating high frequency resources specifically to inference workloads rather than maintaining high frequency for all workloads continuously. This selective application of high performance resources reduces overall power consumption while achieving the necessary speed for inference tasks.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250278314A1CPU Performance Hint for Inference Workloads
Publication Date: 2025.09.04 ADVANCED MICRO DEVICES INC
  • US20250278314A1 patent drawing
  • US20250278314A1 patent drawing
  • US20250278314A1 patent drawing

AI summary

A power manager of an apparatus receives priority and quality-of-service (QOS) parameters (e.g., latency, throughput) for an inference workload. An application, for instance, specifies the priority and QoS parameters for an inference workload to be processed using a central processing unit. The priority and QoS parameters are employed by the power manager as a basis to configure the power setting of the central processing unit. In particular, resource prioritization for central processing units is extended to both real-time and best-effort workloads to satisfy specified QoS parameters for inference workloads.