ML Runtime Tuning of Processing Units for Power-Latency Balance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing power-performance management in computing devices is reactive and does not leverage learned behaviors, leading to inefficiencies and latency in processing unit operations.

Innovation Solution

A machine learning-based optimization model is trained to proactively adjust processing unit settings using activity data, incorporating metrics like utilization and memory usage, to balance performance and power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If reactive power-performance management is used, then system responsiveness is maintained, but latency and inefficiency increase

Engineering Contradiction:
Improvesystem responsivenessVSAvoidlatency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by proactively adjusting processing unit settings before performance degradation occurs. The optimization model continuously monitors activity data and preemptively modifies operational parameters based on learned patterns, preventing latency issues rather than reacting to them after they occur.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements continuous feedback loops where the optimization model monitors runtime activity data, evaluates performance metrics, and adjusts processing unit settings dynamically. This closed-loop feedback mechanism enables the system to maintain responsiveness while minimizing latency through data-driven decisions.

Inventive Principle:
Principle #23Feedback

2Productivity

If processing unit settings are adjusted dynamically, then performance efficiency is improved, but system complexity increases

Engineering Contradiction:
Improveperformance efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The optimization model serves as an intermediary layer between the processing units and the control system. It abstracts the complexity of dynamic settings adjustment by encapsulating the decision-making logic within the model, which processes activity data and generates optimized settings without requiring complex control circuitry or manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system achieves self-service through the optimization model that autonomously monitors runtime activity data, determines optimal settings, and adjusts processing unit parameters without external intervention. This self-managing capability improves performance efficiency while avoiding the complexity of external control mechanisms.

Inventive Principle:
Principle #25Self-service

3Use of energy by moving object

If machine learning models are used for optimization, then power consumption is reduced, but computational overhead increases

Engineering Contradiction:
Improvepower consumptionVSAvoidcomputational overhead
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The system applies partial action by using lightweight machine learning models that perform only the essential optimization functions needed. Rather than implementing comprehensive complex models, the system uses simplified models that provide sufficient optimization benefit while minimizing computational overhead and training requirements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12481542B2Optimizing runtime configuration of processing units using machine learning
Publication Date: 2025.11.25 NVIDIA CORP
  • US12481542B2 patent drawing
  • US12481542B2 patent drawing
  • US12481542B2 patent drawing

AI summary

Disclosed are apparatuses, systems, and techniques that use machine learning techniques for determination and tuning of runtime settings of processing units. In one embodiment, a computing device, which includes one or more processing units, processes, using a machine learning model, a runtime activity data to generate settings for the processing unit(s). The runtime activity data characterizes an execution of a computing application on the processing unit(s). The computing device then modifies, using the generated settings, the execution of the computing application on the processing unit(s).