Multi-Chip Accelerator Workload Balancing with DVFS

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face inefficiencies in controlling workload performance at the single-chip level due to disparities in workload partitioning among multiple accelerator chips, leading to suboptimal power consumption and throughput, especially in multi-chip systems like machine learning and high-performance computing environments.

Innovation Solution

A method and apparatus for controlling workload performance through dynamic voltage and frequency scaling (DVFS) by monitoring and adjusting the performance speed of individual accelerator chips based on performance data, power consumption, and workload partitioning, using a master controller to optimize power distribution and synchronization across multiple chips.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If clock frequency is increased by raising chip voltage, then workload performance is improved, but temperature and power consumption increase and chip longevity shortens

Engineering Contradiction:
Improveworkload throughputVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent implements dynamic voltage and frequency scaling (DVFS) that allows the accelerator chip to operate at different voltage-frequency combinations based on actual workload demands. The system dynamically adjusts operating parameters rather than maintaining fixed high performance, enabling the chip to optimize between power consumption and performance based on real-time conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the operating parameters (voltage and frequency) of the accelerator chip based on workload characteristics. By monitoring metrics such as compute-bound versus memory-bound operations, the system adjusts voltage-frequency combinations to match actual performance needs, avoiding unnecessary power consumption during low-demand periods while maintaining high performance when required.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If clock frequency is increased by raising chip voltage, then workload performance is improved, but temperature increases

Engineering Contradiction:
Improveworkload throughputVSAvoidchip temperature
Core Design Contradiction:
ProductivityVSTemperature

Solution Approach 1:

The system dynamically adjusts voltage and frequency based on thermal conditions and workload demands. When temperature approaches thresholds or workloads are lighter, the system reduces operating parameters to lower heat generation while maintaining acceptable performance levels.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent modifies operating parameters (voltage and frequency) in response to thermal conditions. By reducing voltage-frequency combinations when temperature is elevated or workloads are not compute-intensive, the system manages thermal output while preserving performance during acceptable thermal operating ranges.

Inventive Principle:
Principle #35Parameter changes

3Use of energy by moving object

If DVFS is used to dynamically adjust clock frequency, then power consumption is optimized, but response time to establish new voltage-frequency set point exceeds the period of time needed

Engineering Contradiction:
Improvepower consumptionVSAvoidresponse time for voltage-frequency adjustment
Core Design Contradiction:
Use of energy by moving objectVSLoss of time

Solution Approach 1:

The system pre-configures multiple voltage-frequency combinations that can be quickly switched between based on workload predictions. By having pre-established operating points and using predictive algorithms to anticipate workload changes, the system reduces adjustment latency and avoids the delay of establishing new set points during critical performance windows.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If multiple accelerators work together on a workload, then system capacity increases, but increasing clock speed for one chip does not result in improved throughput when another accelerator is working slower

Engineering Contradiction:
Improvesystem throughputVSAvoidmulti-chip coordination complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system implements feedback mechanisms that monitor the performance of individual accelerators within the multi-chip system. By tracking which accelerators are compute-bound versus memory-bound and adjusting their operating parameters accordingly, the system optimizes overall throughput despite variations in individual chip performance. This feedback-driven coordination allows the system to maintain efficiency across heterogeneous accelerator performance levels.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250298431A1Large-Scale Accelerator System Energy Performance Optimization
Publication Date: 2025.09.25 GOOGLE LLC
  • US20250298431A1 patent drawing
  • US20250298431A1 patent drawing
  • US20250298431A1 patent drawing

AI summary

A method and system for controlling performance of a workload partitioned among a plurality of accelerator chips of a multi-chip system. One or more processors may receive performance speed data for each of the accelerator chips, obtain a model of the partitioned workload, determine a portion of the workload that is either overworked or underworked based on the model of the partitioned workload and the performance speed data for each of the plurality of accelerator chips, and adjust a performance speed of an accelerator chip that performs the portion of the partitioned workload that is either overworked or underworked.