Multi-Chip Accelerator Workload Balancing with DVFS
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face inefficiencies in controlling workload performance at the single-chip level due to disparities in workload partitioning among multiple accelerator chips, leading to suboptimal power consumption and throughput, especially in multi-chip systems like machine learning and high-performance computing environments.
Innovation Solution
A method and apparatus for controlling workload performance through dynamic voltage and frequency scaling (DVFS) by monitoring and adjusting the performance speed of individual accelerator chips based on performance data, power consumption, and workload partitioning, using a master controller to optimize power distribution and synchronization across multiple chips.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If clock frequency is increased by raising chip voltage, then workload performance is improved, but temperature and power consumption increase and chip longevity shortens
Solution Approach 1:
The patent implements dynamic voltage and frequency scaling (DVFS) that allows the accelerator chip to operate at different voltage-frequency combinations based on actual workload demands. The system dynamically adjusts operating parameters rather than maintaining fixed high performance, enabling the chip to optimize between power consumption and performance based on real-time conditions.
Solution Approach 2:
The patent changes the operating parameters (voltage and frequency) of the accelerator chip based on workload characteristics. By monitoring metrics such as compute-bound versus memory-bound operations, the system adjusts voltage-frequency combinations to match actual performance needs, avoiding unnecessary power consumption during low-demand periods while maintaining high performance when required.
2Productivity
If clock frequency is increased by raising chip voltage, then workload performance is improved, but temperature increases
Solution Approach 1:
The system dynamically adjusts voltage and frequency based on thermal conditions and workload demands. When temperature approaches thresholds or workloads are lighter, the system reduces operating parameters to lower heat generation while maintaining acceptable performance levels.
Solution Approach 2:
The patent modifies operating parameters (voltage and frequency) in response to thermal conditions. By reducing voltage-frequency combinations when temperature is elevated or workloads are not compute-intensive, the system manages thermal output while preserving performance during acceptable thermal operating ranges.
3Use of energy by moving object
If DVFS is used to dynamically adjust clock frequency, then power consumption is optimized, but response time to establish new voltage-frequency set point exceeds the period of time needed
Solution Approach 1:
The system pre-configures multiple voltage-frequency combinations that can be quickly switched between based on workload predictions. By having pre-established operating points and using predictive algorithms to anticipate workload changes, the system reduces adjustment latency and avoids the delay of establishing new set points during critical performance windows.
4Productivity
If multiple accelerators work together on a workload, then system capacity increases, but increasing clock speed for one chip does not result in improved throughput when another accelerator is working slower
Solution Approach 1:
The system implements feedback mechanisms that monitor the performance of individual accelerators within the multi-chip system. By tracking which accelerators are compute-bound versus memory-bound and adjusting their operating parameters accordingly, the system optimizes overall throughput despite variations in individual chip performance. This feedback-driven coordination allows the system to maintain efficiency across heterogeneous accelerator performance levels.
Data Source
AI summary
A method and system for controlling performance of a workload partitioned among a plurality of accelerator chips of a multi-chip system. One or more processors may receive performance speed data for each of the accelerator chips, obtain a model of the partitioned workload, determine a portion of the workload that is either overworked or underworked based on the model of the partitioned workload and the performance speed data for each of the plurality of accelerator chips, and adjust a performance speed of an accelerator chip that performs the portion of the partitioned workload that is either overworked or underworked.


