Server GPU Power Control via Dynamic Utilization Level Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing power supply control methods for AI servers with multiple GPUs struggle to manage peak current demands during high-performance computing, leading to potential system shutdowns or restarts while reducing overall computing performance.
Innovation Solution
A power supply control method that divides the system main power supply utilization rate into different levels, with corresponding GPU power control policies to suppress computing capability and reduce power consumption as utilization rates increase, thereby preventing system shutdowns and maintaining performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If the number of GPUs is increased to improve computing performance, then the computing capability of the server is improved, but the current control becomes difficult and the system main power supply may shut down or restart
Solution Approach 1:
The patent segments the power supply control by dividing GPUs into different power consumption groups and implementing separate management for each group. The management unit controls each GPU group independently through separate power supply channels, preventing a single GPU's peak current from affecting the entire system. This segmentation approach maintains system reliability while supporting multiple GPUs.
Solution Approach 2:
The patent applies local quality by providing different power supply control strategies for different GPU groups based on their power consumption characteristics. High-power GPUs receive dedicated power management with higher current thresholds, while low-power GPUs use different control parameters. This localized control optimizes both computing performance and power supply stability.
2Reliability
If a management unit is added to control GPU EDPP, then the system main power supply shutdown or restart is prevented, but the computing performance of the entire AI server is greatly reduced
Solution Approach 1:
The patent implements dynamic power supply control where the management unit adjusts power allocation based on real-time system conditions. When power supply utilization is high, the system dynamically switches to power brake mode for specific GPUs. When utilization is low, GPUs operate at full performance. This dynamic approach maintains reliability while preserving computing performance during normal operation.
Solution Approach 2:
The management unit monitors power supply utilization and automatically adjusts GPU power consumption without external intervention. When detecting high utilization, it triggers power brake signals to specific GPUs to reduce their power consumption. This self-service mechanism prevents system shutdowns while maintaining optimal computing performance during normal conditions.
3Reliability
If a large capacitor is added to control GPU EDPP, then the system main power supply shutdown or restart is prevented, but the device complexity increases
Solution Approach 1:
The patent introduces a management unit as an intermediary between the power supply unit and GPUs. This management unit monitors power supply utilization and controls GPU power consumption through software-based power brake signals, replacing the need for large physical capacitors. The intermediary approach reduces hardware complexity while maintaining power supply stability.
Data Source
AI summary
A power supply control method, system and device for a server are provided. The method includes: dividing a utilization rate of a system main power supply into different levels in advance, and setting a GPU power control policy corresponding to a respective one of the different levels of the utilization rate of the system main power supply, wherein a suppression degree, on a computing capability of GPUs in a system, of the set GPU power control policy increases with the increase in the level of the utilization rate of the system main power supply; acquiring an actual utilization rate of the system main power supply, and determining a target utilization rate level corresponding to the actual utilization rate; and performing power supply control on the GPUs in the system according to the GPU power control policy corresponding to the target utilization rate level.

