GPU cluster starting system based on DPM

By gradually adjusting the GPU cluster power level using a DPM-based dynamic power management unit, the current surge problem during GPU cluster startup was resolved, improving startup stability and security.

CN121658085APending Publication Date: 2026-03-13沐曦科技(成都)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-27
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

When a large-scale GPU cluster starts up, the instantaneous current demand is too high, which can lead to problems such as power supply overload, voltage drop, hardware damage and increased heat dissipation pressure.

Method used

A dynamic power management unit based on DPM is adopted to control the increase of current and voltage by gradually increasing the level of the GPU cluster, avoiding instantaneous current surges and ensuring that current changes are within an acceptable range.

Benefits of technology

It reduces the instantaneous current surge during GPU cluster startup, improving startup stability and safety, and preventing power overload and hardware damage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121658085A_ABST
    Figure CN121658085A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of GPU cluster optimization, in particular to a GPU cluster starting system based on a DPM, the system comprises a GPU cluster, a dynamic power management unit, a processor and a memory storing a computer program, and when the computer program is executed by the processor, the following steps are realized: when a request is received, initializing a gear identifier, and when the request is received, starting the GPU cluster; setting a current gear through the dynamic power supply management unit according to the gear identifier, executing the target use case according to the current gear, when a first preset condition is met, executing the step of updating the gear identifier and returning to execute the step of setting the current gear through the dynamic power supply management unit according to the gear identifier until a second preset condition is met, and completing the starting of the GPU cluster. According to the method, the running frequency and voltage are gradually increased by controlling the DPM gear, so that the starting current of the GPU cluster is gradually increased, the instantaneous current impact during starting of the GPU cluster is reduced, and the stability and safety during starting of the GPU cluster are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of GPU cluster optimization technology, and in particular to a GPU cluster startup system based on DPM. Background Technology

[0002] In the use case of large-scale GPU clusters, when a single GPU chip changes from an idle state to a running state, the power consumption of the GPU chip will instantly rise from a low value to a high value, for example, from 50W to 300W. Since a large-scale GPU cluster contains a large number of GPU chips, when the large-scale GPU cluster changes from an idle state to a running state, the power consumption of all the GPU chips contained in it needs to increase instantaneously, resulting in a significant increase in the demand for instantaneous current.

[0003] The above situations may cause the instantaneous current demand to exceed the power supply system's carrying capacity, leading to power overload, which in turn may cause power outages or voltage drops, resulting in unstable power supply or even power outages. Moreover, when the instantaneous current exceeds the threshold limit set by the circuit design, it will trigger the circuit breaker or fuse to automatically cut off the current to prevent the circuit from overheating or the equipment from being damaged, which may cause the GPU chip to fail to start and run normally. In addition, the simultaneous startup of a large number of GPU chips will cause the voltage of the power supply line to drop sharply, which may affect the performance and stability of other devices on the power supply line. Furthermore, current overload may also cause hardware components to overheat, such as the motherboard, power supply unit, and GPU chip, accelerating hardware aging or even damage. Finally, the simultaneous startup of a large number of GPU chips also means that there is a greater demand for heat dissipation, which increases the pressure on the cooling system.

[0004] Therefore, how to reduce the instantaneous current surge during GPU cluster startup, thereby improving the stability and security of GPU cluster startup, has become an urgent problem to be solved. Summary of the Invention

[0005] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:

[0006] A GPU cluster boot system based on DPM, the system comprising: a GPU cluster, a Dynamic Power Management Unit (DPM), a processor, and a memory storing a computer program, wherein the Dynamic Power Management Unit includes M preset levels {a1, a2, ..., a...} m , ..., a M}, a m For the m-th preset gear in the dynamic power management unit, where m is an integer in the range [1, M], when the computer program is executed by the processor, the following steps are implemented:

[0007] S101, when a request to execute the target use case using the GPU cluster is received, initialize the gear identifier n=1.

[0008] S102, the current gear is set to a through the dynamic power management unit. n a n This refers to the nth preset gear in the dynamic power management unit A.

[0009] S103, execute the target use case in the current gear position, and when the first preset condition is met, execute step S104.

[0010] S104, update n = n + 1, return to execute steps S102 to S103 until the second preset condition is met, and complete the startup of the GPU cluster.

[0011] Compared with the prior art, the present invention has significant advantages. Through the above technical solution, the GPU cluster boot system based on DPM provided by the present invention achieves considerable technological progress and practicality, and has broad industrial application value. It has at least the following advantages:

[0012] This invention provides a GPU cluster boot system based on Dynamic Power Management (DPM). The system includes: a GPU cluster, a Dynamic Power Management (DPM), a processor, and a memory storing a computer program. The DPM includes M preset power levels {a1, a2, ..., a...}. m , ..., a M}, a m For the m-th preset gear in the dynamic power management unit, where m is an integer in the range [1, M], when the computer program is executed by the processor, the following steps are implemented: S101, when a request to execute a target use case using the GPU cluster is received, the gear identifier n = 1 is initialized; S102, the current gear is set to a through the dynamic power management unit. n a n For the nth preset gear in the dynamic power management unit A, in step S103, the target use case is executed with the current gear. When the first preset condition is met, step S104 is executed. In step S104, n = n + 1 is updated, and the execution returns to steps S102 to S103 until the second preset condition is met, and the startup of the GPU cluster is completed.

[0013] It can be seen that when the GPU cluster starts to execute the target use case, the operating frequency and voltage are gradually increased by controlling the DPM level, thereby gradually increasing the current of the GPU cluster during startup. At the same time, it ensures that the instantaneous current change during DPM level switching is not too large, avoiding excessive instantaneous current surge, reducing the instantaneous current surge during GPU cluster startup, and thus improving the stability and security of GPU cluster startup. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a schematic diagram illustrating the process of a computer program being executed by a processor in a GPU cluster startup system based on DPM, as provided in an embodiment of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] This embodiment provides a GPU cluster boot system based on DPM. The system includes: a GPU cluster, a Dynamic Power Management Unit (DPM), a processor, and a memory storing a computer program. The DPM includes M preset power levels {a1, a2, ..., a...}. m , ..., a M}, a m This refers to the m-th preset level in the dynamic power management unit, where m is an integer in the range [1, M]. See [link / reference]. Figure 1 This is a flowchart illustrating the execution of a computer program by a processor in a GPU cluster startup system based on DPM, according to an embodiment of the present invention. When the computer program is executed by the processor, the following steps are implemented:

[0018] S101, When a request to execute the target use case using the GPU cluster is received, initialize the gear identifier n=1;

[0019] S102, the current gear is set to a through the dynamic power management unit. n a nIt is the nth preset gear in the dynamic power management unit A;

[0020] S103. Execute the target use case at the current gear. When the first preset condition is satisfied, execute step S104;

[0021] S104. Update n = n + 1, and return to execute steps S102 to S103 until the second preset condition is satisfied, and complete the startup of the GPU cluster.

[0022] Among them, the dynamic power management unit (Dynamic Power Management, DPM) can be used to run the GPU cluster at a fixed gear. The operating frequencies and voltages corresponding to different gears are different. Accordingly, the currents of the GPU cluster running at different gears are different.

[0023] Specifically, executing the target use case at the current gear may mean that the GPU cluster executes the target use case at the current gear.

[0024] In the prior art, the dynamic power management unit is usually used to control the power consumption of the GPU chip within a set range. In this embodiment, by performing gear switching through the dynamic power management unit, the instantaneous current impact when the GPU cluster changes from the idle state to the running state is dispersed to each moment of gear switching, so that the instantaneous current impact at any moment is maintained within an acceptable range, thereby avoiding excessive instantaneous current impact, reducing the instantaneous current impact during the startup of the GPU cluster, and further improving the stability and security during the startup of the GPU cluster.

[0025] In a specific implementation manner, the first preset condition is: the gear identifier n < M and the execution duration of the current gear is greater than or equal to a preset duration threshold.

[0026] Among them, the gear identifier n < M indicates that the current gear is not the highest gear of the dynamic power management unit. The execution duration of the current gear can be regarded as the delay of gear switching. It can be seen that if the delay is zero, it is equivalent to the dynamic power management unit directly switching from the lowest gear to the highest gear. The lowest gear is usually the gear corresponding to the idle state of the GPU cluster. That is, when the delay is zero, it is the traditional method. Therefore, in this embodiment, the duration threshold should be a number greater than zero.

[0027] Specifically, introducing a delay when controlling the gear switching of the dynamic power management unit is also convenient for adapting to the response time of the power system and ensuring the stability of the power system to supply power to the GPU cluster.

[0028] In a specific implementation manner, the second preset condition is the gear identifier n = M.

[0029] When the gear indicator n = M, it means that the current gear is the highest gear of the dynamic power management unit. At this time, the GPU cluster can execute the target use case according to the current gear without switching gears.

[0030] In one specific implementation, the duration threshold t is obtained by mapping the power system redundancy C through a preset mapping function.

[0031] The power system redundancy C can be obtained from the power system side. In this embodiment, for ease of representation, the power system redundancy C is represented by a normalized value. When the power system redundancy C is 0, it means that the power system has no redundancy, and the power system response time is the longest. When the power system redundancy C is 1, it means that the power system is completely redundant, and the power system response time is the shortest. That is, the larger the power system redundancy, the shorter the power system response time, and the smaller the power system redundancy, the longer the power system response time.

[0032] In one specific implementation, the preset mapping function is:

[0033] t=(F / (e kC-b +1))×T, where F=e -b +1, k is the first adjustment coefficient, b is the second adjustment coefficient, and T is the theoretical maximum delay.

[0034] Where k can be a value greater than one, in this embodiment k is set to 5 and b is set to 0.5. k and b can be used to adjust the slope of the mapping function. The implementer can adjust the values ​​of k and b according to actual needs. T is the theoretical maximum latency. Since the execution efficiency of the target use case will be reduced when the GPU cluster is started using the method of this embodiment and the target use case is executed at a non-maximum level, the theoretical maximum latency needs to be set in order to ensure that the execution efficiency loss of the target use case is within the acceptable range of the user corresponding to the target use case.

[0035] Specifically, when C = 0, t = T, and when C = 1, t approaches 0. Since the mapping result of this mapping function is always greater than zero, it also ensures that the time threshold should be a number greater than zero.

[0036] In one specific implementation, the difference between the current time and the start time of the execution of the target use case in the current gear is taken as the execution duration.

[0037] In one specific implementation, a m Corresponding to reference voltage b m b i <b i+1 , where i is an integer in the range [1, M-1], b ib is the reference voltage corresponding to the i-th preset level in the dynamic power management unit. i+1 The reference voltage is the (i+1)th preset level in the dynamic power management unit.

[0038] In order to ensure that the voltage of the GPU cluster increases gradually during gear switching, constraint b is required. i <b i+1 That is, the reference voltage in the forward gear is less than the reference voltage in the reverse gear.

[0039] In another implementation, the reference frequency corresponding to the i-th preset level in the dynamic power management unit is denoted as q. i Constraint b i ≤b i+1 And q i <q i+1 .

[0040] In one specific implementation, the characteristic is that M = 10.

[0041] In this embodiment, the preset power management unit can be set to 10 levels, and the implementer can adjust the number of preset levels according to the actual situation.

[0042] In this embodiment, when the GPU cluster starts executing the target use case, the operating frequency and voltage are gradually increased by controlling the DPM level, thereby gradually increasing the current of the GPU cluster during startup. At the same time, it is ensured that the instantaneous current change during DPM level switching is not too large, avoiding excessive instantaneous current surges and reducing the instantaneous current surge during GPU cluster startup, thereby improving the stability and security of GPU cluster startup.

[0043] While specific embodiments of the invention have been described in detail by way of example, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the invention. The scope of the invention is defined by the appended claims.

Claims

1. A GPU cluster startup system based on DPM, characterized in that, The system includes: a GPU cluster, a dynamic power management unit (DPM), a processor, and a memory storing a computer program, wherein the dynamic power management unit includes M preset levels {a1, a2, ..., a...} m , ..., a M }, a m For the m-th preset gear in the dynamic power management unit, where m is an integer in the range [1, M], when the computer program is executed by the processor, the following steps are implemented: S101, When receiving a request to execute a target case using the GPU cluster, initialize the gear identifier n = 1; S102, the current gear is set to a through the dynamic power management unit. n a n This refers to the nth preset gear in the dynamic power management unit A; S103, Execute the target case at the current gear. When the first preset condition is met, execute step S104; S104, Update n = n + 1, and return to execute steps S102 to S103 until the second preset condition is met, and complete the startup of the GPU cluster.

2. The GPU cluster startup system based on DPM according to claim 1, characterized in that, The first preset condition is: the gear identifier n < M and the execution duration of the current gear is greater than or equal to a preset duration threshold.

3. The GPU cluster startup system based on DPM according to claim 1, characterized in that, The second preset condition is that the gear identifier n = M.

4. The GPU cluster startup system based on DPM according to claim 2, characterized in that, The duration threshold t is obtained by mapping the power system redundancy C through a preset mapping function.

5. The GPU cluster startup system based on DPM according to claim 4, characterized in that, The preset mapping function is: t=(F / (e kC-b +1))×T, where F=e -b +1, k is the first adjustment coefficient, b is the second adjustment coefficient, and T is the theoretical maximum delay.

6. The GPU cluster startup system based on DPM according to claim 2, characterized in that, Subtract the start time of executing the target case at the current gear from the current time, and use the difference result as the execution duration.

7. The GPU cluster startup system based on DPM according to claim 1, characterized in that, a m Corresponding to reference voltage b m b i <b i+1 , where i is an integer in the range [1, M-1], b i b is the reference voltage corresponding to the i-th preset level in the dynamic power management unit. i+1 The reference voltage is the (i+1)th preset level in the dynamic power management unit.

8. The GPU cluster startup system based on DPM according to claim 1, characterized in that, M=10。