An AI server energy-saving regulation method based on GPU load prediction
Patent Information
- Application Number
- CN202611233555.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-18
AI Technical Summary
[0003]然而,在实际的分布式模型训练场景中存在一个隐蔽的技术缺陷:依赖宏观利用率轮询的被动式供电调控机制无法捕捉底层微架构级别的瞬态阶段切换,导致严重的瞬时能源浪费与无效热量堆积;具体而言,机器学习训练任务的执行特征决定了图形处理器的负载在微秒级时间尺度上呈现剧烈的脉冲式波动
1.本发明通过周期性采集图形处理器内部的线程束调度器分发率、二级缓存未命中率以及节点间互联串行解串器激活特征等微架构运行状态序列,并结合时序预测网络模型输出未来时间周期的图形处理器负载预测张量,能够提前感知底层硬件的瞬态阶段切换;在此基础上,通过向量反向比对触发生成微秒级通信阻尼阶段标识,精准识别图形处理器在密集计算与全局参数归约同步之间的微小阶段切换,解决了利用率轮询机制因采样延迟与粒度不足而无法捕捉微秒级指令流停滞的问题,避免了电压调节模块在通信等待期内持续满载输出无效电压所导致的瞬时能源浪费与无效热量堆积。
Smart Images

Figure CN122776964A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and specifically to an energy-saving control method for AI servers based on GPU load prediction. Background Technology
[0002] With the widespread application of large language models with trillions of parameters in the field of artificial intelligence, data center server clusters bear extremely large matrix tensor operations and multi-node gradient synchronization tasks. In a typical distributed machine learning training process, model training follows an iterative cycle of forward propagation, loss calculation, backpropagation, gradient aggregation, and parameter update. The backpropagation stage generates intensive matrix multiplication operations, while the gradient aggregation stage requires all participating nodes to exchange and synchronize gradient parameters through communication primitives. To support the above-mentioned business scenarios that are both computationally and communication-intensive, servers are equipped with high-power graphics processing unit (GPU) components. Due to the extremely high peak power consumption of GPUs, related technologies typically set static power limits in the baseboard management controller or rely on dynamic voltage and frequency adjustment mechanisms for passive intervention. The execution logic of this technology is as follows: the polling probe in the control panel reads the macroscopic overall utilization rate of the GPU and the core temperature sensor values at fixed time intervals. Only when the macroscopic utilization rate is continuously high or the temperature is close to the heat dissipation limit will the controller issue a command to adjust the power supply voltage.
[0003] However, there is a hidden technical flaw in actual distributed model training scenarios: the passive power supply regulation mechanism that relies on macro-level utilization polling cannot capture the transient phase switching at the underlying micro-architecture level, resulting in serious instantaneous energy waste and ineffective heat accumulation; specifically, the execution characteristics of machine learning training tasks determine that the load of the graphics processor exhibits drastic pulse-like fluctuations on a microsecond time scale. For example, in a data-parallel training task of a large language model with hundreds of billions of parameters using the Adam optimizer, after the backpropagation phase of each training step, all GPUs must execute the AllReduce communication primitive to aggregate gradients. At this time, the computational cores instantly switch from a fully loaded matrix multiplication operation to an idle state waiting for communication data to be ready. This waiting period usually lasts from tens to hundreds of microseconds, but the macroscopic utilization probe still counts this interval as high load due to insufficient sampling granularity. Another example is when using a pipelined parallel model training strategy, pipeline bubbles can occur due to differences in computational load among GPUs at different stages. GPUs that complete the forward propagation of the current micro-batch first will enter an idle waiting state earlier, waiting for the upstream node to complete its computation before receiving input data for the next stage. During this period, its computational core is completely idle, but the power supply voltage remains at a certain level. For example, when the graphics processor performs matrix multiplication, if the required weight data fails to hit the on-chip L2 cache and needs to be reloaded from video memory, the computing core will briefly idle due to data transfer delays. At this time, the overall utilization rate may still show as over 85%, but in reality, the computing core is in a semi-idle state. When the graphics processor switches from intensive computing to the communication stage of waiting for cross-node data synchronization, the underlying thread scheduler will enter a very short pause cycle due to waiting for main memory or bus data. The macroscopic probes of related technologies are limited by sampling delay and granularity and cannot identify such microsecond-level low-level instruction stream pauses, causing the voltage regulation module to still output extremely high invalid voltage at full load during the long communication damping waiting period. This micro-control blind spot causes the cluster to waste huge amounts of power during non-computation cycles and causes excessive thermal fatigue of the power supply circuit. Summary of the Invention
[0004] To address the aforementioned technical issues, this paper provides an AI server energy-saving control method based on GPU load prediction, which solves the problems described above.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for energy-saving control of AI servers based on GPU load prediction includes: The microarchitecture operating state sequence of the graphics processor is periodically collected; wherein, the microarchitecture operating state sequence includes thread bundle scheduler dispatch rate, L2 cache miss rate, and inter-node interconnect serializer deserializer activation characteristics; The microarchitecture running state sequence within a continuous time window is input into a pre-defined time-series prediction network model to obtain deep spatiotemporal correlation features and output a graphics processor load prediction tensor for future time periods. Based on the thread bundle scheduler distribution rate and the node interconnect serial deserializer activation characteristics in the microarchitecture running state sequence, a vector reverse comparison is performed to trigger the generation of a microsecond-level communication damping stage identifier. Based on the graphics processor load prediction tensor, combined with the microsecond-level communication damping stage identifier and the preset thermal resistance-capacitance time constant, the target power supply pulse width modulation parameters are generated. The target power supply pulse width modulation parameters are sent directly to the voltage regulation module at the bottom layer of the graphics processor via the out-of-band management bus, overwriting the current core output voltage and clock frequency.
[0006] Furthermore, the periodic acquisition of the microarchitectural operating state sequence of the graphics processor includes: The thread beam scheduler dispatch rate of successfully issued instructions per unit clock cycle is read from the graphics processor's internal performance counter register via the baseboard management controller. Read the on-chip memory management unit and read the L2 cache miss rate that caused the main memory access wait; Read the physical layer link training status of the high-speed bus interface for peripheral component interconnection, and based on the physical layer link training status, read the activation feature of the inter-node interconnection serial deserializer that is currently active.
[0007] Furthermore, the step of inputting the microarchitecture operating state sequence within a continuous time window into a pre-defined temporal prediction network model to obtain deep spatiotemporal correlation features and outputting a graphics processor load prediction tensor for future time periods includes: Assign corresponding timestamp information to the thread bundle scheduler distribution rate, the secondary cache miss rate, and the inter-node interconnect serial deserializer activation feature; Based on the timestamp information, the thread bundle scheduler distribution rate, the secondary cache miss rate, and the inter-node interconnect serial deserializer activation features are concatenated into a multi-directional state matrix that includes the time direction. The multi-directional state matrix is input into the long short-term memory network to determine the parameters of the forget gate and update gate, and the expected resource usage ratio is generated as the graphics processor load prediction tensor.
[0008] Furthermore, the step of generating target power supply pulse width modulation parameters based on the graphics processor load prediction tensor, combined with the microsecond-level communication damping stage identifier and the preset thermal resistance-capacitance time constant, includes: Obtain the equivalent thermal resistance and equivalent capacitance of the power supply circuit of the target data center server motherboard, and obtain the preset thermal resistance-capacitance time constant, wherein the thermal resistance-capacitance time constant reflects the inherent decay rate of the power supply component due to temperature change. When it is determined that the microsecond-level communication damping stage identifier is in a valid active state, the graphics processor load prediction tensor is mapped to the lowest sustaining voltage level. Based on the thermal resistance-capacitance time constant, a smoothing filter is applied to the lowest sustaining voltage level to generate the target power supply pulse width modulation parameters.
[0009] Furthermore, the step of directly sending the target power supply pulse width modulation parameters to the voltage regulation module at the graphics processor's underlying layer via the out-of-band management bus to overwrite the current core output voltage and clock frequency includes: The target power supply pulse width modulation parameters are encapsulated into a hexadecimal control frame conforming to the power management bus protocol format; The hexadecimal control frame is sent to the digital control pin of the voltage regulation module via the integrated circuit's built-in bus. The multi-phase power supply circuit inside the voltage regulation module is forcibly triggered to perform a phase-switching operation, reducing the number of power supply phases output in parallel, and simultaneously lowering the core clock frequency of the graphics processor.
[0010] Furthermore, after generating the target power supply pulse width modulation parameters, the method further includes: Establish a power token scheduler that spans multiple computing nodes within the cluster; The power token scheduler obtains the current remaining power supply capacity of the first computing node and encapsulates the current remaining power supply capacity into power supply credit points. When the power token scheduler detects that the second computing node has triggered a high-load tensor task and the local power supply of the second computing node has reached the power wall threshold, the power credit of the first computing node is transferred to the second computing node via high-speed Ethernet.
[0011] Furthermore, the step of performing a vector reverse comparison based on the thread bundle scheduler distribution rate and the inter-node interconnect serializer / deserializer activation characteristics in the microarchitecture running state sequence to trigger the generation of a microsecond-level communication damping stage identifier includes: Determine the first-order decreasing derivative of the thread bundle scheduler distribution rate and the first-order increasing derivative of the activation feature of the inter-node interconnect serial deserializer within the continuous acquisition period. When it is determined that the first-order descending derivative is not higher than the preset stagnation lower limit threshold and the first-order ascending derivative is not lower than the preset communication upper limit threshold, it is determined that the underlying tensor core is trapped in a global reduction data synchronization waiting state. The Boolean value of the microsecond-level communication damping stage identifier is set to true and sent to the power decision logic unit.
[0012] Furthermore, the step of inputting the multi-directional state matrix into the Long Short-Term Memory network, determining the forget gate and update gate parameters, and generating the expected resource usage ratio as the graphics processor load prediction tensor includes: The multi-directional state matrix is transformed into a three-dimensional floating-point tensor structure; The three-dimensional floating-point tensor structure is input into the first hidden layer of the long short-term memory network to determine the parameters of the forget gate and update gate, and based on this, the temporal dependency weight matrix for the microarchitecture running state sequence is determined. The time-dependent weight matrix is mapped and scaled through a fully connected layer, and the resulting numerical vector within a preset baseline mapping interval is used as the expected resource occupancy ratio.
[0013] Furthermore, the step of transferring the power credit points of the first computing node to the second computing node via high-speed Ethernet when the power token scheduler detects that the second computing node has triggered a high-load tensor task and the local power supply of the second computing node has reached the power wall threshold includes: Determine the available ampere current value corresponding to the power supply credit score of the first computing node; Send a limit expansion signal containing the available ampere current value to the baseboard management controller of the second computing node; Based on the aforementioned quota expansion signaling, the instantaneous power consumption wall limit of the second computing node is raised to maintain the core frequency from degrading.
[0014] Furthermore, the training process of the time series prediction network model includes: During the model training phase, real power consumption sampling sequences of the graphics processor under different workloads are collected. Determine the mean square deviation between the predicted value output by the Long Short-Term Memory network and the actual power consumption sampling sequence; The preset thermal resistance-capacitance time constant is introduced as a penalty term coefficient into the customized loss function. The gradient is determined based on the customized loss function, and the internal weight parameters of the long short-term memory network are updated by backpropagation.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention periodically collects microarchitectural operating state sequences such as the thread bundle scheduler dispatch rate, L2 cache miss rate, and inter-node interconnect serializer / deserializer activation characteristics within the graphics processing unit (GPU). Combined with a time-series prediction network model, it outputs a GPU load prediction tensor for future time periods, enabling early detection of transient phase transitions in the underlying hardware. Furthermore, by triggering the generation of microsecond-level communication damping phase identifiers through vector back-comparison, it accurately identifies minute phase transitions between intensive computation and global parameter reduction synchronization in the GPU. This solves the problem that the utilization polling mechanism cannot capture microsecond-level instruction stream stagnation due to sampling delay and insufficient granularity. It also avoids the instantaneous energy waste and heat accumulation caused by the voltage regulation module continuously outputting invalid voltage at full load during the communication waiting period.
[0016] 2. This invention generates target power supply pulse width modulation parameters based on the graphics processor load prediction tensor, microsecond-level communication damping stage identifier, and preset thermal resistance-capacitance time constant. These parameters are then directly sent to the graphics processor's underlying voltage regulation module via the out-of-band management bus, enabling active and precise control of the core output voltage and clock frequency. This mechanism differs from the lag of traditional static power walls or passive dynamic voltage and frequency adjustments, improving the real-time performance and matching degree of power supply regulation. While ensuring computing performance, it reduces redundant power consumption during non-computation cycles, alleviates thermal fatigue of the power supply circuit, and extends the service life of the hardware.
[0017] 3. This invention further constructs a power token scheduler that spans multiple computing nodes within the cluster. Through a cross-node dynamic transfer mechanism of power supply credit points, when a local node reaches the power consumption wall threshold, it borrows the remaining power supply capacity from low-load nodes, thereby achieving elastic sharing of cluster-level power supply resources. This design not only improves the overall energy utilization efficiency of the data center, but also avoids frequency degradation and training delays caused by single-node power consumption limitations, ensuring the continuity and stability of large-scale distributed model training tasks. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the steps of the present invention. Detailed Implementation
[0019] The following description is intended to disclose the invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious variations will occur to those skilled in the art.
[0020] Reference Figure 1 As shown, an AI server energy-saving control method based on GPU load prediction includes: Step 101: Periodically collect the microarchitecture operating state sequence of the graphics processor; specifically, the hardware performance counters inside the graphics processor are read by the baseboard management controller at a preset sampling period; the microarchitecture operating state sequence includes the thread bundle scheduler dispatch rate, L2 cache miss rate, and inter-node serializer deserialization activation characteristics; the thread bundle scheduler dispatch rate refers to the proportion of thread bundles that successfully issue instructions within a unit clock cycle, reflecting the busyness of the computing cores; the L2 cache miss rate refers to the proportion of requests that need to access main memory due to L2 cache misses, reflecting the latency pressure of memory access; the inter-node serializer deserialization... Activation characteristics refer to the duty cycle of the physical layer link of the high-speed bus or proprietary interconnect interface of peripheral components in an active transmission state, reflecting the intensity of cross-node communication; the microarchitecture running state sequence may also include hardware underlying operating indicators such as L1 cache read / write latency, tensor core utilization, video memory bandwidth utilization, and register file overflow count; it can be understood that any hardware counter data that can reflect the real-time physical working state of the computing units, storage units and communication units inside the graphics processor can be the above-mentioned microarchitecture running state sequence, of course, it is not limited here that all underlying data generated and collected inside the graphics processor must be included in this sequence; Step 102: Input the microarchitecture running state sequence within a continuous time window into a pre-defined temporal prediction network model to obtain deep spatiotemporal correlation features and output a graphics processor load prediction tensor for future time periods. Specifically, construct a time series input stream from the microarchitecture running state sequence of several consecutive sampling periods in the past; encode the input stream using a long short-term memory network to extract the hidden layer states of hardware state evolution over time; decode the hidden layer states through a fully connected layer to output the expected resource occupancy ratio for several future time steps; the expected resource occupancy ratio is the graphics processor load prediction tensor, used to characterize the computing density and power consumption trend of the graphics processor in the future; in the specific processing, the system will use the historical data collected over a period of time... Historical state data is packaged and input into the neural network. The network performs forward computation using internal weight parameters to extract feature vectors containing the switching patterns between computation and communication. These feature vectors are then processed through linear mapping and nonlinear activation in the output layer to generate continuous resource utilization ratios over a future period. Expected resource utilization ratios can be values such as 0.85, 0.62, or 0.30, representing the expected high, medium, or low load states of the graphics processing unit's computing cores in the near future, respectively. This prediction tensor can be understood as reflecting not only the load at a single future moment but also the continuous trend of load evolution over time. Of course, the prediction time step is not limited to a fixed value; its specific value can be flexibly configured based on the actual hardware response latency and task characteristics. Step 103: Based on the thread bundle scheduler distribution rate and the inter-node interconnect serial deserializer activation characteristics in the microarchitecture running state sequence, perform a vector reverse comparison to trigger the generation of a microsecond-level communication damping stage identifier; specifically, calculate the first derivative of the thread bundle scheduler distribution rate and the first derivative of the inter-node interconnect serial deserializer activation characteristics; when a sharp drop in the thread bundle scheduler distribution rate is detected and the inter-node interconnect serial deserializer activation characteristics remain high or increase, it is determined that the graphics processor is in a global reduction wait state where computation is stagnant but communication is busy; at this time, the generation of a microsecond-level communication damping stage identifier is triggered. This identifier is a Boolean signal used to indicate that the current state is non-computationally intensive. Communication damping period; In the specific determination process, the system uses a differential algorithm to calculate the rate of change between the index value of the current period and the index value of the previous period. If the rate of change of the thread beam scheduler distribution rate is negative and the absolute value is greater than the set decrease rate threshold, and at the same time the rate of change of the activation feature of the inter-node interconnect serial deserializer is positive and the absolute value is greater than the set increase rate threshold, then the reverse comparison is determined to be successful; The microsecond-level communication damping stage identifier can be, for example, a true value or a false value. When it is a true value, the subsequent step-down logic is triggered; when it is a false value, the current power supply state is maintained. It can be understood that this identifier is specifically used to capture the microsecond-level computational window caused by communication primitives in the training of large models; Step 104: Based on the graphics processor load prediction tensor, combined with the microsecond-level communication damping stage identifier and the preset thermal resistance-capacitance time constant, generate the target power supply pulse width modulation parameters. Specifically, obtain the equivalent thermal resistance and equivalent capacitance parameters of the server motherboard power supply circuit, and calculate the thermal resistance-capacitance time constant, which reflects the inertia of the power supply module's temperature response. If the microsecond-level communication damping stage identifier is true, map the graphics processor load prediction tensor to the lowest sustaining voltage level to cover the low load requirements during the communication waiting period. Use the thermal resistance-capacitance time constant as the cutoff frequency parameter of the low-pass filter to smooth voltage jumps, prevent voltage overshoot or undervoltage, and generate the final target power supply. Pulse width modulation parameters; in the specific calculation process, the thermal resistance-capacitance time constant is equal to the product of the equivalent thermal resistance and the equivalent thermal capacity; the system uses this time constant as the smoothing coefficient of the first-order inertial filter to transform the target voltage command, which may originally have a sharp jump, into a voltage trajectory that transitions smoothly over time, and then transforms the smoothed voltage trajectory into a duty cycle sequence of the pulse width modulation signal; the target power supply pulse width modulation parameter can be, for example, a pulse width modulation control signal with a duty cycle of 45%, 30%, or 20%, corresponding to different voltage drop amplitudes; it can be understood that this parameter is the optimal power supply command generated after comprehensively considering future load prediction, current communication status, and hardware thermal safety boundaries. Step 105: The target power supply pulse width modulation parameters are directly sent to the voltage regulation module at the bottom layer of the graphics processor via the out-of-band management bus, overwriting the current core output voltage and clock frequency. Specifically, the target power supply pulse width modulation parameters are encapsulated into a digital control frame conforming to the power management bus protocol. The control frame is sent to the digital multiphase voltage regulation module on the graphics processor board via the integrated circuit's built-in bus. The voltage regulation module adjusts the duty cycle of the pulse width modulation signal according to the received parameters and performs phase-switching operations to reduce the number of parallel power supply phases, thereby reducing the core output voltage and adjusting the clock frequency within microseconds. The target power supply pulse width modulation parameters can be, for example, a voltage control command frame conforming to the power management bus protocol, containing a specific target voltage value. It can be understood that this sending process completely bypasses the operating system and graphics processor firmware, and the board management controller directly performs microsecond-level physical control of the underlying power chip. Of course, the control bus is not limited to the integrated circuit's built-in bus; a serial peripheral interface or dedicated general-purpose input / output pins can also be used for control.
[0021] By using the above technical solutions, fine-grained indicators at the microarchitecture level are used to replace traditional macro-level utilization indicators, which can capture the microsecond-level communication waiting gaps that frequently occur in the training of large models. Combined with the time-series prediction network to predict load changes in advance, and using thermal resistance and capacitance characteristics for smooth regulation, the problem of ineffective energy consumption and heat accumulation caused by the lag of traditional passive voltage regulation is solved, realizing refined and intelligent energy saving of power supply for artificial intelligence servers.
[0022] In some embodiments, the periodic acquisition of a sequence of microarchitectural operating states of the graphics processor includes: Step 201: The performance counter register inside the graphics processor is read via the baseboard management controller to obtain the thread bundle scheduler dispatch rate of successfully issued instructions per unit clock cycle. Specifically, the baseboard management controller accesses the performance monitoring unit inside the graphics processor through the sideband interface; a specific register address is read, which stores the total number of times the scheduler attempted to dispatch instructions and the number of times it successfully dispatched instructions in the past sampling window; the ratio of the number of successful dispatches to the total number of dispatches is calculated to obtain the thread bundle scheduler dispatch rate; the closer this value is to 1, the more saturated the computing cores are; in the specific calculation process, assuming... If the scheduler attempts to distribute instructions a total of 10,000 times within the sampling window, and 8,500 instructions are successfully distributed, then the thread bundle scheduler distribution rate is equal to 8,500 divided by 10,000, which is 0.85. The thread bundle scheduler distribution rate can be, for example, 0.95, 0.70, or 0.40. This metric can be understood as directly reflecting the actual throughput efficiency of the graphics processor's computing core when executing instructions. It is not limited to the calculation method of this metric being the ratio of the number of successful instructions to the total number of instructions; it can also be equivalently converted through the clock cycle utilization rate at the underlying hardware level. Step 202: Read the on-chip memory management unit and read the L2 cache miss rate that causes main memory access wait. Specifically, read the L2 cache access statistics register in the memory management unit; obtain the L2 cache miss count and the total access count; calculate the ratio of the miss count to the total access count to obtain the L2 cache miss rate. A high miss rate usually indicates that the computing core is about to enter an idle cycle due to waiting for data. In the specific calculation process, assuming that the graphics processor accesses the L2 cache 50,000 times in one sampling period, of which 15,000 times require accessing main memory due to misses, the L2 cache miss rate is equal to 15,000 divided by 50,000, and the result is 0.30. The L2 cache miss rate can be, for example, 0.05, 0.45, or 0.80. It can be understood that this indicator reflects the degree to which the bandwidth of the graphics processor's video memory or main memory restricts the computing task. When this value suddenly increases, it means that the computing task is about to be forced to slow down due to data transfer delays. Step 203: Read the physical layer link training status of the peripheral component interconnect high-speed bus interface, and read the activation feature of the inter-node interconnect serial deserializer currently in an active state. Specifically, read the status register of the physical layer of the peripheral component interconnect high-speed bus or proprietary interconnect interface; detect whether the link is in an active transmission state and the throughput count of data packets; calculate the percentage of active link cycles per unit time, and generate the inter-node interconnect serial deserializer activation feature. This feature is used to characterize the data synchronization strength across servers or cards. In the specific calculation process, assuming that the physical layer link is in an active transmission state for a cumulative period of 7 microseconds within a 10-microsecond sampling period, the inter-node interconnect serial deserializer activation feature is equal to 7 divided by 10, and the calculation result is 0.70. The inter-node interconnect serial deserializer activation feature can be, for example, 0.10, 0.60, or 0.95. It can be understood that this indicator reflects the busyness of multi-card interconnection or inter-node communication. When this indicator is close to 1, it indicates that the communication link is in a full-load state. Of course, the statistical dimension of this feature is not limited here, and the effective data packet transmission volume per unit time can also be used for equivalent characterization.
[0023] By using the above technical solution, raw microarchitecture data can be obtained directly from the underlying hardware registers, avoiding the latency and accuracy loss caused by operating system-level driver polling, and ensuring the real-time performance and accuracy of load monitoring.
[0024] In some embodiments, the microarchitecture runtime state sequence within a continuous time window is input into a pre-defined temporal prediction network model to obtain deep spatiotemporal correlation features, and outputs a graphics processor load prediction tensor for future time periods, including: Step 301: Assign corresponding timestamp information to the thread bundle scheduler dispatch rate, L2 cache miss rate, and inter-node interconnect serializer / deserializer activation features. Specifically, at the data acquisition end, assign a high-precision timestamp to each group of microarchitecture state data. Use the timestamps to align multi-source heterogeneous hardware metrics and eliminate phase differences caused by different acquisition paths. In the specific processing, due to the different time consumption of reading different registers, the three metrics at the same moment may be misaligned on the time axis. The system will use the reference clock as the anchor point to map the acquired data onto a unified time grid. The timestamp information can be, for example, an absolute timestamp or a sequence number of a relative sampling period. It can be understood that time alignment is a prerequisite for ensuring that the time series prediction network can correctly learn the causal relationship between metrics. Step 302: Based on the timestamp information, the thread bundle scheduler distribution rate, the L2 cache miss rate, and the inter-node interconnect serializer / deserializer activation features are concatenated into a multi-directional state matrix containing the time direction. Specifically, the three time-aligned feature vectors are arranged along the time axis to construct a two-dimensional matrix containing the time step and feature dimension. Normalization is performed on this matrix to map its values to the range of 0 to 1, eliminating the influence of different physical dimensions. In the specific processing, the system concatenates multiple time-aligned feature vectors in chronological order to form a two-dimensional data matrix. Then, the range scaling formula is used to subtract the minimum value of the feature in the historical sliding window from the original feature value in the matrix, and then divide by the difference between the maximum and minimum values, thereby compressing all values to between 0 and 1. The multi-directional state matrix can be, for example, a floating-point matrix where all elements are distributed between 0 and 1 after normalization. It can be understood that the normalization operation can accelerate the convergence speed of the neural network model and prevent a certain indicator with a large dimension from dominating the gradient update. Step 303: Input the multi-directional state matrix into the Long Short-Term Memory (LSTM) network to determine the parameters of the forget gate and update gate, and generate the expected resource usage ratio as the GPU load prediction tensor. Specifically, the multi-directional state matrix is input into the LTM network; the LTM network's forget gate determines how much past state information to discard based on the current input, and the input gate determines how much new information to store; through the transfer of hidden layer states, the non-linear temporal dependencies between microarchitectural indicators are captured; finally, the output layer outputs a scalar between 0 and 1 through an activation function, as the expected resource usage ratio for future time steps; in the specific calculation process, the forget gate and input gate inside the LTM network will... The network performs a weighted summation of the current input data and the hidden layer state from the previous time step, and generates a control signal through an activation function to determine the update and forgetting of the cell state. Finally, the network transforms the processed hidden state into the expected resource occupancy ratio through the output layer. The expected resource occupancy ratio can be, for example, 0.75, 0.50, or 0.20. It can be understood that the Long Short-Term Memory network effectively solves the gradient vanishing problem in the long sequence training of traditional recurrent neural networks through a gating mechanism, and can accurately remember the historical load fluctuation pattern. Of course, it is not limited here that the time series prediction model must be a Long Short-Term Memory network; gated recurrent units or converter architectures can also be used. By utilizing the powerful temporal memory capabilities of Long Short-Term Memory (LSTM) networks, the switching patterns between computational and communication tasks can be learned from the minute fluctuations in microarchitectural metrics, thereby enabling accurate prediction of future workloads.
[0025] In some embodiments, the target power supply pulse width modulation parameters are generated based on the graphics processor load prediction tensor, combined with the microsecond-level communication damping stage identifier and the preset thermal resistance-capacitance time constant, including: Step 401: Obtain the equivalent thermal resistance and equivalent capacitance of the power supply circuit of the target data center server motherboard to determine the thermal resistance-capacitance time constant. Specifically, obtain the junction-to-environment thermal resistance and thermal capacitance parameters of the power devices in the power supply circuit through offline calibration experiments or by reading the motherboard design parameters. Calculate the product of thermal resistance and thermal capacitance to obtain the thermal resistance-capacitance time constant. This constant characterizes the inertia of temperature change in the power supply circuit, i.e., voltage regulation cannot be faster than the rate of heat accumulation, otherwise it will lead to device damage. In the specific calculation process, assuming that the equivalent thermal resistance of the power supply module is 0.5 degrees Celsius per watt and the equivalent thermal capacitance is 20 joules per degree Celsius obtained through thermal simulation or actual measurement, the thermal resistance-capacitance time constant is equal to 0.5 multiplied by 20, and the calculation result is 10 seconds. The thermal resistance-capacitance time constant can be, for example, 5 seconds, 10 seconds, or 15 seconds. It can be understood that this constant is the physical bottom line for ensuring the thermal safety of the power module, and the execution cycle of any voltage regulation strategy should not be significantly lower than this constant. Step 402: When the microsecond-level communication damping stage flag is determined to be in an active state, the graphics processor load prediction tensor is mapped to the minimum sustaining voltage level. Specifically, if the communication damping stage flag is true, it means that the graphics processor is in a communication waiting period, and the computing core does not require high frequency and high voltage. The voltage frequency lookup table is queried to find the minimum voltage frequency combination required to maintain the current communication bandwidth. The graphics processor load prediction tensor is forcibly clamped to the value corresponding to the minimum level. In the specific processing, assuming that the current prediction tensor indicates a higher future load, but the communication damping stage flag is true, the system queries the preset lookup table and finds that the minimum voltage required to maintain the current communication bandwidth is a specific value. Then, the target voltage corresponding to the prediction tensor is forcibly limited to the minimum value. The minimum sustaining voltage level can be, for example, 0.70 volts, 0.75 volts, or 0.80 volts. It can be understood that this clamping mechanism ensures that during the communication waiting period, the graphics processor can significantly reduce computing power consumption while maintaining the stability of the communication link without interruption. Step 403: Based on the thermal resistor-capacitor time constant, perform smoothing filtering on the lowest sustaining voltage level to generate the target power supply pulse width modulation parameters. Specifically, construct a first-order inertial link filter with its time constant set to the thermal resistor-capacitor time constant. Input the voltage jump command into the filter, and output a voltage trajectory that decays or rises exponentially with time. Convert this trajectory into a duty cycle sequence of a pulse width modulation signal to generate the target power supply pulse width modulation parameters. In the specific processing, the system uses the thermal resistor-capacitor time constant as the smoothing coefficient of the filter to transform the step voltage change that might otherwise cause inductor howling or capacitor overshoot into a smooth transition curve that conforms to the thermal characteristics of the hardware. Then, the smoothed voltage trajectory is converted into a duty cycle sequence of a pulse width modulation signal. The target power supply pulse width modulation parameters can be, for example, a pulse width modulation sequence with a duty cycle that smoothly transitions from 60% to 45%. It can be understood that the smoothing filtering process transforms the violent voltage jump that might otherwise damage the hardware into a smooth adjustment process that conforms to physical laws.
[0026] By introducing the thermal resistance-capacitance time constant as a physical constraint through the above technical solution, it is ensured that the energy-saving control strategy will not damage the hardware due to excessively rapid voltage adjustment, and the synergistic optimization of electrical and thermal characteristics is achieved.
[0027] In some embodiments, the target power supply pulse width modulation parameters are directly sent to the voltage regulation module at the graphics processor's underlying layer via the out-of-band management bus, overwriting the current core output voltage and clock frequency, including: Step 501: Encapsulate the target power supply pulse width modulation parameters into a hexadecimal control frame conforming to the power management bus protocol format. Specifically, according to the power management bus protocol standard, construct a data frame containing the operation command code, data length, and check bits; encode the target voltage value and phase control command into hexadecimal data segments. In the specific processing, assuming the target voltage is 0.75V, the system converts it into the corresponding hexadecimal encoding and encapsulates it into a complete data frame containing the command code, data length, and the hexadecimal data, and finally adds a cyclic redundancy check bit. The hexadecimal control frame can be, for example, a byte sequence such as 0x210x020x2E0xE00xA5. It can be understood that standardized protocol encapsulation ensures seamless communication between the baseboard management controller and power chips from different manufacturers. Step 502: Send the hexadecimal control frame to the digital control pin of the voltage regulation module via the integrated circuit's built-in bus. Specifically, the baseboard management controller, acting as the master device, addresses the slave device address of the graphics processor power supply controller via the integrated circuit's built-in bus. The control frame is written into the power supply controller's command register. In the specific communication process, the baseboard management controller first sends the 7-bit slave device address, then sends the write operation bit. After receiving the slave device's response signal, it sequentially sends the command code and data bytes, and finally sends the stop condition. The integrated circuit's built-in bus can be, for example, an integrated circuit built-in bus, a system management bus, or a power management bus. It can be understood that the out-of-band management bus is independent of the graphics processor's data processing bus, ensuring reliable power supply control commands even if the graphics processor's operating system crashes or freezes. Step 503: Forcefully trigger the multi-phase power supply circuit inside the voltage regulation module to perform a phase-switching operation, reducing the number of power supply phases in parallel output, and simultaneously lowering the core clock frequency of the graphics processor; specifically, after the power supply controller parses the instruction, it shuts down the drive signals of some phases, causing the multi-phase power supply circuit to switch from, for example, a 10-phase working mode to a 4-phase working mode; at the same time, it adjusts the pulse width modulation duty cycle of the remaining phases to match the target voltage; by reducing the number of working phases, it reduces switching losses, and achieves dual energy saving in conjunction with voltage reduction; in the specific control process, when a voltage reduction instruction is received, the phase manager inside the power chip will sequentially shut down the upper and lower bridge arm metal-oxide-semiconductor field-effect transistor drive signals of phases 5 to 10, retaining only the first 4 phases to continue working, and increasing the pulse width modulation duty cycle of these 4 phases from the original 30% to 45% to maintain an output voltage of 0.75V; the multi-phase power supply circuit can, for example, switch from 16 phases to 8 phases, or from 10 phases to 4 phases; it can be understood that the phase-switching operation avoids the serious switching losses and core losses caused by multi-phase parallel connection under light load conditions.
[0028] By using the above technical solution, the underlying power chip can be controlled directly through the external bus, bypassing the operating system, thus achieving a microsecond-level response speed and ensuring that energy-saving actions can be completed even during brief communication gaps.
[0029] In some embodiments, after generating the target power supply pulse width modulation parameters, the method further includes: Step 601: Establish a power token scheduler spanning multiple computing nodes within the cluster. Specifically, deploy a power scheduling service on the cluster management node to maintain a global power token pool. Each computing node acts as a client and maintains a heartbeat connection with this service. During the deployment process, the cluster management node starts a persistent daemon process. This process allocates a global variable in memory as the power token pool and exposes the interface to all computing nodes in the cluster through remote procedure calls or a representational state transition interface. Each computing node sends a heartbeat packet to the service every 100 milliseconds to maintain the connection. The power token scheduler can be, for example, a distributed cluster based on a memory key-value store or a microservice deployed in a container orchestration system. It can be understood that the power token scheduler is the brain of the entire cluster-level energy flow, responsible for the overall coordination and allocation of global power. Step 602: Obtain the current remaining power supply capacity of the first computing node through the power token scheduler, and encapsulate the current remaining power supply capacity into power supply credit points; specifically, monitor the difference between the real-time power consumption of the first computing node and the set power consumption wall; convert the difference into the number of current amperes that can be lent, and map it into power supply credit points and upload it to the power token scheduler; in the specific calculation process, assuming that the power consumption wall limit of the first computing node is 1000 watts, and the current real-time power consumption is 700 watts, then the remaining power supply capacity is 300 watts. If the system sets 10 watts to correspond to 1 credit point, then the power supply credit points that the node can lend are 30; the power supply credit points can be, for example, 10, 30, or 50, etc.; it can be understood that abstracting physical power into integrals facilitates the scheduler to perform fast addition and subtraction operations and allocation. The mapping ratio of the integrals is not limited here, and 1 watt can also correspond to 1 integral point or be nonlinearly mapped according to the voltage level; Step 603: When a high-load tensor task is triggered on the second computing node and the local power supply of the second computing node reaches the power wall threshold, the power credit points of the first computing node are transferred to the second computing node via high-speed Ethernet. Specifically, when the second computing node detects task queuing and is limited by the power wall, it requests credits from the scheduler. The scheduler transfers the credits of the first node to the second node. The second node raises its local instantaneous power wall limit based on the obtained credits. In the specific interaction process, the second computing node sends an application message containing a request for 20 credits to the scheduler via high-speed Ethernet. After the scheduler verifies that the first node's account balance is sufficient, it performs an atomic operation to deduct 20 credits from the first node and add 20 credits to the second node. After receiving the confirmation message, the second node raises its local power wall limit from 1000 watts to 1200 watts. The transfer of power credit points can be achieved, for example, through database transfer transactions or message queue instruction issuance. It can be understood that this cross-node power borrowing mechanism breaks the physical limitation of the single-machine power wall, enabling the cluster to cope with sudden peak demands for large model training.
[0030] In some embodiments, based on the thread bundle scheduler distribution rate and the inter-node interconnect serializer / deserializer activation characteristics in the microarchitecture runtime state sequence, a vector back-comparison is performed to trigger the generation of a microsecond-level communication damping stage identifier, including: Step 701: Determine the first-order decreasing derivative of the thread bundle scheduler distribution rate and the first-order increasing derivative of the activation feature of the inter-node interconnect serial deserializer within the continuous acquisition period. Specifically, the difference between the thread bundle scheduler distribution rate at the current time and the previous time is calculated using a differential algorithm to obtain the descent rate; similarly, the rise rate of the inter-node interconnect serial deserializer activation feature is calculated. In the specific calculation process, assuming the thread bundle scheduler distribution rate at the previous time was 0.90, the current time is 0.70, and the sampling period is 10 microseconds, then the first-order descent derivative is equal to (0.70-0.90) / 10 = -0.02 per microsecond; similarly, if the inter-node interconnect serial deserializer activation feature at the previous time was 0... If the current time is 0.30, then the first ascending derivative is equal to (0.50-0.30) / 10 = 0.02 per microsecond; the first descending derivative can be negative values such as -0.01 or -0.05, and the first ascending derivative can be positive values such as 0.01 or 0.04; it can be understood that derivative calculation can keenly capture the dramatic changing trend of indicators on the microsecond time scale. Here, it is not limited that the difference algorithm must be the first-order forward difference; central difference or backward difference can also be used to improve the calculation accuracy. Step 702: When the first-order descending derivative is not higher than a preset stagnation lower threshold and the first-order ascending derivative is not lower than a preset communication upper threshold, the underlying tensor core is determined to be in a global reduction data synchronization waiting state. Specifically, the stagnation lower threshold is set to -0.5, and the communication upper threshold is set to 0.2. When the characteristics of a sudden drop in computation and a sudden increase in communication are met, it is logically determined that the system has entered the waiting period for synchronization primitives such as global reduction. In the specific determination process, the system compares the calculated first-order descending derivative with the stagnation lower threshold. If -0.02 is less than or equal to -0.5, the threshold is determined. This threshold is only an example; in actual applications, it can be determined based on normalization. The derivative range is set, for example, to -0.01, while the first-order rising derivative 0.02 is greater than or equal to 0.2; similarly, it can actually be set to 0.01, then both conditions are met simultaneously, triggering the judgment logic; the lower threshold for stagnation can be, for example, -0.01, -0.05, etc., and the upper threshold for communication can be, for example, 0.01, 0.05, etc.; it can be understood that the dual threshold joint judgment mechanism effectively eliminates false triggers caused by simply calculating idle time or simply communication fluctuations, and accurately locks the communication damping stage. Of course, the threshold is not limited to a fixed constant here, and can also be adaptively adjusted according to the standard deviation of historical operating data; Step 703: Set the Boolean value of the microsecond-level communication damping stage identifier to true and send it to the power decision logic unit; specifically, generate a high-level signal as an identifier to trigger the subsequent buck logic; in the specific execution process, when the judgment condition of step 702 is met, a flag register inside the power decision logic unit is set to 1, and this flag bit serves as the enable signal for the voltage clamping operation in the subsequent step 402; the microsecond-level communication damping stage identifier can be, for example, a Boolean value of true, a logic high level of 1, or a string of enable; it can be understood that this Boolean identifier is a key bridge connecting the microarchitecture status monitoring and the underlying power supply control. Of course, the data type of the identifier is not limited here, and an enumeration type or an integer status code can also be used to represent different communication damping stages; By utilizing the inverse change characteristics of vector derivatives, the communication damping stage can be accurately identified, avoiding the misjudgment of simple computational idleness as communication waiting, thus improving the accuracy of control.
[0031] In some embodiments, a multi-directional state matrix is input into a long short-term memory network to determine forget gate and update gate parameters, and to generate the expected resource usage ratio as a graphics processor load prediction tensor, including: Step 801: Convert the multi-directional state matrix into a three-dimensional floating-point tensor structure. Specifically, add a batch dimension to convert the matrix into a three-dimensional tensor of [batch size, time step, number of features]. Convert the data type to 32-bit floating-point numbers to adapt to neural network computation. In the specific conversion process, assuming the shape of the original multi-directional state matrix is [100, 3], in order to adapt to the input requirements of the long short-term memory network of the deep learning framework, the system adds a batch dimension to the outermost layer, converts it into a three-dimensional tensor of shape [batch size, 100, 3], and converts the data type to 32-bit floating-point numbers. The three-dimensional floating-point tensor structure can be, for example, a floating-point array of shape [32, 100, 3]. It can be understood that adding the batch dimension is to support the batch parallel computation of the neural network framework, thereby greatly improving the inference speed of the model. Of course, the data type is not limited to 32-bit floating-point numbers; 16-bit or 64-bit floating-point numbers can also be used depending on the hardware computing power. Step 802: Input the three-dimensional floating-point tensor structure into the first hidden layer of the Long Short-Term Memory (LSTM) network to determine the temporal dependency weight matrix for the microarchitecture running state sequence. Specifically, the product of the input tensor and the weight matrix is calculated by matrix multiplication, a bias term is added, and the result is processed by an activation function to update the cell state and hidden state. In the specific calculation process, after the three-dimensional floating-point tensor enters the first hidden layer of the LTM network, the weight matrix inside the network will perform matrix multiplication with the input data. After adding a bias term to the result, it is processed by a non-linear activation function to determine which historical information needs to be forgotten and which new information needs to be updated, ultimately generating new cell states and hidden states. The temporal dependency weight matrix can be, for example, a floating-point matrix with the shape of [number of features, number of hidden layer units]. It can be understood that the weight matrix is the core knowledge learned by the LTM network during training. It determines how the network extracts features useful for future loads from historical microarchitecture states. Of course, the dimension of the weight matrix is not limited here; its size depends on the number of hidden layer units. Step 803: The temporal dependency weight matrix is mapped and scaled through a fully connected layer, and the output is a numerical vector within a preset baseline mapping interval as the expected resource occupancy ratio. Specifically, the hidden state output by the Long Short-Term Memory network is input into the fully connected layer. The output value is compressed to between 0 and 1 using an activation function, which directly corresponds to the predicted percentage of GPU core utilization. In the specific calculation process, the hidden state vector output by the last step of the Long Short-Term Memory network is fed into a fully connected layer. The fully connected layer maps high-dimensional features to low-dimensional output through linear transformation, and then passes through an activation function, such as the Sigmoid function, to strictly limit the output value to the interval between 0 and 1. This output value represents the expected resource occupancy ratio at future time. The expected resource occupancy ratio can be, for example, 0.85, 0.60, or 0.25. It can be understood that the fully connected layer plays the role of feature dimensionality reduction and result mapping, transforming the complex temporal features inside the network into an intuitive load prediction percentage. In some embodiments, when a high-load tensor task is triggered on the second computing node and the local power supply of the second computing node reaches the power wall threshold, transferring the power credit of the first computing node to the second computing node via high-speed Ethernet includes: Step 901: Determine the available ampere current value corresponding to the power supply credit points of the first computing node; specifically, calculate the transferable current limit based on the mapping relationship between power supply credit points and current; in the specific calculation process, assuming that the first computing node has 50 power supply credit points, and the system-set mapping relationship is that 1 credit point corresponds to 10 ampere current, then the transferable available ampere current value of this node is 50 multiplied by 10, that is, 500 amperes; the available ampere current value can be, for example, 100 amperes, 300 amperes, or 500 amperes, etc.; it can be understood that restoring the abstract power supply credit points to the physical current value is to enable the underlying power management hardware to accurately understand and execute power adjustment instructions; Step 902: Send a credit extension signaling message containing the available ampere current value to the baseboard management controller of the second computing node; specifically, the current value is encapsulated in a signaling packet and sent to the second node through a remote procedure call interface; in the specific communication process, the power token scheduler constructs a credit extension signaling packet containing the target node identifier, the available ampere current value, and a timestamp, and sends the signaling packet to the baseboard management controller of the second computing node via high-speed Ethernet; the credit extension signaling message can be, for example, a remote procedure call request message containing the current value of 500 amperes; it can be understood that the credit extension signaling message is the communication carrier for power allocation within the cluster, ensuring the real-time performance and accuracy of cross-node power transfer; Step 903: Based on the quota expansion signaling, raise the instantaneous power limit of the second computing node to maintain the core frequency without frequency degradation. Specifically, after receiving the signaling, the baseboard management controller of the second node modifies the power limit register of the graphics processor, raising the power limit by the corresponding power value, thereby allowing the graphics processor to maintain high-frequency operation without being forced to reduce the frequency during sudden high loads. In the specific execution process, the baseboard management controller of the second node parses the quota expansion signaling, obtains the available current value of 500 amps, calculates the corresponding power increment in combination with the current operating voltage, and then modifies the power limit register inside the graphics processor through the integrated circuit built-in bus, raising the power limit from 1000 watts to 1500 watts. The instantaneous power limit can be raised from 1000 watts to 1200 watts, 1500 watts, or 1800 watts. It can be understood that raising the power limit ensures that the graphics processor has sufficient power supply during the critical tensor calculation stage, avoiding the decrease in calculation frequency due to power limitation, thereby ensuring the overall efficiency of large model training. The above technical solution enables cluster-level energy flow, utilizes the redundant power of low-load nodes to support burst computing of high-load nodes, and improves the overall training throughput of the cluster.
[0032] In some embodiments, the training process of the time-series prediction network model includes: Step 1001: During the model training phase, collect real power consumption sampling sequences of the graphics processing unit (GPU) under different workloads. Specifically, run typical deep learning training tasks and use a high-precision power meter to collect real power consumption data of the GPU as labels. In the specific collection process, run typical artificial intelligence tasks such as large language model training and image recognition training on the server, and use a high-precision power meter with the same sampling period as the microarchitecture state sequence to record the real power consumption data of the GPU in real time, forming a power consumption label sequence that corresponds one-to-one with the microarchitecture state sequence. The real power consumption sampling sequence can be, for example, a power consumption value array containing 10,000 sampling points, with the unit being watts. It can be understood that the real power consumption data is the ground truth for training the time-series prediction network model, and the learning goal of the model is to predict these real power consumption values as accurately as possible. Step 1002: Determine the mean squared deviation between the predicted value output by the Long Short-Term Memory (LSTM) network and the actual power consumption sampling sequence. Specifically, calculate the mean squared error between the predicted power consumption curve and the actual power consumption curve as the base loss value. In the specific calculation process, the predicted power consumption sequence output by the LSTM network is compared point by point with the actual power consumption sequence collected by the high-precision power meter. The squared difference between the predicted value and the actual value at each time point is calculated, and then the average of the squared differences at all time points is calculated to obtain the mean squared deviation. The mean squared deviation can be, for example, 0.05, 0.12, or 0.25. It can be understood that the mean squared deviation reflects the gap between the model prediction result and the actual physical power consumption. The smaller the value, the higher the prediction accuracy of the model. Here, it is not limited that the loss function must be the mean squared error. Mean absolute error or Huber loss, which are more robust to outliers, can also be used. Step 1003: Introduce the thermal resistor-capacitor time constant as a penalty term coefficient into the customized loss function, determine the gradient based on the customized loss function, and backpropagate to update the internal weight parameters of the Long Short-Term Memory network; specifically, construct the loss function, which is the sum of the mean square error and a penalty term; this penalty term is the product of the thermal resistor-capacitor time constant, the voltage change rate, and a weight coefficient; by introducing the thermal resistor-capacitor time constant as a penalty term, the network is forced to learn a smooth prediction result that conforms to the hardware thermal characteristics, avoiding the output of drastically changing prediction values that could lead to hardware instability; update the network weights using gradient descent; in the specific training process, the system first calculates the mean square deviation between the predicted value and the true value, and then... The difference in predicted power consumption between adjacent time steps is calculated as the voltage change rate. The thermal resistance-capacitance time constant is multiplied by the voltage change rate and weight coefficients to obtain a penalty term. The mean square deviation is added to the penalty term to obtain the final customized loss value. Subsequently, the gradient of this loss value with respect to the internal weight parameters of the network is calculated using the backpropagation algorithm, and the weights are updated using the gradient descent method. The customized loss function can be, for example, the mean square deviation plus a penalty term of 0.1. This can be understood as incorporating physical thermal characteristics into the training objective of the neural network, making the prediction model not only accurate but also in line with the physical safety constraints of the hardware, avoiding the model outputting drastic jump prediction values that do not conform to the thermal response characteristics of the power module in pursuit of the ultimate prediction accuracy.
[0033] By incorporating physical thermal characteristics into the training objectives of the neural network through the above technical solution, the prediction model is not only accurate, but also complies with the physical safety constraints of the hardware.
[0034] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A method for energy-saving control of AI servers based on GPU load prediction, characterized in that, include: The microarchitecture operating state sequence of the graphics processor is periodically collected; wherein, the microarchitecture operating state sequence includes thread bundle scheduler dispatch rate, L2 cache miss rate, and inter-node interconnect serializer deserializer activation characteristics; The microarchitecture running state sequence within a continuous time window is input into a pre-defined time-series prediction network model to obtain deep spatiotemporal correlation features and output a graphics processor load prediction tensor for future time periods. Based on the thread bundle scheduler distribution rate and the node interconnect serial deserializer activation characteristics in the microarchitecture running state sequence, a vector reverse comparison is performed to trigger the generation of a microsecond-level communication damping stage identifier. Based on the graphics processor load prediction tensor, combined with the microsecond-level communication damping stage identifier and the preset thermal resistance-capacitance time constant, the target power supply pulse width modulation parameters are generated. The target power supply pulse width modulation parameters are sent directly to the voltage regulation module at the bottom layer of the graphics processor via the out-of-band management bus, overwriting the current core output voltage and clock frequency.
2. The method for energy-saving control of AI servers based on GPU load prediction according to claim 1, characterized in that, The periodic acquisition sequence of the microarchitecture operating state of the graphics processor includes: The thread beam scheduler dispatch rate of successfully issued instructions per unit clock cycle is read from the graphics processor's internal performance counter register via the baseboard management controller. Read the on-chip memory management unit and read the L2 cache miss rate that caused the main memory access wait; Read the physical layer link training status of the high-speed bus interface for peripheral component interconnection, and based on the physical layer link training status, read the activation feature of the inter-node interconnection serial deserializer that is currently active.
3. The method for energy-saving control of AI servers based on GPU load prediction according to claim 1, characterized in that, The step of inputting the microarchitecture operating state sequence within a continuous time window into a pre-set temporal prediction network model to obtain deep spatiotemporal correlation features and outputting a graphics processor load prediction tensor for future time periods includes: Assign corresponding timestamp information to the thread bundle scheduler distribution rate, the secondary cache miss rate, and the inter-node interconnect serial deserializer activation feature; Based on the timestamp information, the thread bundle scheduler distribution rate, the secondary cache miss rate, and the inter-node interconnect serial deserializer activation features are concatenated into a multi-directional state matrix that includes the time direction. The multi-directional state matrix is input into the long short-term memory network to determine the parameters of the forget gate and update gate, and the expected resource usage ratio is generated as the graphics processor load prediction tensor.
4. The method for energy-saving control of AI servers based on GPU load prediction according to claim 3, characterized in that, The step of generating target power supply pulse width modulation parameters based on the graphics processor load prediction tensor, combined with the microsecond-level communication damping stage identifier and the preset thermal resistance-capacitance time constant, includes: Obtain the equivalent thermal resistance and equivalent capacitance of the power supply circuit of the target data center server motherboard, and obtain the preset thermal resistance-capacitance time constant, wherein the thermal resistance-capacitance time constant reflects the inherent decay rate of the power supply component due to temperature change. When it is determined that the microsecond-level communication damping stage identifier is in a valid active state, the graphics processor load prediction tensor is mapped to the lowest sustaining voltage level. Based on the thermal resistance-capacitance time constant, a smoothing filter is applied to the lowest sustaining voltage level to generate the target power supply pulse width modulation parameters.
5. The method for energy-saving control of AI servers based on GPU load prediction according to claim 1, characterized in that, The step of directly sending the target power supply pulse width modulation parameters to the voltage regulation module at the graphics processor's underlying layer via the out-of-band management bus, overwriting the current core output voltage and clock frequency, includes: The target power supply pulse width modulation parameters are encapsulated into a hexadecimal control frame conforming to the power management bus protocol format; The hexadecimal control frame is sent to the digital control pin of the voltage regulation module via the integrated circuit's built-in bus. The multi-phase power supply circuit inside the voltage regulation module is forcibly triggered to perform a phase-switching operation, reducing the number of power supply phases in parallel output, and simultaneously lowering the core clock frequency of the graphics processor to overwrite the current core output voltage and clock frequency.
6. The method for energy-saving control of AI servers based on GPU load prediction according to claim 1, characterized in that, After generating the target power supply pulse width modulation parameters, the method further includes: Establish a power token scheduler that spans multiple computing nodes within the cluster; The power token scheduler obtains the current remaining power supply capacity of the first computing node and encapsulates the current remaining power supply capacity into power supply credit points. When the power token scheduler detects that the second computing node has triggered a high-load tensor task and the local power supply of the second computing node has reached the power wall threshold, the power credit of the first computing node is transferred to the second computing node via high-speed Ethernet.
7. The method for energy-saving control of AI servers based on GPU load prediction according to claim 1, characterized in that, The step of performing a vector back comparison based on the thread bundle scheduler distribution rate and the inter-node interconnect serializer / deserializer activation characteristics in the microarchitecture running state sequence, triggering the generation of a microsecond-level communication damping stage identifier, includes: Determine the first-order decreasing derivative of the thread bundle scheduler distribution rate and the first-order increasing derivative of the activation feature of the inter-node interconnect serial deserializer within the continuous acquisition period. When it is determined that the first-order descending derivative is not higher than the preset stagnation lower limit threshold and the first-order ascending derivative is not lower than the preset communication upper limit threshold, it is determined that the underlying tensor core is trapped in a global reduction data synchronization waiting state. The Boolean value of the microsecond-level communication damping stage identifier is set to true and sent to the power decision logic unit.
8. The method for energy-saving control of AI servers based on GPU load prediction according to claim 3, characterized in that, The step of inputting the multi-directional state matrix into a long short-term memory network, determining the forget gate and update gate parameters, and generating the expected resource usage ratio as the graphics processor load prediction tensor includes: The multi-directional state matrix is transformed into a three-dimensional floating-point tensor structure; The three-dimensional floating-point tensor structure is input into the first hidden layer of the long short-term memory network to determine the parameters of the forget gate and update gate, and based on this, the temporal dependency weight matrix for the microarchitecture running state sequence is determined. The time-dependent weight matrix is mapped and scaled through a fully connected layer, and the resulting numerical vector within a preset baseline mapping interval is used as the expected resource occupancy ratio.
9. The method for energy-saving control of AI servers based on GPU load prediction according to claim 6, characterized in that, When the power token scheduler detects that the second computing node has triggered a high-load tensor task and the local power supply of the second computing node has reached the power wall threshold, the transfer of the power credit points of the first computing node to the second computing node via high-speed Ethernet includes: Determine the available ampere current value corresponding to the power supply credit score of the first computing node; Send a limit expansion signal containing the available ampere current value to the baseboard management controller of the second computing node; Based on the aforementioned quota expansion signaling, the instantaneous power consumption wall limit of the second computing node is raised to maintain the core frequency from degrading.
10. The method for energy-saving control of AI servers based on GPU load prediction according to claim 3, characterized in that, The training process of a time-series prediction network model includes: During the model training phase, real power consumption sampling sequences of the graphics processor under different workloads are collected. Determine the mean square deviation between the predicted value output by the Long Short-Term Memory network and the actual power consumption sampling sequence; The preset thermal resistance-capacitance time constant is introduced as a penalty term coefficient into the customized loss function. The gradient is determined based on the customized loss function, and the internal weight parameters of the long short-term memory network are updated by backpropagation.