Deep learning-based computing power server power consumption optimization method and system

By collecting power consumption and load signals from computing servers and using deep learning for power consumption prediction and matching, the shortcomings of traditional power management methods are solved, and power consumption balance and performance stability of computing server clusters are achieved.

CN122507261APending Publication Date: 2026-08-04YIBIN XINJIYIYUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YIBIN XINJIYIYUN TECH CO LTD
Filing Date
2026-05-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Traditional power management methods for computing servers cannot accurately adjust power consumption according to actual load conditions, leading to resource waste or performance degradation. Furthermore, they lack effective analysis and utilization of the power coupling relationship between computing units, making it difficult to achieve global power balance optimization.

Method used

The system collects power consumption sensing signals and task load signals from the computing server cluster, performs power consumption prediction through a deep learning power drift trend inference network, generates a predicted power consumption drift surface for the computing unit in the next scheduling cycle, filters out candidate computing units that meet the power coupling conditions, performs power matching processing, and generates power balancing scheduling instructions.

Benefits of technology

It enables dynamic and intelligent optimization of the power consumption of computing server clusters, improving the power efficiency of server clusters, reducing operating costs, and ensuring stable performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122507261A_ABST
    Figure CN122507261A_ABST
Patent Text Reader

Abstract

The application provides a kind of power consumption optimization method and system based on deep learning of computing power server, it is related to computing power server power consumption management technical field, first, the power consumption induction signal sequence and task load signal sequence of computing power server cluster operation unit are collected, and are packaged as synchronous monitoring signal array, then power consumption ripple form analysis is carried out, and ripple load correlation feature vector is constructed, then the predicted power consumption drift surface is generated by inputting power consumption drift trend deduction network, according to the predicted power consumption drift surface, a candidate operation unit set is screened and power consumption is matched, to obtain an operation unit power consumption matching allocation scheme, finally, a power consumption balancing scheduling instruction set containing operation unit addressing identification and operation load migration timing is generated and output, to realize the intelligent optimization of computing power server cluster power consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computing server power management technology, and more specifically, to a method and system for optimizing computing server power consumption based on deep learning. Background Technology

[0002] As data centers continue to expand and computing power demands grow, the power consumption of computing server clusters is becoming increasingly prominent, leading not only to a significant increase in operating costs but also putting enormous pressure on energy supply and environmental protection.

[0003] Traditional power management methods for computing servers primarily rely on simple threshold control and static resource allocation. For example, by setting a power consumption cap for the server, measures such as frequency reduction or shutdown are taken to reduce power consumption when the threshold is exceeded. However, these methods have significant limitations. Firstly, they do not fully consider the dynamic power consumption changes of server computing units under different task loads, making it impossible to accurately adjust power consumption based on actual load conditions, easily leading to resource waste or performance degradation. Secondly, static resource allocation methods cannot adapt to complex and ever-changing business scenarios. When faced with sudden high computing power demands, they may fail to allocate computing resources in a timely and reasonable manner, thus affecting the performance and power efficiency of the entire server cluster. Furthermore, existing methods lack effective analysis and utilization of the power coupling relationships between computing units, making it difficult to achieve global power balance optimization. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide a method for optimizing the power consumption of a computing server based on deep learning, the method comprising:

[0005] Collect power consumption sensing signal sequences and task load signal sequences of computing units in a computing server cluster within a continuous time segment, and encapsulate the power consumption sensing signal sequences and the task load signal sequences into a synchronous monitoring signal array with a unified time sequence label;

[0006] The synchronous monitoring signal array is subjected to power ripple morphology analysis processing to extract the power ripple peak-valley timing features of the computing unit and the load switching timing features of the computing load carried by the computing unit, and the power ripple peak-valley timing features and the load switching timing features are constructed into a ripple load correlation feature vector.

[0007] The ripple load associated feature vector is input into a pre-constructed power drift trend inference network. The power drift trend inference network infers the power drift direction of adjacent time segments of the ripple load associated feature vector, and generates the predicted power drift surface of the computing unit in the next scheduling cycle.

[0008] Based on the predicted power consumption drift surface, a set of candidate computing units whose power consumption drift coupling relationship meets the preset coupling screening conditions is selected from the computing unit resource pool of the computing server cluster. Then, power consumption matching processing is performed based on the power consumption demand profile of each computing unit in the candidate computing unit set and the computing load to be allocated, so as to obtain a computing unit power consumption matching allocation scheme.

[0009] Based on the power consumption matching and allocation scheme of the computing unit, a set of power balancing scheduling instructions containing computing unit addressing identifiers and computing load migration timing is generated and output.

[0010] Furthermore, embodiments of the present invention also provide a power consumption optimization system for computing servers based on deep learning, comprising:

[0011] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the aforementioned deep learning-based computing server power optimization method by executing the machine-executable instructions.

[0012] Based on the above, by collecting power consumption sensing signal sequences and task load signal sequences of computing units in a computing server cluster within continuous time segments, and encapsulating them into a synchronous monitoring signal array with a unified time sequence label, the synchronous monitoring signal array is subjected to power consumption ripple morphology analysis processing. The peak-valley time sequence features of power consumption ripple and the load switching time sequence features are extracted, and a ripple-load correlation feature vector is constructed. This deeply explores the intrinsic relationship between computing unit power consumption and load, enabling the capture of the power consumption variation patterns of computing units under different loads. The ripple-load correlation feature vector is input into a pre-constructed power drift trend inference network to generate a predicted power consumption drift surface for the computing unit in the next scheduling cycle, achieving accurate prediction of future power consumption changes for the computing unit. Based on the predicted power consumption drift surface, a candidate computing unit set is selected and power consumption matching processing is performed to obtain a computing unit power consumption matching allocation scheme. This scheme fully considers the power consumption drift coupling relationship between computing units and the load power consumption demand profile, enabling optimal allocation of computing resources and effectively avoiding situations of excessively high local power consumption or idle resources. Finally, based on the power consumption matching and allocation scheme of the computing unit, a set of power balancing scheduling instructions is generated and output, realizing dynamic and intelligent optimization of the power consumption of the computing server cluster, which significantly improves the power consumption efficiency of the server cluster, reduces operating costs, and ensures the performance stability of the server cluster. Attached Figure Description

[0013] Figure 1 This is a schematic diagram of the execution flow of the deep learning-based computing server power consumption optimization method provided in the embodiments of the present invention.

[0014] Figure 2This is a schematic diagram of exemplary hardware and software components of a deep learning-based computing server power optimization system provided in an embodiment of the present invention. Detailed Implementation

[0015] Figure 1 This is a flowchart illustrating a deep learning-based power consumption optimization method for computing servers, as provided in one embodiment of the present invention. A detailed description follows.

[0016] Step S110: Collect the power consumption sensing signal sequence and task load signal sequence of the computing units in the computing server cluster in a continuous time segment, and encapsulate the power consumption sensing signal sequence and the task load signal sequence into a synchronous monitoring signal array with a unified time sequence mark.

[0017] In this embodiment, the computing server cluster comprises K heterogeneous computing units. Each computing unit's power supply motherboard embeds a power consumption sensing probe consisting of a precision shunt resistor and a differential amplifier, as well as a task load monitoring agent running at the operating system kernel level. The power consumption sensing probe obtains the real-time power consumption of the computing unit by directly measuring the instantaneous current and voltage on the power supply bus, while the task load monitoring agent obtains the characteristics of the current computing load by reading the kernel task scheduler status and processor core occupancy registers. The two signals are acquired in parallel at the same sampling frequency and converged to a central data acquisition node for time synchronization and array encapsulation processing. This synchronization and encapsulation mechanism ensures that the power consumption changes and load changes within each signal frame are strictly aligned in time during subsequent power ripple pattern analysis.

[0018] Step S111: Deploy power consumption sensing probes on each computing unit of the computing server cluster. Use the power consumption sensing probes to acquire instantaneous current fluctuation signals and instantaneous voltage fluctuation signals of the power supply lines inside the computing unit according to a preset acquisition granularity. Perform signal aliasing modulation processing on the instantaneous current fluctuation signals and the instantaneous voltage fluctuation signals to generate a power consumption sensing signal sequence that reflects the instantaneous power consumption value of the computing unit at each acquisition point. Each signal point in the power consumption sensing signal sequence carries a timing marker corresponding to the acquisition point.

[0019] The power consumption sensing probe, at the hardware level, consists of a low-inductive precision shunt resistor connected in series with the power supply bus and a differential instrumentation amplifier. The shunt resistor has an extremely small resistance value to minimize the impact of additional power consumption and heat generation on the measurement. The differential instrumentation amplifier amplifies the small voltage difference across the shunt resistor to the full-scale input range of the analog-to-digital converter (ADC), suppressing common-mode noise interference. The ADC quantizes the amplified voltage signal at fixed sampling intervals. Since the transient current change frequency of the arithmetic unit may cover a wide frequency range from millisecond-level loads to microsecond-level clock gating, the sampling granularity needs to be fine enough to capture the detailed shape of the power consumption ripple. Simultaneously, the power supply bus voltage is attenuated to within the ADC's range and synchronously acquired through another voltage divider network. Signal aliasing modulation processing multiplies the instantaneous current and voltage values ​​acquired at the same time to obtain the instantaneous power consumption value at that moment. Each signal point in the power consumption sensing signal sequence is marked with a high-precision timestamp.

[0020] Step S112: Obtain the type identifier of the computing task running on the computing unit and the number of computing cores occupied by the computing task at each acquisition point. Generate the task load status characteristics of the computing unit at each acquisition point according to the type identifier and the number of computing cores. Arrange the task load status characteristics at all acquisition points according to the time sequence marker to generate a task load signal sequence that reflects the change of the task load of the computing unit over time. Each signal point in the task load signal sequence carries a time sequence marker consistent with the acquisition point corresponding to the power consumption sensing signal sequence.

[0021] The task load monitoring agent periodically reads the list of currently running tasks and their attributes through the application programming interface provided by the operating system. Each computational task is assigned a type identifier upon submission to the cluster to distinguish between compute-intensive, memory-intensive, or I / O-intensive load types. The monitoring agent obtains the type identifiers of tasks currently occupying each computation core and counts the total number of cores occupied by tasks of the same type. Task load status characteristics integrate two dimensions of information: qualitative type and quantitative scale. The type dimension describes what kind of task the current computational unit is processing, and the scale dimension describes how much hardware resources are being used to process these tasks. These two dimensions are encoded in vector form and arranged in the order of the collection points to form a task load signal sequence.

[0022] Step S113: Perform timing synchronization verification processing on the power consumption sensing signal sequence and the task load signal sequence; for the power consumption sensing signal sequence, when the timing mark deviation exceeds the preset synchronization tolerance threshold, perform time-stamp-based interpolation correction on the power consumption sensing value of the deviation signal point to obtain the timing deviation-corrected power consumption sensing signal sequence; for the task load signal sequence, when the timing mark deviation exceeds the preset synchronization tolerance threshold, perform alignment correction on the timing mark of the deviation signal point, and correct its task load state characteristics to the same state as the nearest valid timing mark signal point to obtain the timing deviation-corrected task load signal sequence, and perform signal frame segmentation processing according to the same time base to obtain a power consumption sensing signal frame and a task load signal frame containing the same number of signal points and whose timing marks are aligned one by one.

[0023] Since the power consumption sensing probe and the task load monitoring agent are two independent hardware and software paths, their respective crystal clock sources may have slight frequency differences. Furthermore, operating system scheduling delays may cause jitter between the task load signal acquisition time and the preset sampling time. The timing synchronization verification process first calculates the deviation between the acquisition timestamp carried by each signal point and the unified reference clock. For signal points in the power consumption sensing signal whose deviation exceeds the tolerance threshold, since the power consumption signal is a continuously changing analog quantity, a linear interpolation method is used to correct it based on the values ​​of adjacent compliant signal points. For signal points in the task load signal whose deviation exceeds the tolerance threshold, since the load state is a discretely changing categorical quantity, the timestamp is directly aligned to the reference clock, and the load state is corrected to the state of the nearest compliant signal point on the time axis. After correction, the two signals are divided into frame-by-frame aligned power consumption sensing signal frames and task load signal frames with a fixed frame length.

[0024] Step S114: Perform signal frame pairing processing on the power consumption sensing signal frame and the task load signal frame, bind the power consumption sensing signal frame and the task load signal frame with the same timing mark into a synchronous monitoring signal frame unit, and assign a globally unique timing mark to the synchronous monitoring signal frame unit.

[0025] In this embodiment, after the timing synchronization verification and signal frame segmentation in step S113, the power consumption sensing signal frame and the task load signal frame are stored independently in sequence. The time window corresponding to the m-th frame in the power consumption sensing signal frame sequence is completely consistent with the time window corresponding to the m-th frame in the task load signal frame sequence. The signal frame pairing process traverses the two sequences frame by frame, extracts the power consumption sensing signal frame and the task load signal frame with the same frame number, and combines them into a two-dimensional data structure at the data level. The two parts of this data structure store the power consumption signal value sequence and the load state feature sequence, respectively. The number of signal points in the two parts is equal and the timing is aligned point by point. The technical significance of the above pairing and binding is that when the power consumption ripple morphology is analyzed for each frame in the subsequent step S120, it is necessary to access the power consumption waveform information and load state information within the same time window at the same time. The synchronized monitoring signal frame unit after pairing and binding provides the convenience of the above simultaneous access, ensuring that the extracted power consumption ripple peak and valley timing features and load switching timing features correspond to the exact same observation period, eliminating the risk of time misalignment from the data organization level.

[0026] Step S115: Arrange all synchronous monitoring signal frame units in ascending order according to the global timing mark to form a synchronous monitoring signal array, and store the synchronous monitoring signal array in the timing signal buffer storage area for subsequent power consumption ripple mode analysis processing.

[0027] In this embodiment, step S114 assigns a globally unique timing marker to each synchronous monitoring signal frame unit. This timing marker is essentially a monotonically increasing integer sequence number, reflecting the chronological relationship of each frame on the original acquisition timeline. All synchronous monitoring signal frame units are sorted in ascending order according to this timing marker, and the sorted frame unit sequence is the synchronous monitoring signal array. This array is stored in a timing signal buffer storage area implemented by a circular buffer. The capacity of the buffer storage area is fixed. When a new synchronous monitoring signal frame unit arrives and the buffer storage area is full, the old frame with the smallest timing marker is automatically overwritten. The above storage mechanism ensures that the power consumption ripple pattern analysis processing is always based on the monitoring data within the most recent time window, which can capture the latest power consumption dynamic changes of the computing unit and control the storage overhead within a fixed range. At the same time, the ascending order allows subsequent processing to read the signal frame by frame in chronological order, which meets the causality requirements of timing analysis.

[0028] Step S120: Perform power ripple morphology analysis processing on the synchronous monitoring signal array, extract the power ripple peak and valley timing features of the computing unit and the load switching timing features of the computing load carried by the computing unit, and construct the power ripple peak and valley timing features and the load switching timing features into a ripple load correlation feature vector.

[0029] In this embodiment, the power ripple morphology analysis process analyzes the dynamic features of power consumption and load from within each frame of the synchronous monitoring signal frame. The power consumption sensing signal frame records the continuous change curve of the power consumption value of the computing unit in a short period of time. This curve contains various peaks and troughs caused by factors such as load changes, clock frequency adjustments, and power supply noise. By detecting the location and shape parameters of these extreme points, the peak-valley timing features of the power ripple can be extracted. The task load signal frame records the changes in the type and scale of the computing tasks within the same time period. By detecting the time when the load state switches and the differences before and after the switch, the load switching timing features can be extracted. Finally, the two types of features are fused and encoded through a pre-constructed timing feature fusion mapper to generate a ripple load correlation feature vector with fixed dimensions and rich dynamic information, which serves as the input to the downstream power drift trend inference network.

[0030] Step S121: Extract the power consumption sensing signal frame in each synchronous monitoring signal frame unit from the synchronous monitoring signal array, perform sliding window local extremum detection processing on the power consumption sensing signal frame, locate the local power consumption maximum point and local power consumption minimum point of the power consumption sensing signal frame within the sliding window range, the local power consumption maximum point constitutes the peak point set, and the local power consumption minimum point constitutes the valley point set.

[0031] In this embodiment, the power consumption sensing signal frame is a one-dimensional signal sequence. Small fluctuations may exist between adjacent signal points due to measurement noise, and direct point-by-point comparison would detect a large number of meaningless pseudo-extreme points. Therefore, a sliding window local extremum detection method is adopted: a fixed-width window slides point-by-point along the signal sequence, retaining only the signal point with the largest power consumption value within each window as a candidate peak point and the signal point with the smallest power consumption value as a candidate valley point. The sliding step size is smaller than the window width, allowing adjacent windows to overlap, ensuring that no true extremum points located at the window boundaries are missed. All candidate extremum points also need to be merged and deduplicated: if the signal point index distance between two similar candidate extremum points is less than a preset merging distance, only the one whose power consumption value is closer to the window extremum definition is retained, to avoid the same true extremum being repeatedly detected due to window overlap.

[0032] Step S122: Perform neighborhood waveform analysis on each local power consumption maximum point in the peak point set, extract the rising edge duration and falling edge duration corresponding to each local power consumption maximum point, and calculate the peak sharpness parameter based on the ratio of the rising edge duration to the falling edge duration to generate a peak feature descriptor that includes the peak point occurrence sequence, peak amplitude and the peak sharpness parameter.

[0033] In this embodiment, each local power maximum point in the peak point set represents the position and amplitude of a local high point on the power consumption curve. Position and amplitude alone are insufficient to fully describe the morphological characteristics of the peak: some peaks may be steep spikes caused by a sudden increase in computational load, rising and falling rapidly; others may be gentle plateaus caused by continuous execution of batch tasks, rising slowly and remaining at a high level for a longer period. These two types have drastically different effects on power consumption drift; steep spikes may cause instantaneous power overshoot, while gentle plateaus may lead to continuous heat accumulation. Therefore, further analysis of the waveform trend information within the neighborhood of each peak point is needed. Neighborhood waveform analysis uses the peak point as the center, extending a fixed number of signal points to the left and right to form a local observation window. Within this window, the path of power consumption rising from the background level to the peak is traced to the left, and the path of power consumption falling back from the peak level to the background level is traced to the right. The rising edge duration is defined as the time span from the signal point where the power consumption first exceeds the peak amplitude multiplied by a preset threshold to the peak point, reflecting the speed of power consumption increase. The falling edge duration is defined as the time span between the peak point and the signal point where the power consumption first falls below the peak amplitude multiplied by another preset threshold, reflecting the speed at which power consumption falls back. The peak sharpness parameter is the ratio of the falling edge duration to the rising edge duration. A larger ratio indicates a more asymmetrical peak and a more sluggish falling process, while a smaller ratio indicates a more symmetrical peak and similar rising and falling speeds. The peak occurrence timing, peak amplitude, and peak sharpness parameter together form the peak feature descriptor, which fully records the timing, intensity level, and sharpness of the peak.

[0034] Step S123: Perform neighborhood waveform analysis on each local power minima in the valley point set, extract the duration of the falling edge and the duration of the rising edge corresponding to each local power minima, and calculate the valley sharpness parameter based on the ratio of the duration of the falling edge to the duration of the rising edge, generating a valley feature descriptor that includes the valley point occurrence sequence, valley amplitude, and the valley sharpness parameter; mix and sort all peak point occurrence sequences and all valley point occurrence sequences in chronological order to form a unified peak-valley event timestamp sequence.

[0035] The neighborhood waveform analysis of valley points is, in principle, the inverse of that of peak points. A valley represents the location and amplitude of a local low point on the power consumption curve. Some valleys are sharp dips formed by brief idle intervals, while others are gentle valleys formed during task breaks. A fixed number of signal points are extended to the left and right of the valley point to form a neighborhood observation window. Tracing the path of power consumption from a higher level to the valley to the left yields the duration of the falling edge; tracing the path of power consumption from the valley back to a higher level to the right yields the duration of the rising edge. The valley sharpness parameter is the ratio of the duration of the falling edge to the duration of the rising edge. The timing sequences of all peak and valley occurrences are mixed together and arranged chronologically to form a peak-valley event timestamp sequence that completely marks the timeline of all local extreme events on the power consumption curve.

[0036] Step S124: Based on the peak-valley event timestamp sequence, perform statistical processing on the time interval between adjacent peak point timestamps and valley point timestamps to generate a peak-valley time-series spacing distribution histogram, and use the peak-valley time-series spacing distribution histogram as part of the power consumption ripple peak-valley time-series features; at the same time, arrange all peak feature descriptors and valley feature descriptors in the order of their corresponding timestamps to form a peak-valley feature descriptor list, which serves as another part of the power consumption ripple peak-valley time-series features.

[0037] The time interval between adjacent peaks and valleys in the peak-valley event timestamp sequence reflects the local fluctuation frequency characteristics of power consumption ripple. Frequent occurrences of short time intervals between peaks and valleys indicate rapid power consumption fluctuations, suggesting the computing unit may be in an unstable state with frequent load switching. Conversely, generally longer time intervals indicate slower power consumption changes, suggesting the computing unit may be in stable operation or continuous idle state. By dividing the time intervals into numerical intervals, statistically analyzing their frequency, and normalizing the results, a peak-valley time-series spacing distribution histogram is obtained. This histogram statistically describes the density of power consumption ripple fluctuations at different time scales. Furthermore, a peak-valley feature descriptor list arranged in peak-valley event timestamp order preserves the individual morphological information of each local extreme point, complementing the spacing distribution histogram: the histogram provides an overall statistical perspective, while the descriptor list provides a detailed perspective for each point.

[0038] Step S125: Extract the task load signal frame within each synchronous monitoring signal frame unit from the synchronous monitoring signal array, perform point-by-point differential comparison processing on the task load state characteristics between adjacent signal points in the task load signal frame, detect the signal point position where the task load state characteristics change, and mark the signal point position as a load switching event point.

[0039] The task load state feature vector of each signal point in the task load signal frame contains the one-hot encoded type of the currently running computation task and the normalized value of the number of computation cores occupied. When any component of this vector changes between two adjacent signal points, it means that the type of task running on the computation unit has changed or the core usage scale has been adjusted. The above-mentioned load state changes are one of the fundamental causes of power ripple peaks and valleys. Point-by-point differential comparison processing traverses the task load signal frame from front to back, calculates the sum of the absolute values ​​of the differences of each component of the state feature vector for each pair of adjacent signal points, and marks the next signal point as a load switching event point if the sum is greater than zero.

[0040] Step S126: Extract the task load state characteristics before and after the change corresponding to each load switching event point, calculate the difference in the number of computing cores occupied between the before and after states, generate a load switching feature that includes the load switching sequence, the type of the state before the switch, the type of the state after the switch, and the difference in the magnitude of the difference, arrange all load switching features according to the load switching sequence, construct a load switching event sequence that reflects the switching density and the distribution characteristics of the switching magnitude of the computing load in the time dimension, and use the load switching event sequence as the load switching sequence feature.

[0041] The state before the change is taken from the task load state characteristics of the signal point immediately preceding the load switching event point, while the state after the change is taken from the state characteristics of the load switching event point itself. The difference in the number of computing cores reflects the impact of this load switching on the hardware resource usage of the computing unit: the greater the change in the number of cores, the more drastic the power consumption response tends to be. The switching time, task type before and after the switch, and change in resource usage for each switching event are organized into a four-tuple structure. The resulting load switching event sequence, arranged chronologically by switching time, completely describes all load change events and their attributes experienced by the computing unit within the observed time window. This sequence is directly output as the load switching timing feature.

[0042] Step S127: Call the pre-built temporal feature fusion mapper to fuse and encode the peak-valley temporal interval distribution histogram and the peak-valley feature descriptor list, mapping them into a fixed-length ripple feature vector; map the load switching event sequence into a fixed-length load feature vector; and perform vector concatenation processing on the ripple feature vector and the load feature vector to generate the ripple load association feature vector.

[0043] The temporal feature fusion mapper consists of two independent encoding branches. The first branch processes the input after fusing the peak-valley temporal spacing distribution histogram and the peak-valley feature descriptor list. After passing through a fully connected layer and a gated recurrent unit sequence encoding layer, it outputs a fixed-dimensional ripple feature vector. The second branch takes the load switching event sequence as input, with a switching event quadruple at each time step. After being encoded by a gated recurrent unit, the hidden state vector of the last time step is mapped to a fixed-dimensional load feature vector through a fully connected layer. The outputs of the two branches are concatenated along the vector dimension to generate a ripple load-related feature vector.

[0044] Step S130: Input the ripple load associated feature vector into the pre-constructed power drift trend inference network. The power drift trend inference network infers the power drift direction of adjacent time segments of the ripple load associated feature vector and generates the predicted power drift surface of the computing unit in the next scheduling cycle.

[0045] In this embodiment, the power drift trend extrapolation network adopts a deep network architecture consisting of a time-series feature encoding layer, a drift state space modeling layer, a time difference prediction layer, a gradient recursive aggregation layer, a noise adversarial regularization layer, a trend extrapolation expansion layer, and a surface reconstruction layer. Its core idea is to use the ripple load correlation feature vectors extracted from multiple consecutive time segments as time-series inputs, abstracting and refining the evolution trend of power consumption state layer by layer, and finally extrapolating the trend to future time periods to form a continuous predicted power drift surface. This predicted power drift surface describes the trajectory of the power drift amplitude of the computing unit from the current moment to the end of the next scheduling cycle, serving as the core basis for subsequent candidate computing unit selection and power matching allocation.

[0046] Step S131: Input the ripple load associated feature vector into the time-series feature encoding layer of the power consumption drift trend inference network. The time-series feature encoding layer performs time-series dependency encoding on the ripple load associated feature vector to generate a deep time-series representation vector carrying contextual power consumption time-series information.

[0047] The temporal feature encoding layer does not receive the ripple load-related feature vector at a single moment, but rather a sequence of feature vectors from multiple consecutive temporal segments. Its core network structure is a bidirectional long short-term memory network, containing a forward layer that processes the sequence from front to back and a backward layer that processes the sequence from back to front. The forward layer can capture the evolution of power consumption states from the past to the present, while the backward layer can infer the temporal importance of the current feature from the future. After the input feature vector of each temporal segment is encoded in both directions, the forward and backward hidden states are concatenated along the vector dimension, so that the output deep temporal representation vector simultaneously integrates accumulated information from the past and backward reference information from the future, thus possessing stronger temporal awareness.

[0048] Step S132: Input the deep time series representation vector into the drift state space modeling layer of the power consumption drift trend inference network. The drift state space modeling layer uses the internal drift state transition parameter matrix to perform state space projection processing on the deep time series representation vector to generate hidden state feature vectors in the drift state space. The hidden state feature vectors encode the influence mode of power consumption ripple shape and load switching status on power consumption drift in the current time segment.

[0049] The drift-state space modeling layer projects the deep temporal representation vector from the original feature space into a specially designed drift-state space. The physical meaning of each dimension of this space is aligned with the dynamic characteristics of power consumption drift. Internally, it maintains a trainable drift-state transition parameter matrix. A linear transformation maps the input vector to the hidden state space, while a nonlinear activation function introduces complex mapping capabilities. The generated hidden state feature vector is a highly abstract encoding of the ripple pattern, load switching behavior, and coupling relationships within the current time segment.

[0050] Step S133: Input the hidden state feature vector into the time difference prediction layer of the power consumption drift trend inference network. The time difference prediction layer performs time difference operation on the hidden state feature vector to calculate the difference offset vector between the hidden state of the current time segment and the hidden state of the historical time segment. The difference offset vector is used as the original gradient of the power consumption drift direction.

[0051] The essence of drift is the change of state over time. The temporal differential prediction layer subtracts the hidden state feature vector of the current time segment from the hidden state feature vector of the previous time segment element-wise. The resulting differential offset vector represents the direction and magnitude of the change in the hidden power state between two adjacent time segments. The sign of each element in the differential offset vector indicates whether the dimension of the hidden state increases or decreases, and the absolute value indicates the rate of change. This differential offset vector is defined as the original gradient of the power drift direction and is the basis signal for subsequent drift trend aggregation.

[0052] Step S134: Input the original gradient of the power drift direction into the gradient recursive aggregation layer of the power drift trend inference network. The gradient recursive aggregation layer uses an expanding recursive window to recursively aggregate the differential offset vectors corresponding to multiple adjacent time segments to generate an aggregated drift trend vector.

[0053] The differential offset vector of a single pair of adjacent time segments only reflects the drift direction in the short term and is easily affected by local random fluctuations and measurement noise. The gradient recursive aggregation layer adopts an expanding recursive window mechanism, backtracking the differential offset vectors of multiple segments from the current time segment, and assigning an expansion weight that decreases exponentially with the step size for each backtracking step. Differential offset vectors closer to the current time receive a larger weight, while older differential offset vectors have a smaller weight. The aggregated drift trend vector, obtained by weighting and summing all differential offset vectors within the window according to the expansion weight, not only captures the long-term cumulative direction of power consumption drift, but also highlights the dominant role of recent change trends through the weighted decay mechanism.

[0054] Step S135: Input the aggregated drift trend vector into the noise adversarial regularization layer of the power consumption drift trend inference network. The noise adversarial regularization layer performs signal perturbation injection processing and adversarial denoising processing on the aggregated drift trend vector to generate a regularized drift trend vector with noise resistance robustness.

[0055] In real-world operating environments, power consumption monitoring signals inevitably contain acquisition noise and transient environmental disturbances. If the training phase only encounters ideal, noise-free data, trend inference may be biased when faced with noisy data during the inference phase. The noise adversarial regularization layer employs an adversarial training approach: during forward propagation, a random noise signal following a Gaussian distribution is injected into the aggregated drift trend vector. The perturbed vector is then input into a denoising autoencoder branch, which, after encoding compression and decoding, reconstructs and outputs the noise-free clean signal. During the training phase, the denoising autoencoder branch is trained by minimizing the reconstruction error between the clean signal and the unperturbed original signal. Simultaneously, adversarial loss ensures that the network maintains consistency in trend inference even with noise perturbations.

[0056] Step S136: Input the regularized drift trend vector into the trend extrapolation expansion layer of the power consumption drift trend inference network. The trend extrapolation expansion layer projects the regularized drift trend vector along the time axis in the positive direction to the prediction time series interval corresponding to the next scheduling period, generating discrete drift amplitude values ​​for each prediction time series point within the prediction interval.

[0057] The regularized drift trend vector is a compressed global trend representation that needs to be expanded in the time dimension to represent the specific drift magnitude values ​​at each prediction time point within the next scheduling period. The trend extrapolation expansion layer contains a fully connected prediction network whose input is the regularized drift trend vector, and whose output vector dimension equals the total number of discrete prediction time points within the next scheduling period. The multi-layer nonlinear transformation of the fully connected network can learn the complex mapping relationship from the compressed trend representation to the drift magnitude at multiple time points.

[0058] Step S137: Input the discrete drift amplitude value into the surface reconstruction layer of the power consumption drift trend inference network. The surface reconstruction layer performs continuous surface interpolation reconstruction processing on the discrete drift amplitude values ​​of all prediction time series points within the prediction interval to generate a prediction power consumption drift surface that covers the entire prediction time series interval and whose drift amplitude changes continuously.

[0059] The discrete drift amplitude value provides a numerical estimate of the power consumption drift only at a finite number of prediction time points. The surface reconstruction layer uses spline interpolation to continuously fill the gaps between adjacent discrete prediction time points using piecewise polynomial curves, making the prediction results smooth and continuous on the time axis. The generated predicted power consumption drift surface has the time axis as the horizontal axis and the drift amplitude as the vertical axis. The surface trajectory of each computing unit on this surface is the complete predicted image of the unit's power consumption smoothly evolving from the current moment to the end of the next scheduling cycle.

[0060] Step S140: Based on the predicted power consumption drift surface, select a set of candidate computing units in the computing unit resource pool of the computing server cluster whose power consumption drift coupling relationship meets the preset coupling screening conditions, and perform power consumption matching processing based on the power consumption demand profile of each computing unit in the candidate computing unit set and the computing load to be allocated, so as to obtain a computing unit power consumption matching allocation scheme.

[0061] In this embodiment, the computing server cluster operates a large number of computing units simultaneously. When it is necessary to select execution units for a batch of computing loads to be allocated, not just any idle computing unit is suitable. Some computing units may already be in a power consumption increase phase in the next scheduling cycle, and allocating new loads will exacerbate the risk of power overshoot; while other computing units may be in a power consumption decrease phase, and allocating new loads can fill the downward trend and make the overall power consumption more stable. Therefore, it is necessary to perform spatiotemporal alignment and matching between the power consumption trajectory of the computing loads to be allocated in historical operation and the predicted power consumption drift surface of each candidate computing unit, and select the computing units whose power consumption drift trend is most complementary to the load power consumption requirements in terms of time and amplitude to carry the task, thereby achieving balanced optimization of the overall power consumption of the cluster.

[0062] Step S141: Analyze the drift amplitude value of the predicted power consumption drift surface at each predicted timing point in the next scheduling cycle, perform time-series segmented statistical processing on the drift amplitude value, calculate the average drift amplitude value and drift amplitude fluctuation variance of the predicted power consumption drift surface in each segment interval, and generate a surface partition feature set containing partition identifiers and corresponding average drift amplitude values ​​and drift amplitude fluctuation variances.

[0063] In this embodiment, the predicted power drift surface is a continuous curve with time as the horizontal axis and drift amplitude as the vertical axis. The duration of the next scheduling cycle is typically much longer than a single acquisition time segment. Directly matching the entire continuous curve with the historical data of candidate computing units would be computationally intensive and lack local specificity. Time-series segmented statistical processing divides the next scheduling cycle into several continuous and non-overlapping time partitions at fixed time intervals, with each partition corresponding to a sub-period within the scheduling cycle. For the drift amplitude values ​​at all discrete predicted time points within each partition, their arithmetic mean is calculated as the average drift amplitude value for that partition, reflecting whether the overall direction of power drift within that sub-period is upward or downward and its magnitude. Simultaneously, the variance of the drift amplitude values ​​at all predicted time points within that partition relative to the partition's average value is calculated as the drift amplitude fluctuation variance, reflecting the stability of power drift within that sub-period: a smaller fluctuation variance indicates a more stable drift trend, while a larger fluctuation variance indicates a more volatile drift trend. The partition identifier and its corresponding average drift amplitude value and drift amplitude fluctuation variance are packaged to form a surface partition feature set.

[0064] Step S142: Obtain the list of currently active computing units from the computing unit resource pool, and extract the historical power consumption drift surface of each active computing unit within the historical scheduling period. Perform partitioning statistical processing on the historical power consumption drift surface according to the same time partitioning method as the surface partitioning feature set, and generate the historical average drift amplitude value and historical drift amplitude fluctuation variance for each active computing unit corresponding to each partition.

[0065] In this embodiment, the computing unit resource pool maintains a status information table for all online computing units. An active state indicates that the computing unit is currently powered on and the operating system has completed booting, ready to accept task scheduling at any time. A snapshot of the actual power consumption drift surface recorded in the previous scheduling cycle is retrieved from the historical monitoring database based on the unique identifier of each active computing unit. This historical power consumption drift surface is partitioned using the same time partitioning method as in step S141, i.e., the partition boundaries are aligned with the relative time positions of the same scheduling cycle to ensure time consistency in partition-to-partition comparisons. For each partition of each active computing unit, two statistical measures are also calculated: the average drift amplitude and the variance of the drift amplitude fluctuation.

[0066] Step S143: For each active computing unit, the historical average drift amplitude value and the average drift amplitude value of the predicted power consumption drift surface in the corresponding partition are processed by partition-by-partition difference calculation to obtain the average drift amplitude difference of each partition. At the same time, the historical drift amplitude fluctuation variance and the drift amplitude fluctuation variance of the predicted power consumption drift surface in the corresponding partition are processed by partition-by-partition difference calculation to obtain the fluctuation variance difference of each partition.

[0067] In this embodiment, the core criterion for selecting candidate computing units is the degree of matching or complementarity between the historical power consumption drift pattern and the predicted power consumption drift trend of the computing unit. The average drift amplitude difference reflects the absolute deviation between the historical drift direction and amplitude of the computing unit in the sub-period and the expected drift direction and amplitude of the predicted drift surface. The smaller the difference, the more consistent the historical performance of the computing unit in the sub-period is with the current predicted trend. The fluctuation variance difference reflects the degree of difference between the two in drift stability. The smaller the difference, the more similar the historical power consumption fluctuation of the computing unit in the sub-period is to the current predicted fluctuation. The two differences comprehensively measure the degree of coupling from the two dimensions of drift amplitude matching and fluctuation stability matching.

[0068] Step S144: Using the average drift amplitude difference and fluctuation variance difference of each partition as input, call the drift coupling relationship quantization function to perform drift coupling relationship quantization processing on each active computing unit, and calculate the local coupling metric between each active computing unit and the predicted power consumption drift surface in each partition. The local coupling metric is inversely proportional to the average drift amplitude difference and inversely proportional to the fluctuation variance difference.

[0069] In this embodiment, the drift coupling quantization function maps the difference in average drift amplitude and the difference in volatility variance to a local coupling metric. The design principle of this quantization function is that a larger coupling metric indicates a higher degree of matching between the two values ​​in that partition, and vice versa. Therefore, the local coupling metric is inversely proportional to the two differences. The quantization function scales the difference in average drift amplitude using a normalization factor and then weights it with the difference in volatility variance. This weighted sum is then subtracted from an upper limit of a constant. Finally, the result is transformed using a monotonically increasing nonlinear mapping function with a bounded output between 0 and 1 to obtain the local coupling metric. This mapping relationship ensures that the local coupling metric reaches its maximum value when both differences are zero (i.e., a perfect match), and approaches zero when the two differences approach infinity.

[0070] Step S145: Perform a weighted summation on the local coupling metric values ​​of each active computing unit in each partition to generate a global drift coupling comprehensive metric value for each active computing unit. The weight coefficients used in the weighted summation are inversely proportional to the time distance between the partitions at the current moment.

[0071] In this embodiment, the predicted power drift surface covers the entire next scheduling cycle. However, the closer a partition is to the current time, the more reliable its prediction results and the greater its reference value for scheduling decisions. Conversely, partitions farther from the current time are more susceptible to uncertainties. Therefore, a time decay weighting strategy is adopted when synthesizing the coupling metrics of each partition: a weight coefficient is assigned to each partition, which decays as the time distance from the current time increases. This decay can be linear or exponential. The local coupling metric of each partition is multiplied by its corresponding weight coefficient and then summed to obtain the global drift coupling comprehensive metric. This comprehensive metric not only reflects the overall level of coupling but also highlights the priority of recent trend matching through time decay weighting.

[0072] Step S146: Compare the global drift coupling comprehensive metric value of each active computing unit with the coupling metric threshold defined in the preset coupling screening conditions, screen out active computing units whose global drift coupling comprehensive metric value meets the coupling metric threshold, and form a candidate computing unit set by the screened active computing units.

[0073] The coupling metric threshold is set based on the cluster's historical scheduling experience. The selected computing units are those active units whose historical power consumption drift patterns match the current predicted trends with a high degree of accuracy. These units are then used as candidate computing units in the subsequent resource matching process.

[0074] Step S147: Extract the number of available computing cores and available memory capacity of each candidate computing unit in the candidate computing unit set in the next scheduling cycle, construct the available resource configuration vector of the candidate computing unit, evaluate the matching degree between the available resource configuration vector of the candidate computing unit and the resource demand vector of the computing load to be allocated, eliminate candidate computing units whose resource matching degree is lower than the preset resource matching threshold, sort the remaining candidate computing units after resource matching screening in descending order according to the global drift coupling comprehensive metric value, and output the candidate computing unit set.

[0075] In this embodiment, after power consumption coupling matching is passed, basic resource capacity verification is also required. The available resource configuration vector contains two dimensions: the number of available idle computing cores and the available idle memory capacity, both of which are normalized. The resource requirement vector of the computing load to be allocated also contains two normalized dimensions: the minimum number of computing cores required and the minimum memory capacity required. The resource matching degree is equal to the dot product of the two vectors divided by the magnitude of the requirement vector. After removing candidate units whose resource matching degree does not meet the threshold, the final set of candidate computing units is output in descending order of the global drift coupling comprehensive metric.

[0076] Step S148: Obtain historical power consumption trajectory data of the computing load to be allocated within the historical operating cycle, perform time-series segmentation processing on the historical power consumption trajectory data to obtain the stage power consumption sub-trajectory of the computing load to be allocated in different operating stages within the operating cycle, and define the combination of all stage power consumption sub-trajectories as the load power consumption demand profile, perform feature extraction processing on each stage power consumption sub-trajectory in the load power consumption demand profile, and extract the initial power consumption value, peak power consumption value, power consumption rise slope and power consumption duration of each stage power consumption sub-trajectory to form a stage power consumption demand feature vector.

[0077] In this embodiment, the computational load typically follows a fixed power consumption pattern during execution: power consumption gradually increases during the startup phase, remains high during the computationally intensive phase, and fluctuates and decreases during the result write-back phase. The historical power consumption trajectory is divided into different stages according to the inflection point of power change. For each stage, four key parameters describing the power consumption behavior of that stage are extracted: initial power consumption value, peak power consumption value, power consumption rise slope, and power consumption duration. These parameters form the load power consumption demand profile, which is used for stage-level comparison with the historical power supply capabilities of candidate computing units.

[0078] Step S149: Extract the historical power consumption drift surface of each candidate computing unit in the corresponding historical scheduling period from the candidate computing unit set, perform stage alignment mapping on the historical power consumption drift surface, and determine the historical power consumption supply capability characteristics of each candidate computing unit in the time window corresponding to each running stage of the computing load to be allocated. The historical power consumption supply capability characteristics include the historical average power consumption level and historical power consumption fluctuation amplitude in the corresponding time window.

[0079] The stage alignment mapping maps each stage in the load power demand profile to the historical power drift surface of the candidate computing unit in chronological order. Within the mapping time window, the mean and standard deviation of the historical drift surface are calculated as the historical average power consumption level and historical power consumption fluctuation amplitude of that stage.

[0080] Step S1410: Perform stage power matching operation on the historical power supply capability feature of each candidate computing unit and the stage power demand feature vector of the stage corresponding to the computing load to be allocated, and calculate the power supply and demand deviation parameter under each stage. The power supply and demand deviation parameter reflects the degree of difference between the historical power supply capability of the candidate computing unit and the power demand of the load.

[0081] The stage power consumption matching operation determines the inclusion of the historical average power consumption level of each candidate computing unit in each stage within the range formed by the initial and peak power consumption values ​​of the load demand in the same stage. If the historical average power consumption level falls within this range, the supply deviation is small; otherwise, it is large. Simultaneously, a rate matching determination is performed on the historical power consumption fluctuation amplitude and the power consumption rise slope in the same stage as the load demand. The stage power consumption supply and demand deviation parameters are then obtained.

[0082] Step S1411: Serialize and concatenate the power supply and demand deviation parameters of all stages to construct a full-stage power deviation vector corresponding to each candidate computing unit. Perform deviation smoothing and normalization on the full-stage power deviation vector to eliminate the step jump in power deviation between adjacent stages and generate a normalized full-stage power deviation curve. Identify the operating stage corresponding to the deviation peak point in the normalized full-stage power deviation curve. Mark the operating stage where the identified deviation peak point is located as a risk mismatch stage. Perform weighted deduction processing on the number of risk mismatch stages and the deviation magnitude of the risk mismatch stages for each candidate computing unit to obtain the power matching and adaptation score of each candidate computing unit.

[0083] Deviation smoothing and normalization performs a local weighted average of deviation parameters between adjacent stages to eliminate abrupt changes introduced by stage boundary division. Risk mismatch stages refer to those where the peak deviation exceeds a preset mismatch threshold; these stages are high-risk periods where insufficient power supply or overshoot may occur after scheduling. Points corresponding to the number and magnitude of risk mismatch stages are deducted from the full score, with the deduction weight set according to the stage's criticality to task operation.

[0084] Step S1412: Select candidate computing units whose power consumption matching and adaptation scores meet the preset adaptation threshold, bind and pair the selected candidate computing units with the computing load to be allocated, generate a pairing mapping table containing the addressing identifier of the candidate computing units and the identifier of the computing load to be allocated, and for each pair of bound and paired candidate computing units and computing loads to be allocated in the pairing mapping table, allocate the corresponding computing unit occupancy time window according to the power consumption demand time interval of each stage in the load power consumption demand profile, and generate the computing unit power consumption matching and allocation scheme.

[0085] Each record in the pairing mapping table associates the computing load to be allocated with the optimal candidate computing unit selected through power consumption matching. The computing unit occupancy window is strictly allocated according to the power consumption duration of each stage in the load power consumption demand profile, ensuring that the computing unit is reserved throughout the entire operation cycle from start to finish of the load.

[0086] Step S150: Based on the power consumption matching and allocation scheme of the computing unit, generate and output a set of power balancing scheduling instructions that includes the addressing identifier of the computing unit and the timing of computing load migration.

[0087] In this embodiment, the power consumption matching allocation scheme generated in step S140 specifies the target computing unit and the occupied time window for each computing load to be allocated. However, a series of conversions, such as address resolution, network location, data transmission estimation, and instruction orchestration, are still required between the allocation scheme and the instructions that can be actually executed by the scheduled execution nodes. The power consumption balancing scheduling instruction set encapsulates this information, enabling it to be directly distributed to each computing unit and scheduling management node for execution via the network.

[0088] Step S151: Analyze the pairing mapping relationship between the computing unit addressing identifier and the computing load identifier to be allocated in the computing unit power consumption matching allocation scheme, as well as the computing unit occupancy time window corresponding to each pairing mapping relationship, and extract the start timing and end timing of the computing unit addressing identifier, the computing load identifier to be allocated, and the occupancy time window for each pairing mapping relationship.

[0089] The power consumption matching and allocation scheme for computing units stores a pairing mapping table in the form of a data structure. Each record contains four fields: a computing unit addressing identifier to uniquely identify a computing unit within the cluster; a computing load identifier to uniquely identify a computing task waiting to be scheduled; and start and end timings to define the precise time window in which the computing unit is occupied to execute the computing load. The parsing process reads the pairing mapping table record by record, extracting the above four fields as input parameters for subsequent location and instruction generation.

[0090] Step S152: Based on the addressing identifier of the computing unit in each pairing mapping relationship, query the topology database of the computing server cluster to obtain the network addressing path and communication port identifier of the computing unit within the cluster, and generate computing unit location information containing the network addressing path and communication port identifier of the computing unit.

[0091] The topology database of the computing server cluster stores the location information of each computing unit in the cluster's physical network, including the Media Access Control (MAC) address and Internet Protocol (IP) address of the switch port to which the computing unit is connected, and the communication port identifier of the task receiving daemon running on the computing unit. The network addressing path describes the data forwarding path from the scheduling and management node through various levels of switching equipment to the computing unit. The computing unit location information encapsulates the above network reachability information into a structure, allowing subsequent load balancing commands to directly use the IP address and communication port identifier from the location information to establish data transmission connections.

[0092] Step S153: Based on the identifier of the computing load to be assigned in each pairing mapping relationship, query the load storage location mapping table of the computing load storage management node, obtain the storage path of the executable file of the computing load to be assigned and the storage node addressing information of the load dependent dataset, and generate the load source data location information.

[0093] After a computational workload is submitted to the cluster, its executable files and dependent datasets are stored in a specific path on the distributed file system. The workload storage location mapping table, indexed by the workload identifier, records the Internet Protocol address and specific directory path of the storage node for all files within that workload. The workload source data location information contains this storage location information, and during execution, the migration command uses this information to transfer the workload data from the storage node to the target computing unit via the network.

[0094] Step S154: Based on the start timing of the occupied time window of each pairing mapping relationship and the data location information of the load source end, calculate the estimated transmission time required for data transmission between the storage node corresponding to the data location information of the load source end and the computing unit corresponding to the computing unit location information, and reserve the estimated transmission time forward based on the start timing of the occupied time window to obtain the computing load migration start timing.

[0095] Data transfer from one node to another in a cluster network takes time, which depends on the total amount of data to be transferred divided by the available network bandwidth. The load source data location information records the total number of bytes in the executable file and dependent datasets of the load. Combined with the estimated bottleneck bandwidth along the path from the storage node to the target computing unit in the cluster network topology, the estimated transfer time is calculated. If the load migration only begins at the start of the occupied time window, the computing unit will have to wait for the data transfer to complete before starting task execution, resulting in wasted window time. Therefore, the migration start sequence is set to be offset forward by the estimated transfer time from the start of the occupied time window, ensuring that the data arrives at the target computing unit before the window begins.

[0096] Step S155: Based on the computing load migration start timing, and combined with the start timing of the occupied time window and the computing unit location information, construct a single load migration instruction containing the migration start timing, the target computing unit network addressing path, the target computing unit communication port identifier, the load identifier to be migrated, and the load source data location information.

[0097] A single load migration instruction encodes the above information into a structured instruction frame. The instruction frame header contains an instruction type code identifying it as a load migration instruction, and the load portion contains the specific values ​​of each of the aforementioned fields. The communication port identifier tells the source storage node which port to send data to the target computing unit, and the load identifier to be migrated informs the target computing unit which computing task will be received and started.

[0098] Step S156: Generate a corresponding single load migration instruction for each pairing mapping relationship, and arrange all the single load migration instructions in the order of their respective migration start times to form a load migration instruction sequence.

[0099] The load migration instruction sequence, ordered by migration start time, allows the cluster scheduling and management nodes to trigger the execution of each instruction sequentially according to time progression. Furthermore, the ordered sequence facilitates subsequent conflict detection and handling, clearly showing which instructions overlap in execution time.

[0100] Step S157: Perform conflict detection processing on the instructions with overlapping migration time windows in the load migration instruction sequence, identify the overlapping instructions with conflicts, and perform serialization adjustment processing on the migration start timing of the overlapping instructions. Rearrange the execution order of the load migration instructions with conflicts according to the preset priority rules to generate a conflict-free load migration instruction sequence.

[0101] When the migration time windows of two load migration instructions overlap, it means that multiple load data streams need to be transferred from the storage node to the target computing unit within the same time period. If these transmissions pass through a shared network link bottleneck, bandwidth contention may occur, causing all transmissions to slow down or even time out. The conflict detection process iterates through each pair of instructions in the load migration instruction sequence, calculating the length of the time overlap interval between their migration time windows. If the overlap interval length exceeds the tolerance limit, these two instructions are marked as overlapping conflict instructions. Conflict resolution follows a preset priority rule: the priority rule comprehensively considers the urgency of the computing load to be allocated, the power consumption drift coupling metric of the computing unit, and the size of the load data. High-priority instructions maintain their original migration start sequence, while the migration start sequence of low-priority instructions is postponed until after the estimated completion of the high-priority instruction's transmission. The adjusted load migration instruction sequence eliminates network bandwidth contention conflicts.

[0102] Step S158: The conflict-free load migration instruction sequence and the time window occupied by each pairing mapping relationship in the power consumption matching and allocation scheme of the computing unit are encapsulated by the instruction protocol to generate a power balance scheduling instruction set containing the scheduling time window definition of each load migration instruction and the corresponding computing unit. The power balance scheduling instruction set is broadcast to the scheduling execution node of the computing server cluster and all involved computing units. After the broadcast is completed, the instruction confirmation acknowledgment signal returned by each computing unit is received.

[0103] The instruction protocol encapsulation process bundles load migration instructions and their corresponding computational unit time windows into instruction frames conforming to the cluster's internal task scheduling protocol format. Each frame includes fields such as a frame header flag, instruction sequence number, computational unit location information, time window definition, and load migration instruction content. The broadcast delivery method uses the multicast communication channel of the cluster management network to simultaneously push the instruction set to the scheduling execution node and all computational units involved in the matching allocation scheme. Upon receiving the instruction, each computational unit verifies its integrity and returns an acknowledgment signal. The scheduling management node aggregates all acknowledgment signals; if any computational unit fails to return acknowledgment or returns a rejection, an exception handling process is triggered.

[0104] Step S160: After the computing server cluster executes the scheduling operation corresponding to the power balancing scheduling instruction set, it continuously collects the actual power consumption sensing signal sequence and the actual task load signal sequence of each computing unit during the task execution cycle, and encapsulates the actual power consumption sensing signal sequence and the actual task load signal sequence into an actual synchronization monitoring signal array.

[0105] In this embodiment, after the power balancing scheduling instruction set is issued, each computing unit begins receiving load data and initiating task execution according to the scheduling arrangement. During the task execution cycle, the power consumption sensing probes and task load monitoring agents of each computing unit continue to work continuously at a fixed collection granularity, collecting power consumption and load data under actual operating conditions. The actual collected data is processed according to the same timing synchronization verification, signal frame segmentation, frame pairing, and array encapsulation process as steps S110 to S115 to generate an actual synchronization monitoring signal array. This actual synchronization monitoring signal array records the actual power consumption response of the computing unit after the scheduling decision is actually executed, serving as the factual basis for evaluating prediction accuracy and making feedback corrections.

[0106] Step S161: Perform power ripple morphology analysis on the actual synchronous monitoring signal array, extract the actual power ripple peak-valley timing features of each arithmetic unit, and perform segment-by-segment deviation comparison between the actual power ripple peak-valley timing features and the predicted power drift surface to generate the timing deviation distribution between the actual power curve and the predicted power drift surface.

[0107] The power ripple morphology analysis process is identical to steps S120 to S127 for the actual synchronous monitoring signal array, extracting the peak-valley timing characteristics of the actual power ripple. The actual power consumption curve is directly plotted from the actual power consumption sensing signal sequence. The actual power consumption curve and the predicted power consumption drift surface are aligned along the time axis, and the difference between the actual power consumption value and the predicted drift amplitude value is calculated at each predicted timing point. All differences constitute the timing deviation distribution.

[0108] Step S162: Extract the timing points in the timing deviation distribution where the deviation amplitude exceeds the preset deviation tolerance range, and mark the timing points as power prediction deviation anomaly points. Backtrack the computing load type identifier carried by the computing unit and the corresponding load switching event sequence when the power prediction deviation anomaly point appears.

[0109] The deviation tolerance range is a strip-shaped area centered at zero deviation, extending upwards and downwards by a fixed amplitude. When the deviation value at a certain time point exceeds this strip-shaped area, that point is marked as a prediction deviation anomaly. The time of occurrence of this anomaly is traced back from the synchronous monitoring signal array, and the identifier of the computational load type that the computing unit was running at that time is extracted. Simultaneously, the load switching events before and after that time are extracted from the load switching event sequence generated in step S126 to determine what type of load change caused the prediction deviation to exceed the tolerance range.

[0110] Step S163: Based on the occurrence sequence of the power consumption prediction deviation anomaly points and the corresponding computing load type identifier, construct a prediction deviation event record containing deviation amplitude, deviation timing and load type association marker, and append the prediction deviation event record to the prediction deviation historical event database.

[0111] Each prediction deviation event record is a triplet structure containing the deviation magnitude, the absolute timestamp of the anomaly, and the associated load type identifier. The prediction deviation history event repository is persistently stored as a log file in a distributed file system, with a new record appended each time a new event occurs. Over time, this history event repository gradually forms a statistical dataset of prediction deviations covering different load types and time periods.

[0112] Step S164: Perform cluster analysis on the accumulated prediction deviation event records in the prediction deviation historical event database to identify deviation patterns that frequently appear under a specified load type identifier, and encapsulate the statistical features of the deviation patterns into load type deviation fingerprint features.

[0113] A grouped clustering analysis based on load type identifiers is employed. Predicted deviation event records are grouped according to load type identifiers, with the deviation amplitude values ​​within the same group forming a deviation amplitude sample set. Statistical distribution fitting is performed on this sample set, and the mean, standard deviation, and skewness of the deviation amplitudes are extracted as statistical features. If the mean deviation under a certain load type identifier significantly deviates from zero and the standard deviation is small, it indicates that the power consumption drift trend inference network has a systematic directional bias in predicting this type of load; the direction of the bias is indicated by the sign of the mean deviation. The load type deviation fingerprint feature is a fixed-dimensional vector containing the load type identifier, mean deviation, standard deviation deviation, and sample count.

[0114] Step S165: Feed back the load type deviation fingerprint feature to the trend extrapolation expansion layer of the power consumption drift trend inference network, and perform incremental adjustment processing on the trend extrapolation parameters in the trend extrapolation expansion layer so that the adjusted trend extrapolation parameters can reduce the prediction deviation of the computing load corresponding to the load type identifier.

[0115] The trend extrapolation expansion layer is a fully connected prediction network that transforms the regularized drift trend vector into discrete drift amplitude values ​​at future time points. The mean deviation in the load type bias fingerprint indicates the direction and magnitude of the systematic bias. The incremental adjustment process adds a bias correction vector indexed by load type to the output of the trend extrapolation expansion layer. When the predicted load belongs to a load type of a known fingerprint, the mean deviation corresponding to that load type is superimposed on the network output as a correction amount, making the corrected predicted drift amplitude value closer to the actual value.

[0116] Step S170: During the process of the power drift trend inference network inferring the power drift direction of the adjacent time segment of the ripple load associated feature vector, the hidden state feature vector sequence output by the drift state space modeling layer in the power drift trend inference network is recorded simultaneously.

[0117] In this embodiment, the drift state space modeling layer is the core intermediate layer of the power drift trend inference network. Its output hidden state feature vector is a comprehensive abstract encoding of the influence of power ripple pattern and load switching situation on power drift within the current time segment. This vector sequence records the continuous evolution trajectory of the power drift state in the hidden space. When the load characteristics of the computing unit remain stable, the distribution characteristics of the hidden state feature vector also remain relatively stable; when the computing unit experiences task switching, resource contention, or changes in the external environment, the distribution characteristics of the hidden state feature vector will shift accordingly. Therefore, by monitoring the distribution changes of the hidden state feature vector sequence, the non-stationary transition of the computing unit's power drift mode can be detected.

[0118] Step S171: Perform time-series segmentation processing on the hidden state feature vector sequence, dividing the hidden state feature vector sequence into multiple non-overlapping hidden state segments, and perform statistical feature extraction processing on each hidden state segment to extract the hidden state mean vector and hidden state covariance features of each hidden state segment.

[0119] In this embodiment, the hidden state feature vector sequence is divided into several continuous but non-overlapping hidden state segments along the time axis with a fixed segment length. Each segment contains the same number of continuous time-step hidden state feature vectors, reflecting the local distribution of power drift hidden states within a short time period. Two statistical features are calculated for each segment: the hidden state mean vector is the dimensional arithmetic mean of all hidden state feature vectors within the segment, reflecting the average level of power drift states within that time period; the hidden state covariance feature is the covariance matrix formed by the covariances between the dimensions within the segment, reflecting the interconnected changes and uncertainty range of the power drift states between the dimensions within that time period.

[0120] Step S172: Perform difference measurement on the hidden state mean vectors of adjacent hidden state segments, calculate the vector offset distance between the hidden state mean vectors of adjacent hidden state segments, and simultaneously perform difference measurement on the hidden state covariance features of adjacent hidden state segments, calculate the distribution offset distance between the hidden state covariance features of adjacent hidden state segments.

[0121] In this embodiment, the vector offset distance is taken as the L2 norm of the difference between the hidden state mean vectors of two adjacent hidden state segments, i.e., the square root of the sum of the squares of the differences in each dimension. This distance measures the magnitude of the drift state's shift in the latent space from the previous time period to the overall operating point of the next time period. A larger vector offset distance indicates a more significant directional change in the average behavior of power drift. The distribution offset distance is taken as the Frobenius norm of the difference between two covariance matrices, i.e., the square root of the sum of the squares of the differences in each matrix element. This distance measures the degree of change in the fluctuation pattern of the drift state in the latent space. A larger distribution offset distance indicates a greater change in the uncertainty structure of power drift or the coupling relationship between its dimensions.

[0122] Step S173: The vector offset distance and the distribution offset distance are fused to generate a hidden state drift phase transition metric parameter between adjacent hidden state segments. The hidden state drift phase transition metric parameter indicates the severity of the sudden change in the hidden state distribution characteristics between adjacent time segments.

[0123] In this embodiment, the vector offset distance and distribution offset distance describe the changes in the hidden state distribution from two dimensions: mean shift and covariance deformation, respectively. The fusion process first normalizes the two distances, using their respective long-term historical statistical mean as the normalization benchmark. The two normalized distances are then weighted and summed to generate a hidden state drift phase transition metric parameter. This parameter combines the contributions of mean shift and covariance variation: the phase transition metric parameter reaches its maximum value when the hidden states of adjacent segments undergo both overall horizontal migration and fluctuating structural reorganization; when both remain stable, the parameter approaches zero. A higher value of this parameter indicates a more severe non-stationary phase transition in the power consumption drift mode.

[0124] Step S174: Accumulate and calculate the hidden state drift phase transition measurement parameters between multiple consecutive hidden state segments, and construct a hidden state drift phase transition sequence from the multiple consecutive hidden state drift phase transition measurement parameters. In the hidden state drift phase transition sequence, identify the phase transition point where the hidden state drift phase transition measurement parameters exceed the phase transition trigger threshold.

[0125] In this embodiment, the phase transition metric parameters calculated from multiple consecutive adjacent hidden state segments are arranged in their corresponding time order to form a hidden state drift phase transition sequence. This sequence depicts the evolution curve of the power consumption drift mode as it stabilizes or abruptly changes over time. The phase transition trigger threshold is set statistically as the long-term mean of the sequence plus a certain multiple of the long-term standard deviation. When the phase transition metric parameter at a certain point in the sequence exceeds this threshold, that position is considered a phase transition point. The time corresponding to the phase transition point is the critical time node where a significant non-stationary transition of the power consumption drift mode occurs.

[0126] Step S175: Obtain the timing position corresponding to the phase transition point, extract the power consumption sensing signal frame and task load signal frame within a preset window range before and after the timing position from the synchronous monitoring signal array according to the timing position, and perform load mutation event correlation analysis on the extracted signal frames to determine the type of load mutation event that causes the hidden state phase transition.

[0127] In this embodiment, after locating the absolute time coordinates corresponding to the phase transition point, power consumption sensing signal frames and task load signal frames within a preset time window width before and after the phase transition point are extracted from the synchronous monitoring signal array. The load change event correlation analysis first extracts the load switching event sequence within this window according to steps S125 and S126, and then selects the event with the largest load switching amplitude as a phase transition correlation candidate event. If the difference between the occurrence timestamp of this candidate event and the timestamp of the phase transition point is less than a preset correlation time tolerance, then the load switching event is determined to be the main cause of the latent state phase transition, and its pre-switching state type and post-switching state type are extracted as the load change event type for phase transition correlation.

[0128] Step S176: Encode the load mutation event type and the hidden state drift phase transition metric parameter corresponding to the phase transition point into a drift network phase transition trigger signal, and send the drift network phase transition trigger signal to the scheduling management node of the computing server cluster to trigger the scheduling management node to perform scheduling strategy pre-adjustment processing for the load mutation event type.

[0129] The phase transition trigger signal encoding in the drift network includes three fields: the timestamp of the phase transition, the value of the phase transition metric parameter, and the load mutation event type. Upon receiving the signal, the scheduling management node queries the predefined scheduling pre-adjustment rules corresponding to the load mutation event type. The rule base defines different pre-adjustment strategies for different load mutation types, such as increasing the warm-up preparation of idle computing units, adjusting cooling fan speeds, or postponing the issuance of non-urgent tasks. The purpose of pre-adjustment processing is to take proactive mitigation measures before the power consumption anomaly caused by the phase transition amplifies, reducing the impact of the phase transition on the overall power consumption stability of the cluster.

[0130] For example, the method may further include: step S180: acquiring temperature distribution data of the computing unit area collected by the temperature sensing array deployed in the cooling and heat dissipation system of the computing server cluster, arranging the temperature distribution data of the computing unit area according to the acquisition time sequence, and generating regional temperature time sequence change field data.

[0131] In this embodiment, each rack of the computing server cluster is equipped with a multi-channel digital temperature sensor array arranged in layers along the rack height. The spacing between adjacent temperature sensors in the horizontal and vertical directions determines the spatial sampling resolution. Each temperature sensor aggregates temperature data to the rack management controller via a single bus or internal integrated circuit bus protocol. At each acquisition moment, all sensor readings are arranged according to the spatial location of the sensors to form a two-dimensional temperature distribution matrix. Stacking the two-dimensional temperature distribution matrices from multiple consecutive acquisition moments in chronological order creates a three-dimensional regional temperature temporal variation field, with its three dimensions corresponding to the time axis, rack height, and rack depth, respectively. This three-dimensional data field completely records the spatiotemporal distribution evolution of the thermal environment surrounding the computing unit.

[0132] Step S181: Perform spatial temperature gradient calculation processing on the temporal temperature change field data of the region, extract the temperature gradient direction vector and temperature gradient magnitude of each spatial grid point in the computing unit region at each acquisition time, and generate a temperature gradient vector field sequence.

[0133] The temperature at each spatial grid point within the computational unit region forms a two-dimensional gradient vector with respect to the partial derivatives of its temperature with respect to its adjacent grid points in the height and depth directions. The gradient vector points in the direction of the fastest temperature increase at that grid point, and its magnitude represents the temperature increase per unit distance in that direction. For each acquisition moment in the temporal temperature variation field data, the gradient vector is calculated for each grid point of the two-dimensional temperature distribution matrix. The arrangement of the gradient vectors of all grid points at that moment constitutes the temperature gradient vector field. The temperature gradient vector fields from all acquisition moments are arranged chronologically to form a temperature gradient vector field sequence, which describes the changes in the direction and intensity of heat transfer in space over time.

[0134] Step S182: Perform spatiotemporal alignment mapping processing on the temperature gradient vector field sequence and the predicted power consumption drift surface to establish the spatiotemporal correspondence between the drift amplitude value of each predicted time point on the predicted power consumption drift surface and the temperature gradient direction vector and temperature gradient amplitude of the corresponding spatial grid point.

[0135] The time axis of the predicted power consumption drift surface covers the next scheduling cycle, with each predicted timing point corresponding to an absolute timestamp. Each time frame of the temperature gradient vector field sequence also has a corresponding acquisition timestamp. The spatiotemporal alignment mapping uses the timestamp as the alignment key to establish a correspondence between each predicted timing point on the predicted power consumption drift surface and the vector field frame in the temperature gradient vector field sequence with the closest timestamp. In terms of spatial correspondence, each computing unit occupies a fixed physical space position in the rack, binding its predicted timing point drift amplitude value to the temperature gradient vector and amplitude on the corresponding rack space grid point. Through this spatiotemporal alignment, a quantitative correspondence is established between the predicted power consumption change and the local temperature change it may cause, both spatially and temporally.

[0136] Step S183: Based on the spatiotemporal correspondence, call the pre-built power consumption temperature rise coupled response model, perform temperature rise response prediction processing on each predicted time point on the predicted power consumption drift surface, and generate the predicted temperature increment distribution surface corresponding to the predicted power consumption drift surface.

[0137] The power consumption and temperature rise coupled response model describes the thermodynamic response relationship between the increase in power consumption of the computing unit and the increase in ambient air temperature. It is based on the heat conduction and convection equations: the increase in power consumption first translates into a rise in the junction temperature of the computing unit chip. The chip heat is conducted to the heat sink through the thermal resistance of the package. The heat sink then transfers heat to the surrounding air through convection. As the ambient air temperature rises, natural or forced convection occurs driven by the temperature gradient. This model abstracts the above heat transfer path, using the predicted drift amplitude as input to predict the temperature increment of the space surrounding the computing unit at the corresponding time point. The model is called for each predicted time point on the predicted power consumption drift surface to generate the predicted temperature increment value for each grid point in space at that time point. The predicted temperature increment values ​​for all predicted time points and all spatial grid points are arranged sequentially along the time axis to obtain the predicted temperature increment distribution surface. This surface uses time and space as dual independent variables and outputs the predicted temperature increment value.

[0138] Step S184: Match the predicted temperature increment distribution surface with the heat dissipation capacity parameters of the cooling system, and identify overheating risk areas in the predicted temperature increment distribution surface where the predicted temperature increment exceeds the maximum heat dissipation temperature increment limited by the heat dissipation capacity parameters.

[0139] The cooling system is equipped with a rated heat dissipation capacity, which corresponds to a maximum heat dissipation temperature increment value under specific fan speeds and cooling medium temperatures. This maximum heat dissipation temperature increment value represents the system's maximum ability to maintain the temperature rise around the computing unit below this value. The temperature increment value at each spatial grid point in the predicted temperature increment distribution surface at each prediction time is compared point-by-point with this maximum heat dissipation temperature increment value. If the temperature increment value at a spatial grid point exceeds the maximum heat dissipation temperature increment value at a certain prediction time, then that location is considered to have an overheating risk at that time. The spatially connected regions of all overheating risk grid points constitute an overheating risk region. The spatial extent of the overheating risk region and its evolution over time indicate the computing units and their spatial locations that may lead to localized overheating during the execution of the scheduling scheme.

[0140] Step S185: For the computing unit corresponding to the overheating risk area, the occupancy time window of the computing unit is reduced and adjusted in the computing unit power consumption matching and allocation scheme, and the power consumption balancing scheduling instruction set is regenerated based on the reduced occupancy time window.

[0141] For each computing unit located within an overheating risk zone, its originally allocated occupancy time window is retrieved from the computing unit power consumption matching allocation scheme. Based on the start and end times when the temperature increment within the overheating risk zone exceeds the maximum heat dissipation temperature increment, the time interval to be reduced is determined. The reduction adjustment process prioritizes delaying the start time or advancing the end time of the occupancy time window, ensuring that the computing unit does not execute high-power tasks within the overheating risk time interval. If the entire occupancy time window falls within the overheating risk zone and cannot be avoided through local adjustments, the computing unit is removed from the allocation scheme, and its computing load is reallocated to candidate computing units in other non-overheating risk zones. Based on the adjusted occupancy time window and the reallocated pairing mapping relationship, the power balancing scheduling instruction set is regenerated according to steps S150 to S158.

[0142] Step S186: Output the regenerated set of power balancing scheduling instructions to the scheduling execution node of the computing server cluster to update the scheduling arrangement of computing units involving overheating risk areas in the original scheduling plan.

[0143] After receiving the updated power balancing scheduling instruction set, the scheduling execution node replaces the parts of the original, yet-to-be-executed scheduling plans involving overheated computing units with the updated instructions. For tasks already in execution and uninterruptible, the scheduling execution node does not allocate new load after the task is completed, until the overheating risk warning is lifted. The aforementioned scheduling correction mechanism based on temperature prediction feedback extends power scheduling from a single electrical dimension to a thermoelectric coupling dimension, while simultaneously considering power balancing and thermal safety constraints.

[0144] Step S190: Add a power ripple resonance sensing bypass to the drift state space modeling layer inside the power drift trend inference network. Extract the power ripple peak-valley time-series features from the ripple load associated feature vector through the power ripple resonance sensing bypass, and perform frequency domain transformation on the power ripple peak-valley time-series features to generate a power ripple spectrum representation vector.

[0145] In this embodiment, the power ripple of the computing unit may contain multiple adjacent frequency components. When the computing unit operates multiple periodic loads simultaneously or the load switching frequency is close to the power supply regulation frequency, a resonance effect may occur between power ripples of different frequencies, resulting in the actual power fluctuation amplitude being much greater than the simple superposition of a single frequency ripple. Step S190 adds a power ripple resonance sensing bypass to the drift state space modeling layer to specifically detect resonance risk and encode the resonance enhancement effect into the hidden state feature vector. This bypass first extracts the peak-valley time interval distribution histogram and peak-valley feature descriptor list generated in steps S124 and S123 from the ripple load associated feature vector to reconstruct the peak-valley time sequence characteristics of the power ripple. Then, a fast Fourier transform is used to transform the peak-valley time sequence characteristics from the time domain to the frequency domain to obtain the power ripple spectrum representation vector. This frequency domain vector describes the energy distribution of the power ripple at different frequencies.

[0146] Step S191: Perform main frequency component separation processing on the power consumption ripple spectrum characterization vector, extract the main ripple frequency component with the largest amplitude ratio and the secondary ripple frequency component adjacent to the main ripple frequency component in the power consumption ripple spectrum characterization vector, and combine the main ripple frequency component and the secondary ripple frequency component into a ripple resonance frequency pair.

[0147] The frequency components are sorted in descending order of frequency amplitude on the spectral representation vector, and the frequency with the largest amplitude is the primary ripple frequency. After removing the primary ripple frequency, the frequency component with the closest frequency and the largest amplitude among the remaining frequency components is selected as the secondary ripple frequency. If the difference between the primary and secondary frequencies divided by the primary frequency is less than the resonance frequency difference threshold, then the two constitute a pair of ripple resonance frequencies, which poses a risk of resonance amplification.

[0148] Step S192: Input the ripple resonance frequency pair into the frequency beat operation unit in the power consumption ripple resonance sensing bypass, perform frequency beat operation on the main ripple frequency component and the secondary ripple frequency component to generate ripple beat frequency parameters, and construct a ripple beat oscillation waveform based on the ripple beat frequency parameters.

[0149] Frequency beat modulation involves superimposing two signals of different frequencies, resulting in a synthesized signal with an envelope modulation phenomenon where the absolute value of the frequency difference between the two signals is used as the frequency. The ripple beat frequency parameter is taken as the absolute value of the difference between the primary and secondary ripple frequencies. The ripple beat oscillation waveform is constructed using this beat frequency as the frequency and 1 as the amplitude, forming a cosine oscillation waveform sequence whose time range is consistent with the prediction scheduling period.

[0150] Step S193: Obtain the hidden state feature vector corresponding to the current time segment from the drift state space modeling layer, map the hidden state feature vector to the power recovery force representation space of the power drift trend inference network, and generate a power recovery force coefficient vector that reflects the power adjustment inertia characteristics inside the computing unit.

[0151] The hidden state feature vector contains information about the current power consumption state and load status of the computing unit. A fully connected mapping network projects the hidden state feature vector onto the power recovery force representation space. Each dimension of this space corresponds to the power consumption regulation inertia characteristics of the computing unit at different time scales, such as the recovery time constant of the supply voltage and the response delay of the operation frequency adjustment. The components of the power recovery force coefficient vector reflect the damping characteristics of the computing unit recovering to a steady state after being subjected to power consumption disturbances.

[0152] Step S194: Input the ripple beat oscillation waveform and the power recovery coefficient vector into the resonance response simulator in the power ripple resonance sensing bypass, and have the resonance response simulator simulate the resonance amplification response process of the ripple beat oscillation waveform under the constraint of the power recovery coefficient vector to generate the ripple resonance amplification intensity envelope curve.

[0153] The resonant response simulator is a numerical simulation micro-model of a second-order damped oscillation system, whose damping coefficient is determined by the reciprocal of the power consumption restoring force coefficient vector. Using the ripple beat oscillation waveform as the external driving force input, the change in the system's response amplitude over time under this driving force is numerically solved, yielding the ripple resonant amplification intensity envelope curve. When the driving force frequency is close to the system's natural frequency, the response amplitude is amplified; when it is much greater than or much less than the natural frequency, the amplification effect weakens.

[0154] Step S195: The ripple resonance amplification intensity envelope curve and the original power consumption restoring force coefficient vector are superimposed and modulated point by point to generate a modified power consumption restoring force coefficient vector carrying resonance enhancement characteristics, and the modified power consumption restoring force coefficient vector is injected back into the drift state transition parameter matrix of the drift state space modeling layer.

[0155] The values ​​of the resonant amplification intensity envelope curve at each time point are multiplied by the corresponding components in the original power consumption restoring force coefficient vector to obtain the correction coefficient enhanced by the resonant effect. The back-injection process embeds this correction coefficient into the corresponding element of the drift state transition parameter matrix, so that when the state space is projected in subsequent time steps, the transition amplitude of the drift state is affected by the resonant enhancement effect, resulting in a larger predicted power consumption fluctuation.

[0156] Step S196: When the drift state space modeling layer uses the back-injection corrected drift state transition parameter matrix to perform state space projection processing on the deep time series representation vector of subsequent time steps, the corrected drift state transition parameter matrix encodes the ripple resonance amplification effect into the newly generated hidden state feature vector, so that the newly generated hidden state feature vector carries the characterization of the enhanced influence of power ripple resonance on power drift.

[0157] In this embodiment, the core operation of the drift state space modeling layer is a state transition process involving a linear transformation and nonlinear activation. Before the addition of the resonance-aware bypass, it uses the original drift state transition parameter matrix to project the input deep temporal representation vector, generating a hidden state feature vector. When step S195 injects the modified power recovery coefficient vector carrying resonance enhancement characteristics back into the drift state transition parameter matrix, the weights of the power recovery-related dimensions in the parameter matrix are enhanced. This means that when the deep temporal representation vector is input into this layer at subsequent times, the feature components involving power periodic fluctuations and recovery response will obtain greater mapping weights during the state space projection process. The mapped hidden state feature vector has higher activation values ​​in the corresponding dimensions, and these activation values ​​encode the information that "the current computing unit has a ripple resonance risk, and its actual power drift amplitude may be greater than the conventional prediction value." Through the above-mentioned parameter matrix-level correction, the resonance-aware bypass integrates the resonance enhancement effect into the hidden state expression of all subsequent time steps without changing the overall network structure, realizing the organic integration of resonance risk perception and drift inference main process.

[0158] Step S197: The time difference prediction layer performs time difference operation processing based on the hidden state feature vector carrying the power ripple resonance enhancement effect to generate a predicted power drift surface corresponding to the next scheduling cycle and containing the ripple resonance induced drift effect.

[0159] In this embodiment, the hidden state feature vector carrying the resonance enhancement effect characterization propagates forward along subsequent layers of the network. After receiving the hidden state feature vector, the temporal difference prediction layer calculates the difference offset vector between the hidden state of the current time segment and the hidden state of the historical time segment according to the same mechanism as in step S133. Since the hidden state feature vector of the input temporal difference prediction layer has been resonantly enhanced in the power fluctuation related dimensions, the component amplitudes of the calculated difference offset vector in these dimensions are also correspondingly increased. This difference offset vector serves as the original gradient of the power drift direction. After being aggregated by the dilated recursive window of the gradient recursive aggregation layer, denoised and regularized by the noise adversarial regularization layer, fully connected mapping of the trend extrapolation expansion layer, and spline interpolation of the surface reconstruction layer, the generated predicted power drift surface exhibits a larger amplitude fluctuation prediction on the time axis than in the case without resonance perception. In the time interval where resonance occurs, the predicted drift amplitude shows a significant amplification peak; after the resonance effect subsides, the predicted drift amplitude returns to the normal level. The aforementioned predicted power consumption drift surface with resonance-induced drift effect can more realistically reflect the dynamic changes in power consumption of the computing unit when there is a risk of ripple resonance.

[0160] Step S200: Obtain historical scheduling record data of the computing server cluster, extract the pre-scheduling power drift surface snapshot and the post-scheduling actual power drift surface snapshot corresponding to each historical scheduling event from the historical scheduling record data, and construct a set of historical scheduling power drift surface snapshot pairs.

[0161] In this embodiment, the scheduling management node of the computing server cluster persistently stores complete scheduling context data as historical scheduling records each time it performs a load scheduling operation. Each historical scheduling record includes a scheduling event identifier, a timestamp of the scheduling execution, a list of computing units participating in the scheduling, an identifier of the scheduled computing load, and a snapshot of the power drift surface before scheduling on which the scheduling decision is based. After scheduling is completed and a stable observation period has passed, the actual operating data of each computing unit is re-collected and a snapshot of the actual power drift surface after scheduling is generated through a power drift trend extrapolation network. This snapshot is then added to the same historical scheduling record. The snapshots before and after scheduling for the same scheduling event are paired. The previous snapshot records the expectation of future power drift at the time of scheduling decision, and the subsequent snapshot records the actual power drift observed after the actual execution of scheduling, forming a set of historical scheduling power drift surface snapshot pairs.

[0162] Step S201: Perform scheduling intervention effect decoupling processing on each pair of snapshots in the set of historical scheduling power drift surface snapshots, and subtract the power drift surface snapshot before scheduling from the actual power drift surface snapshot after scheduling to obtain the net scheduling utility drift surface for each historical scheduling event. The net scheduling utility drift surface describes the amount of independent power drift change brought about by the scheduling operation.

[0163] In this embodiment, the change in the actual power consumption drift surface after scheduling consists of two superimposed parts: the continuation of the inherent power consumption drift trend of the computing unit itself, and the additional power consumption drift change caused by the scheduling operation as an external intervention. The decoupling processing of the scheduling intervention effect is based on the assumption that, in the short period before and after scheduling execution, without external intervention, the inherent power consumption drift trend of the computing unit remains continuous and changes slowly. Under this assumption, the snapshot of the power consumption drift surface before scheduling is approximated as a prediction of the continuation of the drift trend under no intervention. The drift amplitude value at each prediction time point of the actual power consumption drift surface snapshot after scheduling is subtracted from the drift amplitude value at the corresponding prediction time point of the snapshot before scheduling, and the difference constitutes the net scheduling utility drift surface. Positive values ​​on the surface indicate that the scheduling operation leads to an increase in the power consumption drift amplitude, while negative values ​​indicate that the scheduling operation successfully smooths out the power consumption drift amplitude.

[0164] Step S202: Perform causal association pairing between the scheduling decision parameter vector corresponding to each historical scheduling event and the net scheduling utility drift surface to construct a scheduling utility causal training sample set with the scheduling decision parameter vector as input and the net scheduling utility drift surface as output.

[0165] In this embodiment, each historical scheduling event is recorded by the scheduling management node as a set of scheduling decision parameter vectors during scheduling. These vectors include fields such as the power consumption drift coupling metric of the target computing unit, the one-hot encoding of the type identifier of the load to be scheduled, the normalized peak value of the expected power consumption demand of the load, and the relative position encoding of the scheduling execution time within the scheduling cycle. These scheduling decision parameters are the causal variables that cause the corresponding net scheduling utility drift surface. The scheduling decision parameter vector of each historical scheduling event is used as the input training sample, and the net scheduling utility drift surface corresponding to the same event is used as the output training label to form a causal pairing training sample. The causal pairing samples corresponding to all historical scheduling events are summarized to form a scheduling utility causal training sample set.

[0166] Step S203: Train a scheduling utility counterfactual inference network using the scheduling utility causal training sample set. The front end of the scheduling utility counterfactual inference network is a scheduling parameter encoding layer, which maps the input scheduling decision parameter vector to scheduling intention latent variables. The tail end of the scheduling utility counterfactual inference network is a utility surface generation layer, which generates the corresponding hypothetical net utility drift surface based on the scheduling intention latent variables.

[0167] In this embodiment, the scheduling utility counterfactual inference network adopts an encoder-decoder architecture. The front-end scheduling parameter encoding layer consists of multiple fully connected layers stacked together, each layer containing linear weight matrix multiplication, bias vector addition, and nonlinear activation function transformation. The scheduling parameter encoding layer abstracts and fuses the input scheduling decision parameter vector layer by layer, compressing and mapping it into a low-dimensional scheduling intention latent variable, encoding the underlying logic behind the scheduling decision. The tail-end utility surface generation layer consists of multiple fully connected layers and dimension recombination operations. It receives the scheduling intention latent variable and maps it to a hypothetical net utility drift surface with the same dimensional structure as the net scheduling utility drift surface by expanding its dimensions layer by layer. During training, the mean square error between the hypothetical net utility drift surface and the net scheduling utility drift surface at each prediction time point is calculated. The mean error of all prediction time points is taken as the loss function, and all weight parameters of the two sub-networks are updated through gradient backpropagation.

[0168] Step S204: Obtain the candidate scheduling decision parameter vector corresponding to each candidate operation unit in the current candidate operation unit set, and input each candidate scheduling decision parameter vector into the trained scheduling utility counterfactual inference network to generate the hypothetical net utility drift surface of each candidate operation unit.

[0169] In this embodiment, for each candidate computing unit in the candidate computing unit set selected in step S147, a corresponding candidate scheduling decision parameter vector is constructed. The vector structure is completely consistent with the training samples, and each field is extracted and filled from the candidate computing unit attribute information and the computing load requirement information to be allocated. Each candidate scheduling decision parameter vector is input into the trained scheduling utility counterfactual inference network, and the network forward inference outputs the hypothetical net utility drift surface corresponding to the candidate computing unit. This hypothetical net utility drift surface simulates the power consumption drift change that may occur after scheduling if the candidate computing unit is selected as the scheduling target.

[0170] Step S205: The hypothetical net utility drift surface of each candidate operation unit is superimposed and extrapolated with the predicted power consumption drift surface to generate the counterfactual scheduling predicted power consumption drift surface corresponding to each candidate operation unit. The counterfactual scheduling predicted power consumption drift surface simulates the power consumption drift state that may occur after scheduling if the candidate operation unit is used as the scheduling target.

[0171] In this embodiment, the hypothetical net utility drift surface only reflects the power drift change introduced by the scheduling operation itself, while the predicted power drift surface reflects the power drift trend of the computing unit itself without scheduling intervention. The actual power drift state after scheduling is the result of the superposition of the two. The superposition extrapolation process adds the drift amplitude values ​​of the hypothetical net utility drift surface and the predicted power drift surface at each corresponding prediction time point to generate the predicted power drift surface after counterfactual scheduling. This surface provides a complete hypothetical extrapolation result for each candidate computing unit.

[0172] Step S206: Perform power stability evaluation processing on the predicted power drift surface after counterfactual scheduling for all candidate computing units, extract the global drift amplitude dispersion parameter and the duration parameter of the extreme drift peak for each predicted power drift surface after counterfactual scheduling, and fuse the global drift amplitude dispersion parameter and the duration parameter of the extreme drift peak into a scheduling effect score.

[0173] In this embodiment, the global drift amplitude dispersion parameter is taken as the standard deviation of the drift amplitude values ​​at all predicted time points on the predicted power drift surface after inverse fact scheduling. The smaller the standard deviation, the more stable the power drift state after scheduling. The duration parameter of extreme drift peaks is taken as the longest continuous prediction time step in which the drift amplitude value on the surface continuously exceeds the preset extreme drift threshold. The shorter the duration step, the more effectively the extreme power deviation is suppressed. The two parameters are normalized separately and then weighted and summed to form a scheduling effect score. The higher the score, the better the expected scheduling effect of the candidate scheme.

[0174] Step S207: Select candidate computing units whose scheduling effect scores meet the preset optimization conditions as the final determined power matching target computing units, and update the computing unit power matching allocation scheme according to the final determined power matching target computing units.

[0175] In this embodiment, the preset preferred condition is that the scheduling effect score is the highest among all candidate computing units, or the scheduling effect score exceeds the preferred threshold and the computing unit has the most idle resources. The candidate computing units that meet the conditions are determined as the final power consumption matching target computing units, and their addressing identifiers are used to update the corresponding entries in the computing unit power consumption matching allocation scheme generated in step S140.

[0176] Step S210: Perform full-time domain power consumption signal decomposition processing on the power consumption sensing signal sequence of each computing unit in the computing server cluster over a long period of time. Decompose the power consumption sensing signal sequence of each computing unit into a long-term power consumption trend component, a periodic power consumption fluctuation component, and a residual random power consumption disturbance component. The long-term power consumption trend component reflects the power consumption baseline drift caused by hardware aging of the computing unit. The periodic power consumption fluctuation component reflects the power consumption pattern fluctuation caused by the periodic switching of computing load type of the computing unit. The residual random power consumption disturbance component reflects the power consumption mutation caused by occasional abnormal events.

[0177] In this embodiment, the full-time-domain power consumption signal decomposition processing employs a seasonality and trend decomposition algorithm based on local weighted regression. The long-term power consumption trend component is extracted smoothly through local weighted regression over a large time window, with the window width much larger than the load switching cycle. This filters out short-term fluctuations and periodic components, retaining only monotonic or slowly varying long-scale changes. The periodic power consumption fluctuation component is determined by identifying the autocorrelation peak in the power consumption sequence to determine the dominant cycle length. The arithmetic mean of each cycle segment is then taken after phase alignment as the fluctuation template. This template is periodically repeated on the time axis to obtain the periodic power consumption fluctuation component. The residual random power consumption disturbance component is obtained by subtracting the long-term trend component and the periodic fluctuation component from the original power consumption sequence to obtain the residual.

[0178] Step S211: Perform trend pattern similarity clustering on the long-term power consumption trend components of each computing unit, group computing units with similar long-term power consumption trend component patterns into the same trend similarity group, and extract the mean curve of the long-term power consumption trend component of computing units in each trend similarity group within the same time span as the group trend baseline of the trend similarity group.

[0179] In this embodiment, the dynamic time warping distance is used as a similarity measure between two sequences to measure trend pattern similarity. The dynamic time warping distance is calculated for the time series of long-term power consumption trend components of each pair of computing units, constructing a distance matrix for all computing units. A hierarchical clustering algorithm is then used to group computing units with similar patterns. The group trend baseline is the point-by-point arithmetic mean of the trend component values ​​of all computing units within the group at the same time point.

[0180] Step S212: Perform fluctuation pattern clustering on the periodic power consumption fluctuation components of each computing unit, group computing units with similar fluctuation amplitude and fluctuation frequency of periodic power consumption fluctuation components into the same fluctuation similarity group, and extract the common fluctuation pattern dictionary of periodic power consumption fluctuation components of computing units in each fluctuation similarity group.

[0181] In this embodiment, the periodic power consumption fluctuation component of each computing unit is extracted using Fast Fourier Transform (FFT) to form a fluctuation feature vector, which consists of the dominant frequency component and its corresponding amplitude. K-means clustering is then performed based on the Euclidean distance between the fluctuation feature vectors to divide the computing units into different fluctuation similarity groups. For each fluctuation similarity group, fluctuation templates of all computing units within the group are collected, and clustered again according to waveform morphology. The cluster center is used as an entry in the common fluctuation pattern dictionary, and the probability distribution of this entry within the group is calculated.

[0182] Step S213: Perform perturbation sparse coding processing on the residual random power consumption perturbation components of each computing unit, and extract the sparse representation codebook of perturbation events in the residual random power consumption perturbation components of each computing unit and the perturbation triggering context features corresponding to each perturbation sparse representation codebook.

[0183] In this embodiment, the perturbation sparse coding employs an online dictionary learning algorithm. The residual random power consumption perturbation component is divided into fixed-length signal segments. An overcomplete set of basis functions is iteratively optimized using dictionary learning, such that each signal segment can be approximated by a sparse linear combination of a small number of basis functions in the set. The converged set of basis functions constitutes the sparse representation codebook. The perturbation triggering context features corresponding to each codebook are extracted from the window before and after the perturbation event in the original synchronous monitoring signal array, including the task load state features and power consumption ripple state features at the time of the perturbation.

[0184] Step S214: Perform multi-dimensional heterogeneous model fusion processing on the group trend baseline of the trend similarity group, the common fluctuation pattern dictionary of the fluctuation similarity group, and the perturbation sparse representation codebook of each computing unit to construct a personalized power consumption digital twin model for each computing unit. The personalized power consumption digital twin model can output the complete power consumption response simulation waveform of the computing unit when given any future computing load sequence input.

[0185] In this embodiment, the multidimensional heterogeneous model fusion integrates three hierarchical model components according to a superposition rule. Given a future computing load sequence as input, the digital twin model first matches the most similar fluctuation template in a common fluctuation pattern dictionary based on the load type for each time period in the load sequence, generating a time-series waveform of the periodic power consumption fluctuation component. Then, the group trend baseline is truncated along the time span and superimposed on the fluctuation component to form a baseline trend. Next, based on task switching events in the load sequence, matching perturbation triggering context features are retrieved, and corresponding perturbation patterns are selected from the sparse representation codebook and injected into the corresponding time point to generate a residual random power consumption perturbation component. Finally, the three components are superimposed point-by-point along the time axis to output the complete power consumption response simulation waveform of the computing unit under the future load sequence.

[0186] Step S215: Register the personalized power consumption digital twin model of each computing unit to the digital twin simulation engine of the computing server cluster. The digital twin simulation engine synchronously generates the corresponding virtual scheduling simulation scheme according to the power consumption matching and allocation scheme of the computing unit, and drives the personalized power consumption digital twin model of each computing unit to perform pre-schedule simulation in the virtual scheduling simulation scheme.

[0187] In this embodiment, the digital twin simulation engine, based on the digital twin models of each computing unit, drives the digital twin models of each computing unit to perform simulation step-by-step on the virtual time axis, according to the load migration order and the time window occupied by each computing unit as defined in the computing unit power consumption matching allocation scheme. At each time step, each digital twin model receives the virtual load allocated for that time step as input and outputs the simulated power consumption value for that time step. The power consumption output of all computing units at all simulation time steps constitutes the complete result of the pre-scheduled simulation.

[0188] Step S216: Collect the simulated power response waveforms output by the personalized power digital twin models of each computing unit during the pre-scheduled simulation and deduction process. Perform consistency verification and comparison processing between the simulated power response waveforms and the predicted power drift surface. When there is a deviation between the simulated power response waveforms and the predicted power drift surface that exceeds the preset synchronization deviation tolerance, perform prediction parameter compensation and fine-tuning processing on the prediction output of the power drift trend deduction network.

[0189] In this embodiment, the consistency check comparison calculates the difference point by point between the power consumption value at each predicted time point on the simulated power consumption response waveform of each computing unit and the drift amplitude value at the corresponding time point on the predicted power consumption drift surface, plus the current reference power consumption value. The absolute values ​​of the differences constitute a deviation sequence. If the mean or maximum value of the deviation sequence exceeds the preset synchronization deviation tolerance, it indicates that there is a systematic deviation between the prediction of the power consumption drift trend extrapolation network and the fine simulation results of the digital twin. The prediction parameter compensation fine-tuning process superimposes a bias correction amount at the output of the trend extrapolation layer. This correction amount is set according to the mean deviation, so that the subsequent prediction results converge to the simulation results of the digital twin.

[0190] The power consumption drift trend inference network and the scheduling utility counterfactual inference network involved in this embodiment both involve the specific implementation of artificial intelligence models. The following provides supplementary explanations on the architecture, technical implementation of each module, training data and process, and application methods of the two models.

[0191] The power drift trend inference network employs a seven-layer cascaded deep temporal neural network architecture, consisting of a temporal feature encoding layer, a drift state space modeling layer, a temporal difference prediction layer, a gradient recursive aggregation layer, a noise adversarial regularization layer, a trend extrapolation expansion layer, and a surface reconstruction layer stacked sequentially. The temporal feature encoding layer uses a bidirectional long short-term memory (LSTM) network, with each forward and backward layer containing several LSM units. Each LSM unit contains four components: a forget gate, an input gate, an output gate, and a cell state. The forget gate concatenates the previous time step's hidden state with the current input, performs a linear transformation, and applies a sigmoid activation function to output a forgetting coefficient between 0 and 1. The input gate generates input coefficients and candidate cell states through the same mechanism. The cell state is updated by multiplying the previous cell state by the forgetting coefficient and adding the candidate cell state by the input coefficient. The output gate controls how much information from the cell state after hyperbolic tangent activation is output to the current hidden state. The bidirectional LSM network concatenates the hidden states of the forward and backward layers at the same time step along the vector dimension to form the deep temporal representation vector. The drift state space modeling layer contains a trainable drift state transition parameter matrix. It maps the deep temporal representation vector to the drift hidden state space through linear projection, then weights and fuses it with the hidden state of the previous time step, and outputs the current step hidden state feature vector through a nonlinear activation function. The temporal difference prediction layer subtracts the current step hidden state from the previous step hidden state element by element to obtain the difference offset vector. The gradient recursive aggregation layer maintains an expanding recursive window. Each historical difference offset vector within the window is multiplied by an expansion weight that decays exponentially with the step size, and then summed to obtain the aggregated drift trend vector. The noise adversarial regularization layer contains a noise perturbation injection branch and a denoising autoencoder branch. The denoising autoencoder consists of two subnetworks: encoder dimensionality reduction and decoder dimensionality expansion. During training, it minimizes the reconstruction error between the denoised output and the unperturbed original signal. The trend extrapolation expansion layer is a multi-layer fully connected network that takes the aggregated drift trend vector as input and outputs the discrete drift amplitude values ​​at each prediction time point in the next scheduling cycle. The surface reconstruction layer uses cubic spline interpolation to perform continuous surface interpolation on the discrete values ​​to generate the predicted power drift surface. The entire power drift trend prediction network employs an end-to-end supervised training approach. Training data originates from the synchronous monitoring signal array collected during the historical operation of the computing server cluster, along with the actual labels of the power drift surfaces in the next scheduling cycle. The input is a sequence of ripple load-related feature vectors, and the output is the predicted power drift surface. The optimizer uses an adaptive moment estimation optimizer, with an initial learning rate set to a small positive real number. The batch size is determined based on the number of computing units in the cluster, and the number of training epochs is determined by an early stopping strategy on the validation set. The training loss function is the average of the mean squared errors at each prediction time point between the predicted and actual power drift surfaces.During the model inference stage, the ripple load associated feature vector, which is collected in real time and processed by steps S110 to S127, is directly input into the trained network. The predicted power consumption drift surface is obtained through forward propagation, without the need for additional preprocessing and postprocessing steps.

[0192] The counterfactual inference network for scheduling utility employs an encoder-decoder architecture. The front-end scheduling parameter encoding layer consists of several stacked fully connected layers. Each layer includes linear weight matrix multiplication, bias vector addition, and a modified linear unit activation function. The input is a scheduling decision parameter vector, which is compressed and mapped to low-dimensional scheduling intention latent variables through layer-by-layer nonlinear transformation. The tail-end utility surface generation layer consists of several fully connected layers and a dimension recombination operation. First, the dimensions of the scheduling intention latent variables are progressively expanded through the fully connected layers. Then, the dimension recombination operation reconstructs them into a three-dimensional data structure with the same dimensions as the net scheduling utility drift surface. Its training data comes from a causal training sample set for scheduling utility, constructed from historical scheduling records of the computing server cluster. This sample set contains thousands to tens of thousands of pairs of scheduling decision parameter vectors and net scheduling utility drift surfaces. The input is the scheduling decision parameter vector, and the output is the hypothetical net scheduling utility drift surface. The optimizer employs an adaptive moment estimation optimizer with an initial learning rate set to a small positive real number. The training loss function is the average mean square error of the prediction time points between the hypothetical net utility drift surface and the net scheduling utility drift surface. Training stops once the validation loss converges. During the model inference phase, the candidate scheduling decision parameter vectors of the candidate computational units are input into the trained network, and forward propagation yields the hypothetical net scheduling utility drift surface. This output is directly used for subsequent superposition and inference processing.

[0193] The personalized power consumption digital twin model is not a single neural network, but a hybrid model composed of statistical model components. The long-term power consumption trend component uses a locally weighted regression smoother, employing a weighted least squares method to fit a low-order polynomial to extract the trend within a given time window. The periodic power consumption fluctuation component uses a template extraction method based on autocorrelation period detection and period-aligned averaging. The residual random power consumption perturbation component uses an online dictionary learning algorithm, learning an overcomplete set of basis functions as the perturbation sparse representation codebook through alternating optimization of sparse coding and dictionary updates. These three components are independently constructed from historical power consumption sensing signal sequences, and then fused using superposition rules to build a complete personalized power consumption digital twin model. In model application, the digital twin model receives the future computational load sequence as input, generates fluctuation components by matching fluctuation templates according to the load type, superimposes the trend baseline as the trend component, and injects perturbation components based on perturbation event encoding triggered by load switching events. These three components are then superimposed point-by-point to output the complete power consumption response simulation waveform.

[0194] Based on the same inventive concept, please refer to Figure 2This diagram illustrates a schematic block diagram of a deep learning-based computing server power optimization system provided in an embodiment of this application. The system includes a central processing unit (CPU), a system memory comprising random access memory (RAM) and read-only memory (ROM), and a system bus connecting the system memory and the CPU. The deep learning-based computing server power optimization also includes a basic input / output system to facilitate information transfer between various devices within the computer, and a large-capacity storage device for storing the operating system, applications, and other program modules.

[0195] A basic input / output system includes a display for showing information and input devices such as a mouse and keyboard for user input. Both the display and the input devices are connected to the central processing unit via an input / output controller connected to the system bus. The basic input / output system may also include an input / output controller for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller also provides output to a display screen, printer, or other types of output devices.

[0196] A mass storage device is connected to the central processing unit via a mass storage controller connected to the system bus. The mass storage device and its associated computer-readable medium provide non-volatile storage for power optimization of the deep learning-based computing server. Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. According to various embodiments of this application, the power optimization of the deep learning-based computing server can also be implemented by connecting to a remote computer on a network, such as the Internet. That is, the power optimization of the deep learning-based computing server can be connected to a network via a network interface unit connected to the system bus, or it can be used to connect to other types of networks or remote computer systems.

[0197] The above are merely exemplary embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application shall be included within the protection scope of this application.

Claims

1. A method for optimizing the power consumption of a computing server based on deep learning, characterized in that, The method includes: Collect power consumption sensing signal sequences and task load signal sequences of computing units in a computing server cluster within a continuous time segment, and encapsulate the power consumption sensing signal sequences and the task load signal sequences into a synchronous monitoring signal array with a unified time sequence label; The synchronous monitoring signal array is subjected to power ripple morphology analysis processing to extract the power ripple peak-valley timing features of the computing unit and the load switching timing features of the computing load carried by the computing unit, and the power ripple peak-valley timing features and the load switching timing features are constructed into a ripple load correlation feature vector. The ripple load associated feature vector is input into a pre-constructed power drift trend inference network. The power drift trend inference network infers the power drift direction of adjacent time segments of the ripple load associated feature vector, and generates the predicted power drift surface of the computing unit in the next scheduling cycle. Based on the predicted power consumption drift surface, a set of candidate computing units whose power consumption drift coupling relationship meets the preset coupling screening conditions is selected from the computing unit resource pool of the computing server cluster. Then, power consumption matching processing is performed based on the power consumption demand profile of each computing unit in the candidate computing unit set and the computing load to be allocated, so as to obtain a computing unit power consumption matching allocation scheme. Based on the power consumption matching and allocation scheme of the computing unit, a set of power balancing scheduling instructions containing computing unit addressing identifiers and computing load migration timing is generated and output.

2. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The system collects power consumption sensing signal sequences and task load signal sequences of computing units in the computing server cluster within continuous time segments, and encapsulates these sequences into a synchronous monitoring signal array with a unified time signature, including: A power consumption sensing probe is deployed on each computing unit of the computing server cluster. The power consumption sensing probe acquires the instantaneous current fluctuation signal and instantaneous voltage fluctuation signal of the power supply line inside the computing unit according to a preset acquisition granularity. The instantaneous current fluctuation signal and the instantaneous voltage fluctuation signal are subjected to signal aliasing modulation processing to generate a power consumption sensing signal sequence that reflects the instantaneous power consumption value of the computing unit at each acquisition point. Each signal point in the power consumption sensing signal sequence carries the timing mark corresponding to the acquisition point. The type identifier of the computing task running on the computing unit and the number of computing cores occupied by the computing task at each sampling point are obtained. Based on the type identifier and the number of computing cores, the task load status characteristics of the computing unit at each sampling point are generated. The task load status characteristics at all sampling points are arranged according to the time sequence mark to generate a task load signal sequence that reflects the change of the task load of the computing unit over time. Each signal point in the task load signal sequence carries a time sequence mark that is consistent with the sampling point corresponding to the power consumption sensing signal sequence. The power consumption sensing signal sequence and the task load signal sequence are subjected to timing synchronization verification processing. For the power consumption sensing signal sequence, when the timing mark deviation exceeds the preset synchronization tolerance threshold, the power consumption sensing value of the deviation signal point is corrected by interpolation based on the timestamp to obtain the power consumption sensing signal sequence after timing deviation correction. For the task load signal sequence, when the timing mark deviation exceeds the preset synchronization tolerance threshold, the timing mark of the deviation signal point is aligned and corrected, and its task load state characteristics are corrected to be the same as the state of the nearest valid timing mark signal point to obtain the task load signal sequence after timing deviation correction. The signal frame is then segmented according to the same time base to obtain a power consumption sensing signal frame and a task load signal frame containing the same number of signal points and whose timing marks are aligned one by one. The power consumption sensing signal frame and the task load signal frame are paired by a signal frame pairing process. The power consumption sensing signal frame and the task load signal frame with the same timing mark are bound into a synchronous monitoring signal frame unit, and the synchronous monitoring signal frame unit is given a globally unique timing mark. All synchronous monitoring signal frame units are arranged in ascending order according to the global timing mark to form a synchronous monitoring signal array, and the synchronous monitoring signal array is stored in the timing signal buffer storage area for subsequent power consumption ripple mode analysis and processing.

3. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The process involves performing power ripple pattern analysis on the synchronous monitoring signal array to extract the peak-valley timing features of the power ripple of the computing unit and the load switching timing features of the computing load carried by the computing unit. These power ripple peak-valley timing features and the load switching timing features are then used to construct a ripple load correlation feature vector, including: Extract the power consumption sensing signal frame from each synchronous monitoring signal frame unit from the synchronous monitoring signal array, perform sliding window local extremum detection processing on the power consumption sensing signal frame, locate the local power consumption maximum point and local power consumption minimum point of the power consumption sensing signal frame within the sliding window range, the local power consumption maximum point constitutes the peak point set, the local power consumption minimum point constitutes the valley point set; For each local power consumption maximum point in the peak point set, neighborhood waveform analysis is performed to extract the rising edge duration and falling edge duration corresponding to each local power consumption maximum point. The peak sharpness parameter is calculated based on the ratio of the rising edge duration to the falling edge duration, and a peak feature descriptor containing the peak point occurrence sequence, peak amplitude, and the peak sharpness parameter is generated. For each local power minima in the valley point set, neighborhood waveform analysis is performed to extract the duration of the falling edge and the duration of the rising edge corresponding to each local power minima. The valley sharpness parameter is calculated based on the ratio of the duration of the falling edge to the duration of the rising edge, generating a valley feature descriptor that includes the valley point occurrence sequence, valley amplitude, and the valley sharpness parameter. The occurrence sequences of all peak points and all valley points are mixed and sorted in chronological order to form a unified peak-valley event timestamp sequence. Based on the peak and valley event timestamp sequence, the time interval between adjacent peak point timestamps and valley point timestamps is statistically processed to generate a peak-valley time interval distribution histogram, and the peak-valley time interval distribution histogram is used as part of the power consumption ripple peak-valley time sequence feature; at the same time, all peak feature descriptors and valley feature descriptors are arranged in order of their corresponding timestamps to form a peak-valley feature descriptor list, which is used as another part of the power consumption ripple peak-valley time sequence feature; The valley features are arranged in an alternating sequence according to their respective occurrence times to form a peak-valley alternating feature sequence. The time interval between adjacent peak points and valley points in the peak-valley alternating feature sequence is statistically processed to generate a peak-valley time interval distribution histogram. The peak-valley time interval distribution histogram is used as the power consumption ripple peak-valley time sequence feature. Extract the task load signal frame from each synchronous monitoring signal frame unit from the synchronous monitoring signal array, perform point-by-point differential comparison processing on the task load state characteristics between adjacent signal points in the task load signal frame, detect the signal point position where the task load state characteristics change, and mark the signal point position as a load switching event point. Extract the task load state characteristics before and after the change corresponding to each load switching event point, calculate the difference in the number of computing cores occupied between the before and after states, generate load switching features including load switching sequence, before state type, after state type and the difference, arrange all load switching features according to load switching sequence, construct a load switching event sequence that reflects the switching density and switching amplitude distribution characteristics of computing load in the time dimension, and use the load switching event sequence as the load switching sequence feature; A pre-built temporal feature fusion mapper is invoked to fuse and encode the peak-valley temporal interval distribution histogram and the peak-valley feature descriptor list, mapping them into a fixed-length ripple feature vector; the load switching event sequence is mapped into a fixed-length load feature vector; and the ripple feature vector and the load feature vector are concatenated to generate the ripple load-related feature vector.

4. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The step of inputting the ripple load-related feature vector into a pre-constructed power drift trend inference network, and having the power drift trend inference network infer the power drift direction of adjacent time segments of the ripple load-related feature vector to generate the predicted power drift surface of the computing unit in the next scheduling cycle includes: The ripple load associated feature vector is input into the time-series feature encoding layer of the power consumption drift trend inference network. The time-series feature encoding layer performs time-series dependency encoding on the ripple load associated feature vector to generate a deep time-series representation vector carrying contextual power consumption time-series information. The deep time series representation vector is input into the drift state space modeling layer of the power consumption drift trend inference network. The drift state space modeling layer uses the internal drift state transition parameter matrix to perform state space projection processing on the deep time series representation vector to generate hidden state feature vectors in the drift state space. The hidden state feature vectors encode the influence mode of power consumption ripple shape and load switching status on power consumption drift in the current time segment. The hidden state feature vector is input into the time difference prediction layer of the power consumption drift trend inference network. The time difference prediction layer performs time difference operation on the hidden state feature vector to calculate the difference offset vector between the hidden state of the current time segment and the hidden state of the historical time segment. The difference offset vector is used as the original gradient of the power consumption drift direction. The original gradient of the power drift direction is input into the gradient recursive aggregation layer of the power drift trend inference network. The gradient recursive aggregation layer uses an expanding recursive window to recursively aggregate the differential offset vectors corresponding to multiple adjacent time segments to generate an aggregated drift trend vector. The aggregated drift trend vector describes the cumulative drift direction and drift rate of the power state in multiple consecutive time segments. The aggregated drift trend vector is input into the noise adversarial regularization layer of the power consumption drift trend inference network. The noise adversarial regularization layer performs signal perturbation injection processing and adversarial denoising processing on the aggregated drift trend vector to generate a regularized drift trend vector with noise resistance robustness. The regularized drift trend vector is input into the trend extrapolation expansion layer of the power consumption drift trend inference network. The trend extrapolation expansion layer projects the regularized drift trend vector along the time axis to the prediction time interval corresponding to the next scheduling period, generating discrete drift amplitude values ​​for each prediction time point within the prediction interval. The discrete drift amplitude value is input into the surface reconstruction layer of the power consumption drift trend inference network. The surface reconstruction layer performs continuous surface interpolation reconstruction processing on the discrete drift amplitude values ​​of all prediction time series points within the prediction interval to generate a prediction power consumption drift surface that covers the entire prediction time series interval and whose drift amplitude changes continuously.

5. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The step of selecting a set of candidate computing units from the computing unit resource pool of the computing server cluster whose power drift coupling relationship satisfies preset coupling screening conditions based on the predicted power drift surface includes: The drift amplitude value of the predicted power consumption drift surface at each predicted time point in the next scheduling cycle is analyzed. The drift amplitude value is subjected to time segmentation statistical processing. The average drift amplitude value and drift amplitude fluctuation variance of the predicted power consumption drift surface in each segment interval are calculated. A surface partition feature set containing partition identifiers and corresponding average drift amplitude values ​​and drift amplitude fluctuation variances is generated. Obtain the list of currently active computing units from the computing unit resource pool, and extract the historical power consumption drift surface of each active computing unit within the historical scheduling period. Perform partition statistical processing on the historical power consumption drift surface according to the same time partitioning method as the surface partition feature set, and generate the historical average drift amplitude value and historical drift amplitude fluctuation variance for each active computing unit corresponding to each partition. For each active computing unit, the historical average drift amplitude value and the predicted power consumption drift surface in the corresponding partition are processed by partition-by-partition difference calculation to obtain the average drift amplitude difference of each partition. At the same time, the historical drift amplitude fluctuation variance and the predicted power consumption drift surface in the corresponding partition are processed by partition-by-partition difference calculation to obtain the fluctuation variance difference of each partition. Using the average drift amplitude difference and fluctuation variance difference of each partition as input, the drift coupling relationship quantization function is called to perform drift coupling relationship quantization processing on each active computing unit, and the local coupling metric between each active computing unit and the predicted power consumption drift surface in each partition is calculated. The local coupling metric is inversely proportional to the average drift amplitude difference and inversely proportional to the fluctuation variance difference. The local coupling metric values ​​of each active computing unit in each partition are weighted and summed to generate a global drift coupling comprehensive metric value for each active computing unit. The weighting coefficients used in the weighting and summing process are inversely proportional to the time distance between the partitions at the current moment. The global drift coupling comprehensive metric value of each active computing unit is compared with the coupling metric threshold defined in the preset coupling screening conditions. Active computing units whose global drift coupling comprehensive metric value meets the coupling metric threshold are selected, and the selected active computing units are formed into a candidate computing unit set. Extract the number of available computing cores and available memory capacity for each candidate computing unit in the candidate computing unit set in the next scheduling cycle, construct the available resource configuration vector of the candidate computing unit, evaluate the matching degree between the available resource configuration vector of the candidate computing unit and the resource demand vector of the computing load to be allocated, eliminate candidate computing units whose resource matching degree is lower than the preset resource matching threshold, sort the remaining candidate computing units after resource matching and filtering in descending order according to the global drift coupling comprehensive metric value, and output the candidate computing unit set. Each candidate computing unit in the candidate computing unit set carries the corresponding global drift coupling comprehensive metric value and available resource configuration vector.

6. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The step of performing power consumption matching processing based on the power consumption requirement profiles of each computing unit in the candidate computing unit set and the computing load to be allocated, to obtain a computing unit power consumption matching allocation scheme, includes: The historical power consumption trajectory data of the computing load to be allocated within the historical operating cycle is obtained. The historical power consumption trajectory data is processed by time-series segmentation to obtain the stage power consumption sub-trajectory of the computing load to be allocated in different operating stages within the operating cycle. The combination of all stage power consumption sub-trajectories is defined as the load power consumption demand profile. Feature extraction processing is performed on each stage power consumption sub-trajectory in the load power consumption demand profile to extract the initial power consumption value, peak power consumption value, power consumption rise slope and power consumption duration of each stage power consumption sub-trajectory, forming a stage power consumption demand feature vector. Extract the historical power consumption drift surface of each candidate computing unit in the corresponding historical scheduling period from the candidate computing unit set, perform stage alignment mapping on the historical power consumption drift surface, and determine the historical power consumption supply capability characteristics of each candidate computing unit in the time window corresponding to each running stage of the computing load to be allocated. The historical power consumption supply capability characteristics include the historical average power consumption level and historical power consumption fluctuation amplitude in the corresponding time window. For each candidate computing unit, the historical power supply capability characteristics are matched with the stage power demand feature vector of the stage corresponding to the computing load to be allocated. The power supply and demand deviation parameter under each stage is calculated. The power supply and demand deviation parameter reflects the degree of difference between the historical power supply capability of the candidate computing unit and the power demand of the load. The power supply and demand deviation parameters of all stages are serialized and spliced ​​to construct the full-stage power deviation vector corresponding to each candidate computing unit. The full-stage power deviation vector is then smoothed and normalized to eliminate the step jump in power deviation between adjacent stages, generating a normalized full-stage power deviation curve. The operating stage corresponding to the deviation peak point is identified in the normalized full-stage power deviation curve. The operating stage where the identified deviation peak point is located is marked as a risk mismatch stage. The number of risk mismatch stages and the deviation magnitude of the risk mismatch stages of each candidate computing unit are weighted and deducted to obtain the power matching and adaptation score of each candidate computing unit. Candidate computing units whose power consumption matching and adaptation scores meet a preset adaptation threshold are selected. The selected candidate computing units are bound and paired with the computing load to be allocated, and a pairing mapping table containing the addressing identifier of the candidate computing units and the identifier of the computing load to be allocated is generated. For each pair of bound and paired candidate computing units and computing loads in the pairing mapping table, a corresponding computing unit occupancy time window is allocated according to the power consumption demand time interval of each stage in the load power consumption demand profile, and the computing unit power consumption matching and allocation scheme is generated.

7. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The step of generating and outputting a set of power balancing scheduling instructions, including computing unit addressing identifiers and computing load migration timings, based on the computing unit power consumption matching and allocation scheme, includes: The pairing mapping relationship between the computing unit addressing identifier and the computing load identifier to be allocated in the computing unit power consumption matching allocation scheme is analyzed, as well as the computing unit occupancy time window corresponding to each pairing mapping relationship. The start timing and end timing of the computing unit addressing identifier, the computing load identifier to be allocated, and the occupancy time window of each pairing mapping relationship are extracted. Based on the addressing identifier of the computing unit in each pairing mapping relationship, the topology database of the computing server cluster is queried to obtain the network addressing path and communication port identifier of the computing unit within the cluster, and computing unit location information containing the network addressing path and communication port identifier of the computing unit is generated. Based on the identifier of the computing load to be allocated in each pairing mapping relationship, query the load storage location mapping table of the computing load storage management node, obtain the storage path of the executable file of the computing load to be allocated and the storage node addressing information of the dataset dependent on the load, and generate the load source data location information. Based on the start timing of the occupied time window of each pairing mapping relationship and the data location information of the load source end, calculate the estimated transmission time required for data transmission between the storage node corresponding to the data location information of the load source end and the computing unit corresponding to the computing unit location information, and reserve the estimated transmission time forward based on the start timing of the occupied time window to obtain the computing load migration start timing. Based on the computing load migration start timing, combined with the start timing of the occupied time window and the computing unit location information, a single load migration instruction is constructed, which includes the migration start timing, the target computing unit network addressing path, the target computing unit communication port identifier, the load identifier to be migrated, and the load source data location information. For each pairing mapping relationship, generate a corresponding single load migration instruction, and arrange all the single load migration instructions in the order of their respective migration start times to form a load migration instruction sequence; For load migration instruction sequences with overlapping migration time windows, conflict detection processing is performed to identify conflicting overlapping instructions and to serialize the migration start timing of the overlapping instructions. The conflicting load migration instructions are then rearranged according to a preset priority rule to generate a conflict-free load migration instruction sequence. The conflict-free load migration instruction sequence is encapsulated with the time window of each pairing mapping relationship in the power consumption matching and allocation scheme of the computing unit. This generates a power balancing scheduling instruction set that includes the scheduling time window definition of each load migration instruction and the corresponding computing unit. The power balancing scheduling instruction set is then broadcast to the scheduling execution node of the computing server cluster and all involved computing units. After the broadcast is completed, instruction confirmation acknowledgment signals are received from each computing unit.

8. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The method further includes: After the computing server cluster executes the scheduling operation corresponding to the power balancing scheduling instruction set, it continuously collects the actual power consumption sensing signal sequence and the actual task load signal sequence of each computing unit during the task execution cycle, and encapsulates the actual power consumption sensing signal sequence and the actual task load signal sequence into an actual synchronous monitoring signal array. The actual synchronous monitoring signal array is subjected to power ripple morphology analysis processing to extract the actual power ripple peak and valley timing features of each arithmetic unit. The actual power ripple peak and valley timing features are compared with the predicted power drift surface segment by segment to generate the timing deviation distribution between the actual power curve and the predicted power drift surface. Extract the timing points in the timing deviation distribution where the deviation amplitude exceeds the preset deviation tolerance range, and mark the timing points as power prediction deviation anomaly points. Backtrack the computing load type identifier carried by the computing unit and the corresponding load switching event sequence when the power prediction deviation anomaly point appears. Based on the occurrence sequence of the power consumption prediction deviation anomalies and the corresponding computing load type identifiers, a prediction deviation event record containing deviation amplitude, deviation timing and load type association tags is constructed, and the prediction deviation event record is appended to the prediction deviation historical event database. Cluster analysis is performed on the accumulated prediction deviation event records in the prediction deviation historical event database to identify deviation patterns that frequently appear under a specified load type identifier, and the statistical characteristics of the deviation patterns are encapsulated as load type deviation fingerprint features. The load type deviation fingerprint feature is fed back to the trend extrapolation expansion layer of the power consumption drift trend inference network. The trend extrapolation parameters in the trend extrapolation expansion layer are incrementally adjusted so that the adjusted trend extrapolation parameters can reduce the prediction deviation for the computing load corresponding to the load type identifier.

9. The power consumption optimization method for computing servers based on deep learning according to claim 1, characterized in that, The method further includes: During the process of the power drift trend inference network inferring the power drift direction of the adjacent time segment of the ripple load associated feature vector, the hidden state feature vector sequence output by the drift state space modeling layer in the power drift trend inference network is recorded simultaneously. The hidden state feature vector sequence is subjected to time-series segmentation processing to divide the hidden state feature vector sequence into multiple non-overlapping hidden state segments. Statistical feature extraction processing is performed on each hidden state segment to extract the hidden state mean vector and hidden state covariance features of each hidden state segment. The difference measure is performed on the hidden state mean vectors of adjacent hidden state segments to calculate the vector offset distance between the hidden state mean vectors of adjacent hidden state segments. At the same time, the difference measure is performed on the hidden state covariance features of adjacent hidden state segments to calculate the distribution offset distance of the hidden state covariance features between adjacent hidden state segments. The vector offset distance and the distribution offset distance are fused to generate a hidden state drift phase transition metric parameter between adjacent hidden state segments. The hidden state drift phase transition metric parameter indicates the severity of the sudden change in the hidden state distribution characteristics between adjacent time segments. Accumulate and calculate the hidden state drift phase transition measurement parameters between multiple consecutive hidden state segments, and construct a hidden state drift phase transition sequence from the multiple consecutive hidden state drift phase transition measurement parameters. In the hidden state drift phase transition sequence, identify the phase transition point where the hidden state drift phase transition measurement parameters exceed the phase transition trigger threshold. Obtain the timing position corresponding to the phase transition point, extract the power consumption sensing signal frame and task load signal frame within a preset window range before and after the timing position from the synchronous monitoring signal array based on the timing position, and perform load mutation event correlation analysis on the extracted signal frames to determine the type of load mutation event that leads to the hidden state phase transition. The load mutation event type and the hidden state drift phase transition metric parameters corresponding to the phase transition point are encoded into a drift network phase transition trigger signal. The drift network phase transition trigger signal is sent to the scheduling management node of the computing server cluster to trigger the scheduling management node to perform scheduling strategy pre-adjustment processing for the load mutation event type.

10. A power consumption optimization system for computing servers based on deep learning, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the deep learning-based computing server power optimization method according to any one of claims 1 to 9 by executing the machine-executable instructions.