Intelligent computing power cluster energy consumption prediction modeling method, load prediction method and system
Through the Pearson correlation coefficient method, EMD modal decomposition and t-SNE dimensionality reduction processing, combined with the LSTM model, the lack of energy consumption prediction in data centers is solved, and accurate prediction and resource optimization management of energy consumption of intelligent computing power clusters is realized, energy efficiency is improved and operation costs are reduced.
Patent Information
- Application Number
- CN202510224192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-07-08
AI Technical Summary
The existing data center data flow and energy flow coordinated adjustment lacks computing power and energy consumption analysis and prediction, making it difficult to conduct quantitative analysis and verification, making it difficult to optimize resource allocation and reduce operating costs.
The Pearson correlation coefficient method is used to eliminate factors that have a correlation less than the set threshold. After EMD modal decomposition and t-SNE dimensionality reduction processing, the LSTM model is used to predict the energy consumption of intelligent computing power clusters and build a load prediction model.
It realizes accurate prediction of the energy consumption of intelligent computing power clusters, optimizes resource scheduling, improves energy efficiency, reduces operating costs, and enhances the flexible response capabilities of the power system.
Smart Images

Figure CN120278306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to artificial intelligence, algorithms, and natural language processing, and particularly to an intelligent computing power cluster energy consumption prediction modeling method, a load prediction method, and a system. Background Art
[0002] With the rapid development of big data, cloud computing, and artificial intelligence technologies, intelligent computing power clusters (such as data centers, high-performance computing clusters, etc.) have become key infrastructures to support various applications and services. However, these clusters generate a large amount of energy consumption during operation, and their performance is affected by various factors, such as computing temperature, humidity, air pressure, load, electricity price, etc. Therefore, predicting the performance and energy consumption of intelligent computing power clusters through time series analysis is of great significance for optimizing resource allocation, improving energy efficiency, and reducing operating costs.
[0003] With the implementation of the computing power hub of the integrated big data center collaborative innovation system, it has laid a favorable policy foundation for deepening cross-provincial and cross-regional computing power collaboration, resource access scheduling, promoting computing power allocation and resource scheduling in the intelligent computing center, as well as increasing the proportion of green power consumption and improving the comprehensive energy efficiency level. At present, the online monitoring and analysis system of the computing resource consumption situation in the intelligent computing center and the online monitoring system of the energy supply system in the data center do not operate in coordination. The research on the coordinated regulation of the data flow and energy flow in the data center is still in the theoretical research and analysis stage, lacking the analysis and prediction of the computing power energy consumption in the data center, and it is difficult to quantitatively analyze and deduce and verify the regulation effect of the data center. Summary of the Invention
[0004] To solve the problem that the coordinated regulation of the data flow and energy flow in the data center lacks the analysis and prediction of the computing power energy consumption and it is difficult to quantitatively analyze and deduce and verify the regulation effect of the data center, the present invention provides an intelligent computing power cluster energy consumption prediction modeling method, including:
[0005] Obtain the initial load data of the cluster computing power and the relevant time series data of each influencing factor, and use the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than the set threshold to obtain the first influencing factor;
[0006] Perform EMD modal decomposition on the relevant time series data of the first influencing factor to obtain several IMF characteristic fluctuation sequences;
[0007] Perform t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences;
[0008] Use the dimensionality-reduced IMF characteristic fluctuation sequences and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model.
[0009] Preferably, the obtaining of the initial load data of the cluster computing power and the time series data related to each influencing factor, and the elimination of the influencing factors with a correlation less than a set threshold by using the Pearson correlation coefficient method to obtain the first influencing factors include:
[0010] Select the initial influencing factors of the intelligent computing power cluster load;
[0011] Based on the same time series as the initial load data, collect the data series corresponding to the initial influencing factors;
[0012] Use the Pearson correlation coefficient method to determine the correlation between the initial influencing factors and the intelligent computing power cluster load value;
[0013] Eliminate the influencing factors with a correlation ranking less than the set threshold to obtain the first influencing factors.
[0014] Preferably, the initial influencing factors include one or more of temperature, humidity, air pressure, load, electricity price, task complexity, air pressure, rainfall, and wind speed.
[0015] Preferably, the performing of EMD modal decomposition on the relevant time series data of the first influencing factors to obtain several IMF characteristic fluctuation sequences includes:
[0016] Based on each influencing factor in the first influencing factors, respectively perform the following steps:
[0017] Identify all local extreme points from the time series data corresponding to the selected influencing factors, where the extreme values include maximum values and minimum values;
[0018] Use the identified extreme points to construct upper and lower envelope lines, and calculate the first IMF sequence h1(t) through the mean of the envelope lines;
[0019] Subtract the first IMF sequence h1(t) from the original signal to obtain an updated residual signal; repeat this step until the residual signal becomes a monotonic sequence or meets a preset stopping criterion, thereby obtaining several IMF characteristic fluctuation sequences corresponding to the influencing factors.
[0020] Preferably, the performing of t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences includes:
[0021] Calculate the high-dimensional distribution probability P between two data points in the IMF characteristic fluctuation sequences;
[0022] Take the low-dimensional distribution probability Q between two data points in the initial load data;
[0023] With the goal of minimizing the KL divergence between the high-dimensional distribution P and the low-dimensional distribution Q, the gradient descent method is used to iteratively update the data points in the low-dimensional space to obtain the low-dimensional space data.
[0024] Preferably, after performing t-SNE dimensionality reduction processing on the IMF feature fluctuation sequence, it further includes displaying the clustering and distribution structure of the low-dimensional space data in a two-dimensional or three-dimensional graph.
[0025] Preferably, using the dimensionality-reduced IMF feature fluctuation sequence and the initial load data of the cluster computing power as input data, training the LSTM prediction model to obtain a load prediction model includes:
[0026] Performing standardization processing on the initial load data of the cluster computing power and the dimensionality-reduced IMF feature fluctuation sequence;
[0027] Dividing the standardized load data and feature fluctuation sequence into a training set and a test set;
[0028] Based on the training set, training the LSTM model to obtain a load prediction model;
[0029] Using the test set, evaluating the load prediction model with the mean absolute error, mean square error, root mean square error, and goodness of fit as evaluation indicators. If the evaluation result does not meet the target accuracy rate, return to modify the parameters and retrain the LSTM model until the evaluation result meets the target accuracy rate.
[0030] Based on the same inventive concept, the present invention also provides an intelligent computing power cluster energy consumption prediction modeling system, including:
[0031] An influencing factor elimination module, configured to obtain the initial load data of the cluster computing power and the time series data related to each influencing factor, and use the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than a set threshold to obtain the first influencing factor;
[0032] An EMD modal decomposition module, configured to perform EMD modal decomposition on the time series data of the first influencing factor to obtain a plurality of IMF feature fluctuation sequences;
[0033] A dimensionality reduction processing module, configured to perform t-SNE dimensionality reduction processing on the IMF feature fluctuation sequence;
[0034] A model construction module, configured to use the dimensionality-reduced IMF feature fluctuation sequence and the initial load data of the cluster computing power as input data, and train the LSTM prediction model to obtain a load prediction model.
[0035] Based on the same inventive concept, the present invention also provides an accurate load prediction method under a cluster, including:
[0036] Obtain the first influencing factor sequence for a period of time before the current moment;
[0037] Using the first influencing factor sequence as input, perform energy consumption prediction for each prediction stage by using a load prediction model;
[0038] The load prediction model is constructed as provided by an intelligent computing power cluster energy consumption prediction modeling method of the present invention.
[0039] Preferably, the prediction stage includes energy consumption prediction for the day-ahead stage, the intra-day stage, and the real-time stage; the time scale of the first influencing factor sequence is determined by the prediction stage.
[0040] Based on the same inventive concept, the present invention also provides a precise load prediction system under a cluster, including:
[0041] An acquisition module, configured to obtain the first influencing factor sequence for a period of time before the current moment;
[0042] A prediction module, configured to use the first influencing factor sequence as input and perform energy consumption prediction for each prediction stage by using a load prediction model;
[0043] The load prediction model is constructed as provided by an intelligent computing power cluster energy consumption prediction modeling method of the present invention.
[0044] Based on the same inventive concept, the present invention also provides an electronic device, including: at least one processor and a memory; the memory and the processor are connected by a bus;
[0045] The memory is configured to store one or more programs;
[0046] When the one or more programs are executed by the at least one processor, an intelligent computing power cluster energy consumption prediction modeling method provided by the present invention and / or a precise load prediction method under a cluster provided by the present invention are implemented.
[0047] Based on the same inventive concept, the present invention also provides a readable storage medium, on which an execution program is stored, and when the execution program is executed, an intelligent computing power cluster energy consumption prediction modeling method provided by the present invention and / or a precise load prediction method under a cluster provided by the present invention are implemented.
[0048] Compared with the prior art, the beneficial effects of the present invention are:
[0049] 1. The present invention provides an intelligent computing power cluster energy consumption prediction modeling method and system, including: obtaining the initial load data of the cluster computing power and the relevant time series data of each influencing factor, and using the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than a set threshold to obtain the first influencing factors; performing EMD modal decomposition on the relevant time series data of the first influencing factors to obtain a number of IMF characteristic fluctuation sequences; performing t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences; using the dimensionality-reduced IMF characteristic fluctuation sequences and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model; the load prediction model of the present invention realizes accurate prediction of the energy consumption of the intelligent computing power cluster by considering the time series data flow characteristics of the main influencing factors with high correlation;
[0050] 2. A precise load prediction method under a cluster provided by the present invention includes: obtaining the first influencing factor sequence for a period of time before the current moment; using the first influencing factor sequence as input and using the load prediction model of the present invention to predict the energy consumption in each prediction stage; it can realize precise load prediction under a high-performance computing cluster to promote the computing power allocation and resource scheduling of the intelligent computing center. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a schematic flow chart of an intelligent computing power cluster energy consumption prediction modeling method of the present invention;
[0052] Figure 2 It is a modeling flow chart of the load prediction model in Embodiment 1;
[0053] Figure 3 It is a schematic diagram of the LSTM neuron structure of the present invention;
[0054] Figure 4 It is a schematic diagram of the structure of an intelligent computing power cluster energy consumption prediction modeling system of the present invention;
[0055] Figure 5 It is a schematic diagram of the structure of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] In order to solve the problem that the coordinated regulation of data flow and energy flow in the data center lacks computing power energy consumption analysis and prediction, and it is difficult to quantify and analyze the regulation effect of the data center and verify the performance, and realize the energy consumption reduction and performance prediction requirements of the intelligent computing power cluster, the present invention proposes an intelligent computing power cluster energy consumption prediction modeling method, load prediction method and system. The present invention is based on the intelligent computing power cluster performance and energy consumption prediction technology of time series analysis, and establishes an intelligent computing center load prediction model based on data flow characteristics, which is used to cover the data load, cooling load, self-contained power supply and other diversified resources in the data center. In the control scheme of grid interaction. At the same time, the EMD-t-SNE-LSTM algorithm is used to screen out the main influencing factors such as temperature, humidity, air pressure, load, and electricity price, and the energy consumption is predicted for the performance of the intelligent computing power cluster. The model has a high prediction accuracy through multiple indicators to achieve effective prediction of the energy consumption of the intelligent computing power cluster and optimized management of performance, improve energy efficiency and reduce operating costs. The present invention can be applied to support the energy consumption prediction, load characteristic analysis, grid interaction potential and other directions of intelligent computing power clusters, enhance the flexible response capability of the power system, promote the construction of green data centers and the development of smart grids, and can be widely used to support the energy consumption prediction, load characteristic analysis, grid interaction potential and other directions of intelligent computing power clusters. In order to better understand the present invention, the content of the present invention is further explained below in conjunction with the accompanying drawings and examples of the specification.
[0057] Embodiment 1:
[0058] The present invention proposes a method for predicting energy consumption of intelligent computing clusters, and uses the EMD-t-SNE-LSTM method to build a load prediction model. Figure 1 As shown, including:
[0059] S1. Obtain the initial load data of the cluster computing power and the time series data related to each influencing factor, and use the Pearson correlation coefficient method to eliminate the influencing factors with correlations less than the set threshold to obtain the first influencing factor;
[0060] S2. Perform EMD modal decomposition on the relevant time series data of the first influencing factor to obtain several IMF characteristic fluctuation sequences;
[0061] S3, performing t-SNE dimension reduction processing on the IMF characteristic fluctuation sequence;
[0062] S4. Using the IMF characteristic fluctuation sequence after dimensionality reduction processing and the initial load data of the cluster computing power as input data, the LSTM prediction model is trained to obtain a load prediction model to realize the prediction of the energy consumption of the intelligent computing power cluster.
[0063] Specifically, the load forecasting model in the present invention is specifically constructed as follows: Figure 2 shown.
[0064] Among them, the detailed steps of step S1 to obtain the initial load data of the cluster computing power and the time series data related to each influencing factor, and use the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than the set threshold to obtain the first influencing factor are as follows:
[0065] Use the Pearson correlation coefficient method to analyze the correlation between the influencing factors and the load value, and eliminate the parameters with weak correlation. The calculation formula of the Pearson correlation coefficient R is:
[0066]
[0067] where n is the number of influencing factors, and X o are the current temperature, humidity, wind speed, air pressure, rainfall, etc. respectively. are the average values of the current temperature, humidity, wind speed, air pressure, rainfall, etc. respectively, and Y o is the load value of the current cluster computing power, is the average value of the load of the cluster computing power. R ∈ [-1, 1]. The positive or negative of R represents the positive or negative correlation relationship between variables, and the magnitude of the absolute value is proportional to the correlation of the variables.
[0068]
[0069] Through the Pearson correlation coefficient method for correlation analysis, select the influencing factors with a relatively high degree of correlation with the load value. For example, select the top 5 influencing factors with a relatively high degree of correlation as the first influencing factors. Note that the number of factors selected here can be set according to actual calculation needs. Eliminate the influencing factors with a relatively low degree of correlation. According to the screening by the Pearson correlation coefficient method, the task complexity, air pressure, rainfall, and wind speed are low-correlation influencing factors, so they are eliminated. Retain the influencing factors with high correlation such as temperature, humidity, air pressure, load, and electricity price.
[0070] Step S2: Perform EMD modal decomposition on the time series data related to the first influencing factor to obtain several IMF characteristic fluctuation sequences. The specific process is as follows:
[0071] Empirical mode decomposition (EMD) is an adaptive data analysis technique, especially good at processing complex, non-linear and non-stationary time series data. The core of its self-adaptability lies in being able to automatically perform decomposition according to the inherent time scale characteristics of the data, without preset basis functions or preset decomposition levels, significantly reducing the need for manual intervention.
[0072] The EMD method iteratively resolves the original data into a series of simple oscillatory components with physical significance - Intrinsic Mode Functions (IMFs), and a residual sequence reflecting the long-term trend of the data. This process is carried out directly in the time domain. First, local extreme points in the original signal are identified and extracted, and upper and lower envelope lines are constructed using these points. The first IMF is calculated through the mean of the envelope lines. Subsequently, this IMF is subtracted from the original signal to obtain an updated residual signal. This process is repeated until the residual becomes a monotonic sequence or meets a preset stopping criterion, effectively stripping out the multi-scale features and high-frequency fluctuations in the data, and achieving a step-by-step decomposition of the signal from complex to simple.
[0073] In the first step, using empirical mode decomposition technology, we can obtain the specific characteristics of the factors affecting the computing power of the intelligent cluster at different time periods. This process will identify more refined but increasing numbers of influencing factors.
[0074] The specific decomposition process is as follows (taking environmental temperature as an example):
[0075] 1. Determine extreme points: First, identify all local maxima and minima from the time series X(t) of temperature TEMP; where t is the time variable.
[0076] 2. Construct envelopes: Using these extreme points, construct the upper envelope U(t) and lower envelope L(t) of the signal through cubic spline interpolation.
[0077] 3. Calculate the average envelope: Calculate the average of the upper and lower envelopes to obtain the temperature mean m(t).
[0078]
[0079] 4. Extract IMF: Subtract the average envelope from the original signal to obtain the first IMF sequence h1(t).
[0080] h1(t) = X(t) - m(t)
[0081] 5. Verify and iterate: Check whether h1(t) meets the conditions of the IMF. If it meets, it is the first IMF, denoted as C1(t), and then subtract it from the original signal to obtain a new residual signal R(t).
[0082] R(t) = X(t) - C1(t)
[0083] If it does not meet, repeat steps 1 to 4 until the residual sequence is monotonic or meets the stopping criterion s d ≤0.2 - 0.3. s d represents the standard deviation of the residual sequence, and 0.2 and 0.3 are thresholds. If the change is small enough, the decomposition terminates.
[0084] 6. Decomposition result: Finally, the original signal X(t) is decomposed into several IMFs and a trend term, and the expression is:
[0085]
[0086] where n represents the number of IMFs, C i (t) represents the i-th IMF, and r n is the trend term.
[0087] By gradually extracting different frequency components in the temperature (TEMP), humidity (WET), barometric pressure (BP), load (LOAD), and electricity price (Price) signals according to this process, the original signal is simplified into a series of simple oscillation modes and a trend term, which provides convenience for further analysis.
[0088] After being finely processed by the EMD (Empirical Mode Decomposition) technique, the richness of the data sequence is significantly improved, greatly enhancing the diversity of the feature set. However, this enhancement is accompanied by a sharp increase in the dimension of the input variables, bringing challenges to subsequent processing.
[0089] Therefore, in step S3 for performing t-SNE dimensionality reduction processing on the IMF feature fluctuation sequence, the t-Distributed Stochastic Neighbor Embedding (abbreviated as t-SNE) technique is introduced to perform dimensionality reduction processing on the input variables. While ensuring the prediction accuracy of the LSTM (Long Short-Term Memory) model, optimizing its computational efficiency and effectively resisting the overfitting problem, t-SNE, with its unique non-linear dimensionality reduction ability, significantly reduces the number of input variables on the premise of keeping the key information and local structure of the data unchanged. This process can accelerate the calculation process of the LSTM model and improve the generalization ability of the model through a more compact data representation, ensuring the high accuracy of the prediction results. It also performs well in terms of monotonicity, correlation, and robustness. Achieving the improvement of accuracy and efficiency can effectively optimize the performance of the LSTM model.
[0090] The following is the calculation process of t-SNE:
[0091] Take the feature sequences of the N IMFs (N rows and M columns) signals obtained by EMD decomposition as the high-dimensional data {x1, x2, …, x N} and input them for t-SNE decomposition.
[0092] 1. Calculate the similarity in the high-dimensional space: For each data point x i , calculate the Euclidean distance between it and all other points x j . p ijIndicates that the center point is x i , and the adjacent point is x j when the probability. Since p i|j is symmetric, the distribution probability between two points of its high-dimensional data points is:
[0093]
[0094] 2. Initialize the low-dimensional space: Randomly initialize the same number of points {y1, y2, …, y N} in two-dimensional or three-dimensional space, and use the t-distribution with a degree of freedom of 1. q ij represents the distribution probability of the low-dimensional data points {y1, y2, …, y N}:
[0095]
[0096] 3. Define and optimize the objective function: The goal is to minimize the KL divergence between the high-dimensional distribution P and the low-dimensional distribution Q:
[0097]
[0098] where the smaller the C value, the higher the accuracy. Use the gradient descent method to update the position of y i , and the gradient calculation formula is:
[0099]
[0100] 4. Iterative optimization: Repeatedly execute the gradient descent step to update the points in the low-dimensional space until the convergence condition is reached (such as the gradient is less than the threshold or the maximum number of iterations is reached).
[0101] 5. Result visualization: Plot the optimized low-dimensional space points y i as a two-dimensional or three-dimensional graph to intuitively display the clustering and distribution structure of the data.
[0102] t-SNE can better retain the local features of the original data, enable the data processed by EMD to be effectively and visually represented in the low-dimensional space, and significantly reduce the redundancy and correlation of the time series after EMD decomposition, which is more conducive to subsequent model training and prediction. Thus, the dimensionality-reduced feature sequence can be obtained.
[0103] Step S4 uses the IMF feature fluctuation sequence after dimensionality reduction processing and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model for predicting the energy consumption of the intelligent computing power cluster, including: standardizing the original data and the data after t-SNE dimensionality reduction, dividing the training set and the test set, and inputting them into the LSTM model respectively, and inputting the data after dimensionality reduction into the LSTM model for prediction.
[0104] The input and output data of the model are as follows:
[0105] Input: Temperature (TEMP), Humidity (WET), Barometric Pressure (BP), Load (LOAD), Electricity Price (Price); Output: Model evaluation metrics RMSE, MSE, MAE, R 2
[0106] LSTM is a special architecture of the Recurrent Neural Network (RNN). On this basis, a forget gate, an input gate, and an output gate are added. By introducing a gating mechanism to solve the problems encountered by traditional RNNs in dealing with long-term dependencies, this effectively avoids the gradient explosion and gradient disappearance problems of the RNN network when dealing with long-time data, thus effectively improving the accuracy of model prediction.
[0107] The LSTM neuron structure is as Figure 3 shown, and its forward calculation formula is:
[0108]
[0109] Among them, f t is the activation value of the forget gate, h t-1 is the output of the previous moment, X t is the current input, W f and b f are the weights and biases of the forget gate respectively, W i and b i are the weights and biases of the input gate respectively, W C and b C are the weights and biases of the candidate cell state respectively, σ is the sigmoid activation function (mapping the value between 0 and 1), and tanh is the hyperbolic tangent function. is the candidate unit state, C t is the unit state, i t is the activation value of the input gate, h t is the output of the LSTM, o t is the activation value of the output gate.
[0110] The evaluation of the model includes:
[0111] The mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE), and goodness of fit (R 2 ) are used to evaluate the model. MAE, MSE, RMSE, and R 2 are important indicators for judging the prediction model.
[0112]
[0113] Among them, y i and are the original data and predicted data of the cluster computing power load respectively, is the average value of the cluster computing power load. n is the size of the data volume. The smaller the MAE, MSE, and RMSE, the higher the model accuracy. The closer R 2 is to 1, the better the fitting effect of the model. If the value of the evaluation function does not meet the target accuracy rate, the modified parameters are returned and the LSTM model is retrained.
[0114] Example 2:
[0115] Based on the same inventive concept, the present invention also provides an intelligent computing power cluster energy consumption prediction modeling system, as Figure 4 shown, including:
[0116] An influencing factor elimination module, which is used to obtain the initial load data of the cluster computing power and the time series data related to each influencing factor, and uses the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than the set threshold to obtain the first influencing factor;
[0117] An EMD modal decomposition module, which is used to perform EMD modal decomposition on the time series data related to the first influencing factor to obtain a number of IMF characteristic fluctuation sequences;
[0118] A dimensionality reduction processing module, which is used to perform t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences;
[0119] A model construction module, which is used to use the IMF characteristic fluctuation sequences after dimensionality reduction processing and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model.
[0120] Specifically, each module is used to implement the intelligent computing power cluster energy consumption prediction modeling method in the above embodiment, which will not be elaborated here.
[0121] Example 3:
[0122] Based on the same inventive concept, the present invention also provides an accurate load prediction method under a cluster. The load prediction method of the present invention uses a load prediction model, and this load prediction model is constructed by using the intelligent computing power cluster energy consumption prediction modeling method in the above embodiment. Specifically, it includes:
[0123] Obtain the first influence factor sequence for a period of time before the current moment;
[0124] Using the first influence factor sequence as the input, perform energy consumption prediction for each prediction stage using the load prediction model.
[0125] The load prediction of the present invention can be divided into three prediction stages: day-ahead, intra-day, and real-time. Its time scale is determined by the prediction stage. They are 1h, 15min, and 5min respectively, and the loads for the subsequent 24h, 2h, and 5min are predicted respectively.
[0126] Collect the data sequences of each first influence factor according to the set time scale, and perform load prediction for the three prediction stages of day-ahead, intra-day, and real-time respectively according to the prediction stage.
[0127] Embodiment 4
[0128] To implement the accurate load prediction method under the cluster in the above embodiment, the present invention also provides an accurate load prediction system under the cluster, including:
[0129] An acquisition module, configured to obtain the first influence factor sequence for a period of time before the current moment;
[0130] A prediction module, configured to use the first influence factor sequence as the input and perform energy consumption prediction for each prediction stage using the load prediction model;
[0131] The load prediction model is constructed as provided by an intelligent computing power cluster energy consumption prediction modeling method in the above embodiment of the present invention.
[0132] Embodiment 5
[0133] As Figure 5 shown, the present invention also provides an electronic device, which may be a computer device, a single-chip microcomputer device, an intelligent mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, the processor, and the transceiver component are connected by a bus; the memory can be used to store an execution program, and an exemplary execution program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, and the data can be called and / or modified when the instructions are executed.
[0134] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the storage medium to implement the corresponding method flow or corresponding function, so as to implement the steps of a method for predicting the energy consumption of an intelligent computing power cluster in the above embodiments and / or the steps of a precise load prediction method under a cluster in the above embodiments.
[0135] Embodiment 6
[0136] Based on the same inventive concept, the present invention also provides a readable storage medium, specifically an electronic device-readable storage medium (Memory). The electronic device-readable storage medium is a memory device in the electronic device, used to store programs and data. It can be understood that the storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The storage medium provides a storage space, and this storage space stores the operating system of the terminal. And, one or more instructions suitable for being loaded and executed by the processor are also stored in this storage space. These instructions can be one or more executable programs (including program codes). It should be noted that the storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. By the processor loading and executing one or more instructions stored in the storage medium, the steps of a method for predicting the energy consumption of an intelligent computing power cluster in the above embodiments and / or the steps of a precise load prediction method under a cluster in the above embodiments can be implemented.
[0137] Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0138] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0139] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.
[0140] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realize the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.
[0141] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.
[0142] The above are only embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention awaiting approval.
Claims
1. An intelligent computing power cluster energy consumption prediction modeling method, characterized in that Including: Obtain the initial load data of the cluster computing power and the time series data related to each influencing factor, and use the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than the set threshold to obtain the first influencing factor; Perform EMD modal decomposition on the time series data related to the first influencing factor to obtain a number of IMF characteristic fluctuation sequences; Perform t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences; Use the dimensionality-reduced IMF characteristic fluctuation sequences and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model, so as to realize the prediction of the energy consumption of the intelligent computing power cluster.
2. The modeling method according to claim 1, wherein: The obtaining the initial load data of the cluster computing power and the time series data related to each influencing factor, and using the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than the set threshold to obtain the first influencing factor includes: Select the initial influencing factors of the intelligent computing power cluster load; Based on the same time series as the initial load data, collect the data sequences corresponding to the initial influencing factors; Use the Pearson correlation coefficient method to determine the correlation between the initial influencing factors and the intelligent computing power cluster load value; Eliminate the influencing factors with a correlation ranking less than the set threshold to obtain the first influencing factor.
3. The modeling method according to claim 1 or 2, characterized in that: The initial influencing factors include one or more of temperature, humidity, air pressure, load, electricity price, task complexity, air pressure, rainfall, and wind speed.
4. The modeling method according to claim 1, characterized in that: The performing EMD modal decomposition on the time series data related to the first influencing factor to obtain a number of IMF characteristic fluctuation sequences includes: Perform the following steps respectively based on each influencing factor in the first influencing factor: Identify all local extreme points from the time series data corresponding to the selected influencing factors, and the extreme values include maximum values and minimum values; Use the identified extreme points to construct upper and lower envelope lines, and calculate the first IMF sequence h1(t) through the mean of the envelope lines; Subtract the first IMF sequence h1(t) from the original signal to obtain an updated residual signal; repeat this step until the residual signal becomes a monotonic sequence or meets the preset stop criterion, and then obtain a number of IMF characteristic fluctuation sequences corresponding to the influencing factors.
5. The modeling method according to claim 1, characterized in that: The performing t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences includes: Calculate the high-dimensional distribution probability P between two data points in the IMF characteristic fluctuation sequences; Use the low-dimensional distribution probability Q between two data points in the initial load data; With the goal of minimizing the KL divergence between the high-dimensional distribution P and the low-dimensional distribution Q, use the gradient descent method to iteratively update the data points in the low-dimensional space to obtain the low-dimensional space data.
6. The modeling method according to claim 5, characterized in that: After the performing t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences, it further includes Display the clustering and distribution structure of the low-dimensional space data in a two-dimensional or three-dimensional graph.
7. The modeling method according to claim 1, wherein: The using the dimensionality-reduced IMF characteristic fluctuation sequences and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model includes: Perform standardization processing on the initial load data of the cluster computing power and the dimensionality-reduced IMF characteristic fluctuation sequences; Divide the standardized load data and the characteristic fluctuation sequence into a training set and a test set; Based on the training set, train the LSTM model to obtain a load prediction model; Use the test set to evaluate the load prediction model with the mean absolute error, mean square error, root mean square error, and goodness of fit as evaluation indicators. If the evaluation result does not meet the target accuracy rate, return to modify the parameters and retrain the LSTM model until the evaluation result meets the target accuracy rate.
8. An intelligent computing power cluster energy consumption prediction and modeling system, characterized in that, Include: An influencing factor elimination module, configured to obtain the initial load data of the cluster computing power and the time series data related to each influencing factor, and use the Pearson correlation coefficient method to eliminate the influencing factors with a correlation less than a set threshold to obtain the first influencing factor; An EMD modal decomposition module, configured to perform EMD modal decomposition on the time series data of the first influencing factor to obtain a plurality of IMF characteristic fluctuation sequences; A dimensionality reduction processing module, configured to perform t-SNE dimensionality reduction processing on the IMF characteristic fluctuation sequences; A model construction module, configured to use the IMF characteristic fluctuation sequences after dimensionality reduction processing and the initial load data of the cluster computing power as input data to train the LSTM prediction model to obtain a load prediction model.
9. A precise load forecasting method under a cluster, characterized in that, Include: Obtain the first influencing factor sequence for a period of time before the current moment; Use the first influencing factor sequence as input and use the load prediction model to predict the energy consumption in each prediction stage; The load prediction model is constructed by an intelligent computing power cluster energy consumption prediction modeling method as described in any one of claims 1 to 7.
10. The prediction method according to claim 9, wherein, The prediction stage includes energy consumption prediction in the day-ahead stage, intra-day stage, and real-time stage; the time scale of the first influencing factor sequence is determined by the prediction stage.
11. A precise load forecasting system under a cluster, characterized in that, Include: An acquisition module, configured to obtain the first influencing factor sequence for a period of time before the current moment; A prediction module, configured to use the first influencing factor sequence as input and use the load prediction model to predict the energy consumption in each prediction stage; The modeling is constructed by an intelligent computing power cluster energy consumption prediction modeling method as described in any one of claims 1 to 7.
12. An electronic device, characterized in that, Include: At least one processor and a memory; The memory and the processor are connected by a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, an intelligent computing power cluster energy consumption prediction modeling method as described in any one of claims 1 to 7 and / or a precise load prediction method under a cluster as described in claim 9 are implemented.
13. A readable storage medium, characterized in that, There is an execution program stored thereon, and when the execution program is executed, an intelligent computing power cluster energy consumption prediction modeling method as described in any one of claims 1 to 7 and / or a precise load prediction method under a cluster as described in claim 9 are implemented.