Data center refrigeration energy consumption short-term prediction method based on feature screening-fusion and deep learning
By combining a SEN-MH-BiLSTM-based model with feature screening and deep learning methods, the accuracy and robustness issues in data center cooling energy consumption prediction were solved, achieving high-precision short-term prediction results.
Patent Information
- Application Number
- CN202510515431.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies have problems with insufficient accuracy and robustness in predicting data center cooling energy consumption. In particular, due to the time-varying and complexity of cooling energy consumption, traditional methods find it difficult to effectively capture its changing patterns.
A method based on feature screening, fusion and deep learning is adopted to make predictions through the SEN-MH-BiLSTM model, which includes the compression-excitation network SENet, the bidirectional long short-term memory network BiLSTM and the multi-head self-attention mechanism. It selects highly relevant feature data and performs weighted learning to capture the historical time series characteristics and correlation information of energy consumption data.
The prediction accuracy of data center cooling energy consumption and the generalization ability of the model are improved, which can better explore the dependency between cooling energy consumption and feature sequences and achieve high-precision short-term prediction.
Smart Images

Figure CN120671490A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of refrigeration energy consumption prediction, and specifically to a short-term prediction method for data center refrigeration energy consumption based on feature screening-fusion and deep learning. Background Art
[0002] With the deep integration of emerging technologies such as cloud computing, 5G, artificial intelligence, and big data, data has become a key engine driving high-quality development across various industries. Data centers, as core facilities for information storage and processing, have experienced rapid development. However, with the rapid growth in the scale and number of data centers, their energy consumption is becoming increasingly severe. Statistics show that data center energy consumption accounts for approximately 2% of global electricity consumption, and is growing at a rate of 15-20% annually. It is projected that by 2030, data center energy consumption will account for 8% of global electricity consumption. Therefore, efficient management and scheduling to promote the efficient utilization of data center resources has become a pressing challenge.
[0003] Data center cooling systems consume approximately 40% of their total energy consumption, second only to IT equipment. In this context, short-term forecasting of data center cooling energy consumption can promptly identify system anomalies and failures, providing a basis for decision-making on the proper allocation of electricity and a reference for the construction and development of new, green data centers.
[0004] Extensive research has been conducted on the short-term forecasting of data center cooling energy consumption. Traditional methods for short-term data center energy consumption forecasting primarily include time series analysis and regression analysis. However, these traditional methods require sufficient equipment parameters and operational monitoring data, are limited to analyzing the energy consumption of ultra-large data center infrastructure using air cooling, and lack generalizability. While these methods can reflect energy consumption patterns to a certain extent, data center cooling energy consumption is highly time-varying and complex. Traditional energy consumption forecasting methods often perform poorly in dealing with this complexity, resulting in poor prediction accuracy and robustness.
[0005] In recent years, with the development of artificial intelligence (AI) technology, machine learning and deep learning-based methods have gradually been applied to energy consumption forecasting. AI-based prediction technology for data center energy consumption has attracted widespread attention. However, AI-based prediction technology has the following shortcomings:
[0006] (1) Data center cooling energy consumption is highly time-varying and complex. Traditional energy consumption prediction methods often perform poorly in dealing with this complexity, resulting in poor actual prediction accuracy and robustness.
[0007] (2) The existing data center cooling energy consumption prediction model does not fully consider the input influencing factors and lacks structural design, resulting in low prediction efficiency. Summary of the Invention
[0008] The purpose of the present invention is to provide a method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, comprising the following steps:
[0009] 1) Obtain historical cooling energy consumption and historical cooling energy consumption influencing factors data of the data center; calculate the correlation between each cooling energy consumption influencing factor data and cooling energy consumption, extract cooling energy consumption influencing factor data with correlation greater than a preset threshold as cooling energy consumption feature data, and construct the input feature matrix X;
[0010] 2) Build a short-term prediction model for refrigeration energy consumption based on SEN-MH-BiLSTM, which includes a compression-excitation network SENet, a bidirectional long short-term memory network BiLSTM, and a multi-head self-attention mechanism;
[0011] 3) Using the input feature matrix X as input and the corresponding cooling energy consumption as output, a sample data set is constructed;
[0012] The sample data set is used to train the short-term prediction model of refrigeration energy consumption based on SEN-MH-BiLSTM, and the optimal model for short-term prediction of refrigeration energy consumption is obtained;
[0013] 4) Use the optimal short-term prediction model of cooling energy consumption to predict the cooling energy consumption of the data center at time t in the future.
[0014] Furthermore, the data on factors affecting refrigeration energy consumption include energy consumption data and meteorological data;
[0015] Energy consumption data includes IT energy consumption, auxiliary energy consumption data, and total energy consumption;
[0016] Meteorological data include indoor temperature, indoor humidity, outdoor temperature, outdoor humidity, and wind speed.
[0017] Furthermore, in step 1), the step of calculating the correlation between the data of each refrigeration energy consumption influencing factor and the refrigeration energy consumption includes:
[0018] 1.1) Normalize the historical refrigeration energy consumption and the factors affecting the historical refrigeration energy consumption;
[0019] The normalization formula is expressed as:
[0020]
[0021] Where: x * represents the normalized sample data, x represents the original sample data; x min Represents the minimum value of the sample data set; x max Indicates the maximum value of the sample data set.
[0022] 1.2) Calculate the Pearson correlation coefficient r and Spearman correlation coefficient ρ between the normalized data of factors affecting refrigeration energy consumption and refrigeration energy consumption;
[0023] Among them, the Pearson correlation coefficient r and the Spearman correlation coefficient ρ are as follows:
[0024]
[0025]
[0026] Where x and y are two feature objects; R(x) and R(y) are the ranks of x and y respectively; and are the average rankings of x and y respectively; x g and y g are the g-th observation values of x and y respectively; X g is the g-th observation value in the feature vector X; Y g is the g-th observation value in the feature vector Y; are the expected values of the eigenvectors X and Y respectively; G is the total number of samples;
[0027] 1.3) Remove the data of cooling energy consumption influencing factors whose absolute value of Pearson correlation coefficient r or Spearman correlation coefficient ρ is less than ε1; ε1 is the preset threshold;
[0028] 1.4) Calculate the comprehensive correlation of the remaining cooling energy consumption influencing factors, that is:
[0029]
[0030] Where r k ,p k denote the Pearson and Spearman correlation coefficients between the kth eigenvector and the cooling energy consumption, respectively, w k represents the comprehensive correlation between the kth eigenvector and the cooling energy consumption; r j ,p j denote the Pearson and Spearman correlation coefficients between the jth eigenvector and the cooling energy consumption, respectively;
[0031] 1.5) The cooling energy consumption influencing factor data with a comprehensive correlation greater than ε2 is used as the cooling energy consumption characteristic data; ε2 is a preset threshold.
[0032] Furthermore, in step 1), the input feature matrix X is as follows:
[0033]
[0034] Where: x k =[x k1 x k2 ,xkn ] is the input matrix of the kth input feature at n moments.
[0035] Furthermore, the compression-excitation network SENet includes a compression module, an excitation module, and a recalibration module;
[0036] The compression module performs global average pooling on the input feature matrix X, compressing the multi-dimensional data into a single value, and obtaining:
[0037]
[0038] Where, F sq is the compression module function; z c is the compressed output; X c is the c-th two-dimensional input matrix of input matrix X, where k×n is the spatial dimension of the input data.
[0039] The excitation module generates the corresponding feature weights for the compressed input feature matrix X, namely:
[0040]
[0041] Where σ1, σ2 represent the parameters of the fully connected layer in the excitation module; RELU is the activation function; δ represents the Sigmoid function; s c represents the feature weight generated by the excitation module; σ is the excitation module parameter; F ex is the excitation module function;
[0042] The recalibration module performs a recalibration operation on the data to obtain the output of the compression-excitation network SENet, namely:
[0043]
[0044] Where: X n is the final output feature matrix; s c is the feature weight value.
[0045] Furthermore, the bidirectional long short-term memory network BiLSTM includes a forward LSTM layer and a reverse LSTM layer;
[0046] The input of the forward LSTM layer is the output X of the compression-excitation network SENet. n ;
[0047] The output of the forward LSTM layer As shown below:
[0048]
[0049] Where: b fn 、bin 、b Cn 、b on Represents the bias coefficient of each gate of the LSTM unit; W fn 、W in 、W Cn 、W on Represents the weight coefficient of each gate; tanh and σ represent two different activation functions; LSTM is the calculation process of the traditional LSTM unit, which is used to update the hidden state; represents the hidden state of the forward LSTM at time t; C n (t) is the output of the input gate, f n (t) is the output of the forget gate, i n (t) represents the input of the input gate; The content obtained after updating the input data features; n (t) represents the intermediate output, h t is the final output of the output gate;
[0050] The output of the reverse LSTM layer is shown below:
[0051]
[0052] Where, Represents the hidden state of the backward LSTM at time t.
[0053] Furthermore, the input of the bidirectional long short-term memory network BiLSTM is the output X of the compression-excitation network SENet n , the output is the linear superposition result of the output of the forward LSTM layer and the backward LSTM layer;
[0054] The output of the bidirectional long short-term memory network BiLSTM is as follows:
[0055]
[0056] Where W y is the weight matrix used to linearly combine the hidden layer states; b y is the bias term; Y t Represents the output of BiLSTM.
[0057] Furthermore, the input sequence of the multi-head attention mechanism is Y t =[y1,…,y t ], the output sequence is H=[h1,…,h t ]; H represents the cooling energy consumption prediction output at time t in the future;
[0058] Among them, the output sequence H is as follows:
[0059]
[0060] Where [q1,…,qM] represents multiple query vectors, ⊕ represents vector concatenation; W q ,W k ,W v q i ,k i ,v i The linear mapping weights of , Q is the query matrix, K is the key matrix, and V is the value matrix. dk is the feature dimension, which is normalized to the interval [0,1] by softmax.
[0061] Furthermore, in step 3), the steps of training and testing the SEN-MH-BiLSTM-based refrigeration energy consumption short-term prediction model include:
[0062] The sample data set is used to train the short-term prediction model of refrigeration energy consumption, and the parameters of the prediction model are adjusted according to the loss value obtained from the training, and the weight parameters of each layer of the prediction model are updated until the accuracy of the prediction model is greater than the preset threshold or the number of training times reaches the preset number, and the optimal model for short-term prediction of refrigeration energy consumption is output.
[0063] The technical effects of the present invention are undoubted, and the beneficial effects of the present invention are as follows:
[0064] (1) A combined weighted feature screening method is proposed to calculate the comprehensive correlation between multiple energy consumption, meteorological input variables and cooling energy consumption variables in the data center under multiple evaluation systems. Through the multidimensional correlation analysis of the factors affecting the cooling energy consumption of the data center, the input features of the cooling energy consumption prediction model are efficiently screened.
[0065] (2) A short-term prediction model for data center cooling energy consumption based on SEN-MH-BiLSTM is proposed. The weighted learning of the filtered multi-dimensional input features is performed through the compression-excitation network (SENet), and the correlation information between different input features is integrated; then the bidirectional long short-term memory neural network (BiLSTM) is used to capture the historical time series characteristics of energy consumption data. Finally, the multi-head self-attention mechanism (MHSA) is introduced to block the input information and learn in parallel, which improves the representation and generalization capabilities of the model and makes high-precision predictions.
[0066] The present invention can better explore the dependency relationship between data center cooling energy consumption and characteristic sequences, and effectively improve the prediction accuracy of data center cooling energy consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 is a flow chart of the method;
[0068] Figure 2 is the prediction error under different feature selection schemes;
[0069] Figure 3 To compare the effects of different prediction models; Figure 3 (a)-(c) are the comparison of the 28-day prediction results of the test set, the magnified graph of some peaks in the test set, and the magnified graph of some troughs in the test set, respectively. DETAILED DESCRIPTION
[0070] The present invention will be further described below with reference to the following examples, but it should not be understood that the scope of the present invention is limited to the following examples. Without departing from the above technical ideas of the present invention, various substitutions and modifications can be made according to common technical knowledge and customary means in the art, and all should be included in the scope of protection of the present invention.
[0071] Example 1:
[0072] See also Figure 1-Figure 3 A short-term prediction method for data center cooling energy consumption based on feature screening-fusion and deep learning includes the following steps:
[0073] 1) Obtain historical cooling energy consumption and historical cooling energy consumption influencing factors data of the data center; calculate the correlation between each cooling energy consumption influencing factor data and cooling energy consumption, extract cooling energy consumption influencing factor data with correlation greater than a preset threshold as cooling energy consumption feature data, and construct the input feature matrix X;
[0074] 2) Build a short-term prediction model for refrigeration energy consumption based on SEN-MH-BiLSTM, which includes a compression-excitation network SENet, a bidirectional long short-term memory network BiLSTM, and a multi-head self-attention mechanism;
[0075] 3) Using the input feature matrix X as input and the corresponding cooling energy consumption as output, a sample data set is constructed;
[0076] The sample data set is used to train the short-term prediction model of refrigeration energy consumption based on SEN-MH-BiLSTM, and the optimal model for short-term prediction of refrigeration energy consumption is obtained;
[0077] 4) Use the optimal short-term prediction model of cooling energy consumption to predict the cooling energy consumption of the data center at time t in the future.
[0078] Data on factors affecting refrigeration energy consumption include energy consumption data and meteorological data;
[0079] Energy consumption data includes IT energy consumption, auxiliary energy consumption data, and total energy consumption;
[0080] Meteorological data include indoor temperature, indoor humidity, outdoor temperature, outdoor humidity, and wind speed.
[0081] In step 1), the step of calculating the correlation between the data of each refrigeration energy consumption influencing factor and the refrigeration energy consumption includes:
[0082] 1.1) Normalize the historical refrigeration energy consumption and the factors affecting the historical refrigeration energy consumption;
[0083] The normalization formula is expressed as:
[0084]
[0085] Where: x * represents the normalized sample data, x represents the original sample data; x min Represents the minimum value of the sample data set; x max Indicates the maximum value of the sample data set.
[0086] 1.2) Calculate the Pearson correlation coefficient r and Spearman correlation coefficient ρ between the normalized data of factors affecting refrigeration energy consumption and refrigeration energy consumption;
[0087] Among them, the Pearson correlation coefficient r and the Spearman correlation coefficient ρ are as follows:
[0088]
[0089]
[0090] Where x and y are two feature objects; R(x) and R(y) are the ranks of x and y respectively; and are the average rankings of x and y respectively; x g and y g are the g-th observation values of x and y respectively; X g is the g-th observation value in the feature vector X; Y g is the g-th observation value in the feature vector Y; are the expected values of the eigenvectors X and Y respectively; G is the total number of samples;
[0091] 1.3) Remove the data of cooling energy consumption influencing factors whose absolute value of Pearson correlation coefficient r or Spearman correlation coefficient ρ is less than ε1; ε1 is the preset threshold;
[0092] 1.4) Calculate the comprehensive correlation of the remaining cooling energy consumption influencing factors, that is:
[0093]
[0094] Where r k ,p k denote the Pearson and Spearman correlation coefficients between the kth eigenvector and the cooling energy consumption, respectively, w k represents the comprehensive correlation between the kth eigenvector and the cooling energy consumption; r j ,pj denote the Pearson and Spearman correlation coefficients between the jth eigenvector and the cooling energy consumption, respectively;
[0095] 1.5) The cooling energy consumption influencing factor data with a comprehensive correlation greater than ε2 is used as the cooling energy consumption characteristic data; ε2 is a preset threshold.
[0096] In step 1), the input feature matrix X is as follows:
[0097]
[0098] Where: x k =[x k1 x k2 ,x kn ] is the input matrix of the kth input feature at n moments.
[0099] The compression-excitation network SENet includes a compression module, an excitation module, and a recalibration module;
[0100] The compression module performs global average pooling on the input feature matrix X, compressing the multi-dimensional data into a single value, and obtaining:
[0101]
[0102] Where, F sq is the compression module function; z c is the compressed output; X c is the c-th two-dimensional input matrix of input matrix X, where k×n is the spatial dimension of the input data.
[0103] The excitation module generates the corresponding feature weights for the compressed input feature matrix X, namely:
[0104]
[0105] Where σ1, σ2 represent the parameters of the fully connected layer in the excitation module; RELU is the activation function; δ represents the Sigmoid function; s c represents the feature weight generated by the excitation module; σ is the excitation module parameter; F ex is the excitation module function;
[0106] The recalibration module performs a recalibration operation on the data to obtain the output of the compression-excitation network SENet, namely:
[0107]
[0108] Where: X n is the final output feature matrix; s c is the feature weight value.
[0109] The bidirectional long short-term memory network BiLSTM includes a forward LSTM layer and a reverse LSTM layer;
[0110] The input of the forward LSTM layer is the output X of the compression-excitation network SENet. n ;
[0111] The output of the forward LSTM layer As shown below:
[0112]
[0113] Where: b fn 、b in 、b Cn 、b on Represents the bias coefficient of each gate of the LSTM unit; W fn 、W in 、W Cn 、W on Represents the weight coefficient of each gate; tanh and σ represent two different activation functions; LSTM is the calculation process of the traditional LSTM unit, which is used to update the hidden state; represents the hidden state of the forward LSTM at time t; C n (t) is the output of the input gate, f n (t) is the output of the forget gate, i n (t) represents the input of the input gate; The content obtained after updating the input data features; n (t) represents the intermediate output, h t is the final output of the output gate;
[0114] The output of the reverse LSTM layer is shown below:
[0115]
[0116] Where, Represents the hidden state of the backward LSTM at time t.
[0117] The input of the bidirectional long short-term memory network BiLSTM is the output X of the compression-excitation network SENet n , the output is the linear superposition result of the output of the forward LSTM layer and the backward LSTM layer;
[0118] The output of the bidirectional long short-term memory network BiLSTM is as follows:
[0119]
[0120] Where W yis the weight matrix used to linearly combine the hidden layer states; b y is the bias term; Y t Represents the output of BiLSTM.
[0121] The input sequence of the multi-head attention mechanism is Y t =[y1,…,y t ], the output sequence is H=[h1,…,h t ]; H represents the cooling energy consumption prediction output at time t in the future;
[0122] Among them, the output sequence H is as follows:
[0123]
[0124] Where [q1,…,qM] represents multiple query vectors, ⊕ represents vector concatenation; W q ,W k ,W v q i ,k i ,v i The linear mapping weights of , Q is the query matrix, K is the key matrix, and V is the value matrix. dk is the feature dimension, which is normalized to the interval [0,1] by softmax.
[0125] In step 3), the steps of training and testing the SEN-MH-BiLSTM-based refrigeration energy consumption short-term prediction model include:
[0126] The sample data set is used to train the short-term prediction model of refrigeration energy consumption, and the parameters of the prediction model are adjusted according to the loss value obtained from the training, and the weight parameters of each layer of the prediction model are updated until the accuracy of the prediction model is greater than the preset threshold or the number of training times reaches the preset number, and the optimal model for short-term prediction of refrigeration energy consumption is output.
[0127] Example 2:
[0128] A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, comprising the following steps:
[0129] 1) Obtain historical cooling energy consumption and historical cooling energy consumption influencing factors data of the data center; calculate the correlation between each cooling energy consumption influencing factor data and cooling energy consumption, extract cooling energy consumption influencing factor data with correlation greater than a preset threshold as cooling energy consumption feature data, and construct the input feature matrix X;
[0130] 2) Build a short-term prediction model for refrigeration energy consumption based on SEN-MH-BiLSTM, which includes a compression-excitation network SENet, a bidirectional long short-term memory network BiLSTM, and a multi-head self-attention mechanism;
[0131] 3) Using the input feature matrix X as input and the corresponding cooling energy consumption as output, a sample data set is constructed;
[0132] The sample data set is used to train the short-term prediction model of refrigeration energy consumption based on SEN-MH-BiLSTM, and the optimal model for short-term prediction of refrigeration energy consumption is obtained;
[0133] 4) Use the optimal short-term prediction model of cooling energy consumption to predict the cooling energy consumption of the data center at time t in the future.
[0134] Example 3:
[0135] A method for short-term prediction of data center cooling energy consumption based on feature screening, fusion, and deep learning. The technical content is the same as that of Example 2. Furthermore, the data on factors affecting cooling energy consumption include energy consumption data and meteorological data.
[0136] Energy consumption data includes IT energy consumption, auxiliary energy consumption data, and total energy consumption;
[0137] Meteorological data include indoor temperature, indoor humidity, outdoor temperature, outdoor humidity, and wind speed.
[0138] Example 4:
[0139] A method for short-term prediction of data center cooling energy consumption based on feature screening, fusion, and deep learning, the technical content of which is the same as any one of Examples 2-3, further comprising the step of calculating the correlation between the data of each cooling energy consumption influencing factor and the cooling energy consumption in step 1) comprising:
[0140] 1.1) Normalize the historical refrigeration energy consumption and the factors affecting the historical refrigeration energy consumption;
[0141] The normalization formula is expressed as:
[0142]
[0143] Where: x * represents the normalized sample data, x represents the original sample data; x min Represents the minimum value of the sample data set; x max Indicates the maximum value of the sample data set.
[0144] 1.2) Calculate the Pearson correlation coefficient r and Spearman correlation coefficient ρ between the normalized data of factors affecting refrigeration energy consumption and refrigeration energy consumption;
[0145] Among them, the Pearson correlation coefficient r and the Spearman correlation coefficient ρ are as follows:
[0146]
[0147] Where x and y are two feature objects; R(x) and R(y) are the ranks of x and y respectively; and are the average rankings of x and y respectively; x g and y g are the g-th observation values of x and y respectively; X g is the g-th observation value in the feature vector X; Y g is the g-th observation value in the feature vector Y; are the expected values of the eigenvectors X and Y respectively; G is the total number of samples;
[0148] 1.3) Remove the data of cooling energy consumption influencing factors whose absolute value of Pearson correlation coefficient r or Spearman correlation coefficient ρ is less than ε1; ε1 is the preset threshold;
[0149] 1.4) Calculate the comprehensive correlation of the remaining cooling energy consumption influencing factors, that is:
[0150]
[0151] Where r k ,p k denote the Pearson and Spearman correlation coefficients between the kth eigenvector and the cooling energy consumption, respectively, w k represents the comprehensive correlation between the kth eigenvector and cooling energy consumption;
[0152] 1.5) The cooling energy consumption influencing factor data with a comprehensive correlation greater than ε2 is used as the cooling energy consumption characteristic data; ε2 is a preset threshold.
[0153] Example 5:
[0154] A method for short-term prediction of data center cooling energy consumption based on feature screening, fusion, and deep learning, the technical content of which is the same as any one of Examples 2-4. Further, in step 1), the input feature matrix X is as follows:
[0155]
[0156] Where: x k =[x k1 x k2 ,x kn ] is the input matrix of the kth input feature at n moments.
[0157] Example 6:
[0158] A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, with the same technical content as any one of Examples 2-5, further comprising a compression-excitation network SENet comprising a compression module, an excitation module, and a recalibration module;
[0159] The compression module performs global average pooling on the input feature matrix X, compressing the multi-dimensional data into a single value, and obtaining:
[0160]
[0161] Where, F sq is the compression module function; z c is the compressed output; X c is the c-th two-dimensional input matrix of input matrix X, where k×n is the spatial dimension of the input data.
[0162] The excitation module generates the corresponding feature weights for the compressed input feature matrix X, namely:
[0163]
[0164] Where σ1, σ2 represent the parameters of the fully connected layer in the excitation module; RELU is the activation function; δ represents the Sigmoid function; s c Represents the feature weights generated by the excitation module;
[0165] The recalibration module performs a recalibration operation on the data to obtain the output of the compression-excitation network SENet, namely:
[0166]
[0167] Where: X n is the final output feature matrix; s c is the feature weight value.
[0168] Example 7:
[0169] A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, with the same technical content as any one of Examples 2-6, further comprising a bidirectional long short-term memory network BiLSTM including a forward LSTM layer and a reverse LSTM layer;
[0170] The input of the forward LSTM layer is the output X of the compression-excitation network SENet. n ;
[0171] The output of the forward LSTM layer As shown below:
[0172]
[0173] Where: b fn 、b in 、b Cn 、b on Represents the bias coefficient of each gate of the LSTM unit; Wfn 、W in 、W Cn 、W on Represents the weight coefficient of each gate; tanh and σ represent two different activation functions; LSTM is the calculation process of the traditional LSTM unit, which is used to update the hidden state; represents the hidden state of the forward LSTM at time t;
[0174] The output of the reverse LSTM layer is shown below:
[0175]
[0176] Where, Represents the hidden state of the backward LSTM at time t.
[0177] Example 8:
[0178] A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, the technical content of which is the same as any one of Examples 2-7, further, the input of the bidirectional long short-term memory network BiLSTM is the output X of the compression-excitation network SENet n , the output is the linear superposition result of the output of the forward LSTM layer and the backward LSTM layer;
[0179] The output of the bidirectional long short-term memory network BiLSTM is as follows:
[0180]
[0181] Where W y is the weight matrix used to linearly combine the hidden layer states; b y is the bias term; Y t Represents the output of BiLSTM.
[0182] Example 9:
[0183] A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, the technical content is the same as any one of Examples 2-8, further, the input sequence of the multi-head attention mechanism is Y t =[y1,…,y t ], the output sequence is H=[h1,…,h t ]; H represents the cooling energy consumption prediction output at time t in the future;
[0184] Among them, the output sequence H is as follows:
[0185]
[0186] Where [q1,…,qM] represents multiple query vectors, ⊕ represents vector concatenation; W q ,W k ,W v q i ,k i ,v i The linear mapping weights of , Q is the query matrix, K is the key matrix, and V is the value matrix. dk is the feature dimension, which is normalized to the interval [0,1] by softmax.
[0187] Example 10:
[0188] A method for short-term prediction of data center cooling energy consumption based on feature screening, fusion, and deep learning, the technical content of which is the same as any one of Examples 2-9, further, in step 3), the steps of training and testing the short-term cooling energy consumption prediction model based on SEN-MH-BiLSTM include:
[0189] The sample data set is used to train the short-term prediction model of refrigeration energy consumption, and the parameters of the prediction model are adjusted according to the loss value obtained from the training, and the weight parameters of each layer of the prediction model are updated until the accuracy of the prediction model is greater than the preset threshold or the number of training times reaches the preset number, and the optimal model for short-term prediction of refrigeration energy consumption is output.
[0190] Example 11:
[0191] A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning, comprising the following steps:
[0192] Step S1: Collect the data center's cooling energy consumption, IT energy consumption, auxiliary energy consumption data on an hourly time scale, as well as meteorological data such as indoor and outdoor temperature, humidity, and wind speed as the original data set. Through the weighted comprehensive evaluation system of the Pearson correlation coefficient and the Spearman correlation coefficient, screen out features that are highly correlated with the data center's cooling energy consumption as input variables of the prediction model to avoid the impact of redundant information on prediction accuracy.
[0193] Step S2: Introduce the compression-excitation network (SENet) to learn the correlation information between the data center cooling energy consumption output and multi-dimensional input features and generate weights between each feature vector, so as to achieve multi-dimensional input feature fusion and improve the feature extraction capability of the prediction model.
[0194] Step S3: Based on the SENet network output, a bidirectional long short-term memory network (BiLSTM) is introduced to learn the temporal dependency between cooling energy consumption and historical input data, and efficiently extract the bidirectional time series features of cooling energy consumption sequence data to improve the prediction accuracy of data center cooling energy consumption.
[0195] Step S4: A multi-head self-attention mechanism (MHSA) is introduced after the BiLSTM output. This mechanism dynamically captures the relevance and important information of different locations globally. This multi-head parallel mechanism simultaneously learns the attention of different features, improving the model's representational and generalization capabilities. Finally, a short-term prediction output for cooling energy consumption is obtained.
[0196] Step S5: Multiple groups of training samples of the training set are sequentially input into the SEN-MH-BiLSTM refrigeration energy consumption short-term prediction model, and the prediction model is repeatedly iteratively trained using the training set data until the accuracy meets the requirements or the number of training times reaches the preset number, thereby obtaining the final prediction model.
[0197] The characteristics of step S1 are as follows:
[0198] First, we collected the original feature data set of the data center, including historical energy consumption data (cooling energy consumption, IT equipment energy consumption, auxiliary equipment energy consumption, and total energy consumption) and historical meteorological data (indoor temperature, indoor humidity, outdoor temperature, outdoor humidity, and wind speed). To eliminate the impact of dimensional differences in different data, we normalized the original energy consumption and meteorological data. The normalization formula is expressed as:
[0199]
[0200] Where: x * represents the normalized sample data, x represents the original sample data; x min Represents the minimum value of the sample data set; x max Indicates the maximum value of the sample data set.
[0201] Cooling energy consumption is influenced by multiple factors both inside and outside the data center. Therefore, the appropriate selection of input feature data is crucial for prediction accuracy. This paper proposes a combined weighted feature screening method. By combining the Pearson correlation coefficient and the Spearman correlation coefficient, this method comprehensively evaluates the correlation between cooling energy consumption and multiple factors inside and outside the data center, avoiding the situation where a single screening metric overlooks key features.
[0202] The Pearson correlation coefficient and Spearman correlation coefficient between the eigenvector and the target sequence (cooling energy consumption) are calculated respectively to characterize the degree of correlation between different eigenvectors and cooling energy consumption under different evaluation systems. The calculation formula of the Pearson correlation coefficient is:
[0203]
[0204] Where, X g is the g-th observation value in the feature vector X; Y g is the g-th observation value in the feature vector Y; are the expected values of the eigenvectors X and Y, respectively; G is the total number of samples. The larger the absolute value of the Pearson correlation coefficient r, the stronger the linear correlation between X and Y.
[0205] The Spearman correlation coefficient is used to evaluate the monotonic relationship between two variables. The larger the absolute value, the higher the correlation. The calculation formula of the Spearman correlation coefficient ρ is as follows:
[0206]
[0207] Where x and y are two feature objects; R(x) and R(y) are the ranks of x and y respectively; and are the average rankings of x and y respectively; x g and y g are the g-th observation values of x and y respectively.
[0208] According to the calculation results, weakly correlated feature vectors (with absolute values of correlation coefficients less than 0.3) under the two correlation evaluation indicators are selected and removed to avoid the impact of redundant feature input on the prediction accuracy of refrigeration energy consumption.
[0209] Then, the different weights of the remaining input feature variables in the two evaluation index coefficients are calculated, and the calculated different weight proportion combinations are added together to finally obtain the comprehensive correlation of different feature vectors. The calculation process is as follows:
[0210]
[0211] Where r k ,p k denote the Pearson and Spearman correlation coefficients between the kth eigenvector and the cooling energy consumption, respectively, w k Represents the comprehensive correlation between the kth eigenvector and cooling energy consumption.
[0212] Finally, feature vectors with a comprehensive correlation greater than 0.8 are selected as the final input features of the cooling energy consumption prediction model. An input feature matrix X of size k×n is constructed, with a total of k features and a sliding window width of n. The input feature matrix X of the final prediction model can be expressed as:
[0213]
[0214] Where: x k =[x k1 x k2 ,x kn ] is the input matrix of the kth input feature at n moments, and the dataset is divided into training set, validation set and test set in a ratio of 8:1:1
[0215] The characteristics of step S2 are as follows:
[0216] In traditional neural networks, each eigenvector obtained after convolution is typically processed identically, assigning the same weight. However, in reality, each eigenvector extracts different information and importance after convolution. The squeeze and excitation network (SENet) uses a compression module and an excitation module to learn the importance parameters of each eigenvector, generating weights between each eigenvector, thereby extracting and fusing important information from multiple features.
[0217] First, the input feature matrix X is subjected to global average pooling in the compression module to compress the multidimensional data into a single value. That is, for each time step t, a vector of length k is output, where each element represents the mean of the feature vector. The calculation formula is:
[0218]
[0219] Where, F sq is the compression module function; z c is the compressed output; X c is the c-th two-dimensional input matrix of input matrix X, where k×n is the spatial dimension of the input data.
[0220] Then, the compressed data is imported into the excitation module to learn the nonlinear relationship between the feature vectors and generate the corresponding feature weight for each compressed feature. The calculation formula is:
[0221]
[0222] Where σ1, σ2 represent the parameters of the fully connected layer in the excitation module; RELU is the activation function; δ represents the Sigmoid function; s c Represents the feature weights generated by the excitation module.
[0223] Finally, the data is recalibrated. The feature weights generated by the excitation module are weighted one by one to the original feature vector in a dot product manner to achieve the recalibration of the original feature vector. The weighted output calculation formula is:
[0224]
[0225] Where: X n is the final output feature matrix; s c is the feature weight value, and its value range is [0,1].
[0226] The characteristics of step S3 are as follows:
[0227] While SENet can extract correlation information from multi-dimensional feature vectors, it has limited processing capabilities for time-dependent time series data. To address this issue, this paper introduces a bidirectional long-short-term memory (BiLSTM) network to efficiently mine the inherent connections between current and past data, improving the accuracy of data center cooling energy consumption prediction. The calculation process is as follows:
[0228] First, the output X of the SENet network is n As the forward input data of the BiLSTM network, it is input to the forward LSTM layer to calculate the hidden state sequence h1,h2,…,h from t=1 to t t , get the output of the forward LSTM layer The calculation formula is:
[0229]
[0230] Where: b fn 、b in 、b Cn 、b on Represents the bias coefficient of each gate of the LSTM unit; W fn 、W in 、W Cn 、W on Represents the weight coefficient of each gate; tanh and σ represent two different activation functions; LSTM is the calculation process of the traditional LSTM unit, which is used to update the hidden state; represents the hidden state of the forward LSTM at time t.
[0231] Then input the data in reverse to the reverse LSTM layer, and after getting the reverse output, reverse the output again to get the output of the reverse LSTM layer The calculation formula is as follows:
[0232]
[0233] Where, Represents the hidden state of the backward LSTM at time t.
[0234] Finally, the output of the forward LSTM layer and the output of the reverse LSTM layer are linearly superimposed according to certain weights to obtain the final output result.
[0235]
[0236] Where W y is the weight matrix used to linearly combine the hidden layer states; b y is the bias term; Y t Represents the output of BiLSTM.
[0237] The characteristics of step S4 are as follows:
[0238] Compared with the traditional attention mechanism, the multi-head self-attention mechanism (MHSA) can learn different attention representations in different subspaces in parallel and fuse them, which can better capture data information. Therefore, the multi-head self-attention mechanism is introduced after BiLSTM to more effectively distribute the weights of feature data. The input sequence Y of MHSA is t =[y1,…,y t ], the output sequence is H=[h1,…,h t ]. First, each input y i It is linearly mapped to three different spaces, and the query vector q is obtained by mapping calculation. i , key vector k i and the value vector v i , the mapping calculation formula is:
[0239]
[0240] Where W q ,W k ,W v q i ,k i ,v i The linear mapping weights of , Q is the query matrix, K is the key matrix, and V is the value matrix.
[0241] Then, the scaled dot product is used to calculate the similarity function, and the attention output function is obtained as:
[0242]
[0243] Where dk is the feature dimension, which is normalized to the interval [0, 1] by softmax.
[0244] Finally, a multi-head attention mechanism is introduced to simultaneously select multiple groups of information from the input information using multiple query vectors q, and then concatenate the multiple groups of vectors together to obtain the final cooling energy consumption prediction output H through linear transformation. The specific calculation process is as follows:
[0245]
[0246] Where [q1,…,qM] represents multiple query vectors and ⊕ represents vector concatenation.
[0247] The characteristics of step S5 are as follows:
[0248] The training samples of the training set are input into the SEN-MH-BiLSTM prediction model in sequence. The prediction model is trained using multiple groups of training samples. The parameters of the prediction model are adjusted according to the loss values obtained from the training. The weight parameters of each layer of the prediction model are updated. The training is iterated repeatedly until the model accuracy meets the requirements or the number of training times reaches the preset number, and the final prediction model is obtained.
[0249] Example 12:
[0250] A validation of a short-term prediction method for data center cooling energy consumption based on feature selection-fusion and deep learning is as follows:
[0251] To validate the effectiveness of the proposed short-term data center cooling energy consumption forecasting method, we used historical data from a data center in northwest China for experimental verification. The dataset spanned 290 days (6,960 hours) from January 1 to September 17, 2023, and short-term energy consumption forecasts were performed with a 1-hour time step. Comparative analysis between different forecasting models and the proposed model demonstrated the effectiveness and superiority of the proposed method.
[0252] Calculation effect:
[0253] (1) In order to explore the impact of input feature screening on model prediction accuracy, four different prediction model input feature selection schemes are set, as shown in Table 1. Among them, x1 is cooling energy consumption, x2 is total energy consumption, x3 is IT equipment energy consumption, x4 is other equipment energy consumption, x5 is indoor temperature, x6 is outdoor temperature, x7 is outdoor humidity, x8 is atmospheric pressure, and x9 is wind speed.
[0254] Table 1 Prediction model input feature selection scheme
[0255]
[0256] The above four feature selection schemes are used as input features for the prediction model screening. The overall prediction errors corresponding to the four schemes are as follows: Figure 2 shown.
[0257] Scheme 1, which uses only cooling energy consumption as input to the prediction model, yields the highest average error of the four schemes, at 5.0044%. Scheme 2, which uses all data center power characteristics (including cooling energy consumption, IT equipment energy consumption, other energy consumption, and total energy consumption) as input, yields an average prediction error of 4.3457%, a 13.16% reduction compared to Scheme 1, validating the effectiveness of using power characteristics in this paper. Scheme 3, which uses all power and meteorological time series characteristics as input, yields an average prediction error of 3.6425%, a 16.18% reduction compared to Scheme 2, validating the effectiveness of using meteorological time series characteristics in this paper. Scheme 4, which selects features based on the comprehensive correlation coefficient as input to the prediction model, yields the lowest average prediction error of 3.4421%. Compared to Scheme 3, which uses all power and time series characteristics, this results in a 5.50% reduction in prediction error, validating the effectiveness of using comprehensive correlation analysis to select input features.
[0258] (2) To further illustrate the advantages of the proposed model in the short-term prediction of data center cooling energy consumption, four deep learning models, LSTM, BiLSTM, SENet-LSTM, and SENet-BiLSTM, were selected as comparison models. The LSTM and BiLSTM layers in SEN-LSTM and SEN-BiLSTM have the same parameter settings as the LSTM and BiLSTM models alone. Other parameter settings, such as the number of iterations, training batch size, optimization algorithm, and loss function, for all comparison models are the same as those for SEN-MH-BiLSTM.
[0259] Figure 3 The following graphs plot the actual and predicted energy consumption values for each model in the test set. Compared to the four deep learning models, SEN-MH-BiLSTM demonstrates significant advantages in prediction accuracy. The proposed SEN-MH-BiLSTM model achieves 22.59%, 10.94%, and 11.20% lower prediction errors, respectively, than the averages of all compared models. This demonstrates that the proposed prediction model can better exploit the dependency between data center cooling energy consumption and feature sequences, effectively improving the prediction accuracy of data center cooling energy consumption.
Claims
1. A short-term prediction method for data center cooling energy consumption based on feature screening-fusion and deep learning, characterized by: The following steps are involved: 1) Obtain historical cooling energy consumption and historical cooling energy consumption influencing factors data of the data center; calculate the correlation between each cooling energy consumption influencing factor data and cooling energy consumption, extract cooling energy consumption influencing factor data with correlation greater than a preset threshold as cooling energy consumption feature data, and construct the input feature matrix X; 2) Build a short-term prediction model for refrigeration energy consumption based on SEN-MH-BiLSTM, which includes a compression-excitation network SENet, a bidirectional long short-term memory network BiLSTM, and a multi-head self-attention mechanism; 3) Using the input feature matrix X as input and the corresponding cooling energy consumption as output, a sample data set is constructed. The sample data set is used to train the short-term prediction model of refrigeration energy consumption based on SEN-MH-BiLSTM, and the optimal model for short-term prediction of refrigeration energy consumption is obtained; 4) Use the optimal short-term prediction model of cooling energy consumption to predict the cooling energy consumption of the data center at time t in the future.
2. A method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1, characterized in that: Data on factors affecting refrigeration energy consumption include energy consumption data and meteorological data; Energy consumption data includes IT energy consumption, auxiliary energy consumption data, and total energy consumption; Meteorological data include indoor temperature, indoor humidity, outdoor temperature, outdoor humidity, and wind speed.
3. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1 is characterized in that: In step 1), the step of calculating the correlation between the data of each refrigeration energy consumption influencing factor and the refrigeration energy consumption includes: 1.1) Normalize the historical refrigeration energy consumption and the factors affecting the historical refrigeration energy consumption; The normalization formula is expressed as: Where: x * represents the normalized sample data, x represents the original sample data; x min Represents the minimum value of the sample data set; x max Indicates the maximum value of the sample data set. 1.2) Calculate the Pearson correlation coefficient r and Spearman correlation coefficient ρ between the normalized data of factors affecting refrigeration energy consumption and refrigeration energy consumption; Among them, the Pearson correlation coefficient r and the Spearman correlation coefficient ρ are as follows: Where x and y are two feature objects; R(x) and R(y) are the ranks of x and y respectively; and are the average rankings of x and y respectively; x g and y g are the g-th observation values of x and y respectively; X g is the g-th observation value in the feature vector X; Y g is the g-th observation value in the feature vector Y; are the expected values of the eigenvectors X and Y respectively; G is the total number of samples; 1.3) Remove the data of cooling energy consumption influencing factors whose absolute value of Pearson correlation coefficient r or Spearman correlation coefficient ρ is less than ε1; ε1 is the preset threshold; 1.4) Calculate the comprehensive correlation of the remaining cooling energy consumption influencing factors, that is: Where r k ,p k denote the Pearson and Spearman correlation coefficients between the kth eigenvector and the cooling energy consumption, respectively, w k represents the comprehensive correlation between the kth eigenvector and the cooling energy consumption; r j ,p j denote the Pearson and Spearman correlation coefficients between the jth eigenvector and the cooling energy consumption, respectively; 1.5) The cooling energy consumption influencing factor data with a comprehensive correlation greater than ε2 is used as the cooling energy consumption characteristic data; ε2 is a preset threshold.
4. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1 is characterized in that: In step 1), the input feature matrix X is as follows: Where: x k =[x k1 x k2 ,…x kn ] is the input matrix of the kth input feature at n moments.
5. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1 is characterized in that: The compression-excitation network SENet includes a compression module, an excitation module, and a recalibration module; The compression module performs global average pooling on the input feature matrix X, compressing the multi-dimensional data into a single value, and obtaining: Where, F sq is the compression module function; z c is the compressed output; X c is the c-th two-dimensional input matrix of input matrix X, where k×n is the spatial dimension of the input data. The excitation module generates the corresponding feature weights for the compressed input feature matrix X, namely: Where σ1, σ2 represent the parameters of the fully connected layer in the excitation module; RELU is the activation function; δ represents the Sigmoid function; s c represents the feature weight generated by the excitation module; σ is the excitation module parameter; F ex is the excitation module function; The recalibration module performs a recalibration operation on the data to obtain the output of the compression-excitation network SENet, namely: Where: X n is the final output feature matrix; s c is the feature weight value; F scale It is the recalibration module function.
6. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1 is characterized in that: The bidirectional long short-term memory network BiLSTM includes a forward LSTM layer and a reverse LSTM layer; The input of the forward LSTM layer is the output X of the compression-excitation network SENet. n ; The output of the forward LSTM layer As shown below: Where: b fn 、b in 、b Cn 、b on Represents the bias coefficient of each gate of the LSTM unit; W fn 、W in 、W Cn 、W on Represents the weight coefficient of each gate; tanh and σ represent two different activation functions; LSTM is the calculation process of the traditional LSTM unit, which is used to update the hidden state; represents the hidden state of the forward LSTM at time t; C n (t) is the output of the input gate, f n (t) is the output of the forget gate, i n (t) represents the input of the input gate; The content obtained after updating the input data features; n (t) represents the intermediate output, h t is the final output of the output gate; The output of the reverse LSTM layer is shown below: Where, Represents the hidden state of the backward LSTM at time t.
7. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 6 is characterized in that: The input of the bidirectional long short-term memory network BiLSTM is the output X of the compression-excitation network SENet n , the output is the linear superposition result of the output of the forward LSTM layer and the backward LSTM layer; The output of the bidirectional long short-term memory network BiLSTM is as follows: Where W y is the weight matrix used to linearly combine the hidden layer states; b y is the bias term; Y t Represents the output of BiLSTM.
8. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1 is characterized in that: The input sequence of the multi-head attention mechanism is Y t =[y1,…,y t ], the output sequence is H=[h1,…,h t ]; H represents the cooling energy consumption prediction output at time t in the future; Among them, the output sequence H is as follows: Where [q1,…,qM] represents multiple query vectors, represents vector concatenation; W q ,W k ,W v q i ,k i ,v i The linear mapping weights of , Q is the query matrix, K is the key matrix, and V is the value matrix. dk is the feature dimension, which is normalized to the interval [0,1] by softmax; b y is the bias term.
9. The method for short-term prediction of data center cooling energy consumption based on feature screening-fusion and deep learning according to claim 1 is characterized in that: In step 3), the steps of training and testing the SEN-MH-BiLSTM-based refrigeration energy consumption short-term prediction model include: The sample data set is used to train the short-term prediction model of refrigeration energy consumption, and the parameters of the prediction model are adjusted according to the loss value obtained from the training, and the weight parameters of each layer of the prediction model are updated until the accuracy of the prediction model is greater than the preset threshold or the number of training times reaches the preset number, and the optimal model for short-term prediction of refrigeration energy consumption is output.
Citation Information
Cited By
Multi-dimensional data management method and system based on deep learning
CN121560961A
A multi-dimensional data management method and system based on deep learning
CN121560961B
Two-stage large-scale language model energy consumption analysis method and system facing reasoning task
CN122087759A
Two-stage large language model energy consumption analysis method and system for inference task
CN122087759B