Deep Learning-Based Hot Aisle Prediction Algorithm for Data Center Environments
By using a deep learning-based BiRCNN network model, combined with bidirectional GRU and CNN, the problem of inaccurate hot aisle temperature prediction in data center environmental monitoring systems is solved, achieving fast and accurate hot aisle temperature prediction and optimizing data center operation and maintenance and energy saving.
Patent Information
- Application Number
- CN202211456328.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-21
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2042-11-21
AI Technical Summary
Existing data center environmental monitoring systems cannot reflect the average temperature of the server's environmental hot aisle in a timely and accurate manner, making it impossible for maintenance personnel to effectively judge the server's status.
By employing a deep learning-based BiRCNN network model, combined with bidirectional GRU and CNN, the average temperature of the server's environmental thermal aisle is predicted through data acquisition, processing, and normalization.
It improves the accuracy and speed of hot aisle prediction in data center environments, reduces the workload of operations and maintenance personnel, and optimizes the energy efficiency of data centers.
Smart Images

Figure CN115718687B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hot aisle prediction algorithm technology for server environments, specifically a hot aisle prediction algorithm for data center environments based on deep learning. Background Technology
[0002] By controlling the hot and cold aisles of servers and maintaining appropriate temperature, humidity, and airflow, administrators must monitor these factors to ensure server efficiency and keep the data center running with minimal energy consumption.
[0003] Patent No. 202110604574.X discloses a data center room environment monitoring system, including M server racks, N air conditioning racks, a mobile inspection terminal, and n edge computing terminals arranged on the n movable server racks. M of the M server racks are movable within a predetermined range and spaced apart from the immovable racks. The mobile inspection terminal moves within the data center room environment along a preset inspection route, acquires the processing results from the n edge computing terminals to update the preset inspection route, and adjusts the operating status of some of the N air conditioning racks when passing through the updated preset inspection route. The edge processing terminal controls at least one of the n server racks to move within a predetermined range based on temperature and humidity data acquired by the combined sensors. This invention achieves full-range dynamic inspection and status control of the data center room.
[0004] The network data feature extraction capabilities of the aforementioned environmental monitoring system are not obvious, and the network computing speed is slow, which makes it impossible to reflect the average temperature of the server's environmental hot aisle in a timely and accurate manner. As a result, maintenance personnel cannot judge the status of the server. In view of this, it is necessary to provide a data center environmental hot aisle prediction algorithm based on deep learning. Summary of the Invention
[0005] The purpose of this invention is to provide a deep learning-based hot aisle prediction algorithm for data center environments to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: a deep learning-based hot aisle prediction algorithm for data center environments, comprising the following steps:
[0007] S1: Data Acquisition: Retrieve data from the corresponding server from the monitoring system;
[0008] S2: Data processing: The data obtained in S1 is processed into an average value and then normalized.
[0009] S3: Build the BiRCNN network model;
[0010] S4: Model Training: Input the data processed in S2 into the BiRCNN network model and use it as the target of the model output. Divide the data into training set and test set according to a certain ratio for model training.
[0011] S5: Predicted Data: Save the model training parameters, and input the new time data into the BiRCNN network model to predict the average temperature of the environmental thermal channel corresponding to the server.
[0012] Preferably, the specific method of S1 is as follows: obtain six types of data from the monitoring system: ambient cold aisle temperature, ambient hot aisle temperature, server power, precision air conditioner fan speed, precision air conditioner fan speed setting, and ambient hot aisle temperature after 15 minutes.
[0013] Preferably, the specific method of S2 is as follows: all acquired ambient cold aisle temperatures are processed into an average ambient cold aisle temperature, all acquired ambient hot aisle temperatures are processed into an average ambient hot aisle temperature, all acquired ambient hot aisle temperatures after 15 minutes are processed into an average ambient hot aisle temperature after 15 minutes, and the processed average ambient cold aisle temperature, average ambient hot aisle temperature, average ambient hot aisle temperature after 15 minutes, server power, precision air conditioner fan speed, and precision air conditioner fan speed are normalized.
[0014] Preferably, the BiRCNN network model in S3 includes: an input layer, a BiGRU layer, a convolutional layer, a fully connected layer, and an output layer.
[0015] Preferably, the data input into the input layer includes the average temperature of the ambient cold aisle, the average temperature of the ambient hot aisle, the server power, the speed of the precision air conditioner fan, and the speed of the precision air conditioner fan.
[0016] Preferably, the convolutional layer extracts key information from the time-series data through convolutional kernels.
[0017] Preferably, the fully connected layer further calculates and summarizes the extracted feature data.
[0018] Preferably, the output layer uses the sigmoid activation function to output the average temperature of the ambient thermal aisle corresponding to the server after 15 minutes.
[0019] Preferably, the specific method of S4 is as follows: the average temperature of the ambient cold aisle, the average temperature of the ambient hot aisle, the server power, the speed of the precision air conditioner fan, and the speed of the precision air conditioner fan are used as inputs to the BiRCNN network model, and the average temperature of the ambient hot aisle after 15 minutes is used as the target of the model output. The model is trained by dividing it into training set and test set according to a certain ratio.
[0020] Preferably, the specific method of S5 is as follows: save the model training parameters, input the ambient cold aisle temperature, ambient hot aisle temperature, server power, precision air conditioner fan speed and precision air conditioner fan speed settings into the model, and the average ambient hot aisle temperature corresponding to the server 15 minutes later can be predicted.
[0021] Compared with the prior art, the beneficial effects of the present invention are:
[0022] The BiRCNN network model built in this invention adopts a combination design of bidirectional GRU network and CNN network. While retaining the good predictive ability of bidirectional GRU for time series data, it adds convolutional layers of CNN to better extract key features in time series data. This helps to timely adjust parameters such as supply air set temperature, return air set temperature, compressor speed, and fan speed of precision air conditioners, optimize the energy-saving effect of data center AI group control algorithm, and reduce the pressure on operation and maintenance personnel. Attached Figure Description
[0023] Figure 1 This is a system diagram of the method of the present invention;
[0024] Figure 2 This is a flowchart of the BiRCNN network model of the present invention;
[0025] Figure 3 This is the GRU input / output structure of the present invention;
[0026] Figure 4 This is the basic structure of the GRU unit in this invention;
[0027] Figure 5 This is a basic structural diagram of the BiRNN of this invention;
[0028] Figure 6 This is a performance curve during the training process of this invention;
[0029] Figure 7 This is a schematic diagram of the CNN calculation process of the present invention;
[0030] Figure 8 This is a schematic diagram of the convolutional layer calculation process of the present invention;
[0031] Figure 9 This is a schematic diagram of the maximum pooling calculation process of the present invention;
[0032] Figure 10 This is a schematic diagram of the average pooling calculation process of the present invention;
[0033] Figure 11 This is a schematic diagram of the global average pooling strategy of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] Please see Figure 1-2 This invention provides a deep learning-based hot aisle prediction algorithm for data center environments, comprising the following steps:
[0036] S1: Data Acquisition: Retrieve data from the corresponding server from the monitoring system;
[0037] S2: Data processing: The data obtained in S1 is processed into an average value and then normalized.
[0038] S3: Build the BiRCNN network model;
[0039] S4: Model Training: Input the data processed in S2 into the BiRCNN network model and use it as the target of the model output. Divide the data into training set and test set according to a certain ratio for model training.
[0040] S5: Predicted Data: Save the model training parameters, and input the new time data into the BiRCNN network model to predict the average temperature of the environmental thermal channel corresponding to the server.
[0041] In this embodiment, the specific method of S1 is to obtain six types of data from the monitoring system: ambient cold aisle temperature, ambient hot aisle temperature, server power, precision air conditioner fan speed, precision air conditioner fan speed setting, and ambient hot aisle temperature after 15 minutes.
[0042] In this embodiment, the specific method of S2 is as follows: all the acquired ambient cold aisle temperatures are processed into an average ambient cold aisle temperature, all the acquired ambient hot aisle temperatures are processed into an average ambient hot aisle temperature, all the acquired ambient hot aisle temperatures after 15 minutes are processed into an average ambient hot aisle temperature after 15 minutes, and the processed average ambient cold aisle temperature, average ambient hot aisle temperature, average ambient hot aisle temperature after 15 minutes, server power, precision air conditioner fan speed, and precision air conditioner fan speed are normalized.
[0043] In this embodiment, the BiRCNN network model in S3 includes: an input layer, a BiGRU layer, a convolutional layer, a fully connected layer, and an output layer.
[0044] The Gated Recurrent Unit (GRU) is a variant of the Recurrent Neural Network (RNN). Its biggest difference from LSTM is that it reduces the three gating control structures to two without significantly reducing training efficiency and performance. Furthermore, because GRU has relatively fewer parameters, it can perform better in terms of training speed in some cases and reduces the risk of overfitting.
[0045] Traditional RNNs, due to the connections between neurons in each layer, are prone to gradient explosion or vanishing gradient problems when processing long-term data. The GRU algorithm, through gating design within neurons, controls the retention or forgetting of data within the neurons, thereby solving the gradient explosion or vanishing gradient problems of traditional RNNs to some extent.
[0046] The input and output structure of the GRU is as follows: Figure 3 As shown, it is almost identical to a typical RNN structure. The GRU unit receives information from two directions: the current input X. t The hidden state (H) passed in from the previous node, containing information about the previous node. t-1 For X t With H t-1 After processing, GRU will output Y of the currently hidden node. t The hidden state H passed to the next node t .
[0047] The following is a schematic diagram of a single GRU neuron. Figure 4 As shown, R t The gating mechanism acts as a combination of the forget gate and input gate in LSTM. The GRU unit includes two gating structures: the ResetGate and the UpdateGate. The values of the ResetGate and UpdateGate are controlled within the range [0, 1] by the Sigmoid activation function. A larger ResetGate signal means the neuron retains more information from the current data. A larger UpdateGate signal means the neuron retains more information from the previous input.
[0048] At time step t, the specific calculation steps of the GRU unit are as follows: First, the input of the GRU unit at time step t is the current input X. t and the output H of the previous unit t-1 The data entering the cell will first undergo the following calculations.
[0049] R t =σ(X) t α r +H t-1 β r +br )
[0050] Z t =σ(X) t α z +H t-1 β z +b z )
[0051] α r / z and β r / z b represents a vector of different weight coefficients. r / z Let σ represent different bias vectors, and σ represent the activation function, using the formula:
[0052]
[0053] Processed data R t Will with X t Entering the reset door will restore the hidden state H passed from the previous unit. t-1 A reset signal is obtained by performing a reset calculation. The larger the reset signal value, the more information the current cell retains from the previous cell; conversely, the smaller the reset signal value, the more information the current cell forgets from the previous cell. The calculation method is as follows.
[0054]
[0055] Where α h and β h b represents different weights h Indicates bias, σ * The activation function is expressed using the following formula:
[0056]
[0057] Processed data Z t Will and H t-1 and reset signal Enter the update gate to output the hidden state H of the current unit. t The calculation method is as follows:
[0058]
[0059] GRU and LSTM work similarly, effectively suppressing gradient vanishing or exploding when capturing semantic associations in long sequences, outperforming RNNs. However, GRU still cannot completely solve the gradient vanishing problem, and as a variant of RNN, it shares the same major drawback as the RNN structure itself: it cannot be computed in parallel.
[0060] Bidirectional Recurrent Neural Networks (Bi-RNNs) are an extension of unidirectional RNNs proposed by Schuster et al. RNNs and GRUs have good analytical capabilities for time series data, mainly due to the front-to-back connections between neurons within their neural network structures. However, due to their unidirectional design, both RNNs and GRUs can only utilize past data for analysis and prediction. Bidirectional RNNs, on the other hand, can utilize both past and future data for analysis and prediction.
[0061] The structure of Bi-RNN is as follows: Figure 5 As shown, its main principle is to construct a single layer using two RNN networks. The basic idea is that each training sequence has two recurrent neural networks, one forward and one backward, and both layers are connected to the same output layer. One network runs along the data time series, and the other runs in the reverse direction. Therefore, the number of parameters in a Bi-RNN is twice that of a regular RNN. The data at time t is input into the two RNN networks. The two RNN networks are not directly connected, but they are combined during the output process to form the final output.
[0062] In the hidden layers of Bi-RNN, and For two single-layer RNN networks, the formulas are as follows:
[0063]
[0064]
[0065]
[0066] like Figure 6 As shown, at a given time step t, the mini-batch input (Number of samples is n, number of inputs is d) and the hidden layer activation function is 6. In the architecture of a bidirectional recurrent neural network, the forward hidden state at each time step is... (The number of forward hidden units is h), reverse hidden state (The number of reverse hidden units is h). Calculate the forward hidden state and the reverse hidden state using the following formula:
[0067]
[0068]
[0069] Among them, weight and deviation These are all parameters of the model.
[0070] Hidden states connecting two directions and Get hidden state This state is input into the output layer, and the output layer calculates... (The number of outputs is q):
[0071] O t =H t W hq +b q
[0072] Among them, weight and deviation These are the model parameters for the output layer.
[0073] like Figure 7 As shown, ordinary neural networks directly extract information features from a dataset through fully connected neurons. However, because they can only extract features from one-dimensional vectors, they are prone to losing spatial information. Secondly, too many parameters lead to inefficiency and make model training difficult. Furthermore, a large number of parameters can easily cause network overfitting, affecting the model's prediction performance. Therefore, for data with higher dimensional information, such as images, ordinary neural networks often perform poorly. Convolutional neural networks, by introducing convolution and pooling computational processes, enable neural networks to extract local spatial features, thereby achieving higher accuracy in more advanced data analysis problems. The computational process of a convolutional neural network is shown in the figure. After the image is directly input into the network, it undergoes several stages of convolution and pooling. Subsequently, the results of these operations are provided to one or more fully connected layers, with the final fully connected layer outputting the result.
[0074] A convolutional neural network (CNN) mainly consists of three parts: convolutional layers, pooling layers, and fully connected layers. The convolutional layer is the core layer for building a CNN, handling most of the network's computation. Neurons in a convolutional layer are arranged into feature maps. Each neuron in a feature map has a receptive field, connected to the neighborhood of neurons in the previous layer through a set of trainable weights, also known as a filter bank. A new feature map is computed by convolving the input image with the learned weights, and the convolution result is sent through a non-linear activation function. The weights of all neurons in a feature map are constrained to be equal; however, since different feature maps within the same convolutional layer have different weights, multiple feature values can be extracted at each location. The k-th output feature map Y... k It can be calculated as follows:
[0075] Υ k =f(W k *x)
[0076] In the formula, the input image x represents the input image, the convolutional filter associated with the k-th feature map is denoted by Wk, the multiplication symbol * represents the 2D convolution operator used to calculate the inner product of the filter models at each location in the input image, and f(·) represents the nonlinear activation function. Nonlinear activation functions allow the extraction of nonlinear features. Traditionally, sigmoid and hyperbolic tangent functions are used.
[0077] Convolutional layers extract features from the input data using convolutional kernels. The kernels slide across the input data, and calculations are performed to obtain the output. The calculation method is as follows: Figure 8 As shown, the convolution kernel slides across the input data, performing calculations with the kernel in each region of the input data that is the same size as the kernel, to obtain the output.
[0078] Pooling layers typically follow convolutional layers. Their purpose is to reduce the spatial resolution of feature maps, thereby maintaining spatial invariance to input distortion and translation. The results of convolutional layer operations are then used to further extract key information from the data. Pooling layers reduce the spatial size of the data by employing operations similar to convolution, thus reducing network computation and mitigating overfitting, thereby improving model generalization ability. Common pooling operations include max pooling and average pooling.
[0079] Early approaches used average pooling aggregation layers to propagate the average of the input values across all small neighborhoods of the image to the next layer. However, in more recent models, max pooling aggregation layers propagate the maximum value in the receptive domain to the next layer. Max pooling selects the largest element within each receptive domain:
[0080]
[0081] The output of the pooling operation associated with the k-th feature map is represented by Υ. kij Indicates that x kpq Represents the pooling region The elements contained at position (p, q) represent a receptive field around position (i, j).
[0082] Max pooling extracts features by retaining the maximum value of data within a region, while average pooling extracts features by calculating the average value of data within a region. Figure 9 , 10 The diagrams shown are for max pooling and average pooling with a 2×2 pooling kernel and a step size of 2, respectively.
[0083] To extract more abstract feature representations as the network moves through the data, several convolutional and pooling layers are typically stacked together. Fully connected layers follow these layers to interpret these feature representations and perform high-level inference. The fully connected layers classify images by utilizing the feature vectors of the image obtained from the convolutional and pooling layers.
[0084] During classification, the feature map of the last convolutional layer is vectorized and fed into a fully connected layer, followed by softmax logistic regression. This structure combines convolutional structures with traditional neural network classifiers. It uses convolutional layers as feature extractors to classify the obtained features in a traditional way. However, fully connected layers are prone to overfitting, which affects the generalization ability of the entire network. In addition, the required number of parameters, namely the size and number of channels of the last layer's feature map, needs to be pre-set in the fully connected operation to ensure that the number of training parameters remains unchanged. This contradicts the characteristic of convolutional layers sharing weights and being able to adapt to inputs of different sizes.
[0085] To avoid overfitting and eliminate size limitations, Global Average Pooling (GAP) was proposed to replace traditional fully connected layers. The idea is to generate a feature map for each corresponding class in a single layer for a classification task. For example... Figure 11 As shown, by averaging the channels of each feature map and forming a vector from these averages, the final vector length is made consistent with the number of channels in the last feature map layer, thus solving the problem inherent in fully connected layers. One advantage of global average pooling in fully connected layers is that it strengthens the correspondence between feature maps and classes, making it more suitable for convolutional structures. Therefore, feature maps can be easily interpreted as class confidence maps. Another advantage is that there are no parameters to optimize in global average pooling, thus avoiding overfitting of this layer. Furthermore, global average pooling aggregates spatial information, thus exhibiting stronger robustness to spatial transformations of the input.
[0086] In this embodiment, the data input into the input layer includes the average temperature of the ambient cold aisle, the average temperature of the ambient hot aisle, the server power, the speed of the precision air conditioner fan, and the speed of the precision air conditioner fan.
[0087] In this embodiment, the convolutional layer extracts key information from the time-series data through convolutional kernels.
[0088] In this embodiment, the fully connected layer further calculates and summarizes the extracted feature data.
[0089] In this embodiment, the output layer uses the sigmoid activation function to output the average temperature of the ambient thermal aisle corresponding to the server after 15 minutes.
[0090] In this embodiment, the specific method of S4 is as follows: the average temperature of the ambient cold aisle, the average temperature of the ambient hot aisle, the server power, the speed of the precision air conditioner fan, and the speed of the precision air conditioner fan are used as inputs to the BiRCNN network model, and the average temperature of the ambient hot aisle after 15 minutes is used as the target of the model output. The model is trained by dividing it into training set and test set according to a certain ratio.
[0091] In this embodiment, the specific method of S5 is as follows: save the model training parameters, input the ambient cold aisle temperature, ambient hot aisle temperature, server power, precision air conditioner fan speed and precision air conditioner fan speed settings into the model, and the average ambient hot aisle temperature corresponding to the server 15 minutes later can be predicted.
[0092] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A deep learning-based hot aisle prediction algorithm for data center environments, characterized in that, Includes the following steps: S1: Data Acquisition: Retrieve data from the corresponding server from the monitoring system; S2: Data processing: The data obtained in S1 is processed into an average value and then normalized. S3: Build the BiRCNN network model; S4: Model Training: Input the data processed in S2 into the BiRCNN network model and use it as the target of the model output. Divide the data into training set and test set according to a certain ratio for model training. S5: Predicted data: Save the model training parameters, and input the new time data into the BiRCNN network model to predict the average temperature of the environmental thermal channel corresponding to the server. The specific method of S1 is to obtain six types of data from the monitoring system: ambient cold aisle temperature, ambient hot aisle temperature, server power, precision air conditioner fan speed, precision air conditioner fan speed setting, and ambient hot aisle temperature after 15 minutes. The specific method of S2 is as follows: process all acquired ambient cold aisle temperatures into an average ambient cold aisle temperature, process all acquired ambient hot aisle temperatures into an average ambient hot aisle temperature, process all acquired ambient hot aisle temperatures after 15 minutes into an average ambient hot aisle temperature after 15 minutes, and normalize the processed average ambient cold aisle temperature, average ambient hot aisle temperature, average ambient hot aisle temperature after 15 minutes, server power, precision air conditioner fan speed, and precision air conditioner fan speed. The BiRCNN network model in S3 includes: input layer, BiGRU layer, convolutional layer, fully connected layer, and output layer; The data input into the input layer includes the average temperature of the cold aisle, the average temperature of the hot aisle, the server power, the speed of the precision air conditioner fan, and the speed of the precision air conditioner fan. The convolutional layer extracts key information from the time-series data through convolutional kernels; The fully connected layer further calculates and summarizes the extracted feature data; The output layer uses the sigmoid activation function to output the average temperature of the ambient thermal aisle corresponding to the server after 15 minutes.
2. The deep learning-based data center environment hot aisle prediction algorithm according to claim 1, characterized in that, The specific method of S4 is as follows: the average temperature of the cold aisle, the average temperature of the hot aisle, the server power, the speed of the precision air conditioner fan, and the speed of the precision air conditioner fan are used as inputs to the BiRCNN network model, and the average temperature of the hot aisle after 15 minutes is used as the target of the model output. The model is trained by dividing it into training set and test set according to a certain ratio.
3. The deep learning-based hot aisle prediction algorithm for data center environments according to claim 1, characterized in that, The specific method of S5 is as follows: save the model training parameters, input the ambient cold aisle temperature, ambient hot aisle temperature, server power, precision air conditioner fan speed and precision air conditioner fan speed settings into the model at the new moment, and the average ambient hot aisle temperature corresponding to the server 15 minutes later can be predicted.
Citation Information
Patent Citations
A data center computer room environment monitoring system
CN113311841B
Sea surface temperature prediction method and system based on hybrid learning model
CN115359338A
Softsensor analysis measurement system to provide the output by new progress variable based on collected data from a sort of sensors
KR102407886B1