Data center time sequence prediction method and system based on deep clustering

By adopting a deep clustering-based method in data center scenarios, detecting concept drift and fusing multiple prediction models, the insufficient performance of existing deep learning models in multivariable multi-step time series prediction is solved, and higher prediction accuracy and stability are achieved.

CN120196973AActive Publication Date: 2025-06-24NANJING UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510669013.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

Existing deep learning models perform poorly in data center scenarios, and cannot model multivariable multi-step time series accurately enough, especially when dealing with actual scenarios with conceptual drift.

Method used

Using a deep clustering method, we receive time series data transmitted from IoT devices in the data center in real time, use a clustering algorithm to detect concept drift status, combine feature extraction and clustering models, check the conceptual relative categories of data, and select or fuse multiple prediction models through a prediction algorithm based on ensemble learning to obtain real-time time series prediction results.

Benefits of technology

A multivariable time series prediction framework with considerable generalization is constructed, which can adapt to time series data with concept drift, improve the accuracy and stability of prediction, and solve the problem of poor applicability of traditional models in concept drift scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196973A_ABST
    Figure CN120196973A_ABST
Patent Text Reader

Abstract

The invention discloses a data center time sequence prediction method and system based on deep clustering. The method comprises the following steps: receiving time sequence data transmitted by data center Internet of Things equipment in real time; checking the concept relative category of the data through a clustering algorithm for concept drift condition detection in combination with feature extraction and a clustering model; and selecting one or fusing a plurality of selected different trained prediction models according to the relative categories of the concepts by using a prediction algorithm based on ensemble learning to obtain a real-time time sequence prediction result. According to the method, the concept drift phenomenon frequently occurring in a data center scene can be effectively processed, and the prediction stability is improved. A test result under a real data set proves that the combination of the clustering algorithm for concept drift condition detection and the prediction model obtains excellent performance under a data center scene, and the prediction accuracy and the performance in a concept drift state are superior to those of the existing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data center management, and in particular to a data center time series prediction method and system based on deep clustering. Background Art

[0002] A data center is a large-scale computing infrastructure that operates 24 / 7. Against the backdrop of the rapid development of cloud-based related service applications and artificial intelligence, data centers have been rapidly and massively constructed globally in the past few decades. The growth in the number of data centers has also synchronously brought about continuous growth in energy consumption and carbon emissions of data centers. Improving the energy efficiency of data centers and reducing their carbon emissions have become important issues of concern to all sectors of society.

[0003] Servers, network devices, power distribution devices, and cooling systems are the main sources of electricity consumption in data centers. In response to this, different researchers have proposed solutions that help reduce data centers from various perspectives. For example, optimizing and improving the energy utilization efficiency of IT devices, choosing to build data centers in high-latitude regions, and improving server performance from a hardware perspective. Among the various partial systems of a data center, the parts with the largest energy consumption ratio are the servers and the refrigeration system. The cooling system can account for about 50% of the energy consumption, while the proportion of servers is approximately 26%. Since the cooling system accounts for such a high proportion of the energy consumption in the data center, the research on the data center cooling control system is the key to reducing the energy consumption and carbon emissions of the data center. To reduce the energy consumption of the cooling system, the boundary conditions that need to be ensured are to ensure that relevant indicators such as the temperature and humidity in the server room are within the specified range to ensure the normal operation of the servers, and the operating conditions of the servers always affect the temperature and humidity in the server room. Therefore, the problem of optimizing the energy consumption of the cooling system should be modeled as a constrained optimization problem of a complex time-varying system. To solve this problem, a two-step method of "modeling - optimization" is generally adopted. First, a model-driven or data-driven method is used to perform predictive modeling on the system, and then, taking the output of the predictive modeling as feedback, an optimization algorithm is used. The accuracy of the model prediction directly determines the feasibility and efficiency of subsequent optimization. The main research focus of the present invention is the predictive modeling stage of the system.

[0004] There have already been a certain number of studies on the modeling of data center cooling systems. Initially, researchers mainly used methods including expert knowledge and thermodynamics to model data center cooling systems. This approach has a certain degree of reliability. However, the accuracy of its model is poor because it is difficult to ignore the errors caused by various approximations in a complex system when establishing a system model based on an ideal physical model. In recent years, with the rapid development of Internet of Things technology, sensor components, and information processing technology, the amount of operating state data of data centers that can be obtained has increased geometrically, laying the foundation for the application of data-driven models. Data-driven models generally use methods related to machine learning and deep learning to directly capture the operating patterns of data centers from data. For data-driven modeling of data centers, some studies start from the perspective of regression fitting to capture the relationships between equipment operating conditions, environmental factors, and data center energy consumption. However, the complexity and time-variability of data center systems pose significant challenges to the models used in previous studies. Therefore, it is quite necessary to use more accurate and advanced deep learning models to model data center systems because more accurate modeling can ensure the stability and reliability of subsequent optimizations. Summary of the Invention

[0005] Object of the Invention: Aiming at the deficiencies of the prior art, the object of the present invention is to provide a data center time series prediction method and system based on deep clustering, which solves the problems that the performance of existing deep learning models is not ideal in the data center scenario, cannot accurately model multi-variable and multi-step time series, and cannot handle the modeling problems of actual scenarios with concept drift.

[0006] Technical Solution: To achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a data center time series prediction method based on deep clustering, including the following steps: Real-time receive time series data incoming from Internet of Things devices in the data center; Through a clustering algorithm for concept drift detection, combined with a feature extraction and clustering model, check the relative concept categories of the data; Use a prediction algorithm based on ensemble learning to select one or fuse multiple different pre-trained prediction models according to the relative concept categories to obtain real-time time series prediction results.

[0007] Further, the clustering algorithm distinguishes time series data under different concepts according to the distribution of time series data; the relative concept categories are the clustering results or soft clustering results of the clustering algorithm.

[0008] Further, the pre-training steps of multiple prediction models include: Use a clustering algorithm to divide the training set data; According to the sub-datasets after clustering, one or more prediction models are trained respectively to obtain a set of prediction models corresponding to each clustering result.

[0009] Further, the clustering algorithm is implemented by adding a clusterer to the variational autoencoder. In the variational autoencoder, the latent variable z generated from the original time series data x by the encoder is used to obtain another latent variable y as the clustering label through the clusterer, and the data is jointly reconstructed and generated by the decoder using z and y. Meanwhile, y is reconstructed and generated into a latent variable by the conditional prior generator. 。

[0010] Further, the training loss function of the variational autoencoder with an added clusterer is expressed as: ; where, are the parameters of the decoder, is the probability distribution of the decoder generating the reconstructed original data given the latent variable z and the clustering label y; are the parameters of the conditional prior generator, is the prior distribution of the conditional prior generator generating the reconstructed latent variable conditioned on the clustering label y; are the parameters of the encoder, is the distribution of the encoder inferring the latent variable z based on the given input data x; are the parameters of the clusterer, is the probability distribution of the clusterer predicting the clustering label y based on the latent variable z; p(y) is the prior distribution of the clustering label y, and D KL represents the KL divergence calculation, represents taking the expectation of the sampled z and y.

[0011] Further, the method further includes: continuously updating the training set data through the offline training layer, and respectively updating the clustering algorithm and the prediction model through clustering evaluation and prediction evaluation, and updating them into the online prediction layer.

[0012] Further, the time series prediction is a multivariate multi-step time series prediction. The input variables of the prediction model include the temperature and humidity in the data center computer room, IT load, the working conditions that can be actively adjusted in the cooling system, the monitoring variables that cannot be actively adjusted in the cooling system, and the energy consumption of each part of the cooling system; the prediction variables of the prediction model are a subset of the input variables, including the temperature and humidity in the data center computer room, the monitoring variables that cannot be actively adjusted in the cooling system, and the energy consumption of each part of the cooling system.

[0013] In a second aspect, the present invention provides a data center time series prediction system based on deep clustering, including: The data collection module is used to receive time series data from IoT devices in the data center in real time; A clustering module for checking the relative categories of concepts in data by combining feature extraction with clustering models through a clustering algorithm for concept drift detection; The prediction module uses a prediction algorithm based on ensemble learning to select one or fuse multiple different trained prediction models according to the relative categories of concepts to obtain real-time time series prediction results.

[0014] In a third aspect, the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the data center time series prediction method based on deep clustering are implemented.

[0015] In a fourth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the data center time series prediction method based on deep clustering.

[0016] Beneficial effects: The present invention constructs a multivariate time series prediction framework with considerable generalizability for data center scenarios, which can realize online prediction and offline learning, and can adapt to time series data with concept drift, solving the problem of poor applicability of traditional models for time series data with concept drift. The present invention has conducted sufficient and extensive testing based on data center data sets obtained from real IoT devices, and the results can verify the advanced nature of the data center time series prediction method proposed by the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flow chart of the online prediction method of an embodiment of the present invention.

[0018] Figure 2 It is a structural diagram of a data center time series prediction framework in an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the deep clustering process in an embodiment of the present invention.

[0020] Figure 4 It is a comparison chart of the time series data after VAE reconstruction in an embodiment of the present invention.

[0021] Figure 5 It is a schematic diagram of the time series data clustering result in an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the following provides a detailed description of the embodiments of the present invention with reference to the accompanying drawings: These embodiments are implemented on the premise of the technical solutions of the present invention, and detailed implementation manners and specific operation processes are given. It should be understood that the specific examples described herein are only used to explain the present invention, but the protection scope of the present invention is not limited to the following embodiments.

[0023] As Figure 1 shown, a data center time series prediction method based on deep clustering disclosed in an embodiment of the present invention, the online prediction process mainly includes: real-time receiving time series data transmitted from Internet of Things devices in the data center; through a clustering algorithm for concept drift condition detection, combining feature extraction and a clustering model, checking the relative concept categories of the data; using a prediction algorithm based on ensemble learning, according to the relative concept categories, selecting one or fusing multiple different already trained prediction models, to obtain a real-time time series prediction result.

[0024] The following details the method of the embodiment of the present invention based on the system model of the data center of the Internet of Things. A large number of sensors are deployed at various locations and various devices in the data center, such as power sensors for detecting the operating power of IT devices, air conditioners and other refrigeration devices, temperature and humidity sensors for ensuring the safe operation of the computer room, and flow meter sensors inside the refrigeration system. These sensors form a sensor network, and these sensors collect data regularly through wired or wireless links, and after being processed, upload it to the data center infrastructure management (DCIM) system and store it in the database. The Internet of Things data collected by the sensor network is a digital representation of the actual operation scenario of the data center, and is used for the establishment of a data-driven model based on deep learning for subsequent use.

[0025] This embodiment focuses on studying the prediction models of the IT computer room and the refrigeration system in the data center. The IT computer room is the core functional area of the data center, mainly responsible for data storage and data calculation. The refrigeration equipment adopts the architecture of a centralized water-cooled chilled water air conditioner to dissipate the heat of the IT computer room. Specifically, during continuous computing in the IT computer room, devices such as CPUs and GPUs will generate a large amount of heat. To ensure the safe operation of the data center, the terminal precision air conditioners in the IT computer room need to be continuously turned on to cool the IT load, ensuring that the temperatures of the cold and hot channels of the IT load are maintained within the safe operating range. The low-temperature chilled water inside the air conditioner exchanges heat with the hot air inhaled by the air conditioner, and then flows back to the chiller / heat exchanger through a chilled water pump at the back end of the air conditioner. Through the heat exchange of the chiller / heat exchanger, the chilled water transfers the heat to the cooling water, and the cooled chilled water flows back to the IT computer room to cool the air conditioner again. The cooling water that has absorbed a certain amount of heat in the chiller unit enters the cooling tower under the pressure of the cooling water pump, exchanges heat with the outdoor air, and then flows back to the chiller / heat exchanger for the next heat cycle.

[0026] For the system model according to this embodiment, the problem of predicting and modeling a data center is a multi-variable multi-step time series prediction problem. IoT sensors collect data at regular time intervals, discretizing the state of the data center in time. According to the method for general multi-variable time series prediction modeling problems, the problem is described as follows: , where are the input parameters of the model, and the input is an n-dimensional vector for the past D time steps . are the output parameters of the model, and the output is an m-dimensional vector for the next H time steps .

[0027] To ensure the accuracy of the prediction, we hope that the prediction result of this model is closest to the actual real result. Assume that the predicted output of the model is: , where is the prediction model we expect to construct, L represents a loss function that measures the gap between the predicted output and the real result, and arg min means that we hope to obtain the model parameters when the loss function is minimized. For the selection of the loss function, in time series prediction problems, MSE, MAE, etc. are often used as loss functions.

[0028] In this embodiment, all IoT device measurement points such as the temperature and humidity in the computer room, IT load, the working conditions that can be actively adjusted in the cooling system, the monitoring variables that cannot be actively adjusted in the cooling system, and the energy consumption of each part of the cooling system are regarded as the input variables of the model, while the variables that the model needs to predict are a subset of the input variables, only including the monitoring variables that cannot be actively adjusted in the cooling system, the temperature and humidity in the computer room, and the energy consumption of each part of the cooling system. These predicted variables are used as the optimization objectives and safety boundary conditions for the subsequent optimization part.

[0029] At the same time, in this embodiment, the prediction model does not need to make very long-term predictions because the data center based on IoT can sample regularly and obtain real values regularly to correct prediction errors. However, at the same time, considering that many methods in the optimization part require the system conditions in the next multiple steps to ensure the performance and safety of the normal operation of the system, the future prediction time step in this embodiment cannot be a single step either. Therefore, the future prediction window in this embodiment is set according to the short-term time series prediction problem.

[0030] In this embodiment, a clustering-based time series prediction framework is proposed to address the concept drift problem encountered in time series prediction in the data center scenario. First, the concept drift phenomenon in time series prediction is briefly introduced below, and then the online-offline double-layer prediction framework for solving the concept drift phenomenon is introduced. Finally, the clustering-based concept drift detection and ensemble learning prediction methods used in the prediction framework are described.

[0031] Concept drift is attributed to one of the main reasons for the performance degradation in data-driven systems. Concept drift is a phenomenon in which the statistical characteristics of the target domain change arbitrarily over time. Formally, the definition of concept drift is as follows.

[0032] In this embodiment, the multi-variable multi-step time series prediction problem is brought into the concept drift definition. The input feature is X, and the predicted value Y at a future time is used as the label. In this case, concept drift can be considered as: , where represents the joint distribution of (X, Y) at any time in the interval from 0 to t in time, then represents the joint distribution of (X, Y) at any time after the concept drift occurs at time t.

[0033] Here, it is considered that (X, Y) at any time from 0 to t conforms to the same joint distribution. Therefore, the joint probability at time t and time t + 1 is taken and expressed as: , where represents the joint probability distribution before the concept drift occurs at time t, represents the joint probability distribution at time t + 1 after the concept drift occurs.

[0034] Further disassembling the input and output, substituting X and Y into the specific representation of the time series problem, we can obtain; , where the input feature represents the input from time t - T s + 1 to time t, a total of T s time steps. The output feature represents the predicted output feature from time t + 1 to time t + T p at a total of T p time steps. and represent the input feature and the predicted output feature corresponding to time t + 1 after the concept drift occurs at time t.

[0035] Here, further, we decompose the joint probability into two parts: ; then there is: .

[0036] There are two possible reasons for the above formula to hold: not equal to or not equal to (Of course, it is also possible that both exist simultaneously). For the case where the latter holds, whether the former is equal or not, it will have a huge impact on the performance of the data-driven time series prediction model, because the model itself is to fit as much as possible the process. If changes, then the performance of the model and the prediction will inevitably decline.

[0037] In the data center scenario, the frequency of concept drift is higher than in other scenarios, mainly due to two aspects: IT equipment and the refrigeration system. First, for a machine room with insufficient load, different IT devices will take turns to perform high-heat-consuming computing tasks, and the generation and device allocation of each computing task are randomly generated and unpredictable; on the other hand, some devices in the data center refrigeration system will be turned on or off irregularly, and the devices will take turns to bear the cooling load to extend the device life. At the same time, the number of cooling devices in operation will also be dynamically adjusted according to the results generated by the optimization algorithm. Therefore, in the data center scenario, it is very necessary to consider the impact of concept drift in the prediction modeling process.

[0038] The time series prediction framework proposed in this embodiment is as Figure 2 shown. The time series prediction framework consists of two layers: an online prediction layer and an offline training layer in terms of function. During the operation of each layer, it is completed in two steps: clustering and prediction. The online prediction layer receives the time series data transmitted from the Internet of Things devices in real time. First, it passes through the feature extraction and clustering model for detecting the concept drift situation to check the relative concept categories of the data. Subsequently, the data enters the prediction module. The prediction module uses the ensemble learning algorithm to select and fuse different pre-trained prediction models according to the relative concept categories, and finally obtains the real-time time series prediction result.

[0039] The offline training layer continuously updates the training set data, and updates the clustering algorithm and the time series prediction algorithm for concept drift detection through clustering evaluation and prediction evaluation respectively, and deploys the updated models to the online prediction layer.

[0040] Specifically, since the clustering in this embodiment is an unsupervised evaluation index, the Calinski-Harabasz (CH) index is used to evaluate the clustering results of the algorithm. In the offline training layer, the CH index is used as the evaluation index for the VAE clustering results. The result of the prediction evaluation is consistent with the loss function of the normal prediction algorithm. In this embodiment, RMSE is used as the evaluation index.

[0041] The basis of the time series framework design idea in this embodiment is as follows: In a complete time series dataset that can be used as a training set, there may be many concept drift phenomena, and the distribution of the data can be clustered. We assume that the concepts of time series data are limited, which means that for different time periods, there may be data that conforms to the same distribution, and the running logic of these data is the same. Therefore, the same model should also be used for predictive modeling. The reason why we use clustering instead of classification here is that concept drift in the data center scenario is caused by both known human factors and unknown other factors. For multivariate time series data, it is difficult for humans to directly capture the changing patterns. Therefore, there are no known labels for the specific concept types of time series, and it should be modeled as an unsupervised learning problem.

[0042] Based on this, this time series prediction framework is designed. First, a clustering algorithm for concept drift detection is used to divide the training set data. Then, according to the clustered sub-datasets, the models are trained separately, and finally a set of sub-models corresponding to each clustering result is obtained. When using this time series prediction framework to predict new data subsequently, based on the clustering results or soft clustering results of the clustering algorithm, one or a combination of multiple sub-models that the current data is most likely to belong to are selected for prediction.

[0043] Specifically, during the pre-training process, from the original training set data a time series data with a time window of T p +T S is taken as the overall distribution sampling of where represents the n-dimensional input variable at the i-th moment, N represents the length of the time series, T p is the length of the prediction input time, T S is the length of the prediction output time, and a new dataset is obtained, where m represents the sample size of the new dataset, and there is indicating the new data points obtained after sliding by the time window. For this dataset, a certain clustering algorithm is used to cluster the set of to obtain the clustering result , where k represents the number of clustering clusters. After that, according to the multiple sub-datasets after clustering, the models are trained separately to obtain a set of models corresponding to each clustering result. When using this time series framework subsequently, first, through the clustering algorithm, the clustering category that the current data is most likely to belong to is judged, and the corresponding category model is used for prediction. Of course, an ensemble learning method can also be used to combine models of multiple categories for prediction according to the data and the soft clustering results of each clustering center.

[0044] In this embodiment, in the face of the problem of concept drift in time series, the clustering algorithm will distinguish the time series data under different concepts according to the distribution of the time series. The variational autoencoder (VAE) is an autoencoder based on Bayesian variational inference and is mainly used as a deep generative model, which can effectively learn the distribution of data and remove the noise in the data. The core idea of VAE is to compress the random vector belonging to the high-dimensional space into the latent variable in the low-dimensional space through variational encoding. During the training process, the loss function of the standard VAE is expressed as: ; where ELBO represents the variational lower bound of the log-likelihood of and represents the expectation of the latent variable z sampled from the distribution is the inference distribution (approximate posterior distribution), which represents the approximate probability distribution of the latent variable z given the observed data x, is the joint probability distribution, which describes the joint generation process of the data x and the latent variable z.

[0045] For the computer vision (CV) and natural language processing (NLP) scenarios where VAE is often used, it is crucial to use the decoder to generate data after obtaining the latent variable z. However, in this embodiment, we are more concerned about the latent variable z that contains the information of the original data x. We hope to cluster the original data according to the latent variable z. According to the general method, after completing the training of VAE, traditional clustering methods such as k-means or Gaussian mixture model (GMM) are directly used for the latent variable z to achieve clustering. However, in this embodiment, due to the large amount of time series data samples and the input in the form of a time stream, the clustering center needs to be continuously updated, and the time complexity and space complexity of the traditional method are high. Therefore, it is not feasible for a large number of time series data. Therefore, this embodiment proposes a direct clustering scheme based on neural network in VAE, as shown in Figure 3 .

[0046] In the scheme of this embodiment, the latent variable generated from the original time series data by the encoder is obtained through the clustering device to obtain another latent variable as the clustering label, and the data is jointly reconstructed and generated through the decoder , and at the same time, y is reconstructed and generated into the latent variable through the conditional prior generator. It should be noted here that y represents the discrete variable of the clustering category and can be obtained by fitting during the implementation of the model.

[0047] The important idea of VAE is to use the evidence lower bound as the loss function for model training. In this embodiment, the joint distribution is changed to , and the ELBO (variational lower bound) to be optimized can be expressed as: ; where ELBO still represents The variational lower bound, where p(x, z, y) is the jointly probability distribution after the change of definition, and q(z, y|x) is the new inference distribution (approximate posterior distribution).

[0048] At this time, the joint distribution p(x, z, y) and the approximate posterior q(z, y|x) are decomposed into conditional probabilities, and we have: ; Substituting into the ELBO, we get: .

[0049] Regarding the second and third terms, the following transformations need to be considered: ; This can be used as the loss function during the neural network training process: .

[0050] The above ELBO can be decomposed into three parts, namely , , .

[0051] Among them, the first term can be understood as the reconstruction loss, which is used to measure the gap between the generated x by the model and the real x. The second KL divergence measures the difference between p(z|x) and q(z|y) for each possible y. The role of this term is to encourage that within the same cluster y, the real p(z|x) and the distribution q(z|y) generated by clustering through y are as close as possible, so that y can better represent the clustering effect of z. The third KL divergence measures the difference between q(y|z) and p(y), which is used to control the clustering effect of y, ensuring that y not only generates appropriate clusters according to z, but also conforms to the prior distribution p(y) as much as possible, thus avoiding overfitting or generating unreasonable clustering structures.

[0052] Finally, based on the above content, the loss function can be expressed as: ; where are the parameters of the decoder, is the probability distribution that the decoder generates the reconstructed original data given the latent variable z and the cluster label y, which is used to ensure that the latent variable z and the cluster label y can effectively restore the original information; are the parameters of the conditional prior generator, is the prior distribution that the conditional prior generator generates the reconstructed latent variable conditioned on the cluster label y, which is used to ensure that the latent variables of similar samples are clustered in a specific region; are the parameters of the encoder, is the distribution that the encoder infers the latent variable z according to the given input data x; are the parameters of the clusterer, is the probability distribution that the clusterer predicts the cluster label y according to the latent variable z. DKL Denotes the KL divergence calculation, which is used to measure the difference between two distributions. Denotes taking the expectation of the sampled z and y.

[0053] To prove the technical effect of the present invention, a dataset from the real world is used below to verify the present invention and compare it with other benchmarks.

[0054] In this embodiment, data is collected in real time through Internet of Things devices and uploaded to a database for storage and real-time update through the infrastructure management system of the data center. The data comes from the normal operation data center of a certain communication company in a certain city, and the collection starts at 08:39 on May 31, 2022 and ends at 04:43 on December 8, 2022. The original dataset contains more than 700 data measurement points and more than 250,000 data entries.

[0055] In the data preprocessing process, we process the data separately in terms of the time dimension and the variable dimension. In terms of the time dimension, we first uniformly process the interval time of the data. Since the data collection frequencies of the Internet of Things devices are different, we use linear interpolation and moving window averaging to process the original data, and the time interval of the data is unified to a standard of 1 minute. At the same time, due to the reasons of the Internet of Things devices, there are obvious anomalies or blanks in the data in some time periods. After screening, finally 188,747 valid time point data are obtained.

[0056] In terms of the variable dimension, first, the method of data integration is used to merge the variables with the same or repeated meanings in the variables. After that, some variables that are considered irrelevant or unimportant in the data center modeling process are removed, and the variables are further screened using expert prior knowledge. Finally, dimensionality reduction and feature selection are performed on the data through methods such as recursive feature elimination, and 144 valid variable dimensions are retained.

[0057] In the data center prediction modeling task concerned in this embodiment, since the data center based on Internet of Things devices can update its status in real time online, we do not need a too long prediction time length. In the experiment, we divide it into short-term prediction and long-term prediction to evaluate the results. We set the observation history window length of short-term prediction to 4, the prediction time window length to 16, the observation history window length of long-term prediction to 48, and the prediction time window length to 96.

[0058] To verify the performance of the comparison neural network, we use four common metrics: MAE (Mean Absolute Error), RMSE (Root Mean Square Error), SMAPE (Symmetric Mean Absolute Percentage Error), and MASE (Mean Absolute Scaled Error), and compare them with a group of classical algorithms and advanced models that perform well in time series prediction, including RNN, LSTM, etc.

[0059] In the test phase of the subsequent time series prediction framework, we select various classical models and current advanced time series prediction models as the neural networks required in the concept drift detection clustering algorithm based on VAE, including LSTM, MLP, CNN, Transformer, LightTs, Autoformer, Crossformer, Dlinear, FEDformer, Informer, iTransformer, TimesNet, NS-Transformer, and PatchTST.

[0060] In the neural network model test, the batch size is set to 128, the learning rate is 3e-4, the optimizer used is adam, and the loss function is MSE.

[0061] For the encoder, decoder, clusterer, and conditional prior generator mentioned in this embodiment, a fully connected neural network with 5 layers and 64, 128, 256, 128, and 64 nodes in each layer is used for testing, and the clusterer obtains the final clustering result output through the softmax function.

[0062] Input all the data in the training set into the clustering algorithm for concept drift detection in this embodiment to cluster the data, and the results are as Figure 4 , Figure 5 shown. Figure 4 is a comparison diagram before a certain variable is input and after being reconstructed by VAE, indicating that this VAE-based clustering algorithm can effectively extract the information of the input data and reconstruct it. Figure 5 is a schematic diagram of the clustering result when all variable dimensions and all time dimensions are unfolded together. It can be seen that all the data in the training set is divided into 5 different categories by the clustering algorithm.

[0063] Use the time series prediction framework including the clustering algorithm for concept drift detection on various classical models and current advanced time series prediction models, retrain the model and conduct tests, and the test results are shown in Table 1 and Table 2.

[0064] Table 1 Comparison of Test Results MAE and RMSE

[0065] Table 2 Comparison of Test Results SMAPE and MASE

[0066] It can be found from the above that the framework proposed in the embodiment of the present invention can effectively improve the performance of the prediction model in the case of concept drift when applied to almost all time series prediction models.

[0067] In summary, the present invention uses an algorithm based on deep learning to study the overall prediction modeling of an Internet of Things-based data center. In response to the concept drift occurring in the data center scenario, an online-offline two-layer framework is constructed. A clustering algorithm is used to partition data with different distributions in high-dimensional time series, and multiple models are trained. Finally, an ensemble learning-based method is used to integrate multiple models to obtain the output. Sufficient experiments are carried out on the data set generated in the real scenario, and the experimental results confirm the effectiveness of the present invention in the data center prediction modeling problem.

[0068] Based on the same inventive concept, an embodiment of the present invention also discloses a data center time series prediction system based on deep clustering, including: a data acquisition module for receiving in real time the time series data transmitted from the Internet of Things devices in the data center; a clustering module for checking the relative concept categories of the data by using a clustering algorithm for concept drift condition detection in combination with a feature extraction and clustering model; a prediction module that uses an ensemble learning-based prediction algorithm to select one or fuse multiple different pre-trained prediction models according to the relative concept categories to obtain a real-time time series prediction result. For specific implementation details, refer to the foregoing method embodiment and will not be elaborated herein.

[0069] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps of the data center time series prediction method based on deep clustering are implemented.

[0070] An embodiment of the present invention also discloses a computer program product, including a computer program. When the computer program is executed by the processor, the steps of the data center time series prediction method based on deep clustering are implemented.

[0071] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A data center time series prediction method based on deep clustering, characterized in that It includes the following steps: Receiving in real time the time series data incoming from the Internet of Things devices in the data center; Checking the relative concept categories of the data through a clustering algorithm for concept drift condition detection, in combination with feature extraction and a clustering model; Using a prediction algorithm based on ensemble learning, and selecting one or fusing multiple different pre-trained prediction models according to the relative concept categories to obtain a real-time time series prediction result.

2. The method for predicting the time series of a data center based on deep clustering according to claim 1, wherein The clustering algorithm distinguishes the time series data under different concepts according to the distribution of the time series data; the relative concept category is the clustering result or soft clustering result of the clustering algorithm.

3. The data center time series prediction method based on deep clustering according to claim 1, wherein The pre-training steps of multiple prediction models include: Dividing the training set data using a clustering algorithm; Respectively training one or more prediction models according to the clustered sub-datasets to obtain a set of prediction models corresponding to each clustering result.

4. The data center time series prediction method based on deep clustering according to claim 1, characterized in that The clustering algorithm is implemented by adding a clusterer to the variational autoencoder. In the variational autoencoder, the latent variable z generated from the original time series data x by the encoder is used to obtain another latent variable y as the clustering label through the clusterer, and the data is jointly reconstructed by the decoder using z and y. Meanwhile, y is reconstructed to generate the latent variable through the conditional prior generator. .

5. The method for predicting the time series of a data center based on deep clustering according to claim 4, wherein The training loss function of the variational autoencoder for the clustering enhancer is expressed as: ; Among them, are the parameters of the decoder, is the probability distribution for the decoder to generate the reconstructed original data given the latent variable z and the cluster label y ; are the parameters of the conditional prior generator, is the prior distribution for the conditional prior generator to generate the reconstructed latent variable conditioned on the cluster label y ; are the parameters of the encoder, is the distribution for the encoder to infer the latent variable z from the given input data x; are the parameters of the clusterer, is the probability distribution for the clusterer to predict the cluster label y from the latent variable z; p(y) is the prior distribution of the cluster label y, and D KL represents the KL divergence calculation, represents taking the expectation over the sampled z and y.

6. The data center time series prediction method based on deep clustering according to claim 1, wherein It also includes: Continuously updating the training set data through an offline training layer, and respectively updating the clustering algorithm and the prediction model through clustering evaluation and prediction evaluation, and updating them into the online prediction layer.

7. The method for predicting time series of a data center based on deep clustering according to claim 1, characterized in that The time series prediction is a multivariate multi-step time series prediction. The input variables of the prediction model include the temperature and humidity in the data center computer room, IT load, the working conditions that can be actively adjusted in the cooling system, the monitoring variables that cannot be actively adjusted in the cooling system, and the energy consumption of each part of the cooling system; the prediction variables of the prediction model are a subset of the input variables, including the temperature and humidity in the data center computer room, the monitoring variables that cannot be actively adjusted in the cooling system, and the energy consumption of each part of the cooling system.

8. A data center time series prediction system based on deep clustering, characterized in that, It includes: A data acquisition module for receiving in real time the time series data incoming from the Internet of Things devices in the data center; A clustering module for checking the relative concept categories of the data through a clustering algorithm for concept drift condition detection, in combination with feature extraction and a clustering model; A prediction module that uses a prediction algorithm based on ensemble learning and selects one or fuses multiple different pre-trained prediction models according to the relative concept categories to obtain a real-time time series prediction result.

9. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for data center time series prediction based on deep clustering according to any one of claims 1-7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for data center time series prediction based on deep clustering according to any one of claims 1-7.

Citation Information

Patent Citations

  • Change trend prediction method and device for high-dimensional reproduction concept drift stream data

    CN111797122A

  • Markov process-based time series stream data anomaly detection method

    CN112784896A

  • Spatial-temporal data online anomaly detection method and system based on concept drift self-adaption

    CN118094423A

  • Integrated learning classification method and system with concept drift detection function

    CN119884959A

  • Adaptive retraining of an artificial intelligence model by detecting a data drift, a concept drift, and a model drift

    US20230376825A1