Data center short-term power load prediction method, system, equipment and medium

By combining SARIMA and BiLSTM models and dynamically combining linear and nonlinear predictions, the accuracy problem of data center power load prediction is solved, improving prediction accuracy and adaptability to complex scenarios, and supporting energy conservation and emission reduction in data centers.

CN120933916APending Publication Date: 2025-11-11ELECTRIC POWER PLANNING & ENG INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511043854.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately predict data center electrical loads, especially under conditions of changing business models and declining equipment energy efficiency, resulting in large prediction errors and making it difficult to meet real-time requirements.

Method used

A composite model of SARIMA and BiLSTM is adopted, combining time series analysis and deep learning. The SARIMA model fits the linear part of the data center load, and the BiLSTM model fits the nonlinear part. The model parameters are then optimized using the GWO optimization algorithm to dynamically combine the linear and nonlinear predicted values ​​and output the load prediction value.

Benefits of technology

It improves the accuracy of short-term power load forecasting for data centers, is applicable to complex scenarios, supports the rational arrangement of power supply and energy storage plans, enhances the absorption of new energy sources and the proportion of green electricity in the load, and promotes energy conservation and emission reduction in data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120933916A_ABST
    Figure CN120933916A_ABST
Patent Text Reader

Abstract

The invention relates to the field of data processing, and provides a data center short-term power load prediction method, system and device and a medium, and the method comprises the steps: obtaining a to-be-predicted time period, and obtaining time series data according to the to-be-predicted time period; inputting the time sequence data into a preset SARIMA and BiLSTM composite model to obtain a linear predicted value and a linear residual sequence; normalizing the linear residual sequence to obtain a data set, and generating a nonlinear predicted value and a nonlinear residual sequence according to the data set; and dynamically combining the linear predicted value and the nonlinear predicted value into a load predicted value, and outputting the load predicted value and a nonlinear residual sequence. According to the method, the SARIMA model and the BiLSTM model are combined, the accuracy of data center load prediction can be improved, and the method is suitable for complex data center load prediction tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a method, system, device, and medium for short-term power load forecasting of data centers. Background Technology

[0002] A data center is a facility that centrally houses a large number of servers, storage devices, network equipment, and other computing resources, providing data storage, processing, and management services for various organizations and enterprises. Under the green and low-carbon context, the adoption of distributed energy sources such as wind and solar power, as well as energy storage devices like energy storage, offers data centers more opportunities to reduce energy costs and carbon emissions. Simultaneously, the primary goal of data center energy use is to ensure stable power supply and cooling for servers, requiring a reliable and efficient energy management platform. Accurate forecasting of data center electrical load can help the energy management system make reasonable scheduling decisions, improving the security and economy of data center energy use.

[0003] For a single data center, the annual energy cost can reach tens of millions of yuan. Studying the load energy consumption characteristics is a prerequisite for improving the power utilization efficiency of data centers. However, the electrical load of data centers is highly random, making modeling and prediction difficult.

[0004] Currently, the closest load modeling methods fall into two categories: pure data-driven methods based on time series analysis or feedforward neural networks. While these methods can capture the statistical patterns of historical loads, they neglect the physical mechanisms of energy consumption, making it difficult to accurately reflect the intrinsic drivers such as changes in business load and equipment energy efficiency degradation. Time series methods rely solely on extrapolating historical load values, and when data center business patterns undergo abrupt changes, prediction errors increase significantly. Although feedforward neural networks can handle nonlinear relationships, they cannot model the time-series dependencies in load sequences, are sensitive to short-term fluctuations, and lack the ability to predict long-term trends.

[0005] Physically based modeling methods go to the other extreme. These methods construct mathematical models by analyzing server power consumption models and cooling system thermodynamic equations. While they can reveal the fundamental laws governing energy consumption, the diverse types of data center equipment and complex operating conditions make it difficult to identify model parameters and quantify random interference such as equipment failures and human error. Furthermore, the computational complexity increases exponentially with the changes in airflow organization brought about by the expansion of data centers, making it difficult to meet real-time prediction requirements. Summary of the Invention

[0006] To achieve the above objectives, this invention proposes a method for short-term power load forecasting of data centers, comprising:

[0007] Obtain the time period to be predicted, and obtain time series data based on the time period to be predicted;

[0008] The time series data is input into a preset SARIMA and BiLSTM composite model to obtain linear predicted values ​​and linear residual sequences.

[0009] The linear residual sequence is normalized to obtain a dataset, and nonlinear predicted values ​​and nonlinear residual sequences are generated based on the dataset.

[0010] The linear and nonlinear predicted values ​​are dynamically combined into a load predicted value, and the load predicted value and the nonlinear residual sequence are output.

[0011] In some embodiments, the process of constructing the preset SARIMA and BiLSTM composite model is as follows:

[0012] Construct a pre-defined SARIMA model and a pre-defined BiLSTM model, and combine them into a training model;

[0013] The training model is fine-tuned by calculating the error of historical test data to obtain the final training model as the preset SARIMA and BiLSTM composite model.

[0014] In some embodiments, the process of constructing the preset SARIMA model is as follows:

[0015] Obtain historical data of data center load, group it into training dataset and test dataset, and determine the initial values ​​of SARIMA parameters;

[0016] A first-generation SARIMA model is constructed, and the linear part of the time series in the training dataset is fitted by the first-generation SARIMA model to generate the first predicted value.

[0017] Obtain the measured data corresponding to the first predicted value, and obtain the first residual sequence by subtracting the first predicted value from the measured data;

[0018] The SARIMA parameters are tuned according to grid search and AIC criteria until the residuals generated by two consecutive predictions are less than the first threshold. The SARIMA parameters used in the last prediction are then taken as the optimal values ​​of the SARIMA parameters.

[0019] Construct a pre-defined SARIMA model based on the optimal values ​​of SARIMA parameters.

[0020] In some embodiments, the process of constructing the preset BiLSTM model is as follows:

[0021] The first residual sequence is normalized to obtain the BiLSTM dataset, and the initial values ​​of the BiLSTM parameters are determined.

[0022] Construct an initial BiLSTM model and predict the second residual on the BiLSTM dataset;

[0023] A second predicted value is generated by fitting the nonlinear component in the training dataset using the first-generation BiLSTM model.

[0024] The second residual sequence is obtained by subtracting the second residual from the second predicted value;

[0025] The BiLSTM parameters are tuned according to the GWO optimization algorithm until the second residual is less than the second threshold after two consecutive predictions. The BiLSTM parameters used in the last prediction are then taken as the optimal values ​​of the BiLSTM parameters.

[0026] Construct a pre-defined BiLSTM model based on the optimal values ​​of the BiLSTM parameters.

[0027] In some embodiments, the process of fine-tuning the training model by calculating the error of historical test data to obtain the final training model includes:

[0028] Input the test dataset into the training model to obtain the historical test data error;

[0029] The training model is fine-tuned based on the historical test data to obtain the final training model.

[0030] In some embodiments, the process of dynamically combining the linear and nonlinear forecasts into a load forecast includes:

[0031] The linear and nonlinear predicted values ​​are standardized and concatenated to obtain standardized features.

[0032] The standardized features are input into the conditional encoder of GNF to obtain the latent spatial distribution parameters;

[0033] Dynamic weights are obtained by sampling from the potential spatial distribution parameters;

[0034] The load forecast value is obtained by weighting the linear and nonlinear forecast values ​​according to the dynamic weights.

[0035] In some embodiments, the process of obtaining the time period to be predicted and obtaining time series data based on the time period to be predicted includes:

[0036] Determine the historical data time window based on the time period to be predicted;

[0037] Retrieve historical load data within the historical data time window from the database;

[0038] The time period to be predicted and the historical load data are combined into time series data.

[0039] This invention proposes a short-term power load forecasting system for data centers, comprising:

[0040] The acquisition module is configured to acquire the time period to be predicted and obtain time series data based on the time period to be predicted.

[0041] The first prediction module is configured to input the time series data into a preset SARIMA and BiLSTM composite model to obtain linear prediction values ​​and linear residual sequences.

[0042] The second prediction module is configured to normalize the linear residual sequence to obtain a dataset, and generate nonlinear prediction values ​​and nonlinear residual sequences based on the dataset.

[0043] The combination module is configured to dynamically combine the linear and nonlinear forecast values ​​into a load forecast value, and output the load forecast value and the nonlinear residual sequence.

[0044] This invention proposes a computer device, comprising:

[0045] At least one processor; and a memory storing a computer program that can run on the processor, wherein the processor executes the steps of the data center short-term power load forecasting method when executing the program.

[0046] The present invention proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the data center short-term power load forecasting method.

[0047] The present invention has at least the following beneficial technical effects:

[0048] This invention proposes a method, system, device, and medium for short-term power load forecasting of data centers. The method includes: acquiring a time period to be predicted; obtaining time series data based on the time period; inputting the time series data into a preset SARIMA and BiLSTM composite model to obtain linear predicted values ​​and linear residual sequences; normalizing the linear residual sequences to obtain a dataset; generating nonlinear predicted values ​​and nonlinear residual sequences based on the dataset; dynamically combining the linear and nonlinear predicted values ​​into load predicted values; and outputting the load predicted values ​​and nonlinear residual sequences.

[0049] This invention combines SARIMA and BiLSTM models to improve the accuracy of data center load forecasting, making it suitable for complex data center load forecasting tasks. It aims to effectively enhance the accuracy of short-term data center load forecasting. Based on this, data center power supply systems can rationally arrange power supply and energy storage charging / discharging plans, thereby increasing the local consumption of new energy sources and the proportion of green electricity in the load, and accelerating energy conservation and emission reduction in the field of new data center infrastructure. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.

[0051] Figure 1 A flowchart of a short-term power load forecasting method for data centers provided by the present invention;

[0052] Figure 2 A block diagram of a data center short-term power load forecasting system provided by the present invention;

[0053] Figure 3 A schematic diagram of the structure of an embodiment of the computer device provided by the present invention;

[0054] Figure 4 A schematic diagram of an embodiment of the computer-readable storage medium provided by the present invention;

[0055] Figure 5 This is a flowchart illustrating the model construction and usage in a short-term power load forecasting method for data centers provided by the present invention. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to specific examples and the accompanying drawings.

[0057] It should be noted that all uses of "first" and "second" in the embodiments of the present invention are for the purpose of distinguishing two entities or parameters with the same name but different names. It is clear that "first" and "second" are only for the convenience of expression and should not be construed as limiting the embodiments of the present invention. Subsequent embodiments will not explain this in detail.

[0058] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0059] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0060] Upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0061] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation methods of this disclosure; other methods that comply with relevant laws and regulations may also be applied to the implementation methods of this disclosure.

[0062] Traditional forecasting models are typically used for short-term and long-term load forecasting. One type involves breaking down IT and cooling equipment into multiple components based on their functions, building energy consumption models for each component, and then summing the models to obtain the total energy consumption. This type of method can be used to predict the total energy consumption over a period of time, but it cannot effectively predict load time-series curves. Another type uses deep learning algorithms to describe load variation patterns; however, forecasting models built based on this method require careful selection of appropriate model parameter values ​​to improve forecast accuracy.

[0063] In view of this, the present invention proposes a short-term power load forecasting method for data centers based on SARIMA and BiLSTM. Based on the energy consumption generation mechanism of data centers, data-driven technology is used for modeling, which can effectively improve the accuracy of load modeling and forecasting.

[0064] This invention proposes a short-term power load forecasting method for data centers based on SARIMA and BiLSTM, comprising the following steps: grouping historical load data of data centers, determining initial values ​​of SARIMA model parameters, constructing a SARIMA model, fitting the linear part of the data center load data using the SARIMA model, calculating the residual sequence Δx, normalizing the residual sequence Δx and creating a BiLSTM dataset, constructing a BiLSTM model, using the BiLSTM model to predict the residual Δx, and calculating the quadratic residual Δx. 2x. The GWO (Grey Wolf) optimization algorithm is used to fine-tune the parameters of the BiLSTM model, resulting in a trained model. Historical test data is then used to fine-tune the model parameters to obtain the final trained model. Based on the final trained model, the short-term power load of the data center is predicted. This invention combines SARIMA and BiLSTM models, improving the accuracy of data center load forecasting and making it suitable for complex data center load forecasting tasks. SARIMA parameter tuning is performed using grid search and the AIC criterion. The SARIMA model built based on the tuned parameter values ​​better characterizes the linear and seasonal components of data center power load. The GWO optimization algorithm is used to fine-tune the parameters of the BiLSTM model. Considering that the number of neurons, the number of iterations, and the learning rate have a greater impact on the prediction accuracy of the BiLSTM model, the GWO optimization algorithm is used to optimize these three parameters. The BiLSTM model built based on the optimized parameter values ​​better characterizes the nonlinear components of data center power load.

[0065] This invention aims to effectively improve the accuracy of short-term load forecasting for data centers. Based on this, the data center power supply system can rationally arrange power supply and energy storage charging and discharging plans, thereby improving the local consumption level of new energy sources and the proportion of green electricity in the load, and accelerating the promotion of energy conservation and emission reduction in the field of new infrastructure for data centers.

[0066] To ensure clarity, the complete technical solution of the embodiments of this disclosure will be described in detail below. This invention proposes a method for short-term power load forecasting in data centers; please refer to [link to relevant documentation]. Figure 1 and Figure 5 ,include:

[0067] S1. Obtain the time period to be predicted, and obtain time series data based on the time period to be predicted;

[0068] S2. Input the time series data into a preset SARIMA and BiLSTM composite model to obtain linear predicted values ​​and linear residual sequences;

[0069] S3. Normalize the linear residual sequence to obtain a dataset, and generate nonlinear predicted values ​​and nonlinear residual sequences based on the dataset;

[0070] S4. Dynamically combine the linear and nonlinear predicted values ​​into a load predicted value, and output the load predicted value and the nonlinear residual sequence.

[0071] This invention provides a method for short-term power load forecasting of IT equipment such as servers and auxiliary service equipment such as cooling and lighting in data centers. The invention proposes a short-term power load forecasting method for data centers based on SARIMA and BiLSTM, including the following steps: grouping historical data of data center load, determining initial values ​​of SARIMA model parameters, constructing a SARIMA model, fitting the linear part of the data center load data using the SARIMA model, calculating the residual sequence Δx, normalizing the residual sequence Δx and creating a BiLSTM dataset, constructing a BiLSTM model, using the BiLSTM model to predict the residual Δx, and calculating the quadratic residual Δx. 2 x. The GWO (Grey Wolf) optimization algorithm is used to fine-tune the parameters of the BiLSTM model, resulting in a trained model. Historical test data is then used to fine-tune the model parameters to obtain the final trained model. Based on the final trained model, the short-term power load of the data center is predicted. This invention combines SARIMA and BiLSTM models, improving the accuracy of data center load forecasting and making it suitable for complex data center load forecasting tasks. SARIMA parameter tuning is performed using grid search and the AIC criterion. The SARIMA model built based on the tuned parameter values ​​better characterizes the linear and seasonal components of data center power load. The GWO optimization algorithm is used to fine-tune the parameters of the BiLSTM model. Considering that the number of neurons, the number of iterations, and the learning rate have a greater impact on the prediction accuracy of the BiLSTM model, the GWO optimization algorithm is used to optimize these three parameters. The BiLSTM model built based on the optimized parameter values ​​better characterizes the nonlinear components of data center power load.

[0072] In some embodiments, please refer to Figure 1 and Figure 5 The process of constructing the preset SARIMA and BiLSTM composite model is as follows:

[0073] Construct a pre-defined SARIMA model and a pre-defined BiLSTM model, and combine them into a training model;

[0074] The training model is fine-tuned by calculating the error of historical test data to obtain the final training model as the preset SARIMA and BiLSTM composite model.

[0075] In the complex scenario of data center load forecasting, the pre-set SARIMA model and the pre-set BiLSTM model are organically combined into a training model, and the composite model is constructed by dynamically fine-tuning the error through historical test data. This can effectively integrate the advantages of traditional time series analysis and deep learning, and provide key support for improving prediction accuracy and model robustness.

[0076] The SARIMA model, a classic time series analysis tool, is valuable for its ability to process non-stationary sequences through difference integration and capture the periodic patterns of data center load using a seasonal autoregressive moving average mechanism. Data center load naturally exhibits significant daily and weekly cyclical characteristics—server cluster computing power demands are concentrated during peak daytime business periods, while energy consumption decreases at night due to maintenance tasks or low-load operation; cooling system energy consumption fluctuates with outdoor temperature, forming a seasonal variation with an annual cycle. The SARIMA model, through parameter configuration, can accurately model these patterns with a seasonality order of s=24 corresponding to a daily cycle. However, its limitation lies in its delayed response to nonlinear abrupt changes.

[0077] The BiLSTM model, through its bidirectional long short-term memory network structure, overcomes the time-series modeling bottleneck of traditional recurrent neural networks. Its bidirectional propagation mechanism can simultaneously capture the forward dependencies of the load sequence, the impact of historical loads on the current value and their backward correlations, and the need to correct current predictions due to future changes in operating conditions. Furthermore, the dynamic adjustment function of the gating units (input gate, forget gate, and output gate) gives it a significant advantage in handling short-term fluctuations in data center loads. When a server cluster initiates an AI training task, causing a surge in instantaneous power consumption, BiLSTM can strengthen the feature weights at that moment through the input gate and weaken the influence of irrelevant historical information through the forget gate, thereby more accurately capturing abrupt changes.

[0078] When combining the pre-defined SARIMA and BiLSTM models into the initial training model, a weighted fusion strategy is used to construct the composite output: the SARIMA predictions serve as the base component, reflecting the periodic trend of the load; the BiLSTM predictions serve as the correction component, compensating for nonlinear fluctuations and abrupt changes. The initial weights are determined through cross-validation using historical data; for example, the SARIMA weight is increased during periods of stable business operations, while the BiLSTM contribution is increased during periods of fluctuating business operations. The prediction error of the composite model is calculated using historical test data (covering different seasons, business scenarios, and energy conditions), and the model parameters and fusion weights are dynamically adjusted using the gradient descent algorithm.

[0079] SARIMA ensures the long-term stability of predictions and avoids trend shifts caused by data noise in deep learning models; BiLSTM improves the response speed to short-term mutations and complex nonlinear relationships, making up for the dynamic adaptation defects of traditional time series models; and the fine-tuning mechanism based on historical errors enables the model to continuously learn the evolution of data center operation patterns and maintain the sustainability of prediction accuracy in long-term operation.

[0080] In some embodiments, please refer to Figure 1 and Figure 5 The process of constructing the preset SARIMA model is as follows:

[0081] Obtain historical data of data center load, group it into training dataset and test dataset, and determine the initial values ​​of SARIMA parameters;

[0082] A first-generation SARIMA model is constructed, and the linear part of the time series in the training dataset is fitted by the first-generation SARIMA model to generate the first predicted value.

[0083] Obtain the measured data corresponding to the first predicted value, and obtain the first residual sequence by subtracting the first predicted value from the measured data;

[0084] The SARIMA parameters are tuned according to grid search and AIC criteria until the residuals generated by two consecutive predictions are less than the first threshold. The SARIMA parameters used in the last prediction are then taken as the optimal values ​​of the SARIMA parameters.

[0085] Construct a pre-defined SARIMA model based on the optimal values ​​of SARIMA parameters.

[0086] Figure 5 As shown, the historical load data of the data center is grouped. The collected historical load power data of the data center is split into a model training dataset x and a model test dataset y in a 7:3 ratio.

[0087] Determine the initial values ​​of the SARIMA model parameters. Based on the power load characteristics of the data center, determine the initial value of the seasonal period s. Plot the ACF and PACF diagrams for the model training dataset x, and determine the initial values ​​of the SARIMA model parameters: number of autoregressive terms p, difference frequency d, moving average base q, number of seasonal autoregressive terms P, seasonal difference frequency D, and seasonal moving average base Q.

[0088] Construct a SARIMA model. Based on the established initial values ​​of each parameter, construct a SARIMA model using the statsmodels library.

[0089] The SARIMA model is used to fit the linear portion of the data center load data. The SARIMA model is used to fit the linear and seasonal components of the data center's electricity load within the time period [t1,t2] of the training data, generating the SARIMA model's predicted value x' for the [t1,t2] time period.

[0090] Calculate the residual sequence Δx. The residual sequence Δx is obtained by subtracting the measured data of data center power load during the time period [t1,t2] from the predicted value of the SARIMA model.

[0091] SARIMA parameter tuning was performed using grid search and the AIC criterion. The seven parameters of the SARIMA model were optimized and adjusted using grid search and the AIC criterion. If the difference between the residuals of two predictions was less than a certain threshold, the parameter value used in the last prediction was considered to be the optimal value. The SARIMA model built based on this parameter value can well characterize the linear and seasonal components of the data center's power load.

[0092] The linear portion of historical load data was fitted using the SARIMA model.

[0093] The SARIMA model is described as follows:

[0094]

[0095] In the formula, d and D represent the order of one-step differencing and seasonal differencing, respectively; p and P represent the order of non-seasonal autoregression and seasonal autoregression, respectively; q and Q represent the order of non-seasonal moving average and seasonal moving average, respectively; s represents the seasonal length; and x... t-i x t-ks Represent the historical data values ​​at the i-th and ks-th moments before time point t, respectively. t-j ,∈ t-ls Let φ represent the white noise error values ​​at the i-th and ks-th moments before time t, respectively. i Φ k Let θ represent the non-seasonal autoregressive coefficient and the seasonal autoregressive coefficient, respectively. j Θ l B and B' represent the non-seasonal moving average coefficient and the seasonal moving average coefficient, respectively. s Let x represent the non-seasonal delay operator and the seasonal delay operator, respectively, and let x be the value of x. t-p =B p x t ,

[0096] Calculate the residual sequence Δx, where the formula for calculating the residual sequence Δx is as follows:

[0097] △x=x-x'

[0098] SARIMA parameter tuning, where AIC can be represented as:

[0099] AIC = 2k - 2ln(L)

[0100] In the formula, k is the number of parameters, and L is the likelihood function.

[0101] Determining the initial values ​​of SARIMA parameters is a crucial starting point for model construction. The periodic characteristics of data center load typically manifest as daily fluctuations with a 24-hour cycle and weekly fluctuations with a 7-day cycle; therefore, the seasonality order s is generally set to 24. The initial parameter values ​​can be preliminarily determined by observing the autocorrelation function (ACF) and partial autocorrelation function (PACF) plots: the non-seasonal autoregressive order p and moving average order q are determined by the cutoff positions of PACF and ACF; the seasonal autoregressive order P and moving average order Q are determined by the seasonal lag order and the significance of lags of 24 and 48 periods; and the differencing order d and the seasonal differencing order D are set based on the results of the series stationarity test (ADF test) to ensure that the series satisfies the weak stationarity condition after differencing.

[0102] The core function of the initial SARIMA model, built upon initial values, is to fit the linear component of the training data. This model captures historical load dependencies through an autoregressive term, corrects for random errors through a moving average term, and eliminates trend components in periodic fluctuations through seasonal differencing, ultimately generating the first predicted value. Simultaneously, the measured data corresponding to this predicted value—the actual load values ​​at the same time point in the test set—must be acquired. The difference between the predicted and measured values ​​is used to obtain the first residual sequence. This residual sequence essentially captures the nonlinear fluctuation components that the initial SARIMA model failed to capture, including complex factors such as sudden changes in business traffic, equipment energy efficiency degradation, and fluctuations in distributed energy output, providing crucial feedback signals for subsequent parameter optimization.

[0103] The parameter tuning phase employs an iterative optimization strategy combining grid search and the AIC criterion. Grid search iterates through a pre-defined range of parameter combinations, calculating the AIC (Akaike Information Criterion) value for each parameter group. This criterion strikes a balance between goodness of fit and model complexity; a lower AIC value indicates a better model. After each iteration, the SARIMA model is reconstructed using the current optimal parameters, generating new predicted values ​​and calculating a new residual sequence. Optimization stops when the residuals of two consecutive predictions are both less than a pre-defined first threshold (set based on the data center's actual requirements for prediction accuracy, approximately 5% of the load baseline), and the last used parameters are taken as the optimal values.

[0104] Systematic optimization achieves both accuracy and robustness in linear trend prediction: initial parameter settings incorporate the physical characteristics and statistical regularities of data center load, avoiding the efficiency losses of blind search; the combined use of grid search and the AIC criterion balances model complexity and prediction accuracy, preventing overfitting; the residual-driven iterative optimization mechanism enables the model to dynamically adapt to changes in data center operating modes, such as the addition of new server clusters and cooling system upgrades. This model not only provides a reliable linear benchmark for subsequent composite modeling with BiLSTM, but its residual sequence can also serve as a training target for the nonlinear part, ultimately improving the adaptability and long-term stability of data center load prediction under complex operating conditions.

[0105] In some embodiments, please refer to Figure 1 and Figure 5 The process of constructing the preset BiLSTM model is as follows:

[0106] The first residual sequence is normalized to obtain the BiLSTM dataset, and the initial values ​​of the BiLSTM parameters are determined.

[0107] Construct an initial BiLSTM model and predict the second residual on the BiLSTM dataset;

[0108] A second predicted value is generated by fitting the nonlinear component in the training dataset using the first-generation BiLSTM model.

[0109] The second residual sequence is obtained by subtracting the second residual from the second predicted value;

[0110] The BiLSTM parameters are tuned according to the GWO optimization algorithm until the second residual is less than the second threshold after two consecutive predictions. The BiLSTM parameters used in the last prediction are then taken as the optimal values ​​of the BiLSTM parameters.

[0111] Construct a pre-defined BiLSTM model based on the optimal values ​​of the BiLSTM parameters.

[0112] The residual sequence Δx is normalized, and a BiLSTM dataset is created. The final residual sequence Δx is then normalized, and the normalized data is... Convert to the shape required for the BiLSTM model to serve as the dataset for the BiLSTM model.

[0113] Construct a BiLSTM model. Determine initial values ​​for parameters such as the number of neurons, number of iterations, learning rate, and dropout rate based on empirical practices, and then construct the BiLSTM model using the Keras library.

[0114] Predicting residuals using a BiLSTM model The BiLSTM model is used to fit the nonlinear component of the data center's power load during the time period [t1,t2] of the training data, generating the predicted value x of the BiLSTM model within the time period [t1,t2].

[0115] Calculate the second residual Δ 2 x. The residual data of data center power load during the time period [t1, t2]. The subtraction of the values ​​predicted by the BiLSTM model yields the quadratic residual sequence Δ. 2 x.

[0116] The GWO (Grey Wolf) optimization algorithm was used to fine-tune the parameters of the BiLSTM model. Considering that the number of neurons, the number of iterations, and the learning rate have a greater impact on the prediction accuracy of the BiLSTM model, the GWO algorithm was used to optimize these three parameters. If the difference between the residuals of two predictions is less than a certain threshold, the parameter values ​​used in the last prediction are considered optimal. The BiLSTM model built based on these parameter values ​​can well characterize the nonlinear component of the data center's power load.

[0117] The nonlinear component of historical load data was fitted using a BiLSTM model.

[0118] The BiLSTM model dataset is processed by processing the residual sequence Δx, denoted as z. The output of BiLSTM is obtained according to the following formula, which is composed of the hidden states of the forward and backward LSTMs at each time step t:

[0119]

[0120] For a forward LSTM, at time step t, the hidden state is... and memory unit The calculation is as follows:

[0121]

[0122] For a reverse LSTM, at time step t, the hidden state is... and memory unit The calculation is as follows:

[0123]

[0124]

[0125] In the formula, i, f, o, and c represent the input gate, forget gate, output gate, and memory unit in the LSTM model, respectively, and W, U, and b represent the input weight matrix, hidden state matrix, and bias term of each of the above stages.

[0126] Calculate the second residual Δ 2x, where the quadratic residual sequence Δ 2 The formula for calculating x is expressed as follows:

[0127] △ 2 x=△xx”

[0128] In the formula, x” represents the predicted value obtained by the BiLSTM model.

[0129] BiLSTM model parameter optimization includes three parts: initialization, fitness calculation, and position update.

[0130] The initialization formula is expressed as follows:

[0131] L={L i |i=1,2,...,N}

[0132] In the formula L i ={l i 1,l i 2,...l iD ,} represents the position of the i-th gray wolf.

[0133] The formula for calculating fitness is as follows: f(L) α ), f(L β ), f(L δ )

[0134] In the formula L α L β L and Lδ represent the positions of the best, second best, and third best fitness, respectively, and f(L i ) is the fitness function.

[0135] The position update formula is expressed as follows:

[0136] D α =|C1·L α -L|

[0137] D β =|C2·L β -L|

[0138] D δ =|C3·Lδ-L|

[0139] C i =2·r i

[0140] L1 = L α -A1·D α

[0141] L2 = L β -A2·D β

[0142] L3 = L δ -A3·D δ

[0143] A i =2·a·r i -a

[0144] L(t+1)=(L1+L2+L3) / 3

[0145] In the formula, a is the convergence factor, and as the number of iterations decreases linearly from 2 to 0, r takes a random number between [0,1].

[0146] In the construction of composite models for data center load forecasting, in-depth mining of the first residual sequence output by the SARIMA model is a key step in improving the prediction accuracy of the nonlinear part. Since the first residual sequence contains complex nonlinear characteristics such as sudden changes in business traffic, equipment energy efficiency degradation, and fluctuations in distributed energy output, its numerical range often spans a large range and contains extreme values. Directly inputting it into the neural network can easily lead to vanishing or exploding training gradients. Therefore, the Min-Max normalization method is first used to map the residual sequence to the [0,1] interval. This processing converts each data point into the form (current value - minimum value) / (maximum value - minimum value) by calculating the minimum and maximum values ​​of the sequence. This preserves the distribution characteristics of the original data and eliminates the interference of dimensional differences on the neural network weight updates, providing stable data input for the BiLSTM model.

[0147] Determining the initial values ​​of BiLSTM parameters requires considering the nonlinear characteristics of the data center load and the model complexity. The number of neurons in the input layer is typically related to the size of the time window, which needs to be sufficient to cover the lag effect of load changes—the computing power demand of the server cluster changes with a delay of 1-3 time steps, which is reflected in the cooling system's energy consumption. Therefore, a time window of 4 can be set, corresponding to 1 hour of data, assuming a data sampling interval of 15 minutes, so that the model can capture short-term dependencies. The number of neurons in the hidden layer needs to balance feature extraction capability with the risk of overfitting, initially set to 64. This value allows the gating unit to fully learn the mutation patterns in the residual sequence without reducing training efficiency due to too many parameters. The number of neurons in the output layer is fixed at 1, corresponding to the normalized residual prediction value. The setup of the bidirectional propagation mechanism needs to consider the bidirectional temporal correlation of data center load: the forward LSTM captures the impact of historical residuals on the current value, the cooling system efficiency decline in the first 3 time steps leads to an increase in current energy consumption, the backward LSTM mines the correction needs of future operating condition changes on the current prediction, and the increase in distributed energy output in subsequent time steps can partially offset the current load prediction deviation, and the temporal feature representation is enhanced by splicing bidirectional hidden states.

[0148] The core task of the initial BiLSTM model, built upon initial values, is to fit the nonlinear components of the training data. This model controls the inflow of new information through an input gate, handling sudden spikes in traffic flow corresponding to abrupt changes in the residuals. A forget gate filters historical information, weakening stationary residuals from irrelevant time steps. An output gate generates the current predicted value, ultimately outputting a second predicted value. The difference between this predicted value and the measured normalized residual is then calculated, yielding a secondary residual sequence. This sequence reflects the initial BiLSTM model's ability to capture complex nonlinear features. If significant fluctuations still exist in the secondary residuals, it indicates that the model needs further parameter optimization to improve its adaptability to extreme conditions, such as extreme temperatures causing the cooling system to operate at full load.

[0149] The parameter tuning phase employs the GWO (Grey Wolf) optimization algorithm, which balances global search and local exploitation by simulating the hunting behavior of a grey wolf pack. In the BiLSTM parameter optimization, α, β, and δ wolves represent the current optimal, second-best, and third-best solutions, respectively, while the remaining wolves (ω) update their positions by learning from these three. The wolf positions correspond to parameter combinations, the number of hidden layer neurons, the learning rate, and the batch size. The fitness function is set as the mean squared error (MSE) of the quadratic residuals, and the optimization objective is to minimize the MSE. In each iteration, ω wolves dynamically adjust their search direction based on the positions of α, β, and δ wolves. This avoids getting trapped in local optima, prevents excessively large learning rates that could cause training oscillations, and accelerates convergence to the global optimum. The number of hidden layer neurons is matched to the data complexity. When the quadratic residuals of two consecutive predictions are both less than a preset second threshold (set according to the data center's requirements for the accuracy of nonlinear predictions), and the normalized residuals reach 0.05 standard deviations, optimization stops, and the last parameter is taken as the optimal value.

[0150] The final pre-built BiLSTM model based on optimal parameters offers advantages in achieving high efficiency and robustness in nonlinear prediction through normalization and GWO optimization: normalization eliminates the impact of data distribution differences on training stability; reasonable settings of the time window and hidden layer size balance model complexity and feature extraction capability; the bidirectional propagation mechanism enhances the capture of temporal dependencies; and the global search capability of the GWO algorithm prevents parameters from getting trapped in local optima, while its convergence speed is superior to traditional genetic algorithms, significantly improving optimization efficiency. This model can not only accurately predict nonlinear fluctuations missed by the SARIMA model, but its optimized parameters can also provide a basis for the weight allocation of subsequent composite models. The weights in the composite output are dynamically adjusted based on the prediction error variance of the BiLSTM, ultimately constructing a high-precision load prediction framework that combines linear trend interpretability with adaptability to nonlinear fluctuations.

[0151] In some embodiments, please refer to Figure 1 and Figure 5The process of fine-tuning the training model by calculating the error of historical test data to obtain the final training model includes:

[0152] Input the test dataset into the training model to obtain the historical test data error;

[0153] The training model is fine-tuned based on the historical test data to obtain the final training model.

[0154] The trained model is obtained. A SARIMA model is constructed based on the obtained optimized parameters, and a BiLSTM model is constructed based on the obtained optimized parameters. Together, they form a trained model for short-term load forecasting of data centers.

[0155] The model parameters are fine-tuned using historical test data to obtain the final trained model. The obtained model test dataset y is then substituted into the obtained trained model to obtain the predicted values ​​and prediction errors within the time period [t3, t4] of the test data.

[0156] The final trained model is used to predict the short-term power load of the data center. A time period is selected for prediction, and the resulting trained model obtains the predicted power load of the data center during that period, providing support for optimizing the data center's power supply system.

[0157] In the process of building a composite model for data center load forecasting, inputting the test dataset into the already built training model and calculating the error of historical test data is a key verification step to ensure the model's generalization ability and adaptability to real-world applications. A fine-tuning mechanism based on error feedback can further improve the model's prediction accuracy under complex operating conditions. After the initial parameter optimization of the SARIMA and BiLSTM composite training model is completed, it needs to be deployed on an independent test dataset for full-process verification. This dataset needs to cover a different time period than the training set; the training set contains data from the previous two years, and the test set contains data from the following year, and includes diverse operating scenarios, such as the summer high-temperature period, holiday off-peak periods, and periods of high distributed energy penetration, to test the model's adaptability to unseen data.

[0158] When inputting test data, the data preprocessing workflow must be consistent with that of the training phase: For the SARIMA component, the test sequence must undergo the same differencing process as the training set, including first-order conventional differencing and first-order seasonal differencing, to ensure that the seasonal cycle aligns with the training phase; for the BiLSTM component, the linear prediction residuals generated by SARIMA must be normalized and sliced ​​over the same four time windows to maintain consistency in the input dimensions. The composite model's prediction process employs a hierarchical output mechanism: SARIMA first generates basic linear prediction values, reflecting the periodic trend of the load; BiLSTM then performs nonlinear correction on the SARIMA residuals, generating dynamic adjustment amounts; the final prediction value is a weighted sum of the two, with weights determined through cross-validation using historical data. During stable periods, the SARIMA weight is 0.7, while during volatile periods, the BiLSTM weight increases to 0.6.

[0159] The calculation of historical test data error needs to cover multiple dimensions: absolute error reflects the absolute deviation between the predicted value and the measured value, and can intuitively assess the prediction deviation of the model under extreme conditions; mean square error amplifies the impact of larger errors through the squared term, highlighting the model's ability to capture sudden loads; mean absolute percentage error eliminates the difference in dimensions, making it easier to compare model performance across data centers.

[0160] The fine-tuning mechanism based on historical test data errors employs a dynamic feedback optimization strategy. First, the error distribution characteristics are analyzed: if the error exhibits periodic fluctuations, with larger errors at fixed daily time periods, the seasonal parameters of SARIMA are adjusted, increasing the seasonal difference order; if the error is strongly correlated with business traffic, with error peaks coinciding with business peaks, the input of business characteristic variables to BiLSTM is strengthened, adding auxiliary inputs such as server utilization and task queue length; if the error exhibits a systematic shift, with predicted values ​​consistently lower than measured values, the weighted fusion rules of the composite model are adjusted, increasing the weight of BiLSTM during periods of fluctuation. The fine-tuning process uses mini-batch gradient descent, updating only a subset of key parameters, BiLSTM hidden layer weights, or SARIMA moving average coefficients each time, avoiding model oscillations caused by full parameter updates.

[0161] The final trained model achieved closed-loop optimization of predictive capabilities through test set validation and error feedback: the independent test set eliminated the interference of overfitting in the training data, ensuring that the error metric truly reflects the model's generalization ability; multi-dimensional error analysis accurately located the model's weaknesses, providing clear direction for parameter fine-tuning; the dynamic feedback mechanism enabled the model to continuously learn the evolution of data center operating modes, including energy efficiency degradation due to equipment aging and load pattern changes caused by business structure adjustments, maintaining the sustainability of predictive accuracy in long-term operation. This model not only provides a more reliable load forecast basis for the real-time scheduling of data center energy management systems, but its error-driven optimization framework can also be extended to other time-series forecasting scenarios, such as renewable energy output forecasting and industrial equipment fault early warning, providing a generalized solution for complex system modeling.

[0162] In some embodiments, please refer to Figure 1 and Figure 5 The process of dynamically combining the linear and nonlinear forecast values ​​into a load forecast value includes:

[0163] The linear and nonlinear predicted values ​​are standardized and concatenated to obtain standardized features.

[0164] The standardized features are input into the conditional encoder of GNF to obtain the latent spatial distribution parameters;

[0165] Dynamic weights are obtained by sampling from the potential spatial distribution parameters;

[0166] The load forecast value is obtained by weighting the linear and nonlinear forecast values ​​according to the dynamic weights.

[0167] In the optimization of composite models for data center load forecasting, standardizing and concatenating linear and nonlinear forecasts and constructing a dynamic weighting mechanism is a key innovation to enhance the model's adaptability to complex temporal characteristics. This method integrates the advantages of both forecasting paradigms, retaining the linear model's ability to stably capture periodic trends while strengthening the nonlinear model's response to sudden fluctuations and complex dependencies, ultimately generating forecast results that better reflect actual load changes.

[0168] Standardization is performed on both linear and nonlinear predicted values. Linear predicted values ​​are typically generated by statistical models such as SARIMA, and their numerical range is constrained by the mean and variance of historical data, failing to directly reflect the drastic nature of recent load changes. Nonlinear predicted values ​​are generated by neural networks such as BiLSTM, and their output distribution differs from linear predicted values ​​due to the characteristics of the activation function, with the Sigmoid output ranging from [0,1]. Therefore, the Z-score standardization method is used to calculate the mean and standard deviation of both predicted values, converting each value into the form (current value - mean) / standard deviation. This process not only eliminates the dimensional differences but also ensures that the standardized features have zero mean and unit variance, providing stable input for subsequent latent space modeling.

[0169] After standardized features are input into the conditional encoder of a GNF (Generative Flow Network), the model maps these features to the latent space through multiple layers of nonlinear transform fully connected layers and stacked activation functions. The key to this process lies in the design of the conditional encoder: its input not only includes standardized linear and nonlinear predictions but also incorporates auxiliary variables such as traffic flow and ambient temperature, enabling the latent space distribution to comprehensively reflect the interaction of multiple influencing factors. When a sudden surge in traffic leads to increased cooling system load, the conditional encoder can capture the coordinated changes in linear and nonlinear predictions, thereby generating more accurate latent space parameters.

[0170] The latent space distribution parameters typically include a mean vector and a covariance matrix. The former describes the central location of the latent variables, while the latter characterizes the correlation between variables. When sampling in the latent space, a reparameterization technique is employed: random noise is generated from the standard normal distribution, and through a linear transformation, the product of the mean, the square root of the covariance matrix, and the noise is converted into samples conforming to the target distribution. This technique makes the sampling process differentiable, supporting end-to-end training. The dynamic weights obtained from sampling are essentially linear combinations of the latent variables, and their numerical range is constrained to [0,1] by the Sigmoid function to ensure the rationality of weight allocation. When the latent variable corresponding to the nonlinear prediction value is high, the dynamic weight tends to be 0.7, allowing the nonlinear component to dominate in the final prediction.

[0171] The final load forecast is generated by dynamically weighting and summing the linear and nonlinear forecasts. The advantage of this mechanism lies in its scenario-adaptive weight allocation: during stable operation periods, linear forecasts receive higher weight due to their stable trends; during periods of business fluctuation, nonlinear forecasts dominate because they capture abrupt changes. This adaptive adjustment capability allows the model to automatically balance the contributions of the two forecast paradigms without requiring manual threshold or rule settings, significantly improving forecast accuracy under complex operating conditions.

[0172] In some embodiments, please refer to Figure 1 and Figure 5 The process of obtaining the time period to be predicted and obtaining time series data based on the time period to be predicted includes:

[0173] Determine the historical data time window based on the time period to be predicted;

[0174] Retrieve historical load data within the historical data time window from the database;

[0175] The time period to be predicted and the historical load data are combined into time series data.

[0176] In data center load forecasting, dynamically determining the historical data time window and constructing time series data based on the forecast period is a key preprocessing step to improve the accuracy and adaptability of the model. This method provides more targeted input to the model by accurately matching the correlation features between historical data and the forecast period, thereby enhancing its ability to capture load change patterns.

[0177] The characteristics of the time period to be predicted determine the length and range of the historical data time window. The setting of the time window length needs to comprehensively consider the periodicity of load changes and model complexity: if the data center load exhibits a clear daily periodicity, with high traffic during the day and low traffic at night, the time window is typically set to 7 days to fully cover the fluctuation pattern within a cycle; if there is a weekly periodicity, with weekend load significantly lower than weekday load, the window can be extended to 4 weeks to ensure the model can learn cross-weekly dependencies. The starting point of the time window needs to maintain a reasonable interval from the time period to be predicted: for short-term predictions of the next 24 hours, the window can be immediately adjacent to the time period to be predicted, taking data from the previous 7 days to capture the most recent load change trend; for medium- to long-term predictions of the next 7 days, the window needs to be appropriately shifted forward to take data from the previous 4 weeks to avoid abnormal fluctuations in recent data misleading the model.

[0178] When querying historical load data from the database, it is essential to ensure data quality and integrity, and optimize for the characteristics of data center load data: For time steps with missing values, linear interpolation can be used to fill them in, or the time step can be directly removed and subsequent time indexes adjusted; for outliers, corrections should be made in conjunction with business logic, checking whether they are false alarms caused by equipment failure, or replacing extreme values ​​with the median. Query results should include auxiliary variables related to the load, such as ambient temperature, business traffic, and equipment operating status. If these variables are strongly correlated with the load, they should be used as additional features input to the multidimensional time series model to enhance feature representation capabilities.

[0179] When combining the forecast period with historical load data into time series data, it is necessary to maintain the continuity and consistency of the time dimension. The combined data structure is usually a three-dimensional tensor, with the sample size × time step × number of features: for a single forecast period, the sample size is 1; the time step is determined by the length of the historical window, with 168 steps corresponding to 7 days of data; the number of features includes historical load values ​​and auxiliary variables temperature and flow. If the historical window is 7 days, the sampling interval is 1 hour, and it includes two features, load and temperature, then the shape of the input tensor is 1×168×2. The data for the forecast period needs to be processed separately: for short-term forecasts, it can be used as the target variable to predict the load for the next 24 hours, then the target variable shape is 1×24×1; for autoregressive models, the first few time steps of the forecast period need to be used as initial inputs to gradually generate future forecast values.

[0180] The advantage of this preprocessing method lies in its ability to focus the model on historical information most relevant to the period to be predicted through dynamic time window selection and data combination. When predicting cooling system load during the summer high-temperature period, the model can automatically select historical high-temperature data from the same period as input, avoiding interference from low-temperature data in winter. When predicting server cluster load before a surge in business traffic, the model can learn load response patterns from historical data during periods of sudden traffic changes, thereby generating accurate predictions in advance.

[0181] This invention proposes a short-term power load forecasting system for data centers. Please refer to [link / reference]. Figure 2 ,include:

[0182] The acquisition module 100 is configured to acquire a time period to be predicted and obtain time series data based on the time period to be predicted.

[0183] The first prediction module 200 is configured to input the time series data into a preset SARIMA and BiLSTM composite model to obtain linear prediction values ​​and linear residual sequences.

[0184] The second prediction module 300 is configured to normalize the linear residual sequence to obtain a dataset, and generate nonlinear prediction values ​​and nonlinear residual sequences based on the dataset.

[0185] The combination module 400 is configured to dynamically combine the linear and nonlinear forecast values ​​into a load forecast value, and output the load forecast value and the nonlinear residual sequence.

[0186] This invention aims to effectively improve the accuracy of short-term load forecasting for data centers. Based on this, the data center power supply system can rationally arrange power supply and energy storage charging and discharging plans, thereby improving the local consumption level of new energy sources and the proportion of green electricity in the load, and accelerating the promotion of energy conservation and emission reduction in the field of new infrastructure for data centers.

[0187] Based on the same inventive concept, according to another aspect of the present invention Figure 3 As shown, an embodiment of the present invention also provides a computer device 30, which includes a processor 310 and a memory 320. The memory 320 stores a computer program 321 that can run on the processor. When the processor 310 executes the program, it performs the steps of the method.

[0188] Based on the same inventive concept, according to another aspect of the present invention Figure 4 As shown, embodiments of the present invention also provide a computer-readable storage medium 40, which stores a computer program 410 that executes the above method when executed by a processor.

[0189] Embodiments of the present invention may also include a corresponding computer device. The computer device includes a memory, at least one processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes any of the methods described above when executing the program.

[0190] The memory, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, including the program instructions / modules in this embodiment. The processor executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory, thus implementing the above-described method.

[0191] The memory may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the device. Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In embodiments, the memory may optionally include memory remotely located relative to the processor, which can be connected to the local module via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0192] Finally, it should be noted that those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium for the program can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. The above computer program embodiments can achieve the same or similar effects as any of the corresponding foregoing method embodiments.

[0193] Those skilled in the art will also understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the functionality of various illustrative components, blocks, modules, circuits, and steps has been generally described. Whether this functionality is implemented as software or as hardware depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art can implement the functionality in various ways for each specific application, but such implementation decisions should not be construed as departing from the scope of the embodiments disclosed herein.

[0194] The above are exemplary embodiments disclosed in this invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments of this invention as defined by the claims. The functions, steps, and / or actions of the methods according to the disclosed embodiments described herein do not need to be performed in any particular order. The sequence numbers of the disclosed embodiments of this invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments. Furthermore, although the elements disclosed in the embodiments of this invention may be described or claimed individually, they may be understood as multiple unless explicitly limited to a singular number.

[0195] It should be understood that, as used herein, the singular form “a” is intended to include the plural form as well, unless the context clearly supports an exception. It should also be understood that, as used herein, “and / or” means any and all combinations of one or more of the associated listed items.

[0196] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of the invention (including the claims) is limited to these examples. Within the framework of the invention, technical features of the above embodiments or different embodiments can be combined, and many other variations of different aspects of the invention exist, which are not provided in the details for the sake of brevity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the protection scope of the invention.

Claims

1. A method for short-term power load forecasting of data centers, characterized in that, include: Obtain the time period to be predicted, and obtain time series data based on the time period to be predicted; The time series data is input into a preset SARIMA and BiLSTM composite model to obtain linear predicted values ​​and linear residual sequences. The linear residual sequence is normalized to obtain a dataset, and nonlinear predicted values ​​and nonlinear residual sequences are generated based on the dataset. The linear and nonlinear predicted values ​​are dynamically combined into a load predicted value, and the load predicted value and the nonlinear residual sequence are output.

2. The method for short-term power load forecasting of a data center according to claim 1, characterized in that, The process of constructing the pre-defined SARIMA and BiLSTM composite model is as follows: Construct a pre-defined SARIMA model and a pre-defined BiLSTM model, and combine them into a training model; The training model is fine-tuned by calculating the error of historical test data to obtain the final training model as the preset SARIMA and BiLSTM composite model.

3. The method for short-term power load forecasting of a data center according to claim 2, characterized in that, The process of constructing the preset SARIMA model is as follows: Obtain historical data of data center load, group it into training dataset and test dataset, and determine the initial values ​​of SARIMA parameters; A first-generation SARIMA model is constructed, and the linear part of the time series in the training dataset is fitted by the first-generation SARIMA model to generate the first predicted value. Obtain the measured data corresponding to the first predicted value, and obtain the first residual sequence by subtracting the first predicted value from the measured data; The SARIMA parameters are tuned according to grid search and AIC criteria until the residuals generated by two consecutive predictions are less than the first threshold. The SARIMA parameters used in the last prediction are then taken as the optimal values ​​of the SARIMA parameters. Construct a pre-defined SARIMA model based on the optimal values ​​of SARIMA parameters.

4. The method for short-term power load forecasting of a data center according to claim 3, characterized in that, The process of constructing the preset BiLSTM model is as follows: The first residual sequence is normalized to obtain the BiLSTM dataset, and the initial values ​​of the BiLSTM parameters are determined. Construct an initial BiLSTM model and predict the second residual on the BiLSTM dataset; A second predicted value is generated by fitting the nonlinear component in the training dataset using the first-generation BiLSTM model. The second residual sequence is obtained by subtracting the second residual from the second predicted value; The BiLSTM parameters are tuned according to the GWO optimization algorithm until the second residual is less than the second threshold after two consecutive predictions. The BiLSTM parameters used in the last prediction are then taken as the optimal values ​​of the BiLSTM parameters. Construct a pre-defined BiLSTM model based on the optimal values ​​of the BiLSTM parameters.

5. A method for short-term power load forecasting of a data center according to claim 4, characterized in that, The process of fine-tuning the training model by calculating the error of historical test data to obtain the final training model includes: Input the test dataset into the training model to obtain the historical test data error; The training model is fine-tuned based on the historical test data to obtain the final training model.

6. The method for short-term power load forecasting of a data center according to claim 1, characterized in that, The process of dynamically combining the linear and nonlinear forecast values ​​into a load forecast value includes: The linear and nonlinear predicted values ​​are standardized and concatenated to obtain standardized features. The standardized features are input into the conditional encoder of GNF to obtain the latent spatial distribution parameters; Dynamic weights are obtained by sampling from the potential spatial distribution parameters; The load forecast value is obtained by weighting the linear and nonlinear forecast values ​​according to the dynamic weights.

7. The method for short-term power load forecasting of a data center according to claim 1, characterized in that, The process of obtaining the time period to be predicted and obtaining time series data based on the time period to be predicted includes: Determine the historical data time window based on the time period to be predicted; Retrieve historical load data within the historical data time window from the database; The time period to be predicted and the historical load data are combined into time series data.

8. A short-term power load forecasting system for data centers, characterized in that, include: The acquisition module is configured to acquire the time period to be predicted and obtain time series data based on the time period to be predicted. The first prediction module is configured to input the time series data into a preset SARIMA and BiLSTM composite model to obtain linear prediction values ​​and linear residual sequences. The second prediction module is configured to normalize the linear residual sequence to obtain a dataset, and generate nonlinear prediction values ​​and nonlinear residual sequences based on the dataset. The combination module is configured to dynamically combine the linear and nonlinear forecast values ​​into a load forecast value, and output the load forecast value and the nonlinear residual sequence.

9. A computer device, comprising: At least one processor; The processor also includes a memory storing a computer program that can run on the processor, characterized in that the processor executes the program to perform the steps of the short-term power load forecasting method for a data center as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it performs the steps of the short-term power load forecasting method for data centers as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Intelligent fusion terminal with operation fault accurate prediction function

    CN122153845A