A data center energy-saving time sequence prediction method based on deep reinforcement learning
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2025-05-28
- Publication Date
- 2026-08-07
AI Technical Summary
在空调机房智能运维的场景中,当前数据中心和机房的空调系统能耗高,传统的固定规则和阈值控制方式难以适应复杂动态的运行环境,导致能源浪费严重
[0081] 1. By building components based on the domestically produced MWORKS for modeling, data prediction, and reinforcement learning, we have broken the monopoly of foreign MATLAB and achieved self-reliance and strength in science and technology for our motherland.
Smart Images

Figure CN120730686B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent air conditioning energy-saving control technology, and in particular to a data center energy-saving time-series prediction method based on deep reinforcement learning. Background Technology
[0002] With the frequent bans on foreign software like MATLAB in China, the development of domestic alternatives has become particularly urgent. Domestic researchers are actively responding to the government's strong support for domestic software, exploring its use in scientific research and engineering applications. MWORKS, a domestically developed alternative to MATLAB, has already been applied in multiple fields. Its functionality already meets the needs of some users, especially in system modeling and simulation, where it has demonstrated impressive performance. Through the application of domestic software, enterprises can achieve more refined energy management, improve equipment operating efficiency, and reduce energy waste. Simultaneously, domestic researchers are actively exploring simulation modeling and intelligent control strategies based on domestic software, aiming to achieve or even surpass the performance of foreign software.
[0003] With the development of domestically produced industrial software, it will play an increasingly important role in the industrial ecosystem of intelligent manufacturing and high-end manufacturing. In the scenario of intelligent operation and maintenance of air-conditioned computer rooms, current data center and computer room air conditioning systems have high energy consumption. Traditional fixed rules and threshold control methods are difficult to adapt to complex and dynamic operating environments, leading to serious energy waste. While air-cooled systems are simple in structure and easy to maintain, their energy efficiency is low, while water-cooled systems, although more energy efficient, are complex and require higher levels of operation and maintenance expertise. As data centers expand, fluctuations in computer room temperature and humidity, changes in IT load, and environmental factors make it difficult for traditional control methods to accurately match real-time demands, resulting in frequent start-ups and shutdowns of air conditioning systems, excessive energy consumption, and even affecting the stable operation of equipment. Furthermore, traditional energy management methods lack intelligent predictive capabilities, relying solely on statically set operating parameters and failing to dynamically adjust based on real-time data. This leads to persistently high PUE (Power Usage Effectiveness) and rising operation and maintenance costs year after year, necessitating more advanced energy-saving optimization technologies.
[0004] Existing outdated energy-saving algorithms primarily rely on simple PID control, empirical threshold settings, and limited logic control strategies, failing to adequately consider the complex thermal environment and energy consumption characteristics of data centers. These methods typically start and stop air conditioning based on fixed temperature setpoints, neglecting precise optimization incorporating dynamic environmental factors, leading to over-operation or under-response of air conditioning. Furthermore, while some data centers have introduced limited automatic adjustment strategies, such as timed on / off switching and load balancing, they still lack adaptive capabilities, struggling to cope with IT load fluctuations, external environmental changes, and the demands of multi-device collaborative operation. This static rule-based control approach not only has limited energy-saving effects but may also cause localized overcooling or overheating due to response lag. Therefore, with the increasing demand for intelligent data centers, the limitations of outdated energy-saving algorithms are becoming increasingly apparent, urgently requiring the introduction of AI-driven intelligent energy-saving optimization technologies to achieve more precise and efficient air conditioning energy consumption management.
[0005] Therefore, this invention mainly studies the intelligent energy-saving optimization of air conditioning rooms, which is mainly achieved through physical modeling of the room environment, deep learning time series prediction, and reinforcement learning intelligent control. Through these methods, the research will construct an intelligent air conditioning control system that can dynamically optimize air conditioning operating parameters, ensure temperature and humidity meet standards, and minimize energy consumption. Summary of the Invention
[0006] This invention proposes an intelligent energy-saving control system for air-conditioning computer rooms, constructing a closed-loop control architecture of "simulation-prediction-decision". Relying on the MWORDS scientific computing and system modeling simulation platform, it completes the process of building a reinforcement learning simulation environment for AI air-conditioning energy-saving algorithms, algorithm development, model training, testing and optimization, and integrated application. This supports the development of the domestic scientific computing and system modeling simulation ecosystem and realizes the iterative upgrade of the original air-conditioning energy-saving algorithm.
[0007] The research content is divided into three parts. First, a simulation model of the air conditioning room is established. Partial differential equations are solved simultaneously using historical data from the room to calculate the current temperature distribution and air conditioning energy consumption. Based on this, a visualization simulation system is built to intuitively display the temperature field and energy consumption, providing a foundation for subsequent intelligent control. Second, deep learning time-series prediction models, including but not limited to iTransformer and FFT, are used to predict the room's temperature and energy consumption. By capturing the long-term trends and short-term fluctuations of time-series data, temperature changes are predicted in advance, enabling the system to optimize air conditioning operation strategies based on the predicted data and improve energy efficiency. Finally, reinforcement learning-based intelligent air conditioning control is implemented using the predicted data. Reinforcement learning algorithms, including but not limited to VDN and TD3, are used to control discrete control signals such as the start-up, shutdown, and temperature adjustment of air conditioning equipment, achieving intelligent temperature regulation and forming an end-to-end energy-saving optimization closed loop from virtual simulation and intelligent decision-making to physical execution. The specific steps are as follows:
[0008] I. Establishment of a Simulation Model for the Computer Room Environment
[0009] This invention utilizes Sysplorer, a domestic alternative to MATLAB, on the MWORKS platform to build an air-conditioned room model. The purpose is to build an air-conditioned model and integrate it into the room model. After making decisions through deep learning, the air-conditioned temperature is adjusted by changing the power.
[0010] Firstly, this invention uses a closed-loop air conditioning loop component from the Sysplorer standard library TAThermalSystem (vehicle thermal management model library) to cool the air in the loop. The principle is as follows.
[0011] The Compressor R134a component is used to simulate the process by which a compressor increases the pressure and temperature of the refrigerant through mechanical work. Simultaneously, a power monitoring sensor reads the change in the air conditioning engine's power over time, and the Integrator component integrates the power over time to obtain the air conditioning energy consumption. The compressor's performance equation is described as follows:
[0012] W=∫pdt
[0013]
[0014] h out =h in +Δh comp ,
[0015] Where W represents energy consumption, p represents power, and t represents time. out and p in These are the compressor's inlet and outlet pressures, h. in h out The specific enthalpy of the refrigerant flowing into / out of the compressor. η is the mass flow rate of the refrigerant, η is the compressor efficiency, and Δh is the mass flow rate of the refrigerant. comp It is the increase in specific enthalpy during the compression process. out p in , Δh comp All parameters can be manually set in the model through input. The formula for calculating the compressor efficiency η is:
[0016] η = η vol ×η isentropic ×η mech ,
[0017] Where η vol (volume efficiency), η isentropic (Isoentropy efficiency), η mech(Mechanical efficiency) is a settable parameter, and different emphases are applied depending on the actual situation.
[0018] The Condenser component was then used to simulate the process in the condenser where the refrigerant releases heat and condenses into a liquid through heat exchange with ambient air. The heat exchange equation of the condenser describes the heat transfer between the refrigerant and the air:
[0019]
[0020] Q hot For heat exchange quantity, This refers to the mass flow rate of the refrigerant. The heat exchange efficiency of the condenser is typically described by the logarithmic mean temperature difference.
[0021]
[0022] ΔT1=ΔT ref,in -ΔT air,out ,
[0023] ΔT1=ΔT ref,out -ΔT air,in .
[0024] Where, ΔT ref,in ΔT ref,out ΔT represents the inflow and outflow temperatures of the condensate. air,out ΔT air,in The temperature of the air flowing in and out.
[0025] High-pressure liquid refrigerant passes through the simpleTXV component, where a simulated expansion valve converts it into low-pressure liquid refrigerant through a throttling effect. The throttling principle is as follows:
[0026]
[0027] Where C d Let ρ be the flow coefficient, A be the effective flow area of the expansion valve, ρ be the density of the refrigerant, and ΔP be the pressure difference across the expansion valve. The opening degree y of the expansion valve is controlled by the superheat ΔT. sh :
[0028]
[0029] Where k is the proportionality coefficient, T i is the integration time constant.
[0030] Low-temperature, low-pressure liquid refrigerant is evaporated into a gas after passing through the evaporator assembly (Evaporator R134a), absorbing heat from the surrounding environment to achieve a cooling effect. The evaporator also uses the heat transfer equation and calculates the superheat ΔT. shThis provides key parameters for controlling the opening degree of the expansion valve:
[0031] ΔT sh =ΔT ref,out -T sat ,
[0032] Where T sat This is the saturation temperature of the refrigerant.
[0033] After connecting all components, the set parameters can be changed by the components and transmitted between them to realize the circulation of the air conditioning circuit.
[0034] Air processed by the air conditioner flows out through the air_out1 interface and enters the fan's air inlet a. The fan then sends the air through the air outlet b to the sensor pTSensorAir. The sensor monitors the air temperature and pressure and directs the air to the air resistance airResist. The air resistance simulates flow resistance and directs the air to the room's air inlet airPortIn. After circulating within the room, the air returns to the air conditioner's air inlet air_in1 through the air outlet airPortOut.
[0035] In this process, the present invention uses a PID controller to control the air conditioner's speed and converts it into power:
[0036]
[0037] Where u(t) is the control output (air conditioner speed), e(t) is the error signal (difference between the set value and the measured value), and k p k i k d These are the proportional, integral, and differential coefficients.
[0038] The air conditioner speed is converted from the output of the PID controller to the actual speed via a gain module.
[0039]
[0040] The airflow rate and pressure are determined by the characteristics of the fan and air resistance, as well as the built-in functions of the fan assembly:
[0041]
[0042] Next, by further building a room model, we can simulate the temperature changes after the air conditioner delivers air into the room.
[0043] To build the room model, the air volume is first defined using the airVolumes component from the TYBase library, and the parameters len, hei, and wid are defined as the length, height, and width of the room. The volume parameter V of the component... tot= len × wid × hei.
[0044] The PlaneWall component from the TYThermoFluidSys library is used to define the walls of the room. The solid (material), th (thickness), and surfaceArea (surface area) parameters of the wall are set according to the actual situation. The PlaneWall exchanges heat with airVolumes to simulate the heat conduction between the wall and the air.
[0045] The heat exchange between air and the wall satisfies the following equation:
[0046] Q conv =h×A×(T) air -T wall )
[0047] Among them, Q conv The heat exchange rate between the air and the wall is given by: h = (A / T) * (h / A ... air air temperature T wall This refers to the wall temperature.
[0048] The HeatCapacitor component from TYBase is used as the heat capacity and connected to the wall. The values of C (heat capacity) and T are considered. start The initial temperature was set to simulate the heat storage capacity of the walls. The Adiabatic component from TYThermoFluidSys was used as the adiabatic boundary, which was connected to each of the six walls to prevent excessive heat loss from the boundaries.
[0049] Finally, the heatTransfer component from the TAThermalSystem library and the HeatCapacitor component from TYBase are used to construct a heat transfer system connected to airVolumes to represent the heat transfer process between the room and the outside. HeatCapacitor is used to simulate the external heat capacity. The Aheatloss (heat transfer area) and k (thermal conductivity) in the heatTransfer component can be set according to the actual situation.
[0050] The airVolumes component, used to simulate air volume in the room, uses airPortIn and airPortOut interfaces to simulate air entering and exiting the room, respectively, meaning air exchanges with the outside environment through these ports. These two ports can be used to establish a connection with the indoor air conditioning model, forming a computer room environment simulation model. After connecting all components, indoor temperature changes can be simulated.
[0051] The heat conduction of the wall satisfies the following equation:
[0052]
[0053] Among them, Q cond The heat flux through the wall is given by k, thermal conductivity is given by A, surface area of the wall is given by d, and thickness of the wall is given by T. in With T out This represents the temperature inside and outside the room. It can be used to calculate heat transfer between the room and the outside environment.
[0054] The room model and the indoor air conditioning model can be connected through an airflow interface to form a computer room environment simulation model. After simulating indoor temperature changes, time-series data such as temperature, humidity, and energy consumption can be obtained and imported into a deep learning model.
[0055] In the physical modeling process, 57 features of temperature, humidity and energy consumption data of different air conditioners detected by different detectors are used as parameters for model input. After passing through the algorithm, a predicted target feature sequence is output to the reinforcement learning part for intelligent decision-making.
[0056] II. Time Series Prediction Model Based on Deep Learning
[0057] This invention uses a deep learning-based time series prediction model to predict sensor temperature, humidity, and air conditioning energy consumption. Various deep learning time series prediction methods such as RNN, LSTM, and itransformer can be used. The following uses itransformer as an example.
[0058] First, after wavelet denoising of the sensor temperature input data, the input time-series signal is analyzed in the frequency domain using Fast Fourier Transform (FFT) to construct a physically interpretable signal decomposition architecture. Specifically, amplitude spectrum analysis is used to automatically identify the top k dominant frequency components, which are then classified as high-frequency non-stationary components representing abrupt change modes using a differentiable masking mechanism. The remaining components constitute low-frequency stationary components reflecting long-term trends. This invention innovatively establishes a heterogeneous processing channel: high-frequency components undergo local feature capture via a multilayer perceptron, while low-frequency components are globally dependently modeled using an improved time-series prediction algorithm, forming a complementary feature analysis system. The mathematical principle of FFT is as follows:
[0059]
[0060] Where X(k) represents the transformation result of the complex spectrum at k in the frequency domain, which is the output of the Discrete Fourier Transform (DFT) and represents the amplitude and phase information of the signal at frequency k; x1(r) and x2(r) represent the first and second components of the input signal at time point r; R is the twitch factor, representing the root of unity of the complex number, defined as Rt. N =e -2πi / N .
[0061] This invention addresses the quadratic computation bottleneck of traditional deep learning temporal prediction algorithms (such as Transformer) by employing a locality-sensitive hashing (LSH) mechanism in the encoder layer to achieve attention bucketing computation. Specifically, multiple rounds of hash functions map high-dimensional key vectors to low-dimensional bucket spaces, reducing computational complexity to linear levels while maintaining semantic similarity. Furthermore, the conventional feedforward network is replaced with a Fourier analysis network, utilizing frequency-domain convolutional kernels to enhance the model's ability to extract temporal periodic features. This architecture, in conjunction with the data embedding layer, enables multi-scale representation learning of temporal features from the time domain to the frequency domain. For the input sequence X... :,n Embedded representation Improved time series prediction algorithm block H after L layers l+1 =TrmBlock(H l Iterative processing is performed, and the predicted sequence is finally output through the projection layer. in This is the hidden state matrix.
[0062] To address the characteristics of the decomposed signal components, differentiated processing channels are established: for high-frequency non-stationary components, a residual MLP network with an adaptive activation function is used to capture short-term abrupt change patterns through multi-layer nonlinear transformation; for low-frequency stationary components, a hierarchical attention mechanism based on an improved time-series prediction algorithm is used to model long-range dependencies. After standardization processing of both channels, a dynamic weighted fusion strategy is employed to integrate the prediction results, ultimately outputting a comprehensive prediction value that combines local sensitivity and global consistency.
[0063] III. Intelligent Decision Making Based on Reinforcement Learning Algorithms
[0064] Data prediction models can be used to obtain time-series predictions of temperature, humidity, and energy consumption, which are then fed into the reinforcement learning component for intelligent decision-making. Since temperature regulation is a continuously changing environment, reinforcement learning algorithms such as A3C, VDN, DDPG, and TD3 can be employed to make intelligent decisions regarding air conditioning regulation. The following example uses TD3.
[0065] The state space for reinforcement learning is set as: s = {s1, s2, ... s} 30}, where s1 represents the state of the air conditioner switch, and s2 to s 29 The first 14 dimensions represent the indoor temperature values measured by 14 temperature sensors, and the last 14 dimensions represent the indoor humidity values measured by 14 humidity sensors. The action space is set as a 1*1 vector value, where the action values are 0, 1, 2, and 3, representing turning off the air conditioner, raising the air conditioner temperature by 1 degree Celsius, lowering the air conditioner temperature by 1 degree Celsius, and keeping the temperature unchanged, respectively. The reward function is set as follows:
[0066]
[0067] Where (α,β,γ) are the coefficients of each penalty, empirically initialized as (5.74e-02,2.61e-04,-1.78e1), (T ideal ,T max H ideal H max ,P max ,P min ) represents the custom ideal environmental parameter values in real life, with maximum and minimum environmental parameter values, taken as (25, 30, 50, 60, 157580, 156580) based on experience and field sampling data. (T) ii H ii P) represents the data obtained from the state space, and C represents the data obtained from the state space. switch This is the penalty value for switching air conditioner status. If the air conditioner status is switched, it is defined as 0.01; otherwise, it is 0.
[0068] The TD3 algorithm consists of an Actor-Critic dual network, where the objective of the Actor network is to maximize the Q-value, calculated as follows: Its gradient is
[0069]
[0070] Parameter updates employ a soft update strategy: φ′←τφ+(1-τ)φ′.π φ Let τ be the policy function with network parameters φ, and τ be the soft update parameter with a value of 0.001.
[0071] The Cirtic network uses the Bellman equation to approximate the Q-value y = r + γQ. θ (s′,π φ′ (s′)), where γ is the discount factor with a value of 0.99. The loss function is minimized by the mean squared error function. To update the parameter θ, we still use the soft update strategy: θ′←τθ+(1-τ)θ′.
[0072] In addition, TD3 employs a dual-critic network to calculate the Q-value, preventing overestimation of the Q-value.
[0073]
[0074] At the same time, random noise is introduced into the target action to avoid over-reliance on any single action.
[0075] a′=π φ′ (s′)+∈,∈~clip(N(0,σ)).
[0076] Finally, by training with existing data, the parameter values of the Actor and two Critic networks can be obtained. Given the existing state space, by calling the network parameter model, continuous action values between [-1, 1] are output, and then a mapping function is used...
[0077]
[0078] By mapping the continuous action space to the discrete action space, intelligent decision-making for air conditioning under different actions can be achieved.
[0079] The present invention also provides a computer storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method. Furthermore, the present invention provides an electronic device including a memory and one or more processors, wherein the memory is used to store one or more programs; when the one or more programs are executed by the one or more processors, they implement the above-described method.
[0080] Beneficial effects:
[0081] 1. By building components based on the domestically produced MWORKS for modeling, data prediction, and reinforcement learning, we have broken the monopoly of foreign MATLAB and achieved self-reliance and strength in science and technology for our motherland.
[0082] 2. Sysplorer provides a graphical modeling environment, allowing users to quickly build models by dragging and dropping components, reducing the difficulty of modeling air-conditioned rooms. It has a built-in high-efficiency solver that can quickly process large-scale complex models, improving simulation efficiency. It also shares a working platform with Syslab, making data connection simple and fast.
[0083] 3. In the data prediction section, FFT is used to identify the main frequency components in the frequency domain sequence. These components are then removed from the original time series using filters to obtain a more stable time series. This part is then predicted separately from the remaining part using an MLP model, which further improves the model accuracy.
[0084] 4. In the reinforcement learning part, compared with the traditional DDPG, the TD3 algorithm uses a dual-critic network to calculate the Q value, which prevents overestimation of the results, and introduces random noise into the target action to prevent over-reliance on a certain action, thereby enhancing the accuracy and smoothness of the prediction results. Attached Figure Description
[0085] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings in the following description are merely exemplary, and those skilled in the art can derive other embodiments based on the provided drawings without creative effort.
[0086] Figure 1 Overall flowchart of intelligent operation and maintenance of air conditioning computer room;
[0087] Figure 2 Overall environment simulation model of the computer room;
[0088] Figure 3 Simulation model of the internal environment of the computer room;
[0089] Figure 4 Simulation model of the internal structure of an air conditioner;
[0090] Figure 5 Data prediction model flowchart;
[0091] Figure 6 Flowchart of reinforcement learning algorithm. Detailed Implementation
[0092] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0093] As shown in the figure, this invention provides an intelligent energy-saving control method for air-conditioning rooms based on simulation modeling, deep learning prediction, and reinforcement learning optimization, comprising the following steps:
[0094] Step 1: Building a simulation model of the computer room environment
[0095] Step 1.1: Use Sysplorer to build an air-conditioned room model. After the air-conditioned room is built, set the realexpression parameter value to determine the air-conditioned engine speed and actual power.
[0096] Step 1.2: Use Modelica language to set the ambient temperature T_Amb, ambient humidity phi_Amb, high-pressure side initial pressure P0_high, low-pressure side initial pressure P0_low, compressor discharge initial specific enthalpy h0_high, etc., to set the initial state of the system.
[0097] Step 1.3: Set the proportional coefficient k, maximum opening yMax, minimum opening yMin, etc. of the expansion valve to adjust the operating status of the system.
[0098] Step 1.4: Set the air volume parameter V tot The wall material (solid), thickness (th), surface area, and thermal conductivity, among other parameters, are used in the simulation.
[0099] Step 1.5: After the simulation is complete, the data can be viewed in the power sensor of the EvaporatorR134a component. The energy consumption time series can be obtained by integrating the data over time using the integrator component.
[0100] W=∫pdt.
[0101] Where W represents energy consumption, p represents power, and t represents time;
[0102] Step 1.6: The temperature and humidity time series can be obtained in the pTSensorAir component.
[0103] Step 2: Time Series Prediction Using Deep Learning Step 2.1: Obtain time series data containing 57 features, including temperature, humidity, and air conditioning energy consumption data. First, wavelet denoising is performed on the temperature data collected by the sensors to filter out noise interference. Then, a Fast Fourier Transform (FFT) is performed on the denoised time series signal to convert the time domain signal to the frequency domain, obtaining a complex spectrum. Based on amplitude spectrum analysis, the top k dominant frequency components are automatically identified, and the signal is decomposed into high-frequency non-stationary components and low-frequency stationary components through a differentiable masking mechanism. The high-frequency components represent short-term abrupt change modes, while the low-frequency components reflect long-term trends.
[0104] Step 2.2: Establish a dual-channel processing architecture: Input high-frequency non-stationary components into a multilayer perceptron (MLP) network. The MLP adopts a residual structure with an adaptive activation function and captures local features and short-term mutation patterns through multilayer nonlinear transformation. Input low-frequency stationary components into an improved Transformer model (iTransformer). The model optimizes attention calculation by bucketing through a locality-sensitive hashing mechanism, reducing computational complexity to linear order of magnitude. At the same time, the traditional feedforward network is replaced with a Fourier analysis network, and frequency domain convolution kernels are used to enhance the ability to extract temporal periodic features.
[0105] Step 2.3 Standardization processing; then a dynamic weighted fusion strategy is adopted to adaptively adjust the weights according to the characteristics of the signal components, and integrate the prediction results of the two channels into a comprehensive output sequence. The weights are optimized through learnable parameters to ensure that the prediction results have both local sensitivity and global consistency.
[0106] Step 2.4: Network Training: Using mean squared error (MSE) as the loss function, the parameters of the MLP network and the iTransformer model are jointly optimized through the backpropagation algorithm. After training, the 57-dimensional feature sequence of the target time window is input, and the predicted feature sequence for the future period is output through the above processing flow for subsequent intelligent operation and maintenance system to call.
[0107] Step 3: Building, training, and deploying the reinforcement learning model
[0108] Step 3.1: Algorithm selection and environment setup
[0109] Based on the actual background of the problem, the TD3 algorithm was selected as the reinforcement learning algorithm. A 30-dimensional state space and action space, including temperature, humidity, energy consumption, and air conditioning status, were designed, along with the reward function:
[0110]
[0111] Where (α,β,γ) are the coefficients of each penalty, (T) ideal ,T max H ideal H max ,P max ,P min (T) represents the user-defined ideal environmental parameter values for real-life situations, including maximum and minimum environmental parameter values. i H i P) represents the data obtained from the state space, and C represents the data obtained from the state space. switch This is the penalty value for switching air conditioner status. If the air conditioner status is switched, it is defined as 0.01; otherwise, it is 0.
[0112] Step 3.2: Network parameter training
[0113] Establish a reinforcement learning environment, use random parameters within a specified range to train the network parameters for 1e6 rounds, gradually optimize the network parameters, and obtain the final parameter values of the actor network and two critic networks.
[0114] Step 3.3: Make intelligent decisions
[0115] The ambient temperature, humidity, and air conditioning energy consumption data predicted by the deep learning part are transformed into a 30-dimensional state vector and input into the TD3 environment to obtain the final intelligent decision result, which determines whether the air conditioner is turned on or off or the temperature is adjusted.
[0116] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.
Claims
1. A time-series prediction method for data center energy saving based on deep reinforcement learning, characterized in that, Includes the following steps: Step 1: Build an air-conditioned room model and a computer room environment simulation model using Sysplorer under the MWORKS platform; establish a computer room environment simulation model, solve partial differential equations using historical data of the computer room, calculate the temperature distribution and air conditioning energy consumption in the current space, and build a visualization simulation system to intuitively display the temperature field and energy consumption. Step 2: Temporal prediction using deep learning; By using a deep learning time series prediction model, the temperature and energy consumption of the computer room can be predicted. By capturing the long-term trend and short-term fluctuation of time series data, temperature changes can be predicted in advance, enabling the system to optimize the air conditioning operation strategy based on the prediction data. Step 3: Building, training, and deploying the reinforcement learning model; By employing reinforcement learning algorithms, the start-up, shutdown, and temperature adjustment discrete control signals of air conditioning equipment are controlled to achieve intelligent temperature regulation of the air conditioner, forming an end-to-end energy-saving optimization closed loop from virtual simulation and intelligent decision-making to physical execution; Step 2 specifically includes the following steps: Step 2.1: Obtain time-series data containing multi-dimensional features, including temperature, humidity, and air conditioning energy consumption data; firstly, perform wavelet denoising on the temperature data collected by the sensor to filter out noise interference; then perform a fast Fourier transform on the denoised time-series signal to convert the time-domain signal to the frequency domain and obtain a complex spectrum; based on amplitude spectrum analysis, automatically identify the top k dominant frequency components, and decompose the signal into high-frequency non-stationary components and low-frequency stationary components through a differentiable masking mechanism, where the high-frequency components represent short-term abrupt change modes and the low-frequency components reflect long-term trends; Step 2.2: Establish a dual-channel processing architecture: Input high-frequency non-stationary components into a multilayer perceptron (MLP) network. The MLP uses an adaptive activation function residual structure to capture local features and short-term mutation patterns through multilayer nonlinear transformation. Input low-frequency stationary components into an improved Transformer model, iTransformer. The iTransformer model optimizes attention calculation by binning through a locality-sensitive hashing mechanism, reducing computational complexity to linear order of magnitude. At the same time, the traditional feedforward network is replaced with a Fourier analysis network, and frequency domain convolution kernels are used to enhance the ability to extract temporal periodic features. Step 2.3 Standardization processing; then a dynamic weighted fusion strategy is adopted to adaptively adjust the weights according to the characteristics of the signal components, and integrate the prediction results of the two channels into a comprehensive output sequence. The weights are optimized through learnable parameters to ensure that the prediction results have both local sensitivity and global consistency. Step 2.4: Network Training: Using mean squared error as the loss function, the parameters of the MLP network and the iTransformer model are jointly optimized through backpropagation algorithm. After training, the multidimensional feature sequence of the target time window is input, and the predicted feature sequence for the future time period is output through the above processing flow for subsequent intelligent operation and maintenance system to call.
2. The data center energy-saving time series prediction method based on deep reinforcement learning according to claim 1, characterized in that, Step 1 specifically includes the following steps: Step 1.1: Use Sysplorer to build an air-conditioned room model. After the air-conditioned room is built, set the realexpression parameter value to determine the air-conditioned engine speed and actual power. Step 1.2: Use Modelica language to set the ambient temperature T_Amb, ambient humidity phi_Amb, high-pressure side initial pressure P0_high, low-pressure side initial pressure P0_low, and compressor discharge initial specific enthalpy h0_high to set the initial state of the system; Step 1.3: Set the proportional coefficient k, maximum opening yMax, and minimum opening yMin of the expansion valve to adjust the operating status of the system; Step 1.4: Set air volume parameters Simulations were performed on the wall material solid, thickness th, surface area surfaceArea, and thermal conductivity parameters. Step 1.5: After the simulation is complete, the data can be viewed in the power sensor of the EvaporatorR134a component. The energy consumption time series can be obtained by integrating the data over time using the integrator component. in, Indicates energy consumption. Indicates power, Indicates time; Step 1.6: The temperature and humidity time series can be obtained in the pTSensorAir component.
3. The data center energy-saving time-series prediction method based on deep reinforcement learning according to claim 1, characterized in that, Step 3 specifically includes the following steps: Step 3.1: Algorithm selection and environment setup Based on the actual background of the problem, the TD3 algorithm was selected as the reinforcement learning algorithm. A 30-dimensional state space and action space, including temperature, humidity, energy consumption, and air conditioning status, were designed, along with the reward function: in, Here are the coefficients for each penalty. To customize ideal environmental parameter values, maximum and minimum environmental parameter values for real-life situations, For data obtained from the state space, This is the penalty value for switching air conditioner status. If the air conditioner status is switched, it is defined as 0.01; otherwise, it is 0. Step 3.2: Network parameter training Establish a reinforcement learning environment, use random parameters within a specified range to train the network parameters for 1e6 rounds, gradually optimize the network parameters, and obtain the final parameter values of the actor network and two critic networks. Step 3.3: Make intelligent decisions The ambient temperature, humidity, and air conditioning energy consumption data predicted by the deep learning part are transformed into a 30-dimensional state vector and input into the TD3 environment to obtain the final intelligent decision result, which determines whether the air conditioner is turned on or off or the temperature is adjusted.
4. A computer storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 3.
5. An electronic device, characterized in that, The device includes a memory and one or more processors, the memory being used to store one or more programs; when the one or more programs are executed by the one or more processors, they implement the method as described in any one of claims 1 to 3.
Citation Information
Patent Citations
Self-adaptive multi-band voice mixed emotion perception method
CN118800282A
Remote sensing image domain generalization semantic segmentation method based on multi-scale instance decoupling
CN119992289A