Machine room energy-saving time sequence prediction method based on deep reinforcement learning

By building a simulation-prediction-decision closed-loop control architecture on the MWORKS platform and combining deep learning and reinforcement learning, the problems of high energy consumption and insufficient adaptability in air-conditioning rooms were solved, precise temperature regulation and energy consumption management were achieved, and the energy-saving efficiency and stability of the system were improved.

CN120730686AActive Publication Date: 2025-09-30SOUTHEAST UNIV
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202510697725.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-05-14
Filing Date
2025-05-28
Publication Date
2025-09-30
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Traditional air-conditioning room control methods cannot accurately match dynamic environmental requirements, resulting in excessive energy consumption and a lack of adaptive capabilities. They are unable to cope with IT load fluctuations and changes in the external environment, resulting in energy waste and equipment instability.

Method used

A simulation-prediction-decision closed-loop control architecture based on the MWORKS platform was constructed. Deep learning time series prediction and reinforcement learning algorithms were combined to dynamically optimize air conditioning operating parameters through the computer room environment simulation model, deep learning time series prediction, and reinforcement learning intelligent control, achieving intelligent temperature regulation and energy consumption management.

Benefits of technology

It achieves precise temperature and humidity control in the air-conditioning room, reduces energy consumption, improves the system's adaptability and energy-saving efficiency, and reduces operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120730686A_ABST
    Figure CN120730686A_ABST
Patent Text Reader

Abstract

The invention provides a machine room energy-saving time sequence prediction method based on deep reinforcement learning, and the method mainly comprises the following three parts: building a machine room environment simulation model through Sysplorer under an MWORKS platform, and obtaining the time sequence data of temperature, humidity, energy consumption and the like through simulation modeling; predicting the time sequence data of the environment through a deep learning time sequence prediction method to obtain a final prediction sequence; and finally, inputting a predicted sequence result into reinforcement learning network parameters obtained through training to obtain a final intelligent decision result. The method has the advantages that foreign MATLAB monopoly is broken, the intelligent decision result is smooth, and the accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent air-conditioning energy-saving control technology, and in particular to a computer room energy-saving time series prediction method based on deep reinforcement learning. Background Art

[0002] With the frequent bans of foreign software such as MATLAB in China, the development of alternatives to domestic software has become particularly urgent. Domestic researchers are actively responding to the country's strong support for domestic software and exploring the use of domestic software for scientific research and engineering applications. MWORKS, a domestically developed Matlab alternative, has been used in many fields. Its functionality can already meet the needs of some users, especially in system modeling and simulation, where MWORKS has demonstrated excellent performance. Through the application of domestic software, enterprises can achieve more refined energy consumption management, improve equipment operating efficiency, and reduce energy waste. At the same time, domestic researchers are also actively exploring simulation modeling and intelligent control strategies based on domestic software, in order to achieve or even exceed the performance of foreign software.

[0003] With the development of domestic industrial software, it will play an increasingly important role in the industrial ecosystem of intelligent manufacturing and high-end manufacturing. In the context of intelligent operation and maintenance of air-conditioning rooms, current data center and computer room air conditioning systems consume high energy. Traditional fixed rules and threshold control methods are difficult to adapt to complex and dynamic operating environments, resulting in significant energy waste. While air cooling systems offer a simple structure and easy maintenance, they are energy-inefficient. While water cooling systems offer high energy efficiency, they are complex and require higher levels of operation and maintenance. As data centers expand, fluctuations in room temperature and humidity, IT load variations, and environmental factors impact the efficiency of traditional control methods, making it difficult to accurately match real-time demand. This results in frequent air conditioning system starts and stops, excessive energy consumption, and even impacts stable equipment operation. Furthermore, traditional energy management methods lack intelligent predictive capabilities and rely solely on statically set operating parameters, unable to dynamically adjust based on real-time data. This results in high PUE (power usage effectiveness) and escalating operation and maintenance costs. More advanced energy-saving optimization technologies are urgently needed.

[0004] Existing, outdated energy-saving algorithms primarily rely on simple PID control, empirical threshold settings, and limited logic control strategies, failing to fully account for the complex thermal environment and energy consumption characteristics of computer rooms. These methods typically start and stop air conditioners based on fixed temperature setpoints and fail to accurately optimize based on dynamic environmental factors, resulting in over-operation or under-responsiveness of the air conditioners. Furthermore, while some computer rooms have introduced limited automatic adjustment strategies, such as timed switching and load balancing, these still lack adaptive capabilities and are unable to cope with IT load fluctuations, environmental changes, and the demands of coordinated multi-device operation. This static, rule-based control approach not only has limited energy savings but can also lead to localized overcooling or overheating due to delayed response. Therefore, with the increasing demand for intelligent data centers, the limitations of outdated energy-saving algorithms are becoming increasingly apparent. There is an urgent need to introduce AI-driven intelligent energy-saving optimization technologies to achieve more precise and efficient air conditioner energy management.

[0005] Therefore, this paper focuses on intelligent energy-saving optimization for air-conditioning rooms. This approach involves physical modeling of the room environment, deep learning-based time series prediction, and reinforcement learning-based intelligent control. Through these methods, the research aims to build an intelligent air-conditioning control system that dynamically optimizes air-conditioning operating parameters, ensures temperature and humidity meet standards, and minimizes energy consumption. Summary of the Invention

[0006] This paper proposes an intelligent energy-saving control system for air-conditioning rooms, constructs a "simulation-prediction-decision-making" closed-loop control architecture, and relies on the MWORKS scientific computing and system modeling and simulation platform to complete the reinforcement learning simulation environment construction, algorithm development, model training, testing and tuning, integrated application and other processes of the AI ​​air-conditioning energy-saving algorithm, thereby supporting the development of the domestic scientific computing and system modeling and simulation ecosystem, and realizing the iterative upgrade of the original air-conditioning energy-saving algorithm.

[0007] The research content is divided into three parts. First, a simulation model of the air-conditioning room is established. The partial differential equations are solved by combining the historical data of the room to calculate the temperature distribution and air-conditioning energy consumption in the current space. Based on this, a visual simulation system is built to intuitively display the temperature field and energy consumption, providing a basis for subsequent intelligent control. Secondly, a deep learning time series prediction model is used, including but not limited to time series prediction methods such as iTransformer and FFT, to predict the temperature and energy consumption of the room. By capturing the long-term trends and short-term fluctuations of time series data, temperature changes are predicted in advance, enabling the system to optimize the air-conditioning operation strategy based on the predicted data and improve energy-saving efficiency. Finally, the predicted data is used for intelligent air-conditioning control based on reinforcement learning. Reinforcement learning algorithms, including but not limited to VDN, TD3, etc., are used to control discrete control signals such as the start and stop, temperature adjustment, etc. of the air-conditioning equipment to realize intelligent temperature regulation of the air-conditioning, forming an end-to-end energy-saving optimization closed loop from virtual simulation, intelligent decision-making to physical execution. The specific steps are as follows:

[0008] 1. Establishment of computer room environment simulation model

[0009] This paper uses Sysplorer under the MWORKS platform, a domestic MATLAB alternative, to build an air-conditioned room model. The purpose is to build an air-conditioning model and connect it to the room model. After making decisions through deep learning, the air-conditioning temperature is controlled by changing the power.

[0010] First, the present invention uses the closed-loop air conditioning circuit component in the standard library TAThermalSystem (vehicle thermal management model library) in Sysplorer to cool the air in the circuit. The principle is as follows.

[0011] The CompressorR134a component is used to simulate the process of the compressor increasing the pressure and temperature of the refrigerant through mechanical work. At the same time, the power monitoring sensor is used to read the power change of the air conditioner engine over time. The Integrator component is used to integrate the power over time to finally obtain the air conditioner energy consumption. The performance equation of the compressor is described as follows:

[0012] W=∫pdt

[0013]

[0014] h out =h in +Δh comp ,

[0015] Among them, W represents energy consumption, p represents power, t represents time, p out and p in are the inlet and outlet pressures of the compressor, h in 、h out is the specific enthalpy of the refrigerant flowing into / out of the compressor, is the mass flow rate of refrigerant, η is the compressor efficiency, Δh comp is the specific enthalpy increase during the compression process. out 、p in 、 Δh comp Both can be set manually by inputting into the model. The compressor efficiency η is calculated as follows:

[0016] η=η vol ×η isentropic ×η mech ,

[0017] where η vol (Volumetric efficiency), η isentropic (isentropic efficiency), η mech(Mechanical efficiency) is a parameter that can be set and has different emphases depending on the actual situation.

[0018] The Condenser component is then used to simulate the process in the condenser where the refrigerant releases heat and condenses into liquid through heat exchange with the ambient air. The heat exchange equation of the condenser describes the heat transfer between the refrigerant and the air:

[0019]

[0020] where Q hot is the heat exchange capacity, is the mass flow rate of the refrigerant. The heat exchange efficiency of the condenser is usually described by the logarithmic mean temperature difference:

[0021]

[0022] ΔT1=ΔT ref,in -ΔT air,out ,

[0023] ΔT1=ΔT ref,out -ΔT air,in .

[0024] Where, ΔT ref,in , ΔT ref,out is the inflow and outflow refrigerant temperature, ΔT air,out , ΔT air,in The inflow and outflow air temperatures.

[0025] High-pressure liquid refrigerant passes through the simpleTXV component, and the simulated expansion valve converts the high-pressure liquid refrigerant into low-pressure liquid refrigerant through throttling. The throttling principle is as follows:

[0026]

[0027] Among them C d is the flow coefficient, A is the effective flow area of ​​the expansion valve, ρ is the density of the refrigerant, and ΔP is the pressure difference before and after the expansion valve. The opening y of the expansion valve is controlled by the superheat ΔT sh :

[0028]

[0029] Where k is the proportional coefficient, T i is the integration time constant.

[0030] The low-temperature, low-pressure liquid refrigerant is evaporated into gas after passing through the evaporator group EvaporatorR134a, absorbing the heat of the surrounding environment to achieve a cooling effect. The evaporator also uses the heat transfer equation and calculates the superheat ΔT sh, providing key parameters for controlling the expansion valve opening:

[0031] ΔT sh =ΔT ref,out -T sat ,

[0032] Where T sat is the saturation temperature of the refrigerant.

[0033] After all components are connected, the set parameters can be changed by the components and transmitted between components to realize the circulation of the air conditioning circuit.

[0034] Air conditioned by the air conditioner flows out through the air_out1 interface and into the fan's air inlet a. The fan then sends the air through outlet b to the sensor pTSensorAir. The sensor monitors the air temperature and pressure and sends the air to the air resistance device airResist. The air resistance device simulates flow resistance and sends the air to the room's air inlet airPortIn. After circulating through the room, the air returns to the air conditioner's air inlet air_in1 through the air outlet airPortOut.

[0035] In this process, the present invention controls the speed of the air conditioner through a PID controller and converts it into power:

[0036]

[0037] Where u(t) is the control output (air conditioner speed), e(t) is the error signal (the difference between the set value and the measured value), k p 、k i 、k d are the proportional, integral, and differential coefficients.

[0038] The air conditioner speed is converted from the output of the PID controller to the actual speed through the gain module:

[0039]

[0040] The air flow rate and pressure are determined by the characteristics of the fan and air resistance, as well as the built-in functions of the fan component:

[0041]

[0042] Next, by further building a room model, we can simulate the temperature changes after the air conditioner sends air into the room.

[0043] For the construction of the room model, first use the airVolumes component in the TYBase library to define the air volume, and define the parameters len, hei, and wid as the length, height, and width of the room. The volume parameter V of the component is tot=len×wid×hei.

[0044] The PlaneWall component in the TYThermoFluidSys library is used to define the walls of the room. The parameters of the wall, such as solid (material), th (thickness), and surfaceArea (surface area), are set according to the actual situation. PlaneWall and airVolumes perform heat exchange to simulate heat conduction between the wall and the air.

[0045] The heat exchange between air and wall satisfies the equation:

[0046] Q conv =h×A×(T air -T wall )

[0047] Among them, Q conv is the heat exchange between air and wall, h is the heat transfer coefficient, A is the wall surface area, T air is the air temperature T wall is the wall temperature.

[0048] The HeatCapacitor component in TYBase is used as the heat capacity to connect to the wall, and the C (heat capacity) and T start The initial temperature is set to simulate the heat storage capacity of the wall. The Adiabatic component in TYThermoFluidSys is used as an adiabatic boundary, connecting each of the six walls to prevent excessive heat dissipation from the boundary.

[0049] Finally, we use the heatTransfer component from the TAThermalSystem library and the HeatCapacitor component from TYBase to construct a heat transfer system connected to airVolumes. This represents the heat transfer process between the room and the outside world, with the HeatCapacitor simulating the external heat capacity. The heatTransfer component's Aheatloss (heat transfer area) and k (thermal conductivity) can be set based on actual conditions.

[0050] The airVolumes component, used to simulate the room's air volume, uses the airPortIn and airPortOut air flow interfaces to simulate air entering and exiting the room, respectively. This means that air is exchanged with the outside world through these ports. These two ports connect to the indoor air conditioning model to form a computer room environment simulation model. Once all components are connected, indoor temperature fluctuations can be simulated.

[0051] The heat conduction of the wall satisfies the equation:

[0052]

[0053] Among them, Q cond is the heat flow through the wall, k is the thermal conductivity, A is the wall surface area, d is the wall thickness, T in With T out is the temperature inside and outside the room. It can be used to calculate the heat transfer between the room and the outside world.

[0054] The room model and indoor air conditioning model can be connected through the air flow interface to form a computer room environment simulation model. After simulating indoor temperature changes, time series data such as temperature, humidity, and energy consumption can be obtained for importing into deep learning models.

[0055] In the physical modeling process, 57 features of temperature, humidity detected by different detectors and energy consumption data of different air conditioners are used as model input parameters. After passing through the algorithm, a predicted target feature sequence is output to the reinforcement learning part for intelligent decision-making.

[0056] 2. Time Series Prediction Model Based on Deep Learning

[0057] The present invention adopts a time series prediction model based on deep learning to predict sensor temperature, humidity and air conditioning energy consumption. Various deep learning time series prediction methods such as RNN, LSTM, itransformer, etc. can be used. The following takes itransformer as an example.

[0058] First, after performing wavelet denoising on the sensor temperature of the input data, the input time series signal is analyzed in the frequency domain by fast Fourier transform (FFT) to construct a signal decomposition architecture with physical interpretability. Specifically, the amplitude spectrum analysis is used to automatically identify the first k dominant frequency components, which are delineated as high-frequency non-stationary components that characterize the mutation pattern through a differentiable masking mechanism, and the remaining components constitute low-frequency stationary components that reflect long-term trends. The present invention innovatively establishes a heterogeneous processing channel: the high-frequency component is captured by a multi-layer perceptron for local feature capture, and the low-frequency component is implemented by an improved time series prediction algorithm to implement global dependency modeling, forming a complementary feature analysis system. Among them, the mathematical principle of FFT is as follows:

[0059]

[0060] Where X(k) represents the transformation result at the complex spectrum k in the frequency domain. It is the output of the discrete Fourier transform (DFT) and represents the amplitude and phase information of the signal at frequency k. x1(r) and x2(r) represent the first and second components of the input signal at time point r. R is the rotation factor, which represents the unit root of the complex number and is defined as R. N =e -2πi / N .

[0061] This invention addresses the quadratic computational bottleneck of traditional deep learning time series prediction algorithms (such as Transformer) and uses a local sensitive hashing mechanism in the encoder layer to implement attention bucketing calculations. In specific implementation, high-dimensional key vectors are mapped to low-dimensional bucket spaces through multiple rounds of hash functions, reducing the computational complexity to a linear order while maintaining semantic similarity. Furthermore, the conventional feedforward network is replaced with a Fourier analysis network, and the frequency domain convolution kernel is used to enhance the model's ability to extract time series periodic features. The collaborative design of this architecture and the data embedding layer can realize multi-scale representation learning of time series features from the time domain to the frequency domain. For the input sequence X :,n Embedding representation The improved time series prediction algorithm block H after L layers l+1 =TrmBlock(H l ) iterative processing, and finally output the prediction sequence through the projection layer in is the hidden state matrix.

[0062] Based on the characteristics of the decomposed signal components, differentiated processing channels are established: for high-frequency non-stationary components, a residual MLP network with adaptive activation functions is used to capture short-term mutation patterns through multi-layer nonlinear transformations; for low-frequency stationary components, a hierarchical attention mechanism based on an improved time series prediction algorithm is used to model long-range dependencies. After normalization of both channels, the prediction results are integrated using a dynamic weighted fusion strategy, ultimately outputting a comprehensive prediction value that combines local sensitivity with global consistency.

[0063] 3. Intelligent Decision-Making Based on Reinforcement Learning Algorithms

[0064] The data prediction model generates time series predictions for temperature, humidity, and energy consumption, which are then fed into the reinforcement learning framework for intelligent decision-making. Since temperature regulation is a continuously changing process, the A3C, VDN, DDPG, and TD3 algorithms used in reinforcement learning can be used to make intelligent decisions about air conditioning settings. TD3 is used as an example below.

[0065] The state space of reinforcement learning is set to: s={s1,s2,…s 30}, where s1 represents the state of the air conditioner switch, s2 to s 29 The first 14 dimensions represent the indoor temperature values ​​measured by 14 temperature sensors, and the last 14 dimensions represent the indoor humidity values ​​measured by 14 humidity sensors. The action space is set to a 1*1 vector value, where the action values ​​are 0, 1, 2, and 3, representing turning off the air conditioner, raising the air conditioner by 1 degree Celsius, lowering the air conditioner by 1 degree Celsius, and keeping the temperature unchanged. The reward function is set to:

[0066]

[0067] Among them, (α, β, γ) are the coefficients of each penalty, which are initialized to (5.74e-02, 2.61e-04, -1.78e1) according to experience, (T ideal ,T max ,H ideal ,H max ,P max ,P min ) is the ideal environmental parameter value in real life, the maximum and minimum environmental parameter values ​​are set according to experience and field sampling data (25, 30, 50, 60, 157580, 156580), (T ii ,H ii ,P) is the data obtained from the state space, C switch The penalty value for the air conditioner switching cost is defined as 0.01 if the air conditioner state is switched, and 0 otherwise.

[0068] The TD3 algorithm consists of an Actor-Critic dual network, where the goal of the Actor network is to maximize the Q value, which is calculated as follows: Its gradient is

[0069]

[0070] The parameter update adopts the soft update strategy: φ′←τφ+(1-τ)φ′.π φ is the policy function under the network parameter φ, τ is the soft update parameter, and its value is 0.001

[0071] The Cirtic network uses the Bellman equation to approximate the Q value y = r + γQ θ (s′,π φ′ (s′)), γ is the discount factor, which is set to 0.99. By minimizing the mean square error loss function To update the parameter θ, we still use the soft update strategy θ′←τθ+(1-τ)θ′.

[0072] In addition, TD3 also uses a dual critic network to calculate the Q value to prevent overestimation of the Q value:

[0073]

[0074] At the same time, random noise is introduced into the target action to avoid excessive reliance on a certain action.

[0075] a′=π φ′ (s′)+∈,∈~clip(N(0,σ)).

[0076] Finally, through the training of existing data, the parameter values ​​of the Actor and two Critic networks can be obtained. For the existing state space, by calling the network parameter model, the continuous action value between [-1, 1] is output, and then the continuous action value between [-1, 1] is output through the mapping function.

[0077]

[0078] The continuous action space is mapped to the discrete action space to complete the intelligent decision-making of the air conditioner under different actions.

[0079] The present invention also provides a computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above method. Furthermore, the present invention also provides an electronic device comprising a memory and one or more processors, the memory being configured to store one or more programs; and the one or more programs, when executed by the one or more processors, implementing the above method.

[0080] Beneficial effects:

[0081] 1. Based on the domestic product MWORKS, we built components for modeling, data prediction, and reinforcement learning, breaking the monopoly of foreign MATLAB and achieving self-reliance in China's science and technology.

[0082] 2. Sysplorer provides a graphical modeling environment where users can quickly build models by dragging and dropping components, reducing the difficulty of modeling air-conditioned rooms. It has a built-in efficient solver that can quickly process large-scale complex models, improving simulation efficiency. It also shares a work platform with Syslab, making data connection simple and fast.

[0083] 3. In the data prediction part, FFT is used to identify the main frequency components in the frequency domain sequence, and these components are removed from the original time series using a filter to obtain a more stable time series part. This is then separated from the remaining part and predicted using the MLP model to further improve the model accuracy.

[0084] 4. In the reinforcement learning part, compared with the traditional DDPG algorithm, the TD3 algorithm uses a dual critic network to calculate Q values, preventing overestimation of the results. It also introduces random noise into the target action to prevent over-reliance on a single action, thereby enhancing the accuracy and smoothness of the prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.

[0086] Figure 1 Overall flow chart of intelligent operation and maintenance of air-conditioning rooms;

[0087] Figure 2 Simulation model of the overall environment of the computer room;

[0088] Figure 3 Simulation model of the internal environment of the computer room;

[0089] Figure 4 Simulation model of the internal structure of the air conditioner;

[0090] Figure 5 Data prediction model flow chart;

[0091] Figure 6 Flowchart of the reinforcement learning algorithm. DETAILED DESCRIPTION

[0092] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0093] As shown in the figure, the present invention provides an intelligent energy-saving control method for air-conditioning rooms based on simulation modeling, deep learning prediction, and reinforcement learning optimization, including the following steps:

[0094] Step 1: Building a computer room environment simulation model

[0095] Step 1.1: Use Sysplorer to build an air-conditioned room model. After the air-conditioned room is built, set the realexpression parameter value to determine the air-conditioning engine speed and power.

[0096] Step 1.2: Use Modelica to set the ambient temperature T_Amb, ambient humidity phi_Amb, high-pressure side initial pressure P0_high, low-pressure side initial pressure P0_low, compressor exhaust initial specific enthalpy h0_high, etc. to set the initial state of the system.

[0097] Step 1.3: Set the expansion valve's proportional coefficient k, maximum opening yMax, minimum opening yMin, etc. to adjust the system's operating status.

[0098] Step 1.4: Set the air volume parameter V tot , wall material solid, thickness th, surface area surfaceArea and thermal conductivity and other parameters. Simulation is carried out.

[0099] Step 1.5: After the simulation is complete, you can view the data in the power sensor of the EvaporatorR134a component. Use the integrator component to integrate the time to obtain the energy consumption time series:

[0100] W=∫pdt.

[0101] Where W represents energy consumption, p represents power, and t represents time;

[0102] Step 1.6: The temperature and humidity time series are available in the pTSensorAir component.

[0103] Step 2: Time Series Prediction Using Deep Learning Step 2.1: Obtain time series data containing 57 features, including temperature, humidity, and air conditioning energy consumption data. First, wavelet denoising is performed on the temperature data collected by the sensor to filter out noise interference. Then, a fast Fourier transform (FFT) is performed on the denoised time series signal to convert the time domain signal to the frequency domain to obtain a complex spectrum. Based on amplitude spectrum analysis, the top k dominant frequency components are automatically identified, and a differentiable masking mechanism is used to decompose the signal into high-frequency non-stationary components and low-frequency stationary components. The high-frequency components represent short-term mutation patterns, while the low-frequency components reflect long-term trends.

[0104] Step 2.2: Establish a dual-channel processing architecture: Input the high-frequency non-stationary component into the multi-layer perceptron (MLP) network. The MLP adopts a residual structure with an adaptive activation function and captures local features and short-term mutation patterns through multi-layer nonlinear transformations. Input the low-frequency stationary component into the improved Transformer model (iTransformer). The model optimizes the attention calculation by bucketing through the local sensitive hashing mechanism, reducing the computational complexity to a linear level. At the same time, the traditional feedforward network is replaced by a Fourier analysis network, and the frequency domain convolution kernel is used to enhance the ability to extract time series periodic features.

[0105] Step 2.3 is normalized; a dynamic weighted fusion strategy is then used to adaptively adjust the weights according to the characteristics of the signal components, and the prediction results of the two channels are integrated into a comprehensive output sequence. The weights are optimized through learnable parameters to ensure that the prediction results have both local sensitivity and global consistency.

[0106] Step 2.4: Network training: Using the mean squared error (MSE) as the loss function, the parameters of the MLP network and iTransformer model are jointly optimized through the backpropagation algorithm. After training, the 57-dimensional feature sequence of the target time window is input. The above processing flow outputs the predicted feature sequence for a period of time in the future for subsequent intelligent operation and maintenance system calls.

[0107] Step 3: Establish, train and call the reinforcement learning model

[0108] Step 3.1: Algorithm selection and environment setup

[0109] According to the actual background of the problem, the TD3 algorithm is selected as the reinforcement learning algorithm. A 30-dimensional state space including temperature, humidity, energy consumption, and air conditioning status is designed, the action space, and the reward function are:

[0110]

[0111] Among them, (α, β, γ) are the coefficients of each penalty, (T ideal ,T max ,H ideal ,H max ,P max ,P min ) is the ideal environmental parameter value, maximum and minimum environmental parameter value in real life, (T i ,H i ,P) is the data obtained from the state space, C switch is the penalty value of the air conditioner switching cost. If the air conditioner state is switched, it is defined as 0.01, otherwise it is 0;

[0112] Step 3.2: Network parameter training

[0113] A reinforcement learning environment is established, and 1e6 rounds of network parameter training are performed using random parameters within the specified range. The network parameters are gradually optimized to obtain the final parameter values ​​of the actor network and the two critic networks.

[0114] Step 3.3: Make intelligent decisions

[0115] The ambient temperature, humidity, and air conditioning energy consumption data predicted by the deep learning part are converted into a 30-dimensional state vector and input into the TD3 environment to obtain the final intelligent decision-making result, which determines whether to turn the air conditioner on or off or the temperature adjustment direction.

[0116] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely illustrative of the principles of the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and modifications are intended to fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for predicting energy-saving timing in a computer room based on deep reinforcement learning, characterized in that: The following steps are involved: Step 1: Build a computer room environment simulation model. This model uses the computer room's historical data to solve partial differential equations, calculate the temperature distribution and air conditioning energy consumption in the current space, and construct a visual simulation system to intuitively display the temperature field and energy consumption. Step 2: Time series prediction based on deep learning; Using a deep learning time series prediction model, we predict the temperature and energy consumption of the computer room. By capturing long-term trends and short-term fluctuations in time series data, we can predict temperature changes in advance, enabling the system to optimize air conditioning operation strategies based on the predicted data. Step 3: Establish, train and call the reinforcement learning model; A reinforcement learning algorithm is used to control the start and stop and temperature adjustment discrete control signals of air-conditioning equipment, realize intelligent temperature adjustment of the air conditioner, and form an end-to-end energy-saving optimization closed loop from virtual simulation, intelligent decision-making to physical execution.

2. The method for predicting energy-saving timing of a computer room based on deep reinforcement learning according to claim 1 is characterized in that: Step 1 specifically includes the following steps: Step 1.1: Use Sysplorer to build an air-conditioned room model. After the air-conditioned room is built, set the realexpression parameter values ​​to determine the air-conditioning engine speed and power. Step 1.2: Use Modelica to set the ambient temperature T_Amb, ambient humidity phi_Amb, high-pressure side initial pressure P0_high, low-pressure side initial pressure P0_low, and compressor exhaust initial specific enthalpy h0_high to set the initial state of the system. Step 1.3: Set the proportional coefficient k, maximum opening yMax, and minimum opening yMin of the expansion valve to adjust the operating state of the system; Step 1.4: Set the air volume parameter V tot , wall material solid, thickness th, surface area surfaceArea and thermal conductivity parameters are used for simulation; Step 1.5: After the simulation is complete, the data can be viewed in the power sensor of the EvaporatorR134a component. The energy consumption time series can be obtained by integrating the time using the integrator component: W=∫pdt. Where W represents energy consumption, p represents power, and t represents time; Step 1.6: The temperature and humidity time series can be obtained in the pTSensorAir component.

3. The method for predicting computer room energy saving time series based on deep reinforcement learning according to claim 1 is characterized in that: Step 2 specifically includes the following steps: Step 2.1: Obtain time series data containing multidimensional features, including temperature, humidity, and air conditioning energy consumption data. First, perform wavelet denoising on the temperature data collected by the sensor to filter out noise interference. Then, perform a fast Fourier transform on the denoised time series signal to convert the time domain signal to the frequency domain to obtain a complex spectrum. Based on amplitude spectrum analysis, automatically identify the top k dominant frequency components and decompose the signal into high-frequency non-stationary components and low-frequency stationary components using a differentiable masking mechanism. The high-frequency components represent short-term mutation patterns, while the low-frequency components reflect long-term trends. Step 2.2: Establish a dual-channel processing architecture: Input the high-frequency non-stationary components into the Multi-Layer Perceptron (MLP) network. The MLP uses a residual structure with an adaptive activation function to capture local features and short-term mutation patterns through multi-layer nonlinear transformations. Input the low-frequency stationary components into the improved Transformer model iTransformer. The iTransformer model uses a locality-sensitive hashing mechanism to optimize the attention calculation by bucketing, reducing the computational complexity to a linear level. At the same time, the traditional feedforward network is replaced with a Fourier analysis network, and the frequency domain convolution kernel is used to enhance the ability to extract time series periodic features. Step 2.3: Normalization. A dynamic weighted fusion strategy is then used to adaptively adjust the weights based on the characteristics of the signal components, integrating the prediction results of the two channels into a comprehensive output sequence. The weights are optimized using learnable parameters to ensure that the prediction results have both local sensitivity and global consistency. Step 2.4: Network training: Using mean squared error as the loss function, the parameters of the MLP network and iTransformer model are jointly optimized through the backpropagation algorithm. After training, the multidimensional feature sequence of the target time window is input, and the predicted feature sequence for a period of time in the future is output through the above processing flow for subsequent intelligent operation and maintenance system calls.

4. The method for predicting computer room energy saving time series based on deep reinforcement learning according to claim 1 is characterized in that: Step 3 specifically includes the following steps: Step 3.1: Algorithm selection and environment setup According to the actual background of the problem, the TD3 algorithm is selected as the reinforcement learning algorithm. A 30-dimensional state space including temperature, humidity, energy consumption, and air conditioning status is designed, the action space, and the reward function are: Among them, (α, β, γ) are the coefficients of each penalty, (T ideal ,T max ,H ideal ,H max ,P max ,P min ) is the ideal environmental parameter value, maximum and minimum environmental parameter value in real life, (T i ,H i ,P) is the data obtained from the state space, C switch is the penalty value of the air conditioner switching cost. If the air conditioner state is switched, it is defined as 0.01, otherwise it is 0; Step 3.2: Network parameter training Establish a reinforcement learning environment, use random parameters within the specified range to perform 1e6 rounds of network parameter training, gradually optimize the network parameters, and obtain the final parameter values ​​of the actor network and the two critic networks; Step 3.3: Make intelligent decisions The ambient temperature, humidity, and air conditioning energy consumption data predicted by the deep learning part are converted into a 30-dimensional state vector and input into the TD3 environment to obtain the final intelligent decision result, which determines the air conditioner's switch or temperature adjustment direction.

5. A computer storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

6. An electronic device, characterized in that: The method comprises a memory and one or more processors, wherein the memory is used to store one or more programs; when the one or more programs are executed by the one or more processors, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Operation management method and system for efficient and energy-saving air conditioner room

    CN118517769A

  • Power grid topology optimization method and system based on search sorting

    CN118539441A

  • Self-adaptive multi-band voice mixed emotion perception method

    CN118800282A

  • Building air conditioning system optimization control method and device based on spatio-temporal data

    CN119103659A

  • Data-driven group intelligent management and control method based on reinforcement learning algorithm

    CN119358595A