Data center water chilling unit energy consumption optimization method and system based on LSAC

The operation strategy of the water-cooling unit is optimized through the LSAC algorithm, which solves the problem of energy consumption optimization instability caused by the lack of sensor data and dynamic changes in the environment, and realizes the reduction of energy consumption and system efficiency of the data center water-cooling unit.

CN120430189APending Publication Date: 2025-08-05NANJING UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510604049.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the energy consumption optimization of data center water cooling units, the lack of sensor data and dynamic changes in the environment lead to poor energy consumption optimization results, and the convergence results of the optimization algorithm are unstable.

Method used

Using an LSAC-based method, the energy consumption model of the water-cooling unit is constructed through a data-driven method, a reinforcement learning environment is designed, and the historical time series data is processed using the LSTM network, and the operation strategy of the water-cooling unit is optimized in combination with the SAC algorithm to minimize energy consumption.

Benefits of technology

In the absence of sensor data, it can effectively reduce the energy consumption of the water-cooling unit, improve system stability and optimization effect, and achieve more efficient energy consumption management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430189A_ABST
    Figure CN120430189A_ABST
Patent Text Reader

Abstract

The invention discloses a data center water cooling unit energy consumption optimization method and system based on LSAC, and aims to effectively reduce the operation energy consumption of water cooling equipment through an intelligent algorithm and improve the overall efficiency of the system. An energy consumption optimization framework is designed for a water cooling data center, a data driving method is used for building a water cooling unit energy consumption model to simulate a real environment, the potential problems of untimely feedback and failure of a sensor in a complex dynamic environment are fully considered in the optimization process, an LSAC algorithm is designed to optimize an operation strategy of the water cooling unit, and the energy consumption optimization framework is used for optimizing the energy consumption of the water cooling unit. According to the algorithm, the perception capability of the environment state is enhanced through the LSTM network, and training and deployment of optimization decisions are carried out by using the SAC algorithm. The energy consumption of the water-cooling unit can be effectively reduced under the condition that the data of the data center sensor is missing, the total energy consumption of the water-cooling data center is effectively reduced, and the method has the advantages of being remarkable in optimization effect and high in stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data center energy consumption optimization technology, and in particular to a water cooling unit energy consumption optimization method and system based on an intelligent algorithm. Background Art

[0002] With the rapid development of information technology, data center energy consumption has become increasingly prominent. Water-cooled chillers, as the primary cooling equipment, consume a significant portion of a data center's overall energy consumption. Traditional energy optimization methods often rely on empirical experience and simple control strategies, which are difficult to adapt to complex and dynamic operating environments and result in ineffective energy reduction. Furthermore, sensors in data center environments may experience untimely feedback or malfunction, further impacting energy optimization efforts.

[0003] In recent years, the rapid development of deep learning and reinforcement learning technologies has provided new insights into this problem. By building data-driven energy consumption models and combining them with the decision-making capabilities of reinforcement learning, it is possible to effectively optimize the operating strategies of water-cooling units. However, existing technologies still lack a systematic solution for optimizing the energy consumption of water-cooling units in data centers. Especially in the absence of sensor data, there is an urgent need for intelligent algorithms that can improve energy efficiency and reduce energy consumption without relying on comprehensive observations. Summary of the Invention

[0004] Purpose of the Invention: This invention aims to provide a method and system for optimizing the energy consumption of water-cooled chillers in data centers based on LSAC. This method, through intelligent algorithms, effectively reduces the operating energy consumption of data center water-cooling equipment and improves the overall efficiency of the system. Specifically, this invention aims to address the existing issues of poor energy optimization and unstable convergence of optimization algorithms in the absence of sensor data and in dynamic environmental conditions, thereby achieving stable and efficient energy management.

[0005] Technical solution: To achieve the above-mentioned purpose, the present invention adopts the following technical solution:

[0006] In a first aspect, the present invention provides a method for optimizing energy consumption of a water-cooling unit in a data center based on LSAC, comprising the following steps:

[0007] Collect, clean, and normalize real-time data from sensors in the chiller and its environment;

[0008] Based on the collected data, an energy consumption model of the water-cooling unit is constructed to simulate the energy consumption performance of the equipment under different working conditions;

[0009] Based on the energy consumption model of a water-cooling unit, a reinforcement learning environment is designed. A Markov decision process for optimizing the energy consumption of the water-cooling unit is constructed, as well as a partially observable Markov decision process that partially masks the state space in the Markov decision process. The LSAC algorithm is used to optimize the operation strategy of the water-cooling unit to minimize energy consumption. The LSAC algorithm uses an LSTM network to process historical time series data and embeds the output of the LSTM network into the actor and critic networks of the SAC algorithm.

[0010] Furthermore, a data-driven approach is used in combination with the multi-layer perceptron (MLP) in a deep neural network to construct the energy consumption model of the water-cooling unit. The input of the energy consumption model of the water-cooling unit includes environmental parameters, equipment parameters, and load parameters, and the output is the normalized value of the energy consumption of the water-cooling unit in the current state.

[0011] Preferably, the collected data includes at least the chilled water outlet temperature, chilled water outlet pressure, cooling water outlet temperature, cooling water outlet pressure, evaporator outlet temperature, cold storage percentage, phase voltage, phase current, water pan temperature, cold storage tank temperature, suction temperature and exhaust pressure.

[0012] Preferably, the state space of the Markov decision process, including environmental parameters, equipment parameters and load parameters, is consistent with the input parameters of the constructed water-cooling unit energy consumption model; the action space is the key control parameters selected from the state space; the partially observable Markov decision process masks part of the information in the state space, so that the intelligent agent can only make decisions based on part of the observations.

[0013] Preferably, in the reinforcement learning environment, the reward of the agent after each action is the opposite of the output value of the water-cooling unit energy consumption model in the current state.

[0014] Furthermore, the LSAC algorithm combines the LSTM network structure and the SAC reinforcement learning algorithm to design data preprocessing network actors for the Actor network and the Critic network. summarizer and Q summarizer , the state sequence data of a period of time is first passed through the actor summarizer and Q summarizer The feature representation of the state sequence during this period is generated, and then the feature representation is input into the Actor network and the Critic network to generate the final strategy and action value; a cyclic experience replay mechanism is used to access data during training, and fragments are randomly intercepted from each complete training trajectory to increase the total number and diversity of samples.

[0015] Furthermore, the LSAC algorithm uses three LSTM-based preprocessing networks, namely Actor, Q1, and Q2 preprocessors, and three MLP networks, namely Actor network, action value network Q1, and Q2; the Actor preprocessor uses its output state representation as the input of the policy network, the policy network outputs the mean and standard deviation, and obtains the selected output action and the logarithmic probability of the action based on Gaussian distribution sampling; the action value network Q1 / Q2 concatenates the output of the Q1 / Q2 preprocessor and the action output by the policy network as input, and outputs the corresponding state action value.

[0016] In a second aspect, the present invention provides a data center water cooling unit energy consumption optimization system based on LSAC, comprising:

[0017] The acquisition module is used to collect real-time data from sensors in the water-cooling unit and its environment and to clean and standardize it;

[0018] The water-cooling unit energy consumption modeling module builds a water-cooling unit energy consumption model based on the collected data and simulates the energy consumption performance of the equipment under different working conditions;

[0019] The reinforcement learning environment modeling and energy consumption optimization module is used to design a reinforcement learning environment based on the water-cooling unit energy consumption model, construct a Markov decision process for optimizing the water-cooling unit's energy consumption, and a partially observable Markov decision process that partially masks the state space in the Markov decision process. The LSAC algorithm is used to optimize the water-cooling unit's operating strategy to minimize energy consumption. The LSAC algorithm uses an LSTM network to process historical time series data and embeds the output of the LSTM network into the actor and critic networks in the SAC algorithm.

[0020] In a third aspect, the present invention provides a computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the LSAC-based data center water cooling unit energy consumption optimization method are implemented.

[0021] In a fourth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the steps of the LSAC-based data center water cooling unit energy consumption optimization method.

[0022] Beneficial effects: Compared with the existing technology, the present invention has the following beneficial effects: 1. The present invention makes full use of the water-cooling units in the data center and the sensor data in the environment, and adopts a data-driven approach to model and optimize the energy consumption of the water-cooling units, so that the resulting solution is more generalizable and transferable; 2. The present invention fully considers the uncertain factors in the actual situation. After modeling the optimization process as a Markov decision process, it supplements the corresponding partial observable Markov decision process in the case of missing sensor data, and designs the LSAC optimization algorithm in a targeted manner. This algorithm can effectively reduce the energy consumption of the water-cooling units, and has the advantages of stronger stability and better convergence effect than the original SAC algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a framework diagram of energy consumption optimization of a data center water cooling system according to an embodiment of the present invention.

[0024] Figure 2 Schematic diagram of the network structure of the LSAC algorithm in an embodiment of the present invention.

[0025] Figure 3 This is a comparison chart of the Return convergence process of the SAC algorithm before and after masking variables and the LSAC algorithm after masking variables in the training phase.

[0026] Figure 4 This is a comparison chart of the Return convergence process of the SAC algorithm before and after masking variables and the LSAC algorithm after masking variables in the test phase. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions, and advantages of the present invention more clearly understood, the following detailed description of an embodiment of the present invention is given in conjunction with the accompanying drawings. This embodiment is implemented based on the technical solutions of the present invention, and provides a detailed implementation method and specific operation process. It should be understood that the specific examples described herein are only used to illustrate the present invention, and the scope of protection of the present invention is not limited to the following embodiments.

[0028] This embodiment discloses a new LSAC-based method for optimizing the energy consumption of water-cooling units in data centers, with the goal of effectively reducing the energy consumption of water-cooling units in data centers, improving the energy utilization efficiency of data centers, and thereby effectively reducing the total energy consumption of data centers. The method primarily includes: collecting real-time data from sensors in the water-cooling units and their environment, and performing cleaning and standardization; constructing a water-cooling unit energy consumption model based on the collected data to simulate the energy consumption performance of the equipment under different operating conditions; designing a reinforcement learning environment based on the water-cooling unit energy consumption model, constructing a Markov decision process for optimizing the energy consumption of the water-cooling unit, and a partially observable Markov decision process that partially masks the state space in the Markov decision process, and utilizing the LSAC algorithm to optimize the operating strategy of the water-cooling unit to minimize energy consumption.

[0029] Specifically, we first designed an energy consumption optimization framework for water-cooled data centers. This framework can be divided into four main parts: data acquisition and processing, energy consumption model construction, reinforcement learning environment modeling, and optimization algorithm implementation and deployment. First, we collected real-time data from sensors in the water-cooled chiller and its environment. This data includes, but is not limited to, chilled water outlet temperature, chilled water outlet pressure, cooling water outlet temperature, cooling water outlet pressure, evaporator outlet temperature, cold storage percentage, phase voltage, phase current, water pan temperature, cold storage tank temperature, suction temperature, and discharge pressure. The collected data was cleaned and normalized, addressing missing values and outliers to ensure data quality and reliability. Next, based on this processed data, we built and trained an energy consumption model for the water-cooled chiller. The model's input parameters are the real-time collected data, and its output is the normalized value of the chiller's current energy consumption. Subsequently, we designed a reinforcement learning environment based on the constructed energy consumption model, forming a Markov decision process (MDP) and a partially observable Markov decision process (POMDP) for optimizing the energy consumption of the water-cooled chiller. Finally, we train the designed LSAC algorithm based on the agent's feedback in the reinforcement learning environment. The algorithm then outputs an optimization strategy for deployment in the real-world data center operations and maintenance system. Throughout this cycle, real-time sensor data is continuously updated, and the chiller energy consumption model and network parameters in the LSAC optimization algorithm are updated accordingly to adapt to changes in the data center environment, ensuring model accuracy and the effectiveness of the optimization algorithm. Through this design, we hope to achieve more intelligent and efficient energy management.

[0030] Water-cooled chillers are widely used in data centers and large buildings to provide cooling to maintain stable equipment operation and a comfortable environment. Their core operating principle is to absorb indoor heat through chilled water in the evaporator and then discharge it to the outside through the cooling water system. Compared to air-cooled chillers, water-cooled chillers generally have higher energy efficiency, especially when handling large heat loads. A water-cooled chiller typically consists of a compressor, evaporator, condenser, expansion valve, chilled water pump, and cooling tower. Its operating status is affected by multiple parameters, such as the temperature, pressure, and flow rate of the chilled and cooling water.

[0031] To optimize energy consumption, the operation of water-cooling units requires precise control to ensure consistent and efficient operation under varying loads. Water-cooling units typically account for a significant portion of a data center's overall energy consumption, so optimizing their energy consumption is key to improving overall data center energy efficiency.

[0032] When building an energy consumption model for a water-cooled unit, a comprehensive analysis of its operating status is required. Since the operating status of a water-cooled unit involves multiple physical processes, this makes modeling complex. This embodiment of the present invention takes these factors into consideration and, in order to simplify the analysis, constructs an energy consumption model using a data-driven approach based on collected real-time data. The energy consumption of a water-cooled unit is primarily affected by a series of external and internal factors. The input data we selected includes:

[0033] Environmental parameters: such as outdoor dry bulb temperature, outdoor wet bulb temperature, etc., which have a direct impact on the efficiency of the cooling tower.

[0034] Equipment parameters: including the inlet and outlet temperatures, pressures, and flow rates of chilled water and cooling water. These parameters directly determine the energy efficiency performance of the unit.

[0035] Load parameters: The real-time load conditions of the data center will also affect the energy consumption of the water cooling unit.

[0036] Specifically, we screened a total of 246 important parameters. Through the input of this 246-dimensional data, the model can more accurately capture the energy consumption changes of the water-cooling unit under different operating conditions. Since there may be noise, missing values or outliers in the data, data cleaning and standardization are required before building the model. The goal of data preprocessing is to improve data quality and reduce model errors. The embodiment of the present invention uses interpolation to fill in missing values and identifies and processes outliers through statistical analysis methods. Finally, in order to avoid deviations in model training caused by data of different dimensions, the data is normalized.

[0037] Based on the cleaned and processed data, we used a neural network to construct an energy consumption model for the water-cooling unit. Specifically, we used a multilayer perceptron network with four hidden layers, each with a size of 600, 1000, 500, and 50, respectively. Each perceptron layer consisted of a Linear layer, a BatchNorm1d layer, and an activation function layer, using the ReLU function as the activation function. The network input was a 246-dimensional vector consisting of 246 important parameters, including the chilled water outlet temperature, evaporator outlet temperature, cold storage tank temperature, and exhaust pressure. The output was a 1-dimensional vector, representing the normalized energy consumption of the water-cooling unit.

[0038] In this embodiment of the present invention, a reinforcement learning environment is designed based on the energy consumption model of a water-cooling unit, and a Markov decision process and a partially observable Markov decision process are constructed for optimizing the energy consumption of the water-cooling unit. When constructing a reinforcement learning environment based on the Markov decision process, we must first clarify the system's state, action, reward, and transfer function to describe the energy consumption optimization problem of the water-cooling unit.

[0039] State space: This represents the operating state of the chiller at each moment. In this method, the state is composed of 246 key parameters, including environmental parameters, equipment parameters, and load parameters, which are consistent with the input parameters of the constructed chiller energy consumption model. Therefore, the dimension of the state space is 246. If the chiller energy consumption model is considered the true model, by inputting these parameters, the reinforcement learning agent can obtain comprehensive operating information of the current chiller and observed information about the environment.

[0040] Action Space: To optimize the energy consumption of the water-cooled chiller, the agent needs to adjust the key control variables of the chiller and other equipment at each moment. In this example, we selected 20 parameters that have a significant impact on the energy efficiency of the water-cooled chiller as the control variables in the action space. These 20 parameters were selected from the 246-dimensional state, including: chilled water outlet temperature_x, chilled water outlet pressure_x, chilled water outlet temperature_y, chilled water outlet pressure_y, cooling water outlet temperature_x, cooling water outlet pressure_x, cooling water outlet temperature_y, cooling water outlet pressure_y, evaporator outlet temperature_x, evaporator outlet temperature_y, cold storage percentage_y, phase A voltage_x, water pan temperature_x, water pan temperature_y, chilled water supply pressure_y, cold storage tank temperature_x, phase C voltage_x, phase B current_y, suction temperature, and discharge pressure (x and y are used to distinguish different devices of the same model). The dimension of the action space is 20. By adjusting these parameters, the agent can change the working state of the water cooling unit and thus optimize energy consumption.

[0041] Reward function: This function is used to evaluate the impact of an agent's action on the system at the current moment. In this example, the goal of the reward function is to minimize the total energy consumption of the water-cooling unit. Therefore, we set:

[0042] r t =-E t

[0043] Among them, E t is the energy consumption of the water-cooling unit at time t. Through this reward mechanism, the agent will be guided to take actions that can reduce the energy consumption of the unit.

[0044] The transfer function describes how the chiller's state changes over time. At each moment, the chiller's operating state changes based on its current state and action. Because the chiller's state changes depend not only on the control parameters but also on the external environment and data center load fluctuations, state transitions are complex. To simplify the problem, we assume that while the agent is controlling the device, changes in state variables outside the action space are minimal. These changes are ignored, meaning they remain constant during each round of reinforcement learning training.

[0045] Considering that in actual water-cooling plant optimization problems, the agent cannot directly observe the entire system state due to sensor limitations or data acquisition delays, we supplemented the modeling with a partially observable Markov decision process, masking some information in the state space so that the agent can only make decisions based on a subset of observations.

[0046] The key idea behind the partially observable Markov decision process is that the true state of the system cannot be fully observed, and the agent can only infer the most likely state based on limited observations. Specifically, we mask 16 key parameters from the 246-dimensional state: chilled water outlet temperature_x, chilled water outlet pressure_x, chilled water outlet temperature_y, chilled water outlet pressure_y, cooling water outlet temperature_x, cooling water outlet pressure_x, cooling water outlet temperature_y, cooling water outlet pressure_y, evaporator outlet temperature_x, evaporator outlet temperature_y, cold storage percentage_y, phase A voltage_x, water pan temperature_x, water pan temperature_y, chilled water supply pressure_y, and cold storage tank temperature_x. This simulates scenarios where sensors fail to provide comprehensive information. This allows the agent to infer the complete state based on only partial observations and make control decisions based on this incomplete state information.

[0047] After fully modeling the optimization problem, we designed a loop replay buffer for reinforcement learning that is suitable for processing time series data. This buffer is used in reinforcement learning to store the agent's experience so that it can be randomly sampled from these stored experiences during training, thereby helping to optimize the learning of the policy. We designed a storage structure for storing multi-dimensional observation data, which uses multiple reserved arrays to save the states, actions, rewards, completion flags, and masks collected from the environment. The capacity of each array is defined as capacity, and the specific dimension is set based on the maximum time series length max_episode_len of the interaction between the reinforcement learning agent and the environment. During the storage process, each piece of data stored in the buffer is data of the entire sequence length:

[0048] State storage (self.o): used to store the current state observation data, with dimensions (capacity, max_episode_len+1, o_dim). Among them, o_dim represents the dimension of the state.

[0049] Action storage (self.a): used to store the action of each time step, the dimension is (capacity, max_episode_len, a_dim), where a_dim is the dimension of the action.

[0050] Reward storage (self.r): used to record the reward value of each time step, with a dimension of (capacity, max_episode_len, 1).

[0051] Completion flag storage (self.d): used to record whether each time step is completed, the dimension is (capacity, max_episode_len, 1).

[0052] Mask (self.m): used to mark valid time steps, with a dimension of (capacity, max_episode_len, 1).

[0053] Sequence length storage (self.ep_len): stores the length of each sequence, with a dimension of (capacity).

[0054] The playback buffer in the present invention supports a variety of sampling mechanisms and can extract data from the stored experience as needed for training recurrent neural networks (such as LSTM). The specific implementation method is: when the segment length segment_len is None, the system will directly extract the entire sequence; when segment_len is set to a positive value (the value must be able to divide the sequence length), a segment of the same length will be extracted from the sequence. In particular, in this embodiment, the sequence length is the maximum number of actions performed by the agent in each round of training in the environment, which is set to 40. In order to improve the diversity of sampling and maintain the efficiency of training, we set segment_len to 4, and each sequence will be divided into 10 complete segments of length 4 that can be extracted. In each sampling, the sequence to be sampled is determined according to the Batch Size (value is 32), and then a segment of length 4 is randomly extracted from each sequence.

[0055] Next, an embodiment of the present invention designs an LSAC-based water-cooling unit energy consumption optimization algorithm, which combines the LSTM network structure and the SAC reinforcement learning algorithm.

[0056] The SAC algorithm is a maximum entropy reinforcement learning algorithm based on the Actor-Critic framework. During the training process, in addition to learning a strategy to maximize the expected value of the accumulated reward, it also requires that the entropy of each action output by the strategy be maximized, so as not to miss any useful action.

[0057] In this algorithm, two action-value functions need to be fitted and A policy function π θ Based on the idea of DoubleDQN, in order to alleviate the problem of overestimating the Q value, in the Q network part, a network with a small value is selected when using it:

[0058]

[0059] L Q (ω) is a loss function that measures the action value function Q ω (s t ,a t ) and the target value. E represents the expectation, (s t ,a t ,s t+1 )~R means state s t , action a t and the next state s t+1 is sampled from distribution R. Q ω (s t ,a t ) is in state s t Execute action a t When , the action function with parameter ω estimates the future cumulative reward under this state-action pair. t is in state s t Execute action a t γ is the discount factor, ranging from [0,1], which is used to measure the importance of future rewards. The closer γ is to 1, the more important the future rewards are. ω- (s t+1 ) is the next state s t+1 The value function, which estimates the value of t+1 The future cumulative reward starts from . Further, the sampling distribution in the original expectation changes from (s t ,a t ,s t+1 )~R becomes (s t ,a t+1 ,s t+1 )~R,a t+1 ~π θ (·∣s t+1 ), and V ω- (s t+1 ) is replaced by From the perspective of the change in sampling distribution, we introduce the state s t+1 The strategy π(a t+1 ∣s t+1 ) to sample action a t+1 For V ω- (s t+1 ), we estimate the Q function by taking the minimum value of different parameters And subtract the policy entropy αlogπ(a t+1 ∣s t+1 ), where α is the policy entropy coefficient.

[0060] In the Loss function of the policy network, we also added the action entropy part:

[0061]

[0062] Among them, E represents expectation, s t ~R,a t ~π θ Indicates state s t Sampling from distribution R, a t Based on the current strategy π θ Get π θ (a t ∣s t ) is the action distribution output by the policy network (Gaussian distribution in SAC), θ is the parameter of the policy network, log(π θ (a t ∣s t )) is the logarithmic probability of the current strategy, which represents the logarithm of the probability of choosing the current action. ω (s t ,a t ) is the output of the Q network, indicating that the given state s t and action a t where ω is the parameter of the Q network. α is the entropy weight coefficient, which controls the weight of the entropy term in the loss function. The entropy term encourages the policy to be more random when choosing actions, thereby enhancing exploration.

[0063] In the SAC algorithm, choosing the coefficient of the entropy regularization term is crucial. Different entropy values are required in different states: in states where the optimal action is uncertain, the entropy value should be higher; in states where the optimal action is relatively certain, the entropy value can be lower. To automatically adjust the entropy regularization term, SAC recasts the reinforcement learning objective as a constrained optimization problem: maximizing the expected reward while constraining the mean entropy value to be greater than a specified value.

[0064] In addition, a loss function with an entropy weight coefficient α needs to be designed. When the entropy of the policy is lower than the target value, the training objective will increase the value of α, thereby increasing the importance of the policy's entropy corresponding term in the process of minimizing the loss function mentioned above; when the entropy of the policy is higher than the target value, the training objective will decrease the value of α, thereby making the policy training more focused on value improvement:

[0065]

[0066] Indicates expectation, where s t is a state, obeying the distribution β,a t is an action, given a state st Under the strategy π(a t ∣s t ). -αlogπ(a t ∣s t ) is related to the policy entropy. Policy entropy reflects the uncertainty of the policy. The larger α is, the greater the weight of this term in the loss function, thus emphasizing the importance of policy entropy. The H0 in -αH0 is the target entropy value. This term is also related to α and is used to adjust the impact of α on the loss function based on the relationship between policy entropy and target entropy.

[0067] When we collect historical trajectory data in the loop playback buffer for algorithm training, in order to enable the algorithm to fully consider the correlation between states at different times in this trajectory and the potential characteristics of the states, we design a data preprocessing network actor for the Actor network and the Critic network. summarizer and Q summarizer Specifically, we first pass the state sequence data of a period of time (partial state after masking) through the actor summarizer and Q summarizer To generate the feature representation of the state sequence during this period (the output of the hidden state at each time step), and then input the feature representation into the Actor network and the Critic network to generate the final strategy and action value.

[0068] The Long Short-Term Memory (LSTM) network introduces a "gating" mechanism that uses three gating units (input gate, forget gate, and output gate) to decide which information should be retained and which should be discarded. This enables LSTM to effectively process information over a long period of time while retaining important historical information. We design actors based on the Long Short-Term Memory (LSTM) network. summarizer and Q summarizer , the network structure diagram is shown in Figure 2Specifically, we designed three LSTM-based preprocessing networks: the Actor Preprocessor, the Q1 Preprocessor, and the Q2 Preprocessor, corresponding to the feature representation modules in the figure, as well as three MLP networks: the Actor Network (corresponding to the Policy Network in the figure), the Action Value Network Q1 (corresponding to Q Network 1 in the figure), and the Action Value Network Q2 (corresponding to Q Network 2 in the figure). The Actor Preprocessor uses its output state representation as the input of the Policy Network. The Policy Network has a fully connected layer internal structure that outputs the mean and standard deviation, and selects the output action and the logarithmic probability of the action based on Gaussian distribution sampling. The Action Value Network Q1 has a fully connected layer internal structure that concatenates the output of the Q1 Preprocessor and the action output of the Policy Network as input, and outputs a state action value of 1. The Action Value Network Q2 has a fully connected layer internal structure that concatenates the output of the Q2 Preprocessor and the action output of the Policy Network as input, and outputs a state action value of 2.

[0069] The steps of the LSTM-based water-cooling unit energy consumption optimization algorithm adopted in the embodiment of the present invention are as follows:

[0070] Step 1: Select random network parameters to initialize the Q1 preprocessor Q2 Preprocessor Q1, Q2, copy the same parameters to initialize the Q1 preprocessor target network Q2 Preprocessor Target Network Q1 Target Network Q2 Target Network Select random network parameters to initialize actor, actor preprocessor (actor summarizer ), configure the optimizer for each network, set the initial value of α to 1, and configure its optimizer;

[0071] Step 2: Initialize the cyclic experience replay pool R;

[0072] Step 3: Loop through the following steps:

[0073] Step 3.1: Current observation passes through the actor summarizer Get the observed feature representation, and then pass the feature representation through the actor to get the action a that the agent needs to perform t ;

[0074] Step 3.2: Execute action a t , receive the next state s of the environment t+1 and reward r t , will s t 、a t 、r t 、s t+1 , done and other information are stored in the replay pool;

[0075] Step 3.3: For each training round, loop through the following steps:

[0076] Step 3.31: Randomly sample a batch of state sequences and their corresponding actions, rewards, and next states from the replay pool, including consecutive time steps;

[0077] Step 3.32: Calculate the target Q value based on the sampled data γ is the attenuation factor, set to 0.99, log_prob(a t+1 ) is an actor(actor summarizer (s t+1 )) Output the action log probability, using the target Q value y t and the current Q network's prediction value and Calculate the TD error and minimize the TD loss function and renew Q1, Q2;

[0078] Step 3.33: Update the policy network by minimizing the policy loss. The policy loss consists of a combination of the expected Q value and the log probability of the action, as well as an entropy term: Update actor,actor summarizer ;

[0079] Step 3.34: Adjust the value of α according to the target entropy: L(α) = -α*(log_prob(a t )-action_dim), update α, log_prob(a t ) is an actor(actor summarizer (s t )) Output action log probability, action_dim is the action space dimension, which is 20 in the experiment;

[0080] Step 3.35: Update the two target Q networks using soft update Parameters: τ is set to 0.95.

[0081] This example uses simulation experiments to verify the superiority of the designed method. We designed three sets of experiments: using the traditional SAC algorithm (without masking state variables), using the traditional SAC algorithm (masking 16 key variables in the 246-dimensional state), and using the LSAC algorithm (masking 16 key variables in the 246-dimensional state). The total number of training steps for the algorithm is 100,000.

[0082] like Figure 3 and Figure 4 Experimental results show that when using the traditional SAC algorithm for optimization, after masking 16 key variables, the algorithm's convergence speed slows down compared to the case where state variables are not masked, but the convergence results are comparable. When using the LSAC algorithm, even with 16 key variables masked, the algorithm's optimization results are still significantly improved compared to the traditional SAC algorithm (regardless of whether 16 key variables are masked), and the convergence curve is smoother. In summary, when masking key variables to simulate the missing sensor data in a real environment, the LSAC algorithm we designed is more stable than the traditional SAC algorithm and exhibits excellent optimization performance.

[0083] Based on the same inventive concept, an embodiment of the present invention discloses a data center water-cooling unit energy consumption optimization system based on LSAC, including: an acquisition module, used to collect real-time data from sensors in the water-cooling unit and its environment and perform cleaning and standardization; a water-cooling unit energy consumption modeling module, which constructs a water-cooling unit energy consumption model based on the collected data and simulates the energy consumption performance of the equipment under different working conditions; a reinforcement learning environment modeling and energy consumption optimization module, which is used to design a reinforcement learning environment based on the water-cooling unit energy consumption model, construct a Markov decision process for optimizing the energy consumption of the water-cooling unit, and a partially observable Markov decision process that partially masks the state space in the Markov decision process, and use the LSAC algorithm to optimize the operation strategy of the water-cooling unit to minimize energy consumption; wherein the LSAC algorithm uses an LSTM network to process historical time series data, and embeds the output part of the LSTM network into the Actor and Critic networks in the SAC algorithm.

[0084] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded into the processor, the steps of the LSAC-based data center water cooling unit energy consumption optimization method are implemented.

[0085] An embodiment of the present invention further discloses a computer program product, including a computer program, which, when executed by a processor, implements the steps of the LSAC-based data center water cooling unit energy consumption optimization method.

[0086] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for optimizing energy consumption of water cooling units in a data center based on LSAC, characterized in that: The following steps are involved: Collect, clean, and normalize real-time data from sensors in the chiller and its environment; Based on the collected data, an energy consumption model of the water-cooling unit is constructed to simulate the energy consumption performance of the equipment under different working conditions; Based on the energy consumption model of a water-cooling unit, a reinforcement learning environment is designed. A Markov decision process for optimizing the energy consumption of the water-cooling unit is constructed, as well as a partially observable Markov decision process that partially masks the state space in the Markov decision process. The LSAC algorithm is used to optimize the operation strategy of the water-cooling unit to minimize energy consumption. The LSAC algorithm uses an LSTM network to process historical time series data and embeds the output of the LSTM network into the actor and critic networks of the SAC algorithm.

2. The method for optimizing energy consumption of a water-cooling unit in a data center based on LSAC according to claim 1, characterized in that: The water-cooling unit energy consumption model is constructed using a data-driven approach combined with the multi-layer perceptron (MLP) in a deep neural network. The input of the water-cooling unit energy consumption model includes environmental parameters, equipment parameters, and load parameters, and the output is the normalized value of the water-cooling unit energy consumption in the current state.

3. The method for optimizing energy consumption of a water-cooling unit in a data center based on LSAC according to claim 1, characterized in that: The collected data includes at least the chilled water outlet temperature, chilled water outlet pressure, cooling water outlet temperature, cooling water outlet pressure, evaporator outlet temperature, cold storage percentage, phase voltage, phase current, water pan temperature, cold storage tank temperature, suction temperature and exhaust pressure.

4. The method for optimizing energy consumption of a water-cooling unit in a data center based on LSAC according to claim 1, characterized in that: The state space of the Markov decision process, including environmental parameters, equipment parameters and load parameters, is consistent with the input parameters of the constructed water-cooling unit energy consumption model; The action space is a key control parameter selected from the state space; the partially observable Markov decision process masks part of the information in the state space, so that the agent can only make decisions based on partial observations.

5. The method for optimizing energy consumption of water cooling units in a data center based on LSAC according to claim 1, characterized in that: In the reinforcement learning environment, the reward for each action of the agent is the opposite of the output value of the water-cooling unit energy consumption model in the current state.

6. The method for optimizing energy consumption of a water-cooling unit in a data center based on LSAC according to claim 1, characterized in that: The LSAC algorithm combines the LSTM network structure and the SAC reinforcement learning algorithm to design data preprocessing network actors for the Actor network and the Critic network. summarizer and Q summarizer , the state sequence data of a period of time is first passed through the actor summarizer and Q summarizer The feature representation of the state sequence during this period is generated, and then the feature representation is input into the Actor network and the Critic network to generate the final strategy and action value; a cyclic experience replay mechanism is used to access data during training, and fragments are randomly intercepted from each complete training trajectory to increase the total number and diversity of samples.

7. The method for optimizing energy consumption of water cooling units in a data center based on LSAC according to claim 1, characterized in that: The LSAC algorithm uses three LSTM-based preprocessing networks, namely Actor, Q1, and Q2 preprocessors, and three MLP networks, namely Actor network, action value network Q1, and Q2. The Actor preprocessor uses its output state representation as the input of the policy network. The policy network outputs the mean and standard deviation, and selects the output action and the logarithmic probability of the action based on Gaussian distribution sampling. The action value network Q1 / Q2 concatenates the output of the Q1 / Q2 preprocessor and the action output of the policy network as input, and outputs the corresponding state action value.

8. A data center water cooling unit energy consumption optimization system based on LSAC, characterized in that: include: The acquisition module is used to collect real-time data from sensors in the water-cooling unit and its environment and to clean and standardize it; The water-cooling unit energy consumption modeling module builds a water-cooling unit energy consumption model based on the collected data and simulates the energy consumption performance of the equipment under different working conditions; The reinforcement learning environment modeling and energy consumption optimization module is used to design a reinforcement learning environment based on the water-cooling unit energy consumption model, construct a Markov decision process for optimizing the water-cooling unit's energy consumption, and a partially observable Markov decision process that partially masks the state space in the Markov decision process. The LSAC algorithm is used to optimize the water-cooling unit's operating strategy to minimize energy consumption. The LSAC algorithm uses an LSTM network to process historical time series data and embeds the output of the LSTM network into the actor and critic networks in the SAC algorithm.

9. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the computer program is loaded into a processor, the steps of the LSAC-based data center water cooling unit energy consumption optimization method are implemented according to any one of claims 1 to 7.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the LSAC-based data center water cooling unit energy consumption optimization method are implemented according to any one of claims 1 to 7.