Method and system for optimizing energy consumption of data center refrigeration system based on meta-reinforcement learning

By constructing a Markov decision process for a data center cooling system and introducing a meta-reinforcement learning algorithm with action smoothing and constraint regularization, the problem of poor adaptability of energy consumption optimization strategies for data center cooling systems is solved, achieving intelligent energy-saving operation under complex working conditions and improving system efficiency and stability.

CN120742691BActive Publication Date: 2025-11-11HEFEI UNIV OF TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511205665.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-11
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing energy consumption optimization strategies for data center cooling systems are poorly adaptable and difficult to implement effectively under complex operating conditions and multi-objective collaborative optimization requirements.

Method used

We employ a meta-reinforcement learning-based approach to construct a Markov decision process for a data center cooling system. We introduce the soft actor critic algorithm with action smoothing and constraint regularization (ASCR-SAC) and construct a policy network through a meta-learning mechanism to achieve rapid policy transfer and adaptive optimization.

Benefits of technology

Intelligent energy-saving operation of data center cooling systems was achieved under different operating environments, improving system operating efficiency and energy consumption optimization, and enhancing the stability and adaptability of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120742691B_ABST
    Figure CN120742691B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for optimizing energy consumption in data center cooling systems based on meta-reinforcement learning, relating to the field of energy consumption optimization. First, a Markov decision process based on the operating characteristics of data center cooling systems is constructed. Second, a policy network is built using a meta-reinforcement learning algorithm. To address the requirements of continuity and physical feasibility of control actions in the cooling system, a soft-actor commentator algorithm incorporating action smoothing and constraint regularization is proposed to execute the inner loop policy update of the policy network. Finally, the policy network is used for dynamic optimization control of energy consumption. When a new task is identified, the policy network is fine-tuned by executing the inner loop policy update to generate the corresponding optimal control policy. This invention enables rapid policy migration and adaptive optimization across different operating environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of energy consumption optimization, specifically to a method and system for optimizing energy consumption in a data center cooling system based on meta-reinforcement learning. Background Technology

[0002] With the rapid development of information technology and the widespread application of technologies such as cloud computing, big data, and artificial intelligence, data centers, as core facilities supporting the operation of digital infrastructure, are constantly expanding in scale and increasing in number, resulting in a significant upward trend in overall energy consumption. Cooling systems account for a large proportion of the total energy consumption of data centers, indicating significant potential for optimization.

[0003] Data center cooling system energy consumption optimization refers to a control method that minimizes the total energy consumption of the system by dynamically adjusting the operating parameters of cooling-related equipment (such as partial load rate, pump frequency, fan speed, etc.) under the premise of meeting the data center cooling load and equipment safety operation constraints. In related technologies, this type of optimization typically relies on physical modeling or data-driven methods to construct an energy consumption model, and combines optimization algorithms or intelligent control strategies to determine the optimal control action, aiming to improve the overall system operating efficiency and reduce operating costs. For example, patent CNCN113848711A proposes a data center cooling control algorithm based on reinforcement learning of a security model. Another example is the invention patent CN117031931A, which proposes a data center cold source parameter optimization method, system, and medium based on reinforcement learning.

[0004] The above solutions can reduce system energy consumption and improve operating efficiency to some extent, but they still have many shortcomings when facing complex operating conditions and multi-objective collaborative optimization requirements. Summary of the Invention

[0005] (a) Technical problems to be solved

[0006] To address the shortcomings of existing technologies, this invention provides a method and system for optimizing energy consumption in data center cooling systems based on meta-reinforcement learning, which solves the technical problem of poor adaptability of energy consumption optimization strategies.

[0007] (II) Technical Solution

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] A method for optimizing energy consumption in a data center cooling system based on meta-reinforcement learning, comprising:

[0010] Based on the operating characteristics of data center cooling systems, the energy consumption optimization task of data center cooling systems under different operating environments is described as a corresponding Markov decision process.

[0011] A task set including energy consumption optimization tasks under various operating environments is constructed, and a policy network is constructed using a meta-reinforcement learning algorithm. Among them, an action smoothing regularization term and a feasibility constraint regularization term are introduced through the ASCR-SAC algorithm to perform the inner loop policy update of the policy network.

[0012] Identifying the current task based on the real-time operating environment and generating the optimal control strategy for the current task using the policy network includes:

[0013] If the current task or a similar task does not exist in the task set, the current task is treated as a new task. The policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated using the fine-tuned policy network.

[0014] Preferably, the step of identifying the current task based on the real-time operating environment and generating the optimal control strategy for the current task using the policy network further includes:

[0015] If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

[0016] Preferably, the data center cooling system consists of a chiller unit, a cooling water pump, a chilled water pump, and a cooling tower. The operating characteristics of the data center cooling system refer to the energy consumption model and cooling load balance constraints constructed based on the physical characteristics of each device.

[0017] Preferably, the construction process of the Markov decision process includes:

[0018] (1) Construct the state and action space of the data center cooling system, including:

[0019] exist t The operating environment information observed by the data center cooling system agent during a given time period includes: the required cooling capacity of the data center. Outdoor air density of data centers Cooling tower inlet air entropy chilled water outlet temperature chilled water inlet temperature Cooling water inlet temperature Cooling water outlet temperature chilled water flow rate Cooling water flow rate Cooling tower air volume ;

[0020] definition t The state vector of the data center cooling system for a given time period is:

[0021]

[0022] Among them, uncontrollable state space elements include Controllable state space elements include ;

[0023] exist t Actions that a data center cooling system agent can take during a given time period include: chiller partial load factor. chilled water pump speed ratio Cooling water pump speed ratio Cooling tower fan frequency ;

[0024] definition t The action vector for the data center cooling system during the specified time period is:

[0025]

[0026] (2) Constructing the reward function

[0027] The reward function is defined based on the total energy consumption of the system during a given period, combined with the supply and demand balance of cooling capacity and temperature constraints. t Time-based reward value for:

[0028]

[0029] in, for t Total energy consumption of data center cooling system during the period for t Cooling tower fan power during the period for t chilled water pump power during the period for t Cooling water pump power during the period for t Power consumption of the chiller unit during a given period; for t Penalties for imbalance between supply and demand of data center cooling load during certain time periods; for t Cooling capacity provided by the data center cooling system during a given time period; , Positive weighting coefficients are used to balance the importance of different optimization objectives;

[0030] (3) Constructing the Markov decision process

[0031] Optimize the energy consumption of data center cooling systems under each operating environment. Describe it as a specific Markov decision process:

[0032]

[0033] in, For the task The corresponding Markov decision process; For the task The state space below; For the task Action space; state transition probability Described in the task In the context of the current state and the actions taken In this case, the system transitions to the next state. And receive reward points The probability of; This is the corresponding reward function.

[0034] Preferably, the action smoothing regularization term refers to:

[0035]

[0036] in, This is a motion smoothing penalty, used to encourage smooth changes in motion over consecutive time intervals; The motion smoothing regularization coefficient; , They are respectively t Time period t- Actions during a single period;

[0037] The feasibility constraint regularization term refers to:

[0038]

[0039] in, This is a penalty item for the range of motion, used to limit the output of motion within the allowable operating range of the device; The regularization coefficient for action constraints; For action The j Dimensional components; , For the action number j The maximum and minimum values ​​allowed for each dimension.

[0040] Preferably, the step of constructing the policy network using a meta-reinforcement learning algorithm specifically includes:

[0041] (1) Inner loop strategy update

[0042] For any task in the task set m The policy network is locally updated using the ASCR-SAC algorithm, and the following policy loss function is constructed:

[0043]

[0044] in, For strategy parameters; As expected, For the task m Operational data; The temperature parameter controls the influence of the entropy term on the target value. This represents the logarithmic probability of action sampling, corresponding to the policy entropy; min represents taking the smaller value. , For two Q Network, Subscript , These are the parameters of the two Q-networks mentioned above;

[0045] The strategy uses gradient descent for updating, and the local adaptive update formula is:

[0046]

[0047] in, For the task m The adapted strategy parameters; The learning rate for the inner loop; For gradient operators;

[0048] (2) Update of outer loop parameters

[0049] After completing the local adaptation updates for all tasks in the task set, the following total meta-loss function is constructed:

[0050]

[0051] in, M For a set of tasks;

[0052] For the initial policy parameters Perform optimization updates:

[0053]

[0054] in, The learning rate is denoted as .

[0055] A data center cooling system energy consumption optimization system based on meta-reinforcement learning, comprising:

[0056] The task description module is used to describe the energy consumption optimization task of the data center cooling system under different operating environments as a corresponding Markov decision process based on the operating characteristics of the data center cooling system.

[0057] The network construction module is used to construct a task set that includes energy consumption optimization tasks under various operating environments, and to construct a policy network using a meta-reinforcement learning algorithm. Among them, the ASCR-SAC algorithm is used to introduce action smoothing regularization terms and feasibility constraint regularization terms to perform the inner loop policy update of the policy network.

[0058] A strategy optimization module is used to identify the current task based on the real-time operating environment and generate the optimal control strategy for the current task using the policy network, including:

[0059] If the current task or a similar task does not exist in the task set, the current task is treated as a new task. The policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated using the fine-tuned policy network.

[0060] Preferably, the strategy optimization module, used to identify the current task based on the real-time operating environment and generate the optimal control strategy for the current task using the policy network, further includes:

[0061] If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

[0062] A storage medium storing a computer program for optimizing the energy consumption of a data center cooling system based on meta-reinforcement learning, wherein the computer program causes a computer to execute the data center cooling system energy consumption optimization method as described above.

[0063] An electronic device, comprising:

[0064] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing the data center cooling system energy consumption optimization method as described above.

[0065] (III) Beneficial Effects

[0066] This invention provides a method and system for optimizing energy consumption in data center cooling systems based on meta-reinforcement learning. Compared with existing technologies, it has the following advantages:

[0067] This invention first constructs a Markov decision process based on the operating characteristics of a data center cooling system. Second, a meta-reinforcement learning algorithm is used to construct a policy network. To address the requirements of continuity and physical feasibility of control actions in the cooling system, a soft-actor critic algorithm incorporating action smoothing and constraint regularization is proposed to execute the inner loop policy update of the policy network. Finally, the policy network is used for dynamic optimization control of energy consumption. When a new task is identified, the policy network is fine-tuned by executing the inner loop policy update to generate the corresponding optimal control policy. This invention enables rapid policy migration and adaptive optimization across different operating environments. Attached Figure Description

[0068] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0069] Figure 1 A block diagram illustrating an energy consumption optimization method for a data center cooling system based on meta-reinforcement learning, provided in an embodiment of the present invention.

[0070] Figure 2 A block diagram illustrating another energy consumption optimization method for data center cooling systems based on meta-reinforcement learning, provided in an embodiment of the present invention.

[0071] Figure 3 This is a structural block diagram of a data center cooling system energy consumption optimization system based on meta-reinforcement learning, provided as an embodiment of the present invention. Detailed Implementation

[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0073] This application provides a data center cooling system energy consumption optimization method and system based on meta-reinforcement learning, which solves the technical problem of poor adaptability of energy consumption optimization strategies and realizes intelligent energy-saving operation of data center cooling systems under complex and variable operating conditions.

[0074] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows:

[0075] Traditional reinforcement learning methods suffer from low sample utilization, long training times, and slow convergence during the policy training phase. This is especially true in highly security-sensitive scenarios like data centers, where some historical operational data cannot be obtained or used for training due to data security or privacy compliance requirements, resulting in insufficient sample numbers and low training efficiency. Furthermore, traditional reinforcement learning strategies struggle to adapt to policy transfer between different operating environments, often requiring retraining for new tasks, leading to low practicality and deployment efficiency.

[0076] To address this, embodiments of the present invention provide a data center cooling system energy consumption optimization method based on meta-reinforcement learning, mainly involving the following key points:

[0077] First, a Markov decision model based on the operating characteristics of a data center cooling system is constructed. By defining a reasonable state space, action space, and reward function, the operating state, control variables, and optimization objectives of the system are quantified, providing a model foundation for reinforcement learning.

[0078] Secondly, to address the requirements of the cooling system for the continuity of control actions and physical feasibility, a soft actor-critic (ASCR-SAC) algorithm with action smoothing and constraint regularization is proposed to improve strategy stability and energy consumption optimization.

[0079] Then, based on the ASCR-SAC algorithm, a meta-reinforcement learning real-time control algorithm is constructed. By utilizing a set of multi-task Markov decision processes, the policy can quickly migrate and adapt under different seasons, loads, and climate conditions, thereby enhancing the algorithm's generalization ability and practical deployment efficiency.

[0080] Finally, the established meta-reinforcement learning control strategy is used to dynamically optimize and control the operating parameters of equipment such as chillers, pumps, and cooling towers. Under the premise of meeting the system's cooling load requirements and equipment safety constraints, the total energy consumption of the cooling system is minimized, thereby improving the operating efficiency and intelligence level of the data center cooling system.

[0081] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0082] Example 1:

[0083] like Figure 1 As shown, this embodiment of the invention provides a method for optimizing the energy consumption of a data center cooling system based on meta-reinforcement learning, including:

[0084] S1. Based on the operating characteristics of data center cooling systems, the energy consumption optimization task of data center cooling systems under different operating environments is described as a corresponding Markov decision process.

[0085] S2. Construct a task set that includes energy consumption optimization tasks under various operating environments, and use a meta-reinforcement learning algorithm to construct a policy network; wherein, the ASCR-SAC algorithm is used to introduce an action smoothing regularization term and a feasibility constraint regularization term to perform the inner loop policy update of the policy network.

[0086] S3. Identify the current task based on the real-time operating environment, and generate the optimal control strategy for the current task using the policy network, including:

[0087] S31. If there is no current task or similar task in the task set, the current task is taken as a new task, the policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated using the fine-tuned policy network.

[0088] In this embodiment of the invention, the energy consumption optimization task of a data center cooling system is described as a Markov decision process. An improved SAC algorithm, ASCR-SAC, incorporating action smoothing and constraint regularization terms, and a reinforcement learning control method integrating meta-learning mechanisms are introduced to construct a policy network for generating the optimal control strategy. This allows the control strategy to be continuously optimized during actual operation, enhancing the algorithm's stability and energy-saving performance in dynamic environments.

[0089] In an alternative implementation, such as Figure 2 As shown, the step of identifying the current task based on the real-time operating environment and generating the optimal control strategy for the current task using the policy network further includes:

[0090] S32. If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

[0091] It is understandable that if the current task has participated in the pre-training process of the above policy network, there is no need to fine-tune the policy network at this time, but the corresponding optimal control policy can be generated directly, which further improves the efficiency of regulation.

[0092] Data center cooling systems typically consist of various equipment such as chillers, chilled water pumps, cooling water pumps, and cooling towers. They maintain stable room temperatures through water circulation and heat exchange. Data center cooling systems are characterized by a wide variety of equipment, complex system structures, and significant thermodynamic coupling. Especially when loads change frequently or external weather conditions fluctuate drastically, the interconnectedness of operating states between equipment intensifies, significantly increasing the difficulty of system regulation and making it difficult to guarantee energy efficiency.

[0093] The energy consumption optimization of data center cooling systems, as focused on in this invention, refers to a control method that minimizes the total energy consumption of the system by dynamically adjusting the operating parameters of cooling-related equipment (such as partial load factor, pump frequency, and fan speed) while meeting the constraints of data center cooling load and equipment safe operation. This type of optimization typically relies on physical modeling or data-driven approaches to construct an energy consumption model, and combines this with optimization algorithms or intelligent control strategies to determine the optimal control action, aiming to improve the overall operating efficiency of the system and reduce operating costs.

[0094] Based on this, the following will detail each step of the above scheme:

[0095] In step S1, based on the operating characteristics of the data center cooling system, the energy consumption optimization task of the data center cooling system under different operating environments is described as a corresponding Markov decision process.

[0096] In this step, the data center cooling system consists of a chiller unit, a cooling water pump, a chilled water pump, and a cooling tower. The operating characteristics of the data center cooling system refer to the energy consumption model and cooling load balance constraints constructed based on the physical characteristics of the aforementioned equipment.

[0097] For example, the energy consumption model and cooling load balance constraints constructed from the physical characteristics of each device are as follows:

[0098] (1) Energy consumption model of chiller unit

[0099]

[0100] in, for t Power consumption (kW) of the chiller unit during the period; for t Cooling capacity (kW) provided by the time-period chiller unit; for t Energy efficiency ratio of the time-limited chiller; for t Partial load rate of chiller during the period; , They are respectively t The time period is unaffected by the temperature of chilled water and cooling water; , They are respectively t The ratio of chilled water to cooling water flow rate during a given period; , They are respectively t Time period chilled water and cooling water flow rates (kg / s); , They are respectively t Temperature of chilled water inlet and outlet during the specified time period (°C); for t Cooling water inlet temperature during the period (°C); , They are respectively t Rated inlet and outlet temperatures of chilled water during the specified time period (°C); , They are respectively t Rated cooling water inlet and outlet temperatures (°C) for the specified time period; , These are the rated chilled water and cooling water flow rates (kg / s), respectively. This refers to the rated cooling capacity of the chiller unit. , , The regression coefficient can be determined from the performance curve of the chiller unit.

[0101] (2) Cooling water pump energy consumption model

[0102]

[0103] in, for t Cooling water pump power (kW) for each period; for t Cooling water pump head (m) for a given period; for t Efficiency (%) of cooling water pump during the time period; for t Flow rate of cooling water pump during the period (m) 3 / s); for t Cooling water pump speed ratio during different time periods; , They are respectively t The efficiency of the cooling water pump motor during a given period, the efficiency of the frequency converter (%), and its value are related to... Related; for t The amount of heat transferred by the cooling water pump during a given period (kW); Cooling water density (kg / m³) 3 ); The specific heat capacity of cooling water [kJ / (kg·℃)]; for t Cooling water outlet temperature during the period (°C); for t Cooling water inlet temperature during the period (°C).

[0104] (3) Energy consumption model of chilled water pump

[0105]

[0106] in, fort Power of chilled water pump during a given period (kW); for t Chilled water pump head (m) for a given period; for t Efficiency (%) of chilled water pumps during different time periods; for t Flow rate of chilled water pump during a given period (m³) 3 / s); for t Time period chilled water pump speed ratio; , They are respectively t The efficiency of the chilled water pump motor and the efficiency of the frequency converter (%) during the time period, and their values ​​are related to... Related; for t Heat transferred by the chilled water pump during a given period (kW); The density of chilled water (kg / m³) 3 ); Specific heat capacity of chilled water [kJ / (kg·℃)]; for t Chilled water inlet temperature (°C) during the period; for t Chilled water outlet temperature (°C) during the period.

[0107] (4) Cooling tower energy consumption model

[0108]

[0109] in, for t Cooling tower fan power (kW) for a given period; Efficiency (%) of the cooling tower fan; C This is the pressure loss coefficient; for t Cooling tower air volume per period (m³) 3 / s); for t Air density during the period (kg / m³) 3 ); for t The amount of heat (kW) that the cooling tower needs to handle during a given period; for t Cooling tower outlet air entropy value [kJ / kg] for a given period; for t Entropy of the air entering the cooling tower during a given time period [kJ / kg].

[0110] (5) Cooling load balance constraint

[0111]

[0112] in, for t Cooling capacity (kW) required for the data center during a given time period.

[0113] Next, this step describes the energy consumption optimization task of the data center cooling system under different operating environments as a corresponding Markov decision process, specifically including:

[0114] (1) Construct the state and action space of the data center cooling system, including:

[0115] exist t The operating environment information observed by the data center cooling system agent during a given time period includes: the required cooling capacity of the data center. Outdoor air density of data centers Cooling tower inlet air entropy chilled water outlet temperature chilled water inlet temperature Cooling water inlet temperature Cooling water outlet temperature chilled water flow rate Cooling water flow rate Cooling tower air volume ;

[0116] definition t The state vector of the data center cooling system for a given time period is:

[0117]

[0118] Among them, the state space Some elements in the state space are random and unaffected by the cooling system's actions. Specifically: in the state space... Uncontrollable state space elements include Controllable state space elements include .

[0119] exist t Actions that a data center cooling system agent can take during a given time period include: chiller partial load factor. chilled water pump speed ratio Cooling water pump speed ratio Cooling tower fan frequency ;

[0120] definition t The action vector for the data center cooling system during the specified time period is:

[0121]

[0122] (2) Constructing the reward function

[0123] To optimize the energy consumption control of the data center cooling system, a reward function is designed to guide the agent to minimize overall energy consumption while meeting the data center's cooling load requirements and equipment operating constraints.

[0124] The reward function is defined based on the total energy consumption of the system during a given period, combined with the supply and demand balance of cooling capacity and temperature constraints. t Time-based reward value for:

[0125]

[0126] in, for t Total energy consumption of data center cooling system during the specified time period; for t Penalties for imbalance between supply and demand of data center cooling load during certain time periods; for t Cooling capacity provided by the data center cooling system during a given time period; , A positive weighting coefficient is used to balance the importance of different optimization objectives.

[0127] (3) Constructing the Markov decision process

[0128] As mentioned above, the embodiments of the present invention take into account the diversity of the operating environment of data center cooling systems, such as seasonal changes, different load characteristics, and differences in outdoor climate conditions. Here, multiple Markov decision processes are established to adapt to different operating environments.

[0129] Optimize the energy consumption of data center cooling systems under each operating environment. Describe it as a specific Markov decision process:

[0130]

[0131] in, For the task The corresponding Markov decision process; For the task The state space below; For the task Action space; state transition probability Described in the task In the context of the current state and the actions taken In this case, the system transitions to the next state. And receive reward points The probability of; This is the corresponding reward function.

[0132] For example, the above state transition probability It is expressed by the following formula:

[0133]

[0134] in, The next state after the action is executed; In fact A specific instance, representing the current state given. and perform actions In the case of the next state and rewards The probability distribution. Specifically, It contains two parts of information: one is the given state. and perform actions The system transitions to the next state. The probability; one is in the state Take action below The instant reward value obtained afterwards .

[0135] In step S2, a task set including energy consumption optimization tasks under various operating environments is constructed, and a policy network is constructed using a meta-reinforcement learning algorithm; wherein, an action smoothing regularization term and a feasibility constraint regularization term are introduced through the ASCR-SAC algorithm to perform the inner loop policy update of the policy network.

[0136] It should be noted that the meta-reinforcement learning algorithm introduced in this embodiment of the invention is based on the pre-constructed ASCR-SAC algorithm. Therefore, the aforementioned SCR-SAC algorithm will be introduced first:

[0137] This invention addresses the challenges of high operational continuity, large dynamic load changes, and precise energy consumption optimization in data center cooling systems. It proposes a deep reinforcement learning-based energy consumption optimization algorithm based on Action Smoothing and Constraint Regularization Soft Actor-Critic (ASCR-SAC).

[0138] The ASCR-SAC algorithm, based on the traditional SAC algorithm framework, introduces action smoothing regularization and action range constraint regularization, and further enhances the system's energy consumption optimization capability and training stability through a dynamic temperature parameter adjustment mechanism. This method balances the continuous operation requirements of data center cooling equipment with energy consumption optimization objectives, and can better adapt to the operating characteristics of data center cooling systems.

[0139] The specific process of the ASCR-SAC algorithm is as follows:

[0140] (1) Empirical sampling

[0141] Energy consumption optimization task in data center cooling systems During the process, the refrigeration system interacts with the environment to generate empirical data. The experience data from all tasks is combined to form the experience replay pool D.

[0142] A batch of empirical samples is randomly sampled from the empirical replay pool D:

[0143]

[0144] in, N This represents the batch sample size.

[0145] (2) Action generation and action regularization

[0146] In the ASCR-SAC algorithm, the policy network Based on the current state Output Action To improve the smoothness and rationality of the action output, the following two regularization terms are introduced:

[0147] ①Motion smoothing regularization term

[0148] To encourage smooth changes in movement over consecutive time intervals, a smooth movement penalty is introduced:

[0149]

[0150] in, This is a motion smoothing penalty, used to encourage smooth changes in motion over consecutive time intervals; The motion smoothing regularization coefficient; , These refer to the actions during time period t and time period t-1, respectively.

[0151] ②Action constraint regularization term

[0152] To limit the output of actions to within the device's permissible operating range, an action range penalty is introduced:

[0153]

[0154] in, This is a penalty item for the range of motion, used to limit the output of motion within the allowable operating range of the device; The regularization coefficient for action constraints; For action The j Dimensional components; , For the action number j The maximum and minimum values ​​allowed for each dimension.

[0155] (3) Objectives Q Value Calculation

[0156] Target Q The value can be the current Q Network updates provide a stable target, making the training process more stable.

[0157] In the next state Next, generate new actions from the policy network:

[0158]

[0159] That is, from the strategy distribution China New actions were obtained from sampling. .

[0160] Target Q value The calculation formula is:

[0161]

[0162] in, , For the purpose of delayed updates Q network; This is a reward discount factor, with a value between 0 and 1, used to measure the importance of future rewards; The temperature parameter controls the influence of the entropy term on the target value. This represents the logarithmic probability of action sampling, corresponding to the policy entropy.

[0163] (4) Q Network Update

[0164] Q The network's role is to estimate the expected reward for a given state-action pair. This is achieved by minimizing the current... Q Network and Target Q The mean square error between values Q The network is better able to estimate the expected reward for a given state and action pair.

[0165] Minimize the current Q Network and Target Q Mean squared error loss between values:

[0166]

[0167] in, For the present QNetwork in input state and actions The following prediction Q Values. Updated separately using gradient descent. and The parameter is k, which is the index of the Q network.

[0168] (5) Policy network update

[0169] Policy networks define the behavioral policies of agents, with the optimization goal of maximizing expected reward and action exploration, while also taking into account action smoothness and action legitimacy.

[0170] Policy network parameters By minimizing the policy loss function renew:

[0171]

[0172] in, and These are motion smoothing and motion range regularization terms, respectively.

[0173] (6) Temperature parameter update

[0174] To ensure the agent maintains appropriate exploratory capabilities during training, this invention employs a temperature parameter. Adaptive adjustment mechanism. This involves dynamically adjusting temperature parameters. This makes the policy entropy as close as possible to the target entropy value.

[0175] Temperature parameters By minimizing the temperature loss function renew:

[0176]

[0177] in, The preset target entropy value is used to control the degree of strategy exploration. Temperature parameters are updated via gradient descent. .

[0178] (7) Objective Q Network soft update:

[0179]

[0180] in, It is a soft update coefficient used to smoothly adjust the target. Q The internet helps avoid drastic fluctuations in the learning process.

[0181] Based on the ASCR-SAC algorithm described above, this invention establishes a meta-reinforcement learning real-time control algorithm. This algorithm optimizes the initial policy parameters of the agent through joint training under various cooling task environments, enabling it to quickly adapt to new environments with minimal interactions. This significantly improves the generalization performance and real-time response capability of data center cooling equipment energy consumption optimization. The goal of the algorithm's meta-task is to learn a set of meta-parameters. This enables reinforcement learning agents to handle any new task. In this way, high-performance strategies can be quickly obtained through training with a small amount of data.

[0182] Based on this, this step uses a meta-reinforcement learning algorithm to construct the policy network, specifically including:

[0183] (1) Inner loop strategy update

[0184] For any task in the task set m The policy network is locally updated using the ASCR-SAC algorithm, and the following policy loss function is constructed:

[0185]

[0186] in, For strategy parameters; As expected, For the task m Operational data; The temperature parameter controls the influence of the entropy term on the target value. This represents the logarithmic probability of action sampling, corresponding to the policy entropy; min represents taking the smaller value. , For two Q Network, Subscript , These are the parameters of the two Q-networks mentioned above;

[0187] The strategy uses gradient descent for updating, and the local adaptive update formula is:

[0188]

[0189] in, For the task m The adapted strategy parameters; The learning rate for the inner loop; For gradient operators;

[0190] (2) Update of outer loop parameters

[0191] After completing the local adaptation updates for all tasks in the task set, the following total meta-loss function is constructed:

[0192]

[0193] in, M For a set of tasks;

[0194] For the initial policy parameters Perform optimization updates:

[0195]

[0196] in, The learning rate is denoted as .

[0197] Through continuous iteration, the policy network acquires task migration capabilities, enabling rapid adaptive control when facing new tasks.

[0198] In step S3, the current task is identified based on the real-time operating environment, and the optimal control strategy for the current task is generated using the policy network, including:

[0199] S31. If there is no current task or similar task in the task set, the current task is taken as a new task, the policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated by using the fine-tuned policy network.

[0200] S32. If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

[0201] For example, after updating the policy network parameters based on meta-reinforcement learning, the trained model can be deployed into the actual data center cooling system control framework. Dynamic optimization of the cooling system's energy consumption can be achieved through scheduling and control of the cooling equipment. For instance, the trained meta-reinforcement learning policy model can be deployed to the cooling system's intelligent control platform. When the control system detects that load changes exceed a preset threshold or external environmental parameter changes exceed a set range, it can trigger small-batch data retraining to perform inner-loop fine-tuning and update the policy parameters, maintaining the system's adaptive capability.

[0202] Thus, this embodiment of the invention completes the entire process of the energy consumption optimization method for data center cooling systems based on meta-reinforcement learning.

[0203] Example 2:

[0204] like Figure 3 As shown, this embodiment of the invention provides an energy consumption optimization system for data center cooling systems based on meta-reinforcement learning, comprising:

[0205] The task description module is used to describe the energy consumption optimization task of the data center cooling system under different operating environments as a corresponding Markov decision process based on the operating characteristics of the data center cooling system.

[0206] The network construction module is used to construct a task set that includes energy consumption optimization tasks under various operating environments, and to construct a policy network using a meta-reinforcement learning algorithm. Among them, the ASCR-SAC algorithm is used to introduce action smoothing regularization terms and feasibility constraint regularization terms to perform the inner loop policy update of the policy network.

[0207] A strategy optimization module is used to identify the current task based on the real-time operating environment and generate the optimal control strategy for the current task using the policy network, including:

[0208] If the current task or a similar task does not exist in the task set, the current task is treated as a new task. The policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated using the fine-tuned policy network.

[0209] In an optional implementation, the policy optimization module, configured to identify the current task based on the real-time runtime environment and generate the optimal control policy for the current task using the policy network, further includes:

[0210] If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

[0211] Example 3:

[0212] This invention provides a storage medium storing a computer program for optimizing the energy consumption of a data center cooling system based on meta-reinforcement learning, wherein the computer program causes a computer to execute the data center cooling system energy consumption optimization method as described in Embodiment 1.

[0213] Example 4:

[0214] This invention provides an electronic device, comprising:

[0215] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing energy consumption optimization methods for data center cooling systems as described in Example 1.

[0216] It is understood that the data center cooling system energy consumption optimization system, storage medium and electronic device based on meta-reinforcement learning provided in the embodiments of the present invention correspond to the data center cooling system energy consumption optimization method based on meta-reinforcement learning provided in the embodiments of the present invention. The explanation, examples and beneficial effects of the relevant contents can be referred to the corresponding parts of the data center cooling system energy consumption optimization method, and will not be repeated here.

[0217] In summary, compared with existing technologies, it has the following beneficial effects:

[0218] 1. The embodiments of the present invention enable the control strategy to be continuously optimized in actual operation, thereby enhancing the stability and energy-saving effect of the algorithm in dynamic environments.

[0219] 2. To address the instability of traditional reinforcement learning in continuous action control, this invention proposes the ASCR-SAC algorithm, which introduces action smoothing and constraint regularization terms, to improve the stability and practical feasibility of the control strategy and enhance its deployment effect in physical systems.

[0220] 3. Considering the frequent changes in the data center operating environment and the limited sample acquisition, this embodiment of the invention designs a reinforcement learning control method that integrates meta-learning mechanism. Through multi-task training, it realizes the rapid transfer and adaptation of control strategies, and improves the system's generalization ability to new operating conditions.

[0221] 4. In the design of the reward function for the Markov decision process, the embodiments of the present invention comprehensively consider the two objectives of minimizing energy consumption and balancing the supply and demand of cooling load, and introduce adjustable weight coefficients so that the control strategy can flexibly adjust the focus of the objectives according to different optimization needs, thereby improving the accuracy and practicality of regulation.

[0222] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0223] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing energy consumption in a data center cooling system based on meta-reinforcement learning, characterized in that, include: Based on the operating characteristics of data center cooling systems, the energy consumption optimization task of data center cooling systems under different operating environments is described as a corresponding Markov decision process. A task set including energy consumption optimization tasks under various operating environments is constructed, and a policy network is constructed using a meta-reinforcement learning algorithm. Among them, an action smoothing regularization term and a feasibility constraint regularization term are introduced through the ASCR-SAC algorithm to perform the inner loop policy update of the policy network. Identifying the current task based on the real-time operating environment and generating the optimal control strategy for the current task using the policy network includes: If the current task or a similar task does not exist in the task set, the current task is treated as a new task, the policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated using the fine-tuned policy network. The action smoothing regularization term refers to: in, This is a motion smoothing penalty, designed to encourage smooth changes in motion over consecutive time intervals; The motion smoothing regularization coefficient; , They are respectively t Time period t- Actions during a single period; The feasibility constraint regularization term refers to: in, This is a penalty item for the range of actions, used to limit the output of actions within the allowable operating range of the device; The regularization coefficient for action constraints; For action The j Dimensional components; , For the action number j The maximum and minimum allowed values ​​for each dimension; The construction of the policy network using a meta-reinforcement learning algorithm specifically includes: (1) Inner loop strategy update For any task in the task set m The policy network is locally updated using the ASCR-SAC algorithm, and the following policy loss function is constructed: in, For strategy parameters; As expected, For the task m Operational data; The temperature parameter controls the influence of the entropy term on the target value. This represents the logarithmic probability of action sampling, corresponding to the policy entropy; min represents taking the smaller value. , For two Q Network, Subscript , These are the parameters of the two Q-networks mentioned above; The strategy uses gradient descent for updating, and the local adaptive update formula is: in, For the task m The adapted strategy parameters; The learning rate for the inner loop; For gradient operators; (2) Update of outer loop parameters After completing the local adaptation updates for all tasks in the task set, the following total meta-loss function is constructed: in, M For a set of tasks; For the initial policy parameters Perform optimization updates: in, The learning rate is denoted as .

2. The data center cooling system energy consumption optimization method as described in claim 1, characterized in that, The method of identifying the current task based on the real-time operating environment and generating the optimal control strategy for the current task using the policy network further includes: If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

3. The data center cooling system energy consumption optimization method as described in claim 1, characterized in that, The data center cooling system consists of chiller units, cooling water pumps, chilled water pumps, and cooling towers. The operating characteristics of the data center cooling system refer to the energy consumption model and cooling load balance constraints constructed based on the physical characteristics of each device.

4. The data center cooling system energy consumption optimization method as described in claim 3, characterized in that, The construction process of the Markov decision process includes: (1) Construct the state and action space of the data center cooling system, including: exist t The operating environment information observed by the data center cooling system agent during a given time period includes: the required cooling capacity of the data center. Outdoor air density of data centers Cooling tower inlet air entropy chilled water outlet temperature chilled water inlet temperature Cooling water inlet temperature Cooling water outlet temperature chilled water flow rate Cooling water flow rate Cooling tower air volume ; definition t The state vector of the data center cooling system for a given time period is: Among them, uncontrollable state space elements include Controllable state space elements include ; exist t Actions that a data center cooling system agent can take during a given time period include: chiller partial load factor. chilled water pump speed ratio Cooling water pump speed ratio Cooling tower fan frequency ; definition t The action vector for the data center cooling system during the specified time period is: (2) Constructing the reward function The reward function is defined based on the total energy consumption of the system during a given period, combined with the supply and demand balance of cooling capacity and temperature constraints. t Time-based reward value for: in, for t Total energy consumption of data center cooling system during the period for t Cooling tower fan power during the period for t chilled water pump power during the period for t Cooling water pump power during the period for t Power consumption of the chiller unit during a given period; for t Penalties for imbalance between supply and demand of data center cooling load during certain time periods; for t Cooling capacity provided by the data center cooling system during a given time period; , Positive weighting coefficients are used to balance the importance of different optimization objectives; (3) Constructing the Markov decision process Optimize the energy consumption of data center cooling systems under each operating environment. Describe it as a specific Markov decision process: in, For the task The corresponding Markov decision process; For the task The state space below; For the task Action space; state transition probability Described in the task In the context of the current state and the actions taken In this case, the system transitions to the next state. And receive reward points The probability of; This is the corresponding reward function.

5. A data center cooling system energy consumption optimization system based on meta-reinforcement learning, characterized in that, The method for optimizing the energy consumption of a data center cooling system as described in claim 1 includes: The task description module is used to describe the energy consumption optimization task of the data center cooling system under different operating environments as a corresponding Markov decision process based on the operating characteristics of the data center cooling system. The network construction module is used to construct a task set that includes energy consumption optimization tasks under various operating environments, and to construct a policy network using a meta-reinforcement learning algorithm. Among them, the ASCR-SAC algorithm is used to introduce action smoothing regularization terms and feasibility constraint regularization terms to perform the inner loop policy update of the policy network. A strategy optimization module is used to identify the current task based on the real-time operating environment and generate the optimal control strategy for the current task using the policy network, including: If the current task or a similar task does not exist in the task set, the current task is treated as a new task. The policy network is updated and fine-tuned by executing the inner loop policy, and the optimal control policy for the new task is generated using the fine-tuned policy network.

6. The data center cooling system energy consumption optimization system as described in claim 5, characterized in that, The strategy optimization module is used to identify the current task based on the real-time operating environment and generate the optimal control strategy for the current task using the policy network, and further includes: If the current task or a similar task exists in the task set, the optimal control strategy for the current task is directly generated using the policy network.

7. A storage medium, characterized in that, It stores a computer program for optimizing the energy consumption of a data center cooling system based on meta-reinforcement learning, wherein the computer program causes a computer to execute the data center cooling system energy consumption optimization method as described in any one of claims 1 to 4.

8. An electronic device, characterized in that, include: One or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the programs including methods for performing the data center cooling system energy consumption optimization method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Energy-saving control method for air conditioners of data center based on federal reinforcement learning

    CN113551373A

  • Energy-saving control method and system for refrigerating system of central air conditioner based on deep learning

    CN118856530A