Energy-saving method and device for computer room air conditioner
Patent Information
- Application Number
- CN202310807022.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-03
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-07-03
AI Technical Summary
[0005]目前的机房空调AI节能方法主要采用了静态设定模式,即根据过去一段时间的历史数据以及预测数据进行优化控制,大多采用人工进行控制,无法实现自动调整
[0087]本发明实施例提供的机房空调的节能方法,实时动态巡航监测目标机房的热负荷数据,收集所述目标机房的环境温湿度实时数据,基于强化学习算法,对所述热负荷数据和所述环境温湿度实时数据进行分析,确定所述目标机房的节能策略,基于深度强化学习算法和所述节能策略,确定所述目标机房在当前状态下的最优控制策略,并将所述最优控制策略下发至所述目标机房的空调系统,能够根据机房热负荷动态变化情况,智能调节机房空调系统,实现机房空调的智能化控制和最大限度的节能效果。
Smart Images

Figure CN116963461B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and in particular to an energy-saving method and an energy-saving device for a computer room air conditioner. Background Technology
[0002] With the continuous development of information technology, the operating load of data centers is constantly increasing. Data center air conditioning has become one of the heaviest-loading devices in a data center, consuming about one-third of the total energy. Therefore, how to rationally utilize data center heat load data for control, achieving maximum energy savings while ensuring data center operation, has become a hot topic in the field of data center air conditioning.
[0003] Currently, energy-saving methods for data center air conditioning mainly include the following: traditional timed on / off energy-saving methods, quantitative airflow energy-saving methods, and adaptive temperature energy-saving methods. Traditional timed on / off energy-saving methods suffer from low energy efficiency and inability to adapt to dynamic changes in data center load. Quantitative airflow energy-saving methods require complex aerodynamic simulation design, leading to inconvenience in use and maintenance. Adaptive temperature energy-saving methods are difficult to implement in practical applications or suffer from over-reliance on environmental stability, leading to control failures. Furthermore, traditional air conditioning control methods often rely on temperature sensor feedback signals. While this method is simple, it also has significant drawbacks, such as sudden temperature changes or excessively high or low temperatures in parts of the data center within a short period.
[0004] In recent years, artificial intelligence (AI) technology has developed rapidly, and its impact on energy conservation has become very significant, giving rise to AI-based energy-saving technology for data center air conditioning. Applying AI technology to data center air conditioning control systems can achieve energy savings while maintaining a constant temperature in the data center. Because data center air conditioning is a highly dynamic and complex system encompassing multiple parameters, AI-based energy-saving methods, with their adaptive and intelligent characteristics, can quickly process the ever-changing load data of the data center, making them an important energy-saving approach.
[0005] Current AI energy-saving methods for data center air conditioning mainly adopt a static setting mode, which optimizes control based on historical and predicted data over a period of time. This is mostly done manually and cannot achieve automatic adjustment.
[0006] Therefore, how to achieve AI-powered energy saving based on dynamic cruise of data center heat load, and with self-learning, self-adaptation, and self-adjustment has become an important issue that urgently needs to be addressed. Summary of the Invention
[0007] To address the shortcomings of existing technologies, embodiments of the present invention provide an energy-saving method and an energy-saving device for a computer room air conditioner.
[0008] In a first aspect, embodiments of the present invention provide an energy-saving method for a computer room air conditioner, comprising:
[0009] Real-time dynamic cruise monitoring of the target computer room's heat load data;
[0010] Collect real-time environmental temperature and humidity data of the target computer room;
[0011] Based on reinforcement learning algorithms, the heat load data and the real-time environmental temperature and humidity data are analyzed to determine the energy-saving strategy for the target computer room.
[0012] Based on the deep reinforcement learning algorithm and the energy-saving strategy, the optimal control strategy for the target computer room in the current state is determined, and the optimal control strategy is sent to the air conditioning system of the target computer room.
[0013] Optionally, as described above, the step of analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm to determine the energy-saving strategy for the target computer room includes:
[0014] The security strategy for the target data center is determined based on its maintenance and protection level and power efficiency.
[0015] The security strategy, the heat load data, and the real-time environmental temperature and humidity data are input into a strategy network determined by a reinforcement learning algorithm. The energy-saving strategy for the target computer room is determined from the output of the strategy network.
[0016] Optionally, as described above, the policy network is determined by the following method:
[0017] The air conditioning system of the target computer room is abstracted as an intelligent agent;
[0018] Based on a reinforcement learning-based model-free off-policy algorithm, the agent selects different air conditioning system control actions in different states, where each state includes safety information, heat load information, and environmental information.
[0019] As described above, optionally, the reinforcement learning-based model-free off-policy algorithm enables the agent to select different air conditioning system control actions under different states, including:
[0020] The heat load data, ambient temperature and humidity data, and security policy data of each computer room were used as sample data.
[0021] The sample data is transformed into multiple nodes in the state space, where each node includes: a safe temperature setpoint for the computer room, a predicted power efficiency value for the computer room, and a predicted heat load value for the computer room.
[0022] Define a set of air conditioning system control actions for each node;
[0023] The Actor network is trained using the Actor-Critic method in the Off-policy algorithm, where the input of the Actor network is the nodes in the state space, and the output of the Actor network is the air conditioning system control action corresponding to the node.
[0024] The air conditioning system control actions output by the Actor network are compared with the actual air conditioning system control actions taken, and the value of the air conditioning system control actions output by the Actor network is estimated using a Critic network.
[0025] Use the value to optimize the parameters of the Actor network;
[0026] The optimized Actor network is used as the policy network.
[0027] As described above, optionally, the energy-saving strategy for the target computer room can be determined from the output of the policy network, including:
[0028] The learning network and target network of the reinforcement learning algorithm are used to construct the target value of the operating status reward of the air conditioning system in the target computer room and the hot spot penalty data of the target computer room.
[0029] The score of the output result is determined based on the target value of the running status reward and the hot spot penalty data;
[0030] If the score of the output result is lower than the preset value, the score of the output result is fed back to the policy network so that the policy network can redetermine the new output result;
[0031] The air conditioning system control action corresponding to the highest score in the output results will be used as the energy-saving strategy for the target computer room.
[0032] Optionally, as described above, determining the optimal control strategy for the target computer room in its current state based on the deep reinforcement learning algorithm and the energy-saving strategy includes:
[0033] The indoor environmental indicators of two consecutive time series corresponding to the target computer room, the feedback of the previous time series, and the air conditioning system control action of the previous time series are used as state inputs into the Q network of the deep reinforcement learning algorithm, wherein the air conditioning system control action is determined according to the energy-saving strategy.
[0034] The error gradient is constructed through backpropagation of the error using a Q-network;
[0035] The parameters of the Q-network are updated using the backpropagation algorithm.
[0036] The Q-network is used to predict and select the output action in the current state;
[0037] The epsilon-greedy algorithm is used for optimization to output the optimal control strategy.
[0038] Optionally, as described above, the Q-network is determined in the following manner:
[0039] Determine the various states of the target computer room, including the ambient temperature data, humidity data, and heat load data of the target computer room;
[0040] Determine the control actions of the air conditioning system under various conditions;
[0041] A deep neural network model Q-network is constructed based on the state and the control actions of the air conditioning system.
[0042] The deep neural network model Q-network is trained, and the optimal control policy is learned based on the reward value.
[0043] Alternatively, as described above, the reward value can be determined according to the following formula:
[0044] reward = -(1-comfort) + energy
[0045] Here, comfort refers to the level of comfort, energy is the energy consumption of the air conditioning system, and reward is the reward value.
[0046] Secondly, embodiments of the present invention provide an energy-saving device for a computer room air conditioner, comprising:
[0047] The cruise monitoring module is used for real-time dynamic cruise monitoring of the target computer room's heat load data;
[0048] The collection module is used to collect real-time environmental temperature and humidity data of the target computer room;
[0049] The analysis module is used to analyze the heat load data and the real-time environmental temperature and humidity data based on reinforcement learning algorithms to determine the energy-saving strategy for the target computer room.
[0050] The execution module is used to determine the optimal control strategy for the target computer room in the current state based on the deep reinforcement learning algorithm and the energy-saving strategy, and to send the optimal control strategy to the air conditioning system of the target computer room.
[0051] Optionally, as described in the above apparatus, the analysis module includes:
[0052] The security boundary construction unit is used to determine the security strategy of the target computer room based on the maintenance and protection level of the target computer room and the power usage efficiency of the target computer room.
[0053] The AI energy-saving learning unit is used to input the security policy, the heat load data, and the real-time environmental temperature and humidity data into a policy network determined based on a reinforcement learning algorithm, and to determine the energy-saving policy for the target computer room from the output of the policy network.
[0054] Optionally, as described in the above device, the AI energy-saving learning unit includes:
[0055] Abstract subunit, used to abstract the air conditioning system of the target computer room into an intelligent agent;
[0056] The training subunit is used to enable the agent to select different air conditioning system control actions in different states based on the reinforcement learning model-free off-policy algorithm, where each state includes safety information, heat load information and environmental information.
[0057] As described above, optionally, the training subunit is specifically used for:
[0058] The heat load data, ambient temperature and humidity data, and security policy data of each computer room were used as sample data.
[0059] The sample data is transformed into multiple nodes in the state space, where each node includes: a safe temperature setpoint for the computer room, a predicted power efficiency value for the computer room, and a predicted heat load value for the computer room.
[0060] Define a set of air conditioning system control actions for each node;
[0061] The Actor network is trained using the Actor-Critic method in the Off-policy algorithm, where the input of the Actor network is the nodes in the state space, and the output of the Actor network is the air conditioning system control action corresponding to the node.
[0062] The air conditioning system control actions output by the Actor network are compared with the actual air conditioning system control actions taken, and the value of the air conditioning system control actions output by the Actor network is estimated using a Critic network.
[0063] Use the value to optimize the parameters of the Actor network;
[0064] The optimized Actor network is used as the policy network.
[0065] As described above, optionally, the analysis module 530 further includes: an energy-saving strategy issuing unit, the energy-saving strategy issuing unit specifically used for:
[0066] The learning network and target network of the reinforcement learning algorithm are used to construct the target value of the operating status reward of the air conditioning system in the target computer room and the hot spot penalty data of the target computer room.
[0067] The score of the output result is determined based on the target value of the running status reward and the hot spot penalty data;
[0068] If the score of the output result is lower than the preset value, the score of the output result is fed back to the policy network so that the policy network can redetermine the new output result;
[0069] The air conditioning system control action corresponding to the highest score in the output results will be used as the energy-saving strategy for the target computer room.
[0070] As described above, optionally, the execution module is specifically used for:
[0071] The indoor environmental indicators of two consecutive time series corresponding to the target computer room, the feedback of the previous time series, and the air conditioning system control action of the previous time series are used as state inputs into the Q network of the deep reinforcement learning algorithm, wherein the air conditioning system control action is determined according to the energy-saving strategy.
[0072] The error gradient is constructed through backpropagation of the error using a Q-network;
[0073] The parameters of the Q-network are updated using the backpropagation algorithm.
[0074] The Q-network is used to predict and select the output action in the current state;
[0075] The epsilon-greedy algorithm is used for optimization to output the optimal control strategy.
[0076] As described above, optionally, the execution module is configured to determine the Q network according to the following manner:
[0077] Determine the various states of the target computer room, including the ambient temperature data, humidity data, and heat load data of the target computer room;
[0078] Determine the control actions of the air conditioning system under various conditions;
[0079] A deep neural network model Q-network is constructed based on the state and the control actions of the air conditioning system.
[0080] The deep neural network model Q-network is trained, and the optimal policy is learned based on the reward value.
[0081] As described above, optionally, the execution module is configured to determine the reward value according to the following formula:
[0082] reward = -(1-comfort) + energy
[0083] Here, comfort refers to the level of comfort, energy is the energy consumption of the air conditioning system, and reward is the reward value.
[0084] Thirdly, embodiments of the present invention provide an electronic device, comprising:
[0085] The system includes a memory and a processor, which communicate with each other via a bus. The memory stores program instructions executable by the processor, which can then execute the following methods: real-time dynamic cruise monitoring of the target computer room's heat load data; collection of real-time ambient temperature and humidity data of the target computer room; analysis of the heat load data and the real-time ambient temperature and humidity data based on a reinforcement learning algorithm to determine an energy-saving strategy for the target computer room; and determination of the optimal control strategy for the target computer room in its current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and distribution of the optimal control strategy to the air conditioning system of the target computer room.
[0086] Fourthly, embodiments of the present invention provide a storage medium storing a computer program thereon. When executed by a processor, the computer program implements the following method: real-time dynamic cruise monitoring of heat load data of a target computer room; collecting real-time environmental temperature and humidity data of the target computer room; analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm to determine an energy-saving strategy for the target computer room; determining the optimal control strategy for the target computer room in the current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and distributing the optimal control strategy to the air conditioning system of the target computer room.
[0087] The energy-saving method for air conditioning in a data center provided in this invention involves real-time dynamic monitoring of the heat load data of the target data center, collecting real-time environmental temperature and humidity data of the target data center, analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm, determining the energy-saving strategy for the target data center, determining the optimal control strategy for the target data center under the current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and distributing the optimal control strategy to the air conditioning system of the target data center. This method can intelligently adjust the air conditioning system of the data center according to the dynamic changes in the heat load of the data center, thereby achieving intelligent control of the data center air conditioning and maximizing energy-saving effects. Attached Figure Description
[0088] Figure 1This is a flowchart illustrating the steps of an embodiment of an energy-saving method for a computer room air conditioner according to the present invention;
[0089] Figure 2 This is a schematic diagram of the learning network and the target network in an embodiment of an energy-saving method for air conditioning in a computer room according to the present invention;
[0090] Figure 3 This is a flowchart of an energy-saving method for computer room air conditioning based on the DQN algorithm in an embodiment of the present invention.
[0091] Figure 4 This is a flowchart illustrating the steps of another embodiment of the energy-saving method for air conditioning in a computer room according to the present invention;
[0092] Figure 5 This is a structural block diagram of an embodiment of an energy-saving device for a computer room air conditioner according to the present invention;
[0093] Figure 6 This is a structural block diagram of an embodiment of an electronic device according to the present invention. Detailed Implementation
[0094] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0095] Reference Figure 1 The diagram illustrates a flowchart of an embodiment of an energy-saving method for a computer room air conditioner according to the present invention, which may specifically include the following steps:
[0096] Step S110: Real-time dynamic cruise monitoring of the target computer room's heat load data;
[0097] Specifically, current AI-based energy-saving methods for data center air conditioning mainly adopt a static setting mode, that is, optimizing control based on historical and predicted data over a period of time. This largely relies on manual control and cannot achieve automatic adjustment. To address this issue, this invention provides a novel AI-based energy-saving method for data center air conditioning based on dynamic data center load monitoring. In this embodiment, the heat load data of the target data center is monitored in real-time. The target data center is the data center requiring air conditioning energy saving, and the heat load data includes all heat load data within the target data center, such as the equipment heat load generated by each device, the building heat load generated by the target data center building, and the personnel heat load generated by personnel entering and exiting the target data center.
[0098] To collect this heat load information, real-time dynamic monitoring is conducted on the voltage and current of various devices within the target computer room to compile heat load data. For example, real-time dynamic monitoring monitors the voltage and current values of each rack device within the target computer room, thus providing the overall heat load information. Real-time detection of personnel entry and exit is also performed to determine the personnel heat load data. Real-time monitoring of the target computer room building's heat load data is also conducted, including information such as the building's insulation and the outdoor ambient temperature. By dynamically monitoring the heat sources within the computer room in real-time, this heat load data is aggregated and used as the overall heat load data for the target computer room.
[0099] Step S120: Collect real-time environmental temperature and humidity data of the target computer room;
[0100] Specifically, the target computer room's ambient temperature and humidity are collected in real time using ambient temperature and humidity sensors.
[0101] Step S130: Based on reinforcement learning algorithm, analyze the heat load data and the real-time environmental temperature and humidity data to determine the energy-saving strategy for the target computer room;
[0102] Specifically, based on reinforcement learning algorithms, heat load data and real-time ambient temperature and humidity data are analyzed to output control strategies for the air supply / return of the target computer room air conditioning equipment, such as proportional band setting, fan speed, and compressor frequency.
[0103] Step S140: Based on the deep reinforcement learning algorithm and the energy-saving strategy, determine the optimal control strategy for the target computer room in the current state, and send the optimal control strategy to the air conditioning system of the target computer room.
[0104] Specifically, to prevent safety issues such as uncontrolled air conditioning in the data center due to AI control platform malfunctions, an edge system or server can be set up to receive energy-saving strategies from the target data center. The edge system can be a field supervision unit (FSU) in environmental monitoring. After receiving the energy-saving strategy, the edge system determines the optimal control strategy for the target data center in its current state based on a deep reinforcement learning algorithm and the energy-saving strategy. The optimal control strategy includes switching the air conditioning equipment on and off, adjusting the output frequency of the air conditioning compressor and fan, and adjusting parameters such as temperature and humidity. This optimal control strategy is then sent to the air conditioning system of the target data center. The air conditioning system adjusts the air conditioning temperature and airflow according to the optimal control strategy, thereby achieving the AI energy-saving process in the target data center.
[0105] In practical applications, after obtaining the heat load data and real-time environmental temperature and humidity data of the target computer room, it can be compared with the relevant data from the previous moment to determine whether the strategy needs to be adjusted or optimized, and then proceed with the next step if necessary.
[0106] The energy-saving method for air conditioning in a data center provided in this invention involves real-time dynamic monitoring of the heat load data of the target data center, collecting real-time environmental temperature and humidity data of the target data center, analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm, determining the energy-saving strategy for the target data center, determining the optimal control strategy for the target data center under the current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and distributing the optimal control strategy to the air conditioning system of the target data center. This method can intelligently adjust the air conditioning system of the data center according to the dynamic changes in the heat load of the data center, thereby achieving intelligent control of the data center air conditioning and maximizing energy-saving effects.
[0107] Based on the above embodiments, further, the step of analyzing the heat load data and the real-time environmental temperature and humidity data based on reinforcement learning algorithms to determine the energy-saving strategy for the target computer room includes:
[0108] The security strategy for the target data center is determined based on its maintenance and protection level and power efficiency.
[0109] The security strategy, the heat load data, and the real-time environmental temperature and humidity data are input into a strategy network determined by a reinforcement learning algorithm. The energy-saving strategy for the target computer room is determined from the output of the strategy network.
[0110] Specifically, in practical applications, in order to ensure the safety of the target data center's air conditioning during AI energy-saving control, it is also necessary to determine the target data center's security strategy.
[0111] First, based on Service Level Agreement (SLA) requirements and on-site safety regulations, security boundaries are established to determine the maintenance and protection levels for each target data center. For example, according to the target data center's maintenance procedures and safety management methods, upper and lower limits for temperature are pre-set to ensure that the data center temperature fluctuates within a safe range. Then, combined with the target data center's Power Usage Effectiveness (PUE) requirements, different operational safety strategies are constructed. For example, a safety strategy might be implemented to ensure the target data center's temperature is controlled at 25-30℃ and that the target data center's PUE is no greater than 1.5.
[0112] Subsequently, safety strategies, heat load data, and real-time environmental temperature and humidity data are input into a strategy network determined by a reinforcement learning algorithm. This strategy network is an online self-learning network that learns a neural network through reinforcement learning algorithms and a large amount of training data. It is continuously optimized based on feedback during actual use. The strategy network can output control strategies for the air supply / return air of the target computer room, proportional band settings, fan speed, compressor frequency, etc. Furthermore, the strategy network is trained to output control strategies for air conditioner on / off status, inlet / outlet air temperature and humidity, cold / hot aisle temperature and humidity, supply / return air temperature and humidity setpoints, compressor start / stop status, compressor operating frequency (variable frequency), fan start / stop status, fan operating frequency (variable frequency), voltage, current, active power, electrical energy, valve opening, outdoor temperature and humidity, etc.
[0113] Finally, the energy-saving strategy for the target data center is determined based on the output of the policy network.
[0114] This invention employs an AI-based energy-saving method for data center air conditioning based on dynamic data center load monitoring. This method optimizes and improves data center air conditioning control, and utilizes an online self-learning reinforcement strategy that combines energy saving with security. This approach reduces energy consumption and improves energy utilization while ensuring data center safety.
[0115] Based on the above embodiments, the policy network is further determined by the following method:
[0116] The air conditioning system of the target computer room is abstracted as an intelligent agent;
[0117] Based on a reinforcement learning-based model-free off-policy algorithm, the agent selects different air conditioning system control actions in different states, where each state includes safety information, heat load information, and environmental information.
[0118] Specifically, model-free learning is a model-free learning method in reinforcement learning that does not require explicit understanding of states and transition probabilities between states. Off-policy refers to the inconsistency between the method of updating the state matrix and the method of selecting the policy. This paper employs a model-free off-policy algorithm to formulate energy-saving strategies. It uses the safety information, heat load information, and environmental information of the target data center as states, and the energy-saving strategies implemented for each state as actions, and learns these to formulate energy-saving strategies. Specifically, the air conditioning system of the target data center is abstracted as an agent, enabling it to select different actions in different states to achieve energy-saving goals. Here, the state is the current state of the data center, consisting of safety information, heat load information, and environmental information. Actions include air conditioning system control actions; for example, in a certain state, the agent can choose different actions such as increasing temperature, decreasing humidity, or increasing airflow to maximize energy-saving benefits.
[0119] The embodiments of the present invention employ a reinforcement machine learning algorithm to construct a policy network model, which is used to adaptively handle the relationship between learning states and actions, thereby realizing adaptive intelligent control of the computer room air conditioning.
[0120] Based on the above embodiments, the reinforcement learning-based model-free off-policy algorithm enables the agent to select different air conditioning system control actions under different states, including:
[0121] The heat load data, ambient temperature and humidity data, and security policy data of each computer room were used as sample data.
[0122] The sample data is transformed into multiple nodes in the state space, where each node includes: a safe temperature setpoint for the computer room, a predicted power efficiency value for the computer room, and a predicted heat load value for the computer room.
[0123] Define a set of air conditioning system control actions for each node;
[0124] The Actor network is trained using the Actor-Critic method in the Off-policy algorithm, where the input of the Actor network is the nodes in the state space, and the output of the Actor network is the air conditioning system control action corresponding to the node.
[0125] The air conditioning system control actions output by the Actor network are compared with the actual air conditioning system control actions taken, and the value of the air conditioning system control actions output by the Actor network is estimated using a Critic network.
[0126] Use the value to optimize the parameters of the Actor network;
[0127] The optimized Actor network is used as the policy network.
[0128] Specifically, to train this agent, the Actor-Critic method from the Off-policy algorithm can be used. As the name suggests, Actor-Critic consists of two parts: the actor and the critic. The actor uses a policy function to generate actions and interact with the environment. The critic uses a value function to evaluate the actor's performance and guide the actor's next action.
[0129] Specifically, the following steps can be used to develop an energy-saving strategy for the target data center's air conditioning:
[0130] Step A1: Data Collection: Obtain all information for each computer room from the basic data, including indicators such as temperature, humidity, heat load, and security policies. To simplify the heat load indicators for the computer rooms, the racks can be divided into columns or areas, and the average heat load of a certain column or area can be used as the heat load value for that area. Use the obtained heat load data, ambient temperature and humidity data, and security policy data as sample data.
[0131] Step A2: Then construct the state space: Transform the collected sample data into nodes in the state space, where each node represents a set of similar operating indicators. Nodes include the data center safe temperature setting, PUE prediction value, and data center heat load prediction value.
[0132] Step A3: Define the action space: Define a set of feasible actions for the air conditioning system for each node, such as increasing (decreasing) the temperature, decreasing the humidity, increasing the air volume, etc.
[0133] Step A4: Train the Actor Network: Train a neural network using the Actor-Critic method in the Off-policy algorithm, enabling it to select the optimal action based on the current state to maximize energy savings. The input to the Actor Network is the nodes in the state space, and the output of the Actor Network is the air conditioning system control action corresponding to the node, i.e., the output is the air conditioning system energy-saving strategy.
[0134] Step A5: Validate the strategy: Use the trained neural network to generate an energy-saving strategy and apply it to the computer room air conditioning system to verify its energy-saving effect.
[0135] Specifically, training an Actor network requires a large amount of data to ensure the stability of the training results. Therefore, thorough preparation is needed in data acquisition and processing. Furthermore, during training, attention must be paid to issues such as hyperparameter tuning and loss function selection. After selecting a suitable loss function and hyperparameters, Actor network training can begin. During training, the actions output by the Actor network can be compared with the actual actions taken, and a Critic network can be used to estimate the value of each action. This value estimate can be used to optimize the Actor network's parameters, enabling it to select actions more effectively.
[0136] Once the Actor network is trained, it can be used as a policy network to generate energy-saving strategies. Specifically, the current state of the target data center can be input into the Actor network, and its output action becomes the energy-saving strategy for the target data center. In practical applications, certain adjustments are needed depending on the situation, such as considering the dynamic characteristics of the air conditioning system and changes in the data center's heat load, to make the energy-saving strategy more reasonable.
[0137] In this embodiment of the invention, a model-free off-policy algorithm is used to evaluate the target data center's safe temperature, predicted PUE, and predicted heat load as values. This enables the implementation of output control strategies for the target data center's air conditioning supply / return air, proportional band settings, fan speed, compressor frequency, etc., thereby optimizing and improving the control of the data center's air conditioning. The application of an energy-saving and safety-enhancing online self-learning strategy reduces energy consumption and improves energy utilization while ensuring the safety of the data center.
[0138] Based on the above embodiments, further, the energy-saving strategy for the target data center is determined from the output of the policy network, including:
[0139] The learning network and target network of the reinforcement learning algorithm are used to construct the target value of the operating status reward of the air conditioning system in the target computer room and the hot spot penalty data of the target computer room.
[0140] The score of the output result is determined based on the target value of the running status reward and the hot spot penalty data;
[0141] If the score of the output result is lower than the preset value, the score of the output result is fed back to the policy network so that the policy network can redetermine the new output result;
[0142] The air conditioning system control action corresponding to the highest score in the output results will be used as the energy-saving strategy for the target computer room.
[0143] Specifically, refer to Figure 2 The diagram illustrates a learning network and a target network in an embodiment of an energy-saving method for a computer room air conditioner according to the present invention, as shown below. Figure 2 As shown, the target network is used to tune the learning network, constructing target values for the operating status reward of the air conditioning system in the target computer room and hotspot penalty data for the target computer room. For example, if the output strategy is beneficial to air conditioning energy saving, a reward value is obtained; if it is detrimental to air conditioning energy saving, a penalty value is obtained. Then, the score of the output result of the strategy network for each time is determined. This score can be either a reward value or a penalty value; the reward value can be set to a positive number, and the penalty value can be set to a negative number. If the score of the output result is lower than the preset value, the score of the output result is fed back to the strategy network so that the strategy network can adjust its parameters, redetermine the new output result, and recalculate the score of the new output result. Finally, the air conditioning system control action corresponding to the maximum score in the output result is taken as the energy-saving strategy for the target computer room.
[0144] For example, the on / off rewards for air conditioners 1 and 2 are determined based on the learning network, the target values for the on / off rewards for air conditioners 1 and 2 are determined based on the target network, the parameters of the learning network are revised through the loss function, and finally the action with the maximum reward is selected for output.
[0145] In this embodiment of the invention, the energy-saving strategy output by the policy network is adjusted by learning network and target network to make the output energy-saving strategy more reasonable.
[0146] Based on the above embodiments, further, determining the optimal control strategy for the target data center in its current state based on the deep reinforcement learning algorithm and the energy-saving strategy includes:
[0147] The indoor environmental indicators of two consecutive time series corresponding to the target computer room, the feedback of the previous time series, and the air conditioning system control action of the previous time series are used as state inputs into the Q network of the deep reinforcement learning algorithm, wherein the air conditioning system control action is determined according to the energy-saving strategy.
[0148] The error gradient is constructed through backpropagation of the error using a Q-network;
[0149] The parameters of the Q-network are updated using the backpropagation algorithm.
[0150] The Q-network is used to predict and select the output action in the current state;
[0151] The epsilon-greedy algorithm is used for optimization to output the optimal control strategy.
[0152] Specifically, refer to Figure 3 The diagram illustrates an embodiment of an energy-saving method for computer room air conditioning according to the present invention, specifically a flowchart of an intelligent air conditioning control system based on the DQN algorithm. Figure 3 As shown:
[0153] In this embodiment of the invention, the intelligent air conditioning control problem is transformed into a reinforcement learning problem based on the Deep Q-network (DQN) algorithm in deep reinforcement learning. The edge system obtains energy-saving strategies from the policy network to better determine the adjusted temperature settings. The edge system optimizes the reward by updating the policy model. Based on the DQN algorithm, two consecutive time-series indoor environmental indicators, such as the computer room HVAC system and building environment information (which may include building insulation and outdoor temperature), feedback from the previous time series, and control actions from the previous time series are used as state inputs to the Q-network. An error gradient is constructed through the two Q-networks for backpropagation. The parameters of the Q-network are updated using the backpropagation algorithm, and the Q-network is used to predict and select the optimal output action in the current state. The epsilon-greedy algorithm, a commonly used algorithm in the field of deep reinforcement learning, is then used to output the final control action.
[0154] Specifically, the current status S of the computer room's HVAC system and building environment information is determined.t and set the current state S t The input is fed into the neural network to train the Q-model and obtain the current Q-value (S). t ,a t Based on the current state S t Previous timing state S t-1 and the previous timing control output a t-1 Determine the reward function r t Specifically, the return function r is determined using the following formula (1). t :
[0155] r t =cost(a t-1 S t-1 )+penalty(S t ) Formula (1)
[0156] Where cost is the value function and penalty is the penalty function.
[0157] Then, the current state S is stored in memory. t , reward function r t Previous timing state S t-1 and the previous timing control output a t-1 The previous timing state S stored in memory. t-1 Control output a, Next timing state S t+1 The reward value r is input into the neural network to infer the Q-value and predict the Q-value (S) at the maximum reward value. t+1 ,a t+1 The predicted value is summed with the input value, and combined with the output Q-value of the trained Q-model, the error gradient is obtained using the error backpropagation algorithm. This error gradient is then fed back into the training model to modify the model parameters. Simultaneously, the output Q-value of the trained Q-model is used to control the output 'a' using the epsilon-greedy algorithm. t and utilize control output a t The model continues to infer and optimize in the next time step, and so on, learning and training the model in a loop. It continuously strengthens itself in the field environment, dynamically follows the changes in the heat load of the computer room, environmental changes, and its own cooling capacity, and outputs the optimal control strategy under the current state.
[0158] This invention employs further optimized deep reinforcement learning algorithms and combined strategies to overcome the shortcomings of model-free deep reinforcement learning, significantly accelerating the convergence of precise energy-saving control strategies in the field. By dynamically monitoring the data center load, the system monitors the data center's load status and uses an edge system to adjust the air conditioning operation in a timely manner. It utilizes an online reinforcement strategy combining deep reinforcement learning and secure reinforcement learning to achieve AI-driven energy saving in the data center's air conditioning, reducing energy consumption, lowering data center management costs, and ensuring stable temperature within the data center. This not only achieves energy conservation and consumption reduction in the data center's air conditioning but also significantly improves the data center's energy utilization efficiency.
[0159] Based on the above embodiments, the Q-network is further determined in the following manner:
[0160] Determine the various states of the target computer room, including the ambient temperature data, humidity data, and heat load data of the target computer room;
[0161] Determine the control actions of the air conditioning system under various conditions;
[0162] A deep neural network model Q-network is constructed based on the state and the control actions of the air conditioning system.
[0163] The deep neural network model Q-network is trained, and the optimal control policy is learned based on the reward value.
[0164] Specifically, the Q-network is determined using the following steps:
[0165] Step B1: Determine the status of the target data center: The variables for managing the status include the target data center's ambient temperature, humidity, heat load data, etc.
[0166] Step B2: Determine the air conditioning system control actions for each state. For example, if the control action is to adjust the temperature setting, the temperature setting can be controlled within a safe range, and the action can be represented as the deviation from the current temperature value.
[0167] Step B3: Set a reward signal: Find a suitable reward signal to reward the edge system for making correct decisions. In this embodiment of the invention, the goal is to maximize comfort and minimize energy consumption. Therefore, the reward value can be set according to formula (2):
[0168] reward=-(1 – comfort)+energy formula (2)
[0169] Here, comfort refers to the level of comfort, energy is the energy consumption of the air conditioning system, and reward is the reward value.
[0170] Step B4: Constructing a Neural Network Model (Q-Network): This step constructs a neural network model to learn the relationship between states and actions. A deep neural network model is used because the air conditioning control problem has a large state space, which traditional Q-learning algorithms cannot handle. Deep neural networks can adaptively learn the relationship between states and actions, and relying on powerful approximation techniques, they can approximate any function that can be approximated using deep neural networks.
[0171] Step B5, Training the Model: Train the model so that it can learn the optimal policy based on reward signals. During training, the current state is used as input, and then an action is taken according to the selected policy. Reward signals are received, and these signals are used to update the neural network. By iteratively revising the neural network weights, the optimal policy that achieves maximum comfort and minimum energy consumption can be learned.
[0172] Step B6, Test the model: After installing the trained model, it is also necessary to test its effectiveness in practical applications.
[0173] In the embodiments of this invention, the air conditioning system mainly refers to a precision air conditioning unit with a built-in compressor. The main power-consuming components are the fan and the compressor, and it does not include a chilled water valve. The control parameters include: air conditioner start / stop, fan start / stop (fixed frequency, variable frequency), inlet / outlet air temperature setting, fan minimum speed setting (variable frequency), fan rated speed setting (variable frequency), fan operating speed setting (variable frequency), compressor start / stop (fixed frequency, variable frequency), compressor minimum load setting (variable frequency), compressor maximum load setting (variable frequency), compressor speed setting (variable frequency), etc.
[0174] In this embodiment of the invention, the above data is collected to evaluate the performance of the model and used to predict future temperature settings.
[0175] The energy-saving method for data center air conditioning provided in this invention uses a combined algorithm for multi-level energy-saving control. First, a safety boundary is established based on the operation and maintenance SLA requirements and on-site safety specifications. Then, reinforcement learning algorithms and deep neural networks are used to explore energy-saving strategies under different environments, and the strategy boundary is established in one step. Finally, deep reinforcement learning + security reinforcement learning are used to strengthen the strategy online, ultimately achieving rapid energy saving, safe and controllable operation, and automatic adaptation to changes. As time goes by, the optimization boundary and energy-saving strategy are further strengthened, achieving greater energy savings in the data center.
[0176] Reference Figure 4 The flowchart illustrates another embodiment of the energy-saving method for a computer room air conditioner according to the present invention, which may specifically include:
[0177] The SLA safety boundary is determined using optimization / search algorithms, SLA rules, and knowledge rules. Hotspot and PUE prediction are performed using Long Short-Term Memory (LSTM) / gated recurrent unit (GRU) networks, followed by domain randomized reinforcement learning (RL) to determine the policy boundary. Safety layer data is then determined using the SLA and policy boundaries, and safety reinforcement learning is performed on the penalty function. This data is input into the V network, which evaluates the value of each state. The V network outputs the state values, and the policy network incorporates Soft Actor-Critic (SAC) to output the policy and sampling operations. Q-network reinforcement learning is then used to output the state operation values, thus obtaining the optimal control policy for the target air conditioner.
[0178] The energy-saving method for air conditioning in a data center provided in this invention is based on a model-free off-policy algorithm. It uses the target data center's safe temperature, predicted PUE value, and predicted heat load value as value assessments to determine the energy-saving strategy for the target data center. It implements intelligent control of the air conditioning based on the DQN algorithm. By constructing a neural network model, it adaptively processes the relationship between the learning state and the action to achieve adaptive intelligent control of the data center air conditioning.
[0179] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0180] Reference Figure 5 The diagram shows a structural block diagram of an embodiment of an energy-saving device for a computer room air conditioner according to the present invention, which may specifically include the following modules:
[0181] The cruise monitoring module 510 is used for real-time dynamic cruise monitoring of the heat load data of the target computer room;
[0182] Collection module 520 is used to collect real-time environmental temperature and humidity data of the target computer room;
[0183] Analysis module 530 is used to analyze the heat load data and the real-time environmental temperature and humidity data based on reinforcement learning algorithm to determine the energy-saving strategy for the target computer room;
[0184] The execution module 540 is used to determine the optimal control strategy for the target computer room in the current state based on the deep reinforcement learning algorithm and the energy-saving strategy, and to send the optimal control strategy to the air conditioning system of the target computer room.
[0185] Specifically, the cruise monitoring module 510 may include a computer room equipment heat load acquisition unit, a computer room personnel heat load acquisition unit, an environmental heat load acquisition unit, and a strategy confirmation unit.
[0186] The equipment heat load acquisition unit is used to acquire the voltage and current of the equipment in the target equipment room in real time to calculate the heat load of the equipment. The personnel heat load acquisition unit is used to detect the entry and exit of personnel in the target equipment room in real time to calculate the personnel heat load. The environmental heat load acquisition unit is used to detect the building heat load data of the target equipment room in real time. The strategy confirmation unit is used to summarize and record the heat load changes in the target equipment room in real time, compare the heat load changes with the energy-saving strategy, and provide feedback through the correlation unit on whether the strategy needs to be adjusted or optimized.
[0187] The collection module 520 includes an ambient temperature and humidity sensor, which is used to collect real-time data of the ambient temperature and humidity of the target computer room, and report the changes in the ambient temperature and humidity of the target computer room to the analysis module 530 in real time according to the changes in the energy-saving strategy.
[0188] The analysis module 530 includes a basic data reading unit, a security boundary construction unit, an AI energy-saving learning unit, and an energy-saving strategy delivery unit.
[0189] The basic data reading unit is used by the cruise monitoring module 510 to obtain real-time heat load data of the target computer room, and to obtain real-time temperature and humidity data of the target computer room from the collection module 520. The data is then sent to the AI energy-saving learning unit to build a learning network.
[0190] The security boundary construction unit is used to build different operational security strategies based on the differences in maintenance and protection levels of different target data centers and the PUE operational requirements. For example, the data center temperature is controlled at 25-30℃, and the data center's operational PUE is no greater than 1.5.
[0191] The AI energy-saving learning unit is used to input security policies, heat load data, and real-time environmental temperature and humidity data into a policy network determined based on a reinforcement learning algorithm, and to determine the energy-saving strategy for the target computer room from the output of the policy network.
[0192] Specifically, the AI energy-saving learning unit includes:
[0193] Abstract subunit, used to abstract the air conditioning system of the target computer room into an intelligent agent;
[0194] The training subunit is used to enable the agent to select different air conditioning system control actions in different states based on the reinforcement learning model-free off-policy algorithm, where each state includes safety information, heat load information and environmental information.
[0195] The training subunit is specifically used for:
[0196] All information about the computer room is retrieved from the basic data, including indicators such as temperature, humidity, air volume, and heat load. To simplify the heat load indicators, the computer room racks can be divided into columns or areas, and the average heat load of a certain column or area can be used as the heat load value for that area. The heat load data, ambient temperature and humidity data, and security policy data of each computer room are used as sample data.
[0197] The sample data is transformed into multiple nodes in the state space, where each node includes: a safe temperature setpoint for the computer room, a predicted power efficiency value for the computer room, and a predicted heat load value for the computer room.
[0198] Define a set of air conditioning system control actions for each node, such as increasing (decreasing) temperature, decreasing humidity, increasing air volume, etc.
[0199] The Actor network is trained using the Actor-Critic method in the Off-policy algorithm, where the input of the Actor network is the nodes in the state space, and the output of the Actor network is the air conditioning system control action corresponding to the node.
[0200] The strategy was validated by using a trained Actor network to generate the strategy and applying it to the computer room air conditioning system to verify its energy-saving effect. The air conditioning system control actions output by the Actor network were compared with the actual air conditioning system control actions taken, and the value of the air conditioning system control actions output by the Actor network was estimated using a Critic network.
[0201] Use the value to optimize the parameters of the Actor network;
[0202] The optimized Actor network is used as the policy network.
[0203] In practical applications, the analysis module 530 further includes an energy-saving strategy issuing unit, which is specifically used for:
[0204] The learning network and target network of the reinforcement learning algorithm are used to construct the target value of the operating status reward of the air conditioning system in the target computer room and the hot spot penalty data of the target computer room.
[0205] The score of the output result is determined based on the target value of the running status reward and the hot spot penalty data;
[0206] If the score of the output result is lower than the preset value, the score of the output result is fed back to the policy network so that the policy network can redetermine the new output result;
[0207] The air conditioning system control action corresponding to the highest score in the output results will be used as the energy-saving strategy for the target computer room.
[0208] In practical applications, execution module 540 is specifically used for:
[0209] The indoor environmental indicators of two consecutive time series corresponding to the target computer room, the feedback of the previous time series, and the air conditioning system control action of the previous time series are used as state inputs into the Q network of the deep reinforcement learning algorithm, wherein the air conditioning system control action is determined according to the energy-saving strategy.
[0210] The error gradient is constructed through backpropagation of the error using a Q-network;
[0211] The parameters of the Q-network are updated using the backpropagation algorithm.
[0212] The Q-network is used to predict and select the output action in the current state;
[0213] The epsilon-greedy algorithm is used for optimization to output the optimal control strategy.
[0214] In practical applications, the execution module 540 is used to determine the Q-network in the following manner:
[0215] Determine the various states of the target computer room, including the ambient temperature data, humidity data, and heat load data of the target computer room;
[0216] Determine the control actions of the air conditioning system under various conditions;
[0217] A deep neural network model Q-network is constructed based on the state and the control actions of the air conditioning system.
[0218] The deep neural network model Q network is trained, and the optimal policy is learned based on the reward value. The execution module 540 determines the reward value according to the formula reward = -1 – comfort + energy, where comfort is the level of comfort, energy is the energy consumption of the air conditioning system, and reward is the reward value.
[0219] The energy-saving device for a data center air conditioner provided in this invention provides real-time dynamic monitoring of the heat load data of the target data center, collects real-time environmental temperature and humidity data of the target data center, analyzes the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm, determines the energy-saving strategy for the target data center, determines the optimal control strategy for the target data center under the current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and sends the optimal control strategy to the air conditioning system of the target data center. It can intelligently adjust the data center air conditioning system according to the dynamic changes in the data center heat load, thereby achieving intelligent control of the data center air conditioning and maximizing energy-saving effects.
[0220] As the device embodiment is basically similar to the method embodiment, the description is relatively simple. For relevant details, please refer to the description of the method embodiment. It will not be repeated here.
[0221] Reference Figure 6 The diagram shows a structural block diagram of an embodiment of an electronic device according to the present invention. The device includes: a processor 610, a memory 620, and a bus 630.
[0222] The processor 610 and the memory 620 communicate with each other through the bus 630.
[0223] The processor 610 is used to call program instructions in the memory 620 to execute the methods provided in the above-described method embodiments, such as: real-time dynamic cruise monitoring of the heat load data of the target computer room; collecting real-time environmental temperature and humidity data of the target computer room; analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm to determine the energy-saving strategy of the target computer room; determining the optimal control strategy of the target computer room in the current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and distributing the optimal control strategy to the air conditioning system of the target computer room.
[0224] This invention discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can perform the methods provided in the above-described method embodiments, such as: real-time dynamic cruise monitoring of heat load data of a target computer room; collecting real-time environmental temperature and humidity data of the target computer room; analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm to determine an energy-saving strategy for the target computer room; determining the optimal control strategy for the target computer room in the current state based on a deep reinforcement learning algorithm and the energy-saving strategy, and distributing the optimal control strategy to the air conditioning system of the target computer room.
[0225] This invention provides a non-transitory computer-readable storage medium storing computer instructions that cause a computer to execute the methods provided in the above-described method embodiments. These instructions include, for example: real-time dynamic cruise monitoring of heat load data in a target computer room; collecting real-time environmental temperature and humidity data of the target computer room; analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm to determine an energy-saving strategy for the target computer room; determining the optimal control strategy for the target computer room in its current state based on a deep reinforcement learning algorithm and the energy-saving strategy; and distributing the optimal control strategy to the air conditioning system of the target computer room.
[0226] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0227] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0228] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0229] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0230] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0231] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0232] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0233] The present invention has provided a detailed description of an energy-saving method and an energy-saving device for a computer room air conditioner. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. An energy-saving method for computer room air conditioning, characterized in that, include: Real-time dynamic cruise monitoring of the target computer room's heat load data; Collect real-time environmental temperature and humidity data of the target computer room; Based on reinforcement learning algorithms, the heat load data and the real-time environmental temperature and humidity data are analyzed to determine the energy-saving strategy for the target computer room. Based on the deep reinforcement learning algorithm and the energy-saving strategy, the optimal control strategy for the target computer room in the current state is determined, and the optimal control strategy is sent to the air conditioning system of the target computer room. The step of analyzing the heat load data and the real-time environmental temperature and humidity data based on a reinforcement learning algorithm to determine the energy-saving strategy for the target computer room includes: The security strategy for the target data center is determined based on its maintenance and protection level and power efficiency. The security strategy, the heat load data, and the real-time environmental temperature and humidity data are input into a strategy network determined based on a reinforcement learning algorithm. The energy-saving strategy for the target computer room is determined from the output of the strategy network. The policy network is determined by the following method: The air conditioning system of the target computer room is abstracted as an intelligent agent; Based on a reinforcement learning-based model-free off-policy algorithm, the agent selects different air conditioning system control actions in different states, where each state includes safety information, heat load information, and environmental information.
2. The method according to claim 1, characterized in that, The reinforcement learning-based model-free off-policy algorithm enables the agent to select different air conditioning system control actions under different states, including: The heat load data, ambient temperature and humidity data, and security policy data of each computer room were used as sample data. The sample data is transformed into multiple nodes in the state space, where each node includes: a safe temperature setpoint for the computer room, a predicted power efficiency value for the computer room, and a predicted heat load value for the computer room. Define a set of air conditioning system control actions for each node; The Actor network is trained using the Actor-Critic method in the Off-policy algorithm, where the input of the Actor network is the nodes in the state space, and the output of the Actor network is the air conditioning system control action corresponding to the node. The air conditioning system control actions output by the Actor network are compared with the actual air conditioning system control actions taken, and the value of the air conditioning system control actions output by the Actor network is estimated using a Critic network. Use the value to optimize the parameters of the Actor network; The optimized Actor network is used as the policy network.
3. The method according to claim 2, characterized in that, Determining the energy-saving strategy for the target data center from the output of the policy network includes: The learning network and target network of the reinforcement learning algorithm are used to construct the target value of the operating status reward of the air conditioning system in the target computer room and the hot spot penalty data of the target computer room. The score of the output result is determined based on the target value of the running status reward and the hot spot penalty data; If the score of the output result is lower than the preset value, the score of the output result is fed back to the policy network so that the policy network can redetermine the new output result; The air conditioning system control action corresponding to the highest score in the output results will be used as the energy-saving strategy for the target computer room.
4. The method according to claim 3, characterized in that, The process of determining the optimal control strategy for the target data center in its current state, based on the deep reinforcement learning algorithm and the energy-saving strategy, includes: The indoor environmental indicators of two consecutive time series corresponding to the target computer room, the feedback of the previous time series, and the air conditioning system control action of the previous time series are used as state inputs into the Q network of the deep reinforcement learning algorithm, wherein the air conditioning system control action is determined according to the energy-saving strategy. The error gradient is constructed through backpropagation of the error using a Q-network; The parameters of the Q-network are updated using the backpropagation algorithm. The Q-network is used to predict and select the output action in the current state; The epsilon-greedy algorithm is used for optimization to output the optimal control strategy.
5. The method according to claim 4, characterized in that, The Q network is determined in the following manner: Determine the various states of the target computer room, including the ambient temperature data, humidity data, and heat load data of the target computer room; Determine the control actions of the air conditioning system under various conditions; A deep neural network model Q-network is constructed based on the state and the control actions of the air conditioning system. The deep neural network model Q-network is trained, and the optimal control policy is learned based on the reward value.
6. The method according to claim 5, characterized in that, The reward value is determined according to the following formula: reward=-(1-comfort)+energy Here, comfort refers to the level of comfort, energy is the energy consumption of the air conditioning system, and reward is the reward value.
7. An energy-saving device for a computer room air conditioner, characterized in that, include: The cruise monitoring module is used for real-time dynamic cruise monitoring of the target computer room's heat load data; The collection module is used to collect real-time environmental temperature and humidity data of the target computer room; The analysis module is used to analyze the heat load data and the real-time environmental temperature and humidity data based on reinforcement learning algorithms to determine the energy-saving strategy for the target computer room. The execution module is used to determine the optimal control strategy for the target computer room in the current state based on the deep reinforcement learning algorithm and the energy-saving strategy, and to send the optimal control strategy to the air conditioning system of the target computer room. The analysis module includes: The security boundary construction unit is used to determine the security strategy of the target computer room based on the maintenance and protection level of the target computer room and the power usage efficiency of the target computer room. The AI energy-saving learning unit is used to input the security policy, the heat load data, and the real-time environmental temperature and humidity data into a policy network determined based on a reinforcement learning algorithm, and to determine the energy-saving policy for the target computer room from the output of the policy network. The AI energy-saving learning unit also includes: Abstract subunit, used to abstract the air conditioning system of the target computer room into an intelligent agent; Training subunits for reinforcement learning-based models free Off The policy algorithm enables the agent to select different air conditioning system control actions under different states, where each state includes safety information, heat load information, and environmental information.
8. An electronic device, characterized in that, include: The memory and the processor communicate with each other via a bus; The memory stores program instructions that can be executed by the processor, and the processor can execute the method as described in any one of claims 1 to 6 by calling the program instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for data center machine room control based on reinforcement learning algorithm
CN111126605A