Central air conditioner energy-saving control system based on intelligent algorithm
By constructing a dynamic graph structure and using a time graph network model for prediction, combined with stratified reinforcement learning decision-making, the shortcomings of traditional neural networks in dealing with the problem of explosive growth of combinations are solved, intelligent energy-saving control of the central air-conditioning system is realized, and the system's adaptability and energy-saving effect are improved.
Patent Information
- Application Number
- CN202510541392.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional neural network structures perform poorly when dealing with the problem of explosive growth of combinations, especially when coordinating multiple devices and dealing with differentiated needs, they cannot fully understand the overall synergy of the system, and lack the ability to deal with dynamic changing environments and multiple differentiated needs, so they cannot achieve true intelligent energy saving.
The dynamic graph structure is constructed through the graph building block, and the time graph network model is used to predict the predicted key states of the next moment, and the option-criticist architecture is used to perform hierarchical reinforcement learning, output the target operation frequency combination, and finally convert it into execution instructions.
It realizes accurate state recognition and high-precision state prediction of the central air-conditioning system, significantly reduces decision-making complexity, improves the system's adaptability and energy-saving effects, and can achieve energy consumption optimization while meeting various temperature control needs.
Smart Images

Figure CN120084036A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of energy conservation, and particularly to a central air-conditioning energy conservation control system based on intelligent algorithms. Background Art
[0002] Bio-inspired computational models such as deep learning and reinforcement learning have become cutting-edge methods for solving complex decision-making problems. For example, graph neural networks and long short-term memory networks and gated recurrent units based on memory mechanisms can effectively process structured data and temporal information, thereby learning the internal patterns in complex systems. At the same time, deep reinforcement learning continuously optimizes decision-making strategies through the interaction between the agent and the environment, providing innovative solutions for multi-objective optimization problems. These intelligent algorithms have shown advantages in multiple fields, especially in scenarios that require processing high-dimensional state spaces and large-scale decision-making combinations, enabling a more refined decision-making process.
[0003] However, traditional neural network structures perform poorly in dealing with the problem of combinatorial explosive growth. Especially when the system needs to coordinate multiple devices simultaneously and handle differentiated requirements, the complex topological relationships and mutual influences between devices are difficult to express through conventional neural network structures, resulting in the model being unable to fully understand the overall coordination relationship of the system. At the same time, existing artificial intelligence models lack the ability to handle dynamic changing environments and the coexistence of multiple differentiated requirements, and cannot achieve true intelligent energy conservation.
[0004] Therefore, a central air-conditioning energy conservation control system based on intelligent algorithms is proposed. Summary of the Invention
[0005] The purpose of the present invention is to provide a central air-conditioning energy conservation control system based on intelligent algorithms. By means of a graph construction module, a graph network prediction module, a hierarchical decision-making module, and an instruction execution module, the problem of explosive growth of policy combinations under differentiated requirements is solved, and energy consumption optimization is achieved. First, the topological structure and operating state data of the devices are obtained to construct a dynamic graph structure, where each device serves as a node in the dynamic graph structure, the connection relationship between nodes serves as an edge in the dynamic graph structure, and the operating state data serves as the node feature vector. Then, the dynamic graph structure and candidate frequency combinations at the current moment are received, and a time graph network model is used to predict the predicted key state at the next moment. The option-critic architecture is adopted to perform hierarchical reinforcement learning, and the target operating frequency combination is output according to the representation of the dynamic graph structure at the current moment and the predicted key state at the next moment. Finally, the target operating frequency combination is converted into an execution instruction.
[0006] To achieve the above object, the present invention provides the following technical solutions: A central air-conditioning energy conservation control system based on intelligent algorithms, comprising: The graph construction module obtains the topological structure and operating status data of the device, constructs a dynamic graph structure, takes each device and environmental status as nodes in the dynamic graph structure, the connection relationship between nodes as edges in the dynamic graph structure, and the operating status data as the node feature vector; The graph network prediction module receives the dynamic graph structure and candidate frequency combinations at the current moment, and uses the temporal graph network model to predict the predicted key status at the next moment; The hierarchical decision-making module performs hierarchical reinforcement learning using the option-critic architecture, and outputs the target operating frequency combination according to the dynamic graph structure representation at the current moment and the predicted key status at the next moment; the hierarchical reinforcement learning includes a high-level policy agent and a low-level action agent; The instruction execution module receives the target operating frequency combination and converts it into an execution instruction.
[0007] Preferably, the process of taking the operating status data as the node feature vector includes: collecting the operating status data corresponding to each node; obtaining the static attribute parameters corresponding to the node, where the static attribute parameters refer to the inherent characteristics of the device itself and do not change with the operating status; performing normalization processing on the operating status data and the static attribute parameters; combining the normalized operating status data and the static attribute parameters in the order predefined by the node type to form the node feature vector.
[0008] Preferably, the temporal graph network model specifically includes: The memory layer associates a memory state vector with each node in the dynamic graph structure and updates the memory state vector using a long short-term memory network; The message function calculates the message according to the memory state vector of the interacting node at the previous moment, the node feature vector at the current moment, the edge feature, and the time difference when the nodes interact; The message aggregator aggregates all the messages received by the node at the current moment for updating the memory state vector; The node embedding generation layer combines the memory state vector of the node at the current moment and the node feature vector to generate a node embedding representation; The prediction head receives the node embedding representation and the candidate frequency combination, and outputs the predicted key status at the next moment; the predicted key status includes the predicted value of the outlet water temperature of the cooling circuit and the predicted value of the total power consumption on the cooling side.
[0009] Preferably, the option-critic architecture specifically includes: The shared feature extraction network receives the current state representation as input and outputs the shared feature vector at the current moment; The policy network includes a high-level option policy head and a low-level action policy head. The high-level option policy head outputs the probability distribution of each option based on the shared feature vector, and the low-level action policy head outputs the probability distribution of the candidate frequency combinations for each option, and samples to obtain the target operating frequency combination. The critic network receives the current state representation and outputs the long-term cumulative reward evaluation value of selecting and executing each option in the current state, which is used to guide the parameter update of the policy network and the termination network during the training process of hierarchical reinforcement learning. The termination network receives the shared feature vector at the next moment and outputs the probability of termination of the currently executed option.
[0010] Preferably, the current state representation is composed of the dynamic graph structure at the current moment and the predicted key state at the next moment; the target operating frequency combination is composed of the target operating frequency values of all controlled cooling towers.
[0011] Preferably, the policy network learns to obtain a high-level policy agent and a low-level action agent through the training process of hierarchical reinforcement learning. During the training process of hierarchical reinforcement learning, the calculation method of the reward value includes: Calculating a first reward component according to the predicted value of the total power consumption on the cooling side at the next moment output by the graph network prediction module; Calculating a second reward component according to the deviation between the predicted value of the outlet water temperature of the corresponding cooling loop at the next moment output by the graph network prediction module and the loop target set temperature; Weightedly combining the first reward component and the second reward component according to a preset weight coefficient to obtain the reward value.
[0012] Preferably, the process of converting the target operating frequency combination into an execution instruction includes: extracting the target operating frequency value in the target operating frequency combination; performing anti-normalization processing on each non-zero target operating frequency value and converting it into a specific analog signal according to the frequency converter communication protocol of the corresponding cooling tower; parsing the target operating frequency value of zero into a shutdown instruction for the corresponding cooling tower and converting it into a corresponding digital signal; before generating the final execution instruction, performing a safety logic check, and the safety logic check includes whether the target frequency is within the allowable range, whether the frequency change slope exceeds the limit, and whether the start and stop of the equipment meet the preset interlock protection conditions; sending the instruction passing the safety logic check to the on-site execution mechanism.
[0013] Compared with the prior art, the beneficial effects of the present invention are: 1. Abstract the central air conditioning system as a dynamic graph structure, where each device serves as a node in the graph, the connection relationship between devices serves as an edge, and the operating state data serves as the node feature vector. This structured representation method can accurately capture the complex physical connections and interaction relationships within the system. Through the graph construction module, the system can obtain the device topology structure and operating state data in real time, construct a dynamic graph reflecting the current system state, enabling the entire control system to have a more comprehensive and accurate understanding of the air conditioning network. In the graph network prediction module, a temporal graph network model is used to process the dynamic graph structure, which can learn the complex interaction patterns between devices and their evolution laws over time. This enables the system to accurately predict the changes in future states under given control actions. The graph-based dynamic modeling method can capture the complex interaction effects between different devices more precisely, providing the system with accurate state representation and future prediction capabilities, and laying a foundation for solving the technical problem of the explosive growth of device combinations under different demand scenarios.
[0014] 2. The hierarchical reinforcement learning decision-making effectively solves the problem of combinatorial explosion through the division of labor and cooperation between the high-level policy agent and the low-level action agent. The high-level policy agent is responsible for macro decision-making and selects the most suitable option according to the overall system state; the low-level action agent then performs fine-grained control for specific options and directly outputs the target operating frequency combination of the cooling tower. The hierarchical decision-making mechanism decomposes the high-dimensional frequency combination space into subspaces that are easier to learn, significantly reducing the learning difficulty and decision-making complexity. The system continuously optimizes the policy by interacting with the environment and gradually learns the optimal control strategy for different working conditions. This hierarchical learning method can automatically discover and utilize the hierarchical structure in the problem, has stronger generalization ability for complex scenarios, and the control decision-making is more flexible and efficient, solving the core technical problem of the explosive growth of the control decision-making space under different demand scenarios, enabling the system to achieve energy consumption optimization while meeting various temperature control requirements.
[0015] 3. Deeply integrate the temporal graph network model with the hierarchical reinforcement learning decision-making system to form a closed-loop optimization mechanism. The temporal graph network model can not only capture the static topology structure of the air conditioning system but also handle the temporal evolution of node states and the dynamic interactions between nodes. The system maintains a memory state vector for each node, which is updated through a long short-term memory network, enabling the model to learn the time dependence of device states and output the predicted value of the key state at the next moment, providing forward-looking guidance for decision-making. This prediction ability enables the hierarchical reinforcement learning to evaluate the effects of different candidate actions before actual control and select the optimal control strategy. The calculation of the system reward value is directly based on the prediction results of the graph network, including the predicted energy consumption value and the predicted temperature deviation value, forming a tight coupling between prediction and decision-making, improving the learning efficiency and control safety, realizing the collaborative optimization of prediction and decision-making, enabling the system to cope with complex control scenarios under different demand scenarios with the minimum trial-and-error cost, and enhancing the energy-saving control effect. Brief Description of the Drawings
[0016] Figure 1 FIG. 1 is a schematic structural diagram of a central air-conditioning energy-saving control system based on an intelligent algorithm according to the present invention; Figure 2 FIG. 2 is a schematic structural diagram of a temporal graph network model according to the present invention; Figure 3 FIG. 3 is a schematic structural diagram of an option-critic architecture according to the present invention. Detailed Embodiments
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0018] Please refer to Figures 1 to 3 , the present invention provides a central air-conditioning energy-saving control system based on an intelligent algorithm, and the technical solutions are as follows: Embodiment 1: In this embodiment, a commercial complex building is used as an application scenario. The building includes a data center (requiring the cooling water outlet temperature ≤ 12°C), an office area (comfort air-conditioning, requiring the cooling water outlet temperature to be 18 ± 1°C), and a shopping mall (requiring the cooling water outlet temperature to be 16 ± 2°C), and a total of 6 cooling towers, 4 chiller units, and a distributed water pump system are configured. Traditional control methods are difficult to coordinate multi-objective temperature control requirements and energy consumption optimization, resulting in frequent start-stop of some equipment and large energy consumption fluctuations. The present invention provides a central air-conditioning energy-saving control system based on an intelligent algorithm through the deep integration of hierarchical reinforcement learning and graph neural networks. The system structure is shown in Figure 1 , including: A graph construction module, which acquires the topological structure and operating state data of the equipment, constructs a dynamic graph structure, takes each equipment and environmental state as nodes in the dynamic graph structure, the connection relationship between nodes as edges in the dynamic graph structure, and the operating state data as node feature vectors; A graph network prediction module, which receives the dynamic graph structure and candidate frequency combinations at the current moment, and uses the temporal graph network model to predict the predicted key state at the next moment; A hierarchical decision-making module, which performs hierarchical reinforcement learning using an option-critic architecture, and outputs a target operating frequency combination according to the dynamic graph structure representation at the current moment and the predicted key state at the next moment; the hierarchical reinforcement learning includes a high-level policy agent and a low-level action agent; An instruction execution module, which receives the target operating frequency combination and converts it into an execution instruction.
[0019] Furthermore, the process of using the operating status data as the node feature vector includes: collecting the operating status data corresponding to each node; obtaining the static attribute parameters corresponding to the node, where the static attribute parameters refer to the inherent characteristics of the device itself and do not change with the operating status; performing normalization processing on the operating status data and the static attribute parameters; and combining the normalized operating status data and the static attribute parameters in the order predefined according to the node type to form the node feature vector.
[0020] Table 1 shows the composition of the feature vectors of the key nodes in the dynamic graph construction, covering the operating status data and the static attribute parameters; Table 2 shows the definitions and feature compositions of various types of edges. Table 1 Example of Node Feature Vector Node type Example of operating status data Example of static attribute parameters Cooling tower Start / stop status, frequency, power, inlet / outlet water temperature Rated power, maximum flow rate Chiller Load ratio, chilled water outlet temperature, coefficient of performance Refrigerating capacity, design pressure difference Environmental node Wet bulb temperature, dry bulb temperature, relative humidity None Table 2 Definitions of Main Edge Types Edge type Source node type Destination node type Edge characteristics Cooling water supply Cooling water pump Chiller Pipe diameter, pipe length, flow rate, temperature Cooling water return Chiller Cooling tower Pipe diameter, pipe length, flow rate, temperature Cooling water reflux Cooling tower Cooling water pump Pipe diameter, pipe length, flow rate, temperature Chilled water supply Chilled water pump Terminal equipment Pipe diameter, pipe length, flow rate, temperature Chilled water return Terminal equipment Chiller Pipe diameter, pipe length, flow rate, temperature Environmental impact Environmental node Cooling tower Wet bulb temperature, wind speed, relative humidity Pipe network connection Pipeline node Pipeline node Flow direction, flow rate, pressure drop By combining the operating status data and the static attribute parameters and forming the node feature vector in the predefined order, the system can comprehensively and accurately represent the characteristics and current status of each device. The normalization processing ensures the comparability of data with different dimensions, improves the training stability and prediction accuracy of the graph network model, provides a more accurate state representation for subsequent reinforcement learning decisions, and thus enhances the overall control effect.
[0021] Furthermore, referring to Figure 2 , the time graph network model specifically includes: A memory layer that associates a memory state vector with each node in the dynamic graph structure and updates the memory state vector using a long short-term memory network; A message function that calculates messages based on the memory state vector of the interacting node at the previous moment, the node feature vector at the current moment, the edge feature, and the time difference when the nodes interact; A message aggregator that aggregates all the messages received by the node at the current moment and is used to update the memory state vector; A node embedding generation layer that combines the memory state vector of the node at the current moment with the node feature vector to generate a node embedding representation; A prediction head that receives the node embedding representation and the candidate frequency combination and outputs the predicted key state at the next moment; the predicted key state includes the predicted value of the outlet water temperature of the cooling loop and the predicted value of the total power consumption on the cooling side.
[0022] The time graph network model realizes the accurate modeling of the dynamic characteristics of the system through components such as a memory layer, a message function, a message aggregator, and a node embedding generation layer. The model considers key factors such as node - to - node interactions, time differences, and edge features, can accurately predict the key state at the next moment, provides a reliable simulation environment for reinforcement learning, reduces the trial - and - error cost in actual control, and significantly improves the adaptive ability of the system.
[0023] Further, referring to Figure 3 , the option - critic architecture specifically includes: A shared feature extraction network that receives the current state representation as input and outputs the shared feature vector at the current moment; A policy network, including a high - level option policy head and a low - level action policy head. The high - level option policy head outputs the probability distribution of each option based on the shared feature vector, and the low - level action policy head outputs the probability distribution of candidate frequency combinations for each option, and samples to obtain the target operating frequency combination; A critic network that receives the current state representation and outputs the long - term cumulative reward evaluation value of selecting and executing each option in the current state, which is used to guide the parameter update of the policy network and the termination network during the training process of hierarchical reinforcement learning; A termination network that receives the shared feature vector at the next moment and outputs the probability of termination of the currently executed option.
[0024] The option - critic architecture realizes a complete hierarchical decision - making mechanism through a shared feature extraction network, a hierarchical policy network, a critic network, and a termination network. This architecture can automatically discover and utilize the hierarchical structure in the control problem, significantly reducing the decision - making complexity.
[0025] Further, the current state representation is composed of the dynamic graph structure at the current moment and the predicted key state at the next moment; the target operating frequency combination is composed of the target operating frequency values of all controlled cooling towers.
[0026] Taking the current graph structure and the predicted key state together as the state representation, and combining the target frequency combinations of all controlled devices as the action output, realizes the precise definition of the state representation and the control action, enables the system to have a global view and fine - control ability, can make optimal control decisions based on the current state and future predictions, and effectively cope with complex control scenarios under different demands.
[0027] Further, the policy network obtains a high - level policy agent and a low - level action agent through the training process of hierarchical reinforcement learning. In the training process of hierarchical reinforcement learning, the calculation method of the reward value includes: Calculating the first reward component according to the predicted value of the total cooling - side power consumption at the next moment output by the graph network prediction module; Calculate a second reward component based on the deviation between the predicted value of the outlet water temperature of the corresponding cooling loop at the next moment output by the graph network prediction module and the target set temperature of the loop; Weight and combine the first reward component and the second reward component according to a preset weight coefficient to obtain the reward value.
[0028] Table 3 shows the comparison of the daily average energy consumption between the energy-saving control based on hierarchical reinforcement learning and the traditional energy-saving control of the present invention under the hybrid load condition. The hybrid load operation mode means that there are high, medium, and low load demands simultaneously in different areas of the building. For example, the data center is fully loaded, the office area has a partial load (50%), and the shopping mall has a medium load (70%). It is necessary to coordinate the operation strategies of multiple devices to meet the differentiated temperature control requirements; the energy-saving control based on hierarchical reinforcement learning can effectively reduce energy consumption by decomposing the goal through the high-level strategy and coordinating the device frequencies at the low level.
[0029] Table 3 System energy-saving control effect Control method Operating status Daily average energy consumption Traditional PID control Hybrid load 2855 kWh The present invention Hybrid load 2183 kWh The reward calculation mechanism takes into account two key objectives of energy consumption and temperature control, and achieves multi-objective balance through weighted combination. It enables the system to find the best balance point between energy saving and meeting temperature control requirements, and avoids the problems that may be brought by single-objective optimization.
[0030] Further, the process of converting the target operating frequency combination into an execution instruction includes: extracting the target operating frequency values in the target operating frequency combination; performing anti-normalization processing on each non-zero target operating frequency value and converting it into a specific analog signal according to the communication protocol of the corresponding cooling tower inverter; parsing the target operating frequency value of zero into a shutdown instruction for the corresponding cooling tower and converting it into a corresponding digital signal; before generating the final execution instruction, perform a safety logic check, and the safety logic check includes whether the target frequency is within the allowable range, whether the slope of the frequency change exceeds the limit, and whether the start and stop of the equipment meet the preset interlock protection conditions; send the instruction that passes the safety logic check to the on-site execution mechanism.
[0031] The conversion process from the target frequency combination to the execution instruction takes into account the actual engineering requirements, including anti-normalization processing, protocol conversion, and safety logic check. It ensures the executability and safety of the control instruction, and prevents equipment damage caused by frequency overrun, too fast change, or violation of the interlock protection conditions. The safety logic check mechanism greatly improves the engineering practicability and reliability of the system, enabling the theoretically optimized decision to be safely and effectively applied to the actual equipment control.
[0032] Through the organic combination of a graph construction module, a graph network prediction module, a hierarchical decision-making module, and an instruction execution module, the present invention constructs a closed-loop intelligent control system, effectively solving the technical problem that the explosive growth of the operation strategy combinations of multiple devices under the coexistence of different demands leads to poor energy-saving control effects. The graph-based structured representation provides the system with accurate state recognition capabilities, the temporal graph network model achieves high-precision state prediction, the hierarchical reinforcement learning architecture effectively reduces the decision-making complexity, and the safe instruction conversion ensures safe and reliable control. The deeply integrated architecture enables the system to automatically adapt to various working condition changes, optimize energy consumption while meeting different temperature control requirements, has strong adaptability, high control precision, and better energy-saving effects, and can also continuously optimize control strategies through continuous learning to adapt to equipment performance degradation and system parameter drift.
[0033] Embodiment 2: This embodiment further describes a central air-conditioning energy-saving control system based on an intelligent algorithm, including: A graph construction module, which acquires the topological structure and operating state data of the devices, constructs a dynamic graph structure, takes each device and the environmental state as nodes in the dynamic graph structure, the connection relationship between the nodes as edges in the dynamic graph structure, and the operating state data as the node feature vector; Specifically, a building management system and a dedicated sensor network are used to acquire the topological structure and the original operating state data of the devices, and the original operating state data is cleaned, filtered, and aligned through data preprocessing to obtain the operating state data.
[0034] Furthermore, the process of taking the operating state data as the node feature vector includes: collecting the operating state data corresponding to each node; acquiring the static attribute parameters corresponding to the node, where the static attribute parameters refer to the inherent characteristics of the device itself and do not change with the operating state; performing normalization processing on the operating state data and the static attribute parameters; and combining the normalized operating state data and the static attribute parameters in the predefined order of the node type to form the node feature vector; Specifically, the predefined order combination means arranging different types of parameters in a fixed order to form a feature vector in a unified format, ensuring the consistency and easy processing of different types of node feature vectors. For example, all device node feature vectors include [device ID, device type code, start / stop state...]; and the cooling tower node feature vector also includes [... frequency, inlet water temperature, outlet water temperature, power, maximum flow rate, rated power, rated air volume...].
[0035] A graph network prediction module, which receives the dynamic graph structure and the candidate frequency combination at the current moment, and uses the temporal graph network model to predict the predicted key state at the next moment; Furthermore, the temporal graph network model specifically includes: Memory layer, which associates a memory state vector with each node in the dynamic graph structure and updates the memory state vector using a long short-term memory network; the memory state vector is a hidden representation used in the temporal graph network to capture the historical state of nodes. Message function, which calculates a message based on the memory state vector of the interacting nodes at the previous moment, the node feature vector at the current moment, the edge feature, and the time difference when nodes interact; when a state update or interaction event occurs between two connected nodes, the time difference used by the message function refers to the time interval between the current moment and the last interaction or update moment of these two nodes. For example, cooling tower node A has a frequency adjustment at moment and the connected cooling water pipeline node B receives the impact of this change and changes its state at moment When calculating the message transmitted from node A to node B at moment is used as the time difference of node A, and is used as the time difference of node B. Message aggregator, which aggregates all the messages received by a node at the current moment and is used to update the memory state vector. Node embedding generation layer, which combines the memory state vector of the node at the current moment with the node feature vector to generate a node embedding representation. Prediction head, which receives the node embedding representation and the candidate frequency combination and outputs the predicted key state at the next moment; the predicted key state includes the predicted value of the outlet water temperature of the cooling loop and the predicted value of the total power consumption on the cooling side. Specifically, the candidate frequency combination refers to a set of cooling tower operation frequency schemes to be evaluated, which is generated by the hierarchical decision-making module and input to the graph network prediction module for evaluation.
[0036] Hierarchical decision-making module, which executes hierarchical reinforcement learning using the option-critic architecture and outputs the target operation frequency combination according to the dynamic graph structure representation at the current moment and the predicted key state at the next moment; hierarchical reinforcement learning includes a high-level policy agent and a low-level action agent. Furthermore, the option-critic architecture specifically includes: Shared feature extraction network, which receives the current state representation as input and outputs the shared feature vector at the current moment. Policy network, which includes a high-level option policy head and a low-level action policy head. The high-level option policy head outputs the probability distribution of each option based on the shared feature vector, and the low-level action policy head outputs the probability distribution of the candidate frequency combination for each option, and samples to obtain the target operation frequency combination. The critic network receives the current state representation and outputs the long-term cumulative reward evaluation values for selecting and executing each option in the current state, which is used to guide the parameter updates of the policy network and the termination network during the training process of hierarchical reinforcement learning. An option refers to a continuous sequence of actions and can be regarded as a behavior pattern under a specific goal. In this embodiment, the options correspond to different operating modes of the air conditioning system, and the number of options is a hyperparameter designed by the system and determined through experiments. The termination network receives the shared feature vector at the next moment and outputs the probability of terminating the currently executed option.
[0037] Furthermore, the current state representation is composed of the dynamic graph structure at the current moment and the predicted key state at the next moment; the target operating frequency combination is composed of the target operating frequency values of all controlled cooling towers.
[0038] Furthermore, the policy network learns to obtain a high-level policy agent and a low-level action agent through the training process of hierarchical reinforcement learning. During the training process of hierarchical reinforcement learning, the calculation method of the reward value includes: Calculating a first reward component according to the predicted value of the total cooling-side power consumption at the next moment output by the graph network prediction module; Calculating a second reward component according to the deviation between the predicted value of the outlet water temperature of the corresponding cooling circuit at the next moment output by the graph network prediction module and the loop target set temperature; Weightedly combining the first reward component and the second reward component according to a preset weight coefficient to obtain the reward value; Specifically, during the decision-making process, the low-level action policy head of the policy network will output the probability distribution of all possible cooling tower frequency setting combinations (i.e., candidate frequency combinations) according to the current state and the selected high-level option. The final target operating frequency combination is determined by sampling from the probability distribution. This sampling process enables the agent to tend to select combinations with higher probabilities while also retaining the opportunity to explore other possibilities.
[0039] The instruction execution module receives the target operating frequency combination and converts it into an execution instruction. Further, the process of converting the target operating frequency combination into an execution instruction includes: extracting the target operating frequency value in the target operating frequency combination; performing denormalization processing on each non-zero target operating frequency value and converting it into a specific analog signal according to the frequency converter communication protocol of the corresponding cooling tower; parsing the target operating frequency value of zero into a shutdown instruction for the corresponding cooling tower and converting it into a corresponding digital signal; before generating the final execution instruction, performing a safety logic check, where the safety logic check includes whether the target frequency is within the allowable range, whether the frequency change slope is exceeded, and whether the start and stop of the equipment meet the preset interlock protection conditions; and sending the instruction that passes the safety logic check to the on-site execution mechanism Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A central air-conditioning energy-saving control system based on intelligent algorithm, characterized in that: include: The graph construction module obtains the topological structure and operating status data of the device, constructs a dynamic graph structure, and uses each device and environmental status as a node in the dynamic graph structure, the connection relationship between nodes as an edge in the dynamic graph structure, and the operating status data as a node feature vector; The graph network prediction module receives the dynamic graph structure and candidate frequency combinations at the current moment, and uses the time graph network model to predict the prediction key state at the next moment; The hierarchical decision module uses an option-critic architecture to perform hierarchical reinforcement learning and outputs a target operation frequency combination based on the dynamic graph structure representation at the current moment and the predicted key state at the next moment; Hierarchical reinforcement learning includes high-level policy agents and low-level action agents; The instruction execution module receives the target operating frequency combination and converts it into an execution instruction.
2. According to claim 1, a central air-conditioning energy-saving control system based on intelligent algorithm is characterized in that: The process of using the operating status data as the node feature vector includes: collecting the operating status data corresponding to each node; obtaining the static attribute parameters corresponding to the node, wherein the static attribute parameters refer to the inherent characteristics of the device itself and will not change with the operating status; performing normalization processing on the operating status data and the static attribute parameters; combining the normalized operating status data and the static attribute parameters in an order predefined by the node type to form the node feature vector.
3. According to claim 1, a central air-conditioning energy-saving control system based on intelligent algorithm is characterized in that: The time graph network model specifically includes: The memory layer associates a memory state vector with each node in the dynamic graph structure and uses a long short-term memory network to update the memory state vector; Message function: When nodes interact, the message is calculated based on the memory state vector of the interacting node at the previous moment, the node feature vector at the current moment, the edge feature, and the time difference; Message aggregator, which aggregates all messages received by the node at the current moment and is used to update the memory state vector; The node embedding generation layer combines the node's current memory state vector and node feature vector to generate a node embedding representation; The prediction head receives the node embedding representation and the candidate frequency combination and outputs the predicted key state at the next moment; the predicted key state includes the predicted value of the cooling circuit outlet water temperature and the predicted value of the total power consumption on the cooling side.
4. The central air-conditioning energy-saving control system based on intelligent algorithm according to claim 1 is characterized in that: The option-critic framework specifically includes: A shared feature extraction network receives the current state representation as input and outputs a shared feature vector at the current moment; The strategy network includes a high-level option strategy head and a low-level action strategy head. The high-level option strategy head outputs the probability distribution of each option based on the shared feature vector, and the low-level action strategy head outputs the probability distribution of the candidate frequency combination for each option, and samples the target operation frequency combination. The critic network receives the current state representation and outputs the long-term cumulative reward evaluation value of selecting and executing each option in the current state, which is used to guide the parameter update of the policy network and the termination network during the training process of hierarchical reinforcement learning; The network is terminated, the shared feature vector of the next moment is received, and the probability of termination of the currently executed option is output.
5. The central air-conditioning energy-saving control system based on intelligent algorithm according to claim 4 is characterized in that: The current state representation is composed of the dynamic graph structure at the current moment and the predicted key state at the next moment; the target operating frequency combination is composed of the target operating frequency values of all controlled cooling towers.
6. The central air-conditioning energy-saving control system based on intelligent algorithm according to claim 4 is characterized in that: The strategy network learns to obtain a high-level strategy agent and a low-level action agent through a hierarchical reinforcement learning training process. During the hierarchical reinforcement learning training process, the reward value is calculated in the following manner: Calculating a first reward component according to a predicted value of total power consumption of the cooling side at the next moment output by the graph network prediction module; Calculating a second reward component according to a deviation between a predicted value of the cooling circuit outlet water temperature at the next moment output by the graph network prediction module and a target set temperature of the circuit; The first reward component and the second reward component are weighted and combined according to a preset weight coefficient to obtain the reward value.
7. The central air-conditioning energy-saving control system based on intelligent algorithm according to claim 1 is characterized in that: The process of converting the target operating frequency combination into an execution instruction includes: extracting the target operating frequency value in the target operating frequency combination; performing denormalization processing on each non-zero target operating frequency value, and converting it into a specific analog signal according to the inverter communication protocol of the corresponding cooling tower; parsing the zero target operating frequency value into a shutdown instruction of the corresponding cooling tower, and converting it into a corresponding switch signal; before generating the final execution instruction, performing a safety logic check, the safety logic check includes whether the target frequency is within the allowable range, whether the frequency change slope exceeds the limit, and whether the equipment start and stop meet the preset interlocking protection conditions; and sending the instruction that passes the safety logic check to the on-site actuator.
Citation Information
Patent Citations
Method for selecting optimal operating parameters of refrigerating system based on real-time data
CN117272212A
Air conditioner energy-saving control method and system based on reinforcement learning and digital twinborn model
CN118031385A
Method, system and equipment for optimizing operating parameters of central air conditioner and storage medium
CN118797254A
Chip production logistics optimization scheduling method and system based on deep reinforcement learning
CN119378953A
Artificial neural network machine learning device for predicting air conditioner performance based on standardized learning data and air conditioner performance prediction device using the same
KR102638026B1
Cited By
Data center double-coil cooling system energy consumption control method and system
CN121985521A
A method and system for energy consumption control of a data center double duct cooling system
CN121985521B