HVAC system optimization control method based on deep reinforcement learning

By introducing separate adversarial strategies and long-term memory networks in HVAC systems, combined with deep reinforcement learning and PPO methods, the problem of low control efficiency in complex dynamic environments is solved, and multi-objective optimization and high-rootty control are achieved.

CN120029058APending Publication Date: 2025-05-23BEIJING UNIV OF CIVIL ENG & ARCHITECTURE

Patent Information

Application Number
CN202510110172.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to maintain high efficiency in complex dynamic environments and multi-objective optimization in HVAC system control, especially in the face of sudden weather changes, fluctuations in indoor populations, equipment failures or power interference, reducing control efficiency, waste of energy and degradation of indoor environmental quality.

Method used

The HVAC system optimization control method based on deep reinforcement learning is adopted, and the system's adaptability and control effect in a dynamic environment is enhanced by introducing a separate adversarial strategy and long-term memory network, combining PPO method and dual agent parallel distribution training.

Benefits of technology

Multi-objective balance is achieved, energy consumption, thermal comfort and air quality are optimized, control strategies adapt to dynamic environmental changes, reduce training time, and improve system robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029058A_ABST
    Figure CN120029058A_ABST
Patent Text Reader

Abstract

The invention provides an HVAC system optimization control method based on deep reinforcement learning. Control of a HVAC system is optimized by introducing a separation confrontation strategy and a long short-term memory network. In the training process, a confrontation training strategy is used to enhance the robustness of a main agent, and a confrontation agent is used to simulate changes and disturbances in the environment. According to the method, energy consumption, thermal comfort (PMV) and air quality (CO2 concentration) are comprehensively optimized through a multi-objective reward function, and multi-objective balance is achieved. The long-short-term memory network is used for processing time sequence data and capturing long-term dependency in the system, and the adaptive capacity of a control strategy to dynamic environment changes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of indoor heating technology, and in particular to an HVAC system optimization control method based on deep reinforcement learning. Background Art

[0002] As global energy demand increases and environmental sustainability requirements intensify, optimizing building energy consumption has become a focus of international research. HVAC systems account for a large proportion of building energy use, and their operating efficiency plays a vital role in energy optimization and indoor environmental quality. Traditional HVAC control strategies prioritize energy efficiency, but this focus often ignores indoor air quality and occupant comfort.

[0003] In building energy management, there are many examples of deep reinforcement learning being applied to reduce energy consumption while also improving occupant comfort and air quality. However, when the system encounters sudden weather changes, large fluctuations in the number of people indoors, equipment failures, or power disturbances, these methods are often difficult to respond flexibly, resulting in reduced control efficiency, energy waste, and deterioration in indoor environmental quality. An effective solution is to introduce adversarial training into deep reinforcement learning methods. By introducing interference strategies during the training process, the robustness of deep reinforcement learning methods is enhanced, enabling them to maintain good performance in the face of various interferences. Although adversarial attacks have made some progress, their application in deep reinforcement learning has been limited. Recent studies have shown that deep neural networks are vulnerable to adversarial samples, and even tiny perturbations can cause deep neural networks to make high-confidence incorrect predictions.

[0004] In the field of HVAC system control, many researchers have used a variety of existing technologies to improve the energy efficiency, comfort and air quality of the system. These technologies include on / off control, proportional integral differential control, model predictive control, and deep reinforcement learning. Each method shows different effects in different scenarios. On / off control is one of the earliest HVAC control methods. By setting the temperature threshold, when the indoor temperature exceeds the set range, the system automatically turns on or off the HVAC equipment. This method is simple and easy to use and is suitable for building scenarios with relatively stable environmental conditions. The proportional integral differential control system continuously adjusts the output of the HVAC equipment by adjusting the three parameters of proportion, integration and differentiation to make the indoor temperature close to the preset value. The system can control the indoor ambient temperature more accurately and is widely used in automation control in the industrial and construction fields. Model predictive control is an advanced control strategy that is widely used in energy efficiency optimization in dynamic building environments. It is based on a mathematical model that can predict future system behavior and calculate the optimal control strategy. Model predictive control models are divided into the following three categories: white box models build The physical properties of the building are used to establish a dynamic model of the system, and the laws of thermodynamics are used to predict the behavior of the HVAC system. The black box model relies on a data-driven approach and uses machine learning or deep learning technology to extract the relationship between the system input and output from historical data. The gray box model combines the advantages of white box and black box, using both the theoretical basis of the physical model and data-driven technology to optimize the model parameters. Deep reinforcement learning has gradually become a popular method for optimizing HVAC system control in recent years. By applying deep neural networks, deep reinforcement learning can learn how to make the best decisions in a complex dynamic environment, thereby effectively controlling the system to improve energy efficiency and maintain the comfort of the occupants. The research on deep reinforcement learning is widely used in simulation and actual buildings, showing good performance. Summary of the invention

[0005] An embodiment of the present invention provides an HVAC system optimization control method based on deep reinforcement learning, which is used to solve the problems existing in the prior art.

[0006] In order to achieve the above object, the present invention adopts the following technical scheme.

[0007] A HVAC system optimization control method based on deep reinforcement learning, comprising:

[0008] S1 obtains the state, action and reward function of the HVAC system through the initial control model, and initializes the main agent and the adversarial agent; the reward function includes:

[0009]

[0010] In the formula, R represents the reward function, ω represents the weight, and ω 1represents the weight of air conditioning energy consumption, ω 2 represents the weight of fan energy consumption, ω 3 represents the weight of the PMV index, ω 4 Indicates CO 2 The weights of the levels are as follows: G represents the penalty for high HVAC system energy consumption, the penalty for high fan energy consumption, the penalty for deviation from the optimal PMV range, and the penalty for deviation from the normal CO 2 Function corresponding to the penalty of concentration range;

[0011] G 1 =P AC / Max AC (2)

[0012] In the formula, G 1 represents the energy consumption function of the HVAC system, P AC Indicates the value of air conditioning energy consumption, Max AC Indicates the maximum theoretical energy consumption;

[0013]

[0014] In the formula, G 2 represents the fan energy consumption function, P fan Indicates the value representing the fan energy consumption;

[0015]

[0016] In the formula, G 3 It represents the PMV index function, which is an improved formula for predicting dissatisfaction rate. When PMV is between -0.5 and 0.5, it returns a value close to zero and indicates a comfortable state for the human body. When |PMV| exceeds 0.5, it will be penalized by 20.

[0017]

[0018] In the formula, G 4 Indicates CO 2 Concentration level function, MinCO 2 and MaxCO 2 Theoretically, CO 2 Minimum and maximum concentrations;

[0019] S2 trains the main agent using the PPO method. During the training process, it also interferes with the main agent through adversarial agents and using separation adversarial strategies.

[0020] S3 processes the time series data obtained during the training process of step S2 through a long short-term memory network; the long short-term memory network is set in the policy network of the PPO method to capture the temporal dependency of the HVAC system so that the initial control model can change the control decision according to the environmental changes;

[0021] S4 pass-through

[0022]

[0023] Update the control strategy of the main agent; where, is the ratio of the current strategy probability to the old strategy probability; A t is a generalized advantage estimate, which is used to indicate the relative merits of choosing an action in the current state; the clip function limits the ratio r t In the range of 1-∈ to 1+∈; a represents the action performed at time t, π represents the strategy performed at time t, and each action corresponds to a strategy;

[0024] S5 pass-through

[0025] L value =(r t +γV θ (s t+1 )-V θ (s t )) 2 (7)

[0026] Japanese style

[0027]

[0028] Update the value network of the main agent; where r t represents the immediate reward at time step t; γ is the discount factor, which is used to indicate the importance of future rewards. Its range is between [0,1]. The closer it is to 1, the more important the future rewards are; V θ (s t+1 ) is the value network for the next state s t+1 The predicted value of V θ (s t ) is the value network for the current state s t The predicted value of t +γV θ (s t+1 ) indicates state s t The target value of ; α represents the learning rate; the discount factor γ is a value between 0 and 1, which is used to weight future rewards when calculating returns, indicating the importance of future rewards;

[0029] S6 is expressed by formula (1) and formula

[0030]

[0031] Japanese style

[0032]

[0033] Update the control strategy of the adversarial agent; the control strategy of the adversarial agent includes the switch strategy and the induction strategy; where b t Indicates the switching strategy of the current time step; a t ′ is the induced action generated by the decoy strategy, which is used to induce the main agent to take wrong actions; b in represents injection action; sw represents switch strategy, ad represents adversarial network, and lu represents decoy strategy;

[0034] S7 pass-through

[0035]

[0036] Japanese style

[0037]

[0038] Update the value network of the adversarial agent; where r t ′ represents the reward of the adversarial agent at time step t; φ is the weight of the adversarial agent;

[0039] S8 repeatedly executes steps S2 to S7 multiple times, so that the master agent updates the control strategy, and the antagonistic agent updates the interference behavior to the master agent;

[0040] S9 updates the network weights of the main agent and the adversarial agent through the gradient descent method;

[0041] S10 tests the initial control model obtained by executing step S9;

[0042] S11 repeatedly executes steps S2 to S10, and evaluates and adjusts the main agent and the adversarial agent after each round of training to obtain a target control model;

[0043] S12 uses a target control model to control the HVAC system.

[0044] Preferably, the adversarial agent comprises a recurrent neural network layer, a first fully connected layer, a long short-term memory network layer, a second fully connected layer, and a third fully connected layer;

[0045] The recurrent neural network layer and the first fully connected layer are set in parallel to each other and jointly input data to the long short-term memory network layer; the long short-term memory network layer inputs data to the second fully connected layer and the third fully connected layer respectively set in parallel; the second fully connected layer is used to generate induced error actions and output them to the main intelligent agent; the third fully connected layer is a switch strategy layer, which is used to generate noise disturbances and output them to the main intelligent agent;

[0046] The value network of the adversarial agent is located in the structure of the first fully connected layer.

[0047] As can be seen from the technical solutions provided by the embodiments of the present invention above, the present invention provides an optimized control method for an HVAC system based on deep reinforcement learning, which optimizes the control of the heating, ventilation, and air conditioning system by introducing a separation adversarial strategy and a long short-term memory network. During the training process, an adversarial training strategy is used to enhance the robustness of the main agent, and an adversarial agent is used to simulate changes and perturbations in the environment. This method comprehensively optimizes energy consumption, thermal comfort (PMV), and air quality (CO 2 concentration) through a multi-objective reward function, achieving a multi-objective balance. The long short-term memory network is used to process time series data, capture long-term dependencies in the system, and improve the adaptability of the control strategy to dynamic environmental changes.

[0048] Additional aspects and advantages of the present invention will be given in part in the following description, which will become apparent from the following description, or can be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a control framework structure diagram of an optimized control method for an HVAC system based on deep reinforcement learning provided by the present invention;

[0051] Figure 2 It is a neural network framework diagram of an adversarial agent of an optimized control method for an HVAC system based on deep reinforcement learning provided by the present invention;

[0052] Figure 3 It is a flowchart of a preferred embodiment of an optimized control method for an HVAC system based on deep reinforcement learning provided by the present invention;

[0053] Figure 4 It is a comparison diagram of the effects of the control method of the present invention for an optimized control method for an HVAC system based on deep reinforcement learning and the control method of the prior art. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0055] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.

[0056] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art in the art to which the present invention belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with the meanings in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless defined as herein.

[0057] To facilitate understanding of the embodiments of the present invention, several specific embodiments will be further explained below with reference to the accompanying drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.

[0058] The present invention provides a HVAC system optimization control method based on deep reinforcement learning, which is used to solve the following technical problems existing in the prior art:

[0059] Although the application of existing technologies in HVAC system control has shown excellent control accuracy and energy optimization effects, there are still some shortcomings, especially in complex dynamic environments and multi-objective optimization. Existing methods such as model predictive control and deep reinforcement learning encounter the following problems in practical applications:

[0060] 1. Model predictive control cannot learn from past control operations, so complex calculations are required at each time step, which leads to high computational overhead and delay. Although the system is stable, the computational burden is heavy.

[0061] 2. Deep reinforcement learning methods require a lot of data and computing resources in multi-objective optimization. When the weight ratios between multiple objectives are more complex, training deep reinforcement learning models will be very time-consuming.

[0062] 3. When faced with drastic environmental fluctuations, deep reinforcement learning methods find it difficult to flexibly adjust control strategies, which results in the system's inadequate performance in response to complex changes in external conditions and reduces the robustness of the system.

[0063] In view of this, the objects of the present invention are:

[0064] 1. By combining the separation adversarial strategy with the proximal policy optimization (PPO) method, a method is proposed to decouple the adversarial strategy into a switch strategy and an induction strategy to improve the attack accuracy of the adversarial strategy.

[0065] 2. Insert a long short-term memory network layer between the input layer and the fully connected layer of the method to optimize the processing of time series data from the HVAC system, thereby improving the efficiency of the method.

[0066] 3. Use dual agents to achieve parallel training of action strategies and separation adversarial strategies, use distributed methods for sampling and training, and reduce the training time of the agents.

[0067] 4. Simultaneously optimize and control the energy efficiency, thermal comfort and indoor air quality in the system to achieve multi-objective optimization.

[0068] The present invention provides a HVAC system optimization control method based on deep reinforcement learning, comprising:

[0069] S1 obtains the state, action and reward function of the HVAC system through the initial control model, and initializes the main agent and the adversarial agent;

[0070] S2 trains the main agent using the PPO method. During the training process, it also interferes with the main agent through adversarial agents and using separation adversarial strategies.

[0071] S3 processes the time series data obtained during the training process of step S2 through a long short-term memory network; the long short-term memory network is set in the policy network of the PPO method to capture the temporal dependency of the HVAC system so that the initial control model can change the control decision according to the environmental changes;

[0072] S4 pass-through

[0073]

[0074] Update the control strategy of the main agent; where, is the ratio of the current strategy probability to the old strategy probability; A t is a generalized advantage estimate, which is used to indicate the relative merits of choosing an action in the current state; the clip function limits the ratio r tIn the range of 1-∈ to 1+∈; a represents the action performed at time t, π represents the strategy performed at time t, and each action corresponds to a strategy.

[0075] S5 pass-through

[0076] L value =(r t +γV θ (s t+1 )-V θ (s t )) 2 (7)

[0077] Japanese style

[0078]

[0079] Update the value network of the main agent; where r t represents the immediate reward at time step t; γ is the discount factor, which is used to indicate the importance of future rewards. Its range is between [0,1]. The closer it is to 1, the more important the future rewards are; V θ (s t+1 ) is the value network for the next state s t+1 The predicted value of V θ (s t ) is the value network for the current state s t The predicted value of t +γV θ (s t+1 ) indicates state s t The target value of; α represents the learning rate;

[0080] S6 is expressed by formula (1) and formula

[0081]

[0082] Japanese style

[0083]

[0084] Update the control strategy of the adversarial agent; where b t Indicates the switch decision of the current time step; a t ′ is the induced action generated by the decoy strategy, which is used to induce the main agent to take wrong actions; b in Indicates injection action; sw is the abbreviation of switch, which is used to indicate the switch strategy; ad is the abbreviation of adversarial, which is used to indicate the adversarial network; lu is the abbreviation of lure, which is used to indicate the bait strategy.

[0085] S7 pass-through

[0086]

[0087] Japanese style

[0088]

[0089] Update the value network of the adversarial agent; where r t ′ represents the reward of the adversarial agent at time step t; φ is the weight of the adversarial agent; the discount factor γ is a value between 0 and 1 (i.e., 0≤γ≤1) used to weight future rewards when calculating returns and is used to indicate the importance of future rewards.

[0090] S8 repeatedly executes steps S2 to S7 multiple times, so that the master agent updates the control strategy, and the antagonistic agent updates the interference behavior to the master agent;

[0091] S9 updates the network weights of the main agent and the adversarial agent through the gradient descent method;

[0092] S10 tests the initial control model obtained by executing step S9;

[0093] S11 repeatedly executes steps S2 to S10, and evaluates and adjusts the main agent and the adversarial agent after each round of training to obtain a target control model;

[0094] S12 uses a target control model to control the HVAC system.

[0095] This paper proposes a HVAC system optimization control method based on deep reinforcement learning method. By introducing separation adversarial strategy and long short-term memory network, dual-agent parallel distributed training is adopted to enhance the adaptability and control effect of the system in dynamic environment and reduce the training time. It aims at the shortcomings of traditional deep reinforcement learning method in dealing with environmental disturbances and system changes.

[0096] In the preferred embodiment provided by the present invention, the following improvement scheme is specifically adopted.

[0097] 1. Reward function optimization design

[0098] In the process of optimizing the control of HVAC systems, this paper introduces a multi-objective reward function that combines energy consumption, thermal comfort (Predicted Mean Vote, PMV) and indoor air quality (CO 2 The reward function ensures that the HVAC system minimizes energy consumption while ensuring occupant comfort and air quality.

[0099] The system state S reflects the current physical variables of the environment, including outdoor temperature, indoor temperature, ambient humidity, PMV index, CO 2 concentration, heating and cooling set points, fan speed, indoor air flow rate, energy consumption of air conditioning and fans, and indoor occupancy. The control action A is to adjust the heating and cooling set points and adjust the fan speed. The temperature controller adjusts the air temperature set point of the HVAC system at time t, denoted as T S . T S is modified by a set of discrete values, using T S ={t s1 ,t s2 ,...,t sn The fan speed controller manages the mechanical fan in the ventilation system and also contains discrete actions, represented by V = {v 1 ,v 2 ,...,v n}. Therefore, the action set is represented by A = {(v,t s ):v∈V,t s ∈T S In the present invention, T S The range is from 15℃ to 30℃ and contains 32 discrete values, among which t sn and t sn+1 The difference between them is 0.5℃. The fan speed level is represented by V={0,0.5,1}, where V=1 indicates a fan speed of 1m / s. Therefore, all possible action combinations are similar to (20,0), which means the air conditioner is set to 20℃ and the mechanical fan is turned off. At each time step t, the temperature controller updates the set point of the HVAC system according to the strategy, while the fan speed controller adjusts the mechanical fan speed to ensure that the PMV is within a reasonable range, reduce energy consumption, and regulate CO 2 concentration.

[0100] The state and action constitute the core loop of the interaction between the agent and the environment. The agent perceives the environment through the state and affects the environment through the action. In summary, the state and action of this method are shown in Table 1:

[0101] Table 1 Status and actions of this method

[0102]

[0103] The reward function consists of four components: a penalty for high HVAC system energy consumption, a penalty for high fan energy consumption, a penalty for deviation from the optimal PMV range, and a penalty for deviation from the normal CO 2 The reward function is normalized as shown in formula (1) to formula (5):

[0104]

[0105] Where R represents the reward function, ω represents the weight, 1 represents the weight of air conditioning energy consumption, ω 2 represents the weight of fan energy consumption, ω 3 represents the weight of the PMV index, ω 4 Indicates CO 2 The level weights, G represents the functions corresponding to the four components.

[0106] G 1 =P AC / Max AC (2)

[0107] Among them G 1 represents the energy consumption function of the HVAC system, P AC Indicates the value of air conditioning energy consumption, Max AC Indicates the theoretical maximum energy consumption.

[0108]

[0109] Among them G 2 represents the fan energy consumption function, P fan Indicates the value representing the fan power consumption.

[0110]

[0111] Among them G 3 It represents the PMV index function, which is an improved formula for predicting dissatisfaction rate. When PMV is between -0.5 and 0.5, it returns a value close to zero and indicates the human body's comfort state. When |PMV| exceeds 0.5, it will be penalized by 20.

[0112]

[0113] Among them, G 4 Indicates CO 2 Concentration level function, MinCO 2 and MaxCO 2 Theoretically, CO 2 The minimum and maximum concentrations. 2 When the concentration deviates from the normal range, the function value will increase significantly. 2 When the concentration exceeds 1000ppm, a penalty of 20 will be imposed. This strict penalty mechanism prevents the agent from learning useless cases, such as distinguishing between PMV of -3 and -4, because both cases are unacceptable to the occupants. In order to balance energy consumption and environmental quality, the weight of each item needs to be adjusted to ensure that PMV is between -0.5 and +0.5, CO 2Levels were kept below 1000ppm.

[0114] 2. Control process design

[0115] The experimental framework of the present invention includes a pre-training loop, a control loop, and a learning loop. The present invention uses a dual-agent environment: the main agent is trained using the PPO method, which is suitable for reinforcement learning, while the adversary agent is trained using an adversarial strategy. The dual-agent setting enhances the policy robustness of the main agent through attacks generated by the adversary agent. The control framework structure is as follows Figure 1 shown.

[0116] The pre-training loop is an exploration phase before the experiment, in which actions are randomly selected, allowing the method to collect information about the environment and its possible changes. This helps the method explore a wide area of ​​the state space and avoid premature convergence or getting stuck in a local optimal solution. The present invention sets the number of steps of random sampling to 10,000 steps. Subsequently, training begins, and the policy and value networks are updated through generalized advantage estimation using the data collected from each step. In order to study the robustness of the method, the PPO method was modified to add a separation adversarial strategy and integrate a long short-term memory network to enhance network performance. The long short-term memory network is a specially designed recurrent neural network used to solve the short-term memory limitations encountered by standard recurrent neural networks when processing long sequence data. The structure of the long short-term memory network combined with the PPO method is as follows Figure 1 This is shown in the Learning Cycle section.

[0117] The learning cycle iterates through actions, data collection, model training, and deployment of the learned model during evaluation. Each record contains four different values: the state of the environment when the action was taken, the action performed, the change in state after execution, and the effect of the action evaluated by the reward function.

[0118] Whenever an action is executed, the above four values ​​and the state transition at the next time point are recorded in the experience storage memory pool. It shows how the main agent and the adversarial agent run in parallel loops. The PPO method is an Actor-Critic method, in which the Actor is responsible for outputting the strategy (selecting actions) and the Critic is responsible for estimating the value of the current state. This value estimate helps evaluate the quality of the selected action (evaluated by the generalized advantage estimation model), thereby providing guidance for policy updates. The Critic network is the value network V θ The update needs to be updated by minimizing the value loss. The value loss function is shown in formula (6):

[0119] L value =(r t +γV θ (s t+1 )-V θ(s t )) 2 (6)

[0120] Where: r t Represents the immediate reward at time step t. γ is a discount factor, which is used to indicate the importance of future rewards. Its range is between [0,1]. The closer it is to 1, the more important the future rewards are. V θ (s t+1 ) is the value network for the next state s t+1 The predicted value of V is the expected return for the next step. θ (s t ): is the value network for the current state s t The predicted value of t +γV θ (s t+1 ): indicates state s t The target value (i.e. the current reward r t Plus the discounted return for future states), which is an estimate of future rewards in temporal difference learning. a represents the action performed at time t, π represents the strategy at time t, and each action corresponds to a strategy.

[0121] To minimize the value loss, the weight θ is updated by gradient descent as shown in formula (7):

[0122]

[0123] Among them, α represents the learning rate. The learning rate is a hyperparameter used to control the step size of the model at each parameter update.

[0124] Adversarial Value Network V φ By minimizing the adversarial value loss update, the value loss function is shown in formula (8):

[0125]

[0126] Among them, r t ′ represents the reward of the adversarial agent at time step t. The discount factor γ is a value between 0 and 1 (i.e., 0≤γ≤1) and is used to weight future rewards when calculating returns, and is used to indicate the importance of future rewards. The weight φ is updated by gradient descent, as shown in formula (9):

[0127]

[0128] In formula (8), the value network consists of an input layer, a hidden layer, and an output layer. The input of the value network is the current state vector s t The hidden layer is composed of a multi-layer recurrent neural network. The output of the value network is a scalar, which represents the state value V of the current state.θ (s t ), which is the expected estimate of future cumulative rewards. The prediction error of the value network is back-propagated through the reward signal. Therefore, the optimization objective of the value network is closely related to the reward.

[0129] In the embodiment provided by the present invention, the main agent is set up based on the architecture of the prior art, such as Figure 1 The architecture shown in , which includes an input layer, a long short-term memory network layer, a hidden layer and an output layer arranged in a data flow direction. The specific structure and function can be arranged in a known manner and will not be repeated here.

[0130] In the method of the present invention, a tool called Ray is used to speed up the training. Ray is an open source unified framework for extending artificial intelligence and Python applications (such as machine learning), which provides a computing layer for parallel processing. The principle of this tool is to divide the training work into multiple small tasks and let multiple computing units (called "worker nodes") complete these tasks at the same time. It can be compared to a team collaboration, where everyone performs their duties and then summarizes the results, greatly improving efficiency. Specifically, there is a "coordinator" whose task is to distribute work and collect the results. The working steps are as follows: the coordinator first distributes the data to each worker node. After each worker node receives the data, it interacts with the simulation environment to generate the required training data. Then, each worker node returns the generated data to the coordinator. Finally, the coordinator uses this data to optimize the decision model of the intelligent agent. In this distributed training, you can control how many worker nodes perform tasks simultaneously. By adjusting the 'num_workers' parameter, you can increase or decrease the number of these worker nodes. In this method, 'num_workers' is set to 4, which means that there are 4 worker nodes working simultaneously to train two agents at the same time, which can significantly speed up the training.

[0131] In the embodiment provided by the present invention, the control model of the HVAC system adopts the control model of the prior art, and the specific style and configuration are not described in detail here.

[0132] 3. Strategy optimization design

[0133] Introduction of separate adversarial strategy: This invention uses a new adversarial strategy. By decomposing the adversarial strategy into two parts, the switching strategy and the induction strategy, the disturbance factors in the environment are simulated without affecting the normal operation of the main agent. This strategy enhances the robustness of the main agent to environmental changes and improves the stability of the HVAC system when dealing with complex dynamic environments. The present invention introduces adversarial training and uses two agents: the main agent uses the PPO method strategy to control the HVAC system, while the adversarial agent is trained using a separate adversarial strategy to optimize its anti-interference ability. The original strategy of the main agent is denoted as π, and the strategy of the attacking agent is denoted as π ad The split adversarial strategy decomposes the adversarial approach into two independent sub-strategies: the switch strategy and induction strategies in Decide whether to attack, Determine which action to induce the master agent to take. In the split adversarial attack strategy, the two agents share the same action set but independently sample two actions: b t and a t , where a t is a potential action, b t The method of the present invention updates each sub-strategy independently. The update of the Actor network, i.e., the strategy network, is performed by maximizing the objective function. The objective function of the standard PPO method is shown in formula (10):

[0134]

[0135] in: is the ratio of the current strategy probability to the old strategy probability. t It is a generalized advantage estimate, which is used to express the relative merits of choosing an action in the current state. The clip function limits the ratio r t In the range of 1-∈ to 1+∈, thus preventing the policy update from being too large.

[0136] Due to the separation of the confrontation strategy, the switch strategy is designed and bait strategy This method defines independent ratios for each sub-strategy: switch strategy ratio It is used to measure the necessity of updating the “switching strategy”. The calculation formula is shown in formula (11):

[0137]

[0138] Among them, sw is the abbreviation of switch, which is used to indicate the switch strategy, ad is the abbreviation of adversarial, which is used to indicate the adversarial network, and lu is the abbreviation of lure, which is used to indicate the bait strategy. tIndicates the switch decision of the current time step (i.e., whether to inject disturbance). This ratio reflects the changes before and after the switch strategy is updated, ensuring that the switch strategy does not deviate significantly from the original strategy. Only those steps where disturbances are actually injected are calculated. Its formula is shown in formula (12):

[0139]

[0140] where a t ′ is the induced action generated by the decoy strategy, which is used to induce the main agent to take wrong actions, and b in represents the “injection action”. Only when the switching strategy decides to inject a disturbance (i.e., b t =b in ), otherwise this item is set to 0 and not included in the update. The reason for this is that if the current switch strategy does not choose to inject, the bait strategy will not affect the actual attack behavior, and updating the bait strategy will be invalid. When the adversarial agent decides to inject disturbances at a certain moment through the switching strategy, the bait strategy will select a specific induced action (wrong action) as the interference target. Then, the adversarial agent will inject an adversarial disturbance into the state of the main agent, with the aim of causing the victim agent to misjudge the situation in the current state, thereby taking this erroneous induced action. Noise disturbance is used in this method. In a preferred embodiment, the neural network framework diagram of the adversarial agent embodying the above formula is as shown in Figure 2 As shown. The adversarial agent includes a recurrent neural network layer, a first fully connected layer, a long short-term memory network layer, a second fully connected layer, and a third fully connected layer. The recurrent neural network layer and the first fully connected layer are arranged in parallel to input data to the long short-term memory network layer; the long short-term memory network layer inputs data to the second fully connected layer and the third fully connected layer arranged in parallel respectively; the second fully connected layer is used to generate induced error actions and output them to the main agent, and the third fully connected layer is used to generate noise disturbances and output them to the main agent.

[0141] In some feasible embodiments, the first fully connected layer is set according to the framework of the prior art, and the value network of the adversarial agent is located in the structure of the first fully connected layer. The value network also adopts a known architecture setting, which will not be described in detail.

[0142] The present invention also provides an embodiment to show a preferred implementation process and control effect of the method provided by the present invention. Figure 3 As shown, the execution steps are as follows:

[0143] Step 1: Energyplus is a whole building energy consumption simulation software that can be used to simulate the energy consumption (heating, cooling, ventilation, lighting, plug and process load) and water consumption of buildings. Use Energyplus software to simulate a building with a HVAC system, and compare the indoor temperature, humidity, CO 2 The sensor data such as concentration is used as input data and recorded as state S t At the same time, the corresponding control operation (heating and cooling set points, fan speed) is prepared, which is recorded as action a t The input time series data is used to train the reinforcement learning model.

[0144] Step 2: Initialize the agent and the adversarial agent. The main agent uses the PPO method, while the adversarial agent is based on a separation adversarial strategy to simulate interference with the main agent during training. Set the initial policy network and value network weights of the dual agents, and initialize the experience storage memory pool to store the state, action, and reward data during training.

[0145] Step 3: Pre-training phase. Using the source environment data, the main agent uses the PPO method to perform preliminary training in the HVAC system without adversarial interference to obtain a preliminary control strategy. At this time, the long short-term memory network is combined with the PPO method to capture the long-term dependencies in the time series and improve the model's ability to adapt to dynamic environmental changes. The network is trained by replaying the experience data and looping the training for 10,000 times to obtain a preliminary pre-training model.

[0146] Step 4: Enter the formal training phase. t When , record the current state s t 、Action a t , Reward t and the next state s (t+1) , and store these data in the memory pool. The main agent uses the PPO method to train these data and update the policy network and value network.

[0147] Step 5: Introduce adversarial interference. While the main agent is being trained, start training the adversarial agent. The adversarial agent separates the adversarial strategy and generates interference operations in the environment to simulate the emergencies of the HVAC system in actual applications (such as a sudden increase in the number of people indoors, equipment failure, etc.). The adversarial agent uses an independent strategy network and value network to update its adversarial strategy to maximize the interference effect.

[0148] Step 6: Use long short-term memory networks to process time series data. The long short-term memory network layer is integrated in the policy network of the PPO method to process time series inputs to capture the temporal dependencies of the HVAC system and make smarter control decisions based on changes over long periods of time. This process can improve the stability of the model in long-term environments.

[0149] Step 7: Update the main agent strategy through formula (10). The main agent optimizes the control strategy by maximizing the objective function based on the collected data.

[0150] Step 8: Update the value network of the main agent through formula (6) and formula (7), and update the value network of the main agent through the minimum value loss function.

[0151] Step 9: Update the adversarial agent strategy through formula (10), formula (11) and formula (12). The adversarial agent's strategy network is implemented by maximizing the objective function.

[0152] Step 10: Update the value network of the adversarial agent through formula (8) and formula (9). The value network of the adversarial agent is updated by minimizing the adversarial value loss.

[0153] Step 11: Repeat the above steps. During the entire training process, the main agent and the adversarial agent update their respective strategies and value networks in each round of training. The main agent optimizes the control strategy of the HVAC system in the presence of interference by learning, while the adversarial agent continuously generates new interference to test the robustness of the main agent.

[0154] Step 12: During the training process, the main agent’s goal is achieved by setting the reward function to minimize the energy consumption of the HVAC system while maintaining the indoor PMV and CO 2 concentration.

[0155] Step 13: Use the gradient descent method to update the network weights of the main agent and the adversarial agent, and save the model. After each round of training, calculate the control effect of the main agent under interference and adjust the strategy according to the reward function.

[0156] Step 14: After the model training is completed, the real-time data in the HVAC system is input into the trained main intelligent agent to verify the performance of its control strategy in the actual environment and evaluate its energy efficiency and comfort optimization effects under different interferences.

[0157] Step 15: Repeat the training for several rounds until the model converges. After each training, the main agent and the adversarial agent are evaluated and adjusted to ensure that the final control strategy has high robustness and generalization ability.

[0158] Step 16: End the entire training process and save the final main agent policy and value network. Apply it to the actual HVAC system to achieve intelligent control.

[0159] The deep reinforcement learning control model (DAL-PPO) of the present invention is trained by the main agent and the adversarial agent respectively. This method is trained based on 15-year meteorological data from Shanghai, which belongs to the subtropical monsoon climate. The experiment was carried out in a classroom environment that can accommodate 30 people, and the data from 2010 was used for comparison, and the effectiveness of the separated adversarial strategy in attacking reinforcement learning was evaluated. Subsequently, DAL-PPO was compared with the traditional PPO method in an environment with perturbations. As Figure 4 shown, even in the case of sudden changes in real environmental conditions, excellent control performance can still be maintained. The experimental result data of the average PMV, CO 2 level average, and average total energy consumption are shown in Table 2.

[0160]

[0161] Table 2 DAL-PPO represents the method of the present invention, and the experimental result data is compared with the PPO method

[0162] In summary, the present invention provides an optimized control method for HVAC systems based on deep reinforcement learning. By introducing a separated adversarial strategy and a long short-term memory network, the control of the HVAC system is optimized. During the training process, an adversarial training strategy is used to enhance the robustness of the main agent, and the adversarial agent is used to simulate the changes and perturbations in the environment. This method comprehensively optimizes energy consumption, thermal comfort (PMV), and air quality (CO 2 concentration) through a multi-objective reward function, achieving a multi-objective balance. The long short-term memory network is used to process time series data, capture the long-term dependencies in the system, and improve the adaptability of the control strategy to dynamic environmental changes.

[0163] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.

[0164] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention or certain parts of the embodiments.

[0165] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0166] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed by the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A HVAC system optimization control method based on deep reinforcement learning, characterized in that: include: S1 obtains the state, action and reward function of the HVAC system through the initial control model, and initializes the main agent and the adversarial agent; the reward function includes: Where R represents the reward function, ω represents the weight, ω1 represents the weight of air conditioning energy consumption, ω2 represents the weight of fan energy consumption, ω3 represents the weight of PMV index, ω4 represents the weight of CO2 level, and G represents the function corresponding to the penalty for high energy consumption of HVAC system, the penalty for high energy consumption of fan, the penalty for deviation from the optimal PMV range, and the penalty for deviation from the normal CO2 concentration range; G1=P AC / Max AC (2) Where G1 represents the energy consumption function of the HVAC system, P AC Indicates the value of air conditioning energy consumption, Max AC Indicates the maximum theoretical energy consumption; Where G2 represents the fan energy consumption function, P fan Indicates the value representing the fan energy consumption; Where G3 represents the PMV index function, which is an improved prediction dissatisfaction rate formula. When PMV is between -0.5 and 0.5, it returns a value close to zero and indicates a comfortable state for the human body. When |PMV| exceeds 0.5, a penalty of 20 will be imposed. In the formula, G4 represents the CO2 concentration level function, MinCO2 and MaxCO2 represent the minimum and maximum theoretical CO2 concentrations, respectively; S2 trains the main agent using the PPO method. During the training process, it also interferes with the main agent through adversarial agents and using separation adversarial strategies. S3 processes the time series data obtained during the training process of step S2 through a long short-term memory network; the long short-term memory network is set in the policy network of the PPO method to capture the temporal dependency of the HVAC system so that the initial control model can change the control decision according to the environmental changes; S4 pass-through Update the control strategy of the main agent; where, is the ratio of the current strategy probability to the old strategy probability; A t is a generalized advantage estimate, which is used to indicate the relative merits of choosing an action in the current state; the clip function limits the ratio r t In the range of 1-∈ to 1+∈; a represents the action performed at time t, π represents the strategy performed at time t, and each action corresponds to a strategy; S5 pass-through 50 value =(r t +γV θ (s t+1 )-V θ (s t )) 2 (7) Japanese style Update the value network of the main agent; where r t represents the immediate reward at time step t; γ is the discount factor, which is used to indicate the importance of future rewards. Its range is between [0,1]. The closer it is to 1, the more important the future rewards are; V θ (s t+1 ) is the value network for the next state s t+1 The predicted value of V θ (s t ) is the value network for the current state s t The predicted value of t +γV θ (s t+1 ) indicates state s t The target value of ; α represents the learning rate; the discount factor γ is a value between 0 and 1, which is used to weight future rewards when calculating returns, indicating the importance of future rewards; S6 is expressed by formula (1) and formula Japanese style Update the control strategy of the adversarial agent; the control strategy of the adversarial agent includes the switch strategy and the induction strategy; where b t Indicates the switching strategy of the current time step; a t ′ is the induced action generated by the decoy strategy, which is used to induce the main agent to take wrong actions; b in represents injection action; sw represents switch strategy, ad represents adversarial network, and lu represents decoy strategy; S7 pass-through Japanese style Update the value network of the adversarial agent; where r t ′ represents the reward of the adversarial agent at time step t; φ is the weight of the adversarial agent; S8 repeatedly executes steps S2 to S7 multiple times, so that the master agent updates the control strategy, and the antagonistic agent updates the interference behavior to the master agent; S9 updates the network weights of the main agent and the adversarial agent through the gradient descent method; S10 tests the initial control model obtained by executing step S9; S11 repeatedly executes steps S2 to S10, and evaluates and adjusts the main agent and the adversarial agent after each round of training to obtain a target control model; S12 uses a target control model to control the HVAC system.

2. The method according to claim 1, characterized in that The adversarial agent includes a recurrent neural network layer, a first fully connected layer, a long short-term memory network layer, a second fully connected layer, and a third fully connected layer; The recurrent neural network layer and the first fully connected layer are set in parallel to each other and jointly input data to the long short-term memory network layer; the long short-term memory network layer inputs data to the second fully connected layer and the third fully connected layer respectively set in parallel; the second fully connected layer is used to generate induced error actions and output them to the main intelligent agent; the third fully connected layer is a switch strategy layer, which is used to generate noise disturbances and output them to the main intelligent agent; The value network of the adversarial agent is located in the structure of the first fully connected layer.

Citation Information

Patent Citations

  • Adversarial task-oriented man-machine symbiosis reinforcement learning method and device, computing equipment and storage medium

    CN113688977A

  • Incoming missile defense and confrontation system and method based on deep reinforcement learning

    CN115562007A

  • Heating ventilation air conditioner regulation and control method and device based on reinforcement learning

    CN115950080A

  • Interference strategy sensing method based on generative adversarial imitation learning

    CN116643242A

  • Intelligent killer chain generation method based on reinforcement learning

    CN118364601A

Cited By

  • Intelligent edible mushroom shelter temperature control method and system based on artificial intelligence

    CN122331664A