Building HVAC adaptive coordination control method and device based on deep reinforcement learning

Through the method based on deep reinforcement learning, the model of the HVAC system is constructed and optimized, and the control problem of the HVAC system under complex environments and random disturbances is solved, and efficient and stable energy efficiency optimization and thermal comfort balance are achieved.

CN120010255AActive Publication Date: 2025-05-16TIANJIN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510128677.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-05-16
Estimated Expiration
2045-02-05

AI Technical Summary

Technical Problem

The existing HVAC system control methods are difficult to achieve stability and efficient control under complex environments and random disturbances, resulting in insufficient optimization of energy efficiency and difficult to balance thermal comfort.

Method used

Adaptive coordination control method of building HVAC based on deep reinforcement learning is adopted, and the indoor temperature change model and building thermodynamic model are constructed, combined with the combined simulation framework of EnergyPlus and Python, it is transformed into a Markov decision-making process model, and the deep reinforcement learning algorithm EB-TD3 is used for iterative optimization to achieve adaptive real-time dynamic control.

Benefits of technology

Improve the control efficiency and stability of the HVAC system under complex environments and random disturbances, reduce energy consumption, improve thermal comfort, and achieve the dual needs of building energy efficiency and user comfort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010255A_ABST
    Figure CN120010255A_ABST
Patent Text Reader

Abstract

The invention discloses a building HVAC adaptive coordination control method and device based on deep reinforcement learning. The method comprises the following steps: constructing an indoor temperature change model; describing the indoor temperature change through the indoor temperature change model; building a building thermodynamic model based on the indoor temperature change model; based on the building thermal dynamic model, combining an EnergyPlus and Python joint simulation framework to construct a real-time dynamic control model of the building HVAC system; converting the real-time dynamic control model of the building HVAC system into a Markov decision process model; based on a set expert rule, constructing a deep reinforcement learning algorithm EB-TD3; and carrying out iterative optimization on the Markov decision process model through the deep reinforcement learning algorithm EB-TD3 according to a set optimization target to realize adaptive real-time dynamic control of the building HVAC system. According to the invention, the control efficiency and stability of the HVAC system in a complex environment are improved, the energy consumption is reduced, and the thermal comfort is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building HVAC system optimization control, and in particular to a building HVAC adaptive coordination control method and device based on deep reinforcement learning. Background Art

[0002] With the improvement of building energy efficiency standards and the pursuit of low-carbon goals, building HVAC systems are facing increasingly severe challenges in optimizing energy efficiency and ensuring thermal comfort. Traditional HVAC system control methods mostly rely on fixed rules and experience, and cannot flexibly respond to complex environmental changes and external disturbances, resulting in the failure to fully optimize building energy efficiency and difficulty in balancing energy efficiency and thermal comfort. Especially under complex working conditions, the HVAC system needs to effectively adjust indoor temperature, humidity and other factors to adapt to external weather changes, sudden load fluctuations and the equipment's own operating disturbances. Existing HVAC control strategies usually ignore the impact of random disturbances in the system on the control effect, resulting in poor control accuracy and response speed.

[0003] At present, deep reinforcement learning (DRL), as an intelligent optimization method, can learn the optimal strategy through interaction with the environment and show good performance in many control problems. However, the control of HVAC systems faces great challenges under complex working conditions, especially under the influence of variable indoor and outdoor environmental conditions and random disturbances. How to ensure the stability of the system, improve the convergence speed, and maintain indoor thermal comfort has always been a difficult research point.

[0004] Therefore, how to invent a new adaptive control method for HVAC systems to improve the control efficiency and stability of HVAC systems in complex environments, reduce energy consumption and improve thermal comfort has become an urgent problem to be solved. Summary of the invention

[0005] To this end, the present invention provides a building HVAC adaptive coordination control method and device based on deep reinforcement learning. By combining deep reinforcement learning algorithm and expert knowledge, the control efficiency and stability of the HVAC system in a complex environment are improved, energy consumption is reduced and thermal comfort is improved.

[0006] In order to achieve the above object, the present invention provides the following technical solution: a building HVAC adaptive coordinated control method based on deep reinforcement learning, comprising:

[0007] Constructing an indoor temperature change model; describing the indoor temperature change through the indoor temperature change model; constructing a building thermodynamic model based on the indoor temperature change model;

[0008] Based on the building thermodynamic model, combined with EnergyPlus and Python co-simulation framework, a real-time dynamic control model of the building HVAC system is constructed;

[0009] Converting the real-time dynamic control model of the building HVAC system into a Markov decision process model;

[0010] Based on the set expert rules, a deep reinforcement learning algorithm EB-TD3 is constructed; according to the set optimization goals, the Markov decision process model is iteratively optimized through the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

[0011] As a preferred solution of the building HVAC adaptive coordinated control method based on deep reinforcement learning, the expression of the indoor temperature change model is:

[0012]

[0013] Where δt is the time increment; is the temperature of the current time step; and are the temperatures of the first three time steps respectively; O represents the approximation order, i.e. the error term.

[0014] As a preferred solution of the building HVAC adaptive coordinated control method based on deep reinforcement learning, the expression of the building thermal dynamics model is:

[0015]

[0016] C h =ρ air ×C ρ ×C T

[0017] In the formula, C h is the heat capacity of the indoor air; ρ air is the air density; C ρ is the specific heat capacity of air; C T is the volume of air; C p is the specific heat capacity at constant pressure; N l is the total number of hot zones; N s is the total number of heat exchange surfaces; N z is the total number of fluid branches; h i is the surface heat transfer coefficient; A i is the surface area; T si is the room surface temperature; T h is the room air temperature; is the air mass flow rate entering from the adjacent room; T hi is the temperature of the air in the adjacent room; is the mass flow rate of infiltrated external air; T ∞ is the outside air temperature; is the sum of external heat sources; is the sum of the internal heat sources of the system.

[0019] As a preferred solution for the building HVAC adaptive coordinated control method based on deep reinforcement learning, the core elements of the Markov decision process model include: state space, action space, state transition probability, reward and discount factor;

[0020] The expression of the state space is:

[0021] S={s0,s1,...,s t ,s t+1 ,...}

[0022] Where S is the state space; s t is the state of the agent at time t;

[0023] The expression of the action space is:

[0024] A={a0,a1,...,a n}

[0025] Where A is the action space; a n Indicates a specific control action;

[0026] The state transition probability represents the possibility of transition between states, reflecting the agent's transition from s t Transfer to t+1 process;

[0027] The reward expression is:

[0028] R={r0,r1,...,r t ,...,}

[0029] In the formula, R is the reward; r t The agent is in state s at time t t And perform action a t Rewards received when

[0030] The discount factor is used to balance the importance of immediate rewards and future rewards.

[0031] As a preferred solution of the building HVAC adaptive coordinated control method based on deep reinforcement learning, in the process of iteratively optimizing the Markov decision process model by the deep reinforcement learning algorithm EB-TD3 according to the set optimization target, the iterative optimization steps are:

[0032] Initialize the parameters of the Actor network and the Critic network; based on the same parameters, construct a target network corresponding to the Actor network and the Critic network;

[0033] Acquire the current state through the Actor network, and determine the control action according to the current state; add noise when determining the control action to enhance the exploration ability of the intelligent agent;

[0034] After each control action is executed, the environment provides a reward to the agent and transitions to the next state; the quadruple is stored in the experience replay buffer until the round ends;

[0035] Sampling a small batch of experience from the experience playback buffer; and obtaining a q value of the target network through the small batch of experience calculation;

[0036] According to the loss function of the Critic network, the Critic network is updated;

[0037] Based on the policy gradient theorem, the Actor network is updated;

[0038] According to the soft update strategy, the target network corresponding to the Actor network and the Critic network is updated.

[0039] The present invention also provides a building HVAC adaptive coordination control device based on deep reinforcement learning, based on the above building HVAC adaptive coordination control method based on deep reinforcement learning, including:

[0040] A building thermodynamic model building module is used to build an indoor temperature change model; describe the indoor temperature change through the indoor temperature change model; and build a building thermodynamic model based on the indoor temperature change model;

[0041] A building HVAC system real-time dynamic control model construction module, which is used to construct a building HVAC system real-time dynamic control model based on the building thermodynamic model and in combination with the EnergyPlus and Python joint simulation framework;

[0042] A Markov decision process model conversion module, used to convert the real-time dynamic control model of the building HVAC system into a Markov decision process model;

[0043] The deep reinforcement learning algorithm EB-TD3 construction and iterative optimization module is used to construct the deep reinforcement learning algorithm EB-TD3 based on the set expert rules; according to the set optimization goals, the Markov decision process model is iteratively optimized through the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

[0044] As a preferred solution of the building HVAC adaptive coordination control device based on deep reinforcement learning, in the building thermodynamic model construction module, the expression of the indoor temperature change model is:

[0045]

[0046] Where δt is the time increment; is the temperature of the current time step; and are the temperatures of the first three time steps respectively; O represents the approximation order, i.e. the error term.

[0047] As a preferred solution of the building HVAC adaptive coordination control device based on deep reinforcement learning, in the building thermodynamic model construction module, the expression of the building thermodynamic model is:

[0048]

[0049] C h =ρ air ×C ρ ×C T

[0050] In the formula, C h is the heat capacity of the indoor air; ρ air is the air density; C ρ is the specific heat capacity of air; C T is the volume of air; C p is the specific heat capacity at constant pressure; N l is the total number of hot zones; N s is the total number of heat exchange surfaces; N z is the total number of fluid branches; h i is the surface heat transfer coefficient; A i is the surface area; T si is the room surface temperature; T h is the room air temperature; is the air mass flow rate entering from the adjacent room; T hi is the temperature of the air in the adjacent room; is the mass flow rate of infiltrated external air; T ∞ is the outside air temperature; is the sum of external heat sources; is the sum of the internal heat sources of the system.

[0051] As a preferred solution of the building HVAC adaptive coordination control device based on deep reinforcement learning, in the Markov decision process model conversion module, the core elements of the Markov decision process model include: state space, action space, state transition probability, reward and discount factor;

[0052] The expression of the state space is:

[0053] S={s0,s1,...,s t ,s t+1 ,...}

[0054] Where S is the state space; s t is the state of the agent at time t;

[0055] The expression of the action space is:

[0056] A={a0,a1,...,a n}

[0057] Where A is the action space; a n Indicates a specific control action;

[0058] The state transition probability represents the possibility of transition between states, reflecting the agent's transition from s t Transfer to t+1 process;

[0059] The reward expression is:

[0060] R={r0,r1,...,r t ,...,}

[0061] In the formula, R is the reward; r t The agent is in state s at time t t And perform action a t Rewards received when

[0062] The discount factor is used to balance the importance of immediate rewards and future rewards.

[0063] As a preferred solution for the building HVAC adaptive coordination control device based on deep reinforcement learning, in the deep reinforcement learning algorithm EB-TD3 construction and iterative optimization module, the iterative optimization submodule includes:

[0064] The network initialization submodule is used to initialize the parameters of the Actor network and the Critic network; based on the same parameters, a target network corresponding to the Actor network and the Critic network is constructed;

[0065] A control action determination submodule is used to obtain the current state through the Actor network and determine the control action according to the current state; when determining the control action, noise is added to enhance the exploration ability of the intelligent agent;

[0066] An iterative optimization submodule, which is used for providing a reward to the agent and transitioning to the next state after each control action is executed; storing the quadruple in the experience replay buffer until the round ends;

[0067] The q-value acquisition submodule of the target network is used to sample small batch experience from the experience playback buffer; and obtain the q-value of the target network through the small batch experience calculation;

[0068] A critic network update submodule, used to update the critic network according to the loss function of the critic network;

[0069] Actor network update submodule, used to update the Actor network based on the policy gradient theorem;

[0070] The target network update submodule is used to update the target network corresponding to the Actor network and the Critic network according to the soft update strategy.

[0071] The present invention has the following advantages: the present invention constructs an indoor temperature change model; describes the indoor temperature change through the indoor temperature change model; constructs a building thermodynamic model based on the indoor temperature change model; constructs a real-time dynamic control model of the building HVAC system based on the building thermodynamic model, combined with the EnergyPlus and Python joint simulation framework; converts the real-time dynamic control model of the building HVAC system into a Markov decision process model; constructs a deep reinforcement learning algorithm EB-TD3 based on the set expert rules; and iteratively optimizes the Markov decision process model through the deep reinforcement learning algorithm EB-TD3 according to the set optimization target, so as to realize the adaptive real-time dynamic control of the building HVAC system. The present invention combines the deep reinforcement learning algorithm with expert knowledge, can realize the adaptive adjustment of the HVAC system under complex environment and random disturbance, improve the flexibility and stability of the control strategy, and can achieve the balance between energy saving and comfort, optimize the environmental parameters such as indoor temperature and humidity, ensure that the indoor thermal comfort is maintained while saving energy, and meet the dual needs of building energy efficiency and user comfort. In terms of control strategy, by introducing expert knowledge, the temperature fluctuations caused by the exploration process are reduced, the convergence speed and training stability of the algorithm are improved, and long-term stable operation is ensured. The proposed control strategy can cope with the complex changes in the internal and external environment of the building, such as seasonal fluctuations, climate change, etc., and dynamically adjust the operating parameters of the HVAC system to achieve the best control effect. By combining with simulation platforms such as EnergyPlus, the feasibility of the control method of the present invention in practical applications is verified, and it can effectively improve the operating efficiency of the building HVAC system. BRIEF DESCRIPTION OF THE DRAWINGS

[0072] In order to more clearly illustrate the implementation methods of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the implementation methods or the description of the prior art. Obviously, the drawings in the following description are only exemplary, and for ordinary technicians in this field, other implementation drawings can be derived from the provided drawings without creative work.

[0073] The structures, proportions, sizes, etc. illustrated in this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with the technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantial technical significance. Any structural modification, change in proportion or adjustment of size shall still fall within the scope of the technical contents disclosed in the present invention without affecting the effects and purposes that can be achieved by the present invention.

[0074] Figure 1 This is a schematic diagram of the process flow of the building HVAC adaptive coordination control method based on deep reinforcement learning provided in Example 1 of the present invention;

[0075] Figure 2 A schematic diagram of a building thermodynamic model modeling method in a building HVAC adaptive coordinated control method based on deep reinforcement learning provided in Example 1 of the present invention;

[0076] Figure 3 A schematic diagram of a real-time dynamic control process of the EnergyPlus model in the building HVAC adaptive coordinated control method based on deep reinforcement learning provided in Example 1 of the present invention;

[0077] Figure 4 A schematic diagram of a Markov decision process model in a building HVAC adaptive coordinated control method based on deep reinforcement learning provided in Example 1 of the present invention;

[0078] Figure 5 This is a schematic diagram of the architecture of the building HVAC adaptive coordination control device based on deep reinforcement learning provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0079] The following is a description of the implementation of the present invention by specific embodiments. People familiar with the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0080] Example 1

[0081] See also Figure 1 Embodiment 1 of the present invention provides a building HVAC adaptive coordination control method based on deep reinforcement learning, comprising the following steps:

[0082] S1. Constructing an indoor temperature change model; describing indoor temperature changes through the indoor temperature change model; and constructing a building thermodynamic model based on the indoor temperature change model;

[0083] S2. Based on the building thermodynamic model, combined with EnergyPlus and Python joint simulation framework, a real-time dynamic control model of the building HVAC system is constructed;

[0084] S3, converting the real-time dynamic control model of the building HVAC system into a Markov decision process model;

[0085] S4. Based on the set expert rules, a deep reinforcement learning algorithm EB-TD3 is constructed; according to the set optimization goals, the Markov decision process model is iteratively optimized by the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

[0086] In this embodiment, in step S1, an indoor temperature change model is constructed; indoor temperature changes are described by the indoor temperature change model; and a building thermodynamic model is constructed based on the indoor temperature change model;

[0087] Wherein, the expression of the indoor temperature change model is:

[0088]

[0089] Where δt is the time increment; is the temperature of the current time step; and are the temperatures of the first three time steps respectively; O represents the approximation order, i.e. the error term.

[0090] In this embodiment, a building thermodynamic model is constructed based on the indoor temperature change model;

[0091] Specifically, for a building HVAC system, the modeling method is as follows Figure 2 Sketchup was used to build the geometric model, Openstudio was used to build the thermal zone, and EnergyPlus was used to give the building thermodynamic properties.

[0092] according to Figure 2 It can be seen that there are five main factors affecting the temperature change in the building, namely the total convection heat load of all thermal zones, the convection heat transfer of the room surface, the heat transfer caused by air mixing between different thermal zones, the heat transfer caused by external air infiltration, and the energy input of the air system. Therefore, the expression of the building thermodynamic model is:

[0093]

[0094] C h =ρ air ×C ρ ×C T

[0095] In the formula, C h is the heat capacity of the indoor air; ρ air is the air density; C ρ is the specific heat capacity of air; C T is the volume of air; C p is the specific heat capacity at constant pressure; N l is the total number of hot zones; N sis the total number of heat exchange surfaces; N z is the total number of fluid branches; h i is the surface heat transfer coefficient; A i is the surface area; T si is the room surface temperature; T h is the room air temperature; is the air mass flow rate entering from the adjacent room; T hi is the temperature of the air in the adjacent room; is the mass flow rate of infiltrated external air; T ∞ is the outside air temperature; is the sum of external heat sources; is the sum of the internal heat sources of the system.

[0096] In this embodiment, in step S2, based on the building thermodynamic model, combined with EnergyPlus and Python joint simulation framework, a real-time dynamic control model of the building HVAC system is constructed;

[0097] Among them, EnergyPlus simulation has higher accuracy. However, its real-time dynamic control is limited by the ERL programming language used in the software, which limits the wide applicability of EnergyPlus as a simulation engine. In order to solve this problem, the Python interface pyenergyplus of EnergyPlus is used for secondary development, so that the DRL algorithm can be defined in Python, and the real-time dynamic control of the EnergyPlus model can be realized, which further expands the potential of EnergyPlus as a simulation engine. The specific implementation process is as follows: Figure 3 shown.

[0098] In this embodiment, the modeling information part needs to define various parameters and EnergyPlus model files. The EnergyPlus model file is obtained through the aforementioned modeling process. In the agent definition part, parameters such as meteorological information, simulation duration and simulation time step need to be customized. The meteorological information comes from the data of the typical meteorological year (Typical MeteorologicalYear), which represents the local meteorological characteristics. The simulation duration refers to the time range of the simulation. The simulation time step is δt. In the definition of the DRL algorithm, in order to achieve real-time dynamic control, the following contents need to be defined: control actions, state information, reward functions and expert knowledge. Control actions refer to variables that can be manipulated in the system, such as temperature settings and the switch status of the device. These variables correspond to the actuator setpoints (Actuator Setpoints) in the EnergyPlus software. State information refers to state variables, that is, information obtained by the system through sensors or other means. The reward function represents the optimization goal of the system and is a specific decomposition of the optimization goal. Expert knowledge is used to guide and constrain the exploration behavior of the agent, thereby improving the convergence speed and training stability of the algorithm.

[0099] In the Simulation Process section, the simulation information and the information defined in the Define DRL Algorithm section will be used. The information from the Define DRL Algorithm section will be passed to the Python environment for the control algorithm. The control action information will be converted into actuator setpoints that can be recognized by EnergyPlus. The simulation information will be transferred to EnergyPlus to complete the simulation preparation.

[0100] In this embodiment, in step S3, the real-time dynamic control model of the building HVAC system is converted into a Markov decision process model;

[0101] The core elements of the Markov decision process model include: state space, action space, state transition probability, reward and discount factor;

[0102] The state space includes all potential states that the agent may exist in, and the expression of the state space is:

[0103] S={s0,s1,...,s t ,s t+1 ,...}

[0104] Where S is the state space; s t is the state of the agent at time t;

[0105] The action space includes the set of all possible control actions that the agent can perform, and the expression of the action space is:

[0106] A={a0,a1,...,a n}

[0107] Where A is the action space; a n Indicates a specific control action;

[0108] The state transition probability represents the possibility of transition between states, reflecting the agent's transition from s t Transfer to t+1 process; this process is realized through the real-time dynamic model in EnergyPlus simulation.

[0109] The reward is artificially defined and represents the set of all possible rewards that the agent may obtain. The expression of the reward is:

[0110] R={r0,r1,...,r t ,...,}

[0111] In the formula, R is the reward; r t The agent is in state s at time t t And perform action a t Rewards received when

[0112] The discount factor is used to balance the importance of immediate rewards and future rewards.

[0113] In this embodiment, it is crucial to model the real-time dynamic control problem of the building HVAC system as an MDP. The state variables are initially divided into six categories: time, personnel, temperature, relative humidity, electricity, and weather. The state "time" includes "day type" and "current hour". The state "personnel" includes the number of personnel in each thermal zone. The state "temperature" includes the temperature of each thermal zone. The state "relative humidity" includes the relative humidity information of each thermal zone. The state "electricity" includes the energy consumption of lights, equipment, boilers, chillers, HVAC systems in each thermal zone, and the entire building.

[0114] In the modeling of control actions, the control actions are defined as the heating and cooling temperature set points for each thermal zone, as well as the on / off states of the boiler and chiller. Given that the main goal of the building HVAC control system is to minimize costs while maintaining thermal comfort, the reward function is constructed as a weighted sum of the thermal comfort reward and the electricity price reward, as shown in the following formula:

[0115] f=min(w1*PMV reward +w2*Electricity reward )

[0116] Where f is the reward function corresponding to the optimization objective; w1 is the weighted coefficient of thermal comfort reward; w2 is the weighted coefficient of electricity price reward; PMVreward Reward for thermal comfort; Electricity reward Incentive for electricity prices.

[0117] The predicted mean voting (PMV) model proposed by Fanger is a widely used thermal comfort assessment tool for various building types. Based on this, an MDP for real-time dynamic control of building HVAC systems can be constructed, such as Figure 4 As shown in Figure 2, the subscript 0 represents the initial state, and the subscript t represents the final state.

[0118] In this MDP, the goal of real-time dynamic control of the building HVAC system is to maximize the cumulative reward, as shown in the following equation:

[0119] R t =r t +γr t+1 +γ 2 r t+2 +...

[0120] In the formula, γ is the discount factor.

[0121] The real-time dynamic control of the building HVAC system should be able to implement appropriate control actions based on any given state to maximize the cumulative reward. In this process, the agent adopts the optimal strategy, denoted by π * Accordingly, the general strategy is defined as π. In the optimal strategy π * Under this condition, the cumulative reward of the building HVAC control system is maximized, as shown in the following formula:

[0122]

[0123] Q(s,a) represents the action-value function, that is, in state s t Next, perform action a t The expected cumulative reward after .

[0124] In this embodiment, in step S4, a deep reinforcement learning algorithm EB-TD3 is constructed based on set expert rules; according to the set optimization goals, the Markov decision process model is iteratively optimized by the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

[0125] Specifically, the control process of the building HVAC real-time control system on indoor temperature, thermal comfort and electricity cost is transformed into a Markov decision process, the basis for implementing the DRL algorithm is constructed, and a deep reinforcement learning algorithm based on expert rules is proposed to achieve the best control effect.

[0126] In this embodiment, according to the set optimization goal, the Markov decision process model is iteratively optimized through the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

[0127] Specifically, the iterative optimization steps are:

[0128] S41, initializing the parameters of the Actor network and the Critic network; based on the same parameters, constructing a target network corresponding to the Actor network and the Critic network;

[0129] Specifically, initialize the Actor network and the Critic network μ(s|θ μ ) and Q(s,a|θ Q ), where θ μ and θ Q Represents the parameters of the neural network, and at the same time establishes its corresponding target network μ′ and Q′, whose initial parameter values ​​are the same as those of the Actor and Critic networks. Finally, an experience playback buffer D is constructed to store the experience collected by the agent.

[0130] S42, obtaining the current state through the Actor network, and determining a control action according to the current state; adding noise when determining the control action to enhance the exploration ability of the intelligent agent;

[0131] Specifically, the expression of the control action is:

[0132] a t =μ(s t |θ μ )+N t

[0133] In the formula, a t To control the action; N t For noise.

[0134] S43, after each control action is executed, the environment provides a reward to the agent and transitions to the next state; the quadruple is stored in the experience replay buffer until the round ends;

[0135] Specifically, after each action is performed, the environment will provide the agent with a reward r and transition to the next state s t+1 , the quaternion (s t ,a t ,r t ,s t+1 ) is stored in the experience replay buffer D until the end of the round.

[0136] S44, sampling a small batch of experience from the experience playback buffer; and obtaining the q value of the target network through the small batch of experience calculation;

[0137] Specifically, the q-value calculation formula of the target network is:

[0138]

[0139] In the formula, y represents the target network Q value, r represents the immediate reward, and γ represents the discount factor. Represents the output of the target Critic network, θ' i is the parameter of the i-th Critic target network, π φ' is the Actor target network, φ' represents the network parameters, s' represents the state at the next moment, and ∈ represents the noise used for smoothing. Indicates the minimum value between the two networks.

[0140] S45, updating the Critic network according to the loss function of the Critic network;

[0141] Specifically, the loss function of the Critic network is:

[0142]

[0143] Where L is the loss function; N is the size of the batch taken out; y i is the q value of the target network in the i-th batch.

[0144] S46. Based on the policy gradient theorem, the Actor network is updated;

[0145] Specifically, the update formula is:

[0146]

[0147] In the formula, is the gradient of the loss function; is the gradient of the actor network; is the gradient of the critic network.

[0148] S47: According to the soft update strategy, the target network corresponding to the Actor network and the Critic network is updated.

[0149] Specifically, the update formula is:

[0150] θ Q′ ←τθ Q +(1-τ)θ Q′

[0151] θμ′ ←τθ μ +(1-τ)θ μ′

[0152] Where τ represents the soft update factor.

[0153] In this embodiment, two expert rules are introduced in step S42. The first expert rule processes the reward, that is, using the PMV index of human thermal comfort, and adopts a segmented reward. The range of [-0.5, 0.5] is the most comfortable, the range of abs [0.5, 1.5] is acceptable, and the range above this range is unacceptable. The expert rule establishes a buffer zone between the comfortable temperature and the extreme temperature, so that the reward function has an adaptive characteristic.

[0154] The second expert rule acts on action selection. In unoccupied buildings, more relaxed rules are implemented to allow the agent to freely explore to improve adaptability. In the case of occupancy, the expert rules are strictly enforced. Specifically, when the temperature deviates from the comfort level by a large amount, the boiler and chiller are started, and the temperature setpoint is adjusted to restore the indoor temperature to an acceptable thermal comfort range. These expert rules ensure that the agent guarantees thermal comfort while maintaining the ability to explore. In addition, given the uncertainty of the meeting room schedule and the sparseness of the reward distribution, a fixed temperature setpoint is used during both occupied and unoccupied periods.

[0155] In a possible embodiment, a specific simulation implementation example is provided as follows:

[0156] Parameter settings:

[0157] Assume that the simulation area is a 5km*5km building in Tianjin. The building dimensions are 17 meters, 12 meters, and 6 meters, and include three private offices, two open offices, and a conference room. These areas are divided into four thermal zones. The building envelope includes exterior walls, windows, doors, and interior walls. Its materials and thermal properties are shown in Table 1:

[0158] Material Name Thermal conductivity (W / m·K) Specific heat capacity (J / kg·K) Heavy concrete 1.95 900 Polyurethane 0.0245 1590 Concrete 1.11 920 Wood 0.15 1630 External windows 0.9 /

[0159] Table 1 Material parameters of building envelope

[0160] Single-person offices are designed to accommodate one occupant. The population density in open offices is 3.5 people / m2, and in meeting rooms is 2 people / m2. During the workday, the occupancy rate of these spaces gradually increases from 0.3 to a peak of 0.95 as people arrive, and then drops to 0 when no one is using them. Meeting rooms are used irregularly, with an average of 3 to 4 meetings per week, each lasting 1 to 2 hours.

[0161] Regarding the HVAC system, the rated power of the boiler and chiller is 80kW and 40kW respectively, and their coefficient of performance (COP) is 0.9 and 4.2 respectively. The boiler and chiller are both electrically driven. The supply and return water temperatures as well as the settings of the fans and cooling towers are optimized to improve the system performance. The comparison algorithms in the experiment include the rule-based control algorithm and the benchmark DRL algorithm TD3. The experimental scenarios cover typical winter and summer months.

[0162] Analysis of calculation results:

[0163] The algorithm effect is measured by total cost, electricity price cost, thermal comfort cost, energy consumption, and carbon dioxide emissions. The calculation results are shown in Table 2:

[0164] method season Total cost (yuan) Electricity fee (yuan) Comfort cost (yuan) Energy consumption (kwh) CO2(kg) Rule 19975.74 16493.92 3481.82 18208.07 18153.45 TD3 winter 16597.14 14124.33 2472.81 14119.68 14157.08 EB-TD3 16011.35 14016.61 1994.74 14052.78 14010.62 Rule 3497.16 2885.65 611.51 2793.96 2710.14 TD3 summer 2990.02 2730.17 259.85 2729.53 2647.64 EB-TD3 2618.54 2409.41 209.13 2517.67 2442.14

[0165] From Table 2, we can see that:

[0166] In winter, compared with the rule-based control algorithm, the EB-TD3 algorithm reduced the total cost by 19.84%, improved thermal comfort by 42.71%, reduced HVAC system energy consumption by 22.82%, and reduced monthly CO2 emissions by 4,142.83 kg. Compared with the baseline TD3 algorithm, the EB-TD3 algorithm reduced the total cost by 3.53%, reduced electricity cost by 0.76%, improved thermal comfort by 19.33%, reduced energy consumption by 1.03%, and reduced monthly CO2 emissions by 146.46 kg.

[0167] In summer, compared with the rule-based control method, the EB-TD3 algorithm reduced the overall cost by 25.12%, reduced the electricity cost by 16.50%, improved thermal comfort by 65.80%, reduced the building HVAC system energy consumption by 9.89%, and reduced the monthly carbon dioxide emissions by 268.00 kg. Compared with the baseline TD3 algorithm, the proposed method achieved a 12.42% reduction in total cost, a 11.75% reduction in electricity cost, a 19.52% reduction in thermal comfort cost, a 7.76% reduction in HVAC system energy consumption, and a 205.50 kg reduction in monthly carbon dioxide emissions.

[0168] In summary, the present invention constructs an indoor temperature change model; describes the indoor temperature change through the indoor temperature change model; constructs a building thermodynamic model based on the indoor temperature change model; constructs a real-time dynamic control model of the building HVAC system based on the building thermodynamic model, combined with the EnergyPlus and Python joint simulation framework; converts the real-time dynamic control model of the building HVAC system into a Markov decision process model; constructs a deep reinforcement learning algorithm EB-TD3 based on the set expert rules; and iteratively optimizes the Markov decision process model through the deep reinforcement learning algorithm EB-TD3 according to the set optimization target, so as to realize the adaptive real-time dynamic control of the building HVAC system. The present invention combines the deep reinforcement learning algorithm with expert knowledge, can realize the adaptive adjustment of the HVAC system under complex environment and random disturbance, improve the flexibility and stability of the control strategy, and can achieve the balance between energy saving and comfort, optimize the environmental parameters such as indoor temperature and humidity, ensure that the indoor thermal comfort is maintained while saving energy, and meet the dual needs of building energy efficiency and user comfort. In terms of control strategy, by introducing expert knowledge, the temperature fluctuations caused by the exploration process are reduced, the convergence speed and training stability of the algorithm are improved, and long-term stable operation is ensured. The proposed control strategy can cope with the complex changes in the internal and external environment of the building, such as seasonal fluctuations, climate change, etc., and dynamically adjust the operating parameters of the HVAC system to achieve the best control effect. By combining with simulation platforms such as EnergyPlus, the feasibility of the control method of the present invention in practical applications is verified, and it can effectively improve the operating efficiency of the building HVAC system.

[0169] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or a server. The method of the present embodiment can also be applied in a distributed scenario and completed by multiple devices cooperating with each other. In the case of such a distributed scenario, one of the multiple devices can only perform one or more steps in the method of the embodiment of the present disclosure, and the multiple devices will interact with each other to complete the described method.

[0170] It should be noted that the above describes some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the above embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0171] Example 2

[0172] See also Figure 5Embodiment 2 of the present invention also provides a building HVAC adaptive coordination control device based on deep reinforcement learning, including:

[0173] Building thermodynamic model construction module 001, used to construct an indoor temperature change model; describe the indoor temperature change through the indoor temperature change model; and construct a building thermodynamic model based on the indoor temperature change model;

[0174] A building HVAC system real-time dynamic control model construction module 002 is used to construct a building HVAC system real-time dynamic control model based on the building thermodynamic model and in combination with the EnergyPlus and Python joint simulation framework;

[0175] A Markov decision process model conversion module 003, used to convert the real-time dynamic control model of the building HVAC system into a Markov decision process model;

[0176] The deep reinforcement learning algorithm EB-TD3 construction and iterative optimization module 004 is used to construct the deep reinforcement learning algorithm EB-TD3 based on the set expert rules; according to the set optimization goals, the Markov decision process model is iteratively optimized through the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

[0177] In this embodiment, in the building thermodynamic model construction module 001, the expression of the indoor temperature change model is:

[0178]

[0179] Where δt is the time increment; is the temperature of the current time step; and are the temperatures of the first three time steps respectively; O represents the approximation order, i.e. the error term.

[0180] In this embodiment, in the building thermodynamic model construction module 001, the expression of the building thermodynamic model is:

[0181]

[0182] C h =ρ air ×C ρ ×C T

[0183] In the formula, C h is the heat capacity of the indoor air; ρ air is the air density; C ρ is the specific heat capacity of air; C T is the volume of air; Cp is the specific heat capacity at constant pressure; N l is the total number of hot zones; N s is the total number of heat exchange surfaces; N z is the total number of fluid branches; h i is the surface heat transfer coefficient; A i is the surface area; T si is the room surface temperature; T h is the room air temperature; is the air mass flow rate entering from the adjacent room; T hi is the temperature of the air in the adjacent room; is the mass flow rate of infiltrated external air; T ∞ is the outside air temperature; is the sum of external heat sources; is the sum of the internal heat sources of the system.

[0184] In this embodiment, in the Markov decision process model conversion module 003, the core elements of the Markov decision process model include: state space, action space, state transition probability, reward and discount factor;

[0185] The expression of the state space is:

[0186] S={s0,s1,...,s t ,s t+1 ,...}

[0187] Where S is the state space; s t is the state of the agent at time t;

[0188] The expression of the action space is:

[0189] A={a0,a1,...,a n}

[0190] Where A is the action space; a n Indicates a specific control action;

[0191] The state transition probability represents the possibility of transition between states, reflecting the agent's transition from s t Transfer to t+1 process;

[0192] The reward expression is:

[0193] R={r0,r1,...,r t ,...,}

[0194] In the formula, R is the reward; r t The agent is in state s at time t tAnd perform action a t Rewards received when

[0195] The discount factor is used to balance the importance of immediate rewards and future rewards.

[0196] In this embodiment, in the deep reinforcement learning algorithm EB-TD3 construction and iterative optimization module 004, the iterative optimization submodule includes:

[0197] The network initialization submodule 041 is used to initialize the parameters of the Actor network and the Critic network; based on the same parameters, construct a target network corresponding to the Actor network and the Critic network;

[0198] The control action determination submodule 042 is used to obtain the current state through the Actor network and determine the control action according to the current state; add noise when determining the control action to enhance the exploration ability of the intelligent agent;

[0199] Iterative optimization submodule 043, used for providing rewards to the agent and transitioning to the next state after each control action is executed; storing the quadruple in the experience replay buffer until the round ends;

[0200] The q-value acquisition submodule 044 of the target network is used to sample small batches of experience from the experience playback buffer; and obtain the q-value of the target network through the small batch experience calculation;

[0201] Critic network updating submodule 045, used to update the Critic network according to the loss function of the Critic network;

[0202] Actor network update submodule 046, used to update the Actor network based on the policy gradient theorem;

[0203] The target network updating submodule 047 is used to update the target network corresponding to the Actor network and the Critic network according to the soft update strategy.

[0204] It should be noted that the information interaction, execution process and other contents between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and will not be repeated here.

[0205] Example 3

[0206] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code of a building HVAC adaptive coordination control method based on deep reinforcement learning is stored, and the program code includes instructions for executing the building HVAC adaptive coordination control method based on deep reinforcement learning of embodiment 1 or any possible implementation thereof.

[0207] The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0208] Example 4

[0209] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0210] The processor and the memory communicate with each other via a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the building HVAC adaptive coordination control method based on deep reinforcement learning of Example 1 or any possible implementation thereof.

[0211] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor implemented by reading software codes stored in a memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.

[0212] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from a computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from a website site, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, server or data center.

[0213] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing system, they can be concentrated on a single computing system, or distributed on a network composed of multiple computing systems, and optionally, they can be implemented by a program code executable by a computing system, so that they can be stored in a storage system and executed by the computing system, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0214] Although the present invention has been described in detail above by general description and specific embodiments, it is obvious to those skilled in the art that some modifications or improvements can be made to the present invention. Therefore, these modifications or improvements made without departing from the spirit of the present invention all belong to the scope of protection claimed by the present invention.

Claims

1. A building HVAC adaptive coordination control method based on deep reinforcement learning, characterized in that: include: Constructing an indoor temperature change model; describing the indoor temperature change by using the indoor temperature change model; Constructing a building thermodynamic model based on the indoor temperature change model; Based on the building thermodynamic model, combined with EnergyPlus and Python co-simulation framework, a real-time dynamic control model of the building HVAC system is constructed; Converting the real-time dynamic control model of the building HVAC system into a Markov decision process model; Based on the set expert rules, we build the deep reinforcement learning algorithm EB-TD3; According to the set optimization goal, the Markov decision process model is iteratively optimized through the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

2. The building HVAC adaptive coordination control method based on deep reinforcement learning according to claim 1 is characterized in that: The expression of the indoor temperature change model is: Where δt is the time increment; is the temperature of the current time step; and are the temperatures of the first three time steps respectively; O represents the approximation order, i.e., the error term, indicating that when δt is very small, the approximate error of the derivative is approximately equal to δt 3 Same level.

3. The building HVAC adaptive coordination control method based on deep reinforcement learning according to claim 2 is characterized in that: The expression of the building thermodynamic model is: C h =ρ air ×C ρ ×C T In the formula, C h is the heat capacity of the indoor air; ρ air is the air density; C ρ is the specific heat capacity of air; C T is the volume of air; C p is the specific heat capacity at constant pressure; N l is the total number of hot zones; N s is the total number of heat exchange surfaces; N z is the total number of fluid branches; h i is the surface heat transfer coefficient; A i is the surface area; T si is the room surface temperature; T h is the room air temperature; is the air mass flow rate entering from the adjacent room; T hi is the temperature of the air in the adjacent room; is the mass flow rate of infiltrated external air; T ∞ is the outside air temperature; is the sum of external heat sources; is the sum of the internal heat sources of the system.

4. The building HVAC adaptive coordination control method based on deep reinforcement learning according to claim 3 is characterized in that: The core elements of the Markov decision process model include: state space, action space, state transition probability, reward and discount factor; The expression of the state space is: S={s0,s1,...,s t ,s t+1 ,...} Where S is the state space; s t is the state of the agent at time t; The expression of the action space is: A={a0,a1,...,a n } Where A is the action space; a n Indicates a specific control action; The state transition probability represents the possibility of transition between states, reflecting the agent's transition from s t Transfer to t+1 process; The reward expression is: R={r0,r1,...,r t ,...,} In the formula, R is the reward; r t The agent is in state s at time t t And perform action a t Rewards received when The discount factor is used to balance the importance of immediate rewards and future rewards.

5. The building HVAC adaptive coordination control method based on deep reinforcement learning according to claim 4 is characterized in that: In the process of iteratively optimizing the Markov decision process model by the deep reinforcement learning algorithm EB-TD3 according to the set optimization goal, the iterative optimization steps are: Initialize the parameters of the Actor network and the Critic network; based on the same parameters, construct a target network corresponding to the Actor network and the Critic network; Acquire the current state through the Actor network, and determine the control action according to the current state; add noise when determining the control action to enhance the exploration ability of the intelligent agent; After each control action is executed, the environment provides a reward to the agent and transitions to the next state; the quadruple is stored in the experience replay buffer until the round ends; Sampling a small batch of experience from the experience playback buffer; and obtaining a q value of the target network through the small batch of experience calculation; According to the loss function of the Critic network, the Critic network is updated; Based on the policy gradient theorem, the Actor network is updated; According to the soft update strategy, the target network corresponding to the Actor network and the Critic network is updated.

6. A building HVAC adaptive coordination control device based on deep reinforcement learning, adopting a building HVAC adaptive coordination control method based on deep reinforcement learning according to any one of claims 1 to 5, characterized in that: include: Building thermal dynamics model building module, used to build indoor temperature change model; Describing the indoor temperature change by the indoor temperature change model; Constructing a building thermodynamic model based on the indoor temperature change model; A building HVAC system real-time dynamic control model construction module, which is used to construct a building HVAC system real-time dynamic control model based on the building thermodynamic model and in combination with the EnergyPlus and Python joint simulation framework; A Markov decision process model conversion module, used to convert the real-time dynamic control model of the building HVAC system into a Markov decision process model; The deep reinforcement learning algorithm EB-TD3 construction and iterative optimization module is used to construct the deep reinforcement learning algorithm EB-TD3 based on the set expert rules; According to the set optimization goal, the Markov decision process model is iteratively optimized through the deep reinforcement learning algorithm EB-TD3 to achieve adaptive real-time dynamic control of the building HVAC system.

7. The building HVAC adaptive coordination control device based on deep reinforcement learning according to claim 6 is characterized in that: In the building thermodynamic model construction module, the expression of the indoor temperature change model is: Where δt is the time increment; is the temperature of the current time step; and are the temperatures of the first three time steps respectively; O represents the approximation order, i.e. the error term.

8. The building HVAC adaptive coordination control device based on deep reinforcement learning according to claim 7 is characterized in that: In the building thermodynamic model construction module, the expression of the building thermodynamic model is: In the formula, C h is the heat capacity of the indoor air; ρ air is the air density; C ρ is the specific heat capacity of air; C T is the volume of air; C p is the specific heat capacity at constant pressure; N l is the total number of hot zones; N s is the total number of heat exchange surfaces; N z is the total number of fluid branches; h i is the surface heat transfer coefficient; A i is the surface area; T si is the room surface temperature; T h is the room air temperature; is the air mass flow rate entering from the adjacent room; T hi is the temperature of the air in the adjacent room; is the mass flow rate of infiltrated external air; T ∞ is the outside air temperature; is the sum of external heat sources; is the sum of the internal heat sources of the system.

9. The building HVAC adaptive coordination control device based on deep reinforcement learning according to claim 8, characterized in that: In the Markov decision process model conversion module, the core elements of the Markov decision process model include: state space, action space, state transition probability, reward and discount factor; The expression of the state space is: S={s0,s1,...,s t ,s t+1 ,...} Where S is the state space; s t is the state of the agent at time t; The expression of the action space is: A={a0,a1,...,a n } Where A is the action space; a n Indicates a specific control action; The state transition probability represents the possibility of transition between states, reflecting the agent's transition from s t Transfer to t+1 process; The reward expression is: R={r0,r1,...,r t ,...,} In the formula, R is the reward; r t The agent is in state s at time t t And perform action a t Rewards received when The discount factor is used to balance the importance of immediate rewards and future rewards.

10. The building HVAC adaptive coordination control device based on deep reinforcement learning according to claim 9, characterized in that: In the deep reinforcement learning algorithm EB-TD3 construction and iterative optimization module, the iterative optimization submodule includes: The network initialization submodule is used to initialize the parameters of the Actor network and the Critic network; based on the same parameters, a target network corresponding to the Actor network and the Critic network is constructed; A control action determination submodule is used to obtain the current state through the Actor network and determine the control action according to the current state; when determining the control action, noise is added to enhance the exploration ability of the intelligent agent; An iterative optimization submodule, which is used for providing a reward to the agent and transitioning to the next state after each control action is executed; storing the quadruple in the experience replay buffer until the round ends; The q-value acquisition submodule of the target network is used to sample small batch experience from the experience playback buffer; and obtain the q-value of the target network through the small batch experience calculation; A critic network update submodule, used to update the critic network according to the loss function of the critic network; Actor network update submodule, used to update the Actor network based on the policy gradient theorem; The target network update submodule is used to update the target network corresponding to the Actor network and the Critic network according to the soft update strategy.

Citation Information

Patent Citations

  • Heat pump-floor heating system control method considering user comfort and building heat storage

    CN110543713A

  • Indoor thermal environment control method based on RC model and deep reinforcement learning

    CN116734424A

  • Building integrated energy system optimization method based on deep reinforcement learning

    CN118381729A

  • Methods and systems for training HVAC control using simulated and real experience data

    US20210190364A1

  • Methods and systems for training HVAC control using surrogate model

    US20210191342A1