A bulldozer distance-based imitation reinforcement learning building air conditioning system energy management method and storage medium
Patent Information
- Application Number
- CN202511778455.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-11-28
AI Technical Summary
[0005]因此,本发明解决的技术问题是:现有的强化学习和模仿学习方法存在样本效率低、奖励塑造敏感、延迟反馈,以及如何将模仿学习与强化学习相结合的问题
[0017] The beneficial effects of this invention are as follows: Compared with reinforcement learning alone, the bulldozer distance-based imitation reinforcement learning method provided by this invention exhibits better convergence speed and sample efficiency during training. It also has better reward function and cost-saving effect under the same training steps. In this framework, the bulldozer distance-based imitation learning has an advantage in stability compared with adversarial generative imitation learning and can reduce the burden of hyperparameter tuning. During the online training process, the imitation reinforcement learning can overcome the dependence on expert actions through autonomous exploration in the later stage. Compared with the optimization scheduling strategy with weather and personnel prediction errors as expert actions, the imitation reinforcement learning strategy can produce additional cost-saving effect and maintain the indoor temperature within a comfortable range.
Smart Images

Figure CN121960830B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of building energy system technology, specifically to a building air conditioning system energy management method and storage medium based on bulldozer distance-based imitation reinforcement learning. Background Technology
[0002] The intermittent and unpredictable nature of renewable energy sources such as wind and solar power leads to a mismatch between power generation fluctuations and grid demand, posing a severe challenge to grid reliability. The multi-energy storage characteristics of cold storage building energy systems, including buildings and water tanks, make their energy flexibility an important solution to the above problems. In recent years, with the continuous development of advanced computing technologies such as artificial intelligence, they have also been introduced into the field of building control and energy management to replace traditional manual or rule-based control methods and improve the energy efficiency and flexibility of buildings.
[0003] Reinforcement learning methods, with their advantages of being model-free and dynamic learning, have been applied to building energy management systems. However, the practical application of reinforcement learning in building energy still faces challenges, including low sample efficiency and unstable initial training, dependence on large amounts of training data, sensitivity to reward shaping, and delayed feedback. To address these issues, some research has introduced imitation learning, where agents mimic expert actions (typically rule-based control results or model-based optimization results), reducing reliance on environmental feedback rewards. This avoids the difficulty of reward function design, lowers exploration risk, improves sample efficiency, and accelerates convergence. However, imitation learning still relies on expert knowledge. In the building energy field, due to various uncertainties, model-based optimization and rule-based strategies may still be suboptimal. Existing research lacks methods for combining online training of imitation learning and reinforcement learning, enabling agents to transition to autonomous exploration when imitating suboptimal operating strategies as expert knowledge, and further optimize control actions. This is significant for improving algorithm training efficiency, coping with uncertainty, and achieving adaptive optimization, thereby enhancing the widespread application of artificial intelligence algorithms in the building energy field. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by this invention is that existing reinforcement learning and imitation learning methods suffer from low sample efficiency, sensitivity to reward shaping, delayed feedback, and the problem of how to combine imitation learning with reinforcement learning.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: an imitation reinforcement learning method based on bulldozer distance, including an imitation learning stage in which an imitation learning method based on the original bulldozer distance is introduced; and a reinforcement learning stage in which a dynamic reward function is used to transition the reward composition in agent training from imitation learning to reinforcement learning.
[0007] As a preferred embodiment of the bulldozer distance-based imitation reinforcement learning method of the present invention, the imitation learning stage includes: representing the agent's reward function as the inverse of the original bulldozer distance between the agent's state-action distribution and the expert's state-action distribution, thereby reducing the gap between the agent and the expert during training and establishing an imitation learning method based on the original bulldozer distance.
[0008] As a preferred embodiment of the bulldozer distance-based imitation reinforcement learning method described in this invention, wherein: the original bulldozer distance includes, and the bulldozer distance measure is derived from the distribution... Move into distribution The minimum cost required at that time is expressed as: , in, To minimize the cost of movement, for and The joint distribution set is equivalent to the movement strategy, which represents the movement from position. Move to The probability quality, From arrive Requires moving a distance, For a space, all possibilities A set of point pairs.
[0009] As a preferred embodiment of the bulldozer distance-based imitation reinforcement learning method described in this invention, the agent's state-action distribution includes the empirical distribution of all state-action pairs generated during the agent's interaction with the environment within the imitation learning period. In Markov decision processes, the state-action distribution is defined as follows: , in, For the distribution of agent state and action, To mimic the learning cycle, For the sake of Dirac distribution centered on, For expert state action distribution, This represents the number of expert samples.
[0010] As a preferred embodiment of the bulldozer distance-based imitation reinforcement learning method described in this invention, the reinforcement learning stage includes: a dynamic reward function that smoothly transitions the agent's training focus from imitation learning to autonomous exploration in reinforcement learning; deep exploration of strategies superior to expert actions; and definition of an imitation reinforcement learning algorithm.
[0011] As a preferred embodiment of the bulldozer distance-based imitation reinforcement learning method described in this invention, the dynamic reward function includes a combination of imitation learning reward and environmental reward, and the dynamic reward function is expressed as follows: , in, For dynamic reward function, To encourage learning through rewards, For environmental rewards, i.e., the reward function of reinforcement learning, For activation function, For an experience value, is a coefficient.
[0012] Another objective of this invention is to provide an energy management method for building air conditioning systems based on imitation reinforcement learning, which can solve the problem of current imitation learning and reinforcement learning methods relying on expert knowledge by transitioning from imitation learning to reinforcement learning.
[0013] As a preferred embodiment of the energy management method for building air conditioning systems based on imitation reinforcement learning described in this invention, the method includes: constructing a day-ahead optimization scheduling strategy based on a gray box model of the building air conditioning system; establishing an imitation reinforcement learning method based on bulldozer distance, as described in any one of claims 1 to 6; and constructing an imitation reinforcement learning energy management method for building air conditioning systems.
[0014] As a preferred embodiment of the energy management method for building air conditioning systems based on imitation reinforcement learning described in this invention, the construction of a day-ahead optimization scheduling strategy based on a gray box model of the building air conditioning system includes: establishing an air conditioning system energy consumption model based on a building RC model, a chiller efficiency model, and a water tank energy storage model; identifying parameters based on actual operating data; setting an optimization objective of minimizing operating costs; combining time-of-use electricity price information, energy balance, operational mutual exclusion, and comfort temperature constraints; solving a mixed integer programming problem; and outputting a day-ahead optimization scheduling strategy for system operation.
[0015] As a preferred embodiment of the energy management method for building air conditioning systems based on imitation reinforcement learning described in this invention, the method for constructing an energy management method for building air conditioning systems based on imitation reinforcement learning includes: using the day-ahead optimization scheduling strategy based on the output as expert knowledge for the imitation reinforcement learning method of bulldozer distance; designing the reward function for reinforcement learning based on the actual system and the hyperparameters of the imitation reinforcement algorithm; training the imitation reinforcement learning agent online in an actual or simulated environment; and gradually improving the control performance through multi-day cycles.
[0016] Another object of the present invention is to provide a building air conditioning system energy management storage medium based on imitation reinforcement learning, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the steps of an imitation reinforcement learning method based on bulldozer distance are implemented.
[0017] The beneficial effects of this invention are as follows: Compared with reinforcement learning alone, the bulldozer distance-based imitation reinforcement learning method provided by this invention exhibits better convergence speed and sample efficiency during training. It also has better reward function and cost-saving effect under the same training steps. In this framework, the bulldozer distance-based imitation learning has an advantage in stability compared with adversarial generative imitation learning and can reduce the burden of hyperparameter tuning. During the online training process, the imitation reinforcement learning can overcome the dependence on expert actions through autonomous exploration in the later stage. Compared with the optimization scheduling strategy with weather and personnel prediction errors as expert actions, the imitation reinforcement learning strategy can produce additional cost-saving effect and maintain the indoor temperature within a comfortable range. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is an overall schematic diagram of a bulldozer distance-based imitation reinforcement learning method provided in Embodiment 1 of the present invention.
[0020] Figure 2 The graph shows the change of reward function during the training process of different algorithms for an energy management method for a building air conditioning system based on bulldozer distance imitation reinforcement learning, as provided in Embodiment 2 of the present invention.
[0021] Figure 3 This is an energy storage and release map of the baseline scenario strategy for an energy management method for a building air conditioning system based on bulldozer distance imitation reinforcement learning, provided in Embodiment 2 of the present invention.
[0022] Figure 4 The energy storage and release diagram a is for the optimization scenario strategy of the energy management method of building air conditioning system based on bulldozer distance imitation reinforcement learning provided in Embodiment 2 of the present invention.
[0023] Figure 5 The energy storage and release diagram b is for the optimized scenario strategy of the energy management method for a building air conditioning system based on bulldozer distance imitation reinforcement learning provided in Embodiment 2 of the present invention.
[0024] Figure 6 The energy storage and release diagram c is for the optimized scenario strategy of the energy management method for building air conditioning system based on bulldozer distance imitation reinforcement learning provided in Embodiment 2 of the present invention.
[0025] Figure 7 The energy storage and release diagram d is provided for the optimization scenario strategy of the energy management method of building air conditioning system based on bulldozer distance imitation reinforcement learning provided in Embodiment 2 of the present invention.
[0026] Figure 8 This is an energy storage and release diagram of the imitation reinforcement learning strategy for an energy management method of a building air conditioning system based on bulldozer distance, provided in Embodiment 2 of the present invention.
[0027] Figure 9 Load diagrams for different strategies of a building air conditioning system energy management method based on bulldozer distance imitation reinforcement learning provided in Embodiment 2 of the present invention during peak electricity price periods.
[0028] Figure 10 The different strategy operation effects of the energy management method for building air conditioning system based on bulldozer distance and imitation reinforcement learning provided in Embodiment 2 of the present invention are converted into reinforcement learning reward value maps.
[0029] Figure 11 This is a diagram illustrating the operating costs of different strategies for an energy management method for a building air conditioning system based on bulldozer distance-based imitation reinforcement learning, as provided in Embodiment 2 of the present invention.
[0030] Figure 12 This is an overall flowchart of an energy management method for a building air conditioning system based on bulldozer distance imitation reinforcement learning, provided in Embodiment 3 of the present invention. Detailed Implementation
[0031] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0032] Example 1, referring to Figure 1 As an embodiment of the present invention, an imitation reinforcement learning method based on bulldozer distance is provided. The method includes two stages: an imitation learning stage 100 and a reinforcement learning stage 200. In the imitation learning stage 100, an imitation learning method based on the original bulldozer distance (Wasserstein distance) is introduced. By representing the agent's reward function as the inverse of the original bulldozer distance between the agent's state-action distribution and the expert's state-action distribution, the gap between the agent and the expert is reduced during the training process. In the reinforcement learning stage 200, a dynamic reward function is used to transition the reward composition during the agent's training process from being dominated by imitation learning to being dominated by reinforcement learning.
[0033] Furthermore, the Wasserstein distance measures the distance between data distributions. Move into distribution The minimum cost required at that time is expressed as: , in, To minimize the cost of movement, for and The joint distribution set is equivalent to the movement strategy, which represents the movement from position. Move to The probability quality, From arrive Requires moving a distance, For a space (e.g., the space of real numbers). All possible A set of point pairs.
[0034] It can be seen that the Wasserstein distance is the minimum value of the integral of the movement mass and the movement distance, that is, the minimum cost of movement. In Markov decision processes, the state-action distribution is defined as: , in, For the distribution of agent state and action, To mimic the learning cycle, For the sake of Dirac distribution centered on, For expert state action distribution, This represents the number of expert samples.
[0035] Therefore, the Wasserstein distance is defined as: , , in, For Wasserstein distance, For the distribution of agent state and action, For expert state action distribution, To mimic the learning cycle, For the number of expert samples, For the agent in the first State-action pairs at each time step For the agent in time step state, For the agent in time step The action, The first one provided for experts A sample state-action pair, For experts A demonstration state, For experts A demonstration movement, Let it be a distance function. In this discrete distribution case, the size is A double random matrix, It is a power of the distance.
[0036] Each mass is The distribution center of the agent's state and actions, filled into each center to accommodate The minimum distance of the expert state action distribution for quality; the Wasserstein distance is the minimum value of the distribution movement distance, therefore it needs to be calculated from... The optimal movement strategy is determined as follows: , in, For the optimal movement strategy, For the agent in the first State-action pairs at each time step For the agent in time step state, For the agent in time step The action, The first one provided for experts A sample state-action pair, For experts A demonstration state, For experts A demonstration movement, This is the distance function.
[0037] To obtain the optimal movement strategy, we need to acquire the agent's state-action data throughout the entire training cycle. Since the agent cannot receive immediate rewards, a greedy strategy is used to approximate the optimal strategy, expressed as: , , in, For the greedy strategy An agent's state-action pair, The total number of sample points for the agent. These are expert sample points.
[0038] The meaning of the greedy strategy is that at each training time step ,Will Quality filling to expert state action distribution minimum cost The constraint distribution represents the mass of the movement. The mass of each expert state action point filled is no greater than [missing value]. The greedy strategy is an approximation of the optimal movement strategy, and its lower bound for solving distance is equal to the Wasserstein distance.
[0039] Using a greedy strategy, the agent's immediate reward at each training time step can be obtained, denoted as: , in, To imitate learning rewards, intelligent agents are rewarded. For coefficients, For the agent in the first State-action pairs at each time step For the agent in time step state, For the agent in time step The action, The first one provided for experts A sample state-action pair, For experts A demonstration state, For experts A demonstration movement, Let it be a distance function. This is a transmission plan under a greedy strategy.
[0040] The reward is the negative of the movement cost. By increasing the reward through training, the movement cost between distributions is gradually reduced, that is, the Wasserstein distance is reduced, so that the agent's actions gradually become closer to the expert's actions, thus achieving imitation learning.
[0041] It should be noted that the imitation learning algorithm is not limited to the Wasserstein distance-based imitation learning (PWIL) introduced above. It can also be replaced by other algorithms such as adversarial generative imitation learning (GAIL). The corresponding hyperparameters need to be modified. This invention uses the PWIL algorithm. Compared with the adversarial generative imitation learning (GAIL) algorithm, the PWIL algorithm does not require the establishment of a discriminator model. Instead, it uses Wasserstein distance instead, which reduces the training burden and the difficulty of hyperparameter selection. At the same time, it may improve training stability and is theoretically superior.
[0042] It should also be noted that, in order to transition the agent from imitation learning to reinforcement learning in the later stages of training, thereby avoiding the algorithm's over-reliance on expert actions and enabling it to explore policies superior to expert actions, the dynamic reward function of the Imitation Reinforcement Learning (ImRL) algorithm is defined as follows: , , in, For dynamic reward function, To encourage learning through rewards, For environmental rewards, i.e., the reward function of reinforcement learning, For activation function, For an experience value, is a coefficient.
[0043] The expression indicates that once the periodic reward reaches a certain threshold, the proportion of environmental reward is gradually increased to initiate the reinforcement learning autonomous exploration phase; otherwise, environmental reward has no effect and imitation learning reward is used instead.
[0044] Example 2, refer to Figures 2-11 As an embodiment of the present invention, an energy management method for building air conditioning systems based on bulldozer distance imitation reinforcement learning is provided. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.
[0045] First, the building's air conditioning system terminal adopts a variable air volume (VAV) system. On the energy system side, two chiller units are used as the rated power of the cold source, and each unit is equipped with a water pump. In addition, an energy storage tank is set up to store chilled water produced during the night when the electricity price is low, and release the cooling capacity during the high electricity price period to complete the load transfer.
[0046] An EnergyPlus-Modelica simulation model is built for the building. The EnergyPlus model is used to describe the building's thermal dynamics, while the Modelica Building library is used to model the air conditioning system. The two are jointly simulated using the spawn engine and are used to represent the actual building. The simulation is used to calculate air conditioning load data under the scheduling and control strategy. At the same time, the building and equipment simulation system is used to test and evaluate the effect of the control strategy.
[0047] To meet the need for real-time online training between the scheduling and control algorithm model and the case building, the co-simulation model and the Python computing language are connected through the FMU interface. The scheduling and control signals based on the PyTorch-based simulation reinforcement algorithm can be transmitted to the co-simulation model through the interface to change the building's operating state. At the same time, the operating state parameters of the co-simulation model can be transmitted to the algorithm through the interface, affecting the generation of the scheduling and control signals in the next time step.
[0048] The baseline strategy adopts a rule-based control method. Specifically, the average temperature of the water storage tank is reduced to 5°C during the night when the electricity price is low. Then, during the daytime when the electricity price is highest, the water temperature is released and does not exceed 14°C. Otherwise, the chiller is responsible for cooling, and the water supply temperature is set to 7°C. The water pump operates at the rated flow rate. During the daytime, the fan flow rate is controlled by PID to maintain the room temperature at 26°C.
[0049] The optimization scheduling model solves the mixed integer quadratic programming (MIQP) problem in Gurobi. During the solution process, Gaussian distribution errors are introduced for the room temperature To and the internal heat source Qint to represent the prediction uncertainty of internal and external disturbances to the building load during day-ahead scheduling. By introducing different errors, the optimization results will generate different control actions, thus generating multiple state-action trajectories. To is introduced with Gaussian distributions of 1.5℃, 1℃, and 0.5℃ as the mean and 1℃ as the variance; Qint is introduced with Gaussian distributions of 30%Qint, 20%Qint, and 10%Qint as the mean and 10%Qint as the standard deviation, respectively set as optimization scenarios 1, 2, and 3. The expert knowledge learned through imitation consists of the operating strategies under these three scenarios and their operating states in the environment. Simultaneously, optimization scenario 4, without meteorological and personnel prediction errors, is set as a comparison and reference strategy.
[0050] Table 1 shows the state and action variables of the Markov decision process in reinforcement learning and imitation learning. This study mainly controls the system's operating model, which has 6 discrete action variables with a control step size of 1 hour. The setpoint for the chiller's storage water temperature is set to 5℃, the setpoint for the chiller's supply water temperature is set to 7℃, and the water pump operates at its rated flow rate. The PID controller only controls the AHU fan to maintain the temperature at 22-26.5℃ when the temperature exceeds the limit; otherwise, it operates at the rated flow rate. The discrete PPO algorithm is used for reinforcement learning and imitation learning.
[0051] Table 1. State variables and action variables of the Markov decision process Time (t) 0: Chiller and water tank shut off Electricity price (\lambda_t) 1: Independent cooling unit Room temperature (T_o) 2: The water tank is supplied with cooling separately. Solar radiation (q_s) 3: Cooling storage in water tank Staffing rate in the room (r_{int}) 4: Water tank for cold storage, chiller for cooling Indoor temperature (T_i) 5: Cooling from the water tank and cooling from the chiller Average water tank temperature (T^{st})
[0052] Table 2 compares the hyperparameter settings of different algorithms used in the implementation process, including PPO, PWIL, and Generative Adversarial Learning (GAIL). The hyperparameter settings of different algorithms are basically consistent during the implementation process. The discount factor of the PPO algorithm is set to 1. The purpose is to make the agent fully consider the cumulative effect of all future rewards within a finite time step when making decisions, rather than decaying the long-term rewards. This is because the goal of the policy in this paper is to maintain comfort and reduce the total cost over the entire cycle, rather than the cost of a single time step. Since the reward function of the ImRL algorithm is not directly related to the cost, the discount factor can be reduced to reduce the variance of the reward estimation and improve the stability of the policy gradient update.
[0053] Table 2 Algorithm Hyperparameter Settings
[0054] In the early stages of training, the reward growth trends of the three algorithms were quite similar, and the imitation learning method did not significantly outperform PPO. Figure 2 As shown, the PPO reward increases slowly as training progresses, and fails to surpass the baseline by the end of 500 epochs, indicating limited policy quality. In contrast, PWIL and GAIL achieve a smooth transition to the reinforcement learning phase after approximately 200 epochs, at which point the reward value surpasses the baseline policy. In subsequent training, both imitation learning methods significantly improve performance through autonomous reinforcement exploration, and the final convergence result is close to the ideal policy under the scenario of no prediction error. The hourly cooling capacity of the chiller and the hourly energy storage / discharge of the water tank under different policies are shown below. Figures 3 to 8As shown, due to the limited water tank capacity, the baseline strategy only releases cooling capacity during peak electricity price periods (9-12 AM), using chillers for cooling at other times. Compared to the baseline strategy, the optimized strategy reduces power consumption during peak electricity price periods by increasing cooling storage during the flat electricity price period (12-4 PM) and releasing cooling during the period (4-8 PM). Even after introducing meteorological and human forecasting errors, the optimized strategy remains generally similar to the optimized strategy under ideal forecasting, except that it triggers additional chiller cooling or additional water tank cooling at certain times due to overestimation of load. The reinforcement learning-based strategy, which closely approximates the optimized strategy under ideal forecasting, only increases the water tank's cooling storage capacity for one time during the night.
[0055] A comparative analysis of different strategies on system load transfer, operational efficiency conversion into learning reward values, and operating costs is presented, such as... Figure 9 As shown, compared to the baseline strategy, the optimization strategy and the imitation reinforcement learning strategy mainly achieve this by shifting the peak electricity load from 16:00 to 20:00 to the flat-price period from 12:00 to 16:00. The load shifting amount in other time periods is minimal, such as... Figure 10 and Figure 11 As shown, compared to the baseline strategy, despite the presence of prediction errors, the reward values of the optimized strategies are still higher than those of the baseline strategy, by 8.7%, 8.9%, and 13.9% in scenarios 1, 2, and 3, respectively. The reward value of the optimized strategy without prediction errors is 15.2% higher than the baseline strategy. The ImRL strategy, trained for 500 epochs, has the highest reward value, 15.9% higher than the baseline strategy. The strategy using only the PPO algorithm, however, still has a lower reward value than the baseline after 500 epochs. The running cost results are similar to the reward function results. Optimized scenarios 1, 2, and 3 show cost savings of 8.7%, 8.9%, and 15.3% higher than the baseline scenarios, respectively. The optimized strategy without prediction errors shows the best cost savings, saving 18.3% compared to the baseline strategy. The ImRL strategy trained for 500 epochs saves 17.9% more cost than the baseline, slightly lower than the optimized strategy without prediction errors. The strategy using only the PPO algorithm still has a higher running cost than the baseline after 500 epochs.
[0056] Example 3, referring to Figure 12 As an embodiment of the present invention, a method for energy management of a building air conditioning system based on imitation reinforcement learning is provided, comprising:
[0057] S1: Construct a day-ahead optimization scheduling strategy based on the gray box model of the building air conditioning system.
[0058] Furthermore, the day-ahead optimization scheduling strategy based on the gray box model of the building air conditioning system includes establishing an air conditioning system energy consumption model based on the building RC model, chiller efficiency model, and water tank energy storage model; identifying parameters based on actual operating data; setting an optimization objective of minimizing operating costs; combining time-of-use electricity price information, energy balance, operational mutual exclusion, and comfort temperature constraints; solving a mixed integer programming problem; and outputting the day-ahead optimization scheduling strategy for system operation.
[0059] It should be noted that the RC model is used to describe the building's thermodynamic properties, including three temperature nodes: interior (Ti), intermediate layer of the building envelope (Tm), and exterior envelope (Te). Its dynamic process is described by a discretized heat balance equation, expressed as: , , , in, For indoor air temperature control, For internal thermal mass temperature, For the temperature of the exterior wall, For floor temperature, Outdoor temperature Thermal resistance (K / W), Heat capacity (J / K), The equivalent heat absorption area (m³) For time step, For a moment, To provide cooling capacity for the chiller This is for the cooling capacity of the water tank. As an internal heat source (W), Solar radiation intensity (W / m2).
[0060] It should also be noted that, considering the approximate value of the uniform temperature within the tank energy storage (TES), the energy balance of the TES is expressed as: , , in, The average water temperature in the tank. This refers to the amount of heat dissipated from the water tank to the environment. To store the cooling capacity of the water tank, This is for the cooling capacity of the water tank. The total mass of water in the tank. The specific heat capacity of water, The heat exchange coefficient between the water tank and the environment. The ambient temperature.
[0061] The power consumption of the chiller is expressed as the chiller's coefficient of performance (COP) and cooling capacity: , , , in, For the coefficient of performance of the refrigeration unit, For coefficients, For load rate, To provide cooling capacity for the chiller For chiller power, The power consumption is for cooling the chiller.
[0062] The scheduling optimization strategy, with the objective function of minimizing total electricity cost, is expressed as follows: The strategy for scheduling the air conditioning system operation one day in advance is: , in, For time-of-use electricity pricing, To mimic the learning cycle, To store energy, The power consumption is for cooling the chiller.
[0063] The equality constraints of the optimization model are the building thermal network model, energy consumption and energy storage model, and the inequality constraints include: setting binary variables. , , ∈{0,1} represent the operating modes of whether the chiller provides cooling, whether the water tank stores energy, and whether the water tank releases energy, respectively, as follows: , in, Whether the water tank can store energy To determine whether the water tank releases energy.
[0064] The above indicates that energy storage and release in the water tank cannot occur simultaneously in this system, and the average temperature constraint of the water tank is expressed as: , When the time is 08:00–21:00, the indoor temperature constraint is expressed as: , Relaxation at other times improves the feasibility of the solution.
[0065] It should also be noted that by constructing an RC gray box model, a chiller efficiency model, and a water tank energy storage model, and by developing a day-ahead optimization scheduling strategy, we can solve the technical problems of optimizing scheduling under multiple constraints such as time-of-use pricing, equipment mutual exclusion, and comfort temperature, thereby reducing operating costs and achieving the beneficial effects of reducing energy consumption costs, ensuring indoor thermal comfort, and making full use of energy storage.
[0066] S2: Establish an imitation reinforcement learning method based on bulldozer distance.
[0067] Furthermore, the imitation reinforcement learning method includes two stages. The first is the imitation learning stage 100, and then the transition to the reinforcement learning stage 200 is achieved through the design of a dynamic reward function. The imitation learning introduces an imitation learning method based on the original Wasserstein distance. By representing the agent's reward function as the inverse of the original bulldozer distance between the agent's state-action distribution and the expert's state-action distribution, the gap between the agent and the expert is reduced during training. The dynamic reward function makes the reward composition during the agent's training process transition from imitation learning to reinforcement learning.
[0068] It should be noted that imitation reinforcement learning exhibits better convergence speed and sample efficiency during training, while imitation learning based on Wasserstein distance has advantages in stability and reduces the burden of hyperparameter tuning.
[0069] S3: Constructing a reinforcement learning-based energy management method for building air conditioning systems.
[0070] Furthermore, the construction of an imitation reinforcement learning energy management method for building air conditioning systems includes: an output-based day-ahead optimization scheduling strategy; expert knowledge for the imitation reinforcement learning method based on bulldozer distance; designing the reward function for reinforcement learning and the hyperparameters of the imitation reinforcement algorithm according to the actual system; online training of the imitation reinforcement learning agent in a real or simulated environment; and gradually improving control performance through multi-day cycles.
[0071] It should be noted that Markov decision processes for imitation learning and reinforcement learning are performed separately according to the actual system design. These processes include state variables, action variables, reward functions, and hyperparameters. The imitation and reinforcement learning agent is trained online in a real or simulated environment, and its control performance is gradually improved through multiple days of iteration. The state and action variables are consistent between imitation learning and reinforcement learning, while the reward function of reinforcement learning is more dependent on the actual system. The reward function for imitation learning is expressed as follows: , , in, For environmental rewards, i.e., the reward function of reinforcement learning, for Time-of-use electricity pricing at any given moment The power consumption of the chiller (kWh) The power consumption of the water pump (kWh) The penalty factor is used to penalize violations of room temperature and average tank temperature constraints. The average water temperature in the tank. This refers to the indoor temperature.
[0072] It should also be noted that by designing customized reward functions and hyperparameters, the problem of dynamically optimizing energy consumption and comfort in building air conditioning systems under time-of-use electricity pricing and temperature constraints can be solved, thereby reducing energy consumption, saving operating costs, ensuring indoor comfort, and improving the adaptability and efficiency of system control.
[0073] This embodiment also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements an emergency control system that takes into account the switching of grid-type energy storage parameters as proposed in the above embodiment.
[0074] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0075] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0076] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0077] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0078] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. An energy management method for building air conditioning systems based on imitation reinforcement learning, characterized in that, include, Construct a day-ahead optimization scheduling strategy based on a gray box model of a building air conditioning system; include An energy consumption model for the air conditioning system is established by creating a building RC model, a chiller efficiency model, and a water tank energy storage model. Based on actual operating data, parameters are identified, and an optimization objective of minimizing operating costs is set. Combining time-of-use electricity price information, energy balance, operational mutual exclusion, and comfort temperature constraints, a mixed integer programming problem is solved, and an optimized scheduling strategy for the system's operating days is output. An imitation reinforcement learning method based on bulldozer distance is established, including an imitation learning phase (100). An imitation learning method based on the original bulldozer distance is introduced. By representing the agent reward function as the inverse of the original bulldozer distance between the agent's state-action distribution and the expert's state-action distribution, the gap between the agent and the expert is reduced during the training process. In the reinforcement learning phase (200), the dynamic reward function is used to transition the reward composition of the agent training from imitation learning to reinforcement learning. Construct a simulation reinforcement learning energy management method for building air conditioning systems; This includes output-based day-ahead optimization scheduling strategies, expert knowledge for imitation reinforcement learning methods for bulldozer distance, designing the reward function for reinforcement learning and the hyperparameters of the imitation reinforcement algorithm based on the actual system, training the imitation reinforcement learning agent online in a real or simulated environment, and improving control performance through multi-day cycles. The reinforcement learning phase (200) includes, The dynamic reward function enables the training focus of the agent to smoothly transition from imitation learning to autonomous exploration in reinforcement learning, and to deeply explore strategies that are superior to expert actions, thus defining an imitation reinforcement learning algorithm. The dynamic reward function includes, The dynamic reward function, which combines imitation learning rewards and environmental rewards, is expressed as follows: , in, For dynamic reward function, To encourage learning through rewards, For environmental rewards, i.e., the reward function of reinforcement learning, For activation function, For an experience value, is a coefficient.
2. The energy management method for building air conditioning systems based on imitation reinforcement learning as described in claim 1, characterized in that: The original bulldozer distance includes Bulldozer distance measurement moves data from distribution Move into distribution The minimum cost required at that time is expressed as: , in, To minimize the cost of movement, for and The joint distribution set is equivalent to the movement strategy, which represents the movement from position. Move to The probability quality, From arrive Requires moving a distance, For a space, all possibilities A set of point pairs.
3. The energy management method for building air conditioning systems based on imitation reinforcement learning as described in claim 1, characterized in that: The agent's state and action distribution includes, In a Markov decision process, the empirical distribution of all state-action pairs generated during the agent's interaction with the environment within the imitation learning cycle is defined as follows: , in, For the distribution of agent state and action, To mimic the learning cycle, For the sake of Dirac distribution centered on, For expert state action distribution, This represents the number of expert samples.
4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the energy management method for building air conditioning systems based on imitation reinforcement learning as described in any one of claims 1 to 3.