Near-end strategy optimization method and device for building flexible energy supply system
Through the proximal strategy optimization algorithm and the building virtual energy storage system model, Markov decision-making process model is built to optimize and control the building flexible energy supply system, solving the problems of low computing efficiency and lagging strategy adjustment in the existing technology, and achieving efficient optimization and control of the building flexible energy supply system in dynamic environments.
Patent Information
- Application Number
- CN202510368937.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
AI Technical Summary
The existing optimization control method for flexible building energy supply systems is low in calculation efficiency and lagging in strategy adjustment when dealing with multivariable and nonlinear optimization problems, making it difficult to balance the complexity and accuracy of the optimization model, and it is difficult to achieve optimal economic operation in dynamic meteorological environments.
The near-end strategy optimization (PPO) algorithm is adopted and the building virtual energy storage system (VESS) model is combined with the Markov decision-making process (MDP) model is constructed, and the building flexible energy supply system is optimized and controlled, and real-time and efficient control strategies are output.
Real-time and efficient optimization and control of the building's flexible energy supply system in dynamic environments, reducing electricity costs, ensuring indoor thermal comfort constraints, and improving system flexibility and operating efficiency.
Smart Images

Figure CN120233675A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of optimization control of building energy systems, and particularly relates to a method and device for optimizing a proximal strategy of a building flexible energy supply system. Background Art
[0002] With the acceleration of the urbanization process and the continuous improvement of the living standards of residents, building energy consumption has become an important part of global energy consumption. According to statistical data, the current building operation energy consumption in China accounts for about 21% of the total social energy consumption, and the energy consumption of the air-conditioning system accounts for about 50% of the building operation energy consumption. With the wide access of renewable energy and the development of a new power system, how to improve the economic operation efficiency of the building flexible energy supply system on the premise of ensuring the indoor environmental comfort has become a core problem to be solved urgently. Combining air-conditioning equipment with an electric energy storage system and optimizing the control has become one of the feasible methods to achieve the economic operation of buildings. Although the electric energy storage system is crucial for the building flexible energy supply system, its high installation and maintenance costs, large occupied space and safety problems seriously limit its application in this field. As an innovative flexibility quantification technology, the building virtual energy storage system (VESS) has become a research hotspot in the field of building energy management. This technology utilizes the phase-lag response of the indoor temperature caused by the thermal inertia characteristics of the building envelope. After combining with the time-of-use electricity price mechanism, the building flexible energy supply system can convert photovoltaic electric energy into heat energy and store it in the building envelope during the valley price period, and release the stored heat energy during the peak price period, optimizing the electro-thermal load distribution of the building flexible energy supply system while maintaining the thermal comfort, and effectively improving the economic operation efficiency of the building flexible energy supply system. Combining VESS with the control strategy of the building flexible energy supply system can significantly improve the energy supply-demand matching degree of the building flexible energy supply system and enhance its flexibility and operation efficiency.
[0003] At present, a large number of studies have achieved multi-dimensional quantitative evaluation of VESS by quantifying key parameters such as the charging / discharging power, heat capacity, and state of charge (SOC) of VESS, and explored the optimal control strategy of the building flexible energy supply system combined with VESS by combining optimization control methods such as model predictive control. However, there are two significant limitations in the existing methods: First, the optimal control problem of the building flexible energy supply system involves multi-dimensional control variables such as the purchased power and the charging / discharging power of VESS, as well as complex constraint conditions such as thermal comfort and power balance. When traditional optimization control methods deal with such multi-variable and non-linear optimization problems, there are generally problems of low computational efficiency and lagging strategy adjustment, and it is difficult to balance the complexity and accuracy of the optimization model; at the same time, the actual operating state of the building flexible energy supply system is significantly affected by the change of the outdoor meteorological environment, and the accuracy of the traditional optimization control strategy is restricted by the prediction accuracy of meteorological parameters, and it is difficult to achieve the optimal economic operation under the dynamic meteorological environment.
[0004] With the continuous development of the Internet of Things technology, a large number of sensors and edge devices can provide sufficient and reliable data, thus providing the feasibility for the application of deep reinforcement learning algorithms in the building flexible energy supply system. A large number of studies have used Deep Q-Learning (DQN) and Deep Deterministic Policy Gradient (DDPG) algorithms to construct the optimal control framework of the building flexible energy supply system. However, when DQN and DDPG algorithms deal with multi-variable and non-linear optimal control problems, the dimensionality expansion of their state space will significantly reduce the model convergence efficiency, and their ability to couple and model multiple constraints is insufficient. The soft constraint processing mechanism may lead to constraint violation problems, which in turn affect the convergence speed and robustness of the algorithms. Proximal Policy Optimization (PPO) follows the policy gradient theoretical framework, achieves the balance between policy update stability and exploration efficiency in the continuous control domain, uses the adaptive KL divergence constraint to construct a dynamic penalty function to accurately satisfy multiple types of constraint conditions, and introduces the advantage function variance reduction technology, which can significantly improve the stability of policy gradient estimation. At the same time, it has the advantages of high computational efficiency and stable convergence performance, and can effectively overcome the deficiencies of DQN and DDPG algorithms.
[0005] Therefore, how to invent a proximal policy optimization method for the building flexible energy supply system to achieve real-time and efficient optimal control under the long-term dynamic change of the building flexible energy supply system environment, reduce the electricity cost and ensure the constraint of indoor thermal comfort has become an urgent problem to be solved. Summary of the Invention
[0006] To this end, the present invention provides a method and device for optimizing the proximal strategy of a building flexible energy supply system, constructs a Markov Decision Process (MDP) of the building flexible energy supply system combined with VESS, and further uses the PPO algorithm to optimize and control the building flexible energy supply system combined with VESS, which can achieve real-time and efficient optimization control under long-term dynamic changes in the environment of the building flexible energy supply system, reduce the electricity cost and ensure the constraint of indoor thermal comfort.
[0007] To achieve the above object, the present invention provides the following technical solutions: In the first aspect, a method for optimizing the proximal strategy of a building flexible energy supply system is provided, including:
[0008] Based on the thermal inertia characteristics of the building envelope structure, construct a building virtual energy storage system VESS model, and introduce time-varying parameters through the building virtual energy storage system VESS model;
[0009] According to the set control target, use the time-varying parameters introduced by the building virtual energy storage system VESS model to construct an optimization control model of the building flexible energy supply system combined with VESS;
[0010] Convert the optimization control model of the building flexible energy supply system combined with VESS into a Markov decision process model;
[0011] Optimize the Markov decision process model through the proximal policy optimization algorithm and output the control strategy.
[0012] As a preferred scheme of the method for optimizing the proximal strategy of a building flexible energy supply system, in the process of constructing the building virtual energy storage system VESS model, according to the performance characteristics of the electrical energy storage, introduce three time-varying parameters: virtual power, virtual capacity, and virtual state of charge;
[0013] The relational expression between the three time-varying parameters is:
[0014]
[0015] In the formula, t is the t-th control period; Δt is the control time interval; SOC V (t) is the virtual state of charge of VESS at time t; C V (t) is the virtual capacity of VESS at time t, kWh; P V (t) is the virtual power of VESS at time t, kW.
[0016] As a preferred scheme of the method for optimizing the proximal strategy of a building flexible energy supply system, the objective function expression of the optimization control model of the building flexible energy supply system combined with VESS is:
[0017] C = Ce +C SOC
[0018]
[0019] In the formula, C is the electricity cost of the building flexible energy supply system; C e is the operating cost of the building flexible energy supply system; C SOC is the SOC penalty cost of the VESS; P base (t) is the reference power consumption at time t, kW; P other (t) is the power of other electrical equipment at time t, kW; P PV (t) is the PV power at time t, kW; c(t) is the electricity price at time t, cents / kW; σ is the SOC V over-limit cost coefficient; T in (t) is the indoor temperature of the building at time t, °C; T max is the highest indoor temperature, °C; T min is the lowest indoor temperature, °C.
[0020] As an optimal solution of a proximal strategy optimization method for a building flexible energy supply system, the constraints of the optimized control model of the building flexible energy supply system combined with the VESS include: VESS constraints, PV output power constraints, and electrical power balance constraints of the building flexible energy supply system;
[0021] The expression of the VESS constraint is:
[0022] 0 ≤ SOC V (t) ≤ 1
[0023] -P dismax (t) ≤ P V (t) ≤ P cmax (t)
[0024] In the formula, P dismax (t) is the maximum discharge power of the VESS at the t-th time period, kW; P cmax (t) is the maximum charging power of the VESS at the t-th time period, kW;
[0025] The expression of the PV output power constraint is:
[0026] 0 ≤ P PV (t) ≤ P PVf (t)
[0027] In the formula, P PVf (t) is the maximum predicted output power of the PV at time t, kW;
[0028] The expression of the electrical power balance constraint of the building flexible energy supply system is:
[0029] P com (t) + P PV (t) = P base (t) + P V (t) + P other (t)
[0030] In the formula, P com (t) is the purchased power, in kW.
[0031] As an optimal solution of a proximal strategy optimization method for a building flexible energy supply system, during the process of converting the optimized control model of the building flexible energy supply system combined with VESS into the Markov decision process model, the state variables, control actions, reward function, and transition function are defined;
[0032] The expression of the state variable is:
[0033] s(t) = [SOC V (t), T in (t), T out (t), Q solar (t), c(t)]
[0034] In the formula, s(t) is the state information at the current moment; T out (t) is the outdoor temperature at time t, in °C; Q solar (t) is the solar heat gain power of the building at time t, in kW;
[0035] The expression of the control action is:
[0036] a(t) = [P V (t), P PV (t)]
[0037] In the formula, a(t) represents the action adopted at the current moment;
[0038] The expression of the reward function is:
[0039] r(t) = w1·r elec (t) + w2·r soc (t)
[0040] In the formula, r(t) is the total reward at the current moment; r elec (t) is the electricity cost reward; r soc (t) is the SOC reward of VESS; w1 and w2 are both weight coefficients.
[0041] As an optimal solution of a proximal policy optimization method for a building flexible energy supply system, during the optimization of the Markov decision process model by the proximal policy optimization algorithm, the proximal policy optimization algorithm implements the policy gradient theorem through an Actor-Critic structure; the optimization objective functions of the Actor network and the Critic network in the Actor-Critic structure are as follows:
[0042]
[0043] In the formula, E(t) is the expectation of the loss function at time t; clip(·) is the clipping function; ε is the hyperparameter of clip(·); π θ is the policy of the Actor network under parameter θ; r θ (t) is importance sampling, which measures the difference between the new policy and the old policy before and after update; A(t) is the advantage function; is the policy entropy; β is the entropy regularization term, which is used to balance the utilization of the policy by the control system and the exploration of unknown policies;
[0044] The loss function of the Critic network is:
[0045] L Critic (φ) = E(t)[(V φ (s(t)) - R(t)) 2
[0046] In the formula, φ is the parameter of the Critic network; V φ (s(t)) is the state value function, that is, the judgment of the control system on the quality of the current state; R(t) is the cumulative reward, which for the building flexible energy supply system is the sum of the rewards from the current state to the future state during system operation.
[0047] In the second aspect, the present invention also provides a proximal policy optimization device for a building flexible energy supply system. Based on the above proximal policy optimization method for a building flexible energy supply system, it includes:
[0048] A building virtual energy storage system VESS model construction module, which is used to construct a building virtual energy storage system VESS model based on the thermal inertia characteristics of the building envelope structure, and introduce time-varying parameters through the building virtual energy storage system VESS model;
[0049] A building flexible energy supply system optimization control model construction module combined with VESS, which is used to construct a building flexible energy supply system optimization control model combined with VESS according to the set control target and using the time-varying parameters introduced by the building virtual energy storage system VESS model;
[0050] The Markov decision process model conversion module enables the user to convert the optimized control model of the building flexible energy supply system combined with VESS into a Markov decision process model;
[0051] The Markov decision process model optimization module is used to optimize the Markov decision process model through the proximal policy optimization algorithm and output the control strategy.
[0052] As an optimal solution of a proximal policy optimization device for a building flexible energy supply system, in the building virtual energy storage system VESS model construction module, during the process of constructing the building virtual energy storage system VESS model, according to the performance characteristics of the electrical energy storage, three time-varying parameters, namely virtual power, virtual capacity, and virtual state of charge, are introduced;
[0053] The relational expressions among the three time-varying parameters are:
[0054]
[0055] In the formula, t is the t-th control period; Δt is the control time interval; SOC V (t) is the virtual state of charge of VESS at time t; C V (t) is the virtual capacity of VESS at time t, kWh; P V (t) is the virtual power of VESS at time t, kW.
[0056] As an optimal solution of a proximal policy optimization device for a building flexible energy supply system, in the optimized control model construction module of the building flexible energy supply system combined with VESS, the objective function expression of the optimized control model of the building flexible energy supply system combined with VESS is:
[0057] C = C e + C SOC
[0058]
[0059] In the formula, C is the electricity cost of the building flexible energy supply system; C e is the operating cost of the building flexible energy supply system; C SOC is the SOC penalty cost of VESS; P base (t) is the reference power consumption at time t, kW; P other (t) is the power of other electrical equipment at time t, kW; P PV (t) is the PV power at time t, kW; c(t) is the electricity price at time t, cents / kW; σ is the cost coefficient for SOC V overlimit; T in (t) is the indoor temperature of the building at time t, °C; T maxis the highest indoor temperature, °C; T min is the lowest indoor temperature, °C.
[0060] As an optimal solution of a proximal strategy optimization device for a building flexible energy supply system, in the building flexible energy supply system optimization control model construction module combined with VESS, the constraints of the building flexible energy supply system optimization control model combined with VESS include: VESS constraint, PV output power constraint, and building flexible energy supply system electric power balance constraint;
[0061] The expression of the VESS constraint is:
[0062] 0 ≤ SOC V (t) ≤ 1
[0063] -P dismax (t) ≤ P V (t) ≤ P cmax (t)
[0064] In the formula, P dismax (t) is the maximum discharge power of VESS at the t-th time period, kW; P cmax (t) is the maximum charging power of VESS at the t-th time period, kW;
[0065] The expression of the PV output power constraint is:
[0066] 0 ≤ P PV (t) ≤ P PVf (t)
[0067] In the formula, P PVf (t) is the maximum predicted output power of PV at the t-th time period, kW;
[0068] The expression of the building flexible energy supply system electric power balance constraint is:
[0069] P com (t) + P PV (t) = P base (t) + P V (t) + P other (t)
[0070] In the formula, P com (t) is the purchased electric power, kW.
[0071] As an optimal solution of a proximal strategy optimization device for a building flexible energy supply system, in the Markov decision process model conversion module, in the process of converting the building flexible energy supply system optimization control model combined with VESS into the Markov decision process model, the state variables, control actions, reward function, and transition function are defined;
[0072] The expression of the state variable is as follows:
[0073] s(t) = [SOC V (t), T in (t), T out (t), Q solar (t), c(t)]
[0074] where s(t) is the state information at the current moment; T out (t) is the outdoor temperature in the t period, in °C; Q solar (t) is the building solar heat gain power in the t period, in kW;
[0075] The expression of the control action is as follows:
[0076] a(t) = [P V (t), P PV (t)]
[0077] where a(t) represents the action adopted at the current moment;
[0078] The expression of the reward function is as follows:
[0079] r(t) = w1·r elec (t) + w2·r soc (t)
[0080] where r(t) is the total reward at the current moment; r elec (t) is the electricity cost reward; r soc (t) is the SOC reward of the VESS; w1 and w2 are both weight coefficients.
[0081] As an optimal solution of a proximal policy optimization device for a building flexible energy supply system, in the Markov decision process model optimization module, during the optimization of the Markov decision process model by the proximal policy optimization algorithm, the proximal policy optimization algorithm realizes the policy gradient theorem through the Actor-Critic structure; the optimization objective functions of the Actor network and the Critic network in the Actor-Critic structure are:
[0082]
[0083] where E(t) is the expectation of the loss function at the t moment; clip(·) is the clipping function; ε is the hyperparameter of clip(·); π θ is the policy of the Actor network under the parameter θ; r θ (t) is the importance sampling, which measures the difference between the new policy and the old policy before and after the update; A(t) is the advantage function; is the policy entropy; β is the entropy regularization term used to balance the exploitation of the policy by the control system and the exploration of unknown policies;
[0084] The loss function of the Critic network is:
[0085] L Critic (φ) = E(t)[(V φ (s(t)) - R(t)) 2
[0086] where φ is the Critic network parameter; V φ (s(t)) is the state value function, that is, the judgment of the control system on the quality of the current state; R(t) is the cumulative reward, which for the building flexible energy supply system is the sum of the rewards of the system from the current state to the future state during operation.
[0087] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, in which a program code of a method for optimizing the proximal policy of a building flexible energy supply system is stored, and the program code includes instructions for executing a method for optimizing the proximal policy of a building flexible energy supply system according to the first aspect or any possible implementation manner thereof.
[0088] In a fourth aspect, the present invention provides an electronic device, including: a memory and a processor;
[0089] The processor and the memory communicate with each other through a bus; the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute a method for optimizing the proximal policy of a building flexible energy supply system according to the first aspect or any possible implementation manner thereof.
[0090] The present invention has the following advantages: Based on the thermal inertia characteristics of building envelopes, the present invention constructs a virtual energy storage system (VESS) model for buildings, and introduces time-varying parameters through the VESS model of the building virtual energy storage system; According to the set control objectives, using the time-varying parameters introduced by the VESS model of the building virtual energy storage system, an optimized control model of a building flexible energy supply system combined with VESS is constructed; The optimized control model of the building flexible energy supply system combined with VESS is transformed into a Markov decision process model; The Markov decision process model is optimized by the proximal policy optimization algorithm, and a control strategy is output. The present invention accurately transforms the building flexible energy supply system model combined with VESS into an MDP model, and uses the PPO algorithm to achieve efficient optimized control of the building flexible energy supply system. The present invention realizes the generation of an optimized control strategy at the millisecond level, and the speed is significantly better than the DQN and DDPG methods, reflecting its high efficiency in dealing with the dynamic optimization problems of multi-variable and non-linear building flexible energy supply systems, and can meet the real-time requirements of the optimized control of the building flexible energy supply system; At the same time, in a dynamic meteorological environment, the present invention effectively reduces the operating cost of the building flexible energy supply system, while maintaining a low proportion of room temperature exceeding the limit, reflecting its significant advantages in economy and comfort. Description of the Drawings
[0091] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained according to the provided drawings.
[0092] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0093] Figure 1 It is a schematic flow chart of a proximal policy optimization method for a building flexible energy supply system provided in Embodiment 1 of the present invention;
[0094] Figure 2 It is a schematic diagram of a building flexible energy supply system with VESS in a proximal policy optimization method for a building flexible energy supply system provided in Embodiment 1 of the present invention;
[0095] Figure 3Schematic diagram of electricity price and meteorological data in a possible embodiment provided in Embodiment 1 of the present invention; among them, (a) is the electricity price; (b) is the meteorological data;
[0096] Figure 4 Schematic diagram of the algorithm training situation in a possible embodiment provided in Embodiment 1 of the present invention;
[0097] Figure 5 P under different control methods in a possible embodiment provided in Embodiment 1 of the present invention V Schematic diagram of the change situation;
[0098] Figure 6 SOC under different control methods in a possible embodiment provided in Embodiment 1 of the present invention V Schematic diagram of the change situation;
[0099] Figure 7 Schematic diagram of the architecture of the proximal strategy optimization device for a building flexible energy supply system provided in Embodiment 2 of the present invention. Specific implementation manners
[0100] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0101] Embodiment 1
[0102] Refer to Figure 1 , Embodiment 1 of the present invention provides a proximal strategy optimization method for a building flexible energy supply system, including the following steps:
[0103] S1. Based on the thermal inertia characteristics of the building envelope structure, construct a building virtual energy storage system VESS model, and introduce time-varying parameters through the building virtual energy storage system VESS model;
[0104] S2. According to the set control target, use the time-varying parameters introduced by the building virtual energy storage system VESS model to construct an optimized control model for the building flexible energy supply system combined with VESS;
[0105] S3. Convert the optimized control model for the building flexible energy supply system combined with VESS into a Markov decision process model;
[0106] S4. Optimize the Markov decision process model through the proximal policy optimization algorithm and output the control policy.
[0107] In this embodiment, in step S1, based on the thermal inertia characteristics of the building envelope structure, a building virtual energy storage system (VESS) model is constructed, and time-varying parameters are introduced through the building VESS model.
[0108] Specifically, as Figure 2 shown, the building flexible energy supply system internally includes photovoltaic (PV), inverter air conditioner (IAC), other electrical equipment, smart meters, edge sensors, and a smart control center. During the operation of the building flexible energy supply system, the distribution network (DN) and PV provide the electrical energy required daily; other electrical equipment refers to all electrical equipment in the building flexible energy supply system except IAC; the smart meter is connected to the DN, and dynamically obtains electricity price information in real time and transmits it to the smart control center; the edge sensors monitor the internal operation status of the building and external environmental data, providing basic data support for optimized control; the weather station provides various weather information for the building flexible energy supply system; the IAC is used for building electric heating to ensure that the temperature inside the building is maintained within the user's thermal comfort range; the smart control center can integrate the data of the edge sensors, smart meters, weather stations, PV, and VESS, and combine the PPO algorithm to achieve optimized control of the building flexible energy supply system.
[0109] Taking into account the thermal inertia of the building envelope structure and the dynamic characteristics of indoor and outdoor temperatures comprehensively, the building envelope structure and IAC can be modeled as Figure 2 the building VESS model shown. This model significantly improves the flexibility of optimized control of the building flexible energy supply system by quantifying the virtual electricity storage capacity of the building flexible energy supply system.
[0110] In this embodiment, first, based on the heat exchange mechanism of the building envelope structure, a building heat dissipation power characterization parameter including factors such as outdoor temperature and solar radiation is established; then, through the energy consumption and heating power conversion model of the IAC, the heat storage capacity of the building envelope structure is equivalently converted into electricity storage capacity; finally, the building heat dissipation power parameter and the IAC model are integrated to construct a building VESS model to quantify the electricity storage capacity of the building envelope structure.
[0111] In this embodiment, during the construction of the building VESS model, the present invention introduces three core parameters according to the performance characteristics of electrical energy storage: virtual power, virtual capacity, and virtual state of charge. Specifically, virtual power reflects the dynamic power change of the VESS before and after participating in the optimization control strategy of the building flexible energy supply system, and is an important indicator to measure its ability to respond to scheduling instructions; virtual capacity reveals the maximum storage capacity that the VESS can provide during the optimization scheduling process of the building flexible energy supply system, and is a key parameter to evaluate its contribution to system flexibility; the virtual state of charge directly depicts the instantaneous energy state of the VESS. The relationship among the three time-varying parameters is shown in Equation (1):
[0112]
[0113] In the formula, t is the t-th control period; Δt is the control time interval; SOC V (t) is the virtual state of charge of the VESS at time t; C V (t) is the virtual capacity of the VESS at time t, kWh; P V (t) is the virtual power of the VESS at time t, kW.
[0114] Among them, virtual power: Define the heating power that maintains the indoor temperature at the user-set temperature as the reference heating power, and its calculation is shown in Equation (2):
[0115]
[0116] In the formula, i is the i-th type of building envelope structure, and i takes values from 1 to 4, corresponding to walls, windows, roofs, and doors respectively; S i is the inner surface area of the i-th type of building envelope structure, m 2 ; α in is the heat transfer coefficient of the building inner surface, W / (m 2 ·°C); T in (t) is the indoor temperature of the building at time t, °C; T out (t) is the outdoor temperature at time t, °C; Q solar (t) is the solar heat gain power of the building at time t, kW.
[0117] The main power-consuming equipment of the IAC system is the compressor. To simplify the calculation process, the power consumption of the remaining equipment of the IAC system is ignored, and the relationship between the power consumption of the IAC system and the heating power is simplified to a linear relationship, as shown in Equation (3):
[0118]
[0119] In the formula, Q IAC (t) is the heating power of the IAC at time t, kW; P IAC(t) is the power consumption of the IAC in the t period, kW; a1, b1, a2, and b2 are constant coefficients.
[0120] Substituting Equation (2) into Equation (3), the reference power consumption P base (t) corresponding to Q can be obtained, as shown in Equation (4): base (t) is as follows:
[0121]
[0122] Define the power deviation between the power P IAC at the t moment and P base as the virtual power P V (t), as shown in Equation (5):
[0123]
[0124] When P V (t) > 0, the VESS charges; conversely, the VESS discharges.
[0125] Among them, the virtual capacity:
[0126] In the t period, when the IAC is turned off, the self-discharge power consumption corresponding to the indoor temperature decreasing from T max to T min is defined as the virtual capacity C V (t), and its calculation formula is as shown in Equation (6):
[0127]
[0128] In the formula, t k is the initial moment of the t-th scheduling period, h; Δτ ca (t) is the time required for the indoor temperature to decrease from T max to T min when the IAC is turned off, h.
[0129] Among them, the virtual state of charge:
[0130] In the t period, when the IAC is turned off, the self-discharge power consumption corresponding to the indoor temperature decreasing from the actual indoor temperature T in (t) of the t scheduling period to T min is defined as the actual stored power E V (t) of the VESS, as shown in Equation (7):
[0131]
[0132] In the formula, Δτ E (t) represents that when the IAC is turned off, the indoor temperature decreases from T in (t) to Tmin Required time, h.
[0133] During time period t, the virtual state of charge SOC V (t) is defined as the ratio of E V (t) of VESS to C V (t), as shown in Equation (8):
[0134]
[0135] In the formula, SOC V (t) is the virtual state of charge.
[0136] In this embodiment, in step S2, according to the set control target, a time-varying parameter introduced by the building virtual energy storage system VESS model is used to construct an optimal control model of the building flexible energy supply system combined with VESS;
[0137] Specifically, the optimal control target of the building flexible energy supply system containing VESS is to minimize the electricity cost of the building flexible energy supply system. Its objective function is shown in Equation (9), which includes the operation economic objective function of the building flexible energy supply system and the SOC penalty objective function of VESS. The decision variables of the control model are P V , PV power, and purchased power.
[0138] C = C e + C SOC (9)
[0139]
[0140] In the formula, C is the electricity cost of the building flexible energy supply system; C e is the operation cost of the building flexible energy supply system; C SOC is the SOC penalty cost of VESS; P base (t) is the reference power consumption at time period t, kW; P other (t) is the power of other electrical equipment at time period t, kW; P PV (t) is the PV power at time period t, kW; c(t) is the electricity price at time period t, cents / kW; σ is the cost coefficient for SOC V overlimit; T in (t) is the indoor temperature of the building at time period t, °C; T max is the highest indoor temperature, °C; T min is the lowest indoor temperature, °C.
[0141] In this embodiment, the constraints of the optimal control model of the building flexible energy supply system combined with VESS include: VESS constraints, PV output power constraints, and electrical power balance constraints of the building flexible energy supply system;
[0142] The expression of the VESS constraint is as follows:
[0143] 0 ≤ SOC V (t) ≤ 1(12)
[0144] -P dismax (t) ≤ P V (t) ≤ P cmax (t)(13)
[0145] In the formula, P dismax (t) is the maximum discharge power of the VESS in the t-th time period, in kW; P cmax (t) is the maximum charging power of the VESS in the t-th time period, in kW;
[0146] The expression of the PV output power constraint is as follows:
[0147] 0 ≤ P PV (t) ≤ P PVf (t)(14)
[0148] In the formula, P PVf (t) is the maximum predicted output power of the PV in the t-th time period, in kW;
[0149] The expression of the electric power balance constraint of the building flexible energy supply system is as follows:
[0150] P com (t) + P PV (t) = P base (t) + P V (t) + P other (t)(15)
[0151] In the formula, P com (t) is the purchased electric power, in kW.
[0152] In this embodiment, in step S3, the optimization control model of the building flexible energy supply system combined with VESS is transformed into a Markov decision process model;
[0153] Specifically, the optimization control model of the building flexible energy supply system combined with VESS is transformed into a Markov decision process model (MDP); MDP is the basis of the deep reinforcement learning algorithm, including four components: state variable S, control action A, reward function R, and transition function P. S refers to the set of all information that the controlled object can observe; A represents the set of variables that can be controlled; R reflects the set of rewards for the control actions taken by the agent; P describes the probability set of the changes in the environment where the agent is located. It should be noted that the capital letters used in this invention all represent sets, and the lowercase letters used hereinafter represent the specific values at the corresponding moments in the sets.
[0154] To convert the building flexible energy supply system model combined with VESS into an MDP, the above four variables need to be defined. First, define S. S includes the SOC value of VESS at the current moment, the indoor and outdoor temperatures, solar heat gain, time information, and electricity price information, as shown in Equation (16): V (t), the indoor and outdoor temperatures, solar heat gain, time information, and electricity price information, as shown in Equation (16):
[0155] s(t) = [SOC V (t), T in (t), T out (t), Q solar (t), c(t)] (16)
[0156] In the formula, s(t) is the state information at the current moment; T out (t) is the outdoor temperature in the t period, °C; Q solar (t) is the building solar heat gain power in the t period, kW;
[0157] A is defined as the combination of the power P V (t) of VESS and the photovoltaic power P pv (t). a(t) represents the action adopted at the current moment as shown in Equation (17). The limiting conditions of these two actions are shown in Equations (13) and (14) respectively.
[0158] a(t) = [P V (t), P PV (t)] (17)
[0159] R is the embodiment of the optimization goal. During the training process, the algorithm realizes the final goal by gradually maximizing the cumulative reward. According to the optimization goal (9), the final optimization goal is to reduce the electricity cost on the premise of ensuring thermal comfort. Therefore, R should cover two parts, namely, the electricity cost reward r elec (t) and the SOC reward r soc (t) of VESS. The total reward goal is shown in Equation (18):
[0160] r(t) = w1·r elec (t) + w2·r soc (t) (18)
[0161] In the formula, r(t) is the total reward at the current moment; r elec (t) is the electricity cost reward; r soc (t) is the SOC reward of VESS; w1 and w2 are both weight coefficients.
[0162] Based on the above definitions, the building flexible energy supply system model combined with VESS can be converted into an MDP, and the PPO algorithm can be further used for optimization.
[0163] In this embodiment, in step S4, the proximal policy optimization algorithm is used to optimize the Markov decision process model, and a control policy is output.
[0164] Specifically, on the basis of establishing the MDP model, the PPO algorithm is further used to optimize this process to achieve the purpose of dynamically outputting the optimal control strategy of the building flexible energy supply system in real time. The long-term expected return formula ρ π As shown in Equation (19):
[0165] ρ π = ∑d π (s)∑π(s,a)*r(s,a)(19)
[0166] In the formula, π(s,a) represents the strategy adopted by the control system, that is, the probability that the building flexible energy supply system selects action a in state s; ρ π represents the average reward obtained at each time step after the control strategy interacts with the environment in the long term under the current π; d π (s) refers to the stationary distribution of the entire MDP under π, which describes the probability of each s occurring when the MDP reaches stability; r(s,a) represents the reward obtained by the building flexible energy supply system when taking a in s.
[0167] The ultimate goal of the system control optimization of the building flexible energy supply system is to find the optimal strategy, that is, the control system can take appropriate control actions a in the face of any state s to achieve the maximum long-term expected return ρ π . For this purpose, it is first necessary to initialize the parameter θ as the parameter of π, which can be expressed as f(θ) = π. Since the goal is to maximize ρ π , it is necessary to use the gradient ascent update method to find the derivative of ρ π with respect to θ, as shown in Equation (20):
[0168]
[0169] In the formula, Q π (s,a) is the action value function, that is, the expected reward obtained when the control system takes the control action a in state s.
[0170] In this embodiment, the PPO algorithm adopts the Actor-Critic structure to implement the policy gradient theorem shown in Equation (21), which represents the core update strategy of the Actor network. That is, when the difference between the policies before and after the update is too large, a clipping method needs to be used to ensure that each update can be maintained within a stable range to achieve stable updates. The Actor network, as the policy generation module, is used to formulate the control strategy of the building flexible energy supply system; the Critic network, as the state value evaluation module, is used to evaluate the current state of the building flexible energy supply system. The loss functions of the Actor network and the Critic network are shown in Equations (21) and (22), respectively.
[0171]
[0172] In the formula, E(t) is the expectation of the loss function at time t; clip(·) is the clipping function; ε is the hyperparameter of clip(·); π θ is the policy of the Actor network under parameter θ; r θ (t) is importance sampling, which measures the difference between the new policy and the old policy before and after the update; A(t) is the advantage function; is the policy entropy; β is the entropy regularization term, which is used to balance the exploitation of the policy by the control system and the exploration of unknown policies;
[0173] The loss function of the Critic network is:
[0174] L Critic (φ) = E(t)[(V φ (s(t)) - R(t)) 2 (22)
[0175] In the formula, φ is the parameter of the Critic network; V φ (s(t)) is the state value function, that is, the judgment of the control system on the quality of the current state; R(t) is the cumulative return, which for the building flexible energy supply system is the sum of the rewards from the current state to the future state during system operation.
[0176]
[0177] A(t) = R(t) - V φ (s(t))(24)
[0178]
[0179] Based on the above formulas, namely formulas (21) and (22), the gradient descent method is adopted to reduce the loss function to achieve the update of the strategy. After the algorithm training is completed, when the algorithm is deployed in the building flexible energy supply system, at each moment, the control strategy will receive the state information of the current moment and output the control strategy according to the Actor network.
[0180] In a possible embodiment, a specific strategy optimization example is provided as follows:
[0181] The length, width, and height of the building are selected to be 10 meters, 10 meters, and 3 meters respectively, and the window-wall ratio is 0.2, that is, there is an external window with an area of 6 square meters on each external wall, and the roof is treated as an external wall. In the definition of the reward function, the values of w1 and w2 are both 0.5. The electricity price data used is the winter dynamic electricity price in the eastern region of the United States, which is sourced from the CEIC database, as Figure 3 (a) shown. The scheduling day is selected as a typical day in winter, and the outdoor temperature and solar heat gain data are sourced from the EnergyPlus website, as Figure 3 (b) shown. The control time interval is set to 1h.
[0182] To verify the advantages of the present invention, the present invention is compared with the DQN method and the DDPG method.
[0183] The training situations of the three control methods are as Figure 4 shown, and the control results of the three control methods are shown in Table 1:
[0184] Method Operating cost / cents Proportion of room temperature exceeding limit Strategy generation speed / ms DQN 1147.14 67.87% 349.36 DDPG 875.80 19.89% 172.16 PPO 773.32 15.19% 89.33
[0185] Table 1 Comparison of control results of three control methods
[0186] As can be seen from Table 1, the operating cost of the building flexible energy supply system of the present invention is 773.32 cents, which is 32.6% lower than the DQN method and 11.7% lower than the DDPG method. It can be seen that the present invention has significant economic advantages. The over-limit ratio of the room temperature of the present invention is only 15.19%, which is 51.68% lower than the DQN method and 4.70% lower than the DDPG method. It can be seen that the present invention is beneficial to the comfort of the building interior. In terms of optimizing the control strategy generation speed, the single-step strategy generation time of the PPO method is 89.33 ms, which is 74.43% and 48.11% higher than the DQN and DDPG methods respectively, demonstrating the advantages of the PPO method in terms of computational efficiency and strategy generation speed, and can meet the real-time requirements of the dynamic optimization of the building flexible energy supply system.
[0187] The P V variation under different control methods is as Figure 5 shown. Combining Figure 3Analysis shows that in the low electricity price period (such as 1:00 to 5:00), the IAC output is high and the VESS is charged; while in the high electricity price period (such as 11:00 to 13:00), the IAC output is low and the VESS discharges to compensate for the reduction in the purchased electricity power of the building flexible energy supply system. In terms of the average power throughout the day, the present invention also reaches the lowest. Therefore, the present invention is beneficial to the building flexible energy supply system to save its electricity cost.
[0188] SOC under different control methods V The change situation is as Figure 6 shown. The discrete action control method adopted by the DQN method causes great fluctuations in the SOC. V At 5 o'clock on the same day, the indoor temperature reaches 29.9 °C and the SOC V is 2.2. While at 11 o'clock on the same day, the indoor temperature reaches 14.6 °C and the SOC V is -2.1. This shows that the control accuracy of this method is relatively low and also shows the effectiveness of continuous action control. After further comparing with the DDPG method, the results show that the DDPG method performs more stably in control and the volatility of the SOC V is not high. The SOC of the present invention V has the most stable fluctuations. Among the 24 hours, there are 21 hours when the SOC V is within the constraint range, while the DQN method has 18 hours when the SOC V is within the constraint range. The overall improvement ratio is 16.7%. This also shows the overall improvement effect of the present invention in dealing with SOC V violation and over-limit.
[0189] In summary, based on the thermal inertia characteristics of the building envelope structure, the present invention constructs a model of the building virtual energy storage system VESS, and introduces time-varying parameters through the model of the building virtual energy storage system VESS; according to the set control objectives, using the time-varying parameters introduced by the model of the building virtual energy storage system VESS, constructs an optimized control model of the building flexible energy supply system combined with VESS; transforms the optimized control model of the building flexible energy supply system combined with VESS into a Markov decision process model; optimizes the Markov decision process model through the proximal policy optimization algorithm, and outputs a control strategy. The present invention accurately transforms the building flexible energy supply system model combined with VESS into an MDP model, and uses the PPO algorithm to achieve efficient optimized control of the building flexible energy supply system. The present invention realizes the generation of an optimized control strategy at the millisecond level, and the speed is significantly better than the DQN and DDPG methods, which reflects its high efficiency in dealing with the dynamic optimization problems of multi-variable and non-linear building flexible energy supply systems, and can meet the real-time requirements of the optimized control of the building flexible energy supply system; at the same time, in a dynamic meteorological environment, the present invention effectively reduces the operating cost of the building flexible energy supply system while maintaining a low proportion of room temperature exceeding the limit, which reflects its significant advantages in economy and comfort.
[0190] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario, and multiple devices cooperate with each other to complete it. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0191] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order from those in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0192] Embodiment 2
[0193] See Figure 7 , Embodiment 2 of the present invention also provides a proximal policy optimization device for a building flexible energy supply system, including:
[0194] A building virtual energy storage system VESS model construction module 001, configured to construct a building virtual energy storage system VESS model based on the thermal inertia characteristics of the building envelope structure, and introduce time-varying parameters through the building virtual energy storage system VESS model;
[0195] The construction module 002 of the optimal control model for the building flexible energy supply system combined with VESS is used to construct the optimal control model for the building flexible energy supply system combined with VESS according to the set control objectives and by using the time-varying parameters introduced by the building virtual energy storage system VESS model;
[0196] The Markov decision process model conversion module 003 is used to convert the optimal control model for the building flexible energy supply system combined with VESS into a Markov decision process model;
[0197] The Markov decision process model optimization module 004 is used to optimize the Markov decision process model through the proximal policy optimization algorithm and output the control strategy.
[0198] In this embodiment, in the building virtual energy storage system VESS model construction module 001, during the construction of the building virtual energy storage system VESS model, according to the performance characteristics of the electrical energy storage, three time-varying parameters, namely virtual power, virtual capacity, and virtual state of charge, are introduced;
[0199] The relational expression among the three time-varying parameters is:
[0200]
[0201] In the formula, t is the t-th control period; Δt is the control time interval; SOC V (t) is the virtual state of charge of VESS at time t; C V (t) is the virtual capacity of VESS at time t, kWh; P V (t) is the virtual power of VESS at time t, kW.
[0202] In this embodiment, in the construction module 002 of the optimal control model for the building flexible energy supply system combined with VESS, the objective function expression of the optimal control model for the building flexible energy supply system combined with VESS is:
[0203] C = C e + C SOC
[0204]
[0205] In the formula, C is the electricity cost of the building flexible energy supply system; C e is the operating cost of the building flexible energy supply system; C SOC is the SOC penalty cost of VESS; P base (t) is the reference power consumption at time t, kW; P other (t) is the power of other electrical equipment at time t, kW; P PV(t) is the PV power in period t, in kW; c(t) is the electricity price in period t, in cents / kW; σ is the SOC V Overlimit cost coefficient; T in (t) is the indoor temperature of the building in period t, in °C; T max is the maximum indoor temperature, in °C; T min is the minimum indoor temperature, in °C.
[0206] In this embodiment, in the building flexible energy supply system optimization control model construction module 002 combined with VESS, the constraints of the building flexible energy supply system optimization control model combined with VESS include: VESS constraints, PV output power constraints, and building flexible energy supply system electric power balance constraints;
[0207] The expression of the VESS constraint is:
[0208] 0 ≤ SOC V (t) ≤ 1
[0209] -P dismax (t) ≤ P V (t) ≤ P cmax (t)
[0210] In the formula, P dismax (t) is the maximum discharge power of VESS in the t-th period, in kW; P cmax (t) is the maximum charge power of VESS in the t-th period, in kW;
[0211] The expression of the PV output power constraint is:
[0212] 0 ≤ P PV (t) ≤ P PVf (t)
[0213] In the formula, P PVf (t) is the maximum predicted output power of PV in the t-th period, in kW;
[0214] The expression of the building flexible energy supply system electric power balance constraint is:
[0215] P com (t) + P PV (t) = P base (t) + P V (t) + P other (t)
[0216] In the formula, P com (t) is the purchased electric power, in kW.
[0217] In this embodiment, in the Markov decision process model conversion module 003, during the process of converting the optimized control model of the building flexible energy supply system combined with VESS into the Markov decision process model, the state variables, control actions, reward function, and transition function are defined;
[0218] The expression of the state variable is:
[0219] s(t) = [SOC V (t), T in (t), T out (t), Q solar (t), c(t)]
[0220] In the formula, s(t) is the state information at the current moment; T out (t) is the outdoor temperature in the t period, °C; Q solar (t) is the solar heat gain power of the building in the t period, kW;
[0221] The expression of the control action is:
[0222] a(t) = [P V (t), P PV (t)]
[0223] In the formula, a(t) represents the action adopted at the current moment;
[0224] The expression of the reward function is:
[0225] r(t) = w1·r elec (t) + w2·r soc (t)
[0226] In the formula, r(t) is the total reward at the current moment; r elec (t) is the electricity cost reward; r soc (t) is the SOC reward of VESS; w1 and w2 are both weight coefficients.
[0227] In this embodiment, in the Markov decision process model optimization module 004, during the process of optimizing the Markov decision process model through the proximal policy optimization algorithm, the proximal policy optimization algorithm realizes the policy gradient theorem through the Actor-Critic structure; the optimization objective function of the Actor-Critic structure is:
[0228]
[0229] In the formula, E(t) is the expectation of the loss function at the t moment; clip(·) is the clipping function; ε is the hyperparameter of clip(·); π θThe policy of the Actor network under parameter θ; r θ (t) is importance sampling, which measures the difference between the new policy and the old policy before and after the update; A(t) is the advantage function; is the policy entropy; β is the entropy regularization term, which is used to balance the exploitation of the policy by the control system and the exploration of unknown policies;
[0230] The loss function of the Actor-Critic structure is:
[0231] L Critic (φ) = E(t)[(V φ (s(t)) - R(t)) 2 )
[0232] In the formula, φ is the Critic network parameter; V φ (s(t)) is the state value function, that is, the judgment of the control system on the quality of the current state; R(t) is the cumulative reward, which for the building flexible energy supply system is the sum of the rewards of the system from the current state to the future state during operation.
[0233] It should be noted that the information interaction, execution process, etc. between the above system modules, because they are based on the same concept as the method embodiment in Embodiment 1 of this application, the technical effects brought by them are the same as those of the method embodiment of this application. For specific content, please refer to the description in the method embodiment shown above in this application, and details will not be repeated here.
[0234] Embodiment 3
[0235] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code of a proximal policy optimization method for a building flexible energy supply system is stored, and the program code includes instructions for executing a proximal policy optimization method for a building flexible energy supply system in Embodiment 1 or any possible implementation manner thereof.
[0236] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center integrating one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0237] Embodiment 4
[0238] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0239] The processor and the memory complete communication with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute a proximal strategy optimization method for a building flexible energy supply system according to Embodiment 1 or any possible implementation manner thereof by invoking the program instructions.
[0240] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0241] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.).
[0242] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing system. They can be concentrated on a single computing system or distributed on a network composed of multiple computing systems. Optionally, they can be implemented by program code executable by the computing system, so that they can be stored in a storage system and executed by the computing system. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0243] Although the present invention has been described in detail above with general descriptions and specific embodiments, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of the present invention claimed.
Claims
1. A proximal strategy optimization method for a building flexible energy supply system, characterized in that: include: Based on the thermal inertia characteristics of building envelope structures, a building virtual energy storage system VESS model is constructed, and time-varying parameters are introduced through the building virtual energy storage system VESS model; According to the set control objectives, the time-varying parameters introduced by the building virtual energy storage system (VESS) model are used to construct an optimization control model for the building flexible energy supply system combined with VESS. Converting the optimization control model of the building flexible energy supply system combined with VESS into a Markov decision process model; The Markov decision process model is optimized by a proximal strategy optimization algorithm, and a control strategy is output.
2. The proximal strategy optimization method for a building flexible energy supply system according to claim 1 is characterized in that: In the process of constructing the building virtual energy storage system VESS model, three time-varying parameters, virtual power, virtual capacity and virtual state of charge, are introduced according to the performance characteristics of electric energy storage; The relationship between the three time-varying parameters is expressed as: Where, t is the tth control period; Δt is the control time interval; SOC V (t) is the virtual state of charge of the VESS during period t; C V (t) is the virtual capacity of VESS during period t; P V (t) is the virtual power of VESS during period t.
3. The proximal strategy optimization method for a building flexible energy supply system according to claim 2 is characterized in that: The objective function expression of the optimization control model of the building flexible energy supply system combined with VESS is: C=C e +C SOC Where, C is the electricity cost of the building flexible energy supply system; C e The operating cost of the building's flexible energy supply system; C SOC is the SOC penalty cost of VESS; P base (t) is the reference power consumption during period t; P other (t) is the power of other electrical equipment during period t; P PV (t) is the PV power in period t; c(t) is the electricity price in period t; σ is the SOC V Cross-limit cost coefficient; T in (t) is the indoor temperature of the building during period t; T max is the maximum indoor temperature; T min The lowest indoor temperature.
4. The method for optimizing the proximal strategy of a building flexible energy supply system according to claim 3 is characterized in that: The constraints of the building flexible energy supply system optimization control model combined with VESS include: VESS constraints, PV output power constraints and building flexible energy supply system electric power balance constraints; The expression of the VESS constraint is: 0≤SOC V (t)≤1 -P dismax (t)≤P V (t)≤P cmax (t) Where P dismax (t) is the maximum discharge power of VESS in the tth period; P cmax (t) is the maximum charging power of VESS in period t; The expression of the PV output power constraint is: 0≤P PV (t)≤P PVf (t) Where P PVf (t) is the maximum predicted output power of PV during period t; The expression of the electric power balance constraint of the building flexible energy supply system is: P com (t)+P PV (t)=P base (t)+P V (t)+P other (t) Where P com (t) is the purchased electricity power.
5. The method for optimizing the proximal strategy of a building flexible energy supply system according to claim 4 is characterized in that: In the process of converting the optimization control model of the building flexible energy supply system combined with VESS into the Markov decision process model, the state variables, control actions, reward functions and transfer functions are defined; The expression of the state variable is: s(t)=[SOC V (t),T in (t),T out (t),Q solar (t),c(t)] In the formula, s(t) is the state information at the current moment; T out (t) is the outdoor temperature during period t; Q solar (t) is the solar heat gain of the building during period t; The expression of the control action is: a(t)=[P V (t),P PV (t)] In the formula, a(t) represents the action taken at the current moment; The expression of the reward function is: r(t)=w1·r elec (t)+w2·r soc (t) Where r(t) is the total reward at the current moment; r elec (t) is the electricity cost reward; r soc (t) is the SOC reward of VESS; w1 and w2 are weight coefficients.
6. A proximal strategy optimization method for a building flexible energy supply system according to claim 5, characterized in that: In the process of optimizing the Markov decision process model by the proximal policy optimization algorithm, the proximal policy optimization algorithm implements the policy gradient theorem through the Actor-Critic structure; the optimization objective function of the Actor network and the Critic network in the Actor-Critic structure is: Where E(t) is the expectation of the loss function at time t; clip(·) is the clipping function; ε is the hyperparameter of clip(·); π θ is the strategy of the Actor network under parameter θ; r θ (t) is the importance sampling, which measures the difference between the new strategy and the old strategy before and after the update; A(t) is the advantage function; is the strategy entropy; β is the entropy regularization term, which is used to balance the control system’s use of the strategy and the exploration of unknown strategies; The loss function of the Critic network is: L Critic (φ)=E(t)[(V φ (s(t))-R(t)) 2 ] Where φ is the Critic network parameter; V φ (s(t)) is the state value function, that is, the control system's judgment on the quality of the current state; R(t) is the cumulative return, which is the sum of the rewards from the current state to the future state during the operation of the building flexible energy supply system.
7. A proximal strategy optimization device for a building flexible energy supply system, using a proximal strategy optimization method for a building flexible energy supply system according to any one of claims 1 to 6, characterized in that: include: A building virtual energy storage system VESS model construction module is used to construct a building virtual energy storage system VESS model based on the thermal inertia characteristics of the building envelope structure, and introduce time-varying parameters through the building virtual energy storage system VESS model; The module for constructing an optimal control model of a building flexible energy supply system combined with VESS is used to construct an optimal control model of a building flexible energy supply system combined with VESS according to the set control objectives and by using the time-varying parameters introduced by the building virtual energy storage system VESS model; A Markov decision process model conversion module, in which the user converts the building flexible energy supply system optimization control model combined with VESS into a Markov decision process model; The Markov decision process model optimization module is used to optimize the Markov decision process model through a proximal strategy optimization algorithm and output a control strategy.
8. The proximal strategy optimization device for a building flexible energy supply system according to claim 7 is characterized in that: In the building virtual energy storage system VESS model construction module, in the process of constructing the building virtual energy storage system VESS model, three time-varying parameters of virtual power, virtual capacity and virtual state of charge are introduced according to the performance characteristics of electric energy storage; The relationship between the three time-varying parameters is expressed as: Where, t is the tth control period; Δt is the control time interval; SOC V (t) is the virtual state of charge of the VESS during period t; C V (t) is the virtual capacity of VESS during period t, kWh; P V (t) is the virtual power of VESS during period t.
9. The proximal strategy optimization device for a building flexible energy supply system according to claim 8, characterized in that: In the construction module of the building flexible energy supply system optimization control model combined with VESS, the objective function expression of the building flexible energy supply system optimization control model combined with VESS is: C=C e +C SOC Where, C is the electricity cost of the building flexible energy supply system; C e The operating cost of the building's flexible energy supply system; C SOC is the SOC penalty cost of VESS; P base (t) is the reference power consumption during period t; P other (t) is the power of other electrical equipment during period t; P PV (t) is the PV power in period t, kW; c(t) is the electricity price in period t, cents / kW; σ is SOC V Cross-limit cost coefficient; T in (t) is the indoor temperature of the building during period t; T max is the maximum indoor temperature; T min The lowest indoor temperature.
10. A proximal strategy optimization device for a building flexible energy supply system according to claim 9, characterized in that: In the construction module of the building flexible energy supply system optimization control model combined with VESS, the constraints of the building flexible energy supply system optimization control model combined with VESS include: VESS constraints, PV output power constraints and building flexible energy supply system electric power balance constraints; The expression of the VESS constraint is: 0≤SOC V (t)≤1 -P dismax (t)≤P V (t)≤P cmax (t) Where P dismax (t) is the maximum discharge power of VESS in the tth period; P cmax (t) is the maximum charging power of VESS in period t; The expression of the PV output power constraint is: 0≤P PV (t)≤P PVf (t) Where P PVf (t) is the maximum predicted output power of PV during period t; The expression of the electric power balance constraint of the building flexible energy supply system is: P com (t)+P PV (t)=P base (t)+P V (t)+P other (t) Where P com (t) is the purchased electricity power.
11. A non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores a program code for a proximal strategy optimization method for a building flexible energy supply system, characterized in that: The program code includes instructions for executing a proximal strategy optimization method for a building flexible energy supply system as described in any one of claims 1 to 6.
12. An electronic device comprising: Memory and processor; The processor and the memory communicate with each other via a bus; The memory stores program instructions that can be executed by the processor, and is characterized in that the processor calls the program instructions to execute a proximal strategy optimization method for a building flexible energy supply system as described in any one of claims 1 to 6.
Citation Information
Cited By
Multivariable collaborative optimization method for energy-saving operation strategy of heat pump machine room of public building
CN121276984A