Building energy system double-layer optimization method based on deep reinforcement learning

By adopting a two-layer optimization method of deep reinforcement learning in building energy systems, the hierarchical structure collaboratively optimizes capacity planning and real-time scheduling, the problems of complex modeling and difficult solving in the existing technology are solved, and the dual goals of full-cycle optimization and economic and environmental protection are achieved.

CN120146268APending Publication Date: 2025-06-13TIANJIN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510203105.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When building energy systems deal with complex and variable working conditions, the existing technology has problems of complex modeling and difficulty in solving them, and it is difficult to coordinately optimize long-term equipment investment and short-term dynamic scheduling.

Method used

The two-layer optimization method based on deep reinforcement learning is adopted, through cluster analysis, simulated annealed particle swarm algorithm and multi-agent near-end strategy algorithm, the upper layer processing capacity planning (annual scale), and the lower layer processing real-time scheduling (daily/hour scale), and the equipment capacity configuration and scheduling strategy are coordinated.

Benefits of technology

The "planning-operation" full-cycle optimization of the building energy system has been achieved, which reduces annual value costs, improves the economic and environmental protection of the system, and resolves decision-making conflicts coupled by multiple time scales.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146268A_ABST
    Figure CN120146268A_ABST
Patent Text Reader

Abstract

The invention discloses a building energy system double-layer optimization method based on deep reinforcement learning. The method comprises the following steps: S1, extracting load data of a building energy system by adopting a clustering analysis method to construct an energy coupling model; s2, optimizing the coupled building energy model through a simulated annealing particle swarm optimization algorithm to construct an upper-layer energy capacity configuration model; s3, optimizing the coupled building energy model through a multi-agent near-end strategy algorithm to construct a lower-layer energy strategy scheduling model; s4, the upper-layer energy capacity configuration model obtains equipment capacity configuration information; s5, the lower-layer energy strategy scheduling model outputs an optimal capacity configuration and scheduling strategy of the building energy system according to the equipment capacity configuration information; according to the invention, through algorithm innovation and model refined design, planning-operation full-cycle optimization of the building energy system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of building energy systems, and more specifically to a two-layer optimization method for building energy systems based on deep reinforcement learning. Background Art

[0002] With the progress of efficient power generation technologies and energy storage devices, the role of buildings in the energy system is gradually changing: the traditional building energy system aims to meet the energy demand of buildings, but in future energy consumption scenarios, the building energy system will achieve a "trinity" model of energy production, consumption, and regulation, leading to a new pattern in which the supply-demand relationship changes from "demand determines supply" to "supply guides demand". However, the complex composition of building energy systems and the introduction of a high proportion of renewable energy pose significant challenges to ensuring a scheduling strategy for achieving supply-demand matching. Simplifying the physical models of energy conversion devices is a common approach in the current field of building energy system optimization. However, when the number and types of energy devices increase, continuing to ignore the actual situation where the efficiency or performance coefficient of different devices fluctuates with operating conditions such as the load rate may lead to a large gap between the planning scheme and actual operation. Existing technologies have problems of complex modeling and difficult solution when dealing with such complex variable operating conditions. Therefore, there is an urgent need for an efficient, accurate, and easy-to-use optimization method to achieve the coordinated optimization of the capacity configuration and scheduling strategy of building energy systems. Summary of the Invention

[0003] To solve the problems of the existing technology, the present invention provides a two-layer optimization method for building energy systems based on deep reinforcement learning. This method solves the multi-time scale coupling technical problem in the existing technology of difficult coordinated optimization of long-term equipment investment (upper-layer annual investment cost) and short-term dynamic scheduling (lower-layer operating cost) in building energy systems; in particular, the present invention solves the decision-making conflict caused by the coupling of the two by adopting a hierarchical structure in the building energy system, that is, the upper layer processes capacity planning (annual scale), and the lower layer processes real-time scheduling (daily / hourly scale).

[0004] The present invention adopts the following technical solutions:

[0005] S1. Adopt a clustering analysis method to extract the load data of the building energy system and construct an energy coupling model;

[0006] S2. Optimize the coupled building energy model through a simulated annealing particle swarm algorithm to construct an upper-layer energy capacity configuration model;

[0007] S3. Optimize the coupled building energy model through a multi-agent proximal policy algorithm to construct a lower-layer energy strategy scheduling model;

[0008] S4. The upper-layer energy capacity configuration model obtains equipment capacity configuration information;

[0009] S5. The lower-layer energy strategy scheduling model outputs the optimal capacity configuration and scheduling strategy of the building energy system according to the device capacity configuration information;

[0010] Furthermore, taking five variables of the building's electrical load, heat load, cooling load, photovoltaic power generation per unit value, and wind turbine power generation per unit value as characteristic parameters, after normalization, the K-Means clustering algorithm is used for clustering, and the sum of squared errors is used as the evaluation index of the clustering result to obtain the number of typical days and the typical day load data.

[0011] Furthermore, the solar photovoltaic power generation model is expressed as follows:

[0012]

[0013] where η PV is the photovoltaic power generation per unit value; df PV is the power reduction factor; SI PV is the solar irradiance intensity; tc represents the temperature coefficient; T PV is the photovoltaic surface temperature.

[0014] The wind turbine power generation model is expressed as follows:

[0015]

[0016] where η WT is the wind turbine power generation per unit value; P WT is the rated power of the wind turbine; v, v ci , v r , v co are the current ambient wind speed, cut-in wind speed, rated wind speed, and cut-out wind speed respectively.

[0017] The influence of the device load rate on the actual performance coefficient is expressed as follows:

[0018]

[0019] where η is the actual performance coefficient of the device; b is the constant term; k is the fitting coefficient; PLR is the device load rate.

[0020] The calculation formula for normalization is expressed as follows:

[0021]

[0022] where x i,norm is the normalized sample data; is the average value of the sample data; σ is the standard deviation of the sample data.

[0023] And, the K-Means clustering algorithm is expressed as follows:

[0024]

[0025] Among them, d is the Euclidean distance; x is an individual of the clustering sample; n is the sample dimension, and C i is the clustering center.

[0026] The calculation formula of the sum of squared errors, an evaluation index of the clustering result, is expressed as follows:

[0027]

[0028] Among them, SSE is the sum of squared errors.

[0029] Furthermore, the energy coupling model constructed in step S1 is:

[0030]

[0031] Among them, L is the total output energy matrix of the energy hub; C is the coupling matrix; P is the total input energy matrix of the energy hub.

[0032] Among them, the coupling factor c in the coupling matrix C is composed of the efficiency coefficient and the distribution coefficient of each energy device, and is expressed as follows:

[0033]

[0034] Among them, η is the efficiency coefficient of each energy conversion device; x is the distribution coefficient of the upper-level input energy to the lower level.

[0035] Furthermore, the simulated annealing process is introduced into the particle swarm algorithm to form a particle swarm algorithm based on simulated annealing for the capacity configuration search of the upper layer, and the capacity range limit, the number of particles, the learning rate, the simulated annealing speed, and the inertia weight of each device are set. The specific form is:

[0036]

[0037] Among them, v is the particle velocity; x is the particle position; ω is the inertia weight; c 1 and c 2 are the learning rates; r is a random number between (0, 1).

[0038] Furthermore, the simulated annealing process is used to narrow the search range. The specific form is:

[0039] T = T · v SA

[0040] Among them, T is the annealing temperature; v SA is the annealing speed.

[0041] Furthermore, a lower-layer energy strategy scheduling model is constructed. Among them, the annual scheduling cost of the lower-layer system operation, including the system scheduling cost, equipment expansion cost, and carbon trading cost introduced to limit carbon emissions, is expressed as follows:

[0042]

[0043] Among them, is the annual operation cost; is the annual maintenance cost; C inv is the expansion cost of, ¥; is the annual carbon trading cost; φ is the occurrence probability of a typical day; P is the price; U onoff is the total number of start-stop times of the equipment. Among them, is determined by the carbon quota, actual carbon emissions, and stepped carbon trading price, and is expressed as follows:

[0044]

[0045] Among them, C quota is the carbon emission quota; C actual is the actual carbon emission; C trade is the total carbon trading amount; δ E and δ H are the carbon emission quotas of the power supply equipment and the heat supply equipment respectively; μ is the conversion coefficient for converting the power generation of the cogeneration unit into heat supply; σ E and σ NG are the carbon emission factors of the power grid and natural gas respectively.

[0046] Furthermore, the equipment capacity configuration information is obtained according to the following formula;

[0047]

[0048] Among them: IC is the equipment capacity; rand(min, max) is the equipment capacity randomly generated within the constraint range; TC ann is the annual value cost of the system; F r is the equipment investment correction coefficient; P is the unit capacity cost of the equipment; y is the equipment life; r is the discount rate; is the lower-layer evaluation function;

[0049] Furthermore, the optimal capacity configuration and scheduling strategy of the building energy system are calculated according to the following formula;

[0050]

[0051] Among them, α is the learning rate; Q ω (s, a) is the value function estimated by Critic; β is the learning rate; r is the immediate reward; γ is the discount factor; s′ is the next state.

[0052] Beneficial effects

[0053] 1. By implementing a hierarchical structure in the building energy system, with the upper layer handling capacity planning (annual scale) and the lower layer handling real-time scheduling (daily / hourly scale), the present invention solves the technical problem of the difficulty in the prior art to synergistically optimize the multi-time scale coupling of long-term equipment investment (upper layer annual investment cost) and short-term dynamic scheduling (lower layer operating cost).

[0054] 2. Through the upper layer simulated annealing particle swarm optimization algorithm and the lower layer multi-agent distributed decision-making, the present invention optimizes the coordination of multiple types of equipment including photovoltaic, energy storage, and combined heat and power in the building energy system.

[0055] 3. By guiding the upper layer to invest in low-carbon equipment (such as high-efficiency units) through dynamic carbon cost feedback and integrating the stepped carbon price mechanism into the lower layer objective function, the present invention realizes the dual-objective optimization of economy and environmental protection.

[0056] 4. Through algorithm innovation and refined model design, the present invention realizes the full-cycle optimization of the "planning - operation" of the building energy system, providing a practical solution for low-carbon intelligent buildings. Brief description of the drawings

[0057] Figure 1 It is a schematic diagram of a two-layer optimization method for a building energy system based on deep reinforcement learning according to the present invention.

[0058] Figure 2 It is the electricity price and gas price diagram according to the present invention.

[0059] Figure 3 It is the load demand diagram after clustering based on the K-Means algorithm according to the present invention.

[0060] Figure 4 It is a schematic diagram of constructing an energy hub model according to the present invention. Detailed implementation manners

[0061] The following is an explanation of the present invention in conjunction with the attached Figure 1 ~ attached Figure 4 The present invention is described as follows:

[0062] As Figure 1 shown, the present invention discloses a two-layer optimization method for a building energy system based on deep reinforcement learning, and the method includes:

[0063] S1. The process of extracting the load data of the building energy system by using the clustering analysis method to construct an energy coupling model;

[0064] Based on five variables, namely the electrical load, heat load, cooling load, per-unit value of photovoltaic power generation, and per-unit value of wind turbine power generation of the building, after normalization, the K-Means clustering algorithm is used for clustering, and the sum of squared errors is used as the evaluation index of the clustering result to obtain the number of typical days and typical day load data.

[0065] Furthermore, the solar photovoltaic power generation model is expressed as follows:

[0066]

[0067] Among them, η PV is the per-unit value of photovoltaic power generation; df PV is the power reduction factor; SI PV is the solar irradiance intensity; tc represents the temperature coefficient; T PV is the photovoltaic surface temperature.

[0068] The wind turbine power generation model is expressed as follows:

[0069]

[0070] Among them, η WT is the per-unit value of wind turbine power generation; P WT is the rated power of the wind turbine; v, v ci , v r , v co are the current ambient wind speed, cut-in wind speed, rated wind speed, and cut-out wind speed respectively.

[0071] The influence of the equipment load rate on the actual performance coefficient is expressed as follows:

[0072]

[0073] Among them, η is the actual performance coefficient of the equipment; b is the constant term; k is the fitting coefficient; PLR is the equipment load rate.

[0074] The calculation formula for normalization is expressed as follows:

[0075]

[0076] Among them, x i,norm is the normalized sample data; is the average value of the sample data; σ is the standard deviation of the sample data.

[0077] And, the K-Means clustering algorithm is expressed as follows:

[0078]

[0079] Among them, d is the Euclidean distance; x is the individual of the clustering sample; n is the sample dimension, Ci is the clustering center.

[0080] The calculation formula of the sum of squared errors, an evaluation index of the clustering result, is expressed as follows:

[0081]

[0082] Among them, SSE is the sum of squared errors.

[0083] The energy coupling model constructed in step S1 is:

[0084]

[0085] Among them, L is the total output energy matrix of the energy hub; C is the coupling matrix; P is the total input energy matrix of the energy hub.

[0086] Among them, the coupling factor c in the coupling matrix C is composed of the efficiency coefficient and the distribution coefficient of each energy device, and is expressed as follows:

[0087]

[0088] Among them, η is the efficiency coefficient of each energy conversion device; x is the distribution coefficient of the upper-level input energy to the lower level.

[0089] S2. Optimize the process of constructing the upper-level energy capacity configuration model through the simulated annealing particle swarm algorithm; where:

[0090] Introduce the simulated annealing process into the particle swarm algorithm to form a particle swarm algorithm based on simulated annealing for upper-level capacity configuration search, and set the capacity range limit, the number of particles, the learning rate, the simulated annealing speed, and the inertia weight of each device. The specific form is:

[0091]

[0092] Among them, v is the particle velocity; x is the particle position; ω is the inertia weight; c 1 and c 2 are the learning rates; r is a random number between (0, 1).

[0093] Then, use the simulated annealing process to narrow the search range. The specific form is:

[0094] T = T·v SA

[0095] Among them, T is the annealing temperature; v SA is the annealing speed.

[0096] S3. Optimize the lower-layer energy strategy scheduling model by constructing the coupled building energy model through the multi-agent proximal policy algorithm; among them, the annual scheduling cost of the lower-layer system operation, including system scheduling cost, equipment expansion cost, and carbon trading cost introduced to limit carbon emissions, is expressed as follows:

[0097]

[0098] Among them, is the annual operation cost; is the annual maintenance cost; C inv is the expansion cost of, ¥; is the annual carbon trading cost; φ is the probability of a typical day; P is the price; U onoff is the total number of starts and stops of the equipment. Among them, is determined by carbon quota, actual carbon emissions, and stepped carbon trading price, and is expressed as follows:

[0099]

[0100] Among them, C quota is the carbon emission quota; C actual is the actual carbon emission; C trade is the total carbon trading amount; δ E and δ H are the carbon emission quotas of power supply equipment and heat supply equipment respectively; μ is the conversion coefficient for converting the power generation of a cogeneration unit into heat supply; σ E and σ NG are the carbon emission factors of the power grid and natural gas respectively.

[0101] S4. The upper-layer energy capacity configuration model obtains equipment capacity configuration information according to the following formula;

[0102]

[0103] Among them: IC is the equipment capacity; rand(min, max) is the equipment capacity randomly generated within the constraint range; TC ann is the annual value cost of the system; F r is the equipment investment correction coefficient; P is the unit capacity cost of the equipment; y is the equipment life; r is the discount rate; is the lower-layer evaluation function;

[0104] S5. The lower-layer energy strategy scheduling model calculates and outputs the optimal capacity configuration and scheduling strategy of the building energy system according to the equipment capacity configuration information according to the following formula;

[0105]

[0106] ​

[0107] where α is the learning rate; Q ω (s,a) is the value function estimated by the Critic; β is the learning rate; r is the immediate reward; γ is the discount factor; s′ is the next state.

[0108] Example 2

[0109] Taking the capacity configuration of a certain building energy system as an example, the electricity price and natural gas price are as Figure 2 shown, and the clustering results of the building load using the K-Means algorithm are as Figure 3 shown. The main selected equipment includes combined heat and power units, photovoltaic, wind turbines, electricity storage equipment, heat storage equipment, absorption chillers, and ground source heat pumps. The alternative equipment is electric boilers and electric chillers. The energy equipment parameters are shown in Table 1;

[0110] Table 1 Energy Equipment Parameters

[0111]

[0112]

[0113] As Figure 3 shown, the multi-energy coupling model of the energy system is constructed as follows:

[0114] Step 1: Insert energy nodes at positions involving multiple energy interactions, such as Figure 3 the nodes represented by 1 and 2 in, ensuring that each energy node or energy conversion device only involves the input of a single energy type or the output of a single energy type.

[0115] Step 2: Module the energy hub model into small models according to the longest path of energy flow; since the empty nodes in the energy transmission path will cause the adjacent coupling matrices to be mismatched, that is, mathematically, due to the mismatch of the rows and columns of the adjacent coupling matrices, multiplication operations cannot be performed. At the empty nodes in Figure 3 , virtual nodes with both the efficiency coefficient and the distribution coefficient being 1 are introduced to achieve full coupling of adjacent small models.

[0116] Step 3: By calculating and multiplying the coupling matrices of each small model one by one, the entire energy hub coupling model is obtained as shown below:

[0117]

[0118] Among them:

[0119]

[0120] After obtaining the energy hub coupling matrix, deploy it in the multi-agent proximal policy optimization algorithm environment and debug the hyperparameters of the two-layer optimization algorithm to ensure the algorithm performance, including the hyperparameters of the simulated annealing particle swarm algorithm: the capacity range limit of each device, the number of particles, the learning rate, the simulated annealing speed, and the inertia weight; the hyperparameters of the multi-agent proximal policy optimization algorithm: the learning rate, the discount coefficient, the entropy coefficient, the truncation coefficient, the number of hidden layers of the neural network, and the batch size.

[0121] When the performance of the two-layer optimization algorithm is relatively good, train it repeatedly for many times to ensure the stability of the two-layer optimization algorithm. Otherwise, repeat the above hyperparameter adjustment process.

[0122] Compared with other algorithms, as shown in Table 2, although the particle swarm algorithm is superior to the simulated annealing algorithm in terms of convergence speed, it is more likely to fall into local optimal solutions, which limits its optimization performance. At the same time, the single-layer simulated annealing particle swarm algorithm that combines the two algorithms can find a better solution, but it will increase the training time. In contrast, the two-layer model can significantly reduce the training time. Generally speaking, the proposed two-layer optimization algorithm can reduce the annual value cost by up to 22.55%, and the algorithm convergence speed can be increased by up to 36.9%.

[0123] Table 2 Comparison results of algorithms

[0124]

[0125] Compared with the conventional model that does not consider the fluctuation of the device performance coefficient, as shown in Table 2, the average load rate of the combined heat and power generation in the model considering the fluctuation is higher. Therefore, it has a higher average power generation or heating efficiency and is configured with more capacity. However, due to the more complex working conditions, more energy storage devices need to be configured for regulation. Generally speaking, considering the variable load characteristics of the device can give full play to the device performance more fully, thus avoiding redundant or insufficient capacity configuration of the device, and having higher economy and environmental protection. Among them, the capacities of the electric boiler and the electric chiller as alternative devices are 0.

[0126] Table 3 Device capacity configuration results under different models (unit: kW)

[0127]

[0128]

[0129] Although the present invention has been described above, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many variations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A two-level optimization method for building energy system based on deep reinforcement learning, characterized in that: The method comprises include: S1. Use cluster analysis method to extract load data of building energy system and construct energy coupling model; S2. Optimizing the coupled building energy model by using simulated annealing particle swarm algorithm to construct an upper energy capacity configuration model; Where: v is the particle velocity; x is the particle position; ω is the inertia weight; c1 and c2 are the learning rates; r is a random number between (0,1); S3, optimizing the coupled building energy model through a multi-agent proximal strategy algorithm to construct a lower-level energy strategy scheduling model; in, is the annual operating cost; is the annual maintenance cost; C inv Expansion cost, ¥; is the annual carbon trading cost; φ is the probability of a typical day; P is the price; U onoff The total number of starts and stops of the equipment; S4. The upper-layer energy capacity configuration model obtains equipment capacity configuration information according to the following formula; Where: IC is the device capacity; rand(min,max) is the device capacity within the randomly generated constraint range; TC ann is the annual cost of the system; F r is the equipment investment correction coefficient; P is the equipment unit capacity cost; y is the equipment life; r is the discount rate; is the lower layer evaluation function; S5. The lower-level energy strategy scheduling model calculates and outputs the optimal capacity configuration and scheduling strategy of the building energy system according to the following formula based on the equipment capacity configuration information: Among them, α is the learning rate; Q ω (s,a) is the value function estimated by the Critic; β is the learning rate; r is the immediate reward; γ is the discount factor; s′ is the next state.

2. The two-layer optimization method for building energy system based on deep reinforcement learning according to claim 1 is characterized in that: The step S1 constructs an energy coupling model as follows: Among them, L is the total output energy matrix of the energy hub; C is the coupling matrix; P is the total input energy matrix of the energy hub. Among them, the coupling factor c in the coupling matrix C is composed of the efficiency coefficient and distribution coefficient of each energy device, which is expressed as follows: Among them, η is the efficiency coefficient of each energy conversion device; x is the distribution coefficient of the input energy of the previous level to the next level.

Citation Information

Cited By

  • Existing building regeneration optimization design method based on quantum annealing algorithm

    CN122490682A