Energy collaborative operation method of building group and related device

By combining deep strategy models and deep value networks, the energy response of building clusters is optimized, solving the energy fluctuation problem caused by the intermittency of renewable energy and the uncertainty of demand in building clusters, and reducing energy operating costs.

CN116821723BActive Publication Date: 2026-04-07ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-14
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies cannot effectively solve the energy fluctuation problem caused by the intermittency of renewable energy and the uncertainty of energy demand in integrated energy systems for building complexes, resulting in high energy operating costs and resource waste.

Method used

We use deep strategy models and deep value networks to analyze the energy structure of building clusters, and optimize energy response actions through deep reinforcement training to reduce energy operating costs.

Benefits of technology

By combining deep strategy models and deep value networks, large-scale decision variables were optimized, enabling coordinated energy operation of building clusters and reducing energy operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116821723B_ABST
    Figure CN116821723B_ABST
Patent Text Reader

Abstract

This application discloses an energy coordination method and related apparatus for building clusters. The method includes: inputting the energy structure state into a deep strategy model and outputting energy response actions. The model building process involves: acquiring the historical energy structure state, energy response actions, and energy operating costs of each building; clustering the sample set; extracting the exploration reward function of each building from the clustering results; training the strategy model under the exploration reward function; training the strategy models of all buildings through a deep value network; updating the model parameters of the strategy models of all buildings; and obtaining a deep strategy model for all buildings. It is evident that when dealing with multiple decision variables in a large-scale building cluster, the use of a deep value network for deep reinforcement training of multiple strategy models optimizes the large-scale decision variables. The energy response actions output by the deep strategy model enable the coordinated operation of the integrated energy system of the building cluster, reducing the energy operating costs of the building cluster.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of energy, more particularly, to a building group energy collaborative operation method and related device. BACKGROUND

[0002] With the continuous development of renewable energy technology, more and more distributed renewable energy is connected to building integrated energy systems, which has the demand for collaborative operation. The building integrated energy system is a comprehensive system composed of power grid, heat grid, gas grid, energy conversion device, renewable energy and energy storage device provided for each building in the building group. However, the intermittency of renewable energy and the uncertainty of the energy demand side exacerbate the random fluctuations of both supply and demand of building integrated energy systems, bringing further challenges to collaborative operation optimization.

[0003] At present, when using a unified solution method to solve an uncertain planning model, the situation of not being able to converge due to improper parameter setting is usually encountered, and when using a cluster optimization method, the calculation efficiency is low when solving high-dimensional problems, and the solution is unstable. It can be seen that the traditional method cannot solve the uncertain planning model of the building group integrated energy system, so that the energy fluctuations between buildings are large, resulting in huge energy operation cost of the building group, causing resource waste. SUMMARY

[0004] In view of the above problems, the present application is proposed in order to provide a building group energy collaborative operation method and related device to reduce the energy operation cost of the building group and reduce resource waste.

[0005] In order to achieve the above purpose, the specific scheme is as follows:

[0006] A building group energy collaborative operation method applied to a building group overall planning system, comprising:

[0007] When the energy structure state of a building in the building group changes, input the current energy structure state of each building in the building group into the deep strategy model of each building, and output the energy response action of each building, so as to reduce the energy operation cost of the building group after each building executes its own energy response action;

[0008] The establishment process of the deep strategy model of each building includes:

[0009] Obtain the historical multiple energy structure states of each building in the building group, and the energy response action under each energy structure state, and the energy operation cost after executing the corresponding energy response action under each energy structure state;

[0010] By clustering the sample set of each building and extracting the exploration reward function of the building from the clustering results, each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample.

[0011] Under the exploration reward function of each building, the policy model of that building is trained using the sample set of that building as the training sample.

[0012] Under the supervision of a pre-established deep value network, deep training is performed on the strategy models of each building, and the model parameters of the strategy models of each building are updated to obtain the deep strategy models of each building.

[0013] Optionally, the step of clustering the sample set of each building and extracting the exploration reward function of that building from the clustering results includes:

[0014] For each building, the samples in the sample set of the building are clustered to obtain several sample classes;

[0015] For each sample class of each building, determine the reciprocal of the number of samples in the sample class and the average cost reward of the sample class;

[0016] The exploration reward function for each building is constructed by using the reciprocal of the number of samples in each sample class and the mean of cost reward as proportional parameters.

[0017] Optionally, for each building, the samples in the sample set of the building are clustered to obtain several sample classes, including:

[0018] For each building, calculate the Euclidean distance between each sample in the sample set of the building and its neighboring samples. The neighboring samples of each sample are the samples corresponding to the previous energy structure state of the energy structure state of the sample, or the samples corresponding to the next energy structure state of the energy structure state of the sample.

[0019] In each sample, when the Euclidean distance between each sample and its neighboring samples is less than the preset neighborhood distance, the sample is clustered with the neighboring sample to obtain several sample classes.

[0020] Optionally, under the supervision of a pre-established deep value network, deep training is performed on the policy models of each building to update the model parameters of each building's policy model, resulting in deep policy models for each building, including:

[0021] When the strategy model of each building in the building complex obtains the energy structure state, the energy response actions output by the strategy models of all buildings are obtained.

[0022] Based on the energy response actions output by the individual strategy models of each building, and the shared state variables of the building group, the strategy gradient of the pre-established deep value network is calculated.

[0023] Based on the policy gradient of the deep value network, the model parameters of the policy models of all buildings are updated until the historical energy structure state of all buildings has been obtained by their corresponding policy models. At this point, the policy models of all buildings are determined to be deep policy models.

[0024] Optionally, the step of calculating the policy gradient of the pre-established deep value network based on the energy response actions output by the individual policy models of all buildings and the shared state variables of the building group includes:

[0025] The policy gradient of a pre-built deep value network is calculated using the following formula:

[0026]

[0027] in, Let E be the strategy model for the i-th building, and E be the mathematical expectation. The policy gradient is the policy model of the pre-established deep value network for the i-th building, where N is the total number of buildings in the building group. For a pre-established deep value network, s i Let be the energy structure state of building i, and s be the building group state set, which consists of the energy structure states of each building and the shared state variables of the building group. i Energy response actions for building i.

[0028] An energy collaborative operation device for building complexes, applied to a building complex management system, comprising:

[0029] The energy action output unit is used to input the current energy structure state of each building in the building group into the deep strategy model of each building when the energy structure state of the buildings in the building group changes, and output the energy response action of each building, so as to reduce the energy operation cost of the building group after each building performs its own energy response action.

[0030] The information acquisition unit is used to acquire the historical multiple energy structure states of each building in the building group, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state.

[0031] The reward function exploration unit is used to cluster the sample set of each building and extract the exploration reward function of the building from the clustering results. Each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample.

[0032] The strategy model training unit is used to train the strategy model of each building by using the sample set of that building as training samples under the exploration reward function of each building.

[0033] The model deep reinforcement unit is used to perform deep training on the policy models of each building under the supervision of a pre-established deep value network, update the model parameters of the policy models of each building, and obtain the deep policy models of each building.

[0034] Optionally, the reward function exploration unit includes:

[0035] Clustering unit, used to cluster the samples in the sample set of each building to obtain several sample classes;

[0036] The sample calculation unit is used to determine the reciprocal of the number of samples in each sample class for each building, as well as the average cost reward of the sample class.

[0037] The reward function construction unit is used to construct the exploration reward function for each building by using the inverse of the number of samples in each sample class of each building and the mean of cost reward as proportional parameters for function construction.

[0038] Optionally, the clustering unit includes:

[0039] The first clustering subunit is used to calculate the Euclidean distance between each sample in the sample set of the building and its neighboring samples for each building. The neighboring samples of each sample are the samples corresponding to the previous energy structure state of the energy structure state of the sample, or the samples corresponding to the next energy structure state of the energy structure state of the sample.

[0040] The second clustering subunit is used to cluster each sample with its neighboring sample when the Euclidean distance between each sample and its neighboring sample is less than a preset neighborhood distance, so as to cluster each sample and obtain several sample classes.

[0041] Optionally, the model depth enhancement unit includes:

[0042] The first model deep enhancement subunit is used to obtain the energy response actions output by the strategy models of each building in the building group when the strategy model of each building obtains the energy structure state.

[0043] The second model deep reinforcement subunit is used to calculate the policy gradient of the pre-established deep value network based on the energy response actions output by the policy models of each building and the shared state variables of the building group.

[0044] The third model deep reinforcement subunit is used to update the model parameters of the strategy models of all buildings based on the policy gradient of the deep value network, until the historical energy structure state of all buildings has been obtained by their corresponding strategy models, and then the strategy models of all buildings are determined to be deep policy models.

[0045] Optionally, the second model depth enhancement subunit includes:

[0046] The policy gradient calculation unit is used to calculate the policy gradient of a pre-built deep value network using the following formula:

[0047]

[0048] in, Let E be the strategy model for the i-th building, and E be the mathematical expectation. The policy gradient is the policy model of the pre-established deep value network for the i-th building, where N is the total number of buildings in the building group. For a pre-established deep value network, s i Let be the energy structure state of building i, and s be the building group state set, which consists of the energy structure states of each building and the shared state variables of the building group. i Energy response actions for building i.

[0049] An energy-coordinated operation device for building complexes, including a memory and a processor;

[0050] The memory is used to store programs;

[0051] The processor is used to execute the program to implement the various steps of the energy collaborative operation method for building clusters as described above.

[0052] A storage medium storing a computer program, which, when executed by a processor, implements the various steps of the energy-coordinated operation method for a building complex as described above.

[0053] Using the above technical solution, this application reduces the energy operating cost of the building group by inputting the current energy structure state of each building in the building group into its respective deep strategy model when the energy structure state of the buildings changes, and outputting the energy response actions of each building. The process of establishing the deep strategy model for each building involves: obtaining the historical multiple energy structure states of each building in the building group, the energy response actions under each energy structure state, and the energy operating cost after executing the corresponding energy response actions under each energy structure state; clustering the sample set of each building; and extracting the exploration reward function of the building from the clustering results. Each sample in this set consists of an energy structure state, an energy response action, and a cost reward. The energy response action for each sample corresponds to its specific energy structure state, and the cost reward is based on the energy operating cost of that sample. Furthermore, under the exploration reward function for each building, the sample set of that building is used as training samples to train its strategy model. Under the supervision of a pre-established deep value network, deep training is performed on the strategy models of all buildings, updating the model parameters of each building's strategy model to obtain the deep strategy models for each building. Therefore, when dealing with multiple decision variables in a large-scale building cluster, the use of a deep value network for deep reinforcement training of multiple strategy models optimizes the large-scale decision variables. This makes the energy response actions output by each deep strategy model more capable of enabling the integrated energy system of each building to operate collaboratively, reducing the energy operating cost of the building cluster. Attached Figure Description

[0054] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0055] Figure 1 A schematic diagram illustrating a process for achieving coordinated energy operation of a building complex, provided as an embodiment of this application;

[0056] Figure 2A flowchart illustrating the extraction of a building exploration reward function is provided for an embodiment of this application;

[0057] Figure 3 A schematic diagram of a device structure for realizing coordinated energy operation of a building complex, provided as an embodiment of this application;

[0058] Figure 4 This is a schematic diagram of a device for realizing the coordinated energy operation of a building complex, provided as an embodiment of this application. Detailed Implementation

[0059] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0060] The proposed solution can be implemented based on a terminal with data processing capabilities. This terminal can be a building cluster management system, which can acquire historical data of each building in the building cluster and transmit control signals to each building.

[0061] Next, the energy-coordinated operation method for building clusters in this application may include:

[0062] When the energy structure of a building complex changes, the current energy structure of each building in the complex is input into its respective deep strategy model, and the energy response actions of each building are output. After each building executes its own energy response actions, the energy operating cost of the building complex is reduced.

[0063] Specifically, when external factors cause changes in the energy structure of a building cluster, the building cluster management system can obtain the current energy structure status information of each building and input the energy structure status into the deep strategy model corresponding to that building, so that the deep strategy model can output the energy response actions to be performed in relation to the current energy structure status.

[0064] In this process, after each building performs an energy response action, the energy structure state will change again because the energy response action forces changes in the environment. If the energy response action output by the deep strategy model remains unchanged, then there is no need to perform the energy response action.

[0065] For example, when the energy structure status of a building in a building complex changes, the building complex coordination system can be triggered to obtain the current energy structure status information of all buildings and input the energy structure status of each building into the deep strategy model corresponding to that building. The deep strategy model can analyze the current status of each building in the global building complex and output the energy response actions that the building needs to perform.

[0066] It is understandable that when the energy structure of a building complex changes, it may be due to random fluctuations in supply and demand. This situation will increase the energy operating costs of the building complex. However, after each building performs its own energy response, the operation of energy equipment can be optimized from the perspective of the entire building complex, thereby reducing the energy operating costs of the building complex.

[0067] Next, combined Figure 1 The process of establishing the deep strategy model for each building is described below, and this process may include the following steps:

[0068] Step S110: Obtain the historical multiple energy structure states of each building in the building group, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state.

[0069] Understandably, since each building in a building complex is equipped with a comprehensive building energy system consisting of a power grid, heating network, gas network, energy conversion devices, renewable energy sources, and energy storage equipment, each energy structure state can include one or more of the following: wind turbine output, photovoltaic output, electrical load, heat load, battery state of charge, and temperature. Energy response actions refer to the measures taken under each energy structure state, specifically including one or more of the following: battery charging / discharging power, air conditioning power, and heating power. Energy operating costs are the costs incurred / operated by the building's energy equipment after taking energy response actions for each energy structure state.

[0070] Step S120: Cluster the sample set of each building and extract the exploration reward function of the building from the clustering results.

[0071] Each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of that sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of that sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operating cost of that sample.

[0072] Specifically, the sample set for each building can be represented as M is the number of samples, s i For the i-th energy structure state, a i For the energy response action corresponding to the i-th energy structure state, r i ex To be with a i The corresponding cost reward can be based on clustering algorithms (such as k-means algorithm) to divide the samples into k classes, and extract the exploration reward function based on the novelty of the k classes of samples. This exploration reward function can be used to train the building strategy model to encourage the exploration of better quality states.

[0073] Step S130: Under the exploration reward function of each building, use the sample set of that building as training samples to train the policy model of that building.

[0074] Specifically, the strategy model for each building can accept the input of the energy structure status, analyze the energy structure status, and output energy response actions.

[0075] Understandably, when the policy model observes that the building temperature is low and the heating power needs to be increased, the environment returns a reward based on the state and action. A high reward value indicates a good decision by the policy model, while a low reward value indicates a poor decision. In this way, the policy model learns from training samples to optimize its output policy.

[0076] Step S140: Under the supervision of the pre-established deep value network, perform deep training on the strategy models of each building, update the model parameters of the strategy models of each building, and obtain the deep strategy models of each building.

[0077] Specifically, deep value networks can aim to reduce the overall energy operating costs of a building complex by obtaining the energy response actions output by the strategy models of each building, analyzing the overall energy operating costs of the building complex based on the energy structure status of the current overall environment, and outputting the model parameter adjustment values ​​corresponding to each strategy model to adjust the model parameters of each strategy model.

[0078] For example, when all the building's individual strategy models are trained in depth, each strategy model will output a result after each training sample. The deep value network aims to reduce the overall energy operating cost of the building group. It analyzes the output results, obtains the analysis results, and feeds back the model parameter adjustment information to the corresponding strategy model. This allows the strategy model that receives the model parameter adjustment information to update its model parameters, thereby achieving the effect of optimizing the training.

[0079] The energy coordination operation method for building clusters provided in this embodiment reduces the energy operating cost of the building cluster by inputting the current energy structure state of each building in the cluster into its respective deep strategy model when the energy structure state of the buildings changes. The method outputs energy response actions for each building, thereby reducing the energy operating cost of the building cluster after each building executes its respective energy response action. The process of establishing the deep strategy model for each building involves: acquiring the historical multiple energy structure states of each building in the cluster, the energy response actions under each energy structure state, and the energy operating cost after executing the corresponding energy response action under each energy structure state; clustering the sample set of each building; and extracting the exploration reward function of each building from the clustering results. Each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action for each sample corresponds to its energy structure state, and the cost reward for each sample is the cost reward corresponding to its energy response action, which is based on the energy operating cost of that sample. Furthermore, under the exploration reward function of each building, the sample set of that building is used as training samples to train the strategy model for that building. Under the supervision of a pre-established deep value network, deep training is performed on the strategy models of all buildings, updating the model parameters of each building's strategy model to obtain the deep strategy models for each building. Therefore, when dealing with multiple decision variables in a large-scale building cluster, the use of a deep value network for deep reinforcement training of multiple strategy models optimizes the large-scale decision variables, making the energy response actions output by each deep strategy model more capable of enabling the integrated energy system of each building to operate collaboratively and reducing the energy operating cost of the building cluster.

[0080] In some embodiments of this application, the process of step S120 above, which involves clustering the sample set of each building and extracting the exploration reward function of that building from the clustering results, is described, such as... Figure 2 As shown, the process may include:

[0081] Step S121: For each building, cluster the samples in the building's sample set to obtain several sample classes.

[0082] Specifically, each sample can be numerically processed. For example, if the energy structure consists of wind turbine output and photovoltaic output, the wind turbine output and photovoltaic output can be numerically converted into a single value. If the energy structure does not involve electrical load and heat load, then the values ​​of electrical load and heat load can be recorded as 0. Each sample can represent a multi-dimensional point, and then the samples can be clustered according to the distance between adjacent multi-dimensional points.

[0083] The process of clustering the samples in the sample set of each building to obtain several sample classes may include:

[0084] S1211. For each building, calculate the Euclidean distance between each sample in the building's sample set and its neighboring samples.

[0085] Among them, the neighboring samples of each sample are the samples corresponding to the previous energy structure state of the energy structure state of that sample, or the samples corresponding to the next energy structure state of the energy structure state of that sample.

[0086] For example, if the sample representation of energy structure state A is (10,20,30,0,0), the next energy structure state B of A is (11,19,29,2,0), and the next energy structure state C of B is (100,0,78,20,20), then the Euclidean distance between B and A is [(10-11)]. 2 +(20-19) 2 +(30-29) 2 +(0-2) 2 +(0-0) 2 ] 0.5 The Euclidean distance between B and C is [(11-100)] 2 +(19-0) 2 +(78-29) 2 +(20-2) 2 +(20-0) 2 ] 0.5 .

[0087] S1212. In each sample, when the Euclidean distance between each sample and its neighboring sample is less than the preset neighborhood distance, the sample is clustered with the neighboring sample to cluster each sample and obtain several sample classes.

[0088] Specifically, the preset neighborhood distance can represent the Euclidean distance standard for clustering two adjacent energy structure states, and the preset neighborhood distance can be customized.

[0089] For example, if the preset neighborhood distance is 10, when the sample representation of energy structure state A is (10,20,30,0,0), the next energy structure state B of A is (11,19,29,2,0), and the next energy structure state C of B is (100,0,78,20,20), the Euclidean distance between B and A is [(10-11)]. 2 +(20-19) 2 +(30-29) 2 +(0-2) 2 +(0-0) 2 ] 0.5 If the value is less than 10, then A and B can be clustered, and the Euclidean distance between B and C is [(11-100)]. 2 +(19-0) 2 +(78-29) 2 +(20-2) 2 +(20-0) 2 ] 0.5 If the value is not less than 10, then B and C will not be clustered.

[0090] Step S122: For each sample class of each building, determine the reciprocal of the number of samples in the sample class and the average cost reward of the sample class.

[0091] It is understandable that the reciprocal of the number of samples in a sample class can represent the novelty of that sample class, and the average cost-reward of a sample class can represent the quality of the energy response actions performed by the samples in each sample class in response to the energy structure state.

[0092] Step S123: Construct the exploration reward function for a building by using the reciprocal of the number of samples in each sample class of each building and the mean of the cost reward as the function to construct proportional parameters.

[0093] It is understandable that, since the exploration reward function is proportional to the inverse of the number of samples and the mean of the cost reward, the reward obtained by the energy structure state increases when the novelty of the energy structure state increases, and the reward obtained by the energy structure state also increases when the quality of the energy response action increases.

[0094] In some embodiments of this application, the process of performing deep training on the policy models of each building under the supervision of a pre-established deep value network, updating the model parameters of the policy models of each building, and obtaining the deep policy models of each building is described. This process may include:

[0095] S1. When the strategy model of each building in the building group obtains the energy structure state, obtain the energy response actions output by the strategy model of each building.

[0096] Understandably, each building's strategy model can accept the input of the energy structure status and output energy response actions. The building group coordination system can transmit information to each building to control the strategy model parameters of each building. Therefore, when the strategy model of each building in the building group obtains the energy structure status, the building group coordination system can obtain the energy response actions output by the strategy models of all buildings.

[0097] S2. Calculate the policy gradient of the pre-established deep value network based on the energy response actions output by the individual strategy models of each building and the shared state variables of the building group.

[0098] Specifically, the shared state quantity of a building cluster can represent the state quantity that maintains a consistent state throughout the entire cluster at any given time, such as the electricity price, heating price, and natural gas price in the area where the cluster is located. The policy gradient of a pre-established deep value network can be calculated using the following formula:

[0099]

[0100] in, Let E be the strategy model for the i-th building, and E be the mathematical expectation. The policy gradient is the policy model of the pre-established deep value network for the i-th building, where N is the total number of buildings in the building group. For a pre-established deep value network, s i Let be the energy structure state of building i, and s be the building group state set, which consists of the energy structure states of each building and the shared state variables of the building group. i Energy response actions for building i.

[0101] S3. Based on the policy gradient of the deep value network, update the model parameters of the policy models of all buildings until the historical energy structure state of all buildings has been obtained by their corresponding policy models, and then determine the policy models of all buildings as deep policy models.

[0102] Understandably, the policy models of each building can be trained simultaneously. The deep value network can evaluate whether the policy of each building's policy model has been optimized based on the training samples. If a building's policy model has not been optimized or has even degenerated (degeneration means a decrease in reward) in a single training sample, the deep value network can generate the model parameter adjustment values ​​for that building's policy model and send them to that building. The building's policy model can then change its model parameters to maximize the reward.

[0103] The apparatus for realizing the coordinated operation of building groups energy provided in the embodiments of this application will be described below. The apparatus for realizing the coordinated operation of building groups energy described below can be referred to in correspondence with the method for realizing the coordinated operation of building groups energy described above.

[0104] See Figure 3 , Figure 3 This is a schematic diagram of a device structure for realizing coordinated energy operation of a building complex, as disclosed in an embodiment of this application.

[0105] like Figure 3 As shown, the device may include:

[0106] The energy action output unit 11 is used to input the current energy structure state of each building in the building group into the deep strategy model of each building when the energy structure state of the buildings in the building group changes, and output the energy response action of each building, so as to reduce the energy operation cost of the building group after each building performs its own energy response action.

[0107] The information acquisition unit 12 is used to acquire the historical multiple energy structure states of each building in the building group, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state.

[0108] The reward function exploration unit 13 is used to cluster the sample set of each building and extract the exploration reward function of the building from the clustering results. Each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample.

[0109] The strategy model training unit 14 is used to train the strategy model of each building by using the sample set of that building as the training sample under the exploration reward function of each building.

[0110] The model deep reinforcement unit 15 is used to perform deep training on the strategy models of each building under the supervision of a pre-established deep value network, update the model parameters of the strategy models of each building, and obtain the deep strategy models of each building.

[0111] Optionally, the reward function exploration unit includes:

[0112] Clustering unit, used to cluster the samples in the sample set of each building to obtain several sample classes;

[0113] The sample calculation unit is used to determine the reciprocal of the number of samples in each sample class for each building, as well as the average cost reward of the sample class.

[0114] The reward function construction unit is used to construct the exploration reward function for each building by using the inverse of the number of samples in each sample class of each building and the mean of cost reward as proportional parameters for function construction.

[0115] Optionally, the clustering unit includes:

[0116] The first clustering subunit is used to calculate the Euclidean distance between each sample in the sample set of the building and its neighboring samples for each building. The neighboring samples of each sample are the samples corresponding to the previous energy structure state of the energy structure state of the sample, or the samples corresponding to the next energy structure state of the energy structure state of the sample.

[0117] The second clustering subunit is used to cluster each sample with its neighboring sample when the Euclidean distance between each sample and its neighboring sample is less than a preset neighborhood distance, so as to cluster each sample and obtain several sample classes.

[0118] Optionally, the model depth enhancement unit includes:

[0119] The first model deep enhancement subunit is used to obtain the energy response actions output by the strategy models of each building in the building group when the strategy model of each building obtains the energy structure state.

[0120] The second model deep reinforcement subunit is used to calculate the policy gradient of the pre-established deep value network based on the energy response actions output by the policy models of each building and the shared state variables of the building group.

[0121] The third model deep reinforcement subunit is used to update the model parameters of the strategy models of all buildings based on the policy gradient of the deep value network, until the historical energy structure state of all buildings has been obtained by their corresponding strategy models, and then the strategy models of all buildings are determined to be deep policy models.

[0122] Optionally, the second model depth enhancement subunit includes:

[0123] The policy gradient calculation unit is used to calculate the policy gradient of a pre-built deep value network using the following formula:

[0124]

[0125] in, Let E be the strategy model for the i-th building, and E be the mathematical expectation. The policy gradient is the policy model of the pre-established deep value network for the i-th building, where N is the total number of buildings in the building group. For a pre-established deep value network, s i Let be the energy structure state of building i, and s be the building group state set, which consists of the energy structure states of each building and the shared state variables of the building group. i Energy response actions for building i.

[0126] The energy coordination device for building clusters provided in this application embodiment can be applied to energy coordination equipment for building clusters, such as building cluster management systems. Optionally, Figure 4 This diagram shows the hardware structure of the energy coordination equipment for the building complex, with reference to... Figure 4 The hardware structure of the energy coordination equipment for building clusters may include: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4.

[0127] In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4;

[0128] Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0129] Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device;

[0130] The memory stores a program, which the processor can call. The program is used for:

[0131] When the energy structure status of buildings in a building complex changes, the current energy structure status of each building in the building complex is input into the deep strategy model of each building, and the energy response actions of each building are output. After each building executes its own energy response actions, the energy operating cost of the building complex is reduced.

[0132] The process of establishing the deep strategy model for each building includes:

[0133] The system obtains the historical multiple energy structure states of each building in the building complex, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state.

[0134] By clustering the sample set of each building and extracting the exploration reward function of the building from the clustering results, each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample.

[0135] Under the exploration reward function of each building, the policy model of that building is trained using the sample set of that building as the training sample.

[0136] Under the supervision of a pre-established deep value network, deep training is performed on the strategy models of each building, and the model parameters of the strategy models of each building are updated to obtain the deep strategy models of each building.

[0137] Optionally, the refined and extended functions of the program can be found in the description above.

[0138] This application embodiment also provides a storage medium that can store a program suitable for execution by a processor, the program being used for:

[0139] When the energy structure status of buildings in a building complex changes, the current energy structure status of each building in the building complex is input into the deep strategy model of each building, and the energy response actions of each building are output. After each building executes its own energy response actions, the energy operating cost of the building complex is reduced.

[0140] The process of establishing the deep strategy model for each building includes:

[0141] The system obtains the historical multiple energy structure states of each building in the building complex, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state.

[0142] By clustering the sample set of each building and extracting the exploration reward function of the building from the clustering results, each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample.

[0143] Under the exploration reward function of each building, the policy model of that building is trained using the sample set of that building as the training sample.

[0144] Under the supervision of a pre-established deep value network, deep training is performed on the strategy models of each building, and the model parameters of the strategy models of each building are updated to obtain the deep strategy models of each building.

[0145] Optionally, the refined and extended functions of the program can be found in the description above.

[0146] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0147] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0148] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for coordinated energy operation of a building complex, characterized in that, Applications in building complex management systems include: When the energy structure status of buildings in a building complex changes, the current energy structure status of each building in the building complex is input into the deep strategy model of each building, and the energy response actions of each building are output. After each building executes its own energy response actions, the energy operating cost of the building complex is reduced. The process of establishing the deep strategy model for each building includes: The system obtains the historical multiple energy structure states of each building in the building complex, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state. By clustering the sample set of each building and extracting the exploration reward function of the building from the clustering results, each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample. Under the exploration reward function of each building, the policy model of that building is trained using the sample set of that building as the training sample. Under the supervision of a pre-established deep value network, deep training is performed on the strategy models of each building, and the model parameters of the strategy models of each building are updated to obtain the deep strategy models of each building. The process of clustering the sample set of each building and extracting the exploration reward function for that building from the clustering results includes: For each building, the samples in the sample set of the building are clustered to obtain several sample classes; For each sample class of each building, determine the reciprocal of the number of samples in the sample class and the average cost reward of the sample class; The exploration reward function for each building is constructed by using the reciprocal of the number of samples in each sample class and the mean cost reward as proportional parameters.

2. The method according to claim 1, characterized in that, For each building, the samples in the sample set of that building are clustered to obtain several sample classes, including: For each building, calculate the Euclidean distance between each sample in the sample set of the building and its neighboring samples. The neighboring samples of each sample are the samples corresponding to the previous energy structure state of the energy structure state of the sample, or the samples corresponding to the next energy structure state of the energy structure state of the sample. In each sample, when the Euclidean distance between each sample and its neighboring samples is less than the preset neighborhood distance, the sample is clustered with the neighboring sample to obtain several sample classes.

3. The method according to claim 1, characterized in that, Under the supervision of a pre-established deep value network, deep training is performed on the policy models of each building, updating the model parameters of each building's policy model to obtain the deep policy models of each building, including: When the strategy model of each building in the building complex obtains the energy structure state, the energy response actions output by the strategy models of all buildings are obtained. Based on the energy response actions output by the individual strategy models of each building, and the shared state variables of the building group, the strategy gradient of the pre-established deep value network is calculated. Based on the policy gradient of the deep value network, the model parameters of the policy models of all buildings are updated until the historical energy structure state of all buildings has been obtained by their corresponding policy models. At this point, the policy models of all buildings are determined to be deep policy models.

4. The method according to claim 3, characterized in that, The step of calculating the policy gradient of the pre-established deep value network based on the energy response actions output by the individual strategy models of each building and the shared state variables of the building group includes: The policy gradient of a pre-built deep value network is calculated using the following formula: in, Let i be the strategy model for the i-th building. As a mathematical expectation marker, The policy gradient of the pre-established deep value network for the policy model of the i-th building. The total number of buildings in the building complex. For a pre-established deep value network, Let i be the energy structure state of building i. This is a building cluster state set, which consists of the energy structure state of each building and the shared state variables of the building cluster. Energy response actions for building i.

5. An energy collaborative operation device for a building complex, characterized in that, Applications in building complex management systems include: The energy action output unit is used to input the current energy structure state of each building in the building group into the deep strategy model of each building when the energy structure state of the buildings in the building group changes, and output the energy response action of each building, so as to reduce the energy operation cost of the building group after each building performs its own energy response action. The information acquisition unit is used to acquire the historical multiple energy structure states of each building in the building group, the energy response actions under each energy structure state, and the energy operating cost after performing the corresponding energy response actions under each energy structure state. The reward function exploration unit is used to cluster the sample set of each building and extract the exploration reward function of the building from the clustering results. Each sample in the sample set consists of an energy structure state, an energy response action, and a cost reward. The energy response action of each sample is the energy response action corresponding to the energy structure state of the sample, and the cost reward of each sample is the cost reward corresponding to the energy response action of the sample. The cost reward corresponding to the energy response action of each sample is obtained based on the energy operation cost of the sample. The strategy model training unit is used to train the strategy model of each building by using the sample set of that building as training samples under the exploration reward function of each building. The model deep enhancement unit is used to perform deep training on the strategy models of each building under the supervision of a pre-established deep value network, update the model parameters of the strategy models of each building, and obtain the deep strategy models of each building. The reward function exploration unit includes: Clustering unit, used to cluster the samples in the sample set of each building to obtain several sample classes; The sample calculation unit is used to determine the reciprocal of the number of samples in each sample class for each building, as well as the average cost reward of the sample class. The reward function construction unit is used to construct the exploration reward function for each building by using the inverse of the number of samples in each sample class of each building and the mean of cost reward as proportional parameters for function construction.

6. The apparatus according to claim 5, characterized in that, The clustering unit includes: The first clustering subunit is used to calculate the Euclidean distance between each sample in the sample set of the building and its neighboring samples for each building. The neighboring samples of each sample are the samples corresponding to the previous energy structure state of the energy structure state of the sample, or the samples corresponding to the next energy structure state of the energy structure state of the sample. The second clustering subunit is used to cluster each sample with its neighboring sample when the Euclidean distance between each sample and its neighboring sample is less than a preset neighborhood distance, so as to cluster each sample and obtain several sample classes.

7. An energy collaborative operation device for building complexes, characterized in that, Including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the energy collaborative operation method for building groups as described in any one of claims 1-4.

8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the energy collaborative operation method for building groups as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Privacy-safe building cluster energy consumption collaborative prediction method and system

    CN115409370A