Optimized scheduling method and system for multi-region integrated energy system

By introducing a dual random matrix based on topological relationships between integrated energy systems in the distributed federated learning framework, and collaboratively updating the local optimization model, the problems of data privacy leakage and insufficient real-time in the optimization scheduling of multi-regional integrated energy systems are solved, and more efficient and reliable energy scheduling is achieved.

CN120046941AActive Publication Date: 2025-05-27STATE GRID ZHEJIANG ELECTRIC POWER CO LTD HANGZHOU POWER SUPPLY CO

Patent Information

Application Number
CN202510475157.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-27
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing optimization scheduling method of multi-regional integrated energy systems has problems such as risk of data privacy leakage, strong communication network dependence, and insufficient scheduling delay and real-time.

Method used

In the distributed federated learning framework, a double random matrix constructed based on the topological relationship between comprehensive energy systems is introduced, and the model parameters of the local agent are weighted and aggregated with the model parameters of the neighbor agent, and the local optimization model is synergistically updated.

Benefits of technology

It effectively avoids the risk of data privacy leakage, saves network communication overhead, and improves the real-time and reliability of optimized scheduling of multi-region integrated energy systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046941A_ABST
    Figure CN120046941A_ABST
Patent Text Reader

Abstract

The invention provides an optimal scheduling method and system for a multi-region integrated energy system, and the method comprises the steps: enabling the integrated energy system of each region to be managed by a corresponding local agent, training a preset energy optimal scheduling model through the local agent based on a reinforcement learning idea, and obtaining a local optimization model; the local optimization model is updated according to the encryption parameter data packet of the neighbor intelligent agent and a double random matrix constructed based on the topological relation between the integrated energy systems, and the local optimization model of each local intelligent agent is updated after the local optimization model of each local intelligent agent is updated; and carrying out overall convergence evaluation on all the local optimization models, stopping cooperative training among the local intelligent agents when a preset optimization target is reached, and executing integrated energy system optimization scheduling of the corresponding region according to the local optimization model of each local intelligent agent. According to the method, the risk of data privacy leakage of the integrated energy system in each region can be effectively avoided, and the real-time performance and reliability of optimal scheduling of the multi-region integrated energy system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optimal scheduling of integrated energy systems, and particularly to an optimal scheduling method and system for a multi-region integrated energy system. Background Art

[0002] The growth of global energy demand has made it difficult for the traditional single-energy supply mode to meet the needs of modern society. The integrated energy system can achieve complementary advantages among different energies by integrating multiple energies such as electric energy and natural gas, making up for the technical limitations of single energy. Whether multiple regional integrated energy systems can be reasonably scheduled has also become a key factor affecting the operation stability of the power system.

[0003] In the existing research on the optimal scheduling of multi-region integrated energy systems, although the centralized scheduling method can achieve global optimization through a central server, once the sensitive data such as energy production, energy consumption, and energy trading of the integrated energy systems in each region are centrally stored, they are extremely vulnerable to external attacks or internal leaks, with a relatively high risk of privacy leakage. Moreover, the centralized scheduling is highly dependent on the communication network and the central server. The failure of the central server or the communication link will affect the scheduling efficiency of the entire system, and the increase in the computing burden of the central server will lead to scheduling delays, making it difficult to meet the real-time requirements. Although the introduction of distributed federated reinforcement learning can solve the risk of data privacy leakage in centralized scheduling and avoid the problem of insufficient scheduling efficiency caused by excessive dependence on the central server, in the scheduling optimization that allows each regional integrated energy system to be managed by an independent agent, each agent needs to converge quickly in local training to ensure efficient energy scheduling, and reliable model parameter interaction between agents is required to ensure global scheduling optimization. However, there is a relatively high risk of privacy leakage in the model parameter interaction between agents, and the convergence performance of each agent is also extremely vulnerable to network delay and parameter update rules. Therefore, there is an urgent need to provide an efficient optimal scheduling method that can not only avoid the risk of centralized storage of data of each regional integrated energy system but also improve the real-time performance and reliability of the scheduling of multi-region integrated energy systems. Summary of the Invention

[0004] The object of the present invention is to provide an optimal scheduling method for a multi-region integrated energy system. By introducing an agent interaction coordination training mechanism on the distributed federated learning framework, which uses a double-stochastic matrix constructed based on the topological relationship between integrated energy systems to weighted aggregate the model parameters of each local agent and the model parameters of the corresponding neighbor agents to collaboratively update the local optimization models of each local agent, it can not only effectively avoid the risk of privacy leakage of data of each regional integrated energy system, but also save network communication overhead, improve the real-time performance of the optimal scheduling of multi-region integrated energy systems, and support local agents to take into account individual learning characteristics and collective knowledge sharing, effectively improving the reliability of the optimal scheduling of multi-region integrated energy systems.

[0005] To achieve the above object, an optimization scheduling method and system for a multi-region integrated energy system are provided.

[0006] In a first aspect, an embodiment of the present invention provides an optimization scheduling method for a multi-region integrated energy system. The integrated energy system of each region is managed by a corresponding local agent. The method includes the following steps: Each local agent trains a preset energy optimization scheduling model based on the reinforcement learning idea to obtain a corresponding local optimization model, and updates the local optimization model according to the encrypted parameter data packets transmitted by each neighboring agent and a pre-constructed doubly stochastic matrix; the doubly stochastic matrix is constructed based on the topological relationship between the integrated energy systems of each region; the encrypted parameter data packet includes an agent ID, encrypted model parameters, and an encrypted timestamp; In response to the completion of the update of the local optimization models of all local agents, an overall convergence evaluation is performed on the local optimization models of all local agents. When the corresponding overall convergence evaluation result reaches a preset optimization target, the collaborative training between local agents is stopped, and the integrated energy system of the corresponding region is optimized and scheduled according to the local optimization models of each local agent.

[0007] Further, the state space and action space of the preset energy optimization scheduling model are respectively defined according to the operating state and scheduling strategy of the integrated energy system; the operating state includes electrical load demand, heat load demand, cooling load demand, photovoltaic power generation, energy storage unit state, energy scheduling interval, and future time electricity price; the scheduling strategy includes the electrical output power of a combined cooling, heat, and power unit, the output power of a photovoltaic generator, and the charge and discharge power of an energy storage; the reward function of the preset energy optimization scheduling model is constructed based on the operating cost of the integrated energy system scheduling and the penalty for an agent violating the constraints.

[0008] Further, the step of each local agent training the preset energy optimization scheduling model based on the reinforcement learning idea to obtain a corresponding local optimization model includes: Each local agent obtains the historical data of the integrated energy system of the corresponding region, and preprocesses the historical data of the integrated energy system to obtain a corresponding energy system optimization scheduling data set; Based on the energy system optimization scheduling data set, the preset energy optimization scheduling model is iteratively trained based on the double delayed deterministic policy gradient algorithm to obtain the local optimization model; The model parameters of the local optimization model are encrypted based on a preset encryption algorithm to generate a corresponding encrypted parameter data packet, and the encrypted parameter data packet is sent to each neighboring agent of the local agent.

[0009] Further, the step of encrypting the model parameters of the local optimization model based on a preset encryption algorithm to generate a corresponding encrypted parameter data packet includes: According to a preset parameter sorting rule, arranging the model parameters in the form of column vectors to obtain a corresponding parameter matrix; Performing an encryption operation on the parameter matrix using a homomorphic encryption algorithm to obtain encrypted model parameters; Combining the encrypted model parameters with the corresponding agent ID and encrypted timestamp to generate a data packet to be interacted, and encrypting the data packet to be interacted based on an asymmetric encryption algorithm to generate the encrypted parameter data packet.

[0010] Further, the step of constructing the doubly stochastic matrix includes: Generating an agent adjacency matrix based on the topological relationship between integrated energy systems in each region; Based on the agent adjacency matrix, obtaining the set of neighbor agents of each local agent respectively; Based on the generation principle of the doubly stochastic matrix, generating the doubly stochastic matrix according to the set of neighbor agents of all local agents; the doubly stochastic matrix is expressed as: In the formula, represents the element value of the th row and the th column in the doubly stochastic matrix; and respectively represent the number of neighbor agents of local agents and ; represents the set of neighbor agents of local agent ; represents the total number of local agents.

[0011] Further, the step of updating the local optimization model according to the encrypted parameter data packets transmitted by each received neighbor agent and the pre-constructed doubly stochastic matrix includes: Respectively obtaining the reception timestamps of the encrypted parameter data packets of each neighbor agent, and decrypting each encrypted parameter data packet to obtain the corresponding agent ID, encrypted model parameters, and encrypted timestamp; Based on the reception timestamp and encrypted timestamp of the encrypted parameter data packet corresponding to each agent ID, obtaining the neighbor communication delay between each neighbor agent and the local agent; Based on the neighbor communication delays between each neighbor agent and the local agent, the weight coefficients at the corresponding positions in the double stochastic matrix are corrected to obtain the corresponding local corrected double stochastic matrix; According to the local corrected double stochastic matrix, the encrypted model parameters of each neighbor agent and the local model parameters of the local optimized model are weighted and aggregated to obtain the updated local model parameters, and according to the updated local model parameters, an updated local optimized model is obtained.

[0012] Further, the step of correcting the weight coefficients at the corresponding positions in the double stochastic matrix based on the neighbor communication delays between each neighbor agent and the local agent to obtain the corresponding local corrected double stochastic matrix includes: According to all the neighbor communication delays corresponding to the local agent, the corresponding average neighbor interaction delay is obtained; The delay deviation rates between each neighbor communication delay and the average neighbor interaction delay are calculated respectively; When the delay deviation rate is positive, according to the delay deviation rate and the weight coefficient at the corresponding position of the corresponding neighbor agent in the double stochastic matrix, the corresponding weight coefficient reduction value is calculated, and the weight coefficient at the corresponding position in the double stochastic matrix is adjusted and updated according to the weight coefficient reduction value; According to all the weight coefficient reduction values, the total weight adjustment target for all neighbor agents with negative delay deviation rates is obtained, and the weight coefficients at the corresponding positions of all neighbor agents with negative delay deviation rates in the double stochastic matrix are adjusted and updated according to the total weight adjustment target.

[0013] Further, the updated local model parameters are expressed as: In the formula, and respectively represent the local model parameters of the local agent in the round of collaborative training and the encrypted model parameters of the corresponding neighbor agent ; represents the updated local model parameters of the local agent after the round of collaborative training; represents the weight coefficient of the local model parameters in the local corrected double stochastic matrix; represents the weight coefficient of the encrypted model parameters of the neighbor agent in the local corrected double stochastic matrix; represents the set of neighbor agents of the local agent ;

[0014] Further, the step of overall convergence evaluation of the local optimization models of all local agents includes: Obtain the standard deviation of the test accuracy and the average value of the parameter update amplitude of each local optimization model, and obtain the corresponding model convergence evaluation index according to the standard deviation of the test accuracy and the average value of the parameter update amplitude; the model convergence evaluation index is expressed as: In the formula, represents the model convergence evaluation index of the local optimization model obtained by the th round of collaborative training; represents the standard deviation of the test accuracy of the local optimization models obtained by the previous rounds of collaborative training; represents the average value of the parameter update amplitude of the local optimization models obtained by the previous rounds of collaborative training; and represent the target standard deviation of the test accuracy and the target parameter update amplitude; represents the weight coefficient; Judge whether the model convergence evaluation indexes of all local optimization models reach the preset model convergence threshold respectively. If so, the corresponding model convergence evaluation result is converged; otherwise, it is determined that the corresponding model convergence evaluation result is not converged; Obtain the overall convergence evaluation result according to the model convergence evaluation results of all local optimization models.

[0015] In a second aspect, an embodiment of the present invention provides an optimal scheduling system for a multi-region integrated energy system. The integrated energy systems of each region are managed by corresponding local agents. The system includes: An agent training module, configured to train a preset energy optimal scheduling model based on the reinforcement learning idea by each local agent to obtain a corresponding local optimization model, and update the local optimization model according to the encrypted parameter data packets transmitted by each neighbor agent and a pre-constructed doubly stochastic matrix; the doubly stochastic matrix is constructed based on the topological relationship between the integrated energy systems of each region; the encrypted parameter data packet includes an agent ID, encrypted model parameters, and an encrypted timestamp; An energy optimal scheduling module, configured to, in response to the completion of the update of the local optimization models of all local agents, perform an overall convergence evaluation on the local optimization models of all local agents, and stop the collaborative training between local agents when the corresponding overall convergence evaluation result reaches a preset optimization target, and perform the optimal scheduling of the integrated energy system of the corresponding region according to the local optimization models of each local agent.

[0016] The present invention provides an optimization scheduling method and system for a multi-region integrated energy system. Through the method, the integrated energy systems of each region are managed by corresponding local agents. Each local agent trains a preset energy optimization scheduling model based on the idea of reinforcement learning to obtain a corresponding local optimization model. The local optimization model is updated according to the encrypted parameter data packets received from each neighboring agent, including the agent ID, encrypted model parameters, and encrypted timestamps, and a double stochastic matrix pre-constructed based on the topological relationship between the integrated energy systems of each region. After the local optimization models of all local agents are updated, an overall convergence evaluation is performed on the local optimization models of all local agents. When the corresponding overall convergence evaluation result reaches a preset optimization target, the collaborative training between local agents is stopped, and the integrated energy system optimization scheduling of the corresponding region is executed according to the local optimization models of each local agent. Compared with the prior art, the optimization scheduling method for the multi-region integrated energy system introduces an agent interaction coordination training mechanism that uses a double stochastic matrix constructed based on the topological relationship between integrated energy systems to weighted aggregate the model parameters of each local agent and the model parameters of the corresponding neighboring agent to collaboratively update the local optimization models of each local agent in the distributed federated learning framework. This mechanism can not only effectively avoid the risk of privacy leakage of the data of the integrated energy systems in each region, but also save network communication overhead, improve the real-time performance of the optimization scheduling of the multi-region integrated energy system, and support local agents to take into account both individual learning characteristics and collective knowledge sharing, effectively improving the reliability of the optimization scheduling of the multi-region integrated energy system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is a schematic flowchart of the optimization scheduling method for the multi-region integrated energy system in an embodiment of the present invention; Figure 2 is a schematic structural diagram of the optimization scheduling system for the multi-region integrated energy system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to make the objectives, technical solutions, and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Obviously, the following described embodiments are part of the embodiments of the present invention and are only used to illustrate the present invention, but not to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] The optimal scheduling method for the multi-region integrated energy system provided by the present invention can be understood as a method proposed based on the application status that there is a large risk of privacy leakage in the existing optimal scheduling of the multi-region integrated energy system, and at the same time, it is unable to effectively ensure the real-time performance and reliability of the optimal scheduling of the multi-region integrated energy system. An intelligent agent interaction coordination training mechanism is introduced in the distributed federated learning framework. By using a double stochastic matrix constructed based on the topological relationship between integrated energy systems, the model parameters of each local agent are weighted and aggregated with the model parameters of the corresponding neighbor agents to collaboratively update the local optimization models of each local agent, so as to construct an optimal scheduling method for the multi-region integrated energy system for scheduling and managing each integrated energy system by local agents; the following embodiments will elaborate on the optimal scheduling method for the multi-region integrated energy system of the present invention.

[0020] In one embodiment, as Figure 1 shown, an optimal scheduling method for a multi-region integrated energy system is provided. The integrated energy systems of each region are managed by corresponding local agents. The local agents train local optimization models in a distributed environment. Each local agent can independently learn and make decisions based on local data and experience. The local agents can share key information of other local agents without sending the locally trained local optimization models to the central server, that is, they can efficiently and reliably manage the energy scheduling within the region without a central server; specifically, the optimal scheduling method for the multi-region integrated energy system includes the following steps: S11. Each local agent trains a preset energy optimal scheduling model based on the reinforcement learning idea to obtain a corresponding local optimization model, and updates the local optimization model according to the encrypted parameter data packets transmitted by each received neighbor agent and the pre-constructed double stochastic matrix; the double stochastic matrix is constructed based on the topological relationship between the integrated energy systems of each region; the encrypted parameter data packet includes an agent ID, encrypted model parameters, and an encrypted timestamp, where the agent ID is the unique identifier for distinguishing each local agent, the encrypted model parameters are the encrypted values of the model parameters of the local optimization model corresponding to the agent ID, and the encrypted timestamp is the moment when the agent ID encrypts the model parameters of its local optimization model.

[0021] The preset energy optimization scheduling model in this embodiment can be understood as being constructed based on the scheduling management requirements of the regional integrated energy system, combining nonlinear optimization theory with the idea of reinforcement learning, taking the real-time operation data of the regional integrated energy system as the input, and taking the optimal scheduling strategy of the regional integrated energy system as the output of the quadratic constraint programming problem model for the scheduling of the regional integrated energy system. Specifically, the state space and action space of the preset energy optimization scheduling model are respectively defined according to the operation state and scheduling strategy of the integrated energy system. The operation state includes electricity load demand, heat load demand, cooling load demand, photovoltaic power generation, energy storage unit state, energy scheduling interval, and electricity price at future times, etc. Among them, the electricity price at future times is also obtained based on existing relevant electricity price prediction technologies, and no specific limitation is made here. In order to provide the local agent with an understanding of the constraints of the operation unit, each state included in the operation state needs to be adjusted to the same scale in actual application, that is, scaled to the [0,1] interval through normalization; the corresponding state space is expressed as: In the formula, , , , and respectively represent at time the electricity load demand, heat load demand, cooling load demand, photovoltaic power generation, and energy storage unit state of the th regional integrated energy system;

[0022] The scheduling strategy of the preset energy optimization scheduling model includes the electric output power of the combined cooling, heat and power unit, the output power of the photovoltaic generator, and the charge and discharge power of the energy storage, etc. That is, the corresponding action space can be expressed as: In the formula, represents at time the scheduling strategy of the th regional integrated energy system; respectively represent at time the electric output power of the combined cooling, heat and power unit, the output power of the photovoltaic generator, and the charge and discharge power of the energy storage of the

[0023] In order to facilitate the evaluation of the rationality of the generated energy scheduling strategy, this embodiment preferably constructs the reward function of the preset energy optimization scheduling model based on the operating cost of the integrated energy system scheduling and the penalty for the agent violating the constraints, which is expressed as: In the formula, wherein, and respectively represent the negative value of the operating cost of the th regional integrated energy system at time and the penalty for the agent violating the constraints; is the penalty weight, which can be set according to the actual application requirements; represents the value of the th constraint variable at time and respectively represent the allowable upper limit value and the allowable lower limit value of the th constraint variable; , , and respectively represent the power generation cost of the generator in the th regional integrated energy system at time

[0024] Specifically, the power generation cost of the photovoltaic generator in the th region can be expressed as: wherein, , and are the cost coefficients of the power generation unit, is the output power of the photovoltaic generator in the th region at time

[0025] The electricity purchase cost of the th region trading with the main power grid and other integrated energy systems can be expressed as: In the formula, represents the electricity trading price of the main power grid at time is the power of the th regional integrated energy system trading with the main power grid at time is At time and the Internal trading electricity price of the integrated energy system IES for the integrated energy system trading in the area; is At time and the Trading power of the integrated energy system IES for the integrated energy system trading in the

[0026] Natural gas purchase cost for the i-th area to meet the demand of the combined cooling, heating and power unit Can be expressed as: In the formula, Represents the price of purchasing natural gas from the natural gas network, is Input power of the combined cooling, heating and power unit of the integrated energy system in the -th area at time.

[0027] The Equipment loss cost of the Can be expressed as: Among them, Is the cost per unit capacity loss; Is the equipment Power loss coefficient; is Output power of the equipment of the integrated energy system in the -th area at time. It should be noted that the corresponding scheduling constraint conditions required for the preset energy optimization scheduling model constructed through the above method steps can be determined based on the operation requirements of the actual integrated energy system in combination with the variables involved in the corresponding state space and action space, and no specific limitation is made here.

[0028] After the above-mentioned preset energy optimization scheduling model is started with the multi-area integrated energy system optimization scheduling instruction, it can be distributed to the integrated energy systems in each area based on the distributed federated learning architecture for training the corresponding local agents; specifically, the steps for each local agent to train the preset energy optimization scheduling model based on the reinforcement learning idea to obtain the corresponding local optimization model include:

[0029] Each local agent obtains the historical data of the integrated energy system in the corresponding area, and preprocesses the historical data of the integrated energy system to obtain the corresponding energy system optimal scheduling data set; among them, the historical data of the integrated energy system can be understood as the empirical data of a certain duration collected through the interaction between the agent and the integrated energy system scheduling environment, including the time series data of the operating state and the corresponding time series data of the scheduling strategy; the corresponding preprocessing process can successively include outlier removal, missing value filling, normalization processing, etc. The energy system optimal scheduling data set obtained by preprocessing the historical data of the integrated energy system can be stored in the experience replay buffer to provide reliable training data for network update in subsequent agent training.

[0030] Based on the integrated energy system optimal scheduling data set, the preset energy optimal scheduling model is iteratively trained based on the twin delayed deterministic policy gradient algorithm to obtain the local optimal model; among them, the twin delayed deep deterministic policy gradient algorithm (TD3) is a reinforcement learning algorithm for continuous action space problems. It improves the DDPG algorithm through innovations such as double Critic networks, delayed update of the Actor, and target policy smoothing; in this embodiment, this algorithm is used for the training of local agents in the integrated energy system of each area, which can not only make full use of the historical energy data of each area, but also effectively improve the stability and reliability of the generation of the local optimal model strategy; it should be noted that the specific process of each local agent iteratively training the preset energy optimal scheduling model based on the twin delayed deterministic policy gradient algorithm can be implemented with reference to relevant existing technologies and will not be elaborated here.

[0031] Based on a preset encryption algorithm, the model parameters of the local optimal model are encrypted to generate the corresponding encrypted parameter data packet, and the encrypted parameter data packet is sent to each neighbor agent of the local agent; among them, the model parameters can be understood as the parameters of the key agent network model determined based on actual application requirements for sharing between agents, and can include the weight matrix and bias vector in the agent network.

[0032] In practical applications, after each local agent completes a round of training to obtain the corresponding local optimized model, it is necessary to interact with the neighbor nodes directly connected to it to share the model parameters, so as to achieve the sharing of agent information within a small range and improve the convergence rate and training effect of the model. To avoid the risk of privacy leakage in the interaction of model parameters between local agents, in this embodiment, the model parameters to be shared and some key information convenient for subsequent aggregation of model parameters are preferably packed and encrypted and transmitted to the corresponding neighbor agents. Specifically, the step of encrypting the model parameters of the local optimized model based on a preset encryption algorithm to generate the corresponding encrypted parameter data packet includes: According to the preset parameter sorting rule, arrange the model parameters in the form of column vectors to obtain the corresponding parameter matrix. Among them, the parameter matrix can be understood as a dimensional column vector, where h is the total number of all model parameters. The corresponding preset parameter sorting rule can be understood as the order of arrangement of the model parameters uniformly used by each local agent designed in advance for the convenience of sharing the model parameters between adjacent local agents, and no specific limitation is made here.

[0033] Use the homomorphic encryption algorithm to perform an encryption operation on the parameter matrix to obtain the encrypted model parameters. Among them, the encrypted model parameters can be understood as the encrypted parameter matrix obtained by encrypting each element in the parameter matrix using the homomorphic encryption algorithm respectively.

[0034] Combine the encrypted model parameters with the corresponding agent ID and encrypted timestamp to generate the data packet to be interacted, and encrypt the data packet to be interacted based on the use of the asymmetric encryption algorithm to generate the encrypted parameter data packet. Among them, the arrangement order of the encrypted model parameters, agent ID, and encrypted timestamp in the data packet to be interacted can be set based on the preset communication protocol between agents.

[0035] The double stochastic matrix in this embodiment can be understood as a reference matrix for determining the weight coefficients of all model parameters when each local agent uses the encrypted model parameters of its adjacent neighbor agents to update the local optimized model. To avoid unnecessary waste of communication resources, effectively reduce the network load, and at the same time ensure that each local agent indirectly obtains global information, this embodiment preferably constructs a double stochastic matrix based on the proximity relationship between agents. Specifically, the steps for constructing the double stochastic matrix include: Generate an agent adjacency matrix based on the topological relationship between integrated energy systems in each region. Among them, the agent adjacency matrix can be understood as a matrix representing the direct connection relationship between local agents. For example, if there is a direct connection and interaction relationship between local agent a and local agent b, the matrix elements corresponding to the a-th row and b-th column and the b-th row and a-th column in the agent proximity matrix are both 1. Conversely, if there is no direct connection and interaction relationship between local agent a and local agent b, the matrix elements corresponding to the a-th row and b-th column and the b-th row and a-th column in the agent proximity matrix are both 0.

[0036] Based on the agent adjacency matrix, obtain the neighbor agent sets of each local agent respectively. Among them, the neighbor agent set of a certain local agent can be understood as the set of other agents corresponding to the elements with a value of 1 in the row vector corresponding to the serial number of this local agent in the agent adjacency matrix.

[0037] Based on the principle of generating a doubly stochastic matrix, generate the doubly stochastic matrix according to the neighbor agent sets of all local agents. Among them, the principle of generating a doubly stochastic matrix can be understood as that the sum of the elements in the row vector and column vector of the matrix is 1. To ensure the effectiveness of the weight coefficients used for updating the model parameters of each local agent, in this embodiment, preferably, based on the number of agents in the neighbor agent set of a local agent and the number of agents in the neighbor agent set corresponding to the neighbor agent of this local agent, determine the weight coefficients for updating the local agent with the encrypted model parameters of the neighbor agent. Specifically, the doubly stochastic matrix is expressed as: In the formula, represents the element value of the -th row and -th column in the doubly stochastic matrix; and respectively represent the number of neighbor agents of local agent and local agent ; represents the neighbor agent set of local agent ; represents the total number of local agents.

[0038] In this embodiment, using the doubly stochastic matrix to determine the weight coefficients for updating the local agent with the encrypted model parameters of the neighbor agent can achieve purposeful parameter synchronization based on the relative importance and dependence relationship between agents, so as to adapt to the dynamically changing agent state. At the same time, it can also support each local agent to take into account both individual learning characteristics and collective knowledge sharing to optimize its decision-making process. Furthermore, while ensuring the convergence effect of the local optimization model, it can also ensure the generalization ability and robustness of the local optimization model.

[0039] Specifically, the step of updating the local optimization model according to the encrypted parameter data packets transmitted by each received neighbor agent and the pre-constructed double stochastic matrix includes: Obtain the reception timestamps of the encrypted parameter data packets of each neighbor agent respectively, and decrypt each encrypted parameter data packet to obtain the corresponding agent ID, encrypted model parameters, and encrypted timestamp; among them, decrypting each encrypted parameter data packet can be implemented based on a preset asymmetric encryption algorithm, which will not be elaborated here.

[0040] Based on the reception timestamp and encrypted timestamp of the encrypted parameter data packet corresponding to each agent ID, obtain the neighbor communication delay between each neighbor agent and the local agent; that is, the neighbor communication delay between each neighbor agent and the local agent is the difference between the reception timestamp and the encrypted timestamp of the corresponding encrypted parameter data packet.

[0041] Based on the neighbor communication delay between each neighbor agent and the local agent, correct the weight coefficient at the corresponding position in the double stochastic matrix to obtain the corresponding local corrected double stochastic matrix; among them, the local corrected double stochastic matrix can be understood as further correcting the weight coefficient at the corresponding position in the double stochastic matrix obtained by the foregoing method based on the method of the influence of the communication delay between agents on the convergence efficiency of the local model while considering the relative importance and dependence relationship between agents.

[0042] Specifically, the step of correcting the weight coefficient at the corresponding position in the double stochastic matrix based on the neighbor communication delay between each neighbor agent and the local agent to obtain the corresponding local corrected double stochastic matrix includes: According to all neighbor communication delays corresponding to the local agent, obtain the corresponding average neighbor interaction delay; that is, the average neighbor interaction delay of the local agent is the average of all its neighbor communication delays.

[0043] Calculate the delay deviation rate between each neighbor communication delay and the average neighbor interaction delay respectively; among them, the delay deviation rate can be understood as the ratio obtained by dividing the difference between the neighbor communication delay and the average neighbor interaction delay by the average neighbor interaction delay. If the neighbor communication delay is greater than the average neighbor interaction delay, the corresponding delay deviation rate is positive, otherwise, the corresponding delay deviation rate is negative.

[0044] When the time delay deviation rate is positive, calculate the corresponding weight coefficient reduction value according to the time delay deviation rate and the weight coefficient at the corresponding position of the corresponding neighbor agent in the double stochastic matrix, and adjust and update the weight coefficient at the corresponding position in the double stochastic matrix according to the weight coefficient reduction value; wherein, the weight coefficient reduction value of the neighbor agent can be understood as the product of the time delay deviation rate corresponding to the neighbor agent and its weight coefficient at the corresponding position in the double stochastic matrix; after obtaining the weight coefficient reduction value of the neighbor agent, use the difference between the weight coefficient at the corresponding position in the existing double stochastic matrix and the weight coefficient reduction value as the weight coefficient value at the corresponding position in the updated double stochastic matrix. Through this method, the weight coefficient that affects the convergence efficiency of the local agent model due to large communication delay can be reduced to reduce its contribution to the update of the local optimization model.

[0045] According to all the weight coefficient reduction values, obtain the total weight adjustment target for all neighbor agents with negative time delay deviation rates, and adjust and update the weight coefficients at the corresponding positions of all neighbor agents with negative time delay deviation rates in the double stochastic matrix according to the total weight adjustment target; wherein, the total weight adjustment target can be understood as the cumulative value of the reduction amounts of the weight coefficients of all neighbor agents with positive time delay deviation rates; in order to reasonably allocate the weight coefficients used for updating the local optimization model while maintaining the characteristics of the double stochastic matrix to improve the convergence rate of each local agent model, preferably in this embodiment, the obtained total weight adjustment target is allocated based on the deviation rate ratio of all neighbor agents with negative time delay deviation rates. That is, first calculate the ratio value of the deviation rate of all neighbor agents with negative time delay deviation rates to the sum of all negative time delay deviation rates, and then obtain the weight coefficient increase value of the corresponding neighbor agent based on the product of the ratio value of each neighbor agent with a negative time delay deviation rate and the total weight adjustment target. After obtaining the weight coefficient increase value of the neighbor agent, use the sum of the weight coefficient at the corresponding position in the existing double stochastic matrix and the weight coefficient increase value as the weight coefficient value at the corresponding position in the updated double stochastic matrix. Through this method, the contribution degree of neighbor agents with smaller communication delays to the update of the local optimization model can be improved.

[0046] According to the locally corrected double stochastic matrix, perform weighted aggregation on the encrypted model parameters of each neighbor agent and the local model parameters of the local optimization model to obtain updated local model parameters, and obtain an updated local optimization model according to the updated local model parameters.

[0047] Specifically, the updated local model parameters are expressed as: In the formula, and respectively represent the Local agents in round - robin collaborative training The local model parameters of the local agent and the corresponding neighbor agents The encrypted model parameters; Denote the Updated local model parameters of the local agent after the Round - robin collaborative training; Denote the weight coefficient of the local model parameters in the local modified doubly - stochastic matrix; Denote the weight coefficient of the encrypted model parameters of the neighbor agent In the local modified doubly - stochastic matrix; Denote the set of neighbor agents of the local agent ;

[0048] In this embodiment, on the basis of considering the relative importance and dependence relationship between agents, further considering the impact of communication delay between agents on the convergence efficiency of the local model, a method of modifying the doubly - stochastic matrix based on the communication delay of neighbor agents and then using it to update the local model parameters is adopted. This can ensure that while improving the convergence effect, generalization ability and robustness of the local optimized model, it can also ensure that the update of the local optimized model can adapt to the dynamically changing agent states and network conditions at the same time. Furthermore, it can effectively improve the training efficiency of distributed federated learning and provide a reliable guarantee for the real - time performance and reliability of the optimal scheduling of multi - area integrated energy systems.

[0049] S12. In response to the completion of the update of the local optimized models of all local agents, perform an overall convergence evaluation on the local optimized models of all local agents. When the corresponding overall convergence evaluation result reaches the preset optimization target, stop the collaborative training between local agents, and perform the optimal scheduling of the integrated energy system in the corresponding area according to the local optimized models of each local agent.

[0050] In this embodiment, the overall convergence evaluation can be understood as a comprehensive evaluation of the iterative training convergence performance of the local optimized models of each local agent to determine whether to continue the subsequent collaborative training process between agents. To determine the effectiveness of the model convergence performance evaluation, this embodiment preferably conducts a comprehensive analysis based on two dimensions: test accuracy and parameter update amplitude. Specifically, the step of performing an overall convergence evaluation on the local optimized models of all local agents includes: Obtain the standard deviation of the test accuracy and the average value of the parameter update amplitude of each local optimized model, and based on the standard deviation of the test accuracy and the average value of the parameter update amplitude, obtain the corresponding model convergence evaluation index; The model convergence evaluation index is expressed as: In the formula, Denote the Model convergence evaluation metrics of the locally optimized models obtained through rounds of collaborative training; Indicates the previous Standard deviation of the test accuracy of the locally optimized models obtained through rounds of collaborative training; Indicates the previous Average value of the parameter update amplitude of the locally optimized models obtained through rounds of collaborative training; And Indicates the target standard deviation of the test accuracy and the target parameter update amplitude, which can be set according to actual application requirements and are common metric values for all local agents; Indicates the weight coefficient, which can be set as a fixed constant according to actual needs or can be dynamically set based on a preset dynamic adjustment rule according to actual application requirements; it should be noted that when When this is the case, the standard deviation of the test accuracy and the average value of the parameter update amplitude of each locally optimized model can be calculated using the preset initial standard deviation of the test accuracy and the initial average value of the parameter update amplitude respectively. This initial standard deviation of the test accuracy is less than the target standard deviation of the test accuracy, and the initial average value of the parameter update amplitude is less than the target parameter update amplitude.

[0051] Respectively determine whether the model convergence evaluation metrics of all locally optimized models reach the preset model convergence threshold. If so, the corresponding model convergence evaluation result is determined to be converged; otherwise, the corresponding model convergence evaluation result is determined to be not converged; among them, the preset model convergence threshold can also be set according to actual application requirements and is not specifically limited here.

[0052] Obtain the overall convergence evaluation result based on the model convergence evaluation results of all locally optimized models.

[0053] The preset optimization objective in this embodiment is understood as the iteration training stop condition, which can be set according to actual application requirements. For example, all locally optimized models converge or the proportion of the corresponding model convergence evaluation results being determined to be converged reaches a preset ratio, etc.; when the overall convergence evaluation result reaches the preset optimization objective, the collaborative training between local agents is no longer continued, and it is considered that the local agents corresponding to the current regional integrated energy systems have all been trained and can be used to execute the optimization scheduling of the corresponding regional integrated energy systems; that is, the corresponding optimal scheduling strategy can be obtained by inputting the real-time operation data of the integrated energy systems of each region currently obtained into the corresponding locally optimized models.

[0054] The integrated energy systems of each region provided by the embodiments of the present invention are managed by corresponding local agents. Each local agent trains a preset energy optimization scheduling model based on the idea of reinforcement learning to obtain a corresponding local optimization model, updates the local optimization model according to the encrypted parameter data packets including agent IDs, encrypted model parameters, and encrypted timestamps transmitted by each neighboring agent and a doubly stochastic matrix pre-constructed based on the topological relationship between the integrated energy systems of each region, and after the local optimization models of all local agents are updated, performs an overall convergence evaluation on the local optimization models of all local agents, and when the corresponding overall convergence evaluation result reaches a preset optimization target, stops the collaborative training between local agents, and executes the optimization scheduling of the integrated energy systems of the corresponding regions according to the local optimization models of each local agent. Through the agent interaction and coordination training mechanism that introduces a doubly stochastic matrix constructed based on the topological relationship between integrated energy systems and the network communication delay between agents in the distributed federated learning framework to perform weighted aggregation on the model parameters of each local agent and the model parameters of the corresponding neighboring agents to collaboratively update the local optimization models of each local agent, it can not only effectively avoid the risk of privacy leakage of the data of the integrated energy systems of each region, but also save network communication overhead, improve the convergence efficiency of each agent model, ensure the real-time performance of the optimization scheduling of the multi-region integrated energy system, and also support local agents to take into account both individual learning characteristics and collective knowledge sharing at the same time, effectively improving the reliability and stability of the optimization scheduling of the multi-region integrated energy system.

[0055] It should be noted that although the steps in the above flowchart are shown sequentially according to the indication of the arrows, these steps are not necessarily executed sequentially according to the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders.

[0056] In one embodiment, as Figure 2 shown, an optimization scheduling system for a multi-region integrated energy system is provided. The integrated energy systems of each region are managed by corresponding local agents. The system includes: An agent training module 1, configured to train a preset energy optimization scheduling model based on the idea of reinforcement learning for each local agent to obtain a corresponding local optimization model, and update the local optimization model according to the encrypted parameter data packets transmitted by each neighboring agent and the pre-constructed doubly stochastic matrix; the doubly stochastic matrix is constructed based on the topological relationship between the integrated energy systems of each region; the encrypted parameter data packet includes an agent ID, an encrypted model parameter, and an encrypted timestamp; The energy optimization scheduling module 2 is configured to, in response to the completion of the update of the local optimization models of all local agents, perform an overall convergence evaluation on the local optimization models of all local agents, and when the corresponding overall convergence evaluation result reaches a preset optimization target, stop the collaborative training among the local agents, and perform the optimization scheduling of the integrated energy system in the corresponding area according to the local optimization models of the respective local agents.

[0057] For the specific limitations of the optimization scheduling system of the multi-area integrated energy system, reference can be made to the limitations of the optimization scheduling method of the multi-area integrated energy system in the above text, and the corresponding technical effects can also be equivalently obtained, which will not be elaborated here. Each module in the above optimization scheduling system of the multi-area integrated energy system can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or stored in the memory of the computer device in the form of software, so as to facilitate the processor to call and execute the operations corresponding to the above-mentioned modules.

[0058] In summary, the optimization scheduling method and system of a multi-area integrated energy system provided by the embodiments of the present invention. The optimization scheduling method of the multi-area integrated energy system realizes that the integrated energy systems of the respective areas are managed by the corresponding local agents. Each local agent respectively trains the preset energy optimization scheduling model based on the reinforcement learning idea to obtain the corresponding local optimization model, updates the local optimization model according to the encrypted parameter data packet including the agent ID, encrypted model parameters, and encrypted timestamp transmitted by each neighboring agent and the double stochastic matrix pre-constructed based on the topological relationship among the integrated energy systems of the respective areas, and after the update of the local optimization models of all local agents is completed, performs an overall convergence evaluation on the local optimization models of all local agents, and when the corresponding overall convergence evaluation result reaches a preset optimization target, stops the collaborative training among the local agents, and performs the optimization scheduling of the integrated energy system in the corresponding area according to the local optimization models of the respective local agents. The method introduces an intelligent agent interaction coordination training mechanism that uses a double stochastic matrix constructed based on the topological relationship among the integrated energy systems and the network communication delay among the agents to perform weighted aggregation on the model parameters of each local agent and the model parameters of the corresponding neighboring agents to collaboratively update the local optimization models of each local agent in the distributed federated learning framework. This mechanism can not only effectively avoid the risk of privacy leakage of the data of the integrated energy systems in each area, but also save network communication overhead, improve the convergence efficiency of each agent model, ensure the real-time performance of the optimization scheduling of the multi-area integrated energy system, and support the local agents to simultaneously take into account the individual learning characteristics and collective knowledge sharing, effectively improving the reliability and stability of the optimization scheduling of the multi-area integrated energy system.

[0059] Each embodiment in this specification is described in a progressive manner. For the parts that are the same or similar in each embodiment, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the relevant part of the method embodiment for the relevant content. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0060] The above embodiments only represent several preferred embodiments of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be pointed out that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and substitutions can still be made, and these improvements and substitutions should also be regarded as the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the protection scope of the claims.

Claims

1. An optimization scheduling method for a multi-region integrated energy system, characterized in that: The integrated energy system of each area is managed by a corresponding local intelligent agent, and the method includes the following steps: Each local agent trains the preset energy optimization scheduling model based on the reinforcement learning idea to obtain the corresponding local optimization model, and updates the local optimization model according to the encrypted parameter data packets transmitted by each neighboring agent and the pre-constructed double random matrix; the double random matrix is ​​constructed based on the topological relationship between the integrated energy systems in each region; the encrypted parameter data packet includes the agent ID, encrypted model parameters and encrypted timestamp; In response to the completion of the update of the local optimization model of each local agent, an overall convergence evaluation is performed on the local optimization models of all local agents. When the corresponding overall convergence evaluation result reaches the preset optimization target, the collaborative training between local agents is stopped, and the integrated energy system optimization scheduling of the corresponding area is executed according to the local optimization model of each local agent.

2. The optimization scheduling method of the multi-regional integrated energy system according to claim 1, characterized in that: The state space and action space of the preset energy optimization scheduling model are respectively defined according to the operating state and scheduling strategy of the integrated energy system; the operating state includes electric load demand, thermal load demand, cooling load demand, photovoltaic power generation, energy storage unit status, energy scheduling interval and future electricity price; the scheduling strategy includes the electric output power of the trigeneration unit, the output power of the photovoltaic generator and the energy storage charging and discharging power; the reward function of the preset energy optimization scheduling model is constructed based on the operating cost of the integrated energy system scheduling and the penalty for the intelligent agent violating the constraints.

3. The optimization scheduling method of the multi-regional integrated energy system according to claim 1, characterized in that: The steps of training each local intelligent agent on a preset energy optimization scheduling model based on the reinforcement learning idea to obtain a corresponding local optimization model include: Each local intelligent agent obtains the historical data of the integrated energy system of the corresponding area, and preprocesses the historical data of the integrated energy system to obtain the corresponding energy system optimization scheduling data set; According to the integrated energy system optimization scheduling data set, the preset energy optimization scheduling model is iteratively trained based on a double-delay deterministic policy gradient algorithm to obtain the local optimization model; The model parameters of the local optimization model are encrypted based on a preset encryption algorithm to generate a corresponding encrypted parameter data packet, and the encrypted parameter data packet is sent to each neighboring intelligent agent of the local intelligent agent.

4. The optimization scheduling method of the multi-regional integrated energy system according to claim 3, characterized in that: The step of encrypting the model parameters of the local optimization model based on a preset encryption algorithm to generate a corresponding encrypted parameter data packet includes: According to the preset parameter sorting rules, the model parameters are arranged in the form of column vectors to obtain the corresponding parameter matrix; Using a homomorphic encryption algorithm to perform encryption operation on the parameter matrix to obtain encrypted model parameters; The encrypted model parameters are combined with the corresponding agent ID and the encrypted timestamp to generate a data packet to be interacted with, and the data packet to be interacted with is encrypted based on an asymmetric encryption algorithm to generate the encrypted parameter data packet.

5. The optimization scheduling method of the multi-regional integrated energy system according to claim 1, characterized in that: The steps of constructing the double random matrix include: Generate the agent adjacency matrix based on the topological relationship between the integrated energy systems in each region; Based on the agent adjacency matrix, obtain the neighboring agent sets of each local agent respectively; Based on the principle of double random matrix generation, the double random matrix is ​​generated according to the set of neighboring agents of all local agents; the double random matrix is ​​expressed as: In the formula, represents the first Line The element value of the column; and Represents the local agent and local agents The number of neighbor agents; Represents the local agent The set of neighbor agents; Represents the total number of local agents.

6. The optimization scheduling method of a multi-regional integrated energy system according to claim 1, characterized in that: The step of updating the local optimization model according to the received encrypted parameter data packets transmitted by each neighboring agent and the pre-built double random matrix includes: Obtain the receiving timestamp of the encrypted parameter data packets of each neighboring agent respectively, and decrypt each encrypted parameter data packet to obtain the corresponding agent ID, encrypted model parameters and encrypted timestamp; Based on the receiving timestamp and encryption timestamp of the encrypted parameter data packet corresponding to each agent ID, the neighbor communication delay between each neighbor agent and the local agent is obtained; Based on the neighbor communication delay between each neighbor agent and the local agent, the weight coefficient of the corresponding position in the double random matrix is ​​corrected to obtain the corresponding local corrected double random matrix; According to the local modified double random matrix, the encrypted model parameters of each neighboring intelligent agent are weightedly aggregated with the local model parameters of the local optimization model to obtain updated local model parameters, and based on the updated local model parameters, an updated local optimization model is obtained.

7. The optimization scheduling method of the multi-regional integrated energy system according to claim 6, characterized in that: The step of correcting the weight coefficients of the corresponding positions in the double random matrix based on the neighbor communication delay between each neighbor agent and the local agent to obtain the corresponding local corrected double random matrix includes: According to the communication delays of all neighbors corresponding to the local agent, the corresponding average value of the neighbor interaction delay is obtained; Calculate the delay deviation rate of each neighbor communication delay and the average value of the neighbor interaction delay respectively; When the delay deviation rate is a positive value, the corresponding weight coefficient reduction value is calculated according to the delay deviation rate and the weight coefficient of the corresponding position of the corresponding neighboring agent in the double random matrix, and the weight coefficient of the corresponding position in the double random matrix is ​​reduced and updated according to the weight coefficient reduction value; According to the reduced values ​​of all weight coefficients, the total weight increase target of all neighbor agents with negative delay deviation rates is obtained, and the weight coefficients of the corresponding positions of all neighbor agents with negative delay deviation rates in the double random matrix are increased and updated according to the total weight increase target.

8. The optimization scheduling method of a multi-regional integrated energy system according to claim 6, characterized in that: The updated local model parameters are expressed as: In the formula, and Respectively represent Local Agents in Round Cooperative Training The local model parameters and the corresponding neighbor agents Encrypted model parameters; Indicates Local agent after round of collaborative training Updated local model parameters; represents the weight coefficient of the local model parameters in the local corrected double random matrix; Represents the neighboring agents in the local modified doubly stochastic matrix The weight coefficients of the encryption model parameters; Represents the local agent The set of neighboring agents.

9. The optimization scheduling method of a multi-regional integrated energy system according to claim 1, characterized in that: The step of evaluating the overall convergence of the local optimization models of all local agents comprises: The test accuracy standard deviation and the average parameter update amplitude of each local optimization model are obtained, and the corresponding model convergence evaluation index is obtained according to the test accuracy standard deviation and the average parameter update amplitude; the model convergence evaluation index is expressed as: In the formula, Indicates Model convergence evaluation index of the local optimization model obtained by round-by-round collaborative training; Before The standard deviation of the test accuracy of the local optimization model obtained by rounds of collaborative training; Before The average parameter update amplitude of the local optimization model obtained by rounds of collaborative training; and Represents the target test accuracy standard deviation and target parameter update amplitude; represents the weight coefficient; Determine whether the model convergence evaluation indexes of all local optimization models have reached the preset model convergence threshold. If so, the corresponding model convergence evaluation result is obtained as converged. Otherwise, the corresponding model convergence evaluation result is determined as not converged. The overall convergence evaluation result is obtained according to the model convergence evaluation results of all local optimization models.

10. An optimization dispatching system for a multi-region integrated energy system, characterized in that: The integrated energy system of each area is managed by the corresponding local agent, and the system includes: The agent training module is used for each local agent to train the preset energy optimization scheduling model based on the reinforcement learning idea, obtain the corresponding local optimization model, and update the local optimization model according to the encrypted parameter data packet transmitted by each neighboring agent and the pre-constructed double random matrix; the double random matrix is ​​constructed based on the topological relationship between the integrated energy systems of each region; the encrypted parameter data packet includes the agent ID, encrypted model parameters and encrypted timestamp; The energy optimization scheduling module is used to respond to the completion of the update of the local optimization model of each local intelligent agent, perform an overall convergence evaluation on the local optimization models of all local intelligent agents, and when the corresponding overall convergence evaluation result reaches the preset optimization target, stop the collaborative training between local intelligent agents, and execute the comprehensive energy system optimization scheduling of the corresponding area according to the local optimization model of each local intelligent agent.

Citation Information

Patent Citations

  • Multi-regional power grid collaborative optimization method, system and device and readable storage medium

    CN115333111A

  • Bandwidth-aware decentralized federated learning method and device

    CN116016212A

  • Comprehensive energy system optimal scheduling method and system based on federal reinforcement learning

    CN117151308A

  • Comprehensive energy system economic dispatching model method based on deep reinforcement learning

    CN119273066A

  • Comprehensive energy system model prediction control method based on deep reinforcement learning

    CN119443674A

Cited By

  • Configuration method, device and system of comprehensive energy

    CN120471399A

  • Comprehensive energy system distributed optimization scheduling method and system considering privacy protection

    CN120806575A