Multi-agent collaborative optimization control method and system for regional energy interconnection
By constructing a Markov game model in a multi-region energy interconnection system and using GNN graph neural network to capture the coupling relationship between single agents, and training the PPO algorithm in combination with the federated learning framework, the problem that the difference between single agents in traditional systems is not fully considered, and efficient multi-agent collaborative optimization control is achieved.
Patent Information
- Application Number
- CN202510093303.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional multi-regional energy interconnection systems fail to fully consider the differences between different single agents, resulting in poor coordination among units and low resource utilization efficiency.
A multi-agent collaborative optimization control method for regional energy interconnection is proposed. By dividing the energy interconnection system into multiple single agents, a Markov game model is constructed, and the coupling relationship between single agents is captured using GNN graph neural network, a federated learning framework is introduced, and the model parameters are aggregated and updated through the PPO algorithm training method, and a distributed PPO algorithm training model is constructed.
It improves the accuracy of status representation, optimizes the collaboration between the agents of the multi-region energy interconnection system, and achieves efficient, stable and flexible system operation.
Smart Images

Figure CN120031301A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-agent collaborative optimization control, and more specifically to a multi-agent collaborative optimization control method and system for regional energy interconnection. Background Art
[0002] With the increasing depletion of traditional fossil energy, distributed renewable energy such as solar energy and wind energy has attracted more and more attention. However, renewable energy has the characteristics of intermittency and instability, resulting in a large number of abandoned solar and wind power. With the upgrading of Internet applications, building a regional energy network has become the development direction of energy applications. However, the power grid, cold network, and heat network each have different response characteristics. Obviously, the optimization and regulation of a single network cannot achieve the optimal multi-energy network, nor can it make the most of renewable energy.
[0003] As an effective form of energy management for distributed energy generation, the multi-regional energy interconnection system can effectively improve the comprehensive utilization rate of energy and enhance power supply reliability and power quality.
[0004] Although the traditional multi-regional energy interconnection system has achieved good results in scheduling research, the differences between different single intelligent agents are not fully considered in the model, resulting in poor coordination between units and low resource utilization efficiency. Summary of the invention
[0005] In response to the problems existing in the above-mentioned fields, the present invention proposes a multi-agent collaborative optimization control method and system for regional energy interconnection, which can solve the technical problems of poor coordination between units and low resource utilization efficiency caused by the differences between different single agents, although good results have been achieved in the research on scheduling of traditional multi-regional energy interconnection systems.
[0006] In order to solve the above technical problems, the present invention discloses a multi-agent collaborative optimization control method for regional energy interconnection, comprising the following steps:
[0007] The energy interconnection system is divided into multiple single agents, and the Markov game model is constructed through the system parameters of each single agent;
[0008] According to the Markov game model, each single agent is regarded as a node in the GNN graph neural network graph. The edges between the nodes represent the direct interaction or dependency relationship between the single agents. The coupling relationship between multiple single agents is captured through the GNN graph neural network.
[0009] The federated learning framework is introduced to train the PPO algorithm of each single agent to obtain the PPO algorithm training model parameters of each single agent; according to the coupling relationship between the single agents, the obtained PPO algorithm training model parameters of each single agent are aggregated and updated to build a distributed PPO algorithm training model;
[0010] According to the distributed PPO algorithm training model, the system parameters of each single agent in the energy interconnection system are updated to obtain the control strategy of multi-agent collaborative optimization.
[0011] Preferably, dividing the energy interconnection system into multiple single agents comprises the following steps:
[0012] Based on the operation mechanism of the energy interconnection system, the operation principle of each intelligent agent is studied, and the energy interconnection system of regional energy interconnection is divided into four single intelligent agents, namely, power generation unit, energy storage unit, load unit and energy exchange unit, among which:
[0013] The power generation unit is the main power source of the system, including photovoltaic power generation, wind power generation and diesel generator;
[0014] The energy storage unit is responsible for storing and releasing electric energy, regulating the balance of power supply and demand between power generation and load. The components of the energy storage unit are the energy storage system;
[0015] The load unit, as the power consumption part, is composed of various loads;
[0016] The energy exchange unit is responsible for the power exchange and management between the microgrid and the main grid and within the microgrid to reflect the power exchange efficiency and voltage level.
[0017] Preferably, the step of constructing a Markov game model using the system parameters of each single agent comprises the following steps:
[0018] The Markov game model building process is a mathematical framework for making decisions in an uncertain and dynamically changing environment. The decision-making process consists of six elements, which are <S,O i ,A i ,P,R,γ〉, where:
[0019] S represents the set of all states of the public environment, and the state at time t is s t ∈S;
[0020] O i is the local observation set of agent i, and the local observation of agent i at time t is o i,t ∈O i , the observations of n agents are combined to form the joint observation O = O 1 O 2 ...n ;
[0021] A i is the action set of agent i, and the local action of agent i at time t is a i,t ∈A i , the actions of n agents combine to form a joint action A = A 1 A 2 ...A n ;
[0022] P is the state transition probability, indicating that in state s t The following joint actions are performed to form joint action a t Then the environment enters the next state s t+1 probability;
[0023] R is the reward function, which means that in state s t The next group of agents performs a joint action t After that, the feedback reward given by the environment satisfies s t a t →r t ,r t ∈R;
[0024] γ is the discount factor.
[0025] Preferably, capturing the coupling relationship between multiple single agents through a GNN graph neural network comprises the following steps:
[0026] Each agent is considered as a node in the graph, and the edges between nodes represent the direct interactions or dependencies between agents;
[0027] The feature vector of each node is defined based on its observation space and used as part of the GNN graph neural network input for subsequent information transmission and node update. The information function is used to collect key information from each neighbor node. The process is expressed as:
[0028]
[0029] Where b is the bias term; d is the number of iterations; W d is a learnable weight matrix used to adjust the influence of neighbor information; is the feature vector of node j at the dth iteration;
[0030] The collected key information is normalized and the aggregated information is expressed as:
[0031]
[0032] Where N(i) is the neighbor set of node i;
[0033] Each node uses its current state and aggregated information to update its feature vector. The process is expressed as:
[0034]
[0035] in, is the feature vector of node i at the dth iteration; U x It is the weight matrix of the x-th layer used to update the node’s own feature vector; is the activation function;
[0036] By stacking multiple GNN layers, more complex relationships are captured, and each GNN layer updates the node features to contain more information about its neighbors.
[0037] Preferably, the step of obtaining the PPO algorithm training model parameters of each single agent comprises the following steps:
[0038] Introducing a federated learning framework, under which each single agent uses the PPO algorithm to train locally and uploads the updated model parameters rather than data to the central server; the central server aggregates the received model parameters to form a global model, and then sends the updated global model parameters to the participants of each single agent;
[0039] Each single agent i samples the state, action, reward, and next state quadruple from the local environment: (S i ,A i ,R i ,S i '),in:
[0040] S i is the current state, representing the observation space of each single agent of the power generation unit, energy storage unit, load unit and energy exchange unit;
[0041] A i Indicates that in state S i The actions taken under the above conditions represent the action space representation of each single agent of the power generation unit, energy storage unit, load unit and energy exchange unit;
[0042] R i is the reward function for each single agent of the power generation unit, energy storage unit, load unit and energy exchange unit;
[0043] S i ' is the action A performed by each single agent i The next state to transfer to;
[0044] Use the PPO algorithm to calculate the advantage function A t, to estimate the relative goodness of an action in a certain state, the calculation formula of its advantage function is:
[0045] A t =Q(S i ,A i )-V(S i )
[0046] Among them, Q(S i ,A i ) is the state-action value function, which means that agent i is in state S i Next take action A i The expected return of V(S i ) is the state value function, indicating that in state S i expected return.
[0047] Preferably, the aggregating and updating the acquired PPO algorithm training model parameters of each single agent specifically includes:
[0048] After calculating the advantage function A t After that, the policy parameters are updated, and the clipping objective function is introduced into the PPO algorithm to limit the amplitude of the policy update;
[0049] The optimization objective of the pruned policy loss is:
[0050]
[0051] in, represents the expectation at time step t; θ is the policy network parameter; θ old is the old strategy parameter; ∈ is the cutting function, which is used to limit the strategy update range; S t is the state of the agent; A t is the action taken by the agent; L CLIP (θ) is the policy loss function after pruning; π θ (A t |S t ) indicates that the current strategy is in state S t Next select Action A t probability; For the old strategy in state S t Next select Action A t probability; is the advantage function estimate;
[0052] At the same time, the PPO algorithm optimizes the value function, and the loss function of the value function is:
[0053]
[0054] Among them, L VF(φ) is the value function loss; V φ (S t ) represents the state value function under parameter φ, which is expressed in state S t The expected return under R t represents the cumulative return at time step t;
[0055] The total loss function of the PPO algorithm combines the policy loss and the value function loss, and also includes an entropy regularization term to encourage exploration:
[0056] L(θ,φ)=L CLIP (θ)-c 1 L VF (φ)+c 2 S[π θ ](s t )
[0057] Among them, c 1 is the weight hyperparameter of the value function loss; c 2 is the weight hyperparameter of the entropy regularization term; S[π θ ](s t ) represents the strategy π θ In status t The entropy of .
[0058] Preferably, the obtaining of the control strategy for multi-agent collaborative optimization specifically comprises the following steps:
[0059] After completing the model training of the PPO algorithm, update the model parameters:
[0060]
[0061] in, is the strategy parameter at time step t; κ is the learning rate; represents the objective function gradient;
[0062] Each single agent will update the model parameters of the PPO algorithm Uploaded to the central server, the central server performs weighted average based on the model parameters of the PPO algorithm uploaded by all single agents to form a new global parameter θ t+1 , the weighted average formula is:
[0063]
[0064] Among them, n i is the amount of local data of agent i;
[0065] The central server aggregates the model parameters θ t+1 Update to the new global model parameters, and set the new global model parameters θ t+1The new global model parameters are sent to all single agents, and each single agent uses the new global model parameters to update the system parameters of each single agent in the next round of training to complete the collaborative optimization control between the single agents and obtain the control strategy of multi-agent collaborative optimization.
[0066] Preferably, the multi-agent optimization process is further divided into discrete action strategy optimization and continuous action strategy optimization, which specifically includes the following steps:
[0067] Collect the sampled state, action, and reward data under the current strategy of the multi-agent, and calculate the state-action value function and advantage function:
[0068]
[0069] Among them, s is the multi-agent state; a is the multi-agent action; Q π (s,a) is the state-action value function, which represents the expected return of selecting action a in state s; V π (s) is the state value function;
[0070] For discrete actions, the policy gradient method is used to update the policy parameters, and the policy gradient update formula is:
[0071]
[0072] in, represents the expected value; J(θ) represents the objective function under the strategy parameter θ; represents the gradient of the policy parameter θ; π θ (a|s) is the policy function, which represents the probability of taking action a in state s; θ is the policy parameter;
[0073] For continuous actions, the policy gradient method is used to update the policy parameters, and the entropy regularization term is introduced to encourage exploration. The gradient update formula is:
[0074]
[0075] in, is the entropy regularization coefficient, which is used to balance exploration and utilization; H(π θ (·|s) is the entropy function, which represents the entropy of the strategy in state s. The expression of the entropy function is:
[0076]
[0077] The objective function of the comprehensive optimization process based on discrete action strategy optimization and continuous action strategy optimization is expressed as:
[0078]
[0079] in, Discrete action sets; A collection of continuous actions.
[0080] Preferably, it also includes a multi-agent collaborative optimization control system for regional energy interconnection, including:
[0081] The Markov game model building module is used to divide the energy interconnection system into multiple single agents and build a Markov game model through the system parameters of each single agent;
[0082] The information capture module is used to regard each single agent as a node in the GNN graph neural network graph according to the Markov game model. The edges between the nodes represent the direct interaction or dependency relationship between the single agents, and the coupling relationship between multiple single agents is captured through the GNN graph neural network;
[0083] The distributed PPO algorithm training model construction module is used to introduce the federated learning framework, train the PPO algorithm of each single agent, and obtain the PPO algorithm training model parameters of each single agent; according to the coupling relationship between the single agents, the obtained PPO algorithm training model parameters of each single agent are aggregated and updated to construct a distributed PPO algorithm training model;
[0084] The collaborative optimization control strategy output module is used to update the system parameters of each single agent in the energy interconnection system according to the distributed PPO algorithm training model, and obtain the control strategy of multi-agent collaborative optimization.
[0085] Compared with the prior art, the present invention has the following beneficial effects:
[0086] The present invention proposes a multi-agent collaborative optimization control method. The method can capture the complex relationship and dynamic interaction between the agents in the microgrid group through the constructed Markov game model and the GNN graph neural network, thereby improving the accuracy of state representation. The GNN graph neural network captures the coupling relationship between the single agents of the power generation unit, the energy storage unit, the load unit and the energy exchange unit, and regards each single agent as a node in the GNN graph neural network graph. The edges between the nodes represent the direct interaction or dependency relationship between the single agents. Each GNN layer will update the node features. By stacking multiple such GNN layers, the model can capture more complex relationships, so that it contains more information about its neighbors. The federated learning framework is introduced to aggregate and update the PPO algorithm training model parameters of each single agent to form a global model, and a distributed PPO algorithm training model is constructed to improve the overall performance of the energy interconnection system and the synergy between agents. The method can optimize the collaboration between the agents in the multi-regional energy interconnection system and realize efficient, stable and flexible system operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 This is a flow chart of the multi-agent coordinated optimization control method for regional energy interconnection proposed by the present invention;
[0088] Figure 2 The dispatching results of microgrid 1 on a certain day based on the method proposed in the present invention under three microgrid operating conditions. DETAILED DESCRIPTION
[0089] The following will be combined with the attached embodiment of the present invention Figure 1-Figure 2 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms described in the present invention are only used to describe specific implementation methods and are not used to limit the present invention.
[0090] like Figure 1 As shown, it is a flow chart of a multi-agent collaborative optimization control method for regional energy interconnection proposed by the present invention.
[0091] In order to solve the complex relationships and dynamic interactions among the components in the energy interconnection system, the present invention studies the operating principles of each intelligent agent based on the operating mechanism of the energy interconnection system, divides the system into a multi-agent system consisting of four single agents: power generation unit, energy storage unit, load unit and energy exchange unit, constructs a Markov game model, and uses the GNN graph neural network to capture the complex relationships between the agents.
[0092] Based on the principle of federated learning, each intelligent agent performs PPO algorithm training locally, and at the same time aggregates and updates global parameters through the central server to build a distributed PPO algorithm training model.
[0093] Based on the multi-scale policy gradient optimization method, an optimization model for processing discrete and continuous actions is constructed to further improve the overall performance of the system and the synergy between intelligent agents.
[0094] The multi-agent collaborative optimization control method specifically includes the following steps:
[0095] Step 1: Based on the operation mechanism of the energy interconnection system, the operating principles of each intelligent agent are studied, a Markov game model is constructed, and the GNN graph neural network is used to capture the complex relationships between the intelligent agents.
[0096] The present invention distributes a multi-agent system of multi-region energy interconnection into four single agents, namely a power generation unit, an energy storage unit, a load unit and an energy exchange unit.
[0097] Among them, the power generation unit is the main source of electricity for the system, and its components include photovoltaic power generation, wind power generation, diesel generators, etc. The energy storage unit is responsible for the storage and release of electric energy, and adjusts the balance of power supply and demand between power generation and load. Its composition is mainly composed of energy storage systems. The load unit represents the power consumption part, which is mainly composed of various loads. The energy exchange unit is responsible for the power exchange and management between the microgrid and the main grid and within the microgrid, and can reflect factors such as power exchange efficiency and voltage level.
[0098] The construction process of the Markov game model (MDP) is a mathematical framework for making decisions in an uncertain and dynamically changing environment. It mainly consists of six elements, which are <S,O i ,A i ,P,R,γ represents.
[0099] Among them, S represents the set of all states of the public environment, and the state at time t is s t ∈S;O i is the local observation set of agent i, and the local observation of agent i at time t is o i,t ∈O i , the observations of n agents are combined to form the joint observation O = O 1 O 2 ... n ; A i is the action set of agent i, and the local action of agent i at time t is a i,t ∈A i , the actions of n agents combine to form a joint action A = A 1 A 2 ...A n ; P is the state transition probability, indicating that in state s t The following joint actions are performed to form joint action a t Then the environment enters the next state s t+1 The probability of; R is the reward function, which means that in state s t The next group of agents performs a joint action t After that, the feedback reward given by the environment satisfies s t a t →r t ,r t ∈R; γ is the discount factor.
[0100] Specifically, the Markov game model construction process of the power generation unit includes:
[0101] 1) Observation space O 1 : Describes the local observation quantity of the power generation unit, expressed as:
[0102] O 1 = {Pg ,S g ,C w ,C f}
[0103] Among them, P g is the current power generation; S g The generator is running / stopping state; C w For weather conditions, C f Diesel fuel status.
[0104] 2) Action Space A 1 : Indicates the actions that the power generation unit can take, expressed as:
[0105] A 1 = {P d ,S on / off}
[0106] Among them, P d Indicates generator power regulation; S on / off Indicates starting or stopping the generator.
[0107] 3) Reward function R 1 :
[0108] R 1 =α 1 ×E g -β 1 ×E f -γ 1 ×P std
[0109] Among them, α 1 ,β 1 ,γ 1 are all constants; E g The electrical energy generated by the power generation unit; E f The electrical energy converted from fuel consumption of power generation unit; P std Represents the standard deviation of power fluctuation.
[0110] The process of constructing the Markov game model of the energy storage unit includes:
[0111] 1) Observation space O 2 : Describes the local observation quantity of the energy storage unit, expressed as:
[0112] O 2 = {Q soc ,P c ,P d ,C soc ,C s}
[0113] Among them, Q soc is the battery capacity; Pc is the charging rate; P d is the release rate; C soc is the number of charge and discharge cycles; C s The capacity is attenuated.
[0114] 2) Action Space A 2 : Indicates the actions that the energy storage unit can take, expressed as:
[0115] A 2 ={S c ,R c}
[0116] Among them, S c is the battery charging and discharging status; R c For charge and discharge rate control.
[0117] 3) Reward function R 2 :
[0118] R 2 =α 2 ×E soc +β 2 ×E loss -γ 2 ×E c
[0119] Among them, α 2 ,β 2 ,γ 2 are all constants; E soc The electrical energy stored in the battery; E loss E is the electric energy lost due to charging and discharging efficiency; c The power loss due to battery life loss.
[0120] The Markov game model construction process of the load unit includes:
[0121] 1) Observation space O 3 : Describes the local observation quantity of the load unit, expressed as:
[0122] O 3 = {P n ,P p ,T u ,E c}
[0123] Among them, P n is the real-time power demand; P p is the maximum power demand of the load; T u E is the accumulated usage time of the load; c Accumulates power usage for the load.
[0124] 2) Action Space A3 : Indicates the actions that the load unit can take, expressed as:
[0125] A 3 ={S l ,T a}
[0126] Among them, S l is the load switching state; T a is the load regulation.
[0127] 3) Reward function R 3 :
[0128] R 3 =α 3 ×E m -β 3 ×E u -γ 3 ×E b
[0129] Among them, α 3 ,β 3 ,γ 3 are all constants; E m is the load energy to be satisfied; E u is the unmet load energy; E b The power loss caused by load fluctuation.
[0130] The Markov game model construction process of the energy exchange unit includes:
[0131] 1) Observation space O 4 : Describes the local observation quantity of the energy exchange unit, expressed as:
[0132] O 4 = {P e ,V,F,S}
[0133] Among them, P e is the energy exchange power; V is the line voltage, F is the line frequency; S is the grid-connected status.
[0134] 2) Action Space A 4 : Indicates the actions that the load unit can take, expressed as:
[0135] A 4 = {P a ,S t ,P i}
[0136] Among them, P a Indicates adjustment of energy exchange power; S t Indicates grid-connected state switching; P iIndicates adjusting the inverter output.
[0137] 3) Reward function R 4 :
[0138] R 4 =α 4 ×P en -β 4 ×P e.loss -γ 4 ×P v.loss -χ×P f.loss
[0139] Among them, α 4 , β 4 , γ 4 and χ are both constants; P en is the net exchange power; P e.loss is the switching loss; P v.loss Power loss caused by voltage fluctuation; P f.loss The power loss caused by frequency fluctuation.
[0140] GNN graph neural network is a framework that has emerged in recent years and uses deep learning to directly learn graph structured data. It can well capture the complex relationships and dynamic interactions between various intelligent agents in a microgrid group and improve the accuracy of state representation.
[0141] The present invention regards each single agent as a node in the graph, and the edges between nodes represent the direct interaction or dependency between single agents. The present invention defines the feature vector of each node based on its observation space and uses it as part of the GNN graph neural network input for subsequent information transfer and node update. The information function is used to collect key information from each neighbor node. The process is expressed as follows:
[0142]
[0143] Where b is the bias term; d is the number of iterations; W d is a learnable weight matrix used to adjust the influence of neighbor information; is the feature vector of node j at the dth iteration.
[0144] Since the number of neighbors of different nodes may be different, the collected information is normalized and the aggregated information is expressed as follows:
[0145]
[0146] Among them, N(i) is the neighbor set of node i.
[0147] Each node uses its current state and aggregated information to update its feature vector. The process is expressed as:
[0148]
[0149] in, is the feature vector of node i at the dth iteration; U x It is the weight matrix of the x-th layer used to update the node's own feature vector; is the activation function.
[0150] By stacking multiple such GNN layers, the model is able to capture more complex relationships, with each GNN layer updating the node features so that they contain more information about their neighbors.
[0151] Step 2: Based on the principle of federated learning, each single agent performs PPO algorithm training locally, and aggregates and updates global parameters through the central server to build a distributed PPO algorithm training model.
[0152] In a multi-agent system, each single agent has local data, which may involve sensitive information and cannot be shared directly.
[0153] Based on this, the present invention introduces a federated learning framework, under which each single agent uses the PPO algorithm to train locally and uploads the updated model parameters instead of data to the central server. The central server aggregates the received model parameters to form a global model, and then sends the updated global model parameters to each participant, thereby achieving collaborative optimization control while protecting data privacy.
[0154] Each single agent i samples the state, action, reward, and next state quadruple from the local environment: (S i ,A i ,R i ,S i ′).
[0155] Among them, S i is the current state, which is represented by the observation space of each agent constructed in step 1; A i Indicates that in state S i The action taken under the action space is represented by the action space of each agent constructed in step 1; R i is the reward function of each agent constructed in step 1; S i ' is the execution action A constructed in step 1 i Then transfer to the next state.
[0156] Use the PPO algorithm to calculate the advantage function A t , to estimate the relative goodness of an action in a certain state, the advantage function is calculated as:
[0157] A t =Q(S i ,A i )-V(S i )
[0158] Among them, Q(S i ,A i ) is the state-action value function, which means that agent i is in state S i Next take action A i The expected return of V(S i ) is the state value function, indicating that in state S i expected return.
[0159] After calculating the advantage function, the policy parameters are updated.
[0160] In order to stabilize the training process, a clipping objective function is introduced into the PPO algorithm to limit the amplitude of the policy update. The optimization goal of the clipped objective function is:
[0161]
[0162] in, represents the expectation for time step t; θ is the policy network parameter; θ old is the old strategy parameter; ∈ is the cutting function, which is used to limit the strategy update range; S t is the state of the agent; A t is the action taken by the agent; L CLIP (θ) is the policy loss function after pruning; π θ (A t |S t ) indicates that the current strategy is in state S t Next select Action A t probability; For the old strategy in state S t Next select Action A t probability; is the advantage function estimate.
[0163] While optimizing the strategy, the PPO algorithm also optimizes the value function to improve the accuracy of state value estimation. The loss function of the value function is:
[0164]
[0165] Among them, L VF (φ) is the value function loss; V φ (S t ) represents the state value function under parameter φ, which is expressed in state S t The expected return under R trepresents the cumulative return at time step t.
[0166] The total loss function of the PPO algorithm combines the policy loss and the value function loss, and also includes an entropy regularization term to encourage exploration:
[0167] L(θ,φ)=L CLIP (θ)-c 1 L VF (φ)+c 2 S[π θ ](s t )
[0168] Among them, c 1 is the weight hyperparameter of the value function loss; c 2 is the weight hyperparameter of the entropy regularization term; S[π θ ](s t ) represents the strategy π θ In status t The entropy of .
[0169] After completing the PPO algorithm training locally, update the model parameters:
[0170]
[0171] in, is the strategy parameter at time step t; κ is the learning rate; represents the objective function gradient.
[0172] Each single agent will update the model parameters locally Upload to the central server. The central server performs a weighted average based on the model parameters uploaded by all single agents to form a new global parameter θ t+1 , the weighted average formula is:
[0173]
[0174] Among them, n i is the amount of local data of agent i.
[0175] The central server aggregates the model parameters θ t+1 Update to the new global model parameters, and set the new global model parameters θ t+1 The new global model parameters are sent to all single agents, and each single agent uses the new global model parameters to locally update the system parameters of each single agent in the next round of training.
[0176] Step 3: Based on the multi-scale policy gradient optimization method, build an optimization model for handling discrete and continuous actions to further improve the overall performance of the system and the synergy between intelligent agents.
[0177] The action space of the agent may contain both discrete actions and continuous actions. The multi-scale policy gradient optimization method combines the ability to process discrete actions and continuous actions, so that the optimization process can adapt to complex action spaces. In the present invention, the policy gradient optimization of multi-agents is divided into two parts: discrete action policy optimization and continuous action policy optimization.
[0178] First, the sampled state, action, and reward data under the current strategy of the multi-agent are collected, and the state-action value function and advantage function are calculated based on this:
[0179]
[0180] Among them, s is the multi-agent state; a is the multi-agent action; Q π (s,a) is the state-action value function, which represents the expected return of selecting action a in state s; V π (s) is the state value function.
[0181] For discrete actions, the policy gradient method is used to update the policy parameters, and the policy gradient update formula is:
[0182]
[0183] in, represents the expected value; J(θ) represents the objective function under the strategy parameter θ; represents the gradient of the policy parameter θ; π θ (a|s) is the policy function, which represents the probability of taking action a in state s; θ is the policy parameter.
[0184] For continuous actions, the policy gradient method is used to update the policy parameters, and the entropy regularization term is introduced to encourage exploration. The gradient update formula is:
[0185]
[0186] in, is the entropy regularization coefficient, which is used to balance exploration and utilization; H(π θ (·|s) is the entropy function, which represents the entropy of the strategy in state s. The expression of the entropy function is as follows:
[0187]
[0188] Through the discrete action strategy optimization and continuous action strategy optimization process, the objective function of the comprehensive optimization process can be expressed as:
[0189]
[0190] in, Discrete action sets; A collection of continuous actions.
[0191] The present invention also proposes a multi-agent collaborative optimization control system for regional energy interconnection, comprising:
[0192] The Markov game model building module is used to divide the energy interconnection system into multiple single agents and build a Markov game model through the system parameters of each single agent;
[0193] The information capture module is used to regard each single agent as a node in the GNN graph neural network graph according to the Markov game model. The edges between the nodes represent the direct interaction or dependency relationship between the single agents, and the coupling relationship between multiple single agents is captured through the GNN graph neural network;
[0194] The distributed PPO algorithm training model construction module is used to introduce the federated learning framework, train the PPO algorithm of each single agent, and obtain the PPO algorithm training model parameters of each single agent; according to the coupling relationship between the single agents, the obtained PPO algorithm training model parameters of each single agent are aggregated and updated to construct a distributed PPO algorithm training model;
[0195] The collaborative optimization control strategy output module is used to update the system parameters of each single agent in the energy interconnection system according to the distributed PPO algorithm training model, and obtain the control strategy of multi-agent collaborative optimization.
[0196] The multi-agent collaborative optimization control method for regional energy interconnection proposed in the present invention studies the operating principles of each agent based on the operating mechanism of the energy interconnection system, divides the system into a multi-agent system consisting of four single agents: a power generation unit, an energy storage unit, a load unit, and an energy exchange unit, and constructs a Markov game model. The GNN graph neural network is used to capture the complex relationships between the agents, which can well capture the complex relationships and dynamic interactions between the agents in the microgrid group and improve the accuracy of state representation. Based on the principle of federated learning, each single agent is trained with the PPO algorithm locally, and the global parameters are aggregated and updated through the central server to construct a distributed PPO algorithm training model, form a global model, and construct a distributed PPO algorithm training model to improve the overall performance of the energy interconnection system and the synergy between agents. The method proposed in the present invention can optimize the collaboration between the agents in the multi-regional energy interconnection system under the premise of ensuring data privacy, and realize efficient, stable and flexible system operation.
[0197] Example
[0198] like Figure 2 As shown, the scheduling results of the microgrid 1 on a certain day based on the method proposed in the present invention under three microgrid operating conditions.
[0199] Taking microgrid 1 as an example, the power dispatching in different time periods is as follows:
[0200] During the period of 00:00-02:00, the photovoltaic output is 0, the wind power generation is large, and the load demand is small, so electricity is sold to Microgrid 2.
[0201] During the period of 02:00-05:00, the photovoltaic output is 0, and the wind power generation and gas turbine power generation can meet the load demand, so electricity is sold to Microgrid 3.
[0202] During the period of 05:00-07:00, Microgrid 1 purchases electricity from Microgrid 2 and stores part of the electricity during the period of low load demand.
[0203] During the period of 08:00-14:00, the photovoltaic output increased and the load demand reached the maximum. Microgrid 1 met the demand by increasing the gas turbine output and purchasing electricity from Microgrid 2, and charging and discharging activities were frequent.
[0204] During the period of 14:00-18:00, overall, the photovoltaic output increased and the load demand increased; at this time, the system still did not need to purchase electricity from the upper-level microgrid to maintain the balance of electricity supply and demand.
[0205] During the period of 18:00-20:00, the power price of the power grid is high, and the power balance is mainly maintained by discharging and increasing the output of gas turbines.
[0206] During the period of 20:00-24:00, load demand decreased, and Microgrid 1 reduced its discharge, relying on gas turbines and purchasing a small amount of electricity from Microgrid 3 to maintain system stability.
[0207] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
[0208] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.
Claims
1. A multi-agent collaborative optimization control method for regional energy interconnection, characterized in that: The following steps are involved: The energy interconnection system is divided into multiple single agents, and the Markov game model is constructed through the system parameters of each single agent; According to the Markov game model, each single agent is regarded as a node in the GNN graph neural network graph. The edges between the nodes represent the direct interaction or dependency relationship between the single agents. The coupling relationship between multiple single agents is captured through the GNN graph neural network. The federated learning framework is introduced to train the PPO algorithm of each single agent to obtain the PPO algorithm training model parameters of each single agent; according to the coupling relationship between the single agents, the obtained PPO algorithm training model parameters of each single agent are aggregated and updated to build a distributed PPO algorithm training model; According to the distributed PPO algorithm training model, the system parameters of each single agent in the energy interconnection system are updated to obtain the control strategy of multi-agent collaborative optimization.
2. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 1 is characterized in that: The energy interconnection system is divided into a plurality of single intelligent entities, comprising the following steps: Based on the operation mechanism of the energy interconnection system, the operation principle of each intelligent agent is studied, and the energy interconnection system of regional energy interconnection is divided into four single intelligent agents, namely, power generation unit, energy storage unit, load unit and energy exchange unit, among which: The power generation unit is the main power source of the system, including photovoltaic power generation, wind power generation and diesel generator; The energy storage unit is responsible for storing and releasing electric energy, regulating the balance of power supply and demand between power generation and load. The components of the energy storage unit are the energy storage system; The load unit, as the power consumption part, is composed of various loads; The energy exchange unit is responsible for the power exchange and management between the microgrid and the main grid and within the microgrid to reflect the power exchange efficiency and voltage level.
3. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 2 is characterized in that: The method of constructing a Markov game model through the system parameters of each single agent includes the following steps: The Markov game model building process is a mathematical framework for making decisions in an uncertain and dynamically changing environment. The decision-making process consists of six elements, which are <S,O i ,A i ,P,R,γ>, where: S represents the set of all states of the public environment, and the state at time t is s t ∈S; O i is the local observation set of agent i, and the local observation of agent i at time t is o i,t ∈O i , the observations of n agents are combined to form the joint observation O = O1O2...O n ; A i is the action set of agent i, and the local action of agent i at time t is a i,t ∈A i , the actions of n agents are combined to form a joint action A = A1A2...A n ; P is the state transition probability, indicating that in state s t The following joint actions are performed to form joint action a t Then the environment enters the next state s t+1 probability; R is the reward function, which means that in state s t The next group of agents performs a joint action t After that, the feedback reward given by the environment satisfies s t a t →r t ,r t ∈R; γ is the discount factor.
4. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 3 is characterized in that: The method of capturing the coupling relationship between multiple single agents through the GNN graph neural network includes the following steps: Each agent is considered as a node in the graph, and the edges between nodes represent the direct interactions or dependencies between agents; The feature vector of each node is defined based on its observation space and used as part of the GNN graph neural network input for subsequent information transmission and node update. The information function is used to collect key information from each neighbor node. The process is expressed as: Where b is the bias term; d is the number of iterations; W d is a learnable weight matrix used to adjust the influence of neighbor information; is the feature vector of node j at the dth iteration; The collected key information is normalized and the aggregated information is expressed as: Where N(i) is the neighbor set of node i; Each node uses its current state and aggregated information to update its feature vector. The process is expressed as: in, is the feature vector of node i at the dth iteration; U x It is the weight matrix of the x-th layer used to update the node’s own feature vector; is the activation function; By stacking multiple GNN layers, more complex relationships are captured, and each GNN layer updates the node features to contain more information about its neighbors.
5. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 4 is characterized in that: The method of obtaining the PPO algorithm training model parameters of each single agent includes the following steps: Introducing a federated learning framework, under which each single agent uses the PPO algorithm to train locally and uploads the updated model parameters rather than data to the central server; the central server aggregates the received model parameters to form a global model, and then sends the updated global model parameters to the participants of each single agent; Each single agent i samples the state, action, reward, and next state quadruple from the local environment: (S i ,A i ,R i ,S′ i ),in: S i is the current state, representing the observation space of each single agent of the power generation unit, energy storage unit, load unit and energy exchange unit; A i Indicates that in state S i The actions taken under the above conditions represent the action space representation of each single agent of the power generation unit, energy storage unit, load unit and energy exchange unit; R i is the reward function for each single agent of the power generation unit, energy storage unit, load unit and energy exchange unit; S i ' is the action A performed by each single agent i The next state to transfer to; Using the PPO algorithm to calculate the advantage function A t , to estimate the relative goodness of an action in a certain state, the calculation formula of its advantage function is: A t =Q(S i ,A i )-V(S i ) Among them, Q(S i ,A i ) is the state-action value function, which means that agent i is in state S i Next take action A i The expected return of V(S i ) is the state value function, indicating that in state S i expected return.
6. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 5 is characterized in that: The obtained PPO algorithm training model parameters of each single agent are aggregated and updated, specifically including: After calculating the advantage function A t After that, the policy parameters are updated, and the clipping objective function is introduced into the PPO algorithm to limit the amplitude of the policy update; The optimization objective of the pruned policy loss is: in, represents the expectation at time step t; θ is the policy network parameter; θ old is the old strategy parameter; ∈ is the cutting function, which is used to limit the strategy update range; S t is the state of the agent; A t is the action taken by the agent; L CLIP (θ) is the policy loss function after pruning; π θ (A t |S t ) indicates that the current strategy is in state S t Next select Action A t probability; For the old strategy in state S t Next select Action A t probability; is the advantage function estimate; At the same time, the PPO algorithm optimizes the value function, and the loss function of the value function is: Among them, L VF (φ) is the value function loss; V φ (S t ) represents the state value function under parameter φ, which is expressed in state S t The expected return under R t represents the cumulative return at time step t; The total loss function of the PPO algorithm combines the policy loss and the value function loss, and also includes an entropy regularization term to encourage exploration: L(θ,φ)=L CLIP (θ)-c1L VF (φ)+c2S[π θ ](s t ) Among them, c1 is the weight hyperparameter of the value function loss; c2 is the weight hyperparameter of the entropy regularization term; S[π θ ](s t ) represents the strategy π θ In status t The entropy of .
7. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 6 is characterized in that: The method of obtaining the control strategy of multi-agent collaborative optimization specifically includes the following steps: After completing the model training of the PPO algorithm, update the model parameters: in, is the strategy parameter at time step t; κ is the learning rate; represents the objective function gradient; Each single agent will update the model parameters of the PPO algorithm Uploaded to the central server, the central server performs weighted average based on the model parameters of the PPO algorithm uploaded by all single agents to form a new global parameter θ t+1 , the weighted average formula is: Among them, n i is the amount of local data of agent i; The central server aggregates the model parameters θ t+1 Update to the new global model parameters, and set the new global model parameters θ t+1 The new global model parameters are sent to all single agents, and each single agent uses the new global model parameters to update the system parameters of each single agent in the next round of training to complete the collaborative optimization control between the single agents and obtain the control strategy of multi-agent collaborative optimization.
8. The multi-agent collaborative optimization control method for regional energy interconnection according to claim 7 is characterized in that: It also includes dividing the multi-agent optimization process into discrete action strategy optimization and continuous action strategy optimization, which specifically includes the following steps: Collect the sampled state, action, and reward data under the current strategy of the multi-agent, and calculate the state-action value function and advantage function: Among them, s is the multi-agent state; a is the multi-agent action; Q π (s,a) is the state-action value function, which represents the expected return of selecting action a in state s; V π (s) is the state value function; For discrete actions, the policy gradient method is used to update the policy parameters, and the policy gradient update formula is: in, represents the expected value; J(θ) represents the objective function under the strategy parameter θ; represents the gradient of the policy parameter θ; π θ (a|s) is the policy function, which represents the probability of taking action a in state s; θ is the policy parameter; For continuous actions, the policy gradient method is used to update the policy parameters, and the entropy regularization term is introduced to encourage exploration. The gradient update formula is: in, is the entropy regularization coefficient, which is used to balance exploration and utilization; H(π θ (·|s) is the entropy function, which represents the entropy of the strategy in state s. The expression of the entropy function is: The objective function of the comprehensive optimization process based on discrete action strategy optimization and continuous action strategy optimization is expressed as: in, Discrete action sets; A collection of continuous actions.
9. A multi-agent collaborative optimization control system for regional energy interconnection, characterized in that: include: The Markov game model building module is used to divide the energy interconnection system into multiple single agents and build a Markov game model through the system parameters of each single agent; The information capture module is used to regard each single agent as a node in the GNN graph neural network graph according to the Markov game model. The edges between the nodes represent the direct interaction or dependency relationship between the single agents, and the coupling relationship between multiple single agents is captured through the GNN graph neural network; The distributed PPO algorithm training model construction module is used to introduce the federated learning framework, train the PPO algorithm of each single agent, and obtain the PPO algorithm training model parameters of each single agent; according to the coupling relationship between the single agents, the obtained PPO algorithm training model parameters of each single agent are aggregated and updated to construct a distributed PPO algorithm training model; The collaborative optimization control strategy output module is used to update the system parameters of each single agent in the energy interconnection system according to the distributed PPO algorithm training model, and obtain the control strategy of multi-agent collaborative optimization.
Citation Information
Cited By
Universe water supply scheduling method and system
CN120317634A