A large power grid regional coordinated power flow regulation method based on multi-agent policy gradient model
By constructing a two-layer structure of a multi-agent policy gradient model, the problem of policy exploration and utilization in regional coordinated control of large power grids is solved, achieving efficient distributed control of the power grid and improving control effect and stability.
Patent Information
- Application Number
- CN202310159550.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-24
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2043-02-24
AI Technical Summary
Existing technologies for regional coordinated control of large power grids suffer from high variance and non-convergence issues in strategy exploration and utilization, which limits the control effect.
A two-layer structure based on a multi-agent policy gradient model is adopted. Through centralized training and distributed execution of proto agents and routing agents, a multi-agent policy communication method is constructed to reduce the variance and randomness of policy learning and achieve efficient coordinated control of regional power grids.
It achieves efficient distributed control in large power grids, reduces the variance of model training, and improves the stability and efficiency of control, making it suitable for practical application scenarios of complex power grids.
Smart Images

Figure CN116362377B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart grid technology, and relates to an artificial intelligence technology for distributed power flow control in power networks, particularly a method for regional collaborative power flow control in large power grids based on a multi-agent policy gradient model. Background Technology
[0002] With the increasing proportion of energy in the economic structure, power grids have become one of the most extensive and complex high-dimensional dynamic systems in existence. The close interconnection between large and regional power grids, while transmitting electricity over vast distances, also exacerbates the vulnerability and complexity of power grids, significantly increasing the likelihood of power system failures and making the affected areas more widespread and uncontrollable. This poses a severe challenge to the safe and stable operation of modern power networks. Ensuring the safe, stable, and long-term operation of large power grids has always been a matter of widespread concern in academia and industry. In industry, the stability of the power grid relies heavily on automated devices as a safety barrier. When anomalies occur beyond the handling capacity of these devices, the detection devices report to the power grid dispatching agency. The overall power grid security is ensured through the control knowledge of dispatching experts. However, the response time and processing capacity for anomalies are limited by the knowledge and capabilities of these experts. In academia, the entire power grid is generally considered the research object, with emergency control of large power grids as the application background. Digital means, such as artificial intelligence methods like reinforcement learning, are used to achieve the goal of intelligent control and optimization of complex power grid operation and dispatch. Large-scale power grids have problems such as complex network topology, large action space, and high uncertainty in the operation of the power grid system, which makes model exploration difficult and the value function training variance large.
[0003] In real-world scenarios, large-scale power grids are typically regulated regionally based on administrative units such as prefecture-level cities. For each region, the availability of actual power grid dispatch information is very limited, necessitating the ability to assess the regional power grid's operating environment and make localized decision-making inferences. Within the research context of large-scale interconnected power grids, it is necessary to implement zoned management of large power grids. Rational planning and division of the power system network is beneficial for the safe operation, optimized control, and efficient management of each power grid zone, thereby achieving the stability of large-scale interconnected power grids.
[0004] The literature [Glavic M. Design of a Resistive Brake Controller for Power System Stability Enhancement Using Reinforcement Learning[J].IEEE Transactions on Control Systems Technology,2005,13(5):743-751.] studies the application of reinforcement learning algorithms in instantaneous power angle stability control of power grids. The literature [Xu Y,Zhang W,Liu W,et al. Multiagent-BasedReinforcement Learning for Optimal Reactive Power Dispatch[J].IEEE Transactions on Systems Man & Cybernetics Part C,2012,42(6):1742-1751.] studies a reactive power distribution optimization strategy based on multi-agent reinforcement learning. This method does not require an accurate power grid system model, adopts a model-free reinforcement learning algorithm, and is very effective in power systems of different scales, enabling distributed power grid regulation. The literature [Hossain MJ, Rahnamay-Naeini M. Data-Driven, Multi-Region Distributed State Estimation for Smart Grids[C] / / 2021IEEE PESInnovative Smart Grid Technologies Europe(ISGT Europe).IEEE,2021:1-6.] proposes a distributed state estimation method for multiple regions of the power grid to address the need for low-latency data processing of power system data for real-time wide-area monitoring of smart grids. It identifies regions based on the correlation between geographical distance and the state of power system components and evaluates the performance of the distributed data-driven state estimation method using IEEE 118 test cases.The literature [Cao D, Zhao J, Hu W, et al. Data-driven multi-agent deep reinforcement learning for distribution system decentralized voltage control with high penetration of PVs[J].IEEE Transactions on Smart Grid,2021,12(5):4137-4150.] proposes a multi-agent deep reinforcement learning algorithm that can coordinate the active and reactive power control of photovoltaics with existing static reactive power compensators and battery storage systems, and divide the power grid system into different voltage control regions to achieve better distributed control. The superiority of the proposed method is demonstrated on IEEE 123-node and 342-node systems. North China Electric Power University [Zhao Dongmei, Tao Ran, Ma Taiyi, et al. Active-Reactive Coordination Scheduling Model Based on Multi-Agent Deep Deterministic Strategy Gradient Algorithm[J]. Journal of Electrical Engineering,2021,36(9):1914-1925.] adopts multi-agent technology to intelligently organize various active and reactive power control resources and establish a power grid active-reactive coordination scheduling model. The literature [Tang H,Lv K,Bak-Jensen B,et al.Deep neuralnetwork-based hierarchical learning method for dispatch control of multi-regional power grid[J].Neural Computing and Applications,2022,34(7):5063-5079.] introduces a hierarchical learning optimization method based on deep neural networks to establish an online method to solve the problem of centralized coordination and dispatch of interconnected multi-regional power grids, which can effectively redistribute power resources in large-scale multi-regional power grids.
[0005] It is evident that research based on traditional reinforcement learning algorithms is gradually becoming inadequate for distributed control scenarios with limited information acquisition in power grids. Multi-agent reinforcement learning technology has become an effective approach to solve the problem of regional coordinated control in large power grids. However, when multi-agent reinforcement learning technology is applied to the distributed control of large power grids in multiple regions, there are high variance and non-convergence problems in strategy exploration and utilization, which greatly reduces the control effect. Summary of the Invention
[0006] To overcome the shortcomings of the prior art, the present invention aims to provide a large-scale power grid regional collaborative power flow control method based on a multi-agent policy gradient model. This method utilizes an effective multi-agent policy communication method to reduce the variance and randomness of multi-agent policy learning, thereby improving its application effect in actual power grids.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] A method for regional coordinated power flow control in a large power grid based on a multi-agent policy gradient model includes the following steps:
[0009] Step 1: Perform topological partitioning of the large power grid, dividing the large power grid into multiple control areas, such that the electrical distance within each area is similar and the electrical distance between area power grids is large, and determine the number of area power grids N;
[0010] Step 2: Design the state representation vector S, local observation representation vector O, and action representation vector A for the power network; each regional power grid is regulated by a regional agent.
[0011] Step 3: Design a multi-agent policy gradient model based on the multi-agent proximal policy optimization algorithm. This model consists of two layers of agents, and the local observation representation vector o of each agent in the region is used to... i As the input to the first layer, the output is a specific continuous action space vector, i.e., communication action. All communication actions output from the first layer are mapped and concatenated into a global policy communication message. Policy communication information With the local observation representation vector o i As input to the second layer, the output of the second layer is a continuous action. The final action performed by the regional intelligent agent in the environment is the regulatory action;
[0012] Step 4: Construct a simulated power grid operation environment based on the discretized power grid operation dataset. Interact the model with the simulated power grid operation environment. Each regional agent collects regional sample data. The regional agent obtains the current regional observation information and the final action to be executed by the environment from the simulated power grid operation environment. Each regional agent submits the final action to be executed to the simulated power grid operation environment. The environment provides feedback on global instantaneous reward, the next time state, and whether to end the process.
[0013] Step 5: After each regional agent collects a batch of data, it updates the model parameters and then returns to step 4 to continuously interact with the simulated power grid operating environment and train the multi-agent policy gradient model until the model performance converges.
[0014] Step 6: Based on the trained multi-agent policy gradient model, realize the coordinated regulation and control of the large power grid region.
[0015] Compared with existing technologies, this invention constructs a multi-agent model to interact with the power grid simulation environment. The agents can autonomously learn the collaborative mapping relationship between the real-time operating status of the regional power grid and the control actions, realizing the policy communication capability of the multi-agent in centralized training. This capability has an important impact on the training variance and convergence speed of the model in the multi-agent control scenario. Theory and experiments have proven that this invention can be applied to the actual complex distributed control scenario of the power grid.
[0016] In multi-agent regulation tasks, multiple regional agents coexist in the same power grid environment and jointly influence it. Therefore, each regional agent needs to consider the communication information of other regional agents to coordinate regulation when adjusting its own regulation strategy. This invention considers the policy information between regional agents as the efficient communication information required for multi-agent model training. This invention constructs a two-layer model structure: a protoagent model and a routing agent model, which is a framework of centralized training and distributed execution. In the centralized training phase of multi-agent, the protoagent model can provide policy communication information to the routing agent model, enabling the routing agent model to utilize additional policy information from other regional agents to reduce its own variance in policy exploration and evaluation. Subsequently, the protoagent model can fit the updated policy of the routing agent model online to provide more accurate policy information. The two models interact during training, jointly improving performance. The two-layer model structure constructed in this invention can perform efficient centralized communication with minimal communication cost. After the model training is completed, the two-layer model can converge to the same performance. Therefore, in the distributed execution phase, when each regional agent regulates its own regional power grid, the regional agent only needs to deploy the proto agent model and can ensure high-performance regulation without communication between regional agents. Attached Figure Description
[0017] Figure 1 This is the overall flowchart of the present invention.
[0018] Figure 2 This is a schematic diagram of the power network structure in an embodiment of the present invention.
[0019] Figure 3 This is a structural diagram of the multi-agent policy gradient model in an embodiment of the present invention.
[0020] Figure 4 This is a simulation case of regional power grid division of the IEEE 1888 node large power grid in the embodiments of the present invention.
[0021] Figure 5 This is a performance comparison chart of the algorithm of the present invention with the open-source algorithms IPPO (Independent PPO) and MAPPO (Multi-Agent PPO) in the embodiments of the present invention. Detailed Implementation
[0022] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings and examples.
[0023] In large power grids, access to grid information is limited, making traditional reinforcement learning algorithms insufficient to meet the distributed control requirements. Even with the introduction of multi-agent reinforcement learning methods, the high variance and non-convergence of policy exploration and utilization limit the effectiveness of control.
[0024] Therefore, this invention provides a method for regional collaborative power flow control of large power grids based on a multi-agent policy gradient model. By constructing a two-layer model composed of multiple agents, and utilizing the interactive learning between multi-agent reinforcement learning algorithms and simulated power network environments, each regional agent establishes a mapping relationship between power grid state and control behavior. This provides a feasible means for regional and cross-regional control of large power grids, offering a new perspective and method for studying interconnected large power grids. Furthermore, the invention addresses the non-stationarity problem in multi-agent policy learning by designing an algorithm.
[0025] Specifically, such as Figure 1 As shown, the present invention provides a large-area regional collaborative power flow control method for power grids based on a multi-agent policy gradient model, namely, distributed power flow control of multi-regional power grids, which includes the following steps:
[0026] Step 1: Perform topological partitioning of the large power grid, dividing the large power grid into multiple control areas, such that the electrical distance within each area is similar, while the electrical distance between area power grids is far. Determine the number of area power grids N, and regard each area power grid as an area intelligent agent, that is, each area power grid is controlled by an area intelligent agent.
[0027] Based on the fundamental principle of multi-regional decoupling of large power grids, the shortest path between power grid nodes is taken as the basic electrical distance. According to the community detection theory, the number of regional power grids is determined by the specific power grid scale, and the large power grid is divided into multiple control areas in a graph structure by combining geographical location information.
[0028] In community detection theory, the shortest path between grid nodes is used as a calculation metric to determine edge betweenness. Edge betweenness is defined as the proportion of paths passing through a given edge in all shortest paths in the network. Edge betweenness measures the importance of a node or edge; a higher value indicates greater importance. Regions are divided based on the line with the highest edge betweenness in the grid graph structure. If a line is easily included in the shortest path between any two nodes, it can be considered an important line carrying current transmission between nodes, thus dividing the grid into regions. In real-world grid scenarios, such lines are called tie lines. Modularity assesses the density within the graph structure, thereby determining the number of grid regions after division.
[0029] While community detection theory can divide the topology of a power grid, it does not consider the geographical location information of the grid nodes. Specifically, this invention improves the K-means algorithm by dividing the power grid into regions based on the geographical location information of the nodes and the shortest path of the line connection between the nodes. The first step is to randomly select k grid nodes as initial cluster centers based on the number of regional power grids determined by community detection theory. The second step is to calculate the shortest path distance between the remaining nodes and each cluster center node and group them into the nearest cluster. The third step is to calculate the geographical center of the nodes within each cluster and update the cluster centers. Steps two and three are repeated until the nodes within the cluster are stable.
[0030] In terms of data, based on the IEEE open-source power grid benchmark operating data, we construct various operating modes such as load fluctuation, load mutation, and new energy mutation. For each region, we randomly select a mode to generate power grid operating data, while maintaining the dynamic balance between the source and load sides of the overall power grid, thus constructing power data with diverse operating modes.
[0031] Step 2: Design the state representation vector S, local observation representation vector O, and action representation vector A for the power network.
[0032] The state representation vector S, local observation representation vector O, and action representation vector A of the power network are all continuous spatial variables. The state representation vector S includes the generator power, load power, and node voltage of the generators at the overall grid nodes, as well as the power flow and current values on the lines. The local observation representation vector O includes the generator power, load power, and node voltage of the generators at the regional grid nodes. The action representation vector A is the adjustment value of the current output of the generators. The number of electrical components in different regions is not the same.
[0033] For specific applications, such as power grid structure Figure 2As shown, the number of regional power grids N after division is determined. The number of generators, loads and lines on nodes are determined in different regional power grids. Different regional power grids are controlled by different regional agents. The input of the regional agent is the local observation representation vector of the regional power grid. The regional agent determines the input and output dimensions according to the number of electrical components in the region. The output of the regional agent is a high-dimensional Gaussian distribution with the same dimension as the number of generators in the region. After sampling high-dimensional continuous actions on this distribution, it is multiplied by the ramp rate C of each generator as the action adjustment value of the regional agent in a unit time step.
[0034] The components of the state are explained as follows:
[0035] Generator power output: The active power P generated by each generator at the current moment;
[0036] Load power: The total power (including active power and reactive power) of each load node at the current moment;
[0037] Node voltage: The per-unit voltage value of each node at the current moment;
[0038] Line power flow value: The current value and active power value in each power transmission line at the current moment.
[0039] Step 3: Design a multi-agent policy gradient model based on the multi-agent proximal policy optimization algorithm. This model consists of two layers of agents. In this embodiment of the invention, the two agents are a proto agent and a routing agent, respectively. The local observation representation vector o of each agent is... i As the input to the first layer, i.e., the proto agent, the output is a specific continuous action space vector, called the communication action. The communication actions of all regional agents (i.e., all communication actions output by the proto agent) are mapped and concatenated into a global policy communication message through the communication layer. Policy communication information With the local observation representation vector o i As input to the second layer, the routing agent, the output of the routing agent is a series of actions. And it serves as the final action performed by the regional intelligent agent in the environment, namely the regulatory action.
[0040] In embodiments of the present invention, the proto agent consists only of an Actor policy network, while the routing agent consists of two networks: an Actor network and a Critic network. Figure 3The overall structure of the model shown is determined based on the dimensions of the state representation vector S, local observation representation vector O, and action representation vector A designed in step 2. The input and output dimensions of the Actor network and Critic network for each region's agent are then determined. Specifically, the Actor network of the proto agent uses the local observation representation vector as input, while the Actor and Critic networks of the routing agent both use the local observation representation vector and communication information as input.
[0041] In multi-agent cooperative regulation tasks, each regional agent explores and learns policies by maximizing global rewards. Since all agents learn and explore policies in the same environment, each agent is influenced by the policies of other agents during environmental exploration. During training, agents cannot distinguish the impact of other agents' policies on the environment, leading to higher variance and requiring significant time and computational resources. Therefore, this invention considers treating agent policy behavior as communication information. In the centralized training phase, the proto agent learns through imitation and receives this communication information from the Actor network. The routing agent receives this information and then interacts with the environment, improving the stability of the routing agent's policy learning. Because the proto agent infers the routing agent's policy behavior online through imitation learning, the proto agent and the routing agent have similar performance. Therefore, in the actual execution phase, each regional power grid only needs a proto agent model. Given the local observation representation vector of the regional power grid, it can directly output regulation actions and interact with the environment without inter-regional communication. This is called the centralized training distributed execution method.
[0042] The inference model design method is as follows:
[0043] Step 3.1: Determine the structural parameters of the multi-agent policy gradient model, including the number of agents, the dimension of the input layer, the number of neurons in the hidden layer, the activation function, and the dimension of the output layer.
[0044] Initialize model parameters Where θ and ω represent the Actor parameter vectors of the proto agent and routing agent, respectively. This represents the Critic parameter vector of the routing agent, and the number of agents in the region is N.
[0045] Step 3.2, for each regional agent, use the local observation representation vector o of the current region. i As input to the first-layer proto agent model, the output is the communication action. All proto agent communication actions are mapped and concatenated into a global policy communication message through the communication layer.
[0046] Step 3.3, transmit policy communication information With the local observation representation vector o i As input to the second-layer routing agent model, the routing agent outputs continuous actions. As the final action executed by the regional agent in the environment, the environment performs a power flow calculation after receiving the control actions from all routing agents, and feeds back the global reward value of the entire power grid and the state representation vector at the next moment, thereby enabling the regional agent to reason from regional observations to control actions.
[0047] Step 3.4: During the training phase, the model employs a centralized training distributed execution (CTDE) approach. The proto agent's role is to infer the routing agent's true policy for interacting with the environment. Specifically, it infers the routing agent's policy during the centralized training phase to provide policy communication information to aid the routing agent's training; that is, it can pre-provide policy communication information to help the routing agent train during the centralized training phase. The proto agent's goal is to minimize the distribution of communication actions. The actual action distribution output by the routing agent The KL divergence between them; after receiving local observations and policy communication information, the routing agent outputs the control action. The goal is to interact with the environment and maximize the cumulative reward over each round. Where γ is the discount reward coefficient, γ∈[0,1], t is the current time, n is the nth time, and r k It is a global, real-time reward for the environment. The multi-agent regulation of a large power grid is a team-based cooperative form similar to a football match. All regional agents share the same global reward and learn the cooperative strategies among agents through centralized training to jointly optimize the global goal.
[0048] The update loss function of the Actor network for the Proto agent is as follows:
[0049]
[0050] The update loss function of the Actor network for the routing agent is as follows:
[0051]
[0052]
[0053] The update loss function of the routing agent's Critic network is as follows:
[0054]
[0055]
[0056] In the formula, D KL The KL divergence between distributions is represented by θ, and ω represents the Actor parameter vectors of the proto agent and routing agent, respectively. This represents the Critic parameter vector of the routing agent. R i and P i These represent the i-th routing agent model and the proto agent model, respectively. Let o represent the observation representation vector of the i-th proto agent in the current region's power grid. i Actor Network The output; This indicates that the i-th routing agent receives policy communication information. Then, in the current regional power grid observation representation vector o i Actor Network The output; Let represent the policy network of the i-th routing agent before the update, and ∈ represent the confidence region interval, which is used to measure the optimization of the policy network under a certain confidence region.
[0057] This indicates that the routing agent receives policy communication information. Then, in the current regional power grid observation representation vector o i The Critic network below represents an assessment of current observations in the region; This represents the observation and evaluation after the T-th step of regulation in the regional power grid; It is the multi-step advantage function of the routing agent, where 1:T represents the T-step policy evaluation of the routing agent; y i Let be the multi-step TD objective of the i-th routing agent, which means that the routing agent interacts with the environment in multiple steps, collects multi-step sample data, and then performs policy evaluation.
[0058] Step 3.5: Based on the designed model loss function, calculate the forward propagation model loss using the sampled batch data, and jointly optimize and update the model parameters of the proto agent and routing agent through gradient backpropagation.
[0059] The routing agent's Actor-Critic network is updated using a near-end policy optimization reinforcement learning algorithm, while the proto agent's Actor network uses imitation learning to infer the routing agent's policy behavior online.
[0060]
[0061]
[0062]
[0063] In the formula, and These represent the Actor parameter vectors before and after the update of the j-th proto agent, respectively; and These represent the Actor parameter vectors before and after the update of the j-th routing agent, respectively. and Let represent the Critic parameter vectors before and after the j-th routing agent update; K is a hyperparameter, representing the number of times a batch of training samples can update the network parameters.
[0064] After the above model has converged during training, in the actual distributed execution phase, each regional power grid only needs to deploy a proto agent model to output control actions with only local observation representations as input, so as to quickly respond to abnormal situations in the power grid and achieve the purpose of distributed power flow control of the power grid.
[0065] Step 4: Construct a simulated power grid operation environment based on the discretized power grid operation dataset. In this embodiment of the invention, the open-source computing library pandapower is used as the backend for power grid power flow calculation to construct the simulated power grid operation environment. According to Step 3, the model interacts with the simulated power grid operation environment. Each regional agent collects regional sample data. The regional agent obtains the current regional observation information and the final action to be executed by the environment from the simulated power grid operation environment. Each regional agent submits the final action to be executed to the simulated power grid operation environment for execution. The environment provides feedback on the global instantaneous reward, the next time-instance state, and whether to end the interaction. If the end signal is true, the current round ends, and the power grid state is re-initialized for further interaction; otherwise, the interaction steps are repeated based on the next state.
[0066] Step 5: After each regional agent collects a batch of data, it updates the model parameters and then returns to step 4 to continuously interact with the simulated power grid operating environment and train the multi-agent policy gradient model until the model performance converges.
[0067] Step 6: Based on the trained multi-agent policy gradient model, realize the coordinated regulation and control of the large power grid region.
[0068] After training convergence, the above model, because the proto agent model already considers communication information between regional power grids during the centralized training phase, only needs to deploy the proto agent model in each regional power grid during the actual distributed power grid control phase. No communication between regional power grids is required; with only local observation representations as input, it outputs specific control actions. The control of each regional agent can comprehensively consider the situation of other regional power grids to quickly respond to anomalies in the overall large power grid, achieving the goal of coordinated regional control of the large power grid.
[0069] This invention assumes that when the multi-agent policy gradient model is used for distributed power flow regulation of the power grid, the regulation between regional power grids is parallel. After all regional agents output regional regulation actions, a power flow calculation is required for the entire power grid.
[0070] This invention uses the open-source algorithm PPO as a baseline and proposes the PRPPO (Policy Routing PPO) algorithm based on the two-layer model and centralized policy communication mechanism. The overall process can be summarized as follows:
[0071] Input: Number of iterations T, state set S, observation set O, action set A, number of regional agents N, Actor parameter vectors θ and ω for the proto agent and routing agent, and Critic parameter vector for the routing agent.
[0072] Output: θ, the optimal Actor network parameters for the proto agent;
[0073] Initialization: Multi-agent policy gradient model parameters
[0074] For each iteration, the loop operation is as follows:
[0075] Step 1: Initialize the initial state representation S and obtain the observation representation O of the regional agent;
[0076] For each time step in the current round, perform the following loop operation:
[0077] For each regional agent at the current time step, perform the following operations in a loop:
[0078] Step 2: The Actor network of the proto agent is based on the current local observation representation vector o i Output communication action
[0079] The communication actions of all regional agents are mapped into a global communication message through the communication layer.
[0080] After obtaining the global communication information, the operation is repeated for each regional agent:
[0081] Step 3: Routing agent receives global communication information. With the local observation representation vector o i ;
[0082] Step 4: The routing agent's Actor network outputs realistic actions in interaction with the environment.
[0083] All the real actions of the intelligent agents The splicing is a control action for the entire power grid.
[0084] Step 5: Implement actions for the entire power grid. Obtain global rewards and new state representations;
[0085] Repeat this process until the final state of the current round, obtaining the sequence interaction samples S0, A0, R1, S1, A1, R2, ..., S T-1 A T-1 ,R T ,S T ;
[0086] Construct a loss function based on all interaction samples in a round;
[0087] Step 6 uses the following loss function to update the Actor network parameters of the proto agent:
[0088]
[0089] Step 7 uses the following loss function to update the Actor network parameters of the routing agent:
[0090]
[0091]
[0092] Step 8 uses the following loss function to update the Critic network parameters of the routing agent:
[0093]
[0094]
[0095] Step 9: Perform backpropagation on the loss function and update the parameters.
[0096]
[0097]
[0098]
[0099] Step 10: Return to Step 1 to enter the next round and perform model iteration and interactive updates.
[0100] Using the above-mentioned method of dividing large power grid areas, such as Figure 4 As shown, the IEEE 1888 node power grid is divided into 10 regional power grids as the research object of this invention. Each regional power grid is regarded as an intelligent agent, that is, the entire power grid is controlled by 10 regional intelligent agents in a multi-agent collaborative manner.
[0101] The PRPPO algorithm proposed in this invention is experimentally verified and compared with the open-source algorithms IPPO (Independent PPO) and MAPPO (Multi-Agent PPO). Figure 5 As shown.
[0102] The x-axis represents the number of iterations of model parameters, and the y-axis represents the cumulative reward for each round. This allows for the evaluation of algorithm performance, which gradually converges as the model parameters iterate. Among the three types of algorithms, IPPO employs independent learning, with no communication between agents in the region. This algorithm exhibits significant variance and randomness during training. MAPPO uses a centralized training framework, utilizing global state information as communication information, resulting in high stability during training. PRPPO, the algorithm proposed in this invention, utilizes a two-layer model and employs policy information as communication information during the centralized training phase. It achieves high performance while ensuring training stability, making it the best performing algorithm among the three types.
Claims
1. A method for regional coordinated power flow control in a large power grid based on a multi-agent policy gradient model, characterized in that, Includes the following steps: Step 1: Perform topological partitioning of the large power grid, dividing the large power grid into multiple control areas, such that the electrical distance within each area is similar and the electrical distance between area power grids is large, and determine the number of area power grids N; Step 2: Design the state representation vector S, local observation representation vector O, and action representation vector A for the power network; each regional power grid is regulated by a regional agent. Step 3: Design a multi-agent policy gradient model based on the multi-agent proximal policy optimization algorithm. This model consists of two layers of agents, and the local observation representation vector o of each agent in the region is used to... i As the input to the first layer, the output is a specific continuous action space vector, i.e., communication action. All communication actions output from the first layer are mapped and concatenated into a global policy communication message. Policy communication information With the local observation representation vector o i As input to the second layer, the output of the second layer is a continuous action. The final action performed by the regional intelligent agent in the environment is the regulatory action; Step 4: Construct a simulated power grid operation environment based on the discretized power grid operation dataset. Interact the model with the simulated power grid operation environment. Each regional agent collects regional sample data. The regional agent obtains the current regional observation information and the final action to be executed by the environment from the simulated power grid operation environment. Each regional agent submits the final action to be executed to the simulated power grid operation environment. The environment provides feedback on global instantaneous reward, the next time state, and whether to end the process. Step 5: After each regional agent collects a batch of data, it updates the model parameters and then returns to step 4 to continuously interact with the simulated power grid operating environment and train the multi-agent policy gradient model until the model performance converges. Step 6: Based on the trained multi-agent policy gradient model, realize the coordinated regulation and control of the large power grid region.
2. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 1, characterized in that, In step 1, based on the basic principle of multi-region decoupling of large power grids, the shortest topological path between power grid nodes is used as the basic electrical distance. According to the community discovery theory, the number of regional power grids is determined by the specific power grid scale, and the large power grid is divided into multiple control areas in a graph structure in combination with geographical location information.
3. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 1, characterized in that, In step 2, the state representation vector S, the local observation representation vector O, and the action representation vector A of the power network are all continuous spatial variables. The state representation vector S includes the power generation, load power, and node voltage of generators at the overall grid nodes, as well as the power flow and current values on the lines. The local observation representation vector O includes the power generation, load power, and node voltage of generators at the regional grid nodes. The action representation vector A is the adjustment value of the current output of the generators, and the number of electrical components in different regions is not the same.
4. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 3, characterized in that, Based on the number N of the divided regional power grids, the number of generators, loads, and lines on each node are determined in different regional power grids. Different regional power grids are controlled by different regional agents. The input of the regional agent is the local observation representation vector of the regional power grid, and the input and output dimensions are determined according to the number of electrical components in the region. The output of the regional agent is a high-dimensional Gaussian distribution with the same dimension as the number of generators in the region. After sampling high-dimensional continuous actions on this distribution, the action is multiplied by the ramp rate C of each generator as the action adjustment value of the regional agent in a unit time step.
5. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 1 or 4, characterized in that, In step 3, the first-layer agent is the proto agent, and the second-layer agent is the routing agent. All communication actions output by the proto agent are mapped and concatenated through the communication layer. The proto agent consists only of an Actor policy network, while the routing agent consists of two networks: an Actor network and a Critic network.
6. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 5, characterized in that, The inference design method for the model is as follows: Determine the structural parameters of the model, including the number of agents, the dimension of the input layer, the number of neurons in the hidden layer, the activation function, and the dimension of the output layer; Initialize model parameters Where θ and ω represent the Actor parameter vectors of the proto agent and the routing agent, respectively. The Critic parameter vector represents the routing agent, and the number of agents in the region is N; After receiving the control actions from all routing agents, the environment performs a power flow calculation and feeds back the global reward value of the entire power grid and the state representation vector at the next moment, thereby enabling the regional agent to reason from regional observations to control actions. During the training phase, the model employs a centralized training and distributed execution (CTDE) approach. The proto agent infers the true policy of the routing agent's interaction with the environment, thus providing policy communication information in advance to aid the routing agent's training during the centralized training phase. The proto agent's goal is to minimize... and KL divergence; routing agent output The goal is to interact with the environment and maximize the cumulative reward over each round. Where γ is the discount reward coefficient, γ∈[0,1], t is the current time, n is the nth time, and r k It is the global, immediate reward returned by the environment; The update loss function of the Actor network for the Proto agent is as follows: The update loss function of the Actor network for the routing agent is as follows: The update loss function of the routing agent's Critic network is as follows: In the formula, D KL The KL divergence between distributions is represented by θ, and ω represents the Actor parameter vectors of the proto agent and routing agent, respectively. Represents the Critic parameter vector of the routing agent; R i and P i These represent the i-th routing agent model and the proto agent model, respectively. This represents the local observation representation vector o of the i-th proto agent. i Actor Network The output; This indicates that the i-th routing agent receives policy communication information. Then, in the local observation representation vector o of the current regional power grid i Actor Network The output; This represents the policy network of the i-th routing agent before the update, and ∈ represents the confidence region interval, which is used to measure the optimization of the policy network under a certain confidence region; This indicates that the routing agent receives policy communication information. Then, in the local observation representation vector o of the current regional power grid i The Critic network below represents an assessment of current observations in the region; This represents the observation and evaluation after the T-th step of regulation in the regional power grid; It is the multi-step advantage function of the routing agent, where 1:T represents the T-step policy evaluation of the routing agent; y i For the i-th routing agent, the multi-step TD objective means that the routing agent interacts with the environment in multiple steps, collects multi-step sample data, and then performs policy evaluation. All regional agents share the same global reward, and through centralized training, they learn cooperative strategies among themselves to jointly optimize the global goal. Based on the designed model loss function, the forward propagation model loss is calculated using sampled batch data, and the model parameters of the proto agent and routing agent are jointly optimized and updated through gradient backpropagation.
7. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 6, characterized in that, The routing agent's Actor-Critic network is updated using a near-end policy optimization reinforcement learning algorithm, and the proto agent uses imitation learning to infer the routing agent's policy behavior online, as shown below: In the formula, and These represent the Actor parameter vectors before and after the update of the j-th proto agent, respectively; and These represent the Actor parameter vectors before and after the update of the j-th routing agent, respectively. and Let represent the Critic parameter vectors before and after the update of the j-th routing agent; K is a hyperparameter, representing the number of times a batch of training samples can update the network parameters.
8. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 5, characterized in that, In step 4, a simulated power grid operation environment is constructed based on the discretized power grid operation dataset and pandapower is used as the power grid power flow calculation backend.
9. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 5, characterized in that, In step 4, if the end signal is true, the current round ends and the power grid state is re-initialized for interaction; otherwise, the interaction steps are repeated based on the next state.
10. The large power grid regional cooperative power flow control method based on a multi-agent policy gradient model according to claim 5, characterized in that, In step 6, during distributed grid control, each regional grid only deploys a proto agent model, and there is no need for communication between regional grids. With only local observation representations as input, specific control actions are output.
Citation Information
Patent Citations
Power grid power flow regulation and control decision reasoning method based on depth deterministic strategy gradient network
CN113141012A
Real-time optimal power flow calculation method based on near-end strategy optimization algorithm
CN114566971A