Marl topology control method and system for facilitating renewable energy grid integration
By constructing a multi-agent decision-making problem and topology control method, the overload problem caused by reverse power flow in large-scale power grids was solved, real-time and safe topology control was achieved, and the grid's ability to accept renewable energy was improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- STATE GRID ZHEJIANG ELECTRIC POWER CO LTD QUZHOU POWER SUPPLY CO
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies struggle to achieve real-time, secure, and effective topology control in large-scale power grids, especially in scenarios with a high proportion of distributed energy resources. They cannot simultaneously meet the requirements for network connectivity, equipment constraints, and switching operation frequency, leading to overload and voltage problems caused by reverse power flow.
This paper constructs a topology control problem and a multi-agent decision-making problem under a large-scale power grid. It obtains and preprocesses power grid operation data, divides agents based on topological distance, and constructs the observation, action space and heterogeneous reward and punishment functions for each agent. It adopts a paradigm of centralized training and distributed execution to optimize the agent's strategy and value network, so as to meet the scalability and operational security requirements of large-scale networks.
It achieves the goal of ensuring the effectiveness of multi-agent strategies and value networks while reducing complexity, meeting the scalability requirements of large-scale networks, improving operational security, ensuring network radial structure and equipment operation limitations, effectively solving equipment overload problems, and enhancing the acceptance of renewable energy.
Smart Images

Figure CN122118904A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grid technology, and in particular to a MARL topology control method and system for promoting the grid connection of renewable energy. Background Technology
[0002] With the transformation of the global energy structure, distributed energy resources (DERs), represented by solar and wind power, are increasingly penetrating power distribution networks. While this trend helps achieve clean energy goals, it also poses serious challenges to the stable and safe operation of traditional power distribution networks.
[0003] Specifically, the integration of large-scale distributed energy resources, especially the concentrated output of photovoltaic power generation during peak solar hours, can easily lead to reverse power flows (RPFs), meaning that electricity flows from the user side to the substation. Reverse power flows can cause power flow overload risks in equipment such as lines and transformers, and may also lead to voltage overload problems, thus severely limiting the grid's ability to accommodate renewable energy.
[0004] To address these challenges, an effective technical approach is topology control, also known as Distribution Network Reconfiguration (DNR). This technology dynamically alters the network topology by performing real-time, coordinated opening and closing operations on tie switches and sectionalizing switches within the network. Compared to infrastructure upgrades such as line capacity expansion or new construction, which involve huge investments and lengthy cycles, DNR offers a low-cost, rapidly deployable approach to alleviate line overload, balance network power flow, and thus improve the grid's operational flexibility and its ability to absorb DERs (Distribution Network Adapters). However, in practical applications, especially in complex scenarios with a high proportion of DERs integrated, this technology cannot simultaneously meet the requirements of efficiency, security, and practicality. The core challenge lies in the fact that the control strategy needs to make decisions in a near real-time time while meeting a series of complex and strict operational constraints, such as: First, network connectivity and radial constraints: the network must always be fully connected and maintain a single-power-source, loop-free radial structure at all times; Second, equipment physical limit constraints: the voltage, current, and heat capacity of all equipment (such as lines and transformers) must be within their safe operating limits; Third, switching operation frequency constraints: considering the mechanical wear and service life of switching equipment, the number of operations needs to be limited.
[0005] To overcome these core challenges, various technical paradigms for topology control have emerged in the existing technology, which can be mainly divided into methods based on mathematical optimization, heuristic and metaheuristic methods, and methods based on deep reinforcement learning (DRL). However, these methods also have some limitations in practical applications, mainly including: 1. Mathematical optimization-based methods: such as Mixed Integer Linear Programming (MILP) or Mixed Integer Cone Programming (MICP). These methods can theoretically find high-quality optimal solutions, but their core drawback lies in their extremely high computational complexity. As the scale and complexity of the power grid increase, the solution time grows exponentially, making it difficult to meet the real-time requirements of large-scale online network applications. They are typically only suitable for offline planning.
[0006] 2. Heuristic and metaheuristic methods: such as Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and Harris Hawks Optimization (HHO). These methods are computationally more efficient than purely mathematical optimization methods, but their performance is highly sensitive to hyperparameter adjustments, making them prone to getting trapped in local optima and unable to guarantee optimality of the solution.
[0007] 3. Deep Reinforcement Learning-Based Methods: This is an emerging data-driven decision-making approach that has shown great potential in the field of power grid control. These methods mainly include single-agent DRL methods and methods based on multi-agent deep reinforcement learning (MARL). Traditional single-agent DRL uses a central controller to learn the control strategy for the entire network. When dealing with large-scale power grid topology control problems, it suffers from the "curse of dimensionality." When the number of controllable switches in the network is large, the action space expands dramatically, making model training and convergence extremely difficult, thus rendering it unsuitable for large-scale power grid topology control. MARL-based methods decompose the global control problem into multiple distributed agents. By having multiple distributed agents collaborate to complete the task, the scalability problem of single-agent DRL can be solved. However, applying MARL to solve dynamic topology control problems caused by reverse power flow still presents some challenges: Firstly, handling global constraints is difficult: existing methods struggle to effectively and reliably enforce system-level critical constraints within a distributed decision-making framework, especially the global topology constraint of network radial configuration. Any action that violates this constraint could lead to a large-scale power outage, thus compromising safety. On the other hand, it faces the challenge of high-dimensional discrete action space: when facing a real power grid containing thousands of controllable switches, the ultra-high-dimensional discrete action space formed by their combination is an insurmountable obstacle for existing MARL methods, making it difficult to apply existing technologies to large-scale real systems. Summary of the Invention
[0008] Therefore, the technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a MARL topology control method and system for promoting the grid connection of renewable energy, which can meet the requirements of scalability in large-scale networks, operational security in real-time decision-making, and effectiveness of control strategies.
[0009] To address the aforementioned technical problems, this invention provides a MARL topology control method for promoting grid connection of renewable energy, comprising: We construct topology control and multi-agent decision-making problems under large-scale power grids, acquire and preprocess power grid operation data, and obtain a dataset by classifying and enhancing the power grid operation data according to the reverse power flow index. Based on the dataset, a joint graph capable of capturing the complete operating range of a dynamic power grid is constructed, and agents are partitioned based on topological distance; Construct the observation, action space, and heterogeneous reward and punishment functions for each agent. The observations of the agents include the global power grid state, the power flow analysis results of the entire system, and the agent's local information. The advantage function of each agent is calculated based on its observations, action space, and heterogeneous reward and punishment functions. The policies and value networks of the agents are optimized through a paradigm of centralized training and distributed execution to obtain the optimized multi-agent topology.
[0010] Furthermore, the topology control problem and multi-agent decision-making problem under the large-scale power grid are specifically as follows: The topology control problem of a large-scale power grid is set in a radially operating distribution network. The control objective is to alleviate overload caused by distributed energy sources and balance feeder loads through topology reconfiguration. Multi-agent decision-making problems are solved using tuples. It means that, among them, for A finite set of intelligent agents The set of potential global states for the entire network. For the masked joint action space, For state transition, This is a set of heterogeneous reward and punishment functions for guiding learning; For observation space, This is the discount factor.
[0011] Furthermore, a joint graph capable of capturing the complete operating range of the dynamic power grid is constructed based on the dataset, and agents are partitioned based on topological distance, specifically as follows: By merging multiple samples from the dataset, a joint topology graph that can capture the complete operating range of the dynamic power grid is constructed. This joint topology graph is denoted as... , ,in, Represents the set of vertices. , Let m represent the i-th sample in the dataset, and m represent the number of samples. Let w represent the i-th edge, and w be the set of edge weights; Controllable switches are identified from the vertex set to obtain a controllable switch set. The controllable switch set is clustered to obtain multiple topologically compact regions. Each region is assigned to an agent, which is responsible for the control decision of the corresponding region. When clustering the controllable switch set, the shortest path distance of the controllable switches on the joint topology graph is used as a dissimilarity metric for clustering.
[0012] Furthermore, the heterogeneous reward and punishment function is: , In the formula, Let be the heterogeneous reward and punishment function for the i-th agent at time t. To reflect the overall network performance, The true state at time t Let i be the local cost of the action performed by the i-th agent. Let be the combined action of all agents at time t. Let t be the action of the i-th agent at time t.
[0013] Furthermore, the aforementioned The calculation method is as follows: , In the formula, The preset trend-based reward coefficient, This represents the change in reverse power flow of a single device. The pre-set reward for a successful solution. Calculate the non-convergence penalty for the preset power flow. The pre-defined unblocking penalty, The preset overload penalty coefficient, This is the sum of the per-unit values of all devices in the system whose current or power exceeds the rated value. The preset power outage penalty coefficient, This refers to the number of newly added power outage nodes or loads without power supply caused by the operation. The preset loop-closing penalty coefficient, This represents the number of newly added loop components.
[0014] Furthermore, the aforementioned The calculation method is as follows: , In the formula, The preset power outage penalty coefficient within the area. The number of newly added power-out nodes or unpowered loads within the control area of the i-th agent. The pre-set action is repeated as a penalty. The penalty is set based on the number of preset actions.
[0015] Furthermore, the advantage function is: , In the formula, Let be the advantage function value of the i-th agent at time t. The summation index represents the number of steps to be calculated backwards from the current time t. As a discount factor, For smoothing parameters, For the i-th intelligent agent The timing difference error at each moment; The i-th intelligent agent The method for calculating the timing difference error at time step is as follows: , In the formula, For the i-th intelligent agent Heterogeneous reward and punishment functions at time points, Let i be the value function of the i-th agent. For the i-th intelligent agent Observation of time.
[0016] Furthermore, when optimizing the agent's policy and value network, the optimization objective is to minimize a composite objective function, which includes a truncated agent policy loss, a value function loss, and an entropy reward term.
[0017] Furthermore, the method for calculating the loss of the truncated agent strategy is as follows: , In the formula, The truncated agent policy loss for the i-th agent. Let be the importance sampling rate for the i-th agent at time t. Let be the advantage function value of the i-th agent at time t. For mathematical expectation operators, This is a truncation function. This is the threshold for the cropping range; The method for calculating the value function loss is as follows: , In the formula, Let the loss function be the value function of the i-th agent. For the observation of the i-th agent at time t, Let i be the value function of the i-th agent. Let t be the target reward value for the i-th agent; The entropy reward is calculated as follows: , In the formula, For the entropy reward term of the i-th agent, The entropy coefficient, It is the entropy function. The parameters of the i-th agent are The current policy network at that time.
[0018] The present invention also provides a MARL topology control system for promoting grid connection of renewable energy, comprising: The problem modeling module is used to construct topology control problems and multi-agent decision-making problems under large-scale power grids; The data acquisition module is used to acquire and preprocess power grid operation data, and to classify and enhance the verification of the power grid operation data according to the reverse power flow index to obtain a dataset. The agent partitioning module is used to construct a joint graph that can capture the complete operating range of the dynamic power grid based on the dataset, and to partition agents based on topological distance; A multi-agent modeling module is used to construct the observation, action space, and heterogeneous reward and punishment functions for each agent. The observations of the agents include the global power grid state, the power flow analysis results of the entire system, and the agent's local information. The multi-agent topology optimization module is used to calculate the advantage function of each agent based on the observation, action space and heterogeneous reward and punishment function of each agent. The optimization of the agent's policy and value network is obtained by optimizing the multi-agent topology through a paradigm of centralized training and distributed execution.
[0019] Compared with the prior art, the above-described technical solution of the present invention has the following advantages: This invention constructs a topology control problem and a multi-agent decision-making problem under a large-scale power grid. Based on this, it partitions agents according to topological distance and constructs the observation, action space, and heterogeneous reward / penalty function for each agent. The advantage function of each agent is calculated based on its observation, action space, and heterogeneous reward / penalty function. The optimized multi-agent topology is obtained by optimizing the agents' policies and value network through a paradigm of centralized training and distributed execution. This meets the scalability requirements of large-scale networks, reducing complexity while ensuring the effectiveness of the multi-agent policies and value network. Furthermore, the optimization process comprehensively considers various constraints through heterogeneous reward / penalty functions, improving operational safety (especially satisfying hard constraints such as network radial configuration). Attached Figure Description
[0020] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1 This is a framework diagram of autonomous topology control for a radially operating power distribution network in a preferred embodiment of the present invention.
[0021] Figure 2 This is a flowchart of a method in a preferred embodiment of the present invention.
[0022] Figure 3 This is a flowchart illustrating the masking mechanism in a preferred embodiment of the present invention.
[0023] Figure 4 This is a comparison chart of the ablation experiment training curves of the masking mechanism in the simulation experiment of the preferred embodiment of the present invention.
[0024] Figure 5This is a reverse power flow load diagram of the target transformer before the intelligent agent controls it in a simulation experiment in a preferred embodiment of the present invention.
[0025] Figure 6 This is a reverse power flow load diagram of the target transformer after intelligent agent regulation in a simulation experiment in a preferred embodiment of the present invention. Detailed Implementation
[0026] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0027] This invention discloses a MARL topology control method for promoting grid connection of renewable energy, and constructs a topology control system as follows: Figure 1 The framework shown is for autonomous topology control of distribution networks operating radially, including, for example... Figure 2 The steps shown are as follows: S1: Constructing topology control problems and multi-agent decision-making problems under large-scale power grids.
[0028] The topology control problem of a large-scale power grid is set in a radially operating distribution network containing buses, lines, transformers, normally closed sectionalizing switches and normally open tie switches. The control objective is to alleviate overloads caused by distributed energy sources and balance feeder loads through topology reconfiguration.
[0029] The multi-agent decision problem is modeled as a partially observable Markov Decision Process (POMDP), using tuples. It means that, among them, for A finite set of intelligent agents The set of potential global states for the entire network. For a masked joint action space, the agent starts from a masked joint action space. Select discrete operations to ensure that only valid reconstruction schemes are considered. The state transition is determined by AC power flow simulation. This is a set of heterogeneous reward and punishment functions for guiding learning. , Let i be the heterogeneous reward and punishment function for the i-th agent; The observation space is defined as the space perceived by the i-th agent at time t due to partial observability. The observations in rather than the true state ; This is the discount factor. Choosing this model can facilitate convergence and improve the ability of intelligent agents to process multiple parts of information.
[0030] S2: Acquire and preprocess power grid operation data, and perform hierarchical and enhanced verification of the power grid operation data according to the reverse power flow index to obtain the dataset.
[0031] S2-1: In this embodiment, the acquired power grid operation data originates from a historical operation snapshot of a real substation area (2023-2024). The preprocessing process includes: cleaning the raw data, eliminating inconsistencies, and unifying the component naming and topology structure throughout the dataset.
[0032] S2-2: Screening and Layering of Challenging Scenarios: Applying screening criteria, select snapshots where at least one network component exceeds its rated capacity by a certain percentage due to reverse power flow. Subsequently, calculate the reverse power flow index on the high-voltage side of the main transformer. ,in, It is the reverse power flow. The RPI (Reverse Flow Index) is the rated capacity and can quantify the severity of reverse power flow. Snapshots are graded based on the RPI value.
[0033] S2-3: Scenario Enhancement and Validation: To enhance the diversity of training data, the selected basic operating snapshots were enhanced, including: (1) bounded scaling of net active power injection, (2) adjustment of reactive power injection to maintain a reasonable power factor, and (3) local perturbation of the initial switch configuration while ensuring power flow convergence. Each generated scenario was validated through AC power flow analysis, and only operating snapshots that converged to a feasible solution were retained.
[0034] S3: Construct a joint graph that can capture the complete operating range of the dynamic power grid based on the dataset, and divide the agents based on topological distance.
[0035] S3-1: Through a joint aggregation process, multiple samples (i.e., operational snapshots) in the dataset are merged to construct a joint topology graph that can capture the complete operating range of the dynamic power grid. This joint topology graph is denoted as... , ,in, Represents the set of vertices. , Let m represent the i-th sample in the dataset, and m represent the number of samples. It can describe the power flow section of the power grid; Let w represent the i-th edge, and w be the set of edge weights.
[0036] S3-2: Identify controllable switches from the vertex set to obtain a controllable switch set. Cluster the controllable switch set to obtain multiple topologically compact regions. Each region is assigned to an independent agent, which is responsible for the control decision of the corresponding region. When clustering the controllable switch set, the shortest path distance of the controllable switches on the joint topology graph is used as a dissimilarity metric for clustering.
[0037] S3-2-1: Identify controllable switches from the vertex set to obtain the controllable switch set, denoted as [S3-2-1](S3-2-1). , Let any one of the switch pairs be denoted as . , , , , for Any controllable switch in the system.
[0038] S3-2-2: Using Dijkstra's algorithm, compute the graph. The shortest path distance between all switch pairs is denoted as . .
[0039] S3-2-3: Construct the dissimilarity distance matrix, denoted as [Mathematical structure not provided in the original text]. , Let the value in the p-th row and q-th column of the dissimilarity distance matrix be denoted as . , Initialize K clusters and the center points of K clusters.
[0040] S3-2-4: Based on the dissimilarity distance matrix Each controllable switch is assigned to the nearest center point, and the center point of each cluster is updated to the point within the cluster that minimizes the total dissimilarity.
[0041] S3-2-5: Repeat step S3-2-4 until convergence, obtaining K topologically compact regions, denoted as the set of their center points. Cluster label is .
[0042] S4: Construct the observation, action space, and heterogeneous reward / penalty function for each agent. The agent's observations are constructed to provide global context and local details, including global grid state, system-wide power flow analysis results, and agent-local information. The action space is the set of controllable switching operations within the agent's assigned region. Each agent's control strategy is learned using an Independent Actor-Critic (IAC) architecture. A key feature of this design is that each agent's input, in addition to its private regional information, includes a shared public observation to achieve implicit coordination.
[0043] S4-1: The global power grid state includes state components such as the state of all switches, the state of all selected buses, the state of all selected lines, the state of all selected generators, the state of all selected loads, the state of all selected transformers, the state of all selected external power grids, and a global state repetition flag. The system-wide power flow analysis result includes state components such as global loop component indices (specifically including initial global loop component indices, previous global loop component indices, and current global loop component indices) and global over-limit component indices (specifically including initial global over-limit component indices, previous global over-limit component indices, and current global over-limit component indices). The agent's local information includes state components such as information on unpowered areas, loops, over-limit components, and its own historical action sequence. The observation of the i-th agent at time t (i.e., The state components and dimensional representations of ) are shown in Table 1.
[0044] Table 1. Observational Composition of the Agent
[0045] S4-2: To ensure that only physically and operationally effective actions are considered, apply the following before action selection: Figure 3 The masking mechanism shown is used to effectively prune the action space by setting the original predicted value of inactive actions (i.e., the logit in machine learning) to negative infinity before the softmax function is calculated, thus ensuring the safety of the decision.
[0046] S4-3: The heterogeneous reward and punishment function is: , In the formula, Let be the heterogeneous reward and punishment function for the i-th agent at time t. To reflect the overall network performance, The true state at time t Let i be the local cost of the action performed by the i-th agent. Let be the combined action of all agents at time t. N is the number of agents. Let t be the action of the i-th agent at time t.
[0047] The The calculation method is as follows: , In the formula, Rewards for following trends The preset trend-driven reward coefficient is a preset positive gain coefficient. It represents the change in reverse power flow of a single device, and the sum of the reductions in reverse power flow of all transformers or lines in the network before and after regulation. The pre-set reward for successfully resolving all issues is a positive constant reward. This reward is given when all overload, limit violations, and topology violations are eliminated after network topology adjustments. ; The pre-defined non-convergence penalty for power flow calculation is a negative penalty value. If the network state after the agent takes an action causes the power flow calculation to fail to converge, then this penalty is applied. ; This is a preset disconnection penalty, a negative penalty value. If a topology operation causes unexpected disconnection in the power grid, this penalty is applied. ; As an overload penalty, The preset overload penalty coefficient, This is the sum of per-unit values for all devices in the system whose current or power exceeds the rated value. As a penalty for power outage, The preset power outage penalty coefficient, This refers to the number of newly added power-loss nodes (or unpowered loads) caused by the action; For loop-based penalties, The preset loop-closing penalty coefficient, This represents the number of newly added loop components.
[0048] The The calculation method is as follows: , In the formula, Penalty for power outages in the area The preset power outage penalty coefficient within the area. The number of newly added power-out nodes (or unpowered loads) within the control area of the i-th agent; The pre-defined penalty for repeated actions is a negative penalty value. If the agent's current selected action is a repeated action (or causes state repetition), then this penalty is applied. ; By assigning a pre-defined penalty for the number of actions to each action step, the agent can be motivated to complete the task in the fewest steps.
[0049] In heterogeneous reward and punishment functions, the reward component aims to incentivize agents to strive for desired operational outcomes (such as reducing reverse power flow overload), while the cost and punishment components suppress adverse events (such as safety violations) and inefficient behaviors (such as repetitive actions).
[0050] S5: Calculate the advantage function for each agent based on their observations, action space, and heterogeneous reward / penalty function. Optimize the agent's policy and value network using a centralized training and distributed execution (CTDE) paradigm to obtain the optimized multi-agent topology. The entire optimization process iteratively alternates between the following two stages: Stage 1, parallel data collection and advantage estimation; Stage 2, agent-by-agent policy and value network optimization.
[0051] S5-1: Parallel Data Collection and Advantage Estimation.
[0052] S5-1-1: When collecting data in parallel based on the observation and action spaces of each agent, data is collected by running policy rollouts in multiple parallel instances within the power distribution network environment. For each time step t, the specific process for parallel data collection is as follows: S5-1-1-1: The i-th agent receives its current observation. .
[0053] S5-1-1-2: Through the masking mechanism, the i-th agent calculates the feasibility mask of its action space based on the current state.
[0054] S5-1-1-3: The policy network of the i-th agent is based on observations An action is obtained by sampling from its set of valid actions after masking. The actions of all agents together constitute a joint action. .
[0055] S5-1-1-4: Environmental model based on joint actions The transition to the next state is achieved through an AC power flow simulation. and return an independent reward for the i-th agent. .
[0056] S5-1-1-5: The resulting empirical transition tuple It is stored in the independent experience replay buffer of the i-th agent.
[0057] S5-1-2: The advantage function is calculated using Generalized Advantage Estimation (GAE): , In the formula, Let be the advantage function value of the i-th agent at time t. The summation index represents the number of steps (0, 1, 2, ...) to be calculated from the current time t. As a discount factor, For GAE smoothing parameters, For the i-th intelligent agent Timing difference error at time step.
[0058] The i-th intelligent agent The method for calculating the timing difference error at time step is as follows: , In the formula, For the i-th intelligent agent Heterogeneous reward and punishment functions at time points, Let i be the value function of the i-th agent. For the i-th intelligent agent Observation of time.
[0059] S5-2: Agent-by-agent policy and value network optimization.
[0060] When optimizing the agent's policy and value network, the optimization objective is to minimize a composite objective function, which includes a truncated agent policy loss, a value function loss, and an entropy reward term. The composite objective function is iteratively optimized over multiple rounds on mini-batches of data sampled and shuffled from the experience replay buffer (e.g., using gradient descent).
[0061] The composite objective function is as follows: , In the formula, Let i be the composite objective function of the i-th agent. The truncated agent policy loss for the i-th agent. Used to control the update range of old and new strategies and prevent strategy collapse. Let be the value function loss of the i-th agent, used to measure the accuracy of the value network's predictions. For the entropy reward term of the i-th agent, This is used to encourage strategies to explore and avoid getting trapped in local optima too early.
[0062] The method for calculating the loss of the truncated agent strategy is as follows: , In the formula, The truncated agent policy loss for the i-th agent. Let be the importance sampling rate for the i-th agent at time t, used to measure the difference between the old and new strategies. Let be the advantage function value of the i-th agent at time t. For mathematical expectation operators, This is a truncation function. This is the threshold for the cropping range.
[0063] The method for calculating the value function loss is as follows: , In the formula, Let the loss function be the value function of the i-th agent. For the observation of the i-th agent at time t, Let be the target reward value for the i-th agent at time t.
[0064] The entropy reward is calculated as follows: , In the formula, For the entropy reward term of the i-th agent, The entropy coefficient, It is the entropy function. The parameters of the i-th agent are The current policy network at that time.
[0065] The The calculation method is as follows: , In the formula, The parameters of the i-th agent are The old strategy network of that time.
[0066] S5-3: Training Loop and Model Deployment.
[0067] After the parameters of the policy and value networks are updated, these updated parameters will be broadcast back to each parallel worker for the next round of data collection, i.e., re-execution of S5-1 and S5-2, until the preset number of training iterations is reached and the updates stop, resulting in the optimized results.
[0068] During training, the overall performance of the policy and value networks is periodically evaluated on a fixed set of test scenarios. The policy and value networks with the best performance are used as the optimized multi-agent topology and then deployed in practice.
[0069] This invention also discloses a MARL topology control system for promoting grid connection of renewable energy, comprising: The problem modeling module is used to construct topology control problems and multi-agent decision-making problems under large-scale power grids; The data acquisition module is used to acquire and preprocess power grid operation data, and to classify and enhance the verification of the power grid operation data according to the reverse power flow index to obtain a dataset. The agent partitioning module is used to construct a joint graph that can capture the complete operating range of the dynamic power grid based on the dataset, and to partition agents based on topological distance; A multi-agent modeling module is used to construct the observation, action space, and heterogeneous reward and punishment functions for each agent. The observations of the agents include the global power grid state, the power flow analysis results of the entire system, and the agent's local information. The multi-agent topology optimization module is used to calculate the advantage function of each agent based on the observation, action space and heterogeneous reward and punishment function of each agent. The optimization of the agent's policy and value network is obtained by optimizing the multi-agent topology through a paradigm of centralized training and distributed execution.
[0070] The present invention also discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a MARL topology control method for promoting grid connection of renewable energy.
[0071] The present invention also discloses an apparatus including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements a MARL topology control method for facilitating grid connection of renewable energy.
[0072] This invention constructs a topology control problem and a multi-agent decision-making problem under a large-scale power grid. Based on this, it partitions agents according to topological distance and constructs the observation, action space, and heterogeneous reward / penalty function for each agent. The advantage function of each agent is calculated based on its observation, action space, and heterogeneous reward / penalty function. The optimized multi-agent topology is obtained by optimizing the agents' policies and value network through a paradigm of centralized training and distributed execution. This meets the scalability requirements of large-scale networks, reducing complexity while ensuring the effectiveness of the multi-agent policies and value network. Furthermore, the optimization process comprehensively considers various constraints through heterogeneous reward / penalty functions, improving operational safety (especially satisfying hard constraints such as network radial configuration).
[0073] The control strategy obtained by this invention can maintain the radial structure and connectivity of the network, comply with the operating limitations of all devices, ensure the feasibility of power flow, and meet near real-time latency requirements.
[0074] This invention utilizes a distributed intelligent agent to learn safe and feasible switching operation strategies, effectively addressing operational bottlenecks such as equipment overload caused by high-proportion distributed energy grid integration. Through precise and efficient topology reconfiguration, this invention directly enhances the grid's capacity to accept and absorb renewable energy, providing power system operators with a practical technical tool to effectively reduce wind and solar curtailment, demonstrating significant economic benefits and technological advantages.
[0075] It should be emphasized that the technical solutions described in this invention are not static, and those skilled in the art can make various modifications and extensions according to actual needs. For example: Regarding the extension of state representation: Within the framework of this invention, the representation used to describe the power grid state is not limited to the specific vector form given in the embodiments. Advanced technologies such as Graph Neural Networks (GNNs) can be further integrated to construct a more complex state representation module capable of more deeply perceiving and reasoning about the power grid topology.
[0076] Regarding the expansion of the control action space: The core framework of this invention is not limited to handling discrete switching operations. This framework can be extended to support a hybrid action space, thereby co-optimizing discrete topology switching operations with continuous control commands (e.g., real-time adjustment of distributed energy output, charging and discharging control of energy storage devices, etc.).
[0077] To further demonstrate the beneficial effects of this invention, this embodiment uses a real-world model of a power distribution network in eastern China for simulation evaluation. The network obtained through this invention features highly penetrated distributed energy resources (DERs), a clear historical record of reverse power flow events, and multiple feeders with controllable sectionalizing switches and tie switches. In the simulation experiment, running scenarios for training and testing were generated based on historical archived data from 2023-2024. The multi-agent control framework is implemented using the Stable-Baselines3 library. The standard PPO algorithm is extended to support agent-by-agent update schemes with shared global state observations and area masking actions. Key hyperparameter settings for the agent architecture and training are shown in Table 2.
[0078] Table 2. Main Hyperparameters for Agent Architecture and Training
[0079] Simulation Experiment 1: Validation of the effectiveness of the action masking mechanism (ablation study).
[0080] To verify the importance of the feasibility masking mechanism described in this invention, an agent using the complete model and an identical agent without the masking mechanism were trained in a comparative manner.
[0081] Experimental results are as follows Figure 4 As shown, Figure 4 (a) is a comparison chart for assessing the trend of average rewards. Figure 4 (b) is a comparison chart assessing the trend in average round length. From Figure 4It can be clearly seen that the agent adopting the masking mechanism (the present invention) has a faster learning curve convergence speed and a more stable process. While the agent without masking has extremely large performance fluctuations and difficult convergence in the initial stage of training.
[0082] Conclusion: This comparative experiment proves that the masking mechanism greatly accelerates the learning process and improves stability by preventing the policy from exploring infeasible actions, confirming its crucial role in making complex topology control problems tractable. This directly reflects the inherent security and training efficiency of the solution of the present invention.
[0083] Simulation Experiment 2: Performance comparison and evaluation with the prior art.
[0084] To verify the overall performance advantages of the present invention, in the simulation experiment, the present invention is benchmarked with a centralized single-agent (SA) PPO controller. The test is carried out on two real-world test sets: the standard test set (including 188 scenarios, RPI < 70%) and the difficult test set (including 337 scenarios, 70% < RPI < 80%), and the overload threshold is set to 80%. The experimental results are shown in Table 3.
[0085] Table 3 Performance comparison of different methods on the standard test set and the difficult test set
[0086] As can be seen from Table 3: In terms of effectiveness and efficiency: On both test sets, the success rate of the present invention (99.47% vs 96.26% in the standard set, 82.73% vs 77.68% in the difficult set) is significantly higher than that of the single-agent (SA) baseline. At the same time, the present invention requires fewer average steps to solve faults (2.11 vs 2.90 in the standard set), indicating higher decision-making efficiency. In terms of the reverse power flow mitigation rate, the present invention reaches 100% in the standard set, which is also better than the comparative scheme.
[0087] In terms of the quality of the solution: The present invention can generate a better final system state. Specifically, it shows a higher average load shedding rate (8.06% vs 6.89% in the standard set) and a lower average final maximum load rate (59.03% vs 62.13% in the standard set). A key advantage is that the present invention not only solves the primary overload problem but also can actively correct the pre-existing topology problems in the power grid (such as the average number of closed-loop elements is reduced by 0.754, while the comparative scheme increases).
[0088] In terms of robustness: The scenarios in the difficult test set (such as a large-scale reverse power flow approaching the operating limit) are theoretically not guaranteed to be solved only by topology control. The present invention can still maintain a high success rate under these conditions, highlighting its robustness and practical application advantages relative to the centralized paradigm.
[0089] Simulation Experiment 3: Qualitative Analysis of Control Behavior.
[0090] To more intuitively demonstrate the effects of the control strategy of this invention, through Figure 5 and Figure 6 Presented in all challenging test set scenarios, the reverse power flow load of the target transformer is compared before and after the intervention of this invention. Figure 5 This represents the reverse power flow load of the target transformer before the intelligent agent controls it. Figure 6 This represents the reverse power flow load of the target transformer before the intelligent agent controls it.
[0091] from Figure 5 As can be seen, the initial state exhibits significant overload, manifested by numerous red data points significantly exceeding 80% of the operating threshold plane. The bottom projection also reveals large, continuous areas of severe reverse power flow. From... Figure 6 As can be seen, after intervention by this invention, the number of points exceeding the threshold and the magnitude of these violations were drastically reduced. Correspondingly, the red area in the projection shrank significantly, indicating that the most critical faults were effectively mitigated.
[0092] Conclusion: It can be seen that the present invention can accurately identify and suppress large-scale overloads in complex and extreme scenarios, providing concrete evidence and further proving the beneficial effects of the present invention.
[0093] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0094] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0095] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0096] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0097] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A MARL topology control method for facilitating renewable energy grid integration, characterized by, include: We construct topology control and multi-agent decision-making problems under large-scale power grids, acquire and preprocess power grid operation data, and obtain a dataset by classifying and enhancing the power grid operation data according to the reverse power flow index. Based on the dataset, a joint graph capable of capturing the complete operating range of a dynamic power grid is constructed, and agents are partitioned based on topological distance; Construct the observation, action space, and heterogeneous reward and punishment functions for each agent. The observations of the agents include the global power grid state, the power flow analysis results of the entire system, and the agent's local information. The advantage function of each agent is calculated based on its observations, action space, and heterogeneous reward and punishment functions. The policies and value networks of the agents are optimized through a paradigm of centralized training and distributed execution to obtain the optimized multi-agent topology.
2. The MARL topology control method for promoting grid connection of renewable energy according to claim 1, characterized in that: The topology control problem and multi-agent decision-making problem under large-scale power grids are specifically as follows: The topology control problem of a large-scale power grid is set in a radially operating distribution network. The control objective is to alleviate overload caused by distributed energy sources and balance feeder loads through topology reconfiguration. Multi-agent decision-making problems are solved using tuples. It means that, among them, for A finite set of intelligent agents The set of potential global states for the entire network. For the masked joint action space, For state transition, This is a set of heterogeneous reward and punishment functions for guiding learning; For observation space, This is the discount factor.
3. The MARL topology control method for promoting grid connection of renewable energy according to claim 1, characterized in that: Based on the dataset, a joint graph capable of capturing the complete operating range of a dynamic power grid is constructed, and agents are partitioned based on topological distance, specifically as follows: By merging multiple samples from the dataset, a joint topology graph that can capture the complete operating range of the dynamic power grid is constructed. This joint topology graph is denoted as... , ,in, Represents the set of vertices. , Let m represent the i-th sample in the dataset, and m represent the number of samples. Let w represent the i-th edge, and w be the set of edge weights; Controllable switches are identified from the vertex set to obtain a controllable switch set. The controllable switch set is clustered to obtain multiple topologically compact regions. Each region is assigned to an agent, which is responsible for the control decision of the corresponding region. When clustering the controllable switch set, the shortest path distance of the controllable switches on the joint topology graph is used as a dissimilarity metric for clustering.
4. The MARL topology control method for promoting grid connection of renewable energy according to claim 1, characterized in that: The heterogeneous reward and punishment function is: , In the formula, Let be the heterogeneous reward and punishment function for the i-th agent at time t. To reflect the overall network performance, The true state at time t Let i be the local cost of the action performed by the i-th agent. Let be the combined action of all agents at time t. Let t be the action of the i-th agent at time t.
5. The MARL topology control method for promoting grid connection of renewable energy according to claim 4, characterized in that: The The calculation method is as follows: , In the formula, The preset trend-based reward coefficient, This represents the change in reverse power flow of a single device. The pre-set reward for a successful solution. Calculate the non-convergence penalty for the preset power flow. The pre-defined unblocking penalty, The preset overload penalty coefficient, This is the sum of the per-unit values of all devices in the system whose current or power exceeds the rated value. The preset power outage penalty coefficient, This refers to the number of newly added power outage nodes or loads without power supply caused by the operation. The preset loop-closing penalty coefficient, This represents the number of newly added loop components.
6. The MARL topology control method for promoting grid connection of renewable energy according to claim 4, characterized in that: The The calculation method is as follows: , In the formula, The preset power outage penalty coefficient within the area. The number of newly added power-out nodes or unpowered loads within the control area of the i-th agent. The pre-set action is repeated as a penalty. The penalty is set based on the number of preset actions.
7. The MARL topology control method for promoting grid connection of renewable energy according to claim 1, characterized in that: The advantage function is: , In the formula, Let be the advantage function value of the i-th agent at time t. The summation index represents the number of steps to be calculated backwards from the current time t. As a discount factor, For smoothing parameters, For the i-th intelligent agent The timing difference error at each moment; The i-th intelligent agent The method for calculating the timing difference error at time step is as follows: , In the formula, For the i-th intelligent agent Heterogeneous reward and punishment functions at time points, Let i be the value function of the i-th agent. For the i-th intelligent agent Observation of time.
8. The MARL topology control method for promoting grid connection of renewable energy according to any one of claims 1-7, characterized in that: When optimizing the agent's policy and value network, the optimization objective is to minimize a composite objective function, which includes a truncated agent policy loss, a value function loss, and an entropy reward term.
9. The MARL topology control method for promoting grid connection of renewable energy according to claim 8, characterized in that: The method for calculating the loss of the truncated agent strategy is as follows: , In the formula, The truncated agent policy loss for the i-th agent. Let be the importance sampling rate for the i-th agent at time t. Let be the advantage function value of the i-th agent at time t. For mathematical expectation operators, This is a truncation function. This is the threshold for the cropping range; The method for calculating the value function loss is as follows: , In the formula, Let the loss function be the value function of the i-th agent. For the observation of the i-th agent at time t, Let i be the value function of the i-th agent. Let t be the target reward value for the i-th agent; The entropy reward is calculated as follows: , In the formula, For the entropy reward term of the i-th agent, The entropy coefficient, It is the entropy function. The parameters of the i-th agent are The current policy network at that time.
10. A MARL topology control system for promoting grid connection of renewable energy, characterized in that, include: The problem modeling module is used to construct topology control problems and multi-agent decision-making problems under large-scale power grids; The data acquisition module is used to acquire and preprocess power grid operation data, and to classify and enhance the verification of the power grid operation data according to the reverse power flow index to obtain a dataset. The agent partitioning module is used to construct a joint graph that can capture the complete operating range of the dynamic power grid based on the dataset, and to partition agents based on topological distance; A multi-agent modeling module is used to construct the observation, action space, and heterogeneous reward and punishment functions for each agent. The observations of the agents include the global power grid state, the power flow analysis results of the entire system, and the agent's local information. The multi-agent topology optimization module is used to calculate the advantage function of each agent based on the observation, action space and heterogeneous reward and punishment function of each agent. The optimization of the agent's policy and value network is obtained by optimizing the multi-agent topology through a paradigm of centralized training and distributed execution.