An active power distribution network multi-agent reinforcement learning voltage collaborative control method and device based on a graph attention network and social network analysis, and a medium
Patent Information
- Application Number
- CN202610890065.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]本发明的目的在于提供一种基于图注意力网络与社会网络分析的主动配电网多智能体强化学习电压协同控制方法、设备及介质,用以解决现有分布式电压控制方法在物理邻域约束利用不足、状态相关电压耦合建模不充分、关键智能体贡献难以刻画以及安全区内冗余无功调节较多等问题
[0027]与现有技术相比,本发明的有益效果为提供一种基于图注意力网络与社会网络分析的主动配电网多智能体强化学习电压协同控制方法,有效的降低远端信息干扰和冗余无功调节,提高电压合规性、运行经济性和控制可解释性,具体为:
Smart Images

Figure CN122844175A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of active distribution network voltage / reactive power control technology, specifically to an active distribution network multi-agent reinforcement learning voltage collaborative control method, device, and medium based on graph attention network and social network analysis. Background Technology
[0002] With the rapid integration of distributed photovoltaic (PV), wind power, and flexible loads into the distribution side, the power flow distribution of active distribution networks exhibits stronger uncertainty, time-varying characteristics, and spatial coupling. Affected by random fluctuations in PV output, load changes, and the high impedance ratio structure of distribution lines, the node voltage at the feeder end and in weakly supported areas is prone to exceeding the upper or lower limits, leading to increased reactive power flow, higher network active power losses, and impacting the safe and economical operation of the distribution network.
[0003] Traditional voltage / reactive power control methods typically rely on physical models such as optimal power flow, mixed-integer programming, or sensitivity matrices. These methods require relatively complete network parameters, equipment models, and load / generation forecast information. When large-scale smart inverters are involved in regulation, they also face problems such as high centralized computing burden, strong communication link dependence, and insufficient real-time response.
[0004] In recent years, deep reinforcement learning and multi-agent reinforcement learning have provided new implementation paths for active distributed voltage control in distribution networks. By modeling multiple controlled inverters as cooperative agents, distributed control strategies can be learned within a centralized training and decentralized execution framework. However, existing multi-agent methods still have the following shortcomings: First, if fully connected communication or direct splicing of global states is used, information noise from weakly correlated nodes at remote locations can be easily introduced, leading to a decrease in training efficiency in large-scale networks; Second, the voltage-reactive power coupling relationship between distribution network nodes changes dynamically with the operating state, and fixed adjacency matrices or static edge weights are difficult to accurately represent nonlinear spatial interactions; Third, when all agents are treated with equal weight during training, the contributions of key nodes such as feeder ends or voltage-sensitive areas are easily diluted; Fourth, traditional secondary voltage deviation penalties often require the voltage to remain close to the target voltage. This may induce frequent reactive power fine-tuning even when the voltage is already within a safe range, increasing operational jitter and active power loss.
[0005] Therefore, there is an urgent need for an active distributed voltage control method for distribution networks that can utilize the prior knowledge of the physical topology of the distribution network, adaptively learn the node coupling relationship according to the operating status, and transform the influence of key intelligent agents into training modulation signals, so as to reduce redundant reactive power regulation and network losses while ensuring voltage safety. Summary of the Invention
[0006] The purpose of this invention is to provide an active distribution network multi-agent reinforcement learning voltage collaborative control method, device, and medium based on graph attention network and social network analysis, in order to solve the problems of insufficient utilization of physical neighborhood constraints, insufficient modeling of state-related voltage coupling, difficulty in characterizing the contributions of key agents, and excessive redundant reactive power regulation in the safe zone in existing distributed voltage control methods.
[0007] To address the aforementioned technical problems, this invention proposes a topology-aware multi-agent reinforcement learning-based active distribution network distributed voltage control method based on social network credit modulation. The method includes the following steps: Step 1: Obtain the real-time operating status of the active distribution network, map each controlled inverter as an intelligent agent, use local electrical quantities as local observations of the intelligent agents, and establish a closed-loop voltage control architecture;
[0008] Step 2: Construct a K-nearest neighbor sparse communication graph based on the topological path proximity between the buses where the controlled inverters are located;
[0009] Step 3: Introduce a graph attention network on the K-nearest neighbor sparse communication graph, and adaptively learn the voltage coupling strength related to the state in the neighborhood based on the local observations of the agent to obtain the topology-aware features;
[0010] Step 4: Parse the attention weights of the graph attention network into a directed weighted interaction network, construct a steady-state centrality score based on the in-degree centrality and exponential moving average in social network analysis, and convert the steady-state centrality score into Actor-Critic training modulation weights.
[0011] Step 5: Construct the Deadband-L1 shared reward function, apply voltage-level linear penalties when the voltage approaches or exceeds the statutory safety boundary, and simultaneously consider network active power loss, reactive power output cost and smoothness of operation.
[0012] Step 6: The distributed policy is trained using the MATD3 algorithm, which integrates truncated double Q, Huber loss, attention entropy regularization, and SNA to train modulation weights.
[0013] Step 7: During the execution phase, each agent generates reactive power control actions using local observations and K-nearest neighbor information. These actions are then sent to the corresponding photovoltaic inverters for execution after being constrained and projected by the inverter capability curve.
[0014] Preferably, in step 1, the local observation includes at least the voltage amplitude of the node where the controlled photovoltaic inverter is located, the active power of the node, the reactive power of the node, the instantaneous active power output of the photovoltaic inverter, the rated apparent power of the inverter, and the available reactive power determined by the rated apparent power and the instantaneous active power output.
[0015] Preferably, in step 2, the topological path distance is calculated based on the shortest path length or number of path edges between the buses where the controlled inverters are located along the distribution network topology. Let the intelligent agent... The busbar is Intelligent agents The busbar is The shortest path length or number of path edges between the two in the distribution network topology is calculated as the topological path distance. For each intelligent agent To classify other intelligent agents besides themselves according to Sort the agents in ascending order, select the top K agents as communication neighbors, and retain self-loop connections to obtain a sparse communication edge set.
[0016] Preferably, in step 3, the graph attention network first maps the agent's local observations to latent features through a shared linear transformation, then calculates the original attention score within the neighborhood defined by the K-nearest neighbor sparse communication graph, and obtains the neighborhood attention weights through Softmax normalization.
[0017] Wherein, the neighborhood attention weights are not fully connected and normalized across all agents;
[0018] The graph attention network employs a multi-head attention mechanism, which splices together the neighborhood aggregation results obtained from different attention heads to form the topological perception features of each agent. The topological perception features simultaneously include local operating state information and state-related voltage coupling information within the physically related neighborhood.
[0019] Preferably, in step 4, the in-degree centrality represents the degree to which a certain agent is noticed by other agents when it is a neighbor node, and is obtained by summing the attention weights of all agents that use that agent as an information source; after normalizing the in-degree centrality, the steady-state centrality score is updated using an exponential moving average method.
[0020] The training modulation weights are according to Construction, in which For the first The steady-state centrality score of an agent For the centrality modulation intensity coefficient; when When the weight is 0, each agent has the same training weight. When the value is greater than 0, highly central agents receive higher weights in the Actor-Critic update.
[0021] Preferably, in step 5, the Deadband-L1 reward module of the Deadband-L1 shared reward function is set to a legally defined voltage safety range. Voltage safety corridor is ;
[0022] The voltage penalty term in the Deadband-L1 shared reward function includes: when the node voltage is within a preset voltage safety corridor, no voltage penalty is applied for node voltage deviation; when the node voltage leaves the voltage safety corridor but remains within the legally defined voltage safety range, a deviation penalty is applied according to a first linear weight; when the node voltage exceeds the legally defined voltage safety range, an over-limit penalty is applied according to a second linear weight, and the second linear weight is greater than the first linear weight; when the magnitude of the node voltage exceeding the legally defined voltage safety range exceeds a severe over-limit threshold, a severe over-limit gain factor is introduced to amplify the over-limit penalty.
[0023] Preferably, in step 6, the weighted optimization includes: constructing a target Q value using the truncated double Q mechanism of the dual Critic network, calculating the temporal difference error using Huber loss, introducing the training modulation weights and attention entropy regularization term into the Critic loss, and introducing the training modulation weights into the Actor loss, so that policy learning pays more attention to key agents that have a high informational influence on global voltage regulation.
[0024] Preferably, in step 7, capacity boundary projection is performed on the reactive power control action of each controlled photovoltaic inverter, and the projection range is from... Confirmed, among which For reactive power output, Rated apparent power, It provides instantaneous active power for photovoltaics.
[0025] Another aspect of the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis as described above.
[0026] In another aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis.
[0027] Compared with existing technologies, the beneficial effects of this invention are that it provides an active distribution network multi-agent reinforcement learning voltage collaborative control method based on graph attention network and social network analysis, which effectively reduces remote information interference and redundant reactive power regulation, and improves voltage compliance, operational economy and control interpretability. Specifically:
[0028] 1. By constructing a K-nearest neighbor sparse communication graph through topological path proximity, the information aggregation range of the agent is constrained by the prior physical topology of the power distribution network, which can filter out spatial noise from weakly correlated nodes at the far end, reduce communication redundancy and feature aggregation complexity.
[0029] 2. By using graph attention networks to perform neighborhood aggregation on sparse communication graphs, nonlinear voltage coupling relationships can be adaptively learned based on real-time operating states such as node voltage, power injection, and photovoltaic output, thereby improving the distributed strategy's ability to perceive spatial interactions.
[0030] 3. By mapping attention flow to directed weighted interaction relationships in social networks and constructing steady-state centrality scores using in-degree centrality and exponential moving average, the long-term information influence of key agents can be transformed into interpretable credit modulation weights, thereby alleviating the contribution dilution problem in multi-agent joint training.
[0031] 4. By using the Deadband-L1 shared reward function, the voltage control objective is changed from simple approximation. The shift to low intervention and strong response at risk boundaries within the safety corridor helps reduce redundant reactive actions, reduce action jitter and network active power loss within the safety zone, while maintaining a high voltage compliance rate.
[0032] 5. During the execution phase, there is no need to solve the optimal power flow model online or to explicitly calculate the power flow sensitivity matrix. Each agent can achieve real-time distributed closed-loop control based on local measurements and limited neighborhood communication, which has good engineering scalability. Attached Figure Description
[0033] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this application are illustrated by way of example and not limitation, with the same or corresponding reference numerals denoteing the same or corresponding parts, wherein:
[0034] Figure 1 This is a schematic diagram of the cyber-physical system architecture for distributed voltage control according to the present invention;
[0035] Figure 2 This is a schematic diagram of the overall structure of the MATD3-GS frame of the present invention;
[0036] Figure 3 This is a schematic diagram of the construction of the KNN sparse communication graph with topological proximity constraints according to the present invention;
[0037] Figure 4 This is a schematic diagram of KNN-GAT feature aggregation and SNA centrality credit modulation of the present invention;
[0038] Figure 5This is a schematic diagram of the Deadband-L1 voltage penalty function of the present invention;
[0039] Figure 6 This is a schematic diagram of the closed-loop process of the training and execution phases of the present invention. Detailed Implementation
[0040] The present invention will be further described below with reference to the accompanying drawings and specific embodiments. It should be understood that the following embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention; without departing from the concept of the present invention, those skilled in the art can make adaptive adjustments to the parameter values, network size, training epochs or specific neural network layers.
[0041] This invention relates to the fields of active distribution network voltage / reactive power control, artificial intelligence power system operation optimization, and multi-agent reinforcement learning. The method is applicable to real-time closed-loop voltage control scenarios in active distribution networks where reactive power regulation is collaboratively performed by multiple intelligent inverters under conditions of high-proportion distributed photovoltaic, wind power, and other renewable energy integration.
[0042] Example 1
[0043] like Figure 1 As shown, Figure 1 A schematic diagram of a cyber-physical system architecture for distributed voltage control;
[0044] The active distribution network in this embodiment includes an upstream power grid or relaxed bus, distribution feeders, distributed photovoltaics, conventional loads, controlled photovoltaic inverters, and communication links for information acquisition and transmission. The physical layer is responsible for power flow distribution, voltage state evolution, and reactive power control. The information layer consists of multiple agents corresponding one-to-one with the controlled inverters. The agents generate reactive power control actions based on local measurements and limited neighborhood communication.
[0045] Within a multi-agent reinforcement learning framework, the distributed voltage control of an active distribution network is modeled as a decentralized partially observable Markov decision process (Dec-POMDP). The set of controlled inverters is denoted as: This is mapped to N intelligent agents.
[0046] in, This represents the number of agents corresponding to the controlled inverter. The system state space is denoted as... intelligent agent The local observation space is denoted as At that moment intelligent agent Based on local observations Output reactive power control action The joint action of all agents is denoted as:
[0047]
[0048] The environment evolves to the next state based on the power flow equations of the distribution network, load changes, photovoltaic output fluctuations, and joint actions, and returns to the shared reward. This invention employs a centralized training and decentralized execution structure. During the training phase, the Critic (value network) is allowed to utilize richer joint state and joint action information for value estimation. During the execution phase, each agent relies solely on local observations and KNN neighborhood information to output control actions.
[0049] like Figure 2 As shown, Figure 2 This is a schematic diagram of the overall structure of the MATD3-GS frame;
[0050] The MATD3-GS (Graph-SNA MATD3, a multi-agent dual-delay deep deterministic policy gradient algorithm that integrates graph structure awareness and social network credit modulation) framework in this embodiment mainly includes a topology processor, a GAT encoder, multi-agent actors, dual Critics (dual evaluators), a SNA (Social Network Analysis) analysis engine, a Deadband-L1 (dead-zone-based L1 piecewise linear voltage penalty mechanism) reward module, and an experience replay pool. The topology processor provides physical neighborhood priors, the GAT encoder learns state-related coupling features, the SNA analysis engine generates credit modulation weights based on attention flow, the Deadband-L1 reward module evaluates voltage safety, operational economy, and action smoothness, and the MATD3 optimizer completes the Actor-Critic network update.
[0051] Figure 3 A schematic diagram of constructing a KNN sparse communication graph with topological proximity constraints; Figure 3 The gray lines represent the KNN sparse communication system, used to avoid interference from weak information at distant locations; the double-circle nodes represent controlled inverter agent nodes.
[0052] Figure 4 This is a schematic diagram of feature aggregation and SNA centrality credit modulation in KNN-GAT (Graph Attention Network Based on K Nearest Neighbor Constraints). Figure 4 The upper middle section shows the KNN-GAT neighborhood feature aggregation and action generation path; the lower section shows the SNA centrality credit modulation path based on attention flow. In-degree centrality reflects the degree of attention the agent receives in neighborhood interactions, and steady-state scores are used to modulate Actor-Critic training updates.
[0053] Figure 5This is a schematic diagram of the Deadband-L1 voltage penalty function. The horizontal axis represents the node voltage amplitude, and the vertical axis represents the corresponding penalty value. Within the Deadband voltage safety corridor, node voltage deviations do not incur additional penalties, reducing unnecessary reactive power regulation within the safety zone. When the node voltage deviates from the safety corridor or exceeds the legal voltage boundary, the penalty value increases linearly with the degree of deviation, further intensifying in cases of severe exceedances, thereby achieving a control effect of "low intervention within the safety zone and strong response within the risk zone."
[0054] Figure 6 This is a schematic diagram of the closed-loop process of the training and execution phases of the present invention. Figure 6 In the middle, the upper part is the offline / simulation training closed loop, and the lower part is the online execution closed loop after training is completed; the training objective of the training phase is to integrate KNN-GAT, SNA credit modulation and Deadband-L1 reward to obtain a distributed voltage control strategy; in the execution phase, each agent only relies on local observation and KNN neighborhood information to output reactive power control actions, without relying on online optimal power flow or explicit power flow sensitivity matrix.
[0055] This embodiment provides a multi-agent reinforcement learning-based voltage cooperative control method for active distribution networks based on graph attention networks and social network analysis, including the following steps:
[0056] Step 1: Active Distribution Network Operation Status Acquisition and Agent Modeling: Acquire the active distribution network topology, controlled photovoltaic inverter access nodes, and the real-time operation status of each node. Map each controlled inverter as an agent, use local electrical quantities as local observations of the agent, and establish a closed-loop voltage control architecture consisting of a physical layer and an information layer.
[0057] Specifically, in step 1, the system collects active distribution network operation data at fixed time intervals. For each controlled inverter, the operating state of its bus is mapped to a local observation of the agent, and the local observations of all agents are combined into a joint observation. During the training phase, the global reward and the state at the next time step can be returned from the simulation environment or the historical operating environment; during the execution phase, it is not required that a single agent obtain the complete global state.
[0058] In one embodiment, each agent Local observation It includes at least the voltage amplitude of its bus, the active power of the node, the reactive power of the node, the instantaneous active output of the photovoltaic inverter, the rated apparent power of the inverter, and the available reactive power, which can be expressed as:
[0059]
[0060] in, Represents intelligent agents The busbar at time voltage amplitude, and These represent the active power and reactive power of the nodes, respectively. This indicates the instantaneous active power output of the photovoltaic inverter. This indicates the inverter's rated apparent power. This indicates the currently available reactive power. Due to inverter capacity limitations, the inverter's active and reactive power outputs satisfy the following:
[0061]
[0062] in, Provide reactive power output for photovoltaic inverters;
[0063] Therefore, the inverter at any time The available reactive power range is:
[0064]
[0065] Step 2: Construction of KNN sparse communication graph with topological proximity constraints: Construct a K-nearest neighbor sparse communication graph based on the topological path proximity between the buses where the controlled inverters are located, so that the information interaction of the agents is restricted to the physically related neighborhood.
[0066] Specifically, in step 2, such as Figure 3 As shown, the topological path distance is calculated based on the shortest path length or number of path edges between the buses where the controlled inverters are located along the distribution network topology. Let the intelligent agent... and intelligent agents The busbars located are respectively and The topological path distance between the two is defined as:
[0067]
[0068] in This represents the shortest path length, number of edges, or weighted path distance between two buses along the distribution network topology. For any intelligent agent... To classify other intelligent agents besides themselves according to Sort by size from smallest to largest, and select the one with the shortest topological path distance. Other intelligent agents are considered as communication neighbors. When distances are equal, nodes on the same feeder line are preferred, self-loop connections are preserved, and the intelligent agents are... By adding itself to the neighborhood set, we obtain the KNN neighborhood:
[0069]
[0070] in, Represents intelligent agents The topological distance to other agents Small distance threshold. A sparse communication graph is formed by the KNN neighborhoods of all agents:
[0071]
[0072] Among them, the edge set Defined as:
[0073]
[0074] This KNN graph does not need to be equivalent to an accurate electrical distance map or power flow sensitivity matrix. Its purpose is to provide a lightweight physical inductive bias, reduce the search space of subsequent graph attention mechanisms, and decrease spatial noise and communication redundancy introduced by weakly correlated distant nodes. In an alternative embodiment, a medium-sized system may take... =5, suitable for large-scale systems =10; It can also be adaptively adjusted based on feeder partitions, communication bandwidth, or the number of nodes. If a communication link in a certain area fails, the available neighborhood set for the affected nodes can be recalculated to maintain local control capabilities.
[0075] Step 3, Spatial Feature Aggregation and Distributed Action Generation Based on GAT: A graph attention network is introduced on the K-nearest neighbor sparse communication graph. Based on the agent's local observations, the voltage coupling strength related to the state in the neighborhood is adaptively learned to obtain topology-aware features.
[0076] Specifically, in step 3, such as Figure 4 As shown, the GAT encoder in the KNN sparse communication graph The algorithm aggregates neighborhood features from the local observations of each agent. First, it performs local observations on each agent... Perform a shared linear transformation to obtain the hidden layer features:
[0077]
[0078] in, For learnable weight matrix, For intelligent agents The initial implicit representation. For each edge in the neighborhood. According to the target intelligent agent and Neighbor Intelligent Agent The original attention score is calculated from the hidden layer features:
[0079]
[0080] in, To share attention vectors, This represents a vector concatenation operation. Then, in the agent... KNN neighborhood Perform Softmax normalization internally to obtain neighboring agents. For the target intelligent agent Attention weights:
[0081]
[0082] The above normalization is performed only within the topological neighborhood to prevent fully connected attention from including distant weakly correlated nodes in the aggregation. Furthermore, the GAT encoder can employ a multi-head attention structure. Let the number of attention heads be... , No. The aggregation result of the attention heads is:
[0083]
[0084] in, This represents the normalized attention weight calculated by the m-th attention head. Let be the linear transformation matrix corresponding to the attention head. It is a non-linear activation function. The outputs of multiple attention heads are concatenated to form the agent. Topological sensing features:
[0085]
[0086] Topology-aware features It includes both the agent's own state and state-related coupling information related to voltage regulation within the physically relevant neighborhood. When generating actions, the agent... Actor networks with topology-aware features The input is the reactive power control action of the inverter, and the output corresponds to the reactive power control action of the inverter. The actions generated by all agents constitute a joint reactive power control action, which is applied to the physical layer of the active distribution network.
[0087] The graph attention network employs a multi-head attention mechanism, which splices together the neighborhood aggregation results obtained from different attention heads to form the topological perception features of each agent. The topological perception features simultaneously contain local operating state information and state-related voltage coupling information within the physically related neighborhood.
[0088] Step 4: Centrality credit modulation based on social network analysis: The attention weights of the graph attention network are parsed into a directed weighted interaction network. A steady-state centrality score is constructed based on the in-degree centrality and exponential moving average in social network analysis, and the steady-state centrality score is converted into Actor-Critic training modulation weights.
[0089] Specifically, in step 4, the attention weights of GAT are parsed into a directed weighted interaction network. If a multi-head attention mechanism is used, the attention weights of multiple attention heads can be averaged to obtain the average edge weights for social network analysis. Unlike directly summing the neighborhood attention weights of the same target node, this invention calculates in-degree centrality from the perspective of "being followed by other agents". For any agent... Its in-degree centrality is defined as:
[0090]
[0091] in, Represents intelligent agents When acting as a neighboring node, the target intelligent agent The contribution of attention.
[0092] In-degree centrality represents the degree to which an agent receives attention from other agents when acting as a neighbor node. It is obtained by summing the attention weights of all agents whose information source is that agent. After normalizing the in-degree centrality, an exponential moving average is used to update the steady-state centrality score. If an agent consistently receives high attention from multiple neighboring agents, its in-degree centrality is high, indicating that the agent has significant informational influence in collaborative control. To reduce centrality fluctuations caused by random exploration in the early stages of training, an exponential moving average is applied to the normalized instantaneous centrality to obtain the steady-state centrality score.
[0093]
[0094] in, This represents the normalized instantaneous centrality score. Indicates time The steady-state centrality score, is the exponential moving average smoothing coefficient. Then, the training modulation weights are constructed based on the steady-state centrality score:
[0095]
[0096] in, The central modulation intensity coefficient, For the first The steady-state centrality score of an agent For intelligent agents The training modulated weights, when When the weight is 0, each agent has the same training weight. When the value is greater than 0, highly central agents receive higher weights in the Actor-Critic update. Training modulated weights can be used for both Critic and Actor losses, allowing the training process to focus more on key agents that have a greater impact on global voltage regulation, thereby mitigating the contribution dilution problem in multi-agent joint training.
[0097] Step 5: Construction of Deadband-L1 Shared Reward Function: Construct the Deadband-L1 shared reward function to suppress unnecessary reactive power fine-tuning within the preset voltage safety corridor, apply graded linear penalties when the voltage approaches or exceeds the legal safety boundary, and simultaneously consider network active power loss, reactive power output cost, and smoothness of operation.
[0098] Specifically, in step 5, such as Figure 5 As shown, the Deadband-L1 shared reward function sets the statutory voltage safety range. Preferred And a deadband voltage safety corridor is set up inside it. Preferred When the node voltage is within the Deadband voltage safety corridor, the voltage penalty is zero. When the node voltage leaves the Deadband but has not yet exceeded the legal safety range, the penalty increases linearly with the degree of deviation, and a deviation penalty is applied according to the first linear weight. When the node voltage exceeds the legal safety boundary, a higher-weighted linear penalty is applied, and an over-limit penalty is applied according to the second linear weight, with the second linear weight being greater than the first linear weight. When the over-limit magnitude exceeds the severe over-limit threshold, a gain factor is introduced to further amplify the penalty.
[0099] In a specific form, the Deadband deviation is defined as:
[0100]
[0101] Define the legally defined over-limit deviation as:
[0102]
[0103] node At any moment The Deadband-L1 voltage penalty is defined as follows:
[0104]
[0105] in, The penalty weight for deviation outside the Deadband is... As the legally mandated penalty weight for exceeding limits, and Greater than , This is the severe over-limit gain factor, used to further amplify the penalty when the voltage is severely over-limit.
[0106] The global reward is composed of voltage penalties for all monitored nodes, system active power losses, inverter reactive power output costs, changes in actions at adjacent time points, and global voltage compliance rewards. Taking into account voltage safety, network active power losses, inverter reactive power output costs, action smoothness, and global voltage compliance rewards, a shared reward function is constructed:
[0107]
[0108] in, Indicates global voltage penalty. This indicates the system's active power loss. This indicates that the inverter has no reactive power output. This represents the global voltage compliance indication function. , , , and These represent the weighting coefficients for voltage penalty, active power loss, reactive power output cost, action smoothness, and global compliance reward, respectively. Through the aforementioned Deadband-L1 shared reward function, this invention allows the agent to maintain low intervention when the voltage is within the safe corridor, and provides a stronger risk correction signal when the voltage approaches or exceeds the safe boundary, thereby reducing redundant reactive power actions within the safe zone and balancing voltage safety, action smoothness, and operational economy.
[0109] Step 6, MATD3 training process with SNA weights: The distributed policy is trained by weighted optimization using the MATD3 algorithm which integrates truncated double Q, Huber loss, attention entropy regularization and SNA training modulated weights.
[0110] Specifically, in step 6, such as Figure 6 As shown, a centralized training and decentralized execution structure is adopted during the training phase. Weighted optimization includes: constructing the target Q-value using the truncated double-Q mechanism of the dual-Critic network; calculating the temporal difference error using Huber loss; introducing training modulation weights and attention entropy regularization terms into the Critic loss; and introducing training modulation weights into the Actor loss to make policy learning focus more on key agents with high informational influence on global voltage regulation. Transition samples obtained from environmental sampling are stored in the experience replay pool. When the update condition is met, a mini-batch is sampled from the experience replay pool. The target Actor generates the next state target action, and the dual-target Critics estimate the target Q-values respectively. The smaller of the two values is used to construct a truncated double-Q target to alleviate Q-value overestimation.
[0111] When updating the Critic (evaluator), Huber loss is used to measure the temporal difference error, in order to reduce the impact of large error samples caused by drastic photovoltaic fluctuations, load abrupt changes, or power flow calculation anomalies on training stability. The total Critic loss includes Huber loss weighted by SNA training modulation weights and an attention entropy regularization term; the attention entropy regularization term is used to avoid excessive averaging of GAT weights within the neighborhood.
[0112] During Actor (policy network) updates, a delayed update mechanism is employed, and SNA training modulation weights are introduced into the policy loss to give the policy gradients of high-centrality agents a greater impact in joint policy learning. After the online network update is complete, soft updates are performed on the target Actor and target Critic. Through the above training process, physical neighborhood selection, state-related coupling learning, centrality-credit modulation, and safety-economic reward shaping are unified into end-to-end multi-agent policy learning.
[0113] Step 7, Decentralized Execution and Reactive Power Action Projection: During the execution phase, each agent generates reactive power control actions using only local observations and K-nearest neighbor information. The actions are then projected onto the corresponding photovoltaic inverters for execution after being constrained by the inverter capability curve.
[0114] Specifically, in step 7, during the execution phase, each agent only needs local observations and KNN neighborhood messages to output the original reactive action using the trained Actor network. To ensure physical feasibility, the action is projected onto a feasible range based on the inverter's capability curve before being issued:
[0115]
[0116] in, This represents the boundary trimming function, whose upper and lower boundaries are jointly determined by the inverter's rated apparent power and current active power output. This boundary range is related to the projected range of the capacity boundary. Consistent, This represents the raw reactive action output by the intelligent agent. This represents the executable reactive power command after projection. The projected reactive power command is executed by the corresponding photovoltaic inverter and affects the node voltage through power flow changes, thus forming a closed-loop control process of "physical state perception - neighborhood information interaction - distributed strategy decision-making - reactive power action execution".
[0117] As an example, this method can be trained and validated on IEEE 33-node and IEEE 141-node systems. For the IEEE 33-node system, six controlled agents can be configured and... For the IEEE 141-node system, 22 controlled agents can be configured and selected. The number of GAT attention heads can be set to 4, the hidden dimension can be set to 64, the exponential moving average smoothing coefficient can be set to 0.05, and the Huber loss threshold can be set to 1.0. These parameters are not intended to limit the invention and can be adjusted for different distribution network scales, communication conditions, and inverter arrangements.
[0118] This method includes acquiring the real-time operating status of the active distribution network and establishing a closed-loop voltage control architecture; constructing a K-nearest neighbor sparse communication graph; introducing a graph attention network to obtain topology-aware features; parsing the attention weights of the graph attention network into a directed weighted interaction network, constructing a steady-state centrality score, and converting the steady-state centrality score into Actor-Critic training modulation weights; constructing a Deadband-L1 shared reward function; using the MATD3 algorithm for modulation weights to perform weighted optimization training on the distributed strategy; and in the execution phase, each agent generates reactive power control actions, which are then projected onto the corresponding photovoltaic inverters for execution after being constrained by the inverter capability curve. The technical solution of this invention can solve the technical problems of existing distributed voltage control methods, such as insufficient utilization of physical neighborhood constraints, inadequate modeling of state-dependent voltage coupling, difficulty in characterizing the contributions of key agents, and excessive redundant reactive power regulation within the safe zone.
[0119] As can be seen from the above embodiments, the present invention can improve the safety, economy, smoothness of action and interpretability of distributed voltage control in active distribution networks without relying on online optimal power flow solution and explicit power flow sensitivity matrix. It can extract dynamic coupling features in the physically related neighborhood through KNN-GAT, strengthen the training contribution of key nodes through SNA centrality weight, and suppress redundant reactive power actions in the safe zone through Deadband-L1 reward.
[0120] Implementation effect verification
[0121] This invention has been validated on IEEE 33-node and IEEE 141-node systems and compared with traditional MATD3, MADDPG and ablation models without topological constraints or SNA credit modulation.
[0122] Experimental results show that:
[0123] 1. This invention effectively reduces information interference from weakly correlated distant nodes through a KNN sparse communication mechanism with topological proximity constraints, and can reduce the number of redundant communication connections by about 40% to 60% in the IEEE 141-node system.
[0124] 2. Compared to some benchmark methods, the global voltage compliance rate is significantly improved; compared to MATD3, it is improved by about 8 percentage points in the IEEE 141-node system, and the voltage over-limit rate is significantly reduced.
[0125] 3. By introducing the Deadband-L1 voltage reward mechanism, the fluctuation amplitude of the inverter's reactive power operation is significantly reduced, the number of redundant reactive power adjustments is reduced by about 15% to 30%, and the network active power loss is further reduced.
[0126] 4. By introducing a graph attention flow-based SNA centrality credit modulation mechanism, the contribution of key nodes in training is enhanced, and the algorithm's convergence speed and training stability are superior to the comparative model that does not use credit modulation.
[0127] The above results demonstrate that the present invention can achieve higher voltage safety, operational economy, and distributed control scalability without relying on online optimal power flow solutions and explicit sensitivity matrices.
[0128] Example 2
[0129] This second embodiment provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the topology-aware multi-agent reinforcement learning active distribution network distributed voltage control method based on social network credit modulation as described in the first embodiment.
[0130] Example 3
[0131] This embodiment three provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the topology-sensing multi-agent reinforcement learning active distribution network distributed voltage control method based on social network credit modulation as described in embodiment one.
[0132] Example 4
[0133] This fourth embodiment provides an active distribution network distributed voltage coordination control system, including:
[0134] The data acquisition module is used to acquire real-time operating data of the power distribution network topology and nodes;
[0135] The agent modeling module is used to map the controlled photovoltaic inverter into multiple agents and generate local observations;
[0136] A communication graph construction module is used to construct a KNN sparse communication graph based on topological path distance.
[0137] The GAT encoding and action generation module is used to aggregate neighborhood features on the KNN graph through a graph attention network and output reactive actions.
[0138] The credit modulation module is used to calculate agent centrality based on social network analysis and generate MATD3 training modulation weights.
[0139] The reward module is used to construct the Deadband-L1 shared reward function;
[0140] The training and execution module is used for centralized training of the MATD3 policy network and decentralized execution of reactive power instructions.
[0141] Each module is configured to execute the method as described in Example 1.
[0142] The method provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are merely for the purpose of helping to understand the core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this invention.
Claims
1. A method for active distribution network multi-agent reinforcement learning voltage cooperative control based on graph attention network and social network analysis, characterized in that, The method includes the following steps: Step 1: Obtain the real-time operating status of the active distribution network, map each controlled inverter as an intelligent agent, use local electrical quantities as local observations of the intelligent agents, and establish a closed-loop voltage control architecture; Step 2: Construct a K-nearest neighbor sparse communication graph based on the topological path proximity between the buses where the controlled inverters are located; Step 3: Introduce a graph attention network on the K-nearest neighbor sparse communication graph, and adaptively learn the voltage coupling strength related to the state in the neighborhood based on the local observations of the agent to obtain the topology-aware features; Step 4: Parse the attention weights of the graph attention network into a directed weighted interaction network, construct a steady-state centrality score based on the in-degree centrality and exponential moving average in social network analysis, and convert the steady-state centrality score into the Actor-Critic training modulation weights in the MATD3 algorithm. Step 5: Construct a Deadband-L1 shared reward function that includes Deadband voltage safety corridor segmentation penalty, network active power loss, inverter reactive power output cost, operation smoothing penalty and global voltage compliance reward, and apply voltage graded linear penalty when the voltage approaches or exceeds the statutory safety boundary; Step 6: The distributed policy is trained using the MATD3 algorithm, which integrates truncated double Q, Huber loss, attention entropy regularization, and SNA to train modulation weights. Step 7: After training is completed, the execution phase begins. Each agent generates reactive power control actions using local observations and K-nearest neighbor information. These actions are then projected onto the corresponding photovoltaic inverters after being constrained by the inverter capability curve, thus forming a closed-loop control.
2. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 1, the local observation includes at least the voltage amplitude of the node where the controlled photovoltaic inverter is located, the active power of the node, the reactive power of the node, the instantaneous active power output of the photovoltaic inverter, the rated apparent power of the inverter, and the available reactive power determined by the rated apparent power and the instantaneous active power output.
3. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 2, the topological path distance is calculated based on the shortest path length or number of path edges between the buses where the controlled inverters are located along the distribution network topology. Let the intelligent agent... The busbar is Intelligent agents The busbar is Calculate the shortest path length or number of path edges between the two in the distribution network topology as the topological path distance. For each intelligent agent To classify other intelligent agents besides themselves according to Sort the agents in ascending order, select the top K agents as communication neighbors, and retain self-loop connections to obtain a sparse communication edge set.
4. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 3, the graph attention network first maps the agent's local observations to latent features through a shared linear transformation, then calculates the original attention score within the neighborhood defined by the K-nearest neighbor sparse communication graph, and obtains the neighborhood attention weights through Softmax normalization. Wherein, the neighborhood attention weights are not fully connected and normalized across all agents; The graph attention network employs a multi-head attention mechanism, which splices together the neighborhood aggregation results obtained from different attention heads to form the topological perception features of each agent. The topological perception features simultaneously include local operating state information and state-related voltage coupling information within the physically related neighborhood.
5. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 4, the in-degree centrality represents the degree to which a certain agent is noticed by other agents when it is a neighbor node. It is obtained by summing the attention weights of all agents that use that agent as an information source. After normalizing the in-degree centrality, the steady-state centrality score is updated using an exponential moving average method. The training modulation weights are according to Construction, in which For the first The steady-state centrality score of an agent For the centrality modulation intensity coefficient; when When the weight is 0, each agent has the same training weight. When the value is greater than 0, highly central agents receive higher weights in the Actor-Critic update.
6. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 5, the Deadband-L1 shared reward function is set to a legally defined voltage safety range. Voltage safety corridor is ; The voltage grading linear penalty term includes: when the node voltage is within a preset voltage safety corridor, no voltage penalty is applied for node voltage deviation; when the node voltage leaves the voltage safety corridor but is still within the legal voltage safety range, a deviation penalty is applied according to a first linear weight; when the node voltage exceeds the legal voltage safety range, an over-limit penalty is applied according to a second linear weight, and the second linear weight is greater than the first linear weight; when the magnitude of the node voltage exceeding the legal voltage safety range exceeds a severe over-limit threshold, a severe over-limit gain factor is introduced to amplify the over-limit penalty.
7. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 6, the weighted optimization includes: constructing a target Q value using the truncated double Q mechanism of the dual Critic network, calculating the temporal difference error using Huber loss, introducing the training modulation weights and attention entropy regularization term into the Critic loss, and introducing the training modulation weights into the Actor loss, so that policy learning pays more attention to key agents that have a high informational influence on global voltage regulation.
8. The active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis according to claim 1, characterized in that, In step 7, capacity boundary projection is performed on the reactive power control action of each controlled photovoltaic inverter, and the projection range is from... Confirmed, among which For reactive power output, Rated apparent power, It provides instantaneous active power for photovoltaics.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the active distribution network multi-agent reinforcement learning voltage cooperative control method based on graph attention network and social network analysis as described in any one of claims 1 to 8.