Multi-agent deep reinforcement learning based power distribution network photovoltaic load capacity optimization control method based on graph attention mechanism

By employing a multi-agent deep reinforcement learning method based on graph attention mechanism, the problems of voltage over-limit, reverse power flow, and branch overload in distribution networks with a high proportion of photovoltaic (PV) grids were solved, thereby improving the PV carrying capacity and operational quality of the distribution network.

CN122267710APending Publication Date: 2026-06-23CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA THREE GORGES UNIV
Filing Date
2026-02-02
Publication Date
2026-06-23

Smart Images

  • Figure CN122267710A_ABST
    Figure CN122267710A_ABST
Patent Text Reader

Abstract

A multi-agent deep reinforcement learning method for optimizing the photovoltaic (PV) carrying capacity of distribution networks based on graph attention mechanism includes: acquiring information on the access nodes and access capacity parameters of various controllable devices; establishing a distribution network PV carrying capacity optimization model considering various controllable devices; constructing a MADRL model based on the PV carrying capacity of the distribution network, and defining the state space, action space, reward function, state transition function, and discount factor of the MADRL model's Markov decision process; acquiring the distribution network topology, constructing an adjacency matrix, and constructing a graph attention network based on the state space information and adjacency matrix, embedding deep reinforcement learning (DRL) Actor and Critic networks, and training based on typical daily load and normalized irradiance data; after training, saving the agent weight results to achieve optimal control of the distribution network carrying capacity. This method can explicitly utilize the distribution network graph structure information, enhance the collaborative control capability of multiple devices, and improve the stability and robustness of the strategy in different scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power distribution network dispatching technology, specifically to a multi-agent deep reinforcement learning method for optimizing control of photovoltaic carrying capacity in power distribution networks based on graph attention mechanism. Background Technology

[0002] In recent years, distributed photovoltaic (PV) power has grown rapidly on the distribution network side, significantly changing the power injection method and operating status of the distribution network. With the increase in PV penetration, distribution networks are more prone to problems such as voltage exceeding limits, reverse power flow, line / transformer overload, and increased network losses during operation. This frequently depletes the system's safety margin, resulting in a significant limitation on the available PV capacity. At the same time, PV output and load exhibit significant time-varying and uncertainties, making the carrying capacity boundary highly scenario-dependent. Traditional assessment and control strategies based on a single static operating condition often struggle to balance safety and absorption efficiency.

[0003] In radial distribution networks, the impact of photovoltaic (PV) integration on voltage and power flow is closely related to the integration location, capacity, and line impedance. Multiple integration points form electrical coupling through feeders, and their power injection propagates and superimposes along the network topology, leading to a spatial redistribution of critical constraint nodes and voltage margins. This results in a significant nonlinear and non-additive characteristic of the carrying capacity. Therefore, improving the carrying capacity for high-proportion PV integration requires not only considering local node constraints but also multi-resource coordinated regulation under network topology constraints. This can be achieved through methods such as reactive power regulation using PV inverters, energy storage charging and discharging, and reactive power compensation devices to suppress voltage fluctuations, alleviate power flow congestion, and reduce network losses.

[0004] Existing methods for enhancing the photovoltaic carrying capacity of distribution networks mainly fall into two categories: model-driven methods and data-driven methods. Model-driven methods typically construct optimization models based on power flow equations and operational constraints, and achieve capacity enhancement through reactive power compensation, energy storage configuration, or operational control. While these methods offer strong interpretability, they often rely on precise network parameters and device models. Furthermore, they are prone to forming high-dimensional non-convex optimization problems when multiple time periods, multiple devices, and discrete control variables coexist, resulting in high computational complexity and centralized communication requirements. This makes it difficult to achieve efficient online decision-making under conditions of rapid source-load fluctuations and frequent scenario switching.

[0005] With high-proportion photovoltaic (PV) grid integration, distribution network scheduling is essentially an online sequential decision-making problem. Deep reinforcement learning (DRL) provides a policy learning approach for such problems without explicitly solving complex optimization models. However, in multi-device collaborative control scenarios, it still faces challenges such as multi-agent coupling, high state dimensionality, training instability, and insufficient characterization of topology-related information. In particular, when simply vectorizing the distribution network state and inputting it into the policy network, it is difficult to effectively represent the electrical coupling relationships between nodes and lines and the differences in the importance of key nodes, thus limiting the robustness and transferability of the policy under different topologies and operating scenarios.

[0006] In summary, under the condition of high-penetration distributed photovoltaic access, constraints such as distribution network voltage over-limit, reverse power flow and branch overload are easily triggered. Moreover, source load fluctuations and various controllable resources (photovoltaic inverters, energy storage systems, static var compensators, etc.) have significant time-domain coupling, which makes it difficult for traditional centralized optimization solutions based on mechanism models to balance real-time performance and robustness, and difficult for distributed control to achieve global coordination of multiple devices.

[0007] Therefore, there is an urgent need for a photovoltaic load-bearing capacity optimization control technology that can fully utilize the structural characteristics of the distribution network diagram, be oriented towards multi-device collaboration, and have online adaptive capabilities, so as to improve the photovoltaic absorption level and distribution network operation quality while meeting safety constraints such as voltage and power flow. Summary of the Invention

[0008] This invention proposes a multi-agent deep reinforcement learning (MADRL) method for optimizing the photovoltaic (PV) carrying capacity of distribution networks based on graph attention mechanisms. This method constructs the distribution network operation and multi-device collaborative control process as a Markov decision process, using the distribution network topology as the graph structure input. A graph attention network is introduced to structurally encode node features and electrical-topology coupling relationships. Centralized value assessment and distributed policy execution are achieved within the framework of Multi-Agent Dual-Delay Deterministic Policy Gradient (MATD3), thereby improving the PV carrying capacity of the distribution network while satisfying operational constraints such as voltage and power flow, and taking into account comprehensive operational indicators such as voltage deviation and network losses. Compared with existing technologies, this method can explicitly utilize distribution network graph structure information, enhance multi-device collaborative control capabilities, and improve the stability and robustness of the strategy under different scenarios, making it suitable for improving the PV carrying capacity of distribution networks and online optimization control.

[0009] The technical solution adopted in this invention is as follows:

[0010] A multi-agent deep reinforcement learning-based optimization control method for photovoltaic carrying capacity in power distribution networks, based on graph attention mechanism, includes the following steps: Step 1: Collect typical daily load and solar radiation time series data, and perform irradiance normalization preprocessing on the solar radiation time series data; obtain access node and access capacity parameter information for various controllable devices; Step 2: Establish an optimization model for the photovoltaic carrying capacity of the distribution network, taking into account various controllable equipment; Step 3: Based on the photovoltaic carrying capacity optimization model of the distribution network established in Step 2, construct a MADRL model based on the photovoltaic carrying capacity of the distribution network, and define the state space, action space, reward function, state transition function and discount factor of the Markov decision process of the MADRL. Step 4: Obtain the distribution network topology and construct the adjacency matrix. Based on the state space information and adjacency matrix in Step 3, construct a graph attention network and embed the Actor and Critic networks of deep reinforcement learning (DRL). Train the network based on the typical daily load and normalized irradiance data in Step 1. Step 5: After training is complete, save the agent weight results and conduct experimental tests in the same environment to achieve optimal control of the power distribution network carrying capacity.

[0011] In step 1, the irradiance is normalized using the STC per-unit method, as shown in equation (1): (1); In formula (1): Indicates time t Normalized irradiance; Indicates time t Irradiance collection results; The irradiance results are under standard test conditions. .

[0012] Typical daily load information includes the time of each node in the distribution network. t active load With reactive load .

[0013] In step 1, information on the access nodes and access capacity parameters of various controllable devices is obtained, as shown in Table 1, to construct the power distribution network environment.

[0014]

[0015] In step 2, the objective function in the photovoltaic carrying capacity optimization model of the distribution network is as follows: (2); (3); (4); (5); (6); In the above formula: The objective function is... Indicates the active power of PV; Indicates network active power loss; Indicates the amount of reactive power compensation; Indicates the node voltage deviation; Indicates the number of nodes in the distribution network; This represents the active power provided by distributed photovoltaic power in the distribution network; Indicates the number of branches in the distribution network; This represents the branch resistance between node m and node m-1; and Representing nodes respectively With nodes The active and reactive power of the lines between them; This represents the node voltage at node m-1; and These represent the reactive power compensation and reactive power operation of the static var compensator and the energy storage system, respectively. Represents a node i Node voltage; This is the reference voltage.

[0016] The objective function includes the following constraints: 1) Power flow balance constraints: The power of each node during system operation should satisfy the following formula: (7); In equation (7): Represents a node i The active power provided by the connected distributed photovoltaic system; This represents the reactive power provided by the distributed photovoltaic system connected to node i; Represents a node i The active power consumed by the connected ESS system during charging; Represents a node i The reactive power provided by the connected SVC device; and Representing nodes respectively i The original active load and the original reactive load; Represents a node m The node voltage; Represents a node i With nodes m The voltage phase angle between them; and Represents a node i With nodes m The line conductance and line susceptance between them.

[0017] 2) Node voltage constraints: During operation, each node's voltage must be kept within the constraints: (8); In equation (8): and These represent the minimum and maximum allowable values ​​of the node voltage, respectively. 3) Branch current constraints: During system operation, the current in each branch must be kept within the allowable range: (9); In equation (9): This represents the branch current between node i and node j; and These represent the minimum and maximum allowable values ​​of the branch current, respectively.

[0018] 4) Transformer reverse load rate constraint: The system needs to meet the transformer reverse load rate constraint during operation: (10); In formula (10): This indicates the current reverse load rate of the system; Indicates the rated capacity of the transformer; This represents the reverse load rate. To ensure the normal operation of the transformer under the maximum photovoltaic capacity, the reverse load rate is set to 80% here.

[0019] 5) Constraints of energy storage systems: Active power constraints of energy storage systems: (11); In equation (11): and These represent the active power of the energy storage system at node i during charging and discharging, respectively. and These represent the upper limits of active power for charging and discharging of the energy storage system connected to node i, respectively.

[0020] The state of charge constraints of the energy storage system are as follows: (12); (13); (14); In the above formula: Represents a node i The state of charge of the connected energy storage system ESS during the time period t+1; This indicates the initial state of charge of the energy storage system ESS connected to node i during time period t; and Representing nodes respectively i The charging and discharging switching parameters of the connected energy storage system during time period t are set to 1 when the energy storage system is charging or discharging during time period t, and 0 otherwise. Represents a node i The active power required for charging of the connected ESS system during time period t; and These represent the charging efficiency parameter and the discharging efficiency parameter of the energy storage system, respectively. This represents the rated capacity of the energy storage system connected to node i; and These represent the minimum and maximum states of charge of the energy storage system connected to node i, respectively. Represents a node i The state of charge of the ESS system.

[0021] 6) Reactive power compensation constraints: Static var compensators (SVCs) provide additional reactive power compensation for the reactive power deficiency in a system. The operating characteristics and requirements of SVCs are as follows: (15); (16); (17); In the above formula: This represents the shunt impedance of the static var compensator; and for These represent the internal reactance and capacitance of the static var compensator, respectively. Represents a node i The reactive power provided by the SVC; Represents a node i The rated node voltage; Represents a node i The minimum reactive power provided by the SVC; Represents the node i The maximum reactive power provided by the SVC; This indicates the thyristor firing angle of the static var compensator, ranging from 90° to 180°.

[0022] 7) Capacity and power constraints of photovoltaic inverters: Distributed photovoltaic systems simultaneously input active and reactive power into the system via photovoltaic inverters during operation. Their operating power and capacity must meet the following constraints: (18); (19); In the above formula: Represents a node i The reactive power of distributed photovoltaic power; The rated capacity of the distributed photovoltaic system connected to node i; and Representing nodes respectively i The minimum and maximum active power of the connected distributed photovoltaic system; and Representing nodes respectively i The minimum and maximum reactive power of the connected distributed photovoltaic system.

[0023] In step 3, a Markov decision process for DRL training is constructed, including tuples. ,in, For state space, For action space, For state transition function, For reward function, The discount factors are defined as follows: (1) State space: System status This represents the operating state of the distribution network system within time period t, with the network state at each moment represented by the node characteristic matrix. express, express × A set of dimensional real matrices, where, R Represents the set of real numbers. The number of system nodes. This refers to the node feature dimension. The feature vector contains eight components: node voltage amplitude, node load active and reactive power, node voltage phase angle, equipment type mask, and illumination intensity. The component contents are detailed in the reference [reference needed]. Figure 4 medium state vector part: (20); In equation (20): Let represent the node feature vector of node i during time period t. The meanings of each component are as follows: This represents the negative voltage value of node i during time interval t; This represents the active load power of node i during time period t; This represents the reactive load power of node i during time period t; This represents the voltage phase angle of node i during time interval t; , , These represent the device access masks for the distributed photovoltaic, energy storage system, and static var compensator of node i, respectively. The value is 1 when node i is connected to the corresponding device, and 0 otherwise. This represents the normalized irradiance of node i during time period t; Representing the system Each node at time... Node feature vectors; This indicates the matrix transpose.

[0024] The topology of the distribution network consists of an adjacency matrix. express, express × A set of dimensional real matrices. This adjacency matrix is ​​generated based on path connectivity, with diagonal elements set to 1 to preserve node self-loops, thus each state can be obtained from... This means that the agent extracts the feature relationships between nodes and their neighborhoods through a graph attention network, thereby achieving high-dimensional feature encoding of the state space.

[0025] (2) Action space: Within time interval t, the agent selects a joint action vector based on its current state: (twenty one); In equation (21): Indicates time The joint action vector is obtained by splicing the control actions output by the three device agents in time period t; , , The three components correspond to the adjustment commands of the photovoltaic inverter, the energy storage device, and the static var compensator during the time period t, respectively.

[0026] (3) Reward function: To guide different agents in achieving multi-objective optimization such as voltage stability, loss reduction, and smooth action, a comprehensive reward function is designed: (twenty two); In equation (22): Indicates the agent at time... The comprehensive immediate reward obtained is a weighted sum of multiple sub-rewards or penalties, including voltage constraints, steady-state deviation, smooth operation, and reactive / power factor constraints. The meanings of each sub-item are as follows: This indicates a voltage over-limit penalty item, used to penalize node voltage amplitudes that exceed the allowable range. The greater the degree of over-limit, the greater the penalty. This represents the voltage deviation penalty term, used to measure the degree of deviation of the node voltage from the reference value. The value of each node relative to the reference value is... The greater the deviation, the greater the penalty; This represents the penalty term for amplitude changes, used to suppress excessively rapid changes in action between adjacent time points; the more drastic the change, the greater the penalty.

[0027] This indicates a penalty for frequent reversals in the direction of energy storage system operations. Second-order differential control is used to penalize energy storage system operations; the more frequent the switching between charging and discharging operations of the energy storage system, the greater the penalty.

[0028] This represents the motion jitter penalty term, used to penalize small, high-frequency fluctuations in the control sequence; the more pronounced the jitter, the greater the penalty.

[0029] This represents a reactive power usage penalty, used to limit the reactive power output of the reactive power compensation resources of the PV inverter and SVC. The greater the deviation between the reactive power compensation result and the reactive power load, the greater the penalty.

[0030] This represents a power factor constraint penalty term, used to penalize situations where the power factor falls below a set threshold. The power factor threshold is [value missing]. The worse the power factor, the greater the penalty.

[0031] By combining multiple penalty factors, the agent can optimize voltage distribution while also considering energy balance and device lifespan.

[0032] (4) State transition function: State transitions reflect the dynamic impact of actions on the system. Within each time period t, the environment changes according to the actions. Perform power flow calculations, update network physical quantities such as voltage, current, and losses, and output a new node characteristic matrix. The transition probabilities are expressed as: (twenty three); In equation (23): This represents the distribution network environment state vector for time period t+1. Represents the distribution network environment state vector for time period t; This represents a deterministic mapping from the power flow equations of the distribution network.

[0033] (5) Discount factor: Discount factor Used to measure how much an agent values ​​future rewards. When When the value is close to 0, it indicates that the agent is more focused on the reward at the current moment, and its action adjustment speed is fast, but its ability to adjust across time periods is insufficient. The strategy that approaches 1 hour considers rewards for more time steps, resulting in smoother actions and stronger cross-time-period coordination.

[0034] In step 3, a MADRL model based on the photovoltaic carrying capacity of the distribution network is constructed, which is to model the multi-resource coordinated control process of the distribution network as a Markov decision process. The model can be fully formed by combining the five parts of the Markov process with the network structure given later.

[0035] In step 4, the adjacency matrix is ​​constructed as follows: Distribution network system The state of the directed graph within the time period is ,in, Represents the adjacency matrix. This indicates that there is a branch between nodes i and j. The node feature matrix, Each represents a system Each node's feature vector.

[0036] The graph attention calculation process is as follows: Attention weights of node j to node i: (twenty four); In equation (24): This indicates softmax normalization; Indicates the first l Layer k Under the first attention, the first j The nth adjacent node pair i Attention weight results for each node; Indicates the first l Layer k Under the first attention, the first j The nth adjacent node pair i The attention scores of each node are used to obtain the attention weights after applying softmax normalization to these attention scores. ; Let i represent the set of neighboring nodes of node i.

[0037] Attention score of node j to node i: (25); In equation (25): Indicates the first Layer A learnable parameter vector for each attention head, used to calculate the attention score between adjacent node pairs; Indicates the node With nodes In the Layer The projected feature vectors under each attention head are concatenated for subsequent attention scoring calculation; For nodes In the Layer Projected feature vectors under the size; Represents a node j In the Layer Projected feature vectors under the size; A non-linear function representing attention scoring; Projection process: (26); In equation (26): Indicates the first Layer The linear projection matrix of each attention head; Represents a node In the The input feature vector of the layer; Aggregation process: (27); In equation (27): node In the Layer Weighted aggregated output under each attention head; Indicates nonlinear activation; After combining the attention output, GAT-Actor uses the device-related node embeddings to generate actions; specifically as follows: Due to the uncertainty of source and load in the operation of the distribution network, Gaussian noise is added to the agent's action output, unlike the MATD3 algorithm which relies on the system state. As a strategy basis, the GAT-MATD3 algorithm uses node feature embedding of state during the action generation process. As a basis for strategy updates, that is t During the time period, the first i The actor network of each agent is embedded in the state based on the node features of the current time period. Use local policies Generate deterministic actions : (28); In equation (28): the first The deterministic policy function of an agent It is implemented by an actor network, and the corresponding network parameters are: ; For GAT on the original state The encoded embedding state; Zero-mean Gaussian noise Injecting noise can effectively prevent the algorithm from converging too early and improve the smoothness of the gradient, thus preventing the training results from getting stuck in local optima.

[0038] A batch of empirical tuples is randomly sampled from the empirical replay pool, and the node feature embedding state for the current time period t is obtained using a graph attention network. and the node feature embedding state for the next time period. The target Actor network in Generate actions based on Simultaneously introduce smoothing noise : (29); In equation (29): The target action of the target actor network. For online strategies. The target network noise is not backpropagated in order to construct a stable target action.

[0039] GAT-Critic uses global embedding and action-joint input for centralized evaluation, as detailed below: The GAT-Critic network updates by minimizing the mean squared error between the TD objective and the current Q value. Its loss function is: (30); In equation (30): The loss function for the GAT-Critic network; For the first Each Critic network parameter; this invention employs a dual Critic network, namely... ; Indicates the experience replay pool Calculate the expectation of the transferred samples obtained from the sampling process; express Interaction samples consisting of state over a time period, joint actions, immediate rewards, and the state at the next moment; For parameters The The action value function output by the Critic network; For graph attention encoders The global embedding representation obtained after feature aggregation; Temporal difference target value.

[0040] After several updates to the GAT-Critic network parameters, the GAT-Critic network performs a deterministic policy gradient update. Its core objective is to maximize the value function, i.e., to generate actions that minimize voltage deviation, network loss, and reverse power flow rate. The corresponding gradient expression is as follows: (31); In equation (31): Regarding GAT-actor network parameters The gradient of the policy objective function is used for deterministic gradient policy updates; This represents the objective function of the policy network, i.e., the expected Q value; This indicates the replay pool of experiences. D State samples obtained from sampling s The expectation of the gradient term is calculated; s represents the state vector over time period t; This indicates that the GAT-Critic output is for the first... Actions of an agent The gradient is used to guide the agent's actions to update in the direction of increasing the Q value; The parameter is The action value function output by the Critic network; For the first The policy network corresponding to each agent; This indicates that the policy network output is related to its parameters. The gradient.

[0041] Through backpropagation of the Q-value gradient of the GAT-critic network, the GAT-actor network strategy is adjusted towards the goals of reducing voltage deviation, network loss, and reducing the proportion of reverse power flow.

[0042] In step 4, a graph attention network is constructed and embedded with multi-agent deep reinforcement learning Actor and Critic networks, as detailed below: 4.1: Construction of Graph Attention Networks and Generation of Topology-Aware Vectors Construct an adjacency matrix based on the distribution network topology. And from the node feature matrix ,Will Input the graph attention network, and complete the attention scoring, normalization and weighted aggregation according to equations (24) to (27) to obtain the node embedding matrix: (32); In equation (32): This represents the node embedding matrix within time period t; This represents the topology-aware vector of node i during time period t.

[0043] 4.2: Node Embedding Matrix Embedding GAT-Actor: Let n be the number of device access nodes controlled by each i-th intelligent agent. i Then the local input of the agent is extracted from the corresponding row of the node embedding matrix: (33); In equation (33): This represents the local observation state of the i-th agent in time period t, which serves as the network input of the GAT-actor. This represents the node embedding matrix during time period t.

[0044] GAT-actor network with As input, the output is a continuous control action, generated as follows: (34); In equation (34): This represents the output action vector of the i-th agent during time period t; This represents the actor network parameters of the i-th agent; This represents the exploration noise during time period t.

[0045] 4.3: Node Embedding Matrix Embed GAT-critic: To achieve centralized evaluation, nodes are embedded in a matrix. The global embedding representation is obtained through the pooling operator: (35); In equation (35): This represents the global embedding vector for time interval t, used as the global input to the Critic network. N This indicates the total number of nodes in the distribution network; This represents the attention weight of node i in attention pooling during time period t.

[0046] Value evaluation of a dual-critic network that combines global embedding and joint action input: (36); In equation (36): This represents the policy Q-value output by the Critic network at the final time interval t. This result guides network training and policy convergence. The network employs a dual-Critic structure, and the dual-Critic network outputs two results. and The minimum result is output to prevent excessive iteration and ensure stable training.

[0047] Finally, the GAT-Critic network updates its parameters by minimizing the mean square time difference error shown in equation (30); the GAT-Actor network updates its parameters according to the deterministic policy gradient shown in equation (31).

[0048] In step 5, after training is completed, the agent weight results are saved, as follows: The learnable parameters of the graph attention encoder and the multi-agent Actor-Critic network obtained after training convergence are permanently stored for subsequent direct inference of output control actions under the same topology or similar operating scenarios. The stored weight parameters include: 1) Attention parameters of graph attention network: ; 2) Network parameters for each agent: ; 3) Dual Critic network parameters: , ; By saving the attention weight parameters and network parameters mentioned above, the trained policy network can be directly invoked during the runtime phase to achieve rapid decision-making without the need for repeated training.

[0049] This invention discloses a multi-agent deep reinforcement learning-based method for optimizing the photovoltaic carrying capacity of power distribution networks, with the following technical advantages: 1) This invention explicitly introduces the distribution network topology in the form of an adjacency matrix and uses a graph attention mechanism to weighted aggregate node features. This can adaptively characterize the differences in electrical coupling between nodes and the importance of key constraint nodes, avoiding the loss of topology information caused by simply vectorizing the system state. Thus, it can maintain stable control performance under different access locations, different penetration rates and different operating modes.

[0050] 2) This invention realizes collaborative decision-making of photovoltaic inverters, energy storage systems and reactive power compensation devices under the multi-agent deep reinforcement learning framework. The training phase adopts centralized value assessment and the execution phase adopts distributed strategy output. Under the premise of meeting the operating constraints such as voltage and power flow, it can realize multi-resource complementary regulation, reduce the frequency and amplitude of voltage over-limit, alleviate reverse power flow and branch congestion, thereby improving the photovoltaic carrying capacity of the distribution network and improving operating indicators such as voltage deviation and network loss.

[0051] 3) This invention employs mechanisms such as taking the smaller value of dual Critic and delaying strategy updates to suppress value overestimation and improve training stability; it continuously adapts to changes in load and irradiance distribution and changes in available equipment capacity, reducing dependence on accurate mechanism models and centralized solvers, and meeting the real-time and scalability requirements of online control of distribution networks. Attached Figure Description

[0052] The present invention will be further described below with reference to the accompanying drawings and examples; Figure 1 The results represent typical daily load and irradiance.

[0053] Figure 2 This is a topology diagram of a high-penetration distribution network model for the IEEE-33 node.

[0054] Figure 3 The flowchart shows the attention mechanism.

[0055] Figure 4 Here is the flowchart of the GAT-MATD3 algorithm.

[0056] Figure 5 This is a diagram of the GAT-MATD3 framework.

[0057] Figure 6 The load-bearing capacity results are shown in different scenarios.

[0058] Figure 7 The reactive power results for system Scenario 4 are shown.

[0059] Figure 8 The result is the active power of the system in scenario 4.

[0060] Figure 9 The voltage distribution results at time 12 for different algorithms are shown.

[0061] Figure 10 This is a comparison chart of network losses for different algorithms.

[0062] Figure 11 The training reward values ​​for different algorithms.

[0063] Figure 12 This is a heatmap of GAT attention. Detailed Implementation

[0064] A multi-agent deep reinforcement learning method for optimizing the photovoltaic (PV) carrying capacity of distribution networks, based on graph attention mechanism, addresses issues such as voltage exceedance, reverse power flow, branch overload, and increased network losses caused by high-penetration distributed PV access. This method models the collaborative regulation process of distribution network operation with multiple resources, including PV inverters, energy storage systems, and static var compensators, as a Markov decision process. After collecting typical daily load and irradiance time-series data and performing normalization preprocessing, the method obtains the distribution network topology and constructs an adjacency matrix. Node feature matrices, composed of information such as node voltage, load, phase angle, equipment mask, and irradiance, are input into the graph attention network to achieve weighted aggregation and structured encoding of node features and electrical-topology coupling relationships. Within a multi-agent dual-delay deterministic policy gradient framework, an Actor-Critic network is constructed for centralized value assessment and distributed policy execution. This improves the PV access capacity and reduces voltage deviation and network losses while satisfying operational constraints such as voltage and power flow.

[0065] Example: This simulation uses the pandapower package on the Python platform to model the power distribution network environment, and the tensorflow package to train the DRL network and build the GAT network. The CPU model is AMD Ryzen 9 7945HX with Radeon Graphics.

[0066] Data Description: Irradiance data were collected from photovoltaic power plants in western China, with a granularity of 15 minutes. Data from 2019 and 2020 were used and normalized according to formula (1). Load data were collected from the IEEE 33-node distribution network system. Typical daily load and irradiance data were referenced. Figure 1 .

[0067] As shown below, the multi-agent distribution network capacity optimization scheduling method incorporating graph attention includes the following steps: 1) Distribution network environment modeling: The distribution network used in the simulation is an improved IEEE 33-node distribution network. The system integrates photovoltaic (PV) systems, energy storage systems, and reactive power regulation devices. The connection locations and capacities of various devices are shown in Table 2. The final distribution network system configuration is as follows: Figure 2 .

[0068] To analyze the performance of the proposed method in different scenarios, a control experiment was designed as follows: Control group 1 uses a distributed Critic network, where the Critic of each agent is trained and updated separately based on local information.

[0069] Control group 2 uses a centralized Critic network but does not include a GAT network layer, i.e., the standard MATD3 network structure.

[0070] Control group 3 uses a centralized Critic network combined with a GAT network layer, which is the method of this invention.

[0071] Based on the access devices, the scenarios are divided into the following five types, with scenario 0 being where only PV provides active power. The designs for the other scenarios are as follows: Scenario 1: PV provides both active and reactive power.

[0072] Scenario 2: PV provides both active and reactive power, and reactive power regulation is considered using SVC equipment.

[0073] Scenario 3: PV provides both active and reactive power, and active power regulation is considered for ESS.

[0074] Scenario 4: Simultaneously consider PV, SVC and ESS to regulate the active and reactive power of the system.

[0075]

[0076] 2) DRL parameter settings and GAT network layer settings: This invention employs the GAT-MATD3 algorithm, adding a GAT network layer to the standard MATD3 network structure. Actors use different actor network control output action sets depending on the device type, while Critic employs a centralized Critic network control strategy for optimization. The multi-agent setup method is referenced. Figure 4 It consists of two parts: Actor and Critic.

[0077]

[0078]

[0079]

[0080] 3) Multi-agent training: After completing the setup of the distribution network environment, DRL network, and GAT network layer, the normalized irradiance data from step (1) and the IEEE 33-node distribution network load data are used for training. The algorithm flow is as follows: Figure 3 Training process reference Figure 4 The content includes the environment, intelligent agents, and the replay experience pool.

[0081] 4) Result Evaluation and Verification: Figure 6 The load-bearing capacity comparison experiment of three control experimental groups showed that the control group 3, i.e. the method proposed in this invention, can achieve the best load-bearing capacity results in five simulation experimental scenarios.

[0082] Figure 7 The reactive power dispatching result of the method of the present invention in scenario 4 is that the PV continuously provides reactive power during the day, but due to the capacity constraint of the inverter, the reactive power it can provide first decreases and then increases with the change of irradiance; the SVC and PV work together to compensate for reactive power, so that the total reactive power supply of the system can match the basic reactive load demand and maintain reactive power balance and voltage support capability.

[0083] Figure 8 The active power scheduling result of the method of the present invention in scenario 4 shows that when only PV active power injection is considered, the net load drops to a negative value in some periods and significantly amplifies the peak-valley difference, which poses a risk of reverse power flow. After further introducing ESS active power regulation, the ESS charges and absorbs power when it cannot absorb PV active power, so that the net load turns from negative to positive or close to zero, and achieves peak shaving near the load peak, thereby smoothing out power fluctuations, reducing the risk of reverse power flow and curtailment, and improving photovoltaic absorption and system carrying capacity.

[0084] Figure 9 and Figure 10 The results show that the method of the present invention has the lowest voltage deviation and daytime network loss at time 12 for different algorithms in scenario 4.

[0085] Figure 11 The results show the reward values ​​of different algorithms after 500 training rounds. The results show that, except for the drooping control, the other algorithms basically converge after 300 rounds. In the stable reward performance stage after 450 rounds, the GAT-MATD3 algorithm has the highest reward value per round.

[0086]

[0087] The results in Table 6 show that GAT-MATD3 performs best in terms of load capacity, total network loss, and voltage deviation, indicating that the GAT-MATD3 algorithm improves load capacity while ensuring voltage regulation accuracy.

[0088] Figure 12The results show the attention outcomes of the centralized Critic network and the distributed Actor network after training convergence in Scenario 4. Since the voltage performance has stabilized during training, the attention results in the Critic network are relatively evenly distributed, indicating that the system voltage margin is evenly distributed in Scenario 4, and key constraints are not concentrated on a few nodes. The attention results of the three types of Actor networks show a high correlation with device access. The attention results of access nodes are higher than those of their neighboring nodes. For example, when node 8 accesses a PV device, the attention results are mainly concentrated on node 8 itself. The attention values ​​of corresponding neighboring nodes show a trend of gradually increasing as they approach the device. For example, node 7 is a neighbor of node 8, and the attention results gradually increase from node 6 to node 8. The longitudinal analysis of the heatmap results shows that among nodes 6 to 9, node 8, which accesses a PV device, has the highest average attention result. These results indicate that the attention results of the Actor network are highly correlated with the device access state. Based on this state information, the Actor network can output the action vector most beneficial to its own control of the device.

Claims

1. A multi-agent deep reinforcement learning-based optimization control method for photovoltaic carrying capacity of power distribution networks based on graph attention mechanism, characterized in that... Includes the following steps: Step 1: Collect typical daily load and solar radiation time series data, and perform irradiance normalization preprocessing on the solar radiation time series data; obtain access node and access capacity parameter information for various controllable devices; Step 2: Establish an optimization model for the photovoltaic carrying capacity of the distribution network, taking into account various controllable equipment; Step 3: Based on the photovoltaic carrying capacity optimization model of the distribution network established in Step 2, construct a MADRL model based on the photovoltaic carrying capacity of the distribution network, and define the state space, action space, reward function, state transition function and discount factor of the Markov decision process of the MADRL model. Step 4: Obtain the distribution network topology and construct the adjacency matrix. Based on the state space information and adjacency matrix in Step 3, construct a graph attention network and embed the Actor and Critic networks of deep reinforcement learning (DRL). Train the network based on the typical daily load and normalized irradiance data in Step 1. Step 5: After training is complete, save the agent weight results to achieve optimal control of the power distribution network carrying capacity.

2. The multi-agent deep reinforcement learning method for optimizing and controlling the photovoltaic carrying capacity of a power distribution network based on graph attention mechanism as described in claim 1, characterized in that: In step 2, the objective function in the photovoltaic carrying capacity optimization model of the distribution network is as follows: (2); (3); (4); (5); (6); In the above formula: The objective function is... Indicates the active power of PV; Indicates network active power loss; Indicates the amount of reactive power compensation; Indicates the node voltage deviation; Indicates the number of nodes in the distribution network; This represents the active power provided by distributed photovoltaic power in the distribution network; Indicates the number of branches in the distribution network; This represents the branch resistance between node m and node m-1; and Representing nodes respectively With nodes The active and reactive power of the lines between them; This represents the node voltage at node m-1; and These represent the reactive power compensation and reactive power operation of the static var compensator and the energy storage system, respectively. Represents a node i Node voltage; This is the reference voltage.

3. The multi-agent deep reinforcement learning method for optimizing and controlling the photovoltaic carrying capacity of a power distribution network based on graph attention mechanism as described in claim 2, characterized in that: The objective function includes the following constraints: 1) Power flow balance constraints: The power of each node during system operation should satisfy the following formula: (7); In equation (7): Represents a node i The active power provided by the connected distributed photovoltaic system; This represents the reactive power provided by the distributed photovoltaic system connected to node i; Represents a node i The active power consumed by the connected ESS system during charging; Represents a node i The reactive power provided by the connected SVC device; and Representing nodes respectively i The original active load and the original reactive load; Represents a node m Node voltage; Represents a node i With nodes m The voltage phase angle between them; and Represents a node i With nodes m Line conductance and line susceptance; 2) Node voltage constraints: During operation, each node's voltage must be kept within the constraints. (8); In equation (8): and These represent the minimum and maximum allowable values ​​of the node voltage, respectively. 3) Branch current constraints: During system operation, the current in each branch must be kept within the allowable range: (9); In equation (9): This represents the branch current between node i and node j; and These represent the minimum and maximum allowable values ​​of the branch current, respectively. 4) Transformer reverse load rate constraint: The system needs to meet the transformer reverse load rate constraint during operation: (10); In formula (10): This indicates the current reverse load rate of the system; Indicates the rated capacity of the transformer; Indicates the reverse load rate; 5) Constraints of energy storage systems: Active power constraints of energy storage systems: (11); In equation (11): and These represent the active power of the energy storage system at node i during charging and discharging, respectively. and These represent the upper limits of active power for charging and discharging of the energy storage system connected to node i, respectively. The state of charge constraints of the energy storage system are as follows: (12); (13); (14); In the above formula: Represents a node i The state of charge of the connected energy storage system ESS during the time period t+1; This indicates the initial state of charge of the energy storage system ESS connected to node i during time period t; and Representing nodes respectively i The charging and discharging switching parameters of the connected energy storage system during time period t are set to 1 when the energy storage system is charging or discharging during time period t, and 0 otherwise. Represents a node i The active power required for charging of the connected ESS system during time period t; and These represent the charging efficiency parameter and the discharging efficiency parameter of the energy storage system, respectively. This represents the rated capacity of the energy storage system connected to node i; and These represent the minimum and maximum states of charge of the energy storage system connected to node i, respectively. Represents a node i The state of charge of the ESS system; 6) Reactive power compensation constraint: Static var compensators (SVCs) provide additional reactive power compensation for the reactive power deficiency in a system. The operating characteristics and requirements of SVCs are as follows: (15); (16); (17); In the above formula: This represents the shunt impedance of the static var compensator; With These represent the internal reactance and capacitance of the static var compensator, respectively. Represents a node i The reactive power provided by the SVC; Represents a node i The rated node voltage; Represents a node i The minimum reactive power provided by the SVC; Represents the node i The maximum reactive power provided by the SVC; This indicates the thyristor firing angle of the static var compensator, ranging from 90° to 180°. 7) Capacity and power constraints of photovoltaic inverters: Distributed photovoltaic systems simultaneously input active and reactive power into the system via photovoltaic inverters during operation. Their operating power and capacity must meet the following constraints: (18); (19); In the above formula: Represents a node i The reactive power of distributed photovoltaic power; The rated capacity of the distributed photovoltaic system connected to node i; and Representing nodes respectively i The minimum and maximum active power of the connected distributed photovoltaic system; and Representing nodes respectively i The minimum and maximum reactive power of the connected distributed photovoltaic system.

4. The multi-agent deep reinforcement learning method for optimizing and controlling the photovoltaic carrying capacity of a power distribution network based on graph attention mechanism as described in claim 3, characterized in that: In step 3, a Markov decision process for DRL training is constructed, including tuples. ,in, For state space, For action space, For state transition function, For reward function, The discount factors are defined as follows: (1) State space: System status This represents the operating state of the distribution network system within time period t, with the network state at each moment represented by the node characteristic matrix. express, express × A set of dimensional real matrices, where, R Represents the set of real numbers. The number of system nodes. The node feature dimension; the feature vector contains eight components: node voltage amplitude, node load active and reactive power, node voltage phase angle, equipment type mask, and illumination intensity. The component contents are shown in the state vector in Figure 4. part: (20); In equation (20): Let represent the node feature vector of node i during time period t. The meanings of each component are as follows: This represents the negative voltage value of node i during time interval t; This represents the active load power of node i during time period t; This represents the reactive load power of node i during time period t; This represents the voltage phase angle of node i during time interval t; , , These represent the device access masks for the distributed photovoltaic, energy storage system, and static var compensator of node i, respectively. The value is 1 when node i is connected to the corresponding device, and 0 otherwise. This represents the normalized irradiance of node i during time period t; Representing the system Each node at time... Node feature vectors; Indicates matrix transpose; The topology of the distribution network consists of an adjacency matrix. express, express × A set of dimensional real matrices; this adjacency matrix is ​​generated based on path connectivity, with diagonal elements set to 1 to preserve node self-loops, thus each state can be obtained from... This means that the agent extracts the feature relationships between nodes and their neighborhoods through a graph attention network, thereby achieving high-dimensional feature encoding of the state space; (2) Action space: Within time interval t, the agent selects a joint action vector based on its current state: (21); In equation (21): Indicates time The joint action vector is obtained by splicing the control actions output by the three device agents in time period t; , , The three components correspond to the adjustment commands of the photovoltaic inverter, the energy storage device, and the static var compensator during the time period t, respectively. (3) Reward function: To guide different agents in achieving multi-objective optimization such as voltage stability, loss reduction, and smooth action, a comprehensive reward function is designed: (22); In equation (22): Indicates the agent at time... The comprehensive immediate reward obtained is a weighted sum of multiple sub-rewards or penalties, including voltage constraints, steady-state deviation, smooth operation, and reactive / power factor constraints. The meanings of each sub-item are as follows: This indicates a voltage over-limit penalty item, used to penalize node voltage amplitudes that exceed the allowable range. The greater the degree of over-limit, the greater the penalty. This represents the voltage deviation penalty term, used to measure the degree of deviation of the node voltage from the reference value. The value of each node relative to the reference value is... The greater the deviation, the greater the penalty; This represents a penalty term for amplitude changes, used to suppress excessively rapid changes in action between adjacent time points; the more drastic the change, the greater the penalty. This indicates a penalty for frequent reversals in the direction of energy storage system operations. Second-order differential control is used to penalize energy storage system operations; the more frequent the switching between charging and discharging operations of the energy storage system, the greater the penalty. This represents the motion jitter penalty term, used to penalize small, high-frequency fluctuations in the control sequence; the more pronounced the jitter, the greater the penalty. This represents a reactive power usage penalty, used to limit the reactive power output of the reactive power compensation resources of the PV inverter and SVC. The greater the deviation between the reactive power compensation result and the reactive power load, the greater the penalty. This represents a power factor constraint penalty term, used to penalize situations where the power factor falls below a set threshold. The power factor threshold is [value missing]. The worse the power factor, the greater the penalty. (4) State transition function: State transitions reflect the dynamic impact of actions on the system. Within each time period t, the environment changes according to the actions. Perform power flow calculations, update network physical quantities such as voltage, current, and losses, and output a new node characteristic matrix; the transition probabilities are expressed as: (23); In equation (23): This represents the distribution network environment state vector for time period t+1. Represents the distribution network environment state vector for time period t; This represents a deterministic mapping from the power flow equations of the distribution network; (5) Discount factor: Discount factor Used to measure how much an agent values ​​future returns; when When the value is close to 0, it indicates that the agent is more focused on the reward at the current moment, and its action adjustment speed is fast, but its ability to adjust across time periods is insufficient. The strategy that approaches 1 hour considers rewards for more time steps, resulting in smoother actions and stronger cross-time-period coordination.

5. The multi-agent deep reinforcement learning method for optimizing and controlling the photovoltaic carrying capacity of a power distribution network based on graph attention mechanism according to claim 4, characterized in that: In step 4, the adjacency matrix is ​​constructed as follows: Distribution network system The state of the directed graph within the time period is ,in, Represents the adjacency matrix. This indicates that there is a branch between nodes i and j. The node feature matrix, Each represents a system Each node's feature vector.

6. The multi-agent deep reinforcement learning method for optimizing and controlling the photovoltaic carrying capacity of a power distribution network based on graph attention mechanism as described in claim 5, characterized in that: The graph attention calculation process is as follows: Attention weights of node j to node i: (24); In equation (24): This indicates softmax normalization; Indicates the first l Layer k Under the first attention, the first j The nth adjacent node pair i Attention weight results for each node; Indicates the first l Layer k Under the first attention, the first j The nth adjacent node pair i The attention scores of each node are used to obtain the attention weights after applying softmax normalization to these attention scores. ; Represents the set of neighboring nodes of node i; Attention score of node j to node i: (25); In equation (25): Indicates the first Layer A learnable parameter vector for each attention head, used to calculate the attention score between adjacent node pairs; Indicates that the node With nodes In the Layer The projected feature vectors under each attention head are concatenated for subsequent attention scoring calculation; For nodes In the Layer Projected feature vectors under the size; Represents a node j In the Layer Projected feature vectors under the size; A non-linear function representing attention scoring; Projection process: (26); In equation (26): Indicates the first Layer The linear projection matrix of each attention head; Represents a node In the The input feature vector of the layer; Aggregation process: (27); In equation (27): node In the Layer Weighted aggregated output under each attention head; This indicates nonlinear activation.

7. The multi-agent deep reinforcement learning method for optimizing control of photovoltaic carrying capacity in power distribution networks based on graph attention mechanism as described in claim 6, characterized in that: After combining the attention output, GAT-Actor uses the device-related node embeddings to generate actions; specifically as follows: The GAT-MATD3 algorithm uses node feature embedding of state during the action generation process. As a basis for strategy updates, that is t During the time period, the first i The actor network of each agent is embedded in the state based on the node features of the current time period. Use local policies Generate deterministic actions : (28); In equation (28): the first The deterministic policy function of an agent It is implemented by an actor network, and the corresponding network parameters are: ; For GAT on the original state The encoded embedding state; Zero-mean Gaussian noise ; A batch of empirical tuples is randomly sampled from the empirical replay pool, and the node feature embedding state for the current time period t is obtained using a graph attention network. and the node feature embedding state for the next time period. The target Actor network in Generate actions based on Simultaneously introduce smoothing noise : (29); In equation (29): The target action of the target actor network. For online strategies; The target network noise is not backpropagated in order to construct a stable target action.

8. The multi-agent deep reinforcement learning method for optimizing control of photovoltaic carrying capacity in power distribution networks based on graph attention mechanism according to claim 7, characterized in that: GAT-Critic uses global embedding and action-joint input for centralized evaluation, as detailed below: The GAT-Critic network updates by minimizing the mean squared error between the TD objective and the current Q value. Its loss function is: (30); In equation (30): The loss function for the GAT-Critic network; For the first Each Critic network parameter; this invention employs a dual Critic network, namely... ; Indicates the experience replay pool Calculate the expectation of the transferred samples obtained from the sampling process; express Interaction samples consisting of state over a time period, joint actions, immediate rewards, and the state at the next moment; For parameters The The action value function output by the Critic network; For graph attention encoders The global embedding representation obtained after feature aggregation; Temporal difference target value; After several updates to the GAT-Critic network parameters, the GAT-Critic network performs a deterministic policy gradient update. Its core objective is to maximize the value function, i.e., to generate actions that minimize voltage deviation, network loss, and reverse power flow rate. The corresponding gradient expression is as follows: (31); In equation (31): Regarding GAT-actor network parameters The gradient of the policy objective function is used for deterministic gradient policy updates; This represents the objective function of the policy network, i.e., the expected Q value; This indicates the replay pool of experiences. D State samples obtained by sampling s The expectation of the gradient term is calculated; s represents the state vector over time period t; This indicates that the GAT-Critic output is for the first... Actions of an agent The gradient is used to guide the agent's actions to update in the direction of increasing the Q value; The parameter is The action value function output by the Critic network; For the first The policy network corresponding to each agent; This indicates that the policy network output is related to its parameters. The gradient.

9. The multi-agent deep reinforcement learning method for optimizing control of photovoltaic carrying capacity in power distribution networks based on graph attention mechanism according to claim 8, characterized in that: In step 4, a graph attention network is constructed and embedded with multi-agent deep reinforcement learning Actor and Critic networks, as detailed below: 4.1: Construction of Graph Attention Networks and Generation of Topology-Aware Vectors Construct an adjacency matrix based on the distribution network topology. And from the node feature matrix ,Will Input the graph attention network, and complete the attention scoring, normalization and weighted aggregation according to equations (24) to (27) to obtain the node embedding matrix: (32); In equation (32): This represents the node embedding matrix within time period t; This represents the topology-aware vector of node i during time period t; 4.2: Node Embedding Matrix Embedding GAT-Actor: Let n be the number of device access nodes controlled by each i-th intelligent agent. i Then the local input of the agent is extracted from the corresponding row of the node embedding matrix: (33); In equation (33): This represents the local observation state of the i-th agent in time period t, which serves as the network input of the GAT-actor. This represents the node embedding matrix within time period t; GAT-actor network with As input, the output is a continuous control action, generated as follows: (34); In equation (34): This represents the output action vector of the i-th agent during time period t; This represents the actor network parameters of the i-th agent; This represents the exploration noise over time period t; 4.3: Node Embedding Matrix Embed GAT-critic: To achieve centralized evaluation, nodes are embedded in a matrix. The global embedding representation is obtained through the pooling operator: (35); In equation (35): This represents the global embedding vector at time interval t, used as the global input to the Critic network. N This indicates the total number of nodes in the distribution network; This represents the attention weight of node i in attention pooling during time period t; Value evaluation of a dual-critic network that combines global embedding and joint action input: (36); In equation (36): This represents the policy Q-value output by the Critic network at the final time interval t. This result guides network training and policy convergence. The network employs a dual-Critic structure, and the dual-Critic network outputs two results. and The minimum result is output to prevent excessive iteration and ensure stable training. Finally, the GAT-Critic network updates its parameters by minimizing the mean square time difference error shown in equation (30); the GAT-Actor network updates its parameters according to the deterministic policy gradient shown in equation (31).

10. The multi-agent deep reinforcement learning method for optimizing control of photovoltaic carrying capacity in power distribution networks based on graph attention mechanism according to claim 9, characterized in that: In step 5, the saved weight parameters include: 1) Attention parameters of graph attention network: ; 2) Network parameters for each agent: ; 3) Dual Critic network parameters: , .