A Distribution Network Voltage Regulation Method and System Based on Multi-Agent Policy Gradient

Through distributed multi-agent collaborative optimization, the Actor-Critic network real-time update strategy is used to solve the problem of voltage over limit and network loss in distribution network voltage regulation, and stable and efficient voltage control in complex environments are achieved.

CN120016499BActive Publication Date: 2025-08-05NANJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510458626.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-08-05
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the active distribution network where distributed energy is highly permeable, voltage regulation is difficult to adapt to complex and changeable operating environments. Traditional methods rely on precise physical models, resulting in significant voltage oversight and network loss problems.

Method used

The distributed multi-agent collaborative optimization is adopted, and the power relationship between nodes is described through the DistFlow flow equation, the photovoltaic inverter and energy storage equipment are configured as agents, and the Actor-Critic network is used for real-time policy updates and evaluations to realize intelligent control of photovoltaic inverters and energy storage equipment.

Benefits of technology

It reduces the voltage limit rate and network loss, improves the robustness and adaptability of the algorithm in complex environments, and ensures the stability and optimization effect of voltage regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016499B_ABST
    Figure CN120016499B_ABST
Patent Text Reader

Abstract

The present invention discloses a distribution network voltage control method and system based on multi-agent policy gradients. The method comprises: modeling photovoltaic inverters and energy storage devices to obtain agents, wherein the agents are configured with an actor-critic network architecture; obtaining local observation features of the agents, and inputting the local observation features into an online actor network to obtain the current execution action; evaluating the current execution action through the critic network to obtain the current predicted reward; updating the weight of the online actor network in real time based on the current predicted reward; inputting the current execution action of each agent into a distribution network voltage control model to obtain the system reward and distribution network system status; and updating the weight of the online critic network based on the system reward and distribution network system status. The present invention adopts distributed multi-agent collaborative optimization to reduce voltage over-limit rate and network loss, and improve the robustness and adaptability of the algorithm in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electric energy configuration, and specifically relates to a distribution network voltage control method and system based on multi-agent strategy gradient. Background Art

[0002] With the rapid development of power systems dominated by renewable energy, voltage overshoot and network loss are becoming increasingly prominent issues in active distribution networks with a high penetration of distributed energy resources. Traditional approaches to voltage regulation, driven by a high proportion of distributed photovoltaic integration, rely primarily on on-load voltage regulation by distribution transformers, reactive power regulation by photovoltaic inverters, and the involvement of energy storage devices.

[0003] However, most of these methods rely on accurate physical models of the distribution network. However, in actual applications, the operating environment of the distribution network is complex and changeable, and the parameters change dynamically and frequently. Traditional models are difficult to accurately describe the actual situation, which restricts their practical application effect. Summary of the Invention

[0004] The present invention provides a distribution network voltage control method and system based on multi-agent policy gradient, which adopts distributed multi-agent collaborative optimization to realize intelligent control of photovoltaic inverters and energy storage equipment, reduce voltage over-limit rate and network loss, and improve the robustness and adaptability of the algorithm in complex environments.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] A first aspect of the present invention provides a distribution network voltage control method based on multi-agent policy gradient, comprising:

[0007] The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in a distribution network, and a distribution network voltage control model is established; the nodes in the distribution network are equipped with photovoltaic inverters and / or energy storage devices;

[0008] The photovoltaic inverter and energy storage device are modeled to obtain an intelligent agent, which is configured with an online actor network (also known as an online policy network) and an online critic network (also known as an online evaluation network).

[0009] Obtain the local observation features of each agent and input them into the online Actor network in the corresponding agent to obtain the current execution action; evaluate the current execution action through the Critic network to obtain the current predicted reward; and update the weight of the online Actor network in real time based on the current predicted reward;

[0010] The current execution actions of each agent are input into the distribution network voltage control model to obtain the system reward and distribution network system status; the weight of the online critic network is updated according to the system reward and distribution network system status.

[0011] Furthermore, the DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established, which specifically includes:

[0012] The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and the distribution network voltage control model is established. The expression formula is:

[0013]

[0014]

[0015]

[0016] In the formula, the node and nodes For adjacent nodes, nodes and nodes are adjacent nodes, and Slave nodes Transfer to node Active power and reactive power; and Slave nodes Transfer to node Active power and reactive power; and Node The injected active power and reactive power; and Slave nodes To Node resistance and reactance in circuits; and Node and nodes The voltage amplitude; Represented as a node The set of connected adjacent nodes;

[0017] Add node voltage constraints to the distribution network voltage control model, expressed as:

[0018]

[0019] Where: and Respectively represent the lower and upper limits of the voltage safety range; is the set of all nodes in the power distribution network; For the moment node The voltage amplitude;

[0020] Add the photovoltaic inverter power constraint to the distribution network voltage control model, and the expression formula is:

[0021]

[0022] Where: Indicates at time node Active power of the internal photovoltaic inverter, Indicates at time node Reactive power of the internal photovoltaic inverter, Representative Node The apparent power, is the set of nodes containing photovoltaics in the distribution network; For nodes The maximum active power of the internal photovoltaic inverter;

[0023] Add energy storage state of charge constraints to the distribution network voltage control model, expressed as follows:

[0024]

[0025]

[0026]

[0027] In the formula, is the minimum charge and discharge power of the energy storage device, For The output power of the energy storage device in node i at time, is the maximum charge and discharge power of the energy storage device; is a set of energy storage nodes; and Respectively in Moment and Time Node The capacity of the internal energy storage device; For nodes The charging and discharging efficiency of internal energy storage devices; For nodes Rated capacity of internal energy storage equipment; and are the minimum and maximum capacities of the energy storage device, respectively; is the unit time difference between two adjacent moments.

[0028] Furthermore, the intelligent body is also configured with a target actor network (also called a target policy network) and a target critic network (also called a target evaluation network); the initial weight parameters of the target actor network and the online actor network are the same, and the initial weight parameters of the target critic network and the online critic network are the same;

[0029] The weight of the Critic network is updated according to the system rewards and the distribution network system status, specifically including:

[0030] The next-moment local observation features of each agent are obtained from the distribution network system state. The next-moment local observation features are input into the target Actor network to obtain the agent's next-moment execution action. The next-moment execution action and the next-moment local observation features are input into the target Critic network to obtain the future predicted reward.

[0031] The online critic network is updated based on the future prediction rewards, current prediction rewards, and system rewards accumulated for H times of voltage regulation of the distribution network; the weight parameters of the online actor network are assigned to the target actor network at every set time interval, and the weight parameters of the online critic network are assigned to the target critic network.

[0032] Furthermore, the current actions of each agent are input into the distribution network voltage control model to obtain system rewards, which can be expressed as follows:

[0033]

[0034]

[0035]

[0036] In the formula, For Time Node System rewards; is the power loss in the distribution network; and For Time Node and nodes The voltage exceeds the limit; Representation node Cooperation index; For Time Node Active power of the load; For Time Node Active power output by the internal photovoltaic inverter, For Time Node Active power output by internal energy storage equipment; is the number of nodes in the distribution network; For the moment node The voltage amplitude; is a function that takes its maximum value with 0; and are the upper and lower limits of the voltage amplitude.

[0037] Furthermore, the weights of the online Actor network are updated in real time based on the current predicted rewards, including:

[0038] The training loss function of the online Actor network is constructed based on the current predicted reward, and the expression formula is:

[0039]

[0040]

[0041] In the formula, is the training loss value of the online Actor network; is the number of nodes in the distribution network; For the Critic network based on the node Local observation characteristics and actions The expected reward for the output; Output nodes for online Actor networks actions; is the weight parameter of the online Actor network in node i; is the mapping function representing the online Actor network;

[0042] The parameter iterative gradient of the online Actor network is calculated according to the training loss value of the online Actor network, and the weight parameters of the online Actor network are updated. The expression formula is:

[0043]

[0044] In the formula, Iterate gradients for parameters of online Actor networks; is the learning rate of the online Actor network.

[0045] Furthermore, the online critic network is updated based on the future prediction rewards, current prediction rewards, and system rewards accumulated for H times of distribution network voltage regulation, specifically including:

[0046] Calculate the target reward based on the future predicted reward and the system reward. Calculate the prediction error based on the target reward and the current predicted reward. Sample the prediction error according to the preset noise distribution to obtain the reward noise. Add the reward noise to the system reward to obtain the actual system reward.

[0047] Convert the future predicted rewards and current predicted rewards into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution;

[0048] The training loss value of the online critic network is calculated based on the actual system rewards of the distribution network voltage regulation for H times, the future predicted reward probability distribution, and the current predicted reward probability distribution; the weight parameters of the online critic network are updated according to the training loss value of the online critic network.

[0049] Furthermore, determining the noise distribution of the reward noise specifically includes:

[0050] Obtain historical system rewards, historical work rewards, and historical prediction rewards from the experience pool; calculate historical target rewards based on historical system rewards and historical work rewards, and calculate historical prediction errors using historical target rewards and historical prediction rewards. The expression formula is:

[0051]

[0052] In the formula, Rewards for historical predictions; rewards for historical work; For Time Node Historical system rewards; For Time Node historical forecast errors; For discount reasons;

[0053] Sort the historical system rewards according to the historical prediction error and divide them into M groups. Calculate the variance and amplitude of each group's historical system rewards.

[0054]

[0055]

[0056] In the formula, is the variance of the historical system rewards of the mth group, For nodes Historical system rewards, is the average of the historical system rewards of each node; n is the number of samples in the experience pool; is the maximum variance of historical system rewards for group M; is the magnitude of the historical system reward for the mth group;

[0057] The noise distribution of the reward noise is set according to the variance and amplitude of each group's historical system rewards.

[0058] Furthermore, the training loss value of the online critic network is calculated based on the actual system rewards of the distribution network voltage regulation for H times, the future predicted reward probability distribution, and the current predicted reward probability distribution, specifically including:

[0059] The target reward probability distribution is calculated based on the actual system reward and the future predicted reward probability distribution. The expression formula is:

[0060]

[0061] In the formula, is the target reward probability distribution, Actual system reward for voltage regulation in distribution network; is the expected function; is the discount factor; Predict reward probability distributions for the future;

[0062] The training loss value of the online critic network is calculated based on the cumulative H target reward probability distribution and the current predicted reward probability distribution. The expression formula is:

[0063]

[0064] In the formula, is the target reward probability distribution during the h-th distribution network voltage regulation process; is the current predicted reward probability distribution during the h-th distribution network voltage regulation process; is the cross entropy function, is the training loss value of the online critic network, and h is the voltage control sequence number of the distribution network.

[0065] A second aspect of the present invention provides a distribution network voltage control system based on a multi-agent policy gradient, comprising:

[0066] A model building module uses the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in a distribution network, thereby establishing a voltage control model for the distribution network. The nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices. The photovoltaic inverters and energy storage devices are modeled to obtain an intelligent agent, which is configured with an online actor network and an online critic network.

[0067] The execution module is used to obtain the local observation features of each agent and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action;

[0068] The optimization module is used to evaluate the current execution action through the Critic network to obtain the current predicted reward; update the weight of the online Actor network in real time based on the current predicted reward; input the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and distribution network system status; and update the weight of the online Critic network based on the system reward and distribution network system status.

[0069] The third aspect of the present invention provides an electronic device including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the distribution network voltage control method described in the first aspect of the present invention.

[0070] Compared with the prior art, the present invention has the following beneficial effects:

[0071] The present invention obtains the local observation characteristics of each intelligent agent, inputs the local observation characteristics into the online Actor network in the corresponding intelligent agent to obtain the current execution action, realizes the intelligent control of photovoltaic inverters and energy storage equipment, and regulates the distribution network voltage through multi-agent collaborative regulation, thereby reducing the voltage over-limit rate and network loss.

[0072] The present invention evaluates the currently executed action through a critic network to obtain the current predicted reward; updates the weight of the online actor network in real time based on the current predicted reward; and updates the weight of the online actor network in real time to ensure that the intelligent agent can quickly adjust its strategy to adapt to environmental changes and optimize voltage regulation.

[0073] The present invention inputs the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and the distribution network system status; updates the weight of the online critic network according to the system reward and the distribution network system status, so that the evaluation of the critic network can more accurately estimate the value of the action, maintain the performance stability of the distribution network voltage adjustment, avoid falling into the local optimal solution, and thus guide the actor network to generate more optimal actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 This is a flow chart of a distribution network voltage control method based on multi-agent policy gradient provided in Example 1;

[0075] Figure 2 is a node circuit diagram of the power distribution network provided in Example 1;

[0076] Figure 3This is a voltage limit comparison diagram of various distribution network voltage control methods provided in Example 1;

[0077] Figure 4 This is a network loss comparison diagram of the distribution network voltage control method provided in Example 1;

[0078] Figure 5 This is a graph showing changes in voltage over-limit rate under disturbance conditions provided in Example 1;

[0079] Figure 6 This is a comparison chart of target rewards under different cumulative times provided in Example 1. DETAILED DESCRIPTION

[0080] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and are not intended to limit the scope of protection of the present invention.

[0081] Example 1

[0082] like Figure 1 As shown, this implementation provides a distribution network voltage control method based on multi-agent policy gradient, including:

[0083] The DistFlow power flow equation (also known as the distribution network power flow equation) is used to describe the relationship between active power and reactive power between nodes in the distribution network and to establish a distribution network voltage control model. Specifically, the model includes:

[0084] like Figure 2 The distribution network shown in the figure uses the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in the distribution network. The distribution network voltage control model is established, and the expression formula is:

[0085]

[0086]

[0087]

[0088] In the formula, the node and nodes For adjacent nodes, nodes and nodes are adjacent nodes, and Slave nodes Transfer to node Active power and reactive power; and Slave nodes Transfer to node Active power and reactive power; and Node The injected active power and reactive power; and Slave nodes To Node resistance and reactance in circuits; and Node and nodes The voltage amplitude; Represented as a node The set of connected adjacent nodes;

[0089] Add node voltage constraints to the distribution network voltage control model, expressed as:

[0090]

[0091] Where: and Respectively represent the lower and upper limits of the voltage safety range; is the set of all nodes in the power distribution network; For the moment node The voltage amplitude;

[0092] Add the photovoltaic inverter power constraint to the distribution network voltage control model, and the expression formula is:

[0093]

[0094] Where: Indicates at time node Active power of the internal photovoltaic inverter, Indicates at time node Reactive power of the internal photovoltaic inverter, Representative Node The apparent power, is the set of nodes containing photovoltaics in the distribution network; For nodes The maximum active power of the internal photovoltaic inverter;

[0095] Add energy storage state of charge constraints to the distribution network voltage control model, expressed as follows:

[0096]

[0097]

[0098]

[0099] In the formula, is the minimum charge and discharge power of the energy storage device, For The output power of the energy storage device in node i at time, is the maximum charge and discharge power of the energy storage device; is a set of energy storage nodes; and Respectively in Moment and Time Node The capacity of the internal energy storage device; For nodes The charging and discharging efficiency of internal energy storage devices; For nodes Rated capacity of internal energy storage equipment; and are the minimum and maximum capacities of the energy storage device, respectively; is the unit time difference between two adjacent moments.

[0100] Modeling the distribution network voltage control as a multi-agent Markov game model can more efficiently handle complex constraints and achieve multi-agent collaborative optimization to improve voltage control performance. In the distribution network, nodes are equipped with photovoltaic inverters and / or energy storage devices. Modeling photovoltaic inverters and energy storage devices to obtain agents, such as Figure 1 As shown, the intelligent body is configured with an online Actor network (online policy network) and an online Critic network (online evaluation network) as well as a target Actor network (target policy network) and a target Critic network (target evaluation network); the initial weight parameters of the target Actor network and the online Actor network are the same, and the initial weight parameters of the target Critic network and the online Critic network are the same.

[0101] Obtain the local observation features of each agent, input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action. The action space of the photovoltaic inverter is , the action space of the energy storage device is ; For The active power of the energy storage device in node i at time, For The reactive power of the energy storage device in node i at the moment; the current predicted reward is obtained by evaluating the current execution action through the Critic network.

[0102] The weight of the online Actor network is updated in real time based on the current predicted rewards, including:

[0103] The training loss function of the online Actor network is constructed based on the current predicted reward, and the expression formula is:

[0104]

[0105]

[0106] In the formula, is the training loss value of the online Actor network; is the number of nodes in the distribution network; For the Critic network based on the node Local observation characteristics and actions The expected reward for the output; Output nodes for online Actor networks actions; is the weight parameter of the online Actor network in node i; is the mapping function representing the online Actor network;

[0107] The parameter iteration gradient of the online Actor network is calculated according to the training loss value of the online Actor network, and the weight parameters of the online Actor network are updated. The expression formula is:

[0108]

[0109] In the formula, Iterate gradients for parameters of online Actor networks; is the learning rate of the online Actor network.

[0110] The current actions of each agent are input into the distribution network voltage control model to obtain the system reward, which is expressed as:

[0111]

[0112]

[0113]

[0114] In the formula, For Time Node System rewards; is the power loss in the distribution network; and For Time Node and nodes The voltage exceeds the limit; Representation node Cooperation index; For Time Node Active power of the load; For Time Node Active power output by the internal photovoltaic inverter, For Time Node Active power output by internal energy storage equipment; is the number of nodes in the distribution network; For the moment node The voltage amplitude; is a function that takes its maximum value with 0; and are the upper and lower limits of the voltage amplitude.

[0115] Input the current execution action of each agent into the distribution network voltage control model to obtain the distribution network system state; obtain the local observation features of each agent at the next moment from the distribution network system state, input the local observation features at the next moment into the target Actor network to obtain the agent's next execution action, and input the next moment execution action and the next moment local observation features into the target Critic network to obtain the future prediction reward;

[0116] Setting the noise distribution of reward noise includes:

[0117] Obtain historical system rewards, historical work rewards, and historical prediction rewards from the experience pool; calculate historical target rewards based on historical system rewards and historical work rewards, and calculate historical prediction errors using historical target rewards and historical prediction rewards. The expression formula is:

[0118]

[0119] In the formula, Rewards for historical predictions; rewards for historical work; For Time Node Historical system rewards; For Time Node historical forecast errors; For discount reasons;

[0120] Sort the historical system rewards according to the historical prediction error and divide them into M groups. Calculate the variance and amplitude of each group's historical system rewards.

[0121]

[0122]

[0123] In the formula, is the variance of the historical system rewards of the mth group, For nodes Historical system rewards, is the average of the historical system rewards of each node; n is the number of samples in the experience pool; is the maximum variance of historical system rewards for group M; is the magnitude of the historical system reward for the mth group;

[0124] The noise distribution of the reward noise is set according to the variance and amplitude of each group's historical system rewards.

[0125] The target reward is calculated based on the future predicted reward and the system reward. The prediction error is calculated from the target reward and the current predicted reward. The system reward is grouped based on the prediction error. The reward noise is obtained by sampling according to the preset noise distribution of each group. The reward noise is added to the system reward to obtain the actual system reward. By dynamically adjusting the noise, the exploration ability is enhanced, the local optimum is avoided, and the stability and adaptability of the strategy are improved.

[0126] Convert the future predicted rewards and current predicted rewards into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution, specifically including:

[0127] Using weight vectors To define the future prediction reward probability distribution and the current prediction reward probability distribution, use Represents the possible value of future prediction reward or current prediction reward; For future prediction rewards or current prediction rewards equal to probability;

[0128]

[0129]

[0130] In the formula, and Indicates the maximum and minimum values of future prediction rewards or current prediction rewards; Represents the partition length of discrete values.

[0131] The training loss value of the online critic network is calculated based on the actual system reward of the distribution network voltage regulation H times, the future predicted reward probability distribution, and the current predicted reward probability distribution. Specifically, it includes:

[0132] The target reward probability distribution is calculated based on the actual system reward and the future predicted reward probability distribution. The expression formula is:

[0133]

[0134] In the formula, is the target reward probability distribution, Actual system reward for voltage regulation in distribution network; is the expected function; is the discount factor; Predict reward probability distributions for the future;

[0135] The training loss value of the online critic network is calculated based on the cumulative H target reward probability distribution and the current predicted reward probability distribution. The expression formula is:

[0136]

[0137] In the formula, is the target reward probability distribution during the h-th distribution network voltage regulation process; is the current predicted reward probability distribution during the h-th distribution network voltage regulation process; is the cross entropy function, is the training loss value of the online critic network, and h is the voltage control sequence number of the distribution network.

[0138] The weight parameters of the online critic network are updated according to the training loss value of the online critic network. At set intervals, the weight parameters of the online actor network are assigned to the target actor network, and the weight parameters of the online critic network are assigned to the target critic network.

[0139] In order to verify the effect of this embodiment (abbreviated as A-MADPG algorithm), the adaptive multi-agent deep deterministic policy gradient algorithm (abbreviated as MADDPG algorithm) and the multi-agent soft actor-critic algorithm (MASAC algorithm) are selected for comparative analysis. Figure 3 As shown in the figure, in the early stages of training, the voltage overshoot rate of the distribution network is high. However, as training progresses, the A-MADPG algorithm is able to quickly reduce the voltage overshoot rate from an initial 0.3 to below 0.05, and stabilize it within 0.03 in the later stages. In comparison, the traditional MASAC and MADDPG methods have voltage overshoot rates of approximately 0.05 and 0.06, respectively, under the same conditions.

[0140] like Figure 4As shown, the A-MADPG algorithm rapidly reduces the total network loss in the early stages of training and quickly stabilizes at a low level. In contrast, the MASAC algorithm also rapidly reduces the total network loss in the initial stages, but its rate of reduction is slower than that of A-MADPG. This is because the MASAC strategy's high randomness and requirement for maximizing entropy generally require more time to converge to the optimal solution during training. In the later stages of training, the total network loss is slightly lower than that of A-MADPG, but the difference is very small. The MADDPG algorithm exhibits significant fluctuations throughout training, with the total network loss remaining high and failing to significantly decrease. It also becomes unstable in the later stages (between approximately 0.6 and 0.8). This indicates that the MADDPG algorithm is less effective in optimizing voltage control strategies and reducing network losses, and is easily affected by environmental changes, resulting in poor convergence.

[0141] like Figure 5 As shown in the figure, the A-MADPG algorithm exhibited the lowest voltage overshoot rate throughout the training process and the fastest decline rate. Initially, the voltage overshoot rate was high, but it rapidly declined in the early stages of training and remained near zero in the later stages. This demonstrates that the A-MADPG algorithm can effectively cope with voltage fluctuations under disturbance conditions and maintain voltage stability in the distribution network. The MASAC algorithm also exhibited a high voltage overshoot rate in the initial stages, but its decline was slower than that of the A-MADPG algorithm. Although the voltage overshoot rate decreased and stabilized in the later stages of training, it remained slightly higher than that of the A-MADPG algorithm. In contrast, the MADDPG algorithm exhibited significant voltage overshoot rate fluctuations throughout the training process and remained consistently high, demonstrating significant instability. This result demonstrates that the MADDPG algorithm is unable to effectively control voltage fluctuations under disturbance conditions, resulting in a persistently high voltage overshoot rate, reflecting its limited adaptability in complex environments. This significant difference demonstrates that the A-MADPG algorithm exhibits greater robustness and adaptability in dealing with uncertain disturbances.

[0142] like Figure 6 As shown in the figure, when H = 1, the algorithm converges slowly, and the cumulative reward gradually stabilizes after approximately 4000 training rounds. When H = 3, the convergence speed is significantly accelerated, and the number of training rounds is reduced to 2000. When H = 5, the final reward value is slightly higher than that of H = 1. This shows that moderately increasing H can effectively improve learning efficiency and policy performance. However, experimental results also show that a larger H value is not necessarily better. In additional experiments, when H is further increased to 20, the convergence curve of the cumulative reward value shows significant fluctuations, and the final reward value decreases compared to H = 5. When H is too large, the time span of the reward sequence may cover multiple policy adjustment stages. This makes the long-term rewards inaccurately reflect the effectiveness of the current policy, resulting in the agent being unable to effectively feedback short-term behavior during decision updates.

[0143] To enhance the robustness of the algorithm in uncertain environments, this paper introduces an adaptive noise mechanism into the system reward to account for random and uncertain environmental changes. Specifically, by dynamically adding symmetric noise to the system reward, the agent's exploration capability in low-variance regions is enhanced while preventing premature entrapment in local optima. This mechanism enables the agent to maintain a high level of strategic exploration, facilitating the discovery of potential global optimal strategies. The noise amplitude is dynamically adjusted based on the variance of the region, with lower noise levels in high-variance regions and higher noise levels in low-variance regions. This effectively balances exploration and exploitation, improving the stability and adaptability of the control strategy.

[0144] This embodiment uses a discrete probability distribution to represent all possible reward values for state-action pairs. Instead of relying solely on a single numerical value to evaluate Q-values, the algorithm employs a probability distribution to reflect the diverse possibilities in complex environments. Furthermore, the critic network uses cross-entropy as a loss function, further improving the accuracy and robustness of the value function estimation. This design enables the algorithm to better adapt to the complex real-world power system environment and effectively handle high-dimensional and highly variable scenarios.

[0145] This implementation combines immediate rewards with the cumulative rewards over the next H steps, providing more comprehensive information for estimating the value function. This not only improves the algorithm's ability to assess long-term benefits but also significantly accelerates model convergence. During training, by properly selecting the value of H, a balance can be achieved between short-term behavioral feedback and long-term planning, thereby improving decision-making quality in complex environments.

[0146] Example 2

[0147] This embodiment provides a distribution network voltage control system based on multi-agent policy gradient. The system described in this embodiment can be applied to the method described in Example 1. The distribution network voltage control system includes:

[0148] A model building module uses the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in a distribution network, thereby establishing a voltage control model for the distribution network. The nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices. The photovoltaic inverters and energy storage devices are modeled to obtain an intelligent agent, which is configured with an online actor network and an online critic network.

[0149] The execution module is used to obtain the local observation features of each agent and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action;

[0150] The optimization module is used to evaluate the current execution action through the Critic network to obtain the current predicted reward; update the weight of the online Actor network in real time based on the current predicted reward; input the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and distribution network system status; and update the weight of the online Critic network based on the system reward and distribution network system status.

[0151] The intelligent body is also configured with a target actor network and a target critic network; the initial weight parameters of the target actor network and the online actor network are the same, and the initial weight parameters of the target critic network and the online critic network are the same;

[0152] The optimization module updates the weight of the critic network according to the system reward and the distribution network system status, specifically including:

[0153] The next-moment local observation features of each agent are obtained from the distribution network system state. The next-moment local observation features are input into the target Actor network to obtain the agent's next-moment execution action. The next-moment execution action and the next-moment local observation features are input into the target Critic network to obtain the future predicted reward.

[0154] Calculate the target reward based on the future predicted reward and the system reward. Calculate the prediction error based on the target reward and the current predicted reward. Sample the prediction error according to the preset noise distribution to obtain the reward noise. Add the reward noise to the system reward to obtain the actual system reward.

[0155] Convert the future predicted rewards and current predicted rewards into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution;

[0156] Calculate the training loss value of the online critic network based on the actual system rewards of H cumulative distribution network voltage regulation, the future predicted reward probability distribution, and the current predicted reward probability distribution; update the weight parameters of the online critic network based on the training loss value of the online critic network;

[0157] At set intervals, the weight parameters of the online Actor network are assigned to the target Actor network, and the weight parameters of the online Critic network are assigned to the target Critic network.

[0158] Example 3

[0159] This embodiment provides an electronic device including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the distribution network voltage control method described in Example 1.

[0160] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0161] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0162] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0164] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A distribution network voltage control method based on multi-agent policy gradient, characterized in that: include: The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in a distribution network, and a distribution network voltage control model is established; the nodes in the distribution network are equipped with photovoltaic inverters and / or energy storage devices; Modeling the photovoltaic inverter and energy storage device to obtain an intelligent agent, which is configured with an online actor network and an online critic network; Obtain the local observation features of each agent and input them into the online Actor network in the corresponding agent to obtain the current execution action; evaluate the current execution action through the online Critic network to obtain the current predicted reward; and update the weight of the online Actor network in real time based on the current predicted reward; The current execution action of each intelligent agent is input into the distribution network voltage control model to obtain the system reward and distribution network system status; the intelligent agent is also configured with a target actor network and a target critic network; the initial weight parameters of the target actor network and the online actor network are the same, and the initial weight parameters of the target critic network and the online critic network are the same; The weights of the online critic network and the target critic network are updated according to the system reward and the distribution network system status, specifically including: The next-moment local observation features of each agent are obtained from the distribution network system state, and the next-moment local observation features are input into the target Actor network to obtain the agent's next-moment execution action. The next-moment execution action and the next-moment local observation features are input into the target Critic network to obtain the future predicted reward. Calculate the target reward based on the future predicted reward and the system reward. Calculate the prediction error based on the target reward and the current predicted reward. Sample the prediction error according to the preset noise distribution to obtain the reward noise. Add the reward noise to the system reward to obtain the actual system reward. Convert the future predicted rewards and current predicted rewards into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution; The training loss value of the online critic network is calculated based on the actual system rewards of the distribution network voltage regulation H times, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards; the weight parameters of the online critic network are updated based on the training loss value of the online critic network; At set intervals, the weight parameters of the online Actor network are assigned to the target Actor network, and the weight parameters of the online Critic network are assigned to the target Critic network.

2. The distribution network voltage control method according to claim 1, characterized in that: The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established, which specifically includes: The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and the distribution network voltage control model is established. The expression formula is: ; ; ; In the formula, the node and nodes For adjacent nodes, nodes and nodes are adjacent nodes, and Slave nodes Transfer to node Active power and reactive power; and Slave nodes Transfer to node Active power and reactive power; and Node The injected active power and reactive power; and Slave nodes To Node resistance and reactance in circuits; and Node and nodes The voltage amplitude; Represented as a node The set of connected adjacent nodes; Add node voltage constraints to the distribution network voltage control model, expressed as: ; Where: and Respectively represent the lower and upper limits of the voltage safety range; is the set of all nodes in the power distribution network; For the moment node The voltage amplitude; Add the photovoltaic inverter power constraint to the distribution network voltage control model, and the expression formula is: ; Where: Indicates at time node Active power of the internal photovoltaic inverter, Indicates at time node Reactive power of the internal photovoltaic inverter, Representative Node The apparent power, is the set of nodes containing photovoltaics in the distribution network; For nodes The maximum active power of the internal photovoltaic inverter; Add energy storage state of charge constraints to the distribution network voltage control model, expressed as follows: ; ; ; In the formula, is the minimum charge and discharge power of the energy storage device, For The output power of the energy storage device in node i at time, is the maximum charge and discharge power of the energy storage device; is a set of energy storage nodes; and Respectively in Moment and Time Node The capacity of the internal energy storage device; For nodes The charging and discharging efficiency of internal energy storage devices; For nodes Rated capacity of internal energy storage equipment; and are the minimum and maximum capacities of the energy storage device, respectively; is the unit time difference between two adjacent moments.

3. The distribution network voltage control method according to claim 1, characterized in that: The current actions of each agent are input into the distribution network voltage control model to obtain the system reward, which is expressed as: ; ; ; In the formula, For Time Node System rewards; is the power loss in the distribution network; and For Time Node and nodes The voltage exceeds the limit; Representation node Cooperation index; For Time Node Active power of the load; For Time Node Active power output by the internal photovoltaic inverter, For Time Node Active power output by internal energy storage equipment; is the number of nodes in the distribution network; For the moment node The voltage amplitude; is a function that takes its maximum value with 0; and are the upper and lower limits of the voltage amplitude; is the set of nodes in the distribution network.

4. The method for controlling voltage in a distribution network according to claim 1, wherein: The weight of the online Actor network is updated in real time based on the current predicted rewards, including: The training loss function of the online Actor network is constructed based on the current predicted reward, and the expression formula is: ; ; In the formula, is the training loss value of the online Actor network; is the number of nodes in the distribution network; For the Critic network based on the node Local observation characteristics and perform actions The expected reward for the output; Output nodes for online Actor networks Execution action; For nodes Local observation characteristics of is the weight parameter of the online Actor network in node i; is the mapping function representing the online Actor network; is the set of nodes in the distribution network; The parameter iteration gradient of the online Actor network is calculated according to the training loss value of the online Actor network, and the weight parameters of the online Actor network are updated. The expression formula is: ; In the formula, For nodes Parameter iterative gradient of the intrinsic online Actor network; is the learning rate of the online Actor network.

5. The method for controlling voltage in a distribution network according to claim 1, wherein: Determining the noise distribution of the reward noise specifically includes: Obtain historical system rewards, historical work rewards, and historical prediction rewards from the experience pool; calculate historical target rewards based on historical system rewards and historical work rewards, and calculate historical prediction errors using historical target rewards and historical prediction rewards. The expression formula is: ; In the formula, Rewards for historical predictions; rewards for historical work; For Time Node Historical system rewards; For Time Node historical forecast errors; is the discount factor; Sort the historical system rewards according to the historical prediction error and divide them into M groups. Calculate the variance and amplitude of each group's historical system rewards. ; ; In the formula, is the variance of the historical system rewards of the mth group, For nodes Historical system rewards, is the average of the historical system rewards of each node; n is the number of samples in the experience pool; is the maximum variance of historical system rewards for group M; is the magnitude of the historical system reward for the mth group; The noise distribution of the reward noise is set according to the variance and amplitude of each group's historical system rewards.

6. The method for controlling voltage in a distribution network according to claim 1, wherein: The training loss value of the online critic network is calculated based on the actual system reward of the distribution network voltage regulation H times, the future predicted reward probability distribution, and the current predicted reward probability distribution. Specifically, it includes: The target reward probability distribution is calculated based on the actual system reward and the future predicted reward probability distribution. The expression formula is: ; In the formula, is the target reward probability distribution, Actual system reward for voltage regulation in distribution network; is the expected function; is the discount factor; Predict reward probability distributions for the future; Output nodes for online Actor networks Execution action; For nodes Local observation characteristics of The training loss value of the online critic network is calculated based on the cumulative H target reward probability distribution and the current predicted reward probability distribution. The expression formula is: ; In the formula, is the target reward probability distribution during the h-th distribution network voltage regulation process; is the current predicted reward probability distribution during the h-th distribution network voltage regulation process; is the cross entropy function, is the training loss value of the online critic network, and h is the voltage control sequence number of the distribution network.

7. A distribution network voltage control system based on multi-agent policy gradient, characterized in that: include: A model building module uses the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in a distribution network, thereby establishing a voltage control model for the distribution network. The nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices. The photovoltaic inverters and energy storage devices are modeled to obtain an intelligent agent, which is configured with an online actor network and an online critic network. The execution module is used to obtain the local observation features of each agent and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action; The optimization module is used to evaluate the current execution action through the online critic network to obtain the current predicted reward; update the weight of the online actor network in real time based on the current predicted reward; and input the current execution action of each agent into the distribution network voltage control model to obtain the system reward and distribution network system status; The intelligent body is also configured with a target actor network and a target critic network; the initial weight parameters of the target actor network and the online actor network are the same, and the initial weight parameters of the target critic network and the online critic network are the same; The optimization module updates the weights of the online critic network and the target critic network according to the system reward and the distribution network system status, specifically including: The next-moment local observation features of each agent are obtained from the distribution network system state, and the next-moment local observation features are input into the target Actor network to obtain the agent's next-moment execution action. The next-moment execution action and the next-moment local observation features are input into the target Critic network to obtain the future predicted reward. Calculate the target reward based on the future predicted reward and the system reward. Calculate the prediction error based on the target reward and the current predicted reward. Sample the prediction error according to the preset noise distribution to obtain the reward noise. Add the reward noise to the system reward to obtain the actual system reward. Convert the future predicted rewards and current predicted rewards into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution; The training loss value of the online critic network is calculated based on the actual system rewards of the distribution network voltage regulation H times, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards; the weight parameters of the online critic network are updated based on the training loss value of the online critic network; At set intervals, the weight parameters of the online Actor network are assigned to the target Actor network, and the weight parameters of the online Critic network are assigned to the target Critic network.

8. An electronic device comprising a storage medium and a processor; the storage medium is used to store instructions; and characterized in that: The processor is configured to operate according to the instruction to execute the distribution network voltage control method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Power distribution network voltage autonomous optimization control method and device

    CN113872213A