Power distribution network voltage regulation and control method and system based on multi-agent strategy gradient

By adopting the multi-agent strategy gradient method in the distribution network for collaborative optimization, the problem that traditional methods are difficult to deal with complex distribution network environments is solved, and the effect of reducing voltage overlimit rate and network loss is achieved.

CN120016499AActive Publication Date: 2025-05-16NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510458626.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-16
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the active distribution network with highly penetrating distributed energy, the problem of voltage overload and network loss is prominent, and traditional methods are difficult to accurately describe the complex distribution network environment, which limits its application effect.

Method used

Using a multi-agent strategy gradient method, the intelligent control of photovoltaic inverters and energy storage equipment is realized through distributed multi-agent collaborative optimization, reducing voltage overlimit rate and network loss.

Benefits of technology

It improves the robustness and adaptability of the algorithm in complex environments, reduces the voltage limit rate and network loss, and ensures the performance stability of distribution network voltage regulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120016499A_ABST
    Figure CN120016499A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network voltage regulation and control method and system based on a multi-agent strategy gradient, and the method comprises the steps: carrying out the modeling of a photovoltaic inverter and energy storage equipment, obtaining an agent, and configuring an Actor-Critic network architecture in the agent; obtaining a local observation feature of the intelligent agent, and inputting the local observation feature into an online Actor network to obtain a current execution action; evaluating the current execution action through the Critic network to obtain a current prediction reward; updating the weight of the online Actor network in real time according to the current predicted reward; inputting the current execution action of each agent into a power distribution network voltage control model to obtain a system reward and a power distribution network system state; and updating the weight of the online Critic network according to the system reward and the state of the power distribution network system. According to the method, distributed multi-agent collaborative optimization is adopted, the voltage out-of-limit rate and the network loss are reduced, and the robustness and adaptability of the algorithm in a complex environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electric energy configuration, and specifically relates to a distribution network voltage control method and system based on multi-agent strategy gradient. Background Art

[0002] With the rapid development of the power system dominated by new energy, the voltage over-limit and network loss problems are becoming increasingly prominent in the active distribution network with a high penetration of distributed energy. In order to solve the voltage regulation problem caused by the high proportion of distributed photovoltaic access, the traditional method mainly relies on the on-load voltage regulation of the distribution transformer, the reactive power regulation of the photovoltaic inverter and the participation of energy storage equipment.

[0003] However, most of these methods rely on accurate physical models of distribution networks. But in practical applications, the operating environment of distribution networks is complex and changeable, and parameters change dynamically and frequently. Traditional models are difficult to accurately describe the actual situation, which restricts their actual application effect. Summary of the invention

[0004] The present invention provides a distribution network voltage control method and system based on multi-agent strategy gradient, which adopts distributed multi-agent collaborative optimization to realize intelligent control of photovoltaic inverters and energy storage equipment, reduce voltage over-limit rate and network loss, and improve the robustness and adaptability of the algorithm in complex environments.

[0005] In order to achieve the above object, the technical solution adopted by the present invention is:

[0006] The first aspect of the present invention provides a distribution network voltage control method based on multi-agent strategy gradient, comprising:

[0007] The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established; the nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices;

[0008] Modeling the photovoltaic inverter and energy storage equipment to obtain an intelligent agent, which is configured with an online Actor network (also called an online strategy network) and an online Critic network (also called an online evaluation network);

[0009] Obtain the local observation features of each agent, input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action; evaluate the current execution action through the Critic network to obtain the current predicted reward; update the weight of the online Actor network in real time according to the current predicted reward;

[0010] The current execution actions of each intelligent agent are input into the distribution network voltage control model to obtain the system reward and the distribution network system status; the weight of the online critic network is updated according to the system reward and the distribution network system status.

[0011] Furthermore, the DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established, which specifically includes:

[0012] The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and the distribution network voltage control model is established. The expression formula is:

[0013]

[0014]

[0015]

[0016] In the formula, the node and nodes is the adjacent node, node and nodes are adjacent nodes, and Slave nodes Transfer to Node Active power and reactive power; and Slave nodes Transfer to Node Active power and reactive power; and Node The injected active power and reactive power; and Slave nodes To Node resistance and reactance in circuits; and Node and nodes The voltage amplitude of Represented as a node The set of connected adjacent nodes;

[0017] Add node voltage constraints to the distribution network voltage control model, and the expression formula is:

[0018]

[0019] Where: and They represent the lower and upper limits of the voltage safety range respectively; is the set of all nodes in the power distribution network; For at the moment node The voltage amplitude of

[0020] Add PV inverter power constraints to the distribution network voltage control model, and the expression formula is:

[0021]

[0022] Where: Indicates at time node Active power of the internal photovoltaic inverter, Indicates at time node Reactive power of the internal photovoltaic inverter, Representative Node The apparent power, is the set of nodes containing photovoltaics in the distribution network; For Node The maximum active power of the internal photovoltaic inverter;

[0023] Add energy storage charge state constraints to the distribution network voltage control model, and the expression formula is:

[0024]

[0025]

[0026]

[0027] In the formula, is the minimum charging and discharging power of the energy storage device, For The output power of the energy storage device in node i at time, is the maximum charge and discharge power of the energy storage device; is a set of energy storage nodes; and Respectively in Moment and Time Node The capacity of the internal energy storage device; For Node The charging and discharging efficiency of internal energy storage devices; For Node Rated capacity of internal energy storage equipment; and are the minimum and maximum capacities of the energy storage device, respectively; is the unit time difference between two adjacent moments.

[0028] Furthermore, the intelligent body is also configured with a target Actor network (also called a target policy network) and a target Critic network (also called a target evaluation network); the initial weight parameters of the target Actor network and the online Actor network are the same, and the initial weight parameters of the target Critic network and the online Critic network are the same;

[0029] The weight of the Critic network is updated according to the system reward and the distribution network system status, including:

[0030] The local observation features of each agent at the next moment are obtained from the state of the distribution network system, and the local observation features at the next moment are input into the target Actor network to obtain the agent's execution action at the next moment, and the execution action at the next moment and the local observation features at the next moment are input into the target Critic network to obtain the future prediction reward;

[0031] The online critic network is updated according to the future prediction rewards, current prediction rewards and system rewards accumulated for H times of voltage regulation of the distribution network; the weight parameters of the online actor network are assigned to the target actor network at each set time interval, and the weight parameters of the online critic network are assigned to the target critic network.

[0032] Furthermore, the current actions of each agent are input into the distribution network voltage control model to obtain system rewards, which are expressed as follows:

[0033]

[0034]

[0035]

[0036] In the formula, For Time Node System rewards; is the power loss of the distribution network; and For Time Node and nodes The voltage exceeds the limit; Representation Node Cooperation index; For Time Node Active power of load; For Time Node The active power output of the internal photovoltaic inverter, For Time Node Active power output of internal energy storage device; is the number of nodes in the distribution network; For at the moment node The voltage amplitude of is a function that takes its maximum value with 0; and are the upper and lower limits of the voltage amplitude.

[0037] Furthermore, the weight of the online Actor network is updated in real time based on the current predicted reward, including:

[0038] The training loss function of the online Actor network is constructed based on the current predicted reward, and the expression formula is:

[0039]

[0040]

[0041] In the formula, is the training loss value of the online Actor network; is the number of nodes in the distribution network; For the Critic network, based on the node Local observation characteristics and Action The expected reward for the output; Output nodes for online Actor networks actions; is the weight parameter of the online Actor network in node i; is the mapping function representing the online Actor network;

[0042] The parameter iteration gradient of the online Actor network is calculated according to the training loss value of the online Actor network, and the weight parameter of the online Actor network is updated. The expression formula is:

[0043]

[0044] In the formula, Iterate gradients for parameters of online Actor networks; is the learning rate of the online Actor network.

[0045] Furthermore, the online critic network is updated based on the future prediction rewards, current prediction rewards and system rewards accumulated for H times of voltage regulation in the distribution network, including:

[0046] Calculate the target reward based on the future predicted reward and the system reward, calculate the prediction error based on the target reward and the current predicted reward, sample the prediction error according to the preset noise distribution to obtain the reward noise, and add the reward noise to the system reward to obtain the actual system reward;

[0047] Convert the future predicted reward and the current predicted reward into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution;

[0048] The training loss value of the online critic network is calculated based on the actual system rewards for H cumulative times of voltage regulation in the distribution network, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards; the weight parameters of the online critic network are updated based on the training loss value of the online critic network.

[0049] Further, determining the noise distribution of the reward noise specifically includes:

[0050] Obtain historical system rewards, historical work rewards, and historical prediction rewards from the experience pool; calculate historical target rewards based on historical system rewards and historical work rewards, and calculate historical prediction errors using historical target rewards and historical prediction rewards. The expression formula is:

[0051]

[0052] In the formula, Rewards for historical predictions; rewards for historical work; For Time Node Historical system rewards; For Time Node Historical forecast errors; For discount reasons;

[0053] Sort the historical system rewards according to the historical prediction errors and divide them into M groups, and calculate the variance and amplitude of each group's historical system rewards;

[0054]

[0055]

[0056] In the formula, is the variance of the historical system rewards of the mth group, For Node Historical system rewards, is the average of the historical system rewards of each node; n is the number of samples in the experience pool; is the maximum variance of historical system rewards for group M; is the magnitude of the historical system reward of the mth group;

[0057] The noise distribution of the reward noise is set according to the variance and amplitude of each group's historical system rewards.

[0058] Furthermore, the training loss value of the online critic network is calculated based on the actual system reward of the distribution network voltage regulation for H times, the future predicted reward probability distribution, and the current predicted reward probability distribution, including:

[0059] The target reward probability distribution is calculated based on the actual system reward and the future predicted reward probability distribution. The expression formula is:

[0060]

[0061] In the formula, is the target reward probability distribution, Actual rewards for the system in voltage regulation of distribution networks; is the expected function; is the discount factor; Predict reward probability distributions for the future;

[0062] The training loss value of the online Critic network is calculated based on the cumulative H times target reward probability distribution and the current predicted reward probability distribution. The expression formula is:

[0063]

[0064] In the formula, is the target reward probability distribution in the h-th distribution network voltage regulation process; is the current predicted reward probability distribution in the h-th distribution network voltage regulation process; is the cross entropy function, is the training loss value of the online Critic network, and h is the voltage control sequence number of the distribution network.

[0065] A second aspect of the present invention provides a distribution network voltage control system based on multi-agent strategy gradient, comprising:

[0066] The model building module adopts the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in the distribution network, and establishes a distribution network voltage control model; the nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices; the photovoltaic inverters and energy storage devices are modeled to obtain an intelligent body, and the intelligent body is configured with an online Actor network and an online Critic network;

[0067] The execution module is used to obtain the local observation features of each agent and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action;

[0068] The optimization module is used to evaluate the current execution action through the Critic network to obtain the current predicted reward; update the weight of the online Actor network in real time according to the current predicted reward; input the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and the distribution network system status; update the weight of the online Critic network according to the system reward and the distribution network system status.

[0069] The third aspect of the present invention provides an electronic device including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the distribution network voltage control method described in the first aspect of the present invention.

[0070] Compared with the prior art, the present invention has the following beneficial effects:

[0071] The present invention obtains the local observation characteristics of each intelligent agent, inputs the local observation characteristics into the online Actor network in the corresponding intelligent agent to obtain the current execution action, realizes the intelligent control of photovoltaic inverters and energy storage equipment, and regulates the distribution network voltage through multi-agent collaborative regulation to reduce the voltage over-limit rate and network loss.

[0072] The present invention evaluates the current execution action through the Critic network to obtain the current predicted reward; updates the weight of the online Actor network in real time according to the current predicted reward; and updates the weight of the online Actor network in real time, thereby ensuring that the intelligent agent can quickly adjust its strategy to adapt to environmental changes and optimize voltage regulation.

[0073] The present invention inputs the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and the distribution network system status; updates the weight of the online Critic network according to the system reward and the distribution network system status, so that the evaluation of the Critic network can more accurately estimate the value of the action, maintain the performance stability of the distribution network voltage adjustment, avoid falling into the local optimal solution, and thus guide the Actor network to generate better actions. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 It is a flow chart of a distribution network voltage control method based on multi-agent strategy gradient provided in Example 1;

[0075] Figure 2 is a node line diagram of a power distribution network provided in Example 1;

[0076] Figure 3is a voltage over-limit comparison diagram of various distribution network voltage control methods provided in Example 1;

[0077] Figure 4 This is a network loss comparison diagram of the distribution network voltage control method provided in Example 1;

[0078] Figure 5 is a graph showing a voltage over-limit rate change under disturbance conditions provided in Example 1;

[0079] Figure 6 This is a comparison chart of target rewards under different cumulative times provided in Example 1. DETAILED DESCRIPTION

[0080] The present invention will be further described below in conjunction with the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention, and cannot be used to limit the protection scope of the present invention.

[0081] Example 1

[0082] like Figure 1 As shown, this implementation provides a distribution network voltage control method based on multi-agent strategy gradient, including:

[0083] The DistFlow power flow equation (also known as the distribution network power flow equation) is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established; specifically, it includes:

[0084] like Figure 2 The distribution network shown in the figure uses the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in the distribution network, and establishes a distribution network voltage control model. The expression formula is:

[0085]

[0086]

[0087]

[0088] In the formula, the node and nodes is the adjacent node, node and nodes are adjacent nodes, and Slave nodes Transfer to Node Active power and reactive power; and Slave nodes Transfer to Node Active power and reactive power; and Node The injected active power and reactive power; and Slave nodes To Node resistance and reactance in circuits; and Node and nodes The voltage amplitude of Represented as a node The set of connected adjacent nodes;

[0089] Add node voltage constraints to the distribution network voltage control model, and the expression formula is:

[0090]

[0091] Where: and They represent the lower and upper limits of the voltage safety range respectively; is the set of all nodes in the power distribution network; For at the moment node The voltage amplitude of

[0092] Add PV inverter power constraints to the distribution network voltage control model, and the expression formula is:

[0093]

[0094] Where: Indicates at time node Active power of the internal photovoltaic inverter, Indicates at time node Reactive power of the internal photovoltaic inverter, Representative Node The apparent power, is the set of nodes containing photovoltaics in the distribution network; For Node The maximum active power of the internal photovoltaic inverter;

[0095] Add energy storage charge state constraints to the distribution network voltage control model, and the expression formula is:

[0096]

[0097]

[0098]

[0099] In the formula, is the minimum charging and discharging power of the energy storage device, For The output power of the energy storage device in node i at time, is the maximum charge and discharge power of the energy storage device; is a set of energy storage nodes; and Respectively in Moment and Time Node The capacity of the internal energy storage device; For Node The charging and discharging efficiency of internal energy storage devices; For Node Rated capacity of internal energy storage equipment; and are the minimum and maximum capacities of the energy storage device, respectively; is the unit time difference between two adjacent moments.

[0100] Modeling the voltage control of the distribution network as a multi-agent Markov game model can more efficiently handle complex constraints and achieve multi-agent collaborative optimization to improve voltage control performance; the nodes in the distribution network are equipped with photovoltaic inverters and / or energy storage devices; the photovoltaic inverters and energy storage devices are modeled to obtain agents, such as Figure 1 As shown, the intelligent body is configured with an online Actor network (online policy network) and an online Critic network (online evaluation network) as well as a target Actor network (target policy network) and a target Critic network (target evaluation network); the initial weight parameters of the target Actor network and the online Actor network are the same, and the initial weight parameters of the target Critic network and the online Critic network are the same.

[0101] Obtain the local observation features of each agent, input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action. The action space of the photovoltaic inverter is , the action space of the energy storage device is ; For The active power of the energy storage device in node i at time, For The reactive power of the energy storage device in node i at the moment; the current execution action is evaluated through the Critic network to obtain the current predicted reward.

[0102] The weight of the online Actor network is updated in real time based on the current predicted rewards, including:

[0103] The training loss function of the online Actor network is constructed based on the current predicted reward, and the expression formula is:

[0104]

[0105]

[0106] In the formula, is the training loss value of the online Actor network; is the number of nodes in the distribution network; For the Critic network, based on the node Local observation characteristics and Action The expected reward for the output; Output nodes for online Actor networks actions; is the weight parameter of the online Actor network in node i; is the mapping function representing the online Actor network;

[0107] The parameter iteration gradient of the online Actor network is calculated according to the training loss value of the online Actor network, and the weight parameter of the online Actor network is updated. The expression formula is:

[0108]

[0109] In the formula, Iterate gradients for parameters of online Actor networks; is the learning rate of the online Actor network.

[0110] The current actions of each agent are input into the distribution network voltage control model to obtain the system reward, which is expressed as:

[0111]

[0112]

[0113]

[0114] In the formula, For Time Node System rewards; is the power loss of the distribution network; and For Time Node and nodes The voltage exceeds the limit; Representation Node Cooperation index; For Time Node Active power of load; For Time Node The active power output of the internal photovoltaic inverter, For Time Node Active power output of internal energy storage device; is the number of nodes in the distribution network; For at the moment node The voltage amplitude of is a function that takes its maximum value with 0; and are the upper and lower limits of the voltage amplitude.

[0115] Input the current execution action of each agent into the distribution network voltage control model to obtain the distribution network system state; obtain the local observation features of each agent at the next moment from the distribution network system state, input the local observation features at the next moment into the target Actor network to obtain the agent's execution action at the next moment, and input the execution action at the next moment and the local observation features at the next moment into the target Critic network to obtain the future prediction reward;

[0116] Set the noise distribution of reward noise. The process includes:

[0117] Obtain historical system rewards, historical work rewards, and historical prediction rewards from the experience pool; calculate historical target rewards based on historical system rewards and historical work rewards, and calculate historical prediction errors using historical target rewards and historical prediction rewards. The expression formula is:

[0118]

[0119] In the formula, Rewards for historical predictions; rewards for historical work; For Time Node Historical system rewards; For Time Node Historical forecast errors; For discount reasons;

[0120] Sort the historical system rewards according to the historical prediction errors and divide them into M groups, and calculate the variance and amplitude of each group's historical system rewards;

[0121]

[0122]

[0123] In the formula, is the variance of the historical system rewards of the mth group, For Node Historical system rewards, is the average of the historical system rewards of each node; n is the number of samples in the experience pool; is the maximum variance of historical system rewards for group M; is the magnitude of the historical system reward of the mth group;

[0124] The noise distribution of the reward noise is set according to the variance and amplitude of each group's historical system rewards.

[0125] The target reward is calculated based on the future predicted reward and the system reward. The prediction error is calculated based on the target reward and the current predicted reward. The system reward is grouped according to the prediction error. The reward noise is obtained by sampling according to the preset noise distribution of each group. The reward noise is added to the system reward to obtain the actual system reward. The exploration ability is enhanced by dynamically adjusting the noise, avoiding local optimality, and improving the stability and adaptability of the strategy.

[0126] The future predicted rewards and current predicted rewards are converted into discrete probability distributions to obtain the future predicted rewards probability distribution and the current predicted rewards probability distribution, including:

[0127] Using weight vector To define the future predicted reward probability distribution and the current predicted reward probability distribution, use Represents the possible value of future prediction reward or current prediction reward; The reward for future predictions or the current prediction reward is equal to probability;

[0128]

[0129]

[0130] In the formula, and Indicates the maximum and minimum values ​​of future prediction rewards or current prediction rewards; Represents the partition length of discrete values.

[0131] The training loss value of the online critic network is calculated based on the actual system reward for H cumulative times of voltage regulation in the distribution network, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards. Specifically, it includes:

[0132] The target reward probability distribution is calculated based on the actual system reward and the future predicted reward probability distribution. The expression formula is:

[0133]

[0134] In the formula, is the target reward probability distribution, Actual rewards for the system in voltage regulation of distribution networks; is the expected function; is the discount factor; Predict reward probability distributions for the future;

[0135] The training loss value of the online Critic network is calculated based on the cumulative H times target reward probability distribution and the current predicted reward probability distribution. The expression formula is:

[0136]

[0137] In the formula, is the target reward probability distribution in the h-th distribution network voltage regulation process; is the current predicted reward probability distribution in the h-th distribution network voltage regulation process; is the cross entropy function, is the training loss value of the online Critic network, and h is the voltage control sequence number of the distribution network.

[0138] The weight parameters of the online critic network are updated according to the training loss value of the online critic network. The weight parameters of the online actor network are assigned to the target actor network at set intervals, and the weight parameters of the online critic network are assigned to the target critic network.

[0139] In order to verify the effect of this embodiment (abbreviated as A-MADPG algorithm), the adaptive multi-agent deep deterministic policy gradient algorithm (abbreviated as MADDPG algorithm) and the multi-agent soft actor-critic algorithm (MASAC algorithm) are selected for comparative analysis. Figure 3 As shown in the figure, in the early stage of training, the voltage over-limit rate of the distribution network is high, but as the training progresses, the A-MADPG algorithm can quickly reduce the voltage over-limit rate from the initial 0.3 to below 0.05, and stabilize it within 0.03 in the later stage. In comparison, the voltage over-limit rates of the traditional MASAC and MADDPG methods under the same conditions are about 0.05 and 0.06 respectively.

[0140] like Figure 4As shown in the figure, the A-MADPG algorithm quickly reduced the total network loss in the early stage of training and quickly stabilized at a lower level. In contrast, the MASAC algorithm can also quickly reduce the total network loss in the initial stage, but its decline rate is slower than that of A-MADPG. This is because the high randomness and maximum entropy requirement of the MASAC strategy usually require more time to converge to the optimal solution during training. The total network loss value in the later stage of training is slightly lower than that of A-MADPG, and the difference is very small. The MADDPG algorithm showed obvious fluctuations throughout the training process. The total network loss has been high and has not been significantly reduced. It tends to be unstable in the later stage (about 0.6-0.8). This shows that the MADDPG algorithm is less effective in optimizing voltage control strategies and reducing network losses, and is easily affected by environmental changes, resulting in poor convergence.

[0141] like Figure 5 As shown in the figure, the A-MADPG algorithm exhibits the lowest voltage over-limit rate and the fastest decline rate during the entire training process. At first, the voltage over-limit rate is high, but it drops rapidly in the early stage of training and remains at a level close to zero in the later stage. This shows that the A-MADPG algorithm can effectively cope with voltage fluctuations under disturbance conditions and maintain the voltage stability of the distribution network. The voltage over-limit rate of the MASAC algorithm is also high in the initial stage, but its decline rate is slower than that of the A-MADPG algorithm. Although the voltage over-limit rate decreases and stabilizes in the later stage of training, it still remains at a level slightly higher than that of the A-MADPG algorithm. In contrast. The voltage over-limit rate of the MADDPG algorithm fluctuates greatly during the entire training process and remains at a high level, showing significant instability. This result shows that the MADDPG algorithm cannot effectively control voltage fluctuations under disturbance conditions, and the voltage over-limit rate remains high, reflecting its insufficient adaptability in complex environments. This significant difference shows that the A-MADPG algorithm exhibits stronger robustness and adaptability when dealing with uncertain disturbances.

[0142] like Figure 6 As shown in the figure, when H=1, the convergence speed of the algorithm is slow, and the cumulative reward value gradually stabilizes after about 4000 rounds of training; when H=3, the convergence speed is significantly accelerated, and the number of training rounds is reduced to 2000; when H=5, the final reward value is slightly higher than H=1. It can be seen that moderately increasing the H value can effectively improve learning efficiency and strategy performance. However, the experimental results also show that the larger the H value, the better. In the additional experiment, when H is further increased to 20, the convergence curve of the cumulative reward value fluctuates significantly, and the final reward value decreases compared to H=5. When the H value is too large, the time span of the reward sequence may cover multiple policy adjustment stages. This makes the long-term return unable to accurately reflect the effect of the current strategy, resulting in the inability of the agent to effectively feedback short-term behavior in decision updates.

[0143] In order to improve the robustness of the algorithm in an uncertain environment, the present invention introduces an adaptive noise mechanism in the system reward to cope with random and uncertain environmental changes. Specifically, by dynamically adding symmetric noise to the system reward, the agent's exploration ability in low-variance areas is enhanced, while avoiding falling into local optimal solutions too early. This mechanism enables the agent to maintain a high level of strategic exploration, which helps to discover potential global optimal strategies. The amplitude of the noise is dynamically adjusted according to the variance of the area, with less noise in high-variance areas and more noise in low-variance areas, thereby effectively balancing exploration and development and improving the stability and adaptability of the control strategy.

[0144] This embodiment uses discrete probability distribution to represent all possible reward values ​​for state and action pairs. It no longer relies solely on a single value to evaluate the Q value, but uses probability distribution to reflect multiple possibilities in a complex environment. In addition, the Critic network uses cross entropy as the loss function, which further improves the accuracy and robustness of the value function estimation. This design enables the algorithm to better adapt to the complex actual power system environment and effectively handle high-dimensional and variable scenarios.

[0145] This embodiment provides more comprehensive information for the estimation of the value function by combining the immediate reward and the cumulative reward in the next H steps. It not only improves the algorithm's ability to evaluate long-term benefits, but also significantly accelerates the convergence of the model. During the training process, by reasonably selecting the H value, a balance can be achieved between short-term behavioral feedback and long-term planning, thereby improving the quality of decision-making in complex environments.

[0146] Example 2

[0147] This embodiment provides a distribution network voltage control system based on multi-agent policy gradient. The system described in this embodiment can be applied to the method described in Example 1. The distribution network voltage control system includes:

[0148] The model building module adopts the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in the distribution network, and establishes a distribution network voltage control model; the nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices; the photovoltaic inverters and energy storage devices are modeled to obtain an intelligent body, and the intelligent body is configured with an online Actor network and an online Critic network;

[0149] The execution module is used to obtain the local observation features of each agent and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action;

[0150] The optimization module is used to evaluate the current execution action through the Critic network to obtain the current predicted reward; update the weight of the online Actor network in real time according to the current predicted reward; input the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and the distribution network system status; update the weight of the online Critic network according to the system reward and the distribution network system status.

[0151] The intelligent body is also configured with a target Actor network and a target Critic network; the initial weight parameters of the target Actor network and the online Actor network are the same, and the initial weight parameters of the target Critic network and the online Critic network are the same;

[0152] The optimization module updates the weight of the Critic network according to the system reward and the distribution network system status, specifically including:

[0153] The local observation features of each agent at the next moment are obtained from the state of the distribution network system, and the local observation features at the next moment are input into the target Actor network to obtain the agent's execution action at the next moment, and the execution action at the next moment and the local observation features at the next moment are input into the target Critic network to obtain the future prediction reward;

[0154] Calculate the target reward based on the future predicted reward and the system reward, calculate the prediction error based on the target reward and the current predicted reward, sample the prediction error according to the preset noise distribution to obtain the reward noise, and add the reward noise to the system reward to obtain the actual system reward;

[0155] Convert the future predicted reward and the current predicted reward into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution;

[0156] The training loss value of the online critic network is calculated based on the actual system rewards of the cumulative H distribution network voltage regulation, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards; the weight parameters of the online critic network are updated based on the training loss value of the online critic network;

[0157] At each set time interval, the weight parameters of the online Actor network are assigned to the target Actor network, and the weight parameters of the online Critic network are assigned to the target Critic network.

[0158] Example 3

[0159] This embodiment provides an electronic device including a storage medium and a processor; the storage medium is used to store instructions; the processor is used to operate according to the instructions to execute the distribution network voltage control method described in Example 1.

[0160] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0161] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0162] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0163] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0164] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A distribution network voltage control method based on multi-agent strategy gradient, characterized in that: include: The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established; the nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices; The photovoltaic inverter and energy storage device are modeled to obtain an intelligent agent, which is configured with an online Actor network and an online Critic network. Obtain the local observation features of each agent, and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action; The current predicted reward is obtained by evaluating the current execution action through the Critic network; the weight of the online Actor network is updated in real time according to the current predicted reward; The current execution actions of each intelligent agent are input into the distribution network voltage control model to obtain the system reward and the distribution network system status; the weight of the online critic network is updated according to the system reward and the distribution network system status.

2. The method for controlling voltage in a power distribution network according to claim 1, characterized in that: The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and a distribution network voltage control model is established, which includes: The DistFlow power flow equation is used to describe the relationship between active power and reactive power between nodes in the distribution network, and the distribution network voltage control model is established. The expression formula is: ; ; ; In the formula, the node and nodes is the adjacent node, node and nodes are adjacent nodes, and Slave nodes Transfer to Node Active power and reactive power; and Slave nodes Transfer to Node Active power and reactive power; and Node The injected active power and reactive power; and Slave nodes To Node resistance and reactance in circuits; and Node and nodes The voltage amplitude; Represented as a node The set of connected adjacent nodes; Add node voltage constraints to the distribution network voltage control model, and the expression formula is: ; Where: and They represent the lower and upper limits of the voltage safety range respectively; is the set of all nodes in the power distribution network; For at the moment node The voltage amplitude; Add PV inverter power constraints to the distribution network voltage control model, and the expression formula is: ; Where: Indicates at time node Active power of the internal photovoltaic inverter, Indicates at time node Reactive power of the internal photovoltaic inverter, Representative Node The apparent power, is the set of nodes containing photovoltaics in the distribution network; For Node The maximum active power of the internal photovoltaic inverter; Add energy storage charge state constraints to the distribution network voltage control model, and the expression formula is: ; ; ; In the formula, is the minimum charging and discharging power of the energy storage device, For The output power of the energy storage device in node i at time, is the maximum charge and discharge power of the energy storage device; is a set of energy storage nodes; and Respectively in Moment and Time Node The capacity of the internal energy storage device; For Node The charging and discharging efficiency of internal energy storage devices; For Node Rated capacity of internal energy storage equipment; and are the minimum and maximum capacities of the energy storage device, respectively; is the unit time difference between two adjacent moments.

3. The method for controlling voltage in a power distribution network according to claim 1, characterized in that: The intelligent body is also configured with a target Actor network and a target Critic network; the initial weight parameters of the target Actor network and the online Actor network are the same, and the initial weight parameters of the target Critic network and the online Critic network are the same; The weight of the Critic network is updated according to the system reward and the distribution network system status, including: The local observation features of each agent at the next moment are obtained from the state of the distribution network system, and the local observation features at the next moment are input into the target Actor network to obtain the agent's execution action at the next moment, and the execution action at the next moment and the local observation features at the next moment are input into the target Critic network to obtain the future prediction reward; The online critic network is updated according to the future prediction rewards, current prediction rewards and system rewards accumulated for H times of voltage regulation of the distribution network; the weight parameters of the online actor network are assigned to the target actor network at each set time interval, and the weight parameters of the online critic network are assigned to the target critic network.

4. The method for controlling voltage in a power distribution network according to claim 3, characterized in that: The current actions of each agent are input into the distribution network voltage control model to obtain the system reward, which is expressed as: ; ; ; In the formula, For Time Node System rewards; is the power loss of the distribution network; and For Time Node and nodes The voltage exceeds the limit; Representation Node Cooperation index; For Time Node Active power of load; For Time Node The active power output of the internal photovoltaic inverter, For Time Node Active power output of internal energy storage equipment; is the number of nodes in the distribution network; For at the moment node The voltage amplitude; is a function that takes its maximum value with 0; and are the upper and lower limits of the voltage amplitude.

5. The method for controlling voltage in a power distribution network according to claim 1, characterized in that: The weight of the online Actor network is updated in real time based on the current predicted rewards, including: The training loss function of the online Actor network is constructed based on the current predicted reward, and the expression formula is: ; ; In the formula, is the training loss value of the online Actor network; is the number of nodes in the distribution network; For the Critic network, based on the node Local observation characteristics and Action The expected reward for the output; Output nodes for online Actor networks actions; For Node Local observation characteristics of is the weight parameter of the online Actor network in node i; is the mapping function representing the online Actor network; The parameter iteration gradient of the online Actor network is calculated according to the training loss value of the online Actor network, and the weight parameter of the online Actor network is updated. The expression formula is: ; In the formula, Iterate gradients for parameters of online Actor networks; is the learning rate of the online Actor network.

6. The method for controlling voltage in a distribution network according to claim 3, characterized in that: The online critic network is updated based on the future prediction rewards, current prediction rewards and system rewards accumulated for H times of voltage regulation in the distribution network, including: Calculate the target reward based on the future predicted reward and the system reward, calculate the prediction error based on the target reward and the current predicted reward, sample the prediction error according to the preset noise distribution to obtain the reward noise, and add the reward noise to the system reward to obtain the actual system reward; Convert the future predicted reward and the current predicted reward into discrete probability distributions to obtain the future predicted reward probability distribution and the current predicted reward probability distribution; The training loss value of the online critic network is calculated based on the actual system rewards for H cumulative times of voltage regulation in the distribution network, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards; the weight parameters of the online critic network are updated based on the training loss value of the online critic network.

7. The method for controlling voltage in a power distribution network according to claim 6, characterized in that: Determining the noise distribution of the reward noise specifically includes: Obtain historical system rewards, historical work rewards, and historical prediction rewards from the experience pool; calculate historical target rewards based on historical system rewards and historical work rewards, and calculate historical prediction errors using historical target rewards and historical prediction rewards. The expression formula is: ; In the formula, Rewards for historical predictions; rewards for historical work; For Time Node Historical system rewards; For Time Node Historical forecast errors; is the discount factor; Sort the historical system rewards according to the historical prediction errors and divide them into M groups, and calculate the variance and amplitude of each group's historical system rewards; ; ; In the formula, is the variance of the historical system rewards of the mth group, For Node Historical system rewards, is the average of the historical system rewards of each node; n is the number of samples in the experience pool; is the maximum variance of historical system rewards for group M; is the magnitude of the historical system reward of the mth group; The noise distribution of the reward noise is set according to the variance and amplitude of each group's historical system rewards.

8. The method for voltage control in a distribution network according to claim 6, characterized in that: The training loss value of the online critic network is calculated based on the actual system reward for H cumulative times of voltage regulation in the distribution network, the probability distribution of future predicted rewards, and the probability distribution of current predicted rewards. Specifically, it includes: The target reward probability distribution is calculated based on the actual system reward and the future predicted reward probability distribution. The expression formula is: ; In the formula, is the target reward probability distribution, Actual rewards for the system in voltage regulation of distribution networks; is the expected function; is the discount factor; Predict reward probability distributions for the future; Output nodes for online Actor networks actions; For Node Local observation characteristics of The training loss value of the online Critic network is calculated based on the cumulative H times target reward probability distribution and the current predicted reward probability distribution. The expression formula is: ; In the formula, is the target reward probability distribution in the h-th distribution network voltage regulation process; is the current predicted reward probability distribution in the h-th distribution network voltage regulation process; is the cross entropy function, is the training loss value of the online Critic network, and h is the voltage control sequence number of the distribution network.

9. A distribution network voltage control system based on multi-agent strategy gradient, characterized in that: include: The model building module adopts the DistFlow power flow equation to describe the relationship between active power and reactive power between nodes in the distribution network, and establishes a distribution network voltage control model; the nodes in the distribution network are configured with photovoltaic inverters and / or energy storage devices; the photovoltaic inverters and energy storage devices are modeled to obtain an intelligent body, and the intelligent body is configured with an online Actor network and an online Critic network; The execution module is used to obtain the local observation features of each agent and input the local observation features into the online Actor network in the corresponding agent to obtain the current execution action; The optimization module is used to evaluate the current execution action through the Critic network to obtain the current predicted reward; update the weight of the online Actor network in real time according to the current predicted reward; input the current execution action of each intelligent agent into the distribution network voltage control model to obtain the system reward and the distribution network system status; update the weight of the online Critic network according to the system reward and the distribution network system status.

10. An electronic device comprises a storage medium and a processor; the storage medium is used to store instructions; and characterized in that: The processor is used to operate according to the instruction to execute the distribution network voltage control method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Power distribution network voltage autonomous optimization control method and device

    CN113872213A

  • Power distribution network voltage reactive power optimization method based on graph reinforcement learning

    CN115588998A

  • Voltage out-of-limit control system and method based on multi-agent deep reinforcement learning

    CN119401416A

Cited By

  • Intelligent regulation and control type low-voltage intelligent power distribution control system

    CN120855557A