Power distribution network disturbance self-healing voltage optimization control method, device and system based on multi-agent deep reinforcement learning

By employing a multi-agent deep reinforcement learning approach, the reactive power regulation of photovoltaic inverters is optimized, solving the voltage problem caused by distributed photovoltaic power generation systems and achieving efficient voltage optimization and network loss management in the distribution network.

CN120955688APending Publication Date: 2025-11-14STATE GRID JIANGSU ELECTRIC POWER CO LTD RESEARCH INSTITUTE +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511075773.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional voltage control methods are ill-suited to the randomness and volatility of distributed photovoltaic power generation systems, leading to voltage overruns and power quality issues in the distribution network. Furthermore, the uniform sampling strategy of traditional reinforcement learning results in low learning efficiency.

Method used

A multi-agent deep reinforcement learning approach is adopted to model the distribution network voltage optimization problem as a partitioned Markov decision model, which is then solved using the deep reinforcement learning MGPER-MATD3 model. The reactive power of the photovoltaic inverter is optimized through a priority sampling strategy, and voltage optimization is performed by combining the global reward function and constraints.

Benefits of technology

It improves voltage control effectiveness and experience utilization efficiency, and is suitable for reactive voltage control in distributed distribution networks, achieving more efficient voltage optimization and network loss management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120955688A_ABST
    Figure CN120955688A_ABST
Patent Text Reader

Abstract

The invention discloses a power distribution network disturbance self-healing voltage optimization control method, device and system based on multi-agent deep reinforcement learning. The power distribution network disturbance self-healing voltage optimization control method comprises the following steps: modeling a distributed power distribution network voltage optimization problem as a partition Markov decision model; the partition Markov decision model is solved through a deep reinforcement learning MGPER-MATD3 model, reactive power of each photovoltaic inverter in the distributed power distribution network is obtained, and voltage optimization of the power distribution network is completed; wherein the deep reinforcement learning MGPER-MATD3 model performs sampling according to the priority of each experience during experience sampling, and the priority of each experience is related to the TD error, the experience rare degree and the experience reward feature value. The method has the advantages of good voltage control effect, high experience utilization efficiency and the like, and is suitable for the field of distributed power distribution network reactive voltage control and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of distribution network voltage optimization technology, specifically relating to a distribution network disturbance self-healing voltage optimization control method, device and system based on multi-agent deep reinforcement learning. Background Technology

[0002] With technological advancements, distributed photovoltaic (PV) power generation systems are increasingly being used in smart distribution networks. Distributed PV output exhibits significant randomness and volatility, easily leading to power quality issues such as voltage exceeding limits and fluctuations. This large-scale integration of uncertain load sources poses a severe challenge to the safety and stability of the distribution network. Faced with complex distribution network topologies and variable operating conditions, traditional voltage control methods struggle to meet real-time control requirements.

[0003] In recent years, artificial intelligence technology has been widely applied in the field of power system optimization, showing good results in load forecasting, voltage control, and fault diagnosis. Multi-agent deep reinforcement learning has advantages in distributed decision-making and adaptive learning. However, the uniform sampling strategy of traditional reinforcement learning struggles to effectively identify and utilize key experience samples, leading to low learning efficiency. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a self-healing voltage optimization control method, device, and system for distribution network disturbances based on multi-agent deep reinforcement learning. This method offers advantages such as good voltage control performance and high efficiency in utilizing experience, making it suitable for fields such as reactive power and voltage control in distributed distribution networks.

[0005] To achieve the above-mentioned technical objectives and effects, the present invention is implemented through the following technical solution:

[0006] In a first aspect, the present invention provides a method for self-healing voltage optimization control of distribution network disturbances based on multi-agent deep reinforcement learning, comprising:

[0007] The distributed distribution network voltage optimization problem is modeled as a partitioned Markov decision model. The state space of the partitioned Markov decision model includes the voltage amplitude and voltage phase angle of each node in the distribution network, the reactive power regulation and active power regulation of the photovoltaic inverter, and the active and reactive loads of the distribution network. The action space of the partitioned Markov decision model includes the reactive power regulation action of the photovoltaic inverter. The global reward function of the partitioned Markov decision model corresponds to the control objective of the distributed distribution network voltage optimization problem.

[0008] The partitioned Markov decision model is solved by the deep reinforcement learning MGPER-MATD3 model to obtain the reactive power of each photovoltaic inverter in the distributed distribution network and complete the voltage optimization of the distribution network. In this model, the deep reinforcement learning MGPER-MATD3 model samples according to the priority of each experience when performing experience sampling. The priority of each experience is related to the TD error, experience rarity and experience reward feature value.

[0009] In conjunction with the first aspect, optionally, the control objective of the distributed distribution network voltage optimization problem is:

[0010] ,

[0011] in:

[0012] ,

[0013] ,

[0014] The constraints of the distributed distribution network voltage optimization problem include:

[0015] Conventional power flow constraints:

[0016] ,

[0017] Voltage amplitude constraint:

[0018] ,

[0019] Output constraints of photovoltaic inverter units:

[0020] ,

[0021] ;

[0022] In the formula, It's voltage deviation. It's network overhead. It is the voltage deviation weighting coefficient. It is the network loss weighting coefficient. yes Time step node voltage, It is the rated voltage. yes Time-step branch road electrical conductivity, It is the total number of instruction cycles. The total number of nodes. It is a node and nodes exist The cosine of the voltage phase difference at each time step. for Time step node voltage, and For nodes exist Active and reactive power at each time step; and For nodes Active and reactive loads; and For nodes and nodes exist Voltage amplitude at time step; and For nodes and nodes The conductance and susceptance of the branches between them; Let be the voltage phase difference between node i and node j at time step t; and These are the upper and lower limits of the node voltage; and These are the lower and upper limits of the active power of the photovoltaic inverter; This refers to the reactive power of the photovoltaic inverter. This represents the maximum adjustable reactive power of the photovoltaic inverter. , This refers to the capacity of the photovoltaic inverter; This refers to the active power of the photovoltaic inverter.

[0023] In conjunction with the first aspect, optionally, the mathematical expression of the partitioned Markov decision model is:

[0024] ,

[0025] ,

[0026] ,

[0027] In the formula, It is a state space, representing in Distribution network status parameters obtained from the distribution network at time step; Indicates in Time Step Voltage amplitude at each node; Indicates in Time Step The voltage phase angle of each node, Indicates in Time Step The reactive power regulation of a photovoltaic inverter; Indicates in Time Step The active power regulation of a photovoltaic inverter; For the distribution network system Active and reactive loads at each time step; yes The joint action space of all agents in time step, action space Indicates the first A photovoltaic inverter in The reactive power adjustment action of the time step; This represents the total number of photovoltaic inverters. It is a global reward function used to minimize voltage deviation and network losses while ensuring that the distribution network system complies with constraints; yes Time Step The per-unit voltage value of each node.

[0028] In conjunction with the first aspect, optionally, the sampling strategy of the deep reinforcement learning MGPER-MATD3 model is as follows:

[0029] For each experience in the replay pool, calculate its corresponding priority;

[0030] Normalize all priorities and calculate the sampling probability for each experience.

[0031] In conjunction with the first aspect, optionally, the formula for calculating the priority is:

[0032] ,

[0033] In the formula, for Time step experience priority, yes Time step experience TD error, express Time step experience The rarity of an experience, i.e., the frequency of accessing that experience; express Time step experience Experience reward characteristic value; and These represent the weights, It is a constant. .

[0034] In conjunction with the first aspect, optionally, the aforementioned Time step experience The formula for calculating the TD error is:

[0035] ,

[0036] In the formula, yes Time step experience The instant reward received; It is a discount factor, and its value range is... This is used to measure the importance of future rewards; For experience exist Time step status Take action below Q-value estimation at that time; For experience exist Time step status Take action below Q-value estimation at that time.

[0037] In conjunction with the first aspect, optionally, the aforementioned Time step experience The formula for calculating the rarity of experience is:

[0038] ,

[0039] in:

[0040] ,

[0041] In the formula, For experience Medium experience status The number of times it appears; It is experience The empirical cumulative distribution function of the access frequency of all states; This refers to the access frequency of other experiences in the experience replay pool; Total number of experiences; This is an indicator function; it returns 1 if the condition is true, and 0 otherwise.

[0042] In conjunction with the first aspect, optionally, the aforementioned Time step experience The formula for calculating the experience reward feature value is:

[0043] ,

[0044] In the formula, The reward value threshold, yes Time step experience The instant reward received.

[0045] In conjunction with the first aspect, optionally, solving the partitioned Markov decision model using the deep reinforcement learning MGPER-MATD3 model includes:

[0046] Based on the MATD3 algorithm, the agent continuously interacts with the distribution network to collect experience and store it in the playback pool. Each experience includes the current state of the distribution network, the current action, the current reward, the next state, and the done flag.

[0047] Once the amount of experience in the replay pool reaches a certain threshold, the policy update will begin.

[0048] After each policy update, the TD error, experience rarity, and experience reward characteristics of all experiences are calculated, and the priority of experiences is updated to implement the MGPER mechanism.

[0049] Secondly, the present invention provides a distribution network disturbance self-healing voltage optimization control device based on multi-agent deep reinforcement learning, comprising:

[0050] The modeling module is used to model the distributed distribution network voltage optimization problem as a partitioned Markov decision model. The state space of the partitioned Markov decision model includes the voltage amplitude and voltage phase angle of each node in the distribution network, the reactive power regulation and active power regulation of the photovoltaic inverter, and the active and reactive loads of the distribution network. The action space of the partitioned Markov decision model includes the reactive power regulation action of the photovoltaic inverter. The global reward function of the partitioned Markov decision model corresponds to the control objective of the distributed distribution network voltage optimization problem.

[0051] The optimization module is used to solve the partitioned Markov decision model using the deep reinforcement learning MGPER-MATD3 model to obtain the reactive power of each photovoltaic inverter in the distributed distribution network and complete the distribution network voltage optimization. The deep reinforcement learning MGPER-MATD3 model samples according to the priority of each experience when performing experience sampling. The priority of each experience is related to the TD error, experience rarity, and experience reward feature value.

[0052] Thirdly, the present invention provides a distribution network disturbance self-healing voltage optimization control system based on multi-agent deep reinforcement learning, including a storage medium and a processor.

[0053] The storage medium is used to store instructions;

[0054] The processor is configured to operate according to the instructions to perform the method according to any one of the first aspects.

[0055] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0056] This invention integrates a multi-granularity priority experience replay mechanism into the MATD3 model. Compared with the original MATD3 model and the traditional MADDPG model, this invention has the advantages of better voltage control effect and higher experience utilization efficiency, and is applicable to fields such as reactive power and voltage control in distributed distribution networks. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly described below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein:

[0058] Figure 1 This is a schematic diagram of the data processing principle of a single intelligent agent according to an embodiment of the present invention;

[0059] Figure 2 This is a flowchart illustrating the training process of the MGPER-MATD3 model according to an embodiment of the present invention.

[0060] Figure 3 This is the modified topology of the IEEE 33-node system;

[0061] Figure 4 Typical daily photovoltaic power output and load curves;

[0062] Figure 5 A reward curve for the training process;

[0063] Figure 6 This is a graph showing the per-unit voltage values ​​of some nodes under different control strategies on a typical day. Detailed Implementation

[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0065] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0066] Example 1

[0067] This invention provides a method for self-healing voltage optimization control of distribution network disturbances based on multi-agent deep reinforcement learning, comprising the following steps:

[0068] (1) The distributed distribution network voltage optimization problem is modeled as a partitioned Markov decision model. The state space of the partitioned Markov decision model includes the voltage amplitude and voltage phase angle of each node in the distribution network, the reactive power regulation and active power regulation of the photovoltaic inverter, and the active and reactive loads of the distribution network. The action space of the partitioned Markov decision model includes the reactive power regulation action of the photovoltaic inverter. The global reward function of the partitioned Markov decision model corresponds to the control objective of the distributed distribution network voltage optimization problem.

[0069] (2) Solve the partitioned Markov decision model by using the deep reinforcement learning MGPER-MATD3 model to obtain the reactive power of each photovoltaic inverter in the distributed distribution network and complete the voltage optimization of the distribution network; wherein, when the deep reinforcement learning MGPER-MATD3 model performs experience sampling, it samples according to the priority of each experience, and the priority of each experience is related to the TD error, experience sparseness and experience reward feature value.

[0070] In one specific embodiment of the present invention, the control objective of the distributed distribution network voltage optimization problem is:

[0071] ,

[0072] in:

[0073] ,

[0074] ,

[0075] The constraints of the distributed distribution network voltage optimization problem include:

[0076] Conventional power flow constraints:

[0077] ,

[0078] Voltage amplitude constraint:

[0079] ,

[0080] Output constraints of photovoltaic inverter units:

[0081] ,

[0082] ;

[0083] In the formula, It's voltage deviation. It's network overhead. It is the voltage deviation weighting coefficient. It is the network loss weighting coefficient. yes Time step node voltage, It is the rated voltage. yes Time-step branch road electrical conductivity, It is the total number of instruction cycles. The total number of nodes. It is a node and nodes exist The cosine of the voltage phase difference at each time step. for Time step node voltage, and For nodes exist Active and reactive power at each time step; and For nodes Active and reactive loads; and For nodes and nodes exist Voltage amplitude at time step; and For nodes and nodes The conductance and susceptance of the branches between them; Let be the voltage phase difference between node i and node j at time step t; and These are the upper and lower limits of the node voltage; and These are the lower and upper limits of the active power of the photovoltaic inverter; This refers to the reactive power of the photovoltaic inverter. This represents the maximum adjustable reactive power of the photovoltaic inverter. , The capacity of the photovoltaic inverter; This refers to the active power of the photovoltaic inverter.

[0084] In one specific embodiment of the present invention, the mathematical expression of the partitioned Markov decision model is:

[0085] ,

[0086] ,

[0087] ,

[0088] In the formula, It is a state space, representing in Distribution network status parameters obtained from the distribution network at time step; Indicates in Time Step Voltage amplitude at each node; Indicates in Time Step The voltage phase angle of each node, Indicates in Time Step The reactive power regulation of a photovoltaic inverter; Indicates in Time Step The active power regulation of a photovoltaic inverter; For the distribution network system Active and reactive loads at each time step; yes The joint action space of all intelligent agents in time step, action space Indicates the first A photovoltaic inverter in The reactive power adjustment action of the time step; This represents the total number of photovoltaic inverters. It is a global reward function used to minimize voltage deviation and network losses while ensuring that the distribution network system complies with constraints; yes Time Step The per-unit voltage value of each node.

[0089] In one specific embodiment of the present invention, the sampling strategy of the deep reinforcement learning MGPER-MATD3 model is as follows:

[0090] For each experience in the replay pool, calculate its corresponding priority;

[0091] Normalize all priorities and calculate the sampling probability for each experience.

[0092] In one specific embodiment of the present invention, the priority calculation formula is as follows:

[0093] ,

[0094] In the formula, for Time step experience priority, yes Time step experience TD error, express Time step experience The rarity of an experience, i.e., the frequency of accessing that experience; express Time step experience Experience reward characteristic value; and These represent the weights, It is a constant. .

[0095] In one specific embodiment of the present invention, the Time step experience The formula for calculating the TD error is:

[0096] ,

[0097] In the formula, yes Time step experience The instant reward received; It is a discount factor, and its value range is... This is used to measure the importance of future rewards; For experience exist Time step status Take action below Q-value estimation at that time; For experience exist Time step status Take action below Q-value estimation at that time.

[0098] In one specific embodiment of the present invention, the Time step experience The formula for calculating the rarity of experience is:

[0099] ,

[0100] in:

[0101] ,

[0102] In the formula, For experience Medium experience status The number of times it appears; It is experience The empirical cumulative distribution function of the access frequency of all states; This refers to the access frequency of other experiences in the experience replay pool; Total number of experiences; This is an indicator function; it returns 1 if the condition is true, and 0 otherwise.

[0103] In one specific embodiment of the present invention, the Time step experience The formula for calculating the experience reward feature value is:

[0104] ,

[0105] In the formula, The reward value threshold, yes Time step experience The instant reward received.

[0106] In one specific embodiment of the present invention, solving the partitioned Markov decision model using the deep reinforcement learning MGPER-MATD3 model includes:

[0107] Based on the MATD3 algorithm, the agent continuously interacts with the distribution network to collect experience and store it in the playback pool. Each experience includes the current state of the distribution network, the current action, the current reward, the next state, and the done flag.

[0108] Once the amount of experience in the replay pool reaches a certain threshold, the policy update will begin.

[0109] After each policy update, the TD error, experience rarity, and experience reward characteristics of all experiences are calculated, and the priority of experiences is updated to implement the MGPER mechanism (Multi-Granularity Prioritized Experience ReplayMechanism).

[0110] In the specific implementation process, the partitioned Markov decision model is solved using the deep reinforcement learning MGPER-MATD3 model, as follows: Figure 1 and Figure 2 As shown, the specific steps include:

[0111] (1) Initialization. At the start of training, the target network is initialized as a copy of the parameters of the current actor and critic networks and will be used to calculate the target Q value. Then, the capacity and priority storage method of the experience replay pool are set to prepare for the implementation of a multi-granularity priority experience replay mechanism. The experience replay pool is used to store the interaction data between the agent and the power distribution network. Relatively important experiences can be stored with higher priority in the experience replay pool to ensure they are used more frequently in subsequent training, thereby accelerating the learning speed.

[0112] (2) Experience Collection. The agent continuously interacts with the power distribution network to collect experience, which serves as the sample. First, the agent collects experience at each time step (instantaneous step). The agent selects an action based on the current state, namely adjusting the reactive power of the photovoltaic inverter. Then, the agent executes this action, and the distribution network environment returns a new state and a corresponding reward value based on the action. The new state reflects the grid configuration or state change after the action was executed. Finally, the state, action, reward, new state, and "done" flag of this interaction process are stored in the experience replay pool.

[0113] (3) Strategy Update. After the number of experiences in the experience replay pool reaches a certain threshold, the reactive voltage control strategy update begins. The core steps of the strategy update include priority sampling, target Q value calculation, TD error calculation, and network update. First, a batch of experiences is sampled from the experience replay pool according to the priority of the experiences to obtain the corresponding distribution network status, actions, rewards, next status, and done flag. Then, the target Q value is calculated through the target Critic network, the next Q value is calculated using the next status and actions, and the final target Q value is obtained by combining the current reward and discount factor. Next, the current Q value is calculated through the current Critic network and compared with the target Q value to obtain the TD error. The TD error reflects the difference between the current Q value estimate and the true value. Then, weighted loss calculation is performed to weight the loss of the Critic network so as to pay more attention to important experiences during the update process. Then, the Critic network is updated: the parameters of the Critic network are optimized through backpropagation to reduce prediction errors. After the Critic network is optimized, the Actor network is updated. The loss of the Actor network is achieved by maximizing the Q-value of the Critic network. Finally, the parameters of the target network are updated to ensure that the weights of the target network slowly follow the current network.

[0114] (4) Priority Update. Priorities are calculated based on the priority evaluation mechanism. After each strategy update, the system calculates the TD error, experience scarcity, and experience reward characteristics of all experiences and updates the priority of each experience. The greater the weight of an experience, the greater its impact on the current strategy, thus giving it a higher priority. For distribution networks, a large TD error indicates that the agent's chosen action has not effectively improved grid stability or reduced losses; conversely, a small TD error means that the agent's action has already achieved relatively accurate grid control. Experience scarcity reflects the frequency of occurrence of the state-action pair in the entire experience pool. For states that are only triggered under specific load disturbances, distributed energy access, or fault switching, although the frequency is low, they often correspond to high-risk scenarios, and the agent should focus on learning them to enhance robustness. Experience reward characteristics focus on the reward signal characteristics of the experience in grid control tasks, such as significant voltage recovery and significant loss reduction—high-value feedback. These experiences have clear significance in guiding the strategy towards the optimization goal, and therefore should also be given a higher sampling priority. This priority update process ensures that the priorities in the experience replay pool can dynamically reflect the learning state of the agent. In subsequent sampling, high-importance experiences are used more frequently, which helps the agent to adjust the voltage control strategy more effectively in a short period of time.

[0115] In the specific implementation process, the MGPER-MATD3 model needs to be trained and tested. The steps include:

[0116] (1) Use such as Figure 3 The modified IEEE 33-node example shown is used for training. Figure 3 In this system, PV represents photovoltaic inverters. The reference voltage of the node system is set to 12.66 kV, the per-unit voltage of the root node is 1.0 pu, and the maximum allowable voltage deviation of the grid is ±0.05 pu. Distributed photovoltaic devices with a capacity of 1.5 MW of photovoltaic inverters are installed at nodes 6, 11, 16, 23, 26, and 33. Algorithm training and testing parameters are set, and the specific algorithm parameters are shown in Table 1.

[0117] Table 1 Algorithm Parameter Settings

[0118] .

[0119] (2) Divide the data into training set and test set, and use the training set to train the MGPER-MATD3 model. The reward curve obtained from the training is as follows: Figure 4 The performance of the model was tested using a test set, and detailed data such as voltage deviation and network loss on typical days in the test set are shown in Table 2. The specific photovoltaic output and load curves of the test set are shown below. Figure 5 As shown. Figure 6 This is a per-unit voltage curve of some nodes under different control strategies on a typical day. The highest voltage in the original system data for that day occurred at node 17, and the lowest voltage at node 18. After adjustment with MGPER-MATD3, the highest voltage still occurred at node 17, and the lowest voltage at node 25. Another node, 33, was also selected. The per-unit voltage curves of nodes 17, 18, 25, and 33 were compared.

[0120] Table 2 Dataset Partitioning

[0121]

[0122] The results show that the MGPER-MATD3 model proposed in this invention is significantly better than the MADDPG model in controlling the average voltage deviation, and slightly better than the MATD3 model. In terms of controlling network loss, it is slightly worse than the MADDPG model, and slightly better than the MATD3 model. Considering all factors, the MGPER-MATD3 model has the best reactive power voltage optimization control performance, but there is still room for improvement in terms of controlling network loss.

[0123] Example 2

[0124] Based on the same inventive concept as Embodiment 1, this embodiment of the invention provides a distribution network disturbance self-healing voltage optimization control device based on multi-agent deep reinforcement learning, comprising:

[0125] The modeling module is used to model the distributed distribution network voltage optimization problem as a partitioned Markov decision model. The state space of the partitioned Markov decision model includes the voltage amplitude and voltage phase angle of each node in the distribution network, the reactive power regulation and active power regulation of the photovoltaic inverter, and the active and reactive loads of the distribution network. The action space of the partitioned Markov decision model includes the reactive power regulation action of the photovoltaic inverter. The global reward function of the partitioned Markov decision model corresponds to the control objective of the distributed distribution network voltage optimization problem.

[0126] The optimization module is used to solve the partitioned Markov decision model through a deep reinforcement learning MGPER-MATD3 network to obtain the reactive power of each photovoltaic inverter in the distributed distribution network and complete the voltage optimization of the distribution network. The deep reinforcement learning MGPER-MATD3 network samples according to the priority of each experience when performing experience sampling. The priority of each experience is related to the TD error, experience rarity, and experience reward feature value.

[0127] Example 3

[0128] Based on the same inventive concept as in Embodiment 1, this embodiment of the invention provides a distribution network disturbance self-healing voltage optimization control system based on multi-agent deep reinforcement learning, including a storage medium and a processor.

[0129] The storage medium is used to store instructions;

[0130] The processor is configured to operate according to the instructions to execute the method according to any one of Embodiment 1.

[0131] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0133] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0134] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0135] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

[0136] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A method for self-healing voltage optimization control of distribution network disturbances based on multi-agent deep reinforcement learning, characterized in that, include: The distributed distribution network voltage optimization problem is modeled as a partitioned Markov decision model. The state space of the partitioned Markov decision model includes the voltage amplitude and voltage phase angle of each node in the distribution network, the reactive power regulation and active power regulation of the photovoltaic inverter, and the active and reactive loads of the distribution network. The action space of the partitioned Markov decision model includes the reactive power regulation action of the photovoltaic inverter. The global reward function of the partitioned Markov decision model corresponds to the control objective of the distributed distribution network voltage optimization problem. The partitioned Markov decision model is solved by using the deep reinforcement learning MGPER-MATD3 model to obtain the reactive power of each photovoltaic inverter in the distributed distribution network, thus completing the voltage optimization of the distribution network. The deep reinforcement learning MGPER-MATD3 model samples experiences based on their priority during the sampling process. The priority of each experience is related to the TD error, experience rarity, and experience reward feature value.

2. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 1, characterized in that: The control objective of the distributed distribution network voltage optimization problem is: , in: , , The constraints of the distributed distribution network voltage optimization problem include: Conventional power flow constraints: , Voltage amplitude constraint: , Output constraints of photovoltaic inverter units: , ; In the formula, It's voltage deviation. It's network overhead. It is the voltage deviation weighting coefficient. It is the network loss weighting coefficient. yes Time step node voltage, It is the rated voltage. yes Time-step branch road electrical conductivity, It is the total number of instruction cycles. The total number of nodes. It is a node and nodes exist The cosine of the voltage phase difference at each time step. for Time step node voltage, and For nodes exist Active and reactive power at each time step; and For nodes Active and reactive loads; and For nodes and nodes exist Voltage amplitude at time step; and For nodes and nodes The conductance and susceptance of the branches between them; Let be the voltage phase difference between node i and node j at time step t; and These are the upper and lower limits of the node voltage; and These are the lower and upper limits of the active power of the photovoltaic inverter; This refers to the reactive power of the photovoltaic inverter. This represents the maximum adjustable reactive power of the photovoltaic inverter. , The capacity of the photovoltaic inverter; This refers to the active power of the photovoltaic inverter.

3. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 2, characterized in that: The mathematical expression for the partitioned Markov decision model is: , , , In the formula, It is a state space, representing in Distribution network status parameters obtained from the distribution network at time step; Indicates in Time Step Voltage amplitude at each node; Indicates in Time Step The voltage phase angle of each node, Indicates in Time Step The reactive power regulation of a photovoltaic inverter; Indicates in Time Step The active power regulation of a photovoltaic inverter; For the distribution network system Active and reactive loads at each time step; yes The joint action space of all intelligent agents in time step, action space Indicates the first A photovoltaic inverter in The reactive power adjustment action of the time step; This represents the total number of photovoltaic inverters. It is a global reward function used to minimize voltage deviation and network losses while ensuring that the distribution network system complies with constraints; yes Time Step The per-unit voltage value of each node.

4. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 1, characterized in that: The sampling strategy of the deep reinforcement learning MGPER-MATD3 model is as follows: For each experience in the replay pool, calculate its corresponding priority; Normalize all priorities and calculate the sampling probability for each experience.

5. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 4, characterized in that: The formula for calculating the priority is: , In the formula, for Time step experience priority, yes Time step experience TD error, express Time step experience The rarity of an experience, i.e., the frequency of accessing that experience; express Time step experience Experience reward characteristic value; and These represent the weights, It is a constant. .

6. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 4, characterized in that: The Time step experience The formula for calculating the TD error is: , In the formula, yes Time step experience The instant reward received; It is a discount factor, and its value range is... This is used to measure the importance of future rewards; For experience exist Time step status Take action below Q-value estimation at that time; For experience exist Time step status Take action below Q-value estimation at that time.

7. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 4, characterized in that: The Time step experience The formula for calculating the rarity of experience is: , in: , In the formula, For experience Medium experience status The number of times it appears; It is experience The empirical cumulative distribution function of the access frequency of all states; This refers to the access frequency of other experiences in the experience replay pool; Total number of experiences; This is an indicator function; it returns 1 if the condition is true, and 0 otherwise.

8. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 4, characterized in that: The Time step experience The formula for calculating the experience reward feature value is: , In the formula, The reward value threshold, yes Time step experience The instant reward received.

9. The self-healing voltage optimization control method for distribution network disturbances based on multi-agent deep reinforcement learning according to claim 1, characterized in that: The step of solving the partitioned Markov decision model using the deep reinforcement learning MGPER-MATD3 model includes: Based on the MATD3 algorithm, the agent continuously interacts with the distribution network to collect experience and store it in the playback pool. Each experience includes the current state of the distribution network, the current action, the current reward, the next state, and the done flag. Once the amount of experience in the replay pool reaches a certain threshold, the policy update will begin. After each policy update, the TD error, experience rarity, and experience reward characteristics of all experiences are calculated, and the priority of experiences is updated to implement the MGPER mechanism.

10. A power distribution network disturbance self-healing voltage optimization control device based on multi-agent deep reinforcement learning, characterized in that, include: The modeling module is used to model the distributed distribution network voltage optimization problem as a partitioned Markov decision model. The state space of the partitioned Markov decision model includes the voltage amplitude and voltage phase angle of each node in the distribution network, the reactive power regulation and active power regulation of the photovoltaic inverter, and the active and reactive loads of the distribution network. The action space of the partitioned Markov decision model includes the reactive power regulation action of the photovoltaic inverter. The global reward function of the partitioned Markov decision model corresponds to the control objective of the distributed distribution network voltage optimization problem. The optimization module is used to solve the partitioned Markov decision model using the deep reinforcement learning MGPER-MATD3 model to obtain the reactive power of each photovoltaic inverter in the distributed distribution network and complete the distribution network voltage optimization. The deep reinforcement learning MGPER-MATD3 model samples according to the priority of each experience when performing experience sampling. The priority of each experience is related to the TD error, experience rarity, and experience reward feature value.

11. A self-healing voltage optimization control system for power distribution network disturbances based on multi-agent deep reinforcement learning, characterized in that, Including storage media and processor; The storage medium is used to store instructions; The processor is configured to operate according to the instructions to perform the method according to any one of claims 1-9.

Citation Information

Cited By

  • Distributed photovoltaic scheduling method based on multi-agent consensus optimization

    CN121436615A