Distributed photovoltaic inverter and avc system soft coordination power distribution network voltage control method

CN122801313APending Publication Date: 2026-09-22STATE GRID ANHUI ELECTRIC POWER CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611118604.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-27
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

现有研究多采用“硬协同”策略,即对AVC系统与逆变器统一建模并联合优化,但这种方式改变了AVC系统的原有控制规则,实施成本高、落地困难,并可能破坏多级AVC系统的协调关系

Benefits of technology

1、本发明在保持AVC系统原有闭环控制逻辑、九区图判据以及有载调压变压器OLTC和电容器组CBs动作规则不变的基础上,引入了分布式光伏逆变器参与配电网电压调节,不直接改变AVC系统的控制规则,也不直接替代AVC系统输出有载调压变压器OLTC和电容器组CBs控制指令。相比于现有对AVC系统和逆变器进行统一建模、联合优化的硬协同方式,本发明能够在不破坏既有AVC系统运行规程和多级协调关系的前提下实现协同控制,降低了系统改造成本,提高了工程应用的可实施性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801313A_ABST
    Figure CN122801313A_ABST
Patent Text Reader

Abstract

This invention discloses a distribution network voltage control method based on soft coordination between distributed photovoltaic inverters and AVC systems, comprising: 1. Establishing an AVC-in-the-loop operation model based on the nine-zone diagram method, maintaining the original rule-based control logic of the AVC system; 2. Constructing a soft-coordinated control mechanism between the AVC system and the distributed photovoltaic inverter, influencing the distribution network voltage state by adjusting the reactive power output of the inverter; 3. Modeling the control process of the distributed photovoltaic inverter as a Markov decision process, and designing a reward function that includes voltage deviation, network losses, and the number of AVC system equipment actions; 4. Training the inverter reactive power control strategy using a soft actor-commentator algorithm, and verifying it through distribution network examples and AVC-in-the-loop systems. This invention can achieve coordinated voltage control between distributed photovoltaic inverters and AVC systems without changing the original operating rules of the AVC system, thereby reducing distribution network losses, reducing the number of equipment actions, and improving voltage operation stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system automatic control and distributed energy access, and in particular to a distribution network voltage control method based on soft coordination between photovoltaic inverters and automatic voltage control (AVC) systems, belonging to the field of distribution network voltage-reactive power optimization control technology. Background Technology

[0002] Given the depletion of fossil fuels and global warming, the use of renewable energy has been considered a crucial step in reducing carbon emissions. Distributed photovoltaic (PV), as one of the main forms of solar energy utilization, has experienced rapid development due to its advantages such as renewability and zero pollution. With the large-scale integration of distributed energy sources into distribution systems, distributed photovoltaic (PV) exhibits characteristics such as "rapid growth in installed capacity" and "a significant increase in the proportion of medium- and low-voltage grid connection." However, the volatility and intermittency of renewable energy sources also pose significant challenges to the safe operation of distribution networks, leading to frequent problems such as voltage exceeding limits and voltage oscillations. How to effectively solve these problems has become an important issue for the safe operation of current distribution systems.

[0003] In medium- and low-voltage distribution networks, automatic voltage control (AVC) systems are typically based on SCADA platforms for centralized management of grid voltage. AVC systems generally use substation bus voltage qualification rates and equipment operating procedures as constraints, adjusting on-load tap changing transformers (OLTCs) and capacitor banks (CBs) to achieve comprehensive goals such as minimizing network losses and maximizing voltage qualification rates. Traditional AVC systems often employ a nine-zone diagram method for voltage-reactive power control, using zone-based adjustment strategies to ensure that voltage and reactive power remain within safe operating ranges.

[0004] However, in distribution systems with a high proportion of photovoltaic grid connection, the reverse power flow generated by numerous active devices complicates network power flow, making it difficult for traditional AVC systems relying solely on the regulation methods of on-load tap changers (OLTCs) and capacitor banks (CBs) at the power plant side to completely solve the voltage over-limit problem at some nodes. Furthermore, due to the limited number of mechanical operations of OLTCs and CBs, their regulation capabilities are insufficient to adapt to the frequent voltage fluctuations caused by high-proportion photovoltaic integration. In addition, the nine-zone diagram method relies on empirically set regional parameters, while the random fluctuations in renewable energy output increase the difficulty of parameter tuning, easily leading to frequent AVC system operations and affecting equipment lifespan.

[0005] Distributed photovoltaic inverters (PV inverters) possess the ability to smoothly regulate reactive power and have advantages such as fast response speed and flexible scheduling, thus they are widely used in voltage-reactive power control in distribution networks. Existing research mostly adopts a "hard collaboration" strategy, which involves modeling and jointly optimizing the AVC system and the inverter. However, this approach changes the original control rules of the AVC system, resulting in high implementation costs, difficulties in implementation, and potential disruption of the coordination relationship between multi-level AVC systems. Summary of the Invention

[0006] The purpose of this invention is to provide a distribution network voltage control method based on soft collaboration between photovoltaic inverters and AVC systems, in order to achieve coordinated voltage regulation between distributed inverters and AVC systems without changing the original operating rules of the AVC system, thereby reducing the number of equipment operations, reducing network losses and improving voltage stability.

[0007] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The present invention provides a distribution network voltage control method based on the soft coordination of distributed photovoltaic inverters and AVC systems, characterized by the following steps: Step 1: Obtain the current operating data of the distribution network, the operating data of the distributed photovoltaic inverters, and the equipment status data of the AVC system, and construct the current operating status of the distribution network based on the operating data and equipment status data; Step 2: Model the process of distributed photovoltaic inverters participating in distribution network voltage control as a Markov decision process, thereby constructing the action space and reward function of the distributed photovoltaic inverter reactive power control agent, which is used to conduct offline training on the distributed photovoltaic inverter reactive power control agent to obtain the trained distributed photovoltaic inverter reactive power control agent. Step 3: Input the current operating status of the distribution network into the trained distributed photovoltaic inverter reactive power control agent, and output the reactive power output ratio of each distributed photovoltaic inverter at the current time to generate reactive power control commands for each distributed photovoltaic inverter at the current time. Step 4: After each distributed photovoltaic inverter executes the reactive power control command at the current moment, it performs power flow calculation or real-time measurement on the distribution network to obtain the power flow status of the distribution network after reactive power adjustment by the distributed photovoltaic inverter at the current moment, including: the bus voltage on the low-voltage side of the main transformer, the power factor on the high-voltage side, and the voltage amplitude of each node. Step 5: Based on the low-voltage bus voltage and high-voltage power factor of the main transformer after reactive power regulation by the distributed photovoltaic inverter at the current moment, the AVC system determines the operating area of ​​the AVC system at the current moment according to the nine-zone diagram method. This is used to generate the tap adjustment action of the on-load tap changer (OLTC) and the switching action of the capacitor bank (CBs) at the current moment. Step 6: Based on the tap changer adjustment action of the on-load tap changer (OLTC) and the switching action of the capacitor bank (CBs) at the current moment, update the equipment status data of the AVC system at the current moment, and combine it with the reactive power, load fluctuation and photovoltaic output fluctuation of the distributed photovoltaic inverter at the current moment to obtain the operating status of the distribution network at the next moment. Step 7: Take the operating state of the distribution network at the next moment as the new operating state of the distribution network at the current moment, and return to Step 3 for rolling control, thereby realizing the distribution network voltage control of soft collaboration between distributed photovoltaic inverters and AVC system.

[0008] The characteristic of the distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems described in this invention is that step one includes: Step 1.1, obtain the first Each node at the current moment active power reactive power and voltage amplitude Thus constructing the current moment The set of active power of nodes Current moment The set of nodal reactive power and the current moment Set of node voltage amplitudes ,in, Indicates the number of nodes in the distribution network. ; Step 1.2: Obtain the current time value of the on-load tap-changing transformer (OLTC). tap position The capacitor bank CBs at the current time Number of throw groups On-load tap-changing transformer (OLTC) at the current moment Time interval from the last action The capacitor bank CBs at the current time Time interval from the last action On-load tap-changing transformers (OLTCs) as of the current time Number of actions As of the current time, the capacitor bank CBs The number of actions is Thus, the distribution network is constructed at the current moment. Operating status .

[0009] Furthermore, the action space in step two includes: Step 2.1: Assume the number of distributed photovoltaic inverters is... , No. At the current moment, a distributed photovoltaic inverter The reactive power output ratio is Thus constructing the current moment action ;in, , ; Step 2.2, let the first... The apparent capacity of a distributed photovoltaic inverter is At the current moment The active power is Therefore, the first equation can be used to calculate the second equation. At the current moment, a distributed photovoltaic inverter upper limit of reactive power ; (1) Step 2.3: Based on the upper limit of reactive power Use equation (2) to determine the first At the current moment, a distributed photovoltaic inverter Actual reactive power The output range; (2) Step 2.4, according to the first At the current moment, a distributed photovoltaic inverter reactive power output ratio The actual reactive power is calculated using equation (3). Thus, the current time is obtained. set of reactive power control quantities ; (3) Furthermore, the reward function in step two includes: Step 3.1: Let the set of branches in the distribution network be... , No. The node and the first Branches between nodes At the present moment The branch current is branch road The branch resistance is Thus, equation (4) is used to obtain the current state of the distribution network at the current time. Reward function for network loss : (4) Step 3.2, set the upper limit of the node voltage as follows: The lower limit of the node voltage is And based on the j-th node at the current time voltage amplitude Calculate the first using equation (5) Each node at the current moment The voltage exceeds the limit ; (5) Step 3.3: Calculate the current state of the distribution network using equation (6). Reward function for the voltage over-limit portion : (6) Step 3.4: Let the end time of a single round be... The maximum number of operations allowed for an on-load tap-changing transformer (OLTC) is [number]. The minimum allowable number of operations for an on-load tap-changing transformer (OLTC) is [number]. The maximum allowed number of operations for capacitor banks (CBs) is [number]. The lower limit of the number of operations allowed for the capacitor bank CBs is The operating penalty coefficient for an on-load tap-changing transformer (OLTC) is: The penalty coefficient for the action of the capacitor bank (CBs) is Therefore, equation (7) and equation (8) are used to calculate the distribution network at the current time. The reward function for the penalty part of the on-load tap-changing transformer (OLTC) operation count. The reward function for the penalty part of the number of actions of capacitor banks. : (7) (8) Step 3.5: Construct a distributed photovoltaic inverter reactive power control intelligent agent at the current moment using equation (9). reward function : (9) In equation (9), and These are the weighting coefficients for the network loss term and the voltage over-limit term, respectively.

[0010] Furthermore, in step two, the offline training of the reactive power control agent for the distributed photovoltaic inverter includes: Step 4.1: Constructing the reactive power control intelligent agent for distributed photovoltaic inverters includes: policy network. First Evaluation Network Second evaluation network First objective evaluation network Second objective evaluation network ,in, For the parameters of the policy network, and These are the parameters of the first evaluation network and the parameters of the second evaluation network, respectively. and These are the parameters of the first objective evaluation network and the second objective evaluation network, respectively. Step 4.2, Set the experience replay pool as follows and the number of samples in the small batch is ;initialization ; Step 4.3: Set the current time in a single round. Distribution network operating status Input Policy Network Process the data and output the current time. The action distribution parameters are sampled to obtain the current time. action ; Step 4.4: Based on the current time action Calculate the current state of each distributed photovoltaic inverter. The actual reactive power, and the current moment The actual reactive power input is processed in the power flow calculation model of the distribution network to obtain the current reactive power of the distributed photovoltaic inverter. Power flow status of the distribution network after reactive power regulation; Step 4.5: The AVC system determines the current status of the distributed photovoltaic inverter based on the data from the current time. The power flow state of the distribution network after reactive power regulation is generated according to the nine-zone diagram method, showing the AVC system at the current moment. The control actions are used to update the current time of the on-load tap changer (OLTC). The tap position and the number of capacitor banks (CBs) switched on are used to obtain the next moment. Distribution network operating status ; Step 4.6: Calculate the current time. Reward value and with , and Together they constitute the current moment samples Then, it is stored in the experience replay pool. Middle; General Assign to Then, return to step 4.3 until the experience replay pool is reached. Until the number of samples reaches the preset value; Step 4.7, from the experience replay pool Randomly selected 1 sample; of which, the first bar sample ,in, For the first The state of the sample For the first The actions of the sample For the first Rewards for each sample For the first The next state of the sample ; Step 4.8, Input Policy Network Processing and sampling are performed to obtain the first... The next action of the sample Therefore, the first equation (10) is used to calculate the second equation. Target evaluation value of each sample : (10) In equation (10), Discount factor; Step 4.9: Construct the first... using equation (11) A rating network loss function and minimize the loss function To achieve the goal, update the first evaluation network separately. Second evaluation network : (11) Step 4.10: Construct a policy network using equation (12). loss function Parameters used to update the policy network : (12) In equation (12), Indicates the updated number An evaluation network; The coefficient of the entropy regularization term; Step 4.11: Use equation (13) to evaluate the parameters of the network for the first objective. Parameters of the second objective evaluation network Perform a soft update: (13) In equation (13), For the soft update coefficients of the target network; For assignment; For the first The parameters of the target evaluation network; Step 4.12: Repeat steps 4.3 to 4.11 to perform offline training on the reactive power control agent of the distributed photovoltaic inverter until the maximum number of training rounds is reached or the reward convergence condition is met, and the trained reactive power control agent of the distributed photovoltaic inverter is obtained.

[0011] Furthermore, step three involves calculating the reactive power output ratio of each distributed photovoltaic inverter at the current moment. The actual reactive power, formed at the current moment set of reactive power control quantities ; and thus according to Send the current time to each distributed photovoltaic inverter The reactive power control command causes each distributed photovoltaic inverter to execute the current time... Reactive power regulation.

[0012] Furthermore, step five includes: Step 6.1: Assume the upper limit of the bus voltage on the low-voltage side of the main transformer is... The lower limit of the bus voltage on the low-voltage side of the main transformer is The upper limit of the power factor on the high-voltage side is The lower limit of the power factor on the high-voltage side is ; Step 6.2: Construct a nine-zone operating plane for the AVC system, using the low-voltage bus voltage of the main transformer as the ordinate and the high-voltage power factor as the abscissa, and then... , , and The nine-zone map operating plane is divided into nine operating areas; Step 6.3: Based on the current state of the distributed photovoltaic inverter... Voltage of the low-voltage side bus of the main transformer after reactive power regulation and high voltage side power factor Determine the AVC system at the current moment. Operating area ; Step 6.4: Generate the AVC system at the current moment. Control actions ;in, This indicates the current time of the on-load tap-changing transformer (OLTC). The tap adjustment amount, This indicates the capacitor bank CBs at the current time. The amount of cutting adjustment; When the time interval between the last operation of the on-load tap changer (OLTC) or capacitor bank (CBs) is less than the corresponding minimum operating interval, the corresponding equipment is prohibited from operating, and the corresponding adjustment amount is set to 0. when In Inner and In At that time, the order and At the current moment No tap change or switching operations are performed to maintain the on-load tap changer (OLTC) at the current time. The tap position and the number of capacitor banks (CBs) switched on remain unchanged; when Below At that time, the AVC system is in the current moment First, execute the throwing and cutting action, that is, according to The number of capacitor banks (CBs) switched on and off is determined at the current time. Then perform the tap adjustment action, that is, according to Adjust the tap position of the on-load tap changer (OLTC) to increase the bus voltage; when Higher than At that time, the AVC system is in the current moment First, perform the tap adjustment action, that is, according to Lowering the tap position of the on-load tap changer (OLTC) if... Still higher And then at the current moment Execute the switching adjustment action, that is, according to The number of capacitor banks switched on and off is reduced to lower the bus voltage; when Below or higher At that time, according to Generate capacitor banks CBs at the current time The switching action, and the on-load tap-changing transformer OLTC at the current moment. The tap adjustment action together constitutes the AVC system at the current moment. Control actions .

[0013] Furthermore, step six includes: Step 7.1: Use equation (14) to obtain the distribution network at the next moment. Tap position of on-load tap changing transformer (OLTC) and the number of switching groups of capacitor banks (CBs) : (14) Step 7.2, according to , , Load fluctuations and photovoltaic output fluctuations are used to perform power flow calculations on the distribution network to obtain the distribution network's power flow at the next time step. The set of active power of nodes Node reactive power set and node voltage amplitude set ; thereby constructing the next moment Distribution network operating status .

[0014] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program supporting the processor in performing the method described therein, and the processor is configured to execute the program stored in the memory.

[0015] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program is executed by a processor to perform the steps of the method described thereon.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention, while maintaining the original closed-loop control logic, nine-zone diagram criteria, and operating rules of the on-load tap changer (OLTC) and capacitor bank (CBs) of the AVC system, introduces distributed photovoltaic inverters to participate in distribution network voltage regulation. It does not directly change the control rules of the AVC system, nor does it directly replace the AVC system's output control commands for the OLTC and capacitor bank (CBs). Compared to existing hard-coordination methods that unify modeling and jointly optimize the AVC system and inverters, this invention achieves collaborative control without disrupting the existing AVC system operating procedures and multi-level coordination relationships, reducing system modification costs and improving the feasibility of engineering applications.

[0017] 2. This invention utilizes the fast response speed and continuously adjustable reactive power of distributed photovoltaic inverters. By adjusting the reactive power output of the inverters, it changes the voltage at distribution network nodes and the operating status of power plants, enabling the AVC system to generate more reasonable voltage regulation and switching actions under the original nine-zone control logic. This alleviates the problem of frequent operations and high regulation pressure in traditional AVC systems that rely solely on on-load tap changers (OLTCs) and capacitor banks (CBs) to cope with high-proportion photovoltaic fluctuations. It reduces the number of operations of OLTCs and CBs, lowers mechanical equipment wear, and extends equipment lifespan.

[0018] 3. This invention models the reactive power and voltage control process of distributed photovoltaic inverters as a Markov decision process, and uses the power state of the distribution network, the voltage state of nodes, and the state of AVC equipment as the agent's state information. Network losses, voltage deviation, and the number of AVC equipment actions are all incorporated into the reward function. Through this design, the agent can simultaneously consider voltage safety, operational economy, and equipment action costs when learning control strategies, thereby improving the problem that traditional voltage-reactive power control methods struggle to simultaneously coordinate voltage quality, network losses, and equipment lifespan, and enhancing the overall effectiveness of distribution network voltage control.

[0019] 4. This invention employs deep reinforcement learning to train the reactive power control strategy of distributed photovoltaic inverters, enabling the control strategy to adaptively adjust the inverter's reactive power output based on photovoltaic power output fluctuations, load changes, and AVC device status. Compared to control methods relying on fixed rules or static optimization models, this invention enhances the distribution network's responsiveness to source-load uncertainties and rapid voltage fluctuations, reduces distribution network losses, minimizes node voltage offset, and improves voltage qualification rate and operational stability. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of an automatic voltage control system; Figure 2 It is a nine-zone map; Figure 3 This is a schematic diagram of the power distribution network topology according to an embodiment of the present invention; Figure 4 This is a test result diagram of the AVC system (transformer 2); Figure 5 This is a test result diagram of the AVC system (transformer 3); Figure 6 This is a comparison chart of the training reward curves of different reinforcement learning algorithms. Detailed Implementation

[0021] In this embodiment, a distribution network with distributed photovoltaic (PV) access is used as the object. A distribution network voltage control method with soft collaboration between distributed PV inverters and an AVC system is constructed. In this method, the AVC system generates control commands based on the collected information such as the high-voltage side power factor, low-voltage side bus voltage, tap position, and the number of capacitor banks switched on and off. While maintaining the original nine-zone diagram control logic, equipment action rules, and closed-loop operation mode of the AVC system, before the AVC system takes action, the reactive power control agent of the distributed PV inverter is first trained to adjust the voltage of each distributed PV inverter. The reactive power output of the inverter changes the voltage-reactive power operating state input to the AVC system, causing favorable changes in the distribution network voltage state and the substation side operating state. The AVC system then generates tap change actions for the on-load tap changer (OLTC) and switching actions for the capacitor banks (CBs) according to its existing rules. This achieves a new "soft coordination" control strategy that coordinates with the distributed photovoltaic (PV) inverter without altering the original operating logic of the AVC system. This enables soft coordination of distribution network voltage control between the distributed PV inverter and the AVC system, improving the flexibility and economy of voltage control. Therefore, "soft coordination" means that the distributed PV inverter does not directly replace the AVC system in issuing commands to the OLTC or capacitor banks, but rather indirectly guides the AVC system to make more reasonable control actions within its existing logic through reactive power output. Specifically, for example... Figure 1 As shown, the method includes the following steps: Step 1: Obtain the current operating data of the distribution network, the operating data of the distributed photovoltaic inverters, and the equipment status data of the AVC system, and construct the current operating status of the distribution network based on the operating data and equipment status data.

[0022] In practice, distribution network operation data can be obtained from SCADA systems, distribution automation master stations, smart terminals, power flow simulation programs, or real-time measurement devices; the operation data of distributed photovoltaic inverters should at least include the inverter's current active power, available capacity, or apparent capacity; the equipment status data of the AVC system should at least include the tap position of the on-load tap changer (OLTC), the number of capacitor banks (CBs) switched on and off, the time interval since the last action, and the cumulative number of actions.

[0023] Step 1.1, obtain the first Each node at the current moment active power reactive power and voltage amplitude Thus constructing the current moment The set of active power of nodes Current moment The set of nodal reactive power and the current moment Set of node voltage amplitudes ,in, Indicates the number of nodes in the distribution network. .

[0024] Step 1.2: Obtain the current time value of the on-load tap-changing transformer (OLTC). tap position The capacitor bank CBs at the current time Number of throw groups On-load tap-changing transformer (OLTC) at the current moment Time interval from the last action The capacitor bank CBs at the current time Time interval from the last action On-load tap-changing transformers (OLTCs) as of the current time Number of actions As of the current time, the capacitor bank CBs The number of actions is Thus, the distribution network is constructed at the current moment. Operating status .

[0025] Will , , and The reason for including the operating status is that the on-load tap changer (OLTC) and capacitor banks (CBs) in the AVC system are devices with constraints on operating intervals, lifetimes, and number of operations. If the agent only observes the node voltage and ignores the device status, it may learn a strategy that frequently relies on mechanical device regulation. By explicitly adding the above device status, the reactive power control agent of the distributed photovoltaic inverter can perceive the cost of AVC device operations during training, and thus tend to prioritize the use of the continuous reactive power regulation capability of the distributed photovoltaic inverter within feasible limits.

[0026] In a specific example, such as Figure 3 As shown, this embodiment employs a 33-node test system, which includes 5 distributed photovoltaic (PV) access nodes. To reduce the state dimensionality, the distribution network is divided into 5 sub-regions based on the PV node locations and feeder structure. The sum of the active and reactive power of the nodes within each sub-region is used as the region's power state, and the maximum and minimum voltage values ​​within each sub-region are used as the region's voltage state. This regionalization process retains voltage exceedance risk information while reducing the input dimensionality of the neural network, thus improving offline training efficiency.

[0027] Step 2: Model the process of distributed photovoltaic inverters participating in distribution network voltage control as a Markov decision process, thereby constructing the action space and reward function of the distributed photovoltaic inverter reactive power control agent, which is used to conduct offline training on the distributed photovoltaic inverter reactive power control agent to obtain the trained distributed photovoltaic inverter reactive power control agent.

[0028] In the Markov decision-making process, the state is the distribution network operating state. The action is the reactive power output ratio of the distributed photovoltaic inverter. The reward is a reward function that comprehensively considers network loss, voltage over-limit, and the number of actions of AVC system equipment. The state transition is jointly determined by "distributed photovoltaic inverter reactive power regulation - distribution network power flow calculation or real-time measurement - AVC system nine-zone diagram action - load and photovoltaic fluctuation update". The advantage of this modeling is that the agent does not need to rewrite the AVC system rules, but rather learns how to adjust the inverter reactive power output before the original response of the AVC system.

[0029] Step 2.1: Assume the number of distributed photovoltaic inverters is... , No. At the current moment, a distributed photovoltaic inverter The reactive power output ratio is Thus constructing the current moment action ;in, , .

[0030] Step 2.2, let the first... The apparent capacity of a distributed photovoltaic inverter is At the current moment The active power is Therefore, the first equation can be used to calculate the second equation. At the current moment, a distributed photovoltaic inverter upper limit of reactive power ; (1) Step 2.3: Based on the upper limit of reactive power Use equation (2) to determine the first At the current moment, a distributed photovoltaic inverter Actual reactive power The output range; (2) Step 2.4, according to the first At the current moment, a distributed photovoltaic inverter reactive power output ratio The actual reactive power is calculated using equation (3). Thus, the current time is obtained. set of reactive power control quantities ; (3) Step 3: Construct the reward function, including: network loss part, voltage over-limit part, penalty part for the number of times the on-load tap changer (OLTC) operates, and penalty part for the number of times the capacitor bank (CBs) operates.

[0031] The reward function guides the reactive power control agent of the distributed photovoltaic inverter to learn a control strategy that "ensures voltage safety, reduces operating losses, and minimizes mechanical actions in the AVC system." Since reinforcement learning typically aims to maximize cumulative rewards, this embodiment sets network losses, voltage exceedances, and equipment action deviations as negative rewards, enabling the agent to actively avoid these unfavorable operating states during training.

[0032] Step 3.1: Use equation (4) to obtain the current time of the distribution network. Reward function for network loss : (4) In equation (4), let the set of branches in the distribution network be... , No. The node and the first Branches between nodes At the present moment The branch current is branch road The branch resistance is .

[0033] Step 3.2: Calculate the first step using equation (5). Each node at the current moment The voltage exceeds the limit ; (5) In equation (5), the upper limit of the node voltage is assumed to be... The lower limit of the node voltage is And based on the j-th node at the current time voltage amplitude .

[0034] Step 3.3: Calculate the current state of the distribution network using equation (6). Reward function for the voltage over-limit portion : (6) Step 3.4: Let the end time of a single round be... The maximum number of operations allowed for an on-load tap-changing transformer (OLTC) is [number]. The minimum allowable number of operations for an on-load tap-changing transformer (OLTC) is [number]. The maximum allowed number of operations for capacitor banks (CBs) is [number]. The lower limit of the number of operations allowed for the capacitor bank CBs is The operating penalty coefficient for an on-load tap-changing transformer (OLTC) is: The penalty coefficient for the action of the capacitor bank (CBs) is Therefore, equation (7) and equation (8) are used to calculate the distribution network at the current time. The reward function for the penalty part of the on-load tap-changing transformer (OLTC) operation count. The reward function for the penalty part of the number of actions of capacitor banks. : (7) (8) The purpose of setting upper and lower limits for the number of actions is to prevent AVC system equipment from over-operating or losing its necessary regulation capability due to excessive suppression. If the number of OLTC or CBs actions exceeds the upper limit, it indicates that the mechanical equipment is operating too frequently, which can easily increase equipment wear; if the number of actions is below the lower limit, it may indicate insufficient voltage regulation or that the regulation responsibility has been unreasonably transferred to the distributed photovoltaic inverter. Therefore, interval-type penalties can help the soft coordination strategy achieve a balance between "protecting AVC equipment" and "ensuring voltage regulation capability".

[0035] Step 3.5: Construct a distributed photovoltaic inverter reactive power control intelligent agent at the current moment using equation (9). reward function : (9) In equation (9), and These are the weighting coefficients for the network loss term and the voltage over-limit term, respectively.

[0036] Step 4: Offline training of the reactive power control agent for the distributed photovoltaic inverter; This embodiment uses the soft actor-critic algorithm to train the reactive power control agent for the distributed photovoltaic inverter. This algorithm is suitable for continuous action spaces, can directly output the reactive power output ratio of each distributed photovoltaic inverter, and enhances the policy exploration capability through entropy regularization to avoid the policy from getting trapped in local optima too early.

[0037] Step 4.1: Constructing the reactive power control intelligent agent for distributed photovoltaic inverters includes: policy network. First Evaluation Network Second evaluation network First objective evaluation network Second objective evaluation network ,in, For the parameters of the policy network, and These are the parameters of the first evaluation network and the parameters of the second evaluation network, respectively. and These are the parameters of the first objective evaluation network and the second objective evaluation network, respectively.

[0038] Step 4.2, Set the experience replay pool as follows and the number of samples in the small batch is ;initialization ; Step 4.3: Set the current time in a single round. Distribution network operating status Input Policy Network Process the data and output the current time. The action distribution parameters are sampled to obtain the current time. action .

[0039] Step 4.4: Based on the current time action Calculate the current state of each distributed photovoltaic inverter. The actual reactive power, and the current moment The actual reactive power input is processed in the power flow calculation model of the distribution network to obtain the current reactive power of the distributed photovoltaic inverter. Power flow status of the distribution network after reactive power regulation; Step 4.5: The AVC system determines the current status of the distributed photovoltaic inverter based on the data from the current time. The power flow state of the distribution network after reactive power regulation is generated according to the nine-zone diagram method, showing the AVC system at the current moment. The control actions are used to update the current time of the on-load tap changer (OLTC). The tap position and the number of capacitor banks (CBs) switched on are used to obtain the next moment. Distribution network operating status .

[0040] Step 4.6: Calculate the current time. Reward value and with , and Together they constitute the current moment samples Then, it is stored in the experience replay pool. Middle; General Assign to Then, return to step 4.3 until the experience replay pool is reached. Until the number of samples reaches the preset value; In this embodiment, a control cycle of 5 minutes and a training round of 24 hours can be used, resulting in 288 control moments per round. This time granularity reflects the changes in photovoltaic output and load within a day, while also matching the requirements for equipment operation intervals in the actual operation of the AVC system.

[0041] Step 4.7, from the experience replay pool Randomly selected 1 sample; of which, the first bar sample ,in, For the first The state of the sample For the first The actions of the sample For the first Rewards for each sample For the first The next state of the sample .

[0042] Step 4.8, Input Policy Network Processing and sampling are performed to obtain the first... The next action of the sample Therefore, the first equation (10) is used to calculate the second equation. Target evaluation value of each sample : (10) In equation (10), This is the discount factor.

[0043] Step 4.9: Construct the first... using equation (11) A rating network loss function and minimize the loss function To achieve the goal, update the first evaluation network separately. Second evaluation network : (11) Step 4.10: Construct a policy network using equation (12). loss function Parameters used to update the policy network : (12) In equation (12), Indicates the updated number An evaluation network; is the coefficient of the entropy regularization term.

[0044] Step 4.11: Use equation (13) to evaluate the parameters of the network for the first objective. Parameters of the second objective evaluation network Perform a soft update: (13) In equation (13), For the soft update coefficients of the target network; For assignment; For the first The parameters of the target evaluation network; Step 4.12: Repeat steps 4.3 to 4.11 to perform offline training on the reactive power control agent of the distributed photovoltaic inverter until the maximum number of training rounds is reached or the reward convergence condition is met, and the trained reactive power control agent of the distributed photovoltaic inverter is obtained.

[0045] Step 5: Input the current operating status of the distribution network into the trained distributed photovoltaic inverter reactive power control agent, and output the reactive power output ratio of each distributed photovoltaic inverter at the current moment. This information is used to generate reactive power control commands for each distributed photovoltaic inverter at the current moment, thereby realizing online rolling control.

[0046] When running online, the reactive power control agent of the distributed photovoltaic inverter receives the current state. Post-output Then convert according to equations (1) to (3) to get The system then sends reactive power control commands to each distributed photovoltaic inverter via the inverter communication interface or the dispatch control system. These commands can either absorb or generate reactive power, with the specific direction determined by positive or negative signs, and can be configured according to the agreed-upon positive reactive power direction in the project.

[0047] Step 6: After each distributed photovoltaic inverter executes the reactive power control command at the current moment, it performs power flow calculation or real-time measurement on the distribution network to obtain the power flow status of the distribution network after reactive power adjustment by the distributed photovoltaic inverter at the current moment, including: the bus voltage on the low-voltage side of the main transformer, the power factor on the high-voltage side, and the voltage amplitude of each node.

[0048] The AVC system determines the operating zone of the main transformer at the current moment based on the low-voltage bus voltage and high-voltage power factor of the main transformer after reactive power regulation by the distributed photovoltaic inverter, according to the nine-zone diagram method. This zone is used to generate the tap changer operation of the on-load tap changer (OLTC) and the switching operation of the capacitor bank (CBs) at the current moment.

[0049] like Figure 2 As shown, the nine-zone diagram uses the low-voltage bus voltage of the main transformer as the vertical axis and the high-voltage power factor as the horizontal axis, dividing the system into nine operating zones based on the upper and lower limits of voltage and power factor. When a device is located in the middle zone, it indicates that both voltage and power factor meet the operating requirements, and the AVC system does not need to operate. When located in other zones, the AVC system determines, based on the zone location, whether to adjust the OLTC, switch CBS, or simultaneously adjust both types of equipment.

[0050] Step 6.1: Assume the upper limit of the bus voltage on the low-voltage side of the main transformer is... The lower limit of the bus voltage on the low-voltage side of the main transformer is The upper limit of the power factor on the high-voltage side is The lower limit of the power factor on the high-voltage side is ; Step 6.2: Construct a nine-zone operating plane for the AVC system, using the low-voltage bus voltage of the main transformer as the ordinate and the high-voltage power factor as the abscissa, and then... , , and The nine-zone map operating plane is divided into nine operating areas.

[0051] Step 6.3: Based on the current state of the distributed photovoltaic inverter... Voltage of the low-voltage side bus of the main transformer after reactive power regulation and high voltage side power factor Determine the AVC system at the current moment. Operating area ; Step 6.4: Generate the AVC system at the current moment. Control actions ;in, This indicates the current time of the on-load tap-changing transformer (OLTC). The tap adjustment amount, This indicates the capacitor bank CBs at the current time. The amount of cutting adjustment.

[0052] When the time interval between the last operation of the on-load tap changer (OLTC) or capacitor bank (CBs) is less than the corresponding minimum operating interval, the corresponding equipment is prohibited from operating, and the corresponding adjustment amount is set to 0. when In Inner and In At that time, the order and At the current moment No tap change or switching operations are performed to maintain the on-load tap changer (OLTC) at the current time. The tap position and the number of capacitor banks (CBs) switched on and off remain unchanged.

[0053] when Below At that time, the AVC system is currently... First, execute the throwing and cutting action, that is, according to The number of capacitor banks (CBs) switched on and off is determined at the current time. Then perform the tap adjustment action, that is, according to Adjust the tap position of the on-load tap changer (OLTC) to increase the bus voltage.

[0054] when Higher than At that time, the AVC system is currently... First, perform the tap adjustment action, that is, according to Lowering the tap position of the on-load tap changer (OLTC) if... Still higher And then at the current moment Execute the switching adjustment action, that is, according to The number of capacitor banks switched off is reduced to lower the bus voltage.

[0055] when Below or higher At that time, according to Generate capacitor banks CBs at the current time The switching action, and the on-load tap-changing transformer OLTC at the current moment. The tap adjustment action together constitutes the AVC system at the current moment. Control actions .

[0056] Step 7: Based on the tap changer operation of the on-load tap changer (OLTC) and the switching operation of the capacitor bank (CBs) at the current moment, update the equipment status data of the AVC system at the current moment, and combine it with the reactive power, load fluctuation, and photovoltaic output fluctuation of the distributed photovoltaic inverter at the current moment to obtain the operating status of the distribution network at the next moment.

[0057] Step 7.1: Use equation (14) to obtain the distribution network at the next moment. Tap position of on-load tap changing transformer (OLTC) and the number of switching groups of capacitor banks (CBs) : (2) Step 7.2, according to , , Load fluctuations and photovoltaic output fluctuations are used to perform power flow calculations on the distribution network to obtain the distribution network's power flow at the next time step. The set of active power of nodes Node reactive power set and node voltage amplitude set ; thereby constructing the next moment Distribution network operating status .

[0058] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.

[0059] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.

[0060] Example verification and effect description; To verify the feasibility and effectiveness of the present invention, this embodiment constructs as follows: Figure 3 The 33-node test system shown uses a 5-minute control time granularity and a 24-hour training cycle; and sets up a typical daily operation scenario including distributed photovoltaic access.

[0061] like Figure 4 and Figure 5 As shown, when only the AVC system is used for voltage-reactive power regulation without coordinated control of distributed photovoltaic inverters, the low-voltage bus voltage and high-voltage power factor of the main transformer can be controlled within the preset range. However, the on-load tap-changing transformer (OLTC) and capacitor banks (CBs) operate frequently throughout the day, especially during periods of rapid changes in photovoltaic output or load ramp-up, easily leading to frequent tap changes or switching. This result indicates that relying solely on the AVC system can maintain the plant-side indicators, but it is difficult to fully suppress the operational pressure on equipment under high-proportion photovoltaic integration.

[0062] In training the reactive power control strategy for distributed photovoltaic inverters, this embodiment compares five algorithms: the Soft Actor-Commentator Algorithm (SAC), the Independent Deep Deterministic Policy Gradient Algorithm (IDDPG), the Shapley Value-Based Deep Deterministic Policy Gradient Algorithm (SQDDPG), the Multi-Agent Dual-Delay Deep Deterministic Policy Gradient Algorithm (MATD3), and the Multi-Agent Deep Deterministic Policy Gradient Algorithm (MADDPG). Figure 6 As shown, the training reward curves of different reinforcement learning algorithms are significantly different. Among them, the SAC algorithm can obtain a higher and more stable cumulative reward after training convergence, indicating that it can achieve a good comprehensive balance between voltage deviation, network loss and the number of AVC device actions.

[0063] 1) The Soft Actor-Critic (SAC) algorithm is a stochastic policy gradient algorithm for off-track policies. Based on the maximum entropy reinforcement learning framework, it considers both reward maximization and policy entropy maximization during policy optimization, thus balancing exploration ability and control performance in the continuous action space. It is suitable for the continuous reactive power regulation scenario of distributed photovoltaic inverters.

[0064] 2) The Independent Deep Deterministic Policy Gradient (IDDPG) algorithm is a decentralized multi-agent deep reinforcement learning algorithm. In this algorithm, each agent is trained using the deep deterministic policy gradient method and outputs control actions mainly based on its own local observation information. It is suitable for distributed control scenarios where each distributed photovoltaic inverter makes independent decisions.

[0065] 3) The Shapley Q-value Deep Deterministic Policy Gradient (SQDDPG) algorithm is an algorithm that introduces cooperative game theory on the basis of the deep deterministic policy gradient algorithm. It evaluates the contribution of different agents to the global control objective through Shapley value, thereby improving the contribution allocation problem in the multi-agent cooperative training process.

[0066] 4) The Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) algorithm is an extension of the Twin Delayed Deep Deterministic Policy Gradient algorithm for multi-agent control scenarios. This algorithm reduces the risk of overestimation of evaluation values ​​and improves the stability of continuous action control policy training through mechanisms such as dual evaluation networks and delayed policy updates.

[0067] 5) The Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm is a multi-agent deep reinforcement learning algorithm that employs a centralized training and distributed execution framework. During the training phase, the algorithm uses the state and action information of multiple agents for centralized evaluation; during the execution phase, each agent independently outputs control actions based on its own observation information, making it suitable for the collaborative control problem of multiple distributed photovoltaic inverters.

[0068] Based on the combined training reward curve and test results, the SAC algorithm exhibits the best overall performance in this voltage-reactive power control scenario. Its maximum entropy learning mechanism and stochastic strategy enable it to better balance exploration and utilization, achieve effective collaboration with the AVC system, and minimize network losses and the number of control actions while ensuring voltage quality.

[0069] In addition to the training data, a typical day was selected to test five different algorithms, and their voltage-reactive power control effect and soft coordination effect were calculated. The voltage-reactive power control effect was measured by total grid loss, average voltage deviation, voltage qualification rate, and average power factor. The test results are shown in Table 1. To measure the soft coordination effect between the distributed photovoltaic inverter and the AVC system, the number of on-load tap changer (OLTC) operations and capacitor bank (CBs) operations in the AVC system, as well as the total reactive power output and average current of the distributed photovoltaic inverter, were measured. The test results are shown in Table 2. Furthermore, this invention sets up a scenario where voltage-reactive power regulation is performed solely by the AVC system, and compares this scenario with the various algorithms.

[0070] Table 1 Comparison of Voltage-Reactive Power Control Effects Table 2 Comparison of Synergistic Effects As shown in Table 1, the total network loss corresponding to the SAC algorithm is 36.96MW, which is a reduction of approximately 39.23% compared to 60.82MW in the AVC-only scenario; the average voltage deviation decreased from 0.0392pu in the AVC-only scenario to 0.0237pu, a reduction of approximately 39.54%; and the voltage qualification rate increased from 0.738 to 0.997. This demonstrates that the present invention, through the soft coordination of reactive power regulation of distributed photovoltaic inverters and the original control logic of the AVC system, can significantly improve voltage quality and reduce network losses.

[0071] As shown in Table 2, under the SAC algorithm control, the OLTC operates 9 times and the CBs operate 10 times, totaling 19 times; in the AVC scenario alone, the total number of operations is 28. Therefore, the total number of operations in the AVC system is reduced by approximately 32.14% after adopting the method of this invention. Meanwhile, the average current of the SAC algorithm is 20.97A, which is about 8.23% lower than the 22.85A in the AVC scenario alone. This indicates that the method reduces mechanical equipment operations without excessively increasing inverter current to achieve control effects, demonstrating good operational economy.

[0072] In summary, this embodiment demonstrates that, without altering the original nine-zone diagram criteria, the OLTC tap changer adjustment rules, and the CBs switching rules of the AVC system, the present invention enables the AVC system to achieve a better input state under the original control logic by using a trained distributed photovoltaic inverter reactive power control agent to perform rolling adjustment of the inverter's reactive power output. This reduces the number of equipment actions, lowers distribution network losses, reduces node voltage deviation, and improves voltage operation stability.

[0073] It should be noted that the 33-node test system, 5 photovoltaic nodes, 5-minute sampling period, 24-hour single round, and specific numerical results described above are all examples illustrating the implementation of this invention. This invention is also applicable to distribution network voltage control scenarios with other node sizes, other distributed photovoltaic access locations, other AVC equipment configurations, and other load / photovoltaic forecast data conditions.

Claims

1. A distribution network voltage control method for soft coordination between distributed photovoltaic inverters and AVC systems, characterized in that, The procedure is as follows: Step 1: Obtain the current operating data of the distribution network, the operating data of the distributed photovoltaic inverters, and the equipment status data of the AVC system, and construct the current operating status of the distribution network based on the operating data and equipment status data; Step 2: Model the process of distributed photovoltaic inverters participating in distribution network voltage control as a Markov decision process, thereby constructing the action space and reward function of the distributed photovoltaic inverter reactive power control agent, which is used to conduct offline training on the distributed photovoltaic inverter reactive power control agent to obtain the trained distributed photovoltaic inverter reactive power control agent. Step 3: Input the current operating status of the distribution network into the trained distributed photovoltaic inverter reactive power control agent, and output the reactive power output ratio of each distributed photovoltaic inverter at the current time to generate reactive power control commands for each distributed photovoltaic inverter at the current time. Step 4: After each distributed photovoltaic inverter executes the reactive power control command at the current moment, it performs power flow calculation or real-time measurement on the distribution network to obtain the power flow status of the distribution network after reactive power adjustment by the distributed photovoltaic inverter at the current moment, including: the bus voltage on the low-voltage side of the main transformer, the power factor on the high-voltage side, and the voltage amplitude of each node. Step 5: Based on the low-voltage bus voltage and high-voltage power factor of the main transformer after reactive power regulation by the distributed photovoltaic inverter at the current moment, the AVC system determines the operating area of ​​the AVC system at the current moment according to the nine-zone diagram method. This is used to generate the tap adjustment action of the on-load tap changer (OLTC) and the switching action of the capacitor bank (CBs) at the current moment. Step 6: Based on the tap changer adjustment action of the on-load tap changer (OLTC) and the switching action of the capacitor bank (CBs) at the current moment, update the equipment status data of the AVC system at the current moment, and combine it with the reactive power, load fluctuation and photovoltaic output fluctuation of the distributed photovoltaic inverter at the current moment to obtain the operating status of the distribution network at the next moment. Step 7: Take the operating state of the distribution network at the next moment as the new operating state of the distribution network at the current moment, and return to Step 3 for rolling control, thereby realizing the distribution network voltage control of soft collaboration between distributed photovoltaic inverters and AVC system.

2. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 1, characterized in that, Step one includes: Step 1.1, obtain the first Each node at the current moment active power reactive power and voltage amplitude Thus constructing the current moment The set of active power of nodes Current moment The set of nodal reactive power and the current moment Set of node voltage amplitudes ,in, Indicates the number of nodes in the distribution network. ; Step 1.2: Obtain the current time value of the on-load tap-changing transformer (OLTC). tap position The capacitor bank CBs at the current time Number of throw groups On-load tap-changing transformer (OLTC) at the current moment Time interval from the last action The capacitor bank CBs at the current time Time interval from the last action On-load tap-changing transformers (OLTCs) as of the current time Number of actions As of the current time, the capacitor bank CBs The number of actions is Thus, the distribution network is constructed at the current moment. Operating status .

3. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 1, characterized in that, The action space in step two includes: Step 2.1: Assume the number of distributed photovoltaic inverters is... , No. At the current moment, a distributed photovoltaic inverter The reactive power output ratio is Thus constructing the current moment action ;in, , ; Step 2.2, let the first... The apparent capacity of a distributed photovoltaic inverter is At the current moment The active power is Therefore, the first equation can be used to calculate the second equation. At the current moment, a distributed photovoltaic inverter upper limit of reactive power ; (1) Step 2.3: Based on the upper limit of reactive power Use equation (2) to determine the first At the current moment, a distributed photovoltaic inverter Actual reactive power The output range; (2) Step 2.4, according to the first At the current moment, a distributed photovoltaic inverter reactive power output ratio The actual reactive power is calculated using equation (3). Thus, the current time is obtained. set of reactive power control quantities ; (3)。 4. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 1, characterized in that, The reward function in step two includes: Step 3.1: Let the set of branches in the distribution network be... , No. The node and the first Branches between nodes At the present moment The branch current is branch road The branch resistance is Thus, equation (4) is used to obtain the current state of the distribution network at the current time. Reward function for network loss : (4) Step 3.2, set the upper limit of the node voltage as follows: The lower limit of the node voltage is And based on the j-th node at the current time voltage amplitude Calculate the first using equation (5) Each node at the current moment The voltage exceeds the limit ; (5) Step 3.3: Calculate the current state of the distribution network using equation (6). Reward function for the voltage over-limit portion : (6) Step 3.4: Let the end time of a single round be... The maximum number of operations allowed for an on-load tap-changing transformer (OLTC) is [number]. The minimum allowable number of operations for an on-load tap-changing transformer (OLTC) is [number]. The maximum allowed number of operations for capacitor banks (CBs) is [number]. The lower limit of the number of operations allowed for the capacitor bank CBs is The operating penalty coefficient for an on-load tap-changing transformer (OLTC) is: The penalty coefficient for the action of the capacitor bank (CBs) is Therefore, equation (7) and equation (8) are used to calculate the distribution network at the current time. The reward function for the penalty part of the on-load tap-changing transformer (OLTC) operation count. The reward function for the penalty part of the number of actions of capacitor banks. : (7) (8) Step 3.5: Construct a distributed photovoltaic inverter reactive power control intelligent agent at the current moment using equation (9). reward function : (9) In equation (9), and These are the weighting coefficients for the network loss term and the voltage over-limit term, respectively.

5. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 1, characterized in that, Step two, which involves offline training of the reactive power control agent for the distributed photovoltaic inverter, includes: Step 4.1: Constructing the reactive power control intelligent agent for distributed photovoltaic inverters includes: policy network. First Evaluation Network Second evaluation network First objective evaluation network Second objective evaluation network ,in, For the parameters of the policy network, and These are the parameters of the first evaluation network and the parameters of the second evaluation network, respectively. and These are the parameters of the first objective evaluation network and the second objective evaluation network, respectively. Step 4.2, Set the experience replay pool as follows and the number of samples in the small batch is ;initialization ; Step 4.3: Set the current time in a single round. Distribution network operating status Input Policy Network Process the data and output the current time. The action distribution parameters are sampled to obtain the current time. action ; Step 4.4: Based on the current time action Calculate the current state of each distributed photovoltaic inverter. The actual reactive power, and the current moment The actual reactive power input is processed in the power flow calculation model of the distribution network to obtain the current reactive power of the distributed photovoltaic inverter. Power flow status of the distribution network after reactive power regulation; Step 4.5: The AVC system determines the current status of the distributed photovoltaic inverter based on the data from the current time. The power flow state of the distribution network after reactive power regulation is generated according to the nine-zone diagram method, showing the AVC system at the current moment. The control actions are used to update the current time of the on-load tap changer (OLTC). The tap position and the number of capacitor banks (CBs) switched on are used to obtain the next moment. Distribution network operating status ; Step 4.6: Calculate the current time. Reward value and with , and Together they constitute the current moment samples Then, it is stored in the experience replay pool. Middle; General Assign to Then, return to step 4.3 until the experience replay pool is reached. Until the number of samples reaches the preset value; Step 4.7, from the experience replay pool Randomly selected 1 sample; of which, the first bar sample ,in, For the first The state of the sample For the first The actions of the sample For the first Rewards for each sample For the first The next state of the sample ; Step 4.8, Input Policy Network Processing and sampling are performed to obtain the first... The next action of the sample Therefore, the first equation (10) is used to calculate the second equation. Target evaluation value of each sample : (10) In equation (10), Discount factor; Step 4.9: Construct the first... using equation (11) A rating network loss function and minimize the loss function To achieve the goal, update the first evaluation network separately. Second evaluation network : (11) Step 4.10: Construct a policy network using equation (12). loss function Parameters used to update the policy network : (12) In equation (12), Indicates the updated number An evaluation network; The coefficient of the entropy regularization term; Step 4.11: Use equation (13) to evaluate the parameters of the network for the first objective. Parameters of the second objective evaluation network Perform a soft update: (13) In equation (13), For the soft update coefficients of the target network; For assignment; For the first The parameters of the target evaluation network; Step 4.12: Repeat steps 4.3 to 4.11 to perform offline training on the reactive power control agent of the distributed photovoltaic inverter until the maximum number of training rounds is reached or the reward convergence condition is met, and the trained reactive power control agent of the distributed photovoltaic inverter is obtained.

6. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 1, characterized in that, Step three involves calculating the reactive power output ratio of each distributed photovoltaic inverter at the current moment. The actual reactive power, formed at the current moment set of reactive power control quantities ; and thus according to Send the current time to each distributed photovoltaic inverter The reactive power control command causes each distributed photovoltaic inverter to execute the current time... Reactive power regulation.

7. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 6, characterized in that, Step five includes: Step 6.1: Assume the upper limit of the bus voltage on the low-voltage side of the main transformer is... The lower limit of the bus voltage on the low-voltage side of the main transformer is The upper limit of the power factor on the high-voltage side is The lower limit of the power factor on the high-voltage side is ; Step 6.2: Construct a nine-zone operating plane for the AVC system, using the low-voltage bus voltage of the main transformer as the ordinate and the high-voltage power factor as the abscissa, and then... , , and The nine-zone map operating plane is divided into nine operating areas; Step 6.3: Based on the current state of the distributed photovoltaic inverter... Voltage of the low-voltage side bus of the main transformer after reactive power regulation and high voltage side power factor Determine the AVC system at the current moment. Operating area ; Step 6.4: Generate the AVC system at the current moment. Control actions ;in, This indicates the current time of the on-load tap-changing transformer (OLTC). The tap adjustment amount, This indicates the capacitor bank CBs at the current time. The amount of cutting adjustment; When the time interval between the last operation of the on-load tap changer (OLTC) or capacitor bank (CBs) is less than the corresponding minimum operating interval, the corresponding equipment is prohibited from operating, and the corresponding adjustment amount is set to 0. when In Inner and In At that time, the order and At the current moment No tap change or switching operations are performed to maintain the on-load tap changer (OLTC) at the current time. The tap position and the number of capacitor banks (CBs) switched on remain unchanged; when Below At that time, the AVC system is in the current moment First, execute the throwing and cutting action, that is, according to The number of capacitor banks (CBs) switched on and off is determined at the current time. Then perform the tap adjustment action, that is, according to Adjust the tap position of the on-load tap changer (OLTC) to increase the bus voltage; when Higher than At that time, the AVC system is in the current moment First, perform the tap adjustment action, that is, according to Lowering the tap position of the on-load tap changer (OLTC) if... Still higher And then at the current moment Execute the switching adjustment action, that is, according to The number of capacitor banks switched on and off is reduced to lower the bus voltage; when Below or higher At that time, according to Generate capacitor banks CBs at the current time The switching action, and the on-load tap-changing transformer OLTC at the current moment. The tap adjustment action together constitutes the AVC system at the current moment. Control actions .

8. The distribution network voltage control method for soft collaboration between distributed photovoltaic inverters and AVC systems according to claim 7, characterized in that, Step six includes: Step 7.1: Use equation (14) to obtain the distribution network at the next moment. Tap position of on-load tap changing transformer (OLTC) and the number of switching groups of capacitor banks (CBs) : (14) Step 7.2, according to , , Load fluctuations and photovoltaic output fluctuations are used to perform power flow calculations on the distribution network to obtain the distribution network's power flow at the next time step. The set of active power of nodes Node reactive power set and node voltage amplitude set ; thereby constructing the next moment Distribution network operating status 。 9. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports a processor in executing the method of any one of claims 1-8, the processor being configured to execute the program stored in the memory.

10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method according to any one of claims 1-8.