Droop control method and system for photovoltaic inverter in active power distribution network

By constructing the sag control function of the photovoltaic inverter and optimizing the reactive power output strategy using the Markov decision-making process, the problem of lack of systematic coordination of the sag control of the photovoltaic inverter in the active distribution network is solved, and the optimal voltage control of the whole system is achieved.

CN119966008AActive Publication Date: 2025-05-09NORTHEAST DIANLI UNIVERSITY

Patent Information

Application Number
CN202510436364.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
2045-04-09

AI Technical Summary

Technical Problem

In the active distribution network, the sag control of photovoltaic inverters lacks systematic coordination, which makes it impossible to meet the optimal voltage control of the entire system.

Method used

By constructing the sag control function of the photovoltaic inverter, the Markov decision-making process is used to optimize the output strategy of reactive power to minimize the active power loss of the active distribution network.

Benefits of technology

The target sag control strategy of photovoltaic inverter is realized, the voltage-reactive control efficiency is improved, and the optimal voltage control needs of the entire system are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119966008A_ABST
    Figure CN119966008A_ABST
Patent Text Reader

Abstract

The invention provides a droop control method and system for a photovoltaic inverter in an active power distribution network. According to the implementation scheme, a droop control function of the photovoltaic inverter is constructed; taking a linear droop region voltage lower limit, a linear droop region voltage upper limit, a dead zone region voltage upper limit, a dead zone region voltage lower limit and basic reactive power in the droop control function as decision variables, and taking active power loss minimization of the active power distribution network as a target to construct a target function; based on the objective function and constraint conditions of the objective function, defining a state space, an action space and a reward function of a Markov decision process; and based on the reward function, performing observation learning on the state space for each action in the action space, and outputting the action with the maximum reward value as a target droop control strategy of the photovoltaic inverter. By adopting the technical scheme of the invention, the droop control strategy of the photovoltaic inverter can be quickly and accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electric power voltage regulation, and in particular to a droop control method and system for a photovoltaic inverter in an active power distribution network. Background Art

[0002] With the continuous expansion of renewable energy power generation, the participation of renewable energy in power system voltage regulation has become an inevitable trend. As a growing, affordable and promising renewable energy, distributed photovoltaic (PV) has been widely used in active distribution network (ADN), which helps to reduce the use of fossil fuels and achieve a sustainable ecological civilization environment. However, due to the intermittent and uncertain nature of PV power generation and the high impedance ratio of ADN, the deep integration of PV and ADN has brought many operational challenges to ADN, such as excessive power loss, voltage over-limit and rapid voltage fluctuation. In order to meet these challenges, various voltage-reactive power control (Volt / Var control, VVC) methods have been applied in ADNs for voltage regulation and power loss reduction.

[0003] Conventional control devices, including on-load tap changers (OLTCs), capacitor banks (CBs), and voltage regulators (VRs), are used to perform VVC. They are mechanical devices that become less sensitive in mitigating the rapid changes of PVs in ADNs. In contrast, as an electronic device, photovoltaic inverters (PVs) are able to perform VVC by providing responsive and flexible reactive power adjustment. Therefore, adopting the droop control strategy of PV inverters to perform VCC has become an effective control strategy. Each inverter-based PV system independently adjusts its reactive power according to the locally measured voltage, with marginal operating costs. The conventional default droop control advocated by the IEEE 1547.8 standard is one of the most basic voltage control strategies. However, the lack of systematic coordination between the default droop controls will result in the PV inverter-based VVC method failing to meet the optimality of the entire system. Summary of the invention

[0004] The present invention provides a droop control method and system for a photovoltaic inverter in an active power distribution network, which can solve at least one of the above technical problems.

[0005] According to one aspect of the present invention, a droop control method for a photovoltaic inverter in an active power distribution network is provided, comprising: When the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at the maximum capacity; when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power as a requirement, and a droop control function of the photovoltaic inverter is constructed; Taking the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and basic reactive power in the droop control function as decision variables, and taking minimizing the active power loss of the active distribution network as the goal, constructing an objective function; Based on the objective function and the constraints of the objective function, define the state space, action space and reward function of the Markov decision process; Based on the reward function, the state space is observed and learned for each action in the action space, and the action with the maximum reward value is output as the target droop control strategy of the photovoltaic inverter.

[0006] According to another aspect of the present invention, there is provided a droop control device for a photovoltaic inverter in an active power distribution network, comprising: A droop control function determination module is used to construct a droop control function of the photovoltaic inverter based on the requirement that when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop area, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity, when the bus voltage is higher than the upper voltage limit of the linear droop area, the photovoltaic inverter absorbs reactive power at the maximum capacity, and when the bus voltage is between the lower voltage limit of the dead zone area and the upper voltage limit of the dead zone area, the photovoltaic inverter outputs basic reactive power; An objective function construction module is used to construct an objective function with the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and the basic reactive power in the droop control function as decision variables and with the minimization of the active power loss of the active distribution network as the goal; A redefinition module, for defining a state space, an action space and a reward function of a Markov decision process based on the objective function and the constraints of the objective function; A control strategy determination module is used to observe and learn the state space for each action in the action space based on the reward function and output the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.

[0007] According to another aspect of the present invention, a droop control system for a photovoltaic inverter in an active power distribution network is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the processor, and the processor obtains the instructions from the memory and executes the instructions, so that the processor can execute the droop control method for a photovoltaic inverter in an active power distribution network described in any one of the embodiments of the present invention.

[0008] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to provide to a computer and instruct the computer to execute the droop control method of a photovoltaic inverter in an active power distribution network described in any one of the embodiments of the present invention.

[0009] The technical solution of the present invention is adopted, when the bus voltage in the active distribution network is lower than the lower limit of the linear droop region voltage, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity, when the bus voltage is higher than the upper limit of the linear droop region voltage, the photovoltaic inverter absorbs reactive power at the maximum capacity, and when the bus voltage is between the lower limit of the dead zone region voltage and the upper limit of the dead zone region voltage, the photovoltaic inverter outputs basic reactive power as a requirement, and the droop control function of the photovoltaic inverter is constructed. In this way, the droop control curve of the photovoltaic inverter can be accurately described. Then, the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and the basic reactive power in the droop control function are used as decision variables, and the objective function is constructed with the goal of minimizing the active power loss of the active distribution network. In this way, an optimization problem based on the droop control function with the goal of minimizing network consumption can be constructed. Based on the objective function and the constraints of the objective function, the state space, action space and reward function of the Markov decision process are defined. In this way, the optimization problem is redefined as a Markov decision process, which is convenient for the subsequent solution of the objective function. In the Markov decision process solution process, based on the reward function, the state space is observed and learned for each action in the action space, and the action with the largest reward value is output as the target droop control strategy of the photovoltaic inverter, which can improve the solution efficiency and accuracy, so as to obtain the target droop control strategy of the photovoltaic inverter.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention. Figure 1is a flow chart of a droop control method for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention; Figure 2 is a schematic diagram of a droop control function according to an embodiment of the present invention; Figure 3 is a schematic diagram of a population iteration process according to an embodiment of the present invention; Figure 4 is a schematic diagram of an active power distribution network according to an embodiment of the present invention; Figure 5 is a structural block diagram of a droop control device for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention; Figure 6 The block diagram is a block diagram of an electronic device for implementing the method according to the embodiment of the present invention. DETAILED DESCRIPTION

[0012] The following is a description of exemplary embodiments of the present invention in conjunction with the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0013] Figure 1 The present invention is a flowchart of a droop control method for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention.

[0014] like Figure 1 As shown, the droop control method of the photovoltaic inverter in the active power distribution network may include: S110, when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity, when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at the maximum capacity, when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power as a requirement, and constructs a droop control function of the photovoltaic inverter; S120, taking the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and the basic reactive power in the droop control function as decision variables, and taking minimizing the active power loss of the active distribution network as the goal, constructing an objective function; S130, based on the objective function and the constraints of the objective function, define the state space, action space and reward function of the Markov decision process; S140, based on the reward function, observe and learn the state space for each action in the action space and output the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.

[0015] Understandably, there should be different types of droop control functions for PV inverters on different buses in the active distribution network to improve the efficiency of voltage / var control (VVC). Therefore, comprehensive optimization of different types of PV inverters can be considered.

[0016] Exemplarily, the above droop control function is: in, Indicates the active distribution network The photovoltaic inverter of each node is Reactive power at the moment, It is Nodes in The bus voltage observed at any moment, It is Nodes in The lower voltage limit of the linear droop region of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the linear droop region of the photovoltaic inverter at the moment, It is Nodes in The lower voltage limit of the dead zone of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the dead zone of the photovoltaic inverter at the moment, and It is Nodes in The slopes of the two linear droop regions of the PV inverter at time It is Nodes in The basic reactive power of the photovoltaic inverter at the moment, It is Nodes in The maximum output reactive power of the PV inverter at that moment.

[0017] Figure 2 is a schematic diagram of a droop control function according to an embodiment of the present invention.

[0018] like Figure 2As shown, the droop control function can be composed of five piecewise functions with different conditions. It includes piecewise functions corresponding to two maximum reactive power output areas, two linear droop areas and dead zone areas. Among them, Figure 2 The D length area in the dead zone is the dead zone, the areas on both sides of the dead zone are linear droop areas, and the other ends of the two linear droop areas are connected to the maximum reactive power output area. In other words, it is a logic-based function that selects the corresponding piecewise function to calculate the output reactive power based on the bus voltage. In the function model of this example, in order to make full use of the capacity, when the bus voltage is lower than the lower limit of the linear droop area voltage Or higher than the upper voltage limit of the linear droop region When the inverter generates or absorbs reactive power at its maximum capacity, the inverter will maintain its output as the set value, i.e. the basic reactive power. If the bus voltage is in the two slope regions, the linear droop region, the inverter reactive output is calculated based on the set point and the droop slope. and , the degree of change in the inverter's output reactive power when responding to a certain voltage change should be optimized.

[0019] in, , ,in, It indicates the output reactive power change of the photovoltaic inverter in the first linear droop region per unit time. It represents the bus voltage change observed per unit time for the first linear droop region. It indicates the change in output reactive power per unit time of the photovoltaic inverter in the second linear droop region. It represents the bus voltage change observed per unit time for the second linear droop region.

[0020] Exemplarily, to ensure the safe and stable operation of the photovoltaic inverter, the safe operation constraints of the photovoltaic inverter are as follows: ; in, yes Time Node The maximum reactive output power corresponding to the droop control function of the photovoltaic inverter on Is a node The rated apparent power of the PV inverter on Is a node The active power of the PV inverter on.

[0021] Exemplarily, with the primary constraint of ensuring that the bus voltage is within the safety regulations, and with the goal of minimizing network losses, a mathematical model of voltage-reactive power control of the distribution network based on a complete droop control strategy is established. Specifically, the optimal power flow of the active distribution network is considered, and the parameters of the droop control function are used as decision variables to establish a mathematical model of voltage-reactive power control of the active distribution network with the goal of minimizing network losses.

[0022] Exemplarily, the objective function corresponding to the data model may be as follows: ; in, is the objective function, yes Timeline The active power loss on the droop control function , , , and These parameters are decision variables.

[0023] Exemplarily, the active power loss of the active power distribution network is as follows: ; in, It is a line The resistance, It is a line exist The current at the moment, To optimize the cycle, It is a collection of lines in the active distribution network.

[0024] Exemplarily, the constraints of the objective function include power flow constraints of the active power distribution network and safe operation constraints of the photovoltaic inverter.

[0025] For example, the power flow constraint condition of the active power distribution network may be: ; ; in, yes Always install on the node The actual active power output by the photovoltaic inverter, yes Time Node The active power required by the load, yes Time Node The actual component of the upper voltage phasor, yes Time Node The imaginary component of the upper voltage phasor, yes Timeline The actual component of the upper line admittance, yes Timeline The imaginary component of the upper line admittance, yes Always install on the node The actual reactive power output by the PV inverter based on the droop control function, yes Time Node The reactive power required by the load, yes Time Node The actual component of the upper voltage phasor, yes Time Node The imaginary component of the upper voltage phasor.

[0026] In addition, the node voltage also needs to be set within a safe range to ensure the safe operation of the active distribution network and photovoltaic inverters.

[0027] Exemplarily, the full operation constraint of the photovoltaic inverter is that the output reactive power of the photovoltaic inverter is within the maximum output value range of the photovoltaic inverter, and .

[0028] In one embodiment, based on the objective function and the constraints of the objective function, the state space, action space and reward function of the Markov decision process are defined, including: determining the state space of the Markov decision process based on the load active power, load reactive power, node voltage and photovoltaic output active power of the node where the photovoltaic inverter is located in the active distribution network; determining the action space of the Markov decision process based on the decision variables in the objective function; determining the reward function of the Markov decision process based on the constraints of the objective function and the dependent variables in the objective function.

[0029] In this example, taking into account the uncertainty of real-time photovoltaic output, the difficulty in obtaining the parameter structure of the active distribution network system, and the difficulty in accurately modeling the power flow model, the state space, action space and appropriate reward function in the Markov decision process are defined according to the mathematical model of voltage-reactive power optimization of the active distribution network. The voltage-reactive power control problem of the active distribution network can be restated as a Markov decision process to overcome the above problems.

[0030] For example, the state space in the Markov decision process is defined, and the information that can reflect the operation status of the distribution network in the active distribution network power flow model is expressed as the state space ,as follows: ; in, yes Time Node The active power required by the load, that is, the active power of the load mentioned above, yes Time Node The reactive power required by the load, that is, the reactive power of the load mentioned above, yes Nodes observed at all times The voltage on the node, that is, the voltage of the node mentioned above, yes Always install on the node The actual output active power of the photovoltaic inverter is the above photovoltaic output active power.

[0031] For example, the action space in the Markov decision process is the variable in the droop control function. As shown in the following formula: .

[0032] Exemplarily, the reward function in the Markov decision process is composed of the objective function and constraints in the mathematical model, as follows: ; ; ; in, yes The single-step reward value of the agent at the moment, , , is the reward or penalty coefficient for each target in the reward function, is the total reward value, yes The degree to which the node voltage exceeds the limit in the active distribution network at any time. The photovoltaic inverter A penalty function that always runs safely, yes Timeline Active power loss on It is the node set of the active distribution network.

[0033] in, ,in, The value is and , Indicates the voltage lower limit, Indicates the upper voltage limit, Representation Node exist Voltage at the moment.

[0034] In one embodiment, based on a reward function, the state space is observed and learned for each action in the action space, and the action with the largest reward value is output as the target droop control strategy of the photovoltaic inverter, including: based on the state space, simulating the environment of the active distribution network; based on the action space, determining the initial population; starting from the initial population, performing the following iterative operations on the population: based on the fitness of each individual in the population, selecting the first K elite individuals in the fitness ranking from the population, adding noise to the first K elite individuals, sampling individuals from the first K elite individuals after adding noise, and adding the sampling results to the population, using an intelligent agent to interactively learn each individual in the population after adding the sampling results and the environment of the active distribution network, and using a reward function to calculate the learning results of the intelligent agent to obtain the fitness of each individual in the population after adding the sampling results; when the number of iterations of the population reaches a preset threshold, stop performing the above iterative operations on the population, and output the individual with the highest fitness in the population as the target droop control strategy of the photovoltaic inverter. Wherein, K is an integer greater than 1.

[0035] It can be understood that the population after adding the sampling results in each iteration is used as the population of the next iteration. The fitness of each individual in the population after adding the sampling results is the fitness of the corresponding individual in the population of the next iteration.

[0036] Figure 3 Schematic diagram of a population iteration process according to an embodiment of the present invention.

[0037] like Figure 3 As shown in FIG. 1 , this example uses an evolutionary strategy-based soft actor-critic (SAC) (ES-SAC) to learn the Markov decision process and obtain the final droop control strategy.

[0038] In this example, an evolutionary strategy (ES) for policy search is integrated to enable the SAC (Soft Actor-Critic) agent to learn additional experience and improve sample efficiency. In addition, the multivariate normal distribution is updated based on the elite population to shift the behavior distribution toward areas with higher fitness.

[0039] Exemplarily, the SAC algorithm is a model-free, action-evaluation, non-policy algorithm. This example uses its powerful deep neural networks (DNNs) to approximate the proposed droop control function with dead zone, and then uses system information to obtain the optimal droop control function for each inverter. On the one hand, compared with the policy-based deep reinforcement learning (DRL) algorithm, the policy-based SAC algorithm can sample from the playback buffer to reduce the variance of experience. On the other hand, SAC has a strong exploration capability and is more suitable for high-dimensional complex VVC environments.

[0040] SAC considers a broader maximum entropy objective, which promotes the selection of random strategies by increasing the expected entropy of the strategy above a specified value. To further explain the SAC algorithm, we first introduce the entropy regularized reinforcement learning algorithm. SAC maximizes the trade-off between the entropy of the strategy and the expected return given. The specific principle is as follows: ; ; in, is the entropy function, which is used to measure the randomness of the strategy; Indicates in status Take action The probability of , which can also be called strategy or strategy probability; is the action variable, express The action of the moment, express The state of the moment, express Thus, since the entropy function can quantify the amount of information about taking different actions in a specified state, that is, the expected entropy, by taking all possible actions The information amount is summed up to get the strategy in state The entropy under , which can reflect the determinism or stochasticity of the policy.

[0041] Among them, the hyperparameters is the trade-off coefficient, which is used to adjust the randomness of the optimal strategy. Its value range is ; represents the learned policy network, represents the policy network before learning, is the reward function, is Adaptability at all times, represents the soft Bellman equation, the result of which can be considered as an expected value.

[0042] in, Indicates the current state Execute the current action The next state obtained The reward value obtained when .

[0043] For example, SAC simultaneously learns two Function: First function and second function , and learn a policy network The parameters of the policy network are , the parameters of the evaluation network are , The value is or Understandably, the first function The corresponding evaluation network parameters are ,second function The corresponding evaluation network parameters are .

[0044] Entropy Regularization Function Aims to minimize the soft Bellman equation, as follows: ; in, is the fitness function, Indicates in status Next select action hour The function value of a function, , Indicates in status Next select action hour The function value of a function, Indicates in status Next select action The probability of express Moment of action.

[0045] Among them, the two after entropy regularization The function is used for the SAC algorithm. The SAC algorithm is based on two Get the minimum value in the function Therefore, in SAC The loss function corresponding to the function is as follows: .

[0046] in, . In getting reward and the new state Afterwards, the trajectory parameters are combined Stored in experience pool Among them, the experience pool Includes multiple experiences, each of which corresponds to a trajectory parameter combination.

[0047] in, It can be expressed as follows: ; in, , which means that in the policy network Searchable status in Next action , yes The reward value at that moment.

[0048] This example uses the reparameterization technique to optimize the policy. During the policy search process, Gaussian noise is added to SAC and action samples are obtained. ,as follows: ; in, Indicates that the parameter is In the policy network of and noise The action samples obtained by searching under the condition of represents the hyperbolic tangent function; It is to use DNNs method in the state The following parameters are The mean generated by the policy network; It is to use DNNs method in the state The following parameters are The covariance generated by the policy network processing; is from a standard normal distribution Random noise sampled in .

[0049] As shown below, by using a single-step gradient descent method to update function: .

[0050] in, express The loss function of the function gradient.

[0051] As shown below, the strategy is updated using a single-step method using the gradient ascent method: ; in, Indicates that from the experience pool A small batch of experience randomly drawn from Indicates that the parameter is In the policy network of The action samples obtained by searching under the condition , Indicates in status Select Action Sample hour The function value of a function, Indicates in status Select Action Sample The probability of Representation Policy Network The loss function gradient.

[0052] Finally, the SAC objective evaluates the network parameters Updated as follows: ; in, is a tiny constant used to update the weights in the evaluation network. Represents the evaluation network parameters target value.

[0053] Exemplarily, for the ES algorithm, in each iteration, the ES algorithm adds multivariate normal noise to the current population, selects the best performing individual, updates the parameters of the multivariate normal noise, and repeats this process until it finds the best strategy that best suits its problem.

[0054] For the policy network in the above example It can also be understood as an action entity, and can be iteratively updated using the following examples.

[0055] This example converts the multivariate Gaussian distribution Applied to sampling from action groups, where is the mean vector, , is the covariance matrix, is the variance, is the identity matrix. Before training, the multivariate Gaussian distribution should be initialized as follows: .

[0056] Then, since the optimal covariance matrix is ​​usually unknown, ES starts from the current Generation overall population distribution Sampling action individuals , and gradually update it. In other words, this distribution is obtained by New candidate action individuals are extracted from each generation to guide the strategy search. After all candidate action individuals are evaluated, the population distribution is updated in each generation. The sampling process can be described as follows: ; in, Represents the index of the ES-based action individual, Indicates ES-based action individuals; Represents the number of ES-based action individuals in the action group.

[0057] In this example, for each iteration, the action behavior subject based on SAC Will be attached to the ES-based action behavior body Therefore, all action individuals are integrated into a population with a total number of Then, all the action individuals in the population interact with the VVC simulation environment, and their respective objective function values ​​F (fitness) are evaluated by the agent's evaluation network, which represents the random reward of the time step in the agent's interaction. At the same time, the ES can provide information flow to the SAC agent by sending its trajectory to the replay buffer of the SAC agent. The replay buffer provides the small batch experience required for the update of the action network and the evaluation network of the SAC agent.

[0058] Next, ES directly selects the elite action individual with the highest fitness to obtain the elite population, where the number of action individuals in the elite group is K. And the K fittest individuals in the elite group are used to update the multivariate Gaussian distribution , from which the next generation is sampled after adding updated noise to prevent premature convergence. Among them, the multivariate Gaussian distribution The update is as follows: ; ; in, Indicates updated , for the next iteration, Indicates updated , for the next iteration. Indicates adding noise The number of actions under the condition Indicated in Daidi The fitness of an individual action.

[0059] The SAC agent then interacts with the VVC simulation environment. The evaluation network randomly samples a mini-batch of experience from the playback buffer and updates its network parameters using the single-step gradient descent method in the aforementioned SAC algorithm. The action network is then updated using the evaluation network and the mini-batch experience, i.e., the action network is updated using the policy gradient calculated using the single-step gradient ascent method in the aforementioned SAC algorithm.

[0060] Finally, the updated evaluation network, action network, and multivariate Gaussian distribution will be used for the next generation until the optimal strategy is explored.

[0061] In this example, the main idea behind the ES-SAC algorithm is to combine the heuristic method of ES to generate various experiences. At the same time, SAC uses its powerful gradient-based method to learn from them. The SAC method after integrating ES enables pure DRL to learn various strategies and improve exploration capabilities.

[0062] Based on the principle of the above-mentioned improved algorithm, in order to optimize the optimal droop control curve for each photovoltaic, the present invention uses an improved deep reinforcement learning algorithm to solve the Markov decision process. Solving the Markov decision process can be regarded as an interaction process between the DRL agent and the environment, which can be specifically described as follows: the state space in the Markov decision process corresponds to the observations that the DRL agent can obtain from the environment; the action space in the Markov decision process corresponds to the actions output by the DRL agent based on the rewards fed back by the environment; the reward function in the Markov decision process corresponds to the goal that the DRL agent needs to learn, and in the process of interacting with the environment, the reward value is maximized. The present invention is based on the continuous interaction between the improved ES-SAC agent and the active distribution network environment, so that the ES-SAC agent learns the optimal strategy, which corresponds to the optimal droop control parameters of each photovoltaic inverter, and then obtains the optimal droop control function based on the optimal parameters. The specific steps are as follows: Step 1: The active power generation power of photovoltaics, the active power and reactive power of loads, and the node voltage in the active distribution network system are used as the state space of the ES-SAC algorithm reinforcement learning; the four voltage inflection points and the basic reactive power setting value in the droop control strategy of the photovoltaic inverter are used as the action space of the ES-SAC algorithm reinforcement learning; and the objective function and operation constraints of the system are used as the reward function of the ES-SAC algorithm reinforcement learning.

[0063] Step 2: Design the neural network parameters inside the ES-SAC agent to obtain the improved ES-SAC agent.

[0064] Step 3, simulate the interactive learning process between the active distribution network and the intelligent agent, and interactively generate experience data and put it into the experience replay pool. At the same time, the improved ES-SAC intelligent agent can extract experience data from the experience pool and update the internal neural network parameters based on this, and generate new strategies to act on the distribution network environment.

[0065] Step 4, repeat step 3 until the maximum number of training rounds is reached, and ensure that the reward shows an increasing convergence trend during the training process, so as to obtain a trained improved ES-SAC intelligent agent, and then based on the trained ES-SAC, the optimal active distribution network control strategy can be generated in real time.

[0066] This example is based on the continuous interaction between the improved ES-SAC agent and the active distribution network environment, so that the ES-SAC agent can learn the optimal strategy, which is the optimal droop control parameter corresponding to each photovoltaic inverter, and then obtain the optimal droop control function based on the optimal parameters.

[0067] Compared with the prior art, the beneficial effects of the embodiments of the present invention are: the embodiments of the present invention construct a complete droop control function, study the voltage-reactive power control model of the active distribution network based on the complete droop control function, and use an improved deep reinforcement learning algorithm to optimize the droop control strategy, explore the superiority of the improved deep reinforcement learning algorithm in solving the voltage-reactive power control problem, and can enable photovoltaic inverters to participate in active distribution network voltage regulation.

[0068] Figure 4 It is a structural schematic diagram of an active power distribution network according to an embodiment of the present invention.

[0069] Take the IEEE-33 node system in the power field as an example. Figure 4 As shown, one inverter-based PV is installed at each of nodes 4, 10, 13, 18, 21, 25, 27, 29 and 33. Figure 4 The PV location and the network topology including 33 nodes are shown. The capacity of each PV inverter does not exceed 110% of its rated power. The reference voltage of the IEEE-33 node system is 12.66kV, and the set reference power is 10MVA. The node voltage per unit value of the relaxed bus is 1p.u., and the safe operating range of the node voltage is 0.95pu-1.05pu. The method provided in this example is applied to the active distribution network, and relevant verification experiments can be carried out.

[0070] Figure 5 It is a structural block diagram of a droop control device for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention.

[0071] like Figure 5As shown, the droop control device of the photovoltaic inverter in the active power distribution network may include: The droop control function determination module 510 is used to construct a droop control function of the photovoltaic inverter based on the requirement that when the bus voltage in the active distribution network is lower than the lower limit of the linear droop region voltage, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity, when the bus voltage is higher than the upper limit of the linear droop region voltage, the photovoltaic inverter absorbs reactive power at the maximum capacity, and when the bus voltage is between the lower limit of the dead zone region voltage and the upper limit of the dead zone region voltage, the photovoltaic inverter outputs basic reactive power; An objective function construction module 520 is used to construct an objective function with the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and the basic reactive power in the droop control function as decision variables and with minimizing the active power loss of the active distribution network as the goal; A redefinition module 530, for defining a state space, an action space and a reward function of a Markov decision process based on the objective function and the constraints of the objective function; The control strategy determination module 540 is used to observe and learn the state space for each action in the action space based on the reward function and output the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.

[0072] In one embodiment, the redefinition module includes: A state space determination unit, configured to determine a state space of a Markov decision process based on load active power, load reactive power, node voltage, and photovoltaic output active power of a node where the photovoltaic inverter is located in the active power distribution network; an action space determination unit, configured to determine the action space of the Markov decision process based on the decision variables in the objective function; A reward function determination unit is used to determine the reward function of the Markov decision process based on the constraint conditions of the objective function and the dependent variable in the objective function.

[0073] In one embodiment, the control strategy determination module includes: An environment determination unit, configured to simulate an environment of the active power distribution network based on the state space; A population initialization unit, used for determining an initial population based on the action space; A population iteration unit is used to perform the following iterative operations on the population starting from the initial population: based on the fitness of each individual in the population, select the top K elite individuals in fitness ranking from the population, add noise to the top K elite individuals, sample individuals from the top K elite individuals after adding noise, and add the sampling results to the population, use an intelligent agent to interactively learn between each individual in the population after adding the sampling results and the environment of the active power distribution network, and use the reward function to calculate the learning results of the intelligent agent to obtain the fitness of each individual in the population after adding the sampling results; A strategy determination unit is used to stop executing the iteration operation on the population when the number of iterations of the population reaches a preset threshold, and output the individual with the highest fitness in the population as the target droop control strategy of the photovoltaic inverter.

[0074] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, reference can be made to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0075] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0076] According to an embodiment of the present invention, the present invention also provides a system and a readable storage medium.

[0077] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0078] like Figure 6As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 to a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0079] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0080] The computing unit 801 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 801 performs the various methods and processes described above, such as a droop control method for a photovoltaic inverter in an active power distribution network. For example, in some embodiments, the droop control method for a photovoltaic inverter in an active power distribution network may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the droop control method for a photovoltaic inverter in an active power distribution network described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the droop control method for the photovoltaic inverter in the active power distribution network in any other appropriate manner (for example, by means of firmware).

[0081] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0082] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0083] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0084] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0085] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0086] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0087] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.

[0088] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A droop control method for a photovoltaic inverter in an active power distribution network, characterized in that: include: When the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at the maximum capacity; when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power as a requirement, and a droop control function of the photovoltaic inverter is constructed; Taking the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and basic reactive power in the droop control function as decision variables, and taking minimizing the active power loss of the active distribution network as the goal, constructing an objective function; Based on the objective function and the constraints of the objective function, define the state space, action space and reward function of the Markov decision process; Based on the reward function, the state space is observed and learned for each action in the action space, and the action with the maximum reward value is output as the target droop control strategy of the photovoltaic inverter.

2. The method according to claim 1, characterized in that: The droop control function is: in, Indicates the active power distribution network The photovoltaic inverter of each node is Reactive power at the moment, It is Nodes in The bus voltage observed at any moment, It is Nodes in The lower voltage limit of the linear droop region of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the linear droop region of the photovoltaic inverter at the moment, It is Nodes in The lower voltage limit of the dead zone of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the dead zone of the photovoltaic inverter at the moment, and It is Nodes in The slopes of the two linear droop regions of the PV inverter at time It is Nodes in The basic reactive power of the photovoltaic inverter at the moment, It is Nodes in The maximum output reactive power of the PV inverter at that moment.

3. The method according to claim 1, characterized in that The step of defining the state space, action space and reward function of the Markov decision process based on the objective function and the constraints of the objective function includes: Determine a state space of a Markov decision process based on load active power, load reactive power, node voltage, and photovoltaic output active power of a node where the photovoltaic inverter is located in the active power distribution network; Determining an action space of the Markov decision process based on the decision variables in the objective function; A reward function of the Markov decision process is determined based on the constraints of the objective function and the dependent variables in the objective function.

4. The method according to claim 3, characterized in that: Based on the reward function, observing and learning the state space for each action in the action space and outputting an action with the maximum reward value as a target droop control strategy for the photovoltaic inverter includes: Based on the state space, simulating the environment of the active power distribution network; Based on the action space, determining an initial population; Starting from the initial population, the following iterative operations are performed on the population: based on the fitness of each individual in the population, the top K elite individuals ranked by fitness are selected from the population, noise is added to the top K elite individuals, individual sampling is performed from the top K elite individuals after the noise is added, and the sampling results are added to the population, using an intelligent agent to interactively learn between each individual in the population after the sampling results are added and the environment of the active power distribution network, and the learning results of the intelligent agent are calculated using the reward function to obtain the fitness of each individual in the population after the sampling results are added; When the number of iterations of the population reaches a preset threshold, the iteration operation on the population is stopped, and the individual with the highest fitness in the population is output as the target droop control strategy of the photovoltaic inverter.

5. The method according to claim 1, characterized in that The constraint conditions of the objective function include the power flow constraint conditions of the active power distribution network and the safe operation constraint conditions of the photovoltaic inverter.

6. A droop control device for a photovoltaic inverter in an active power distribution network, characterized in that: include: A droop control function determination module is used to construct a droop control function of the photovoltaic inverter based on the requirement that when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop area, the photovoltaic inverter in the active distribution network generates reactive power at the maximum capacity, when the bus voltage is higher than the upper voltage limit of the linear droop area, the photovoltaic inverter absorbs reactive power at the maximum capacity, and when the bus voltage is between the lower voltage limit of the dead zone area and the upper voltage limit of the dead zone area, the photovoltaic inverter outputs basic reactive power; An objective function construction module is used to construct an objective function with the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit and the basic reactive power in the droop control function as decision variables and with the minimization of the active power loss of the active distribution network as the goal; A redefinition module, for defining a state space, an action space and a reward function of a Markov decision process based on the objective function and the constraints of the objective function; A control strategy determination module is used to observe and learn the state space for each action in the action space based on the reward function and output the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.

7. The device according to claim 6, characterized in that The redefinition module includes: A state space determination unit, configured to determine a state space of a Markov decision process based on load active power, load reactive power, node voltage, and photovoltaic output active power of a node where the photovoltaic inverter is located in the active power distribution network; an action space determination unit, configured to determine the action space of the Markov decision process based on the decision variables in the objective function; A reward function determination unit is used to determine the reward function of the Markov decision process based on the constraint conditions of the objective function and the dependent variable in the objective function.

8. The device according to claim 7, characterized in that The control strategy determination module comprises: An environment determination unit, configured to simulate an environment of the active power distribution network based on the state space; A population initialization unit, used for determining an initial population based on the action space; A population iteration unit is used to perform the following iterative operations on the population starting from the initial population: based on the fitness of each individual in the population, select the top K elite individuals in fitness ranking from the population, add noise to the top K elite individuals, sample individuals from the top K elite individuals after adding noise, and add the sampling results to the population, use an intelligent agent to interactively learn between each individual in the population after adding the sampling results and the environment of the active power distribution network, and use the reward function to calculate the learning results of the intelligent agent to obtain the fitness of each individual in the population after adding the sampling results; A strategy determination unit is used to stop executing the iteration operation on the population when the number of iterations of the population reaches a preset threshold, and output the individual with the highest fitness in the population as the target droop control strategy of the photovoltaic inverter.

9. A droop control system for a photovoltaic inverter in an active power distribution network, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the processor, and the processor is used to obtain the instructions from the memory and execute the instructions, so that the processor can execute the droop control method of the photovoltaic inverter in the active distribution network described in any one of claims 1-5.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to be provided to a computer to instruct the computer to execute the droop control method for a photovoltaic inverter in an active power distribution network according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for selecting deep network parameter sparse threshold based on Gaussian distribution estimation

    CN111488981A

  • Photovoltaic inverter control method based on environment adaptive droop control

    CN117318141A

  • Photovoltaic inverter coordinated voltage reactive power control method and system and storage medium

    CN118508537A

  • Distributed photovoltaic reactive voltage control method and system

    CN118630776A

Cited By

  • Inverter adaptive power distribution method and system based on environment perception

    CN120933981A

  • Environment perception based inverter adaptive power distribution method and system

    CN120933981B