Droop Control Method and System for Photovoltaic Inverters in Active Distribution Networks
By constructing the sag control function of the photovoltaic inverter and an improved deep reinforcement learning algorithm, the reactive power adjustment of the photovoltaic inverter is optimized, and the problems of voltage fluctuations and power losses in the active distribution network are solved, and more efficient voltage regulation and power optimization are achieved.
Patent Information
- Application Number
- CN202510436364.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-09
AI Technical Summary
In active distribution networks, the intermittent and uncertainty of distributed photovoltaic power generation lead to voltage fluctuations and power loss. The traditional photovoltaic inverter control strategy lacks systematic coordination and cannot meet the optimal voltage regulation of the entire system.
The sag control function of the photovoltaic inverter is constructed, and the state space, action space and reward functions of the Markov decision-making process are defined. The sag control strategy is optimized using an improved deep reinforcement learning algorithm, and the reactive power adjustment of the photovoltaic inverter is combined with the ES-SAC algorithm to minimize active power loss.
The voltage-reactive control efficiency of the photovoltaic inverter is improved, the sagging control strategy of the photovoltaic inverter is optimized, and the power loss and voltage fluctuations of the active distribution network are reduced.
Smart Images

Figure CN119966008B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power voltage regulation, and in particular to a droop control method and system for a photovoltaic inverter in an active power distribution network. Background Art
[0002] With the continuous expansion of renewable energy power generation, their participation in power system voltage regulation has become an inevitable trend. Distributed photovoltaics (PV), a growing, affordable, and promising renewable energy source, have been widely adopted in active distribution networks (ADNs), helping to reduce fossil fuel use and achieve a sustainable ecological civilization. However, due to the intermittent and uncertain nature of PV power generation and the high impedance ratio of ADNs, the deep integration of PV and ADNs presents many operational challenges, such as excessive power loss, voltage over-limit, and rapid voltage fluctuations. To address these challenges, various voltage / var control (VVC) methods have been applied in ADNs to regulate voltage and reduce power losses.
[0003] Traditional control devices, including on-load tap changers (OLTCs), capacitor banks (CBs), and voltage regulators (VRs), are used to implement VVC. These mechanical devices are less sensitive to the rapid variations in PV power in ADNs. In contrast, photovoltaic inverters (PVs), as electronic devices, are capable of implementing VVC by providing responsive and flexible reactive power adjustment. Therefore, employing PV inverter-based droop control strategies to implement VCC has become an effective control strategy. Each inverter-based PV system independently adjusts its reactive power based on locally measured voltage, with marginal operating costs. The conventional default droop control advocated by the IEEE 1547.8 standard is one of the most basic voltage control strategies. However, the lack of systematic coordination among default droop controls results in PV inverter-based VVC approaches failing to achieve system-wide optimality. Summary of the Invention
[0004] The present invention provides a droop control method and system for a photovoltaic inverter in an active power distribution network, which can solve at least one of the above technical problems.
[0005] According to one aspect of the present invention, a droop control method for a photovoltaic inverter in an active power distribution network is provided, comprising:
[0006] A droop control function for the photovoltaic inverter is constructed based on the following requirements: when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at maximum capacity; when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power;
[0007] An objective function is constructed by taking the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit, and basic reactive power in the droop control function as decision variables and minimizing the active power loss of the active distribution network as the goal;
[0008] Based on the objective function and the constraints of the objective function, defining a state space, an action space, and a reward function of a Markov decision process;
[0009] Based on the reward function, observation learning is performed on the state space for each action in the action space, and an action with the maximum reward value is output as a target droop control strategy for the photovoltaic inverter.
[0010] According to another aspect of the present invention, there is provided a droop control device for a photovoltaic inverter in an active power distribution network, comprising:
[0011] a droop control function determination module, configured to construct a droop control function for the photovoltaic inverter based on the following requirements: when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at maximum capacity; and when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power;
[0012] an objective function construction module, configured to construct an objective function using the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit, and basic reactive power in the droop control function as decision variables and minimizing the active power loss of the active distribution network as a goal;
[0013] A redefinition module, configured to define a state space, an action space, and a reward function of a Markov decision process based on the objective function and the constraints of the objective function;
[0014] A control strategy determination module is used to observe and learn the state space for each action in the action space based on the reward function and output the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.
[0015] According to another aspect of the present invention, a droop control system for a photovoltaic inverter in an active power distribution network is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the processor, and the processor retrieves the instructions from the memory and executes the instructions, so that the processor can execute the droop control method for a photovoltaic inverter in an active power distribution network described in any one of the embodiments of the present invention.
[0016] According to another aspect of the present invention, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to be provided to a computer and instruct the computer to execute the droop control method for a photovoltaic inverter in an active power distribution network described in any one of the embodiments of the present invention.
[0017] Using the technical solution of the present invention, a droop control function for the photovoltaic inverter is constructed, with the requirements that when the bus voltage in the active distribution network is below the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at maximum capacity; when the bus voltage is above the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at maximum capacity; and when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power. This accurately describes the droop control curve of the photovoltaic inverter. Then, using the lower voltage limit of the linear droop region, the upper voltage limit of the linear droop region, the upper voltage limit of the dead zone region, the lower voltage limit of the dead zone region, and the basic reactive power in the droop control function as decision variables, and minimizing the active power loss of the active distribution network as the objective, an objective function is constructed. This allows an optimization problem based on the droop control function to be constructed with the objective of minimizing network power consumption. Based on the objective function and its constraints, the state space, action space, and reward function of the Markov decision process are defined. This redefines the optimization problem as a Markov decision process, facilitating the subsequent solution of the objective function. During the Markov decision process, based on the reward function, observations are learned about the state space for each action in the action space, and the action with the maximum reward is output as the target droop control strategy for the photovoltaic inverter. This improves solution efficiency and accuracy, ultimately resulting in the target droop control strategy for the photovoltaic inverter.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are provided for a better understanding of the present invention and do not constitute a limitation of the present invention.
[0020] Figure 1 is a flow chart of a droop control method for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of a droop control function according to an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of a population iteration process according to an embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of an active power distribution network according to an embodiment of the present invention;
[0024] Figure 5 This is a structural block diagram of a droop control device for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention;
[0025] Figure 6 is a block diagram of an electronic device for implementing the method according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The following description of exemplary embodiments of the present invention is made in conjunction with the accompanying drawings, and various details of the embodiments of the present invention are included to facilitate understanding. These details should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] Figure 1 The flowchart of the droop control method of the photovoltaic inverter in the active power distribution network according to one embodiment of the present invention is shown.
[0028] like Figure 1 As shown, the droop control method of the photovoltaic inverter in the active power distribution network may include:
[0029] S110, constructing a droop control function for the photovoltaic inverter with the requirements that when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at its maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at its maximum capacity; and when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power;
[0030] S120, constructing an objective function with the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit, and basic reactive power in the droop control function as decision variables and minimizing the active power loss of the active distribution network as the goal;
[0031] S130, based on the objective function and the constraints of the objective function, defining the state space, action space and reward function of the Markov decision process;
[0032] S140 , based on the reward function, observing and learning the state space for each action in the action space and outputting the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.
[0033] Understandably, different types of droop control functions should be implemented for PV inverters on different buses in an active distribution network to improve the efficiency of voltage / var control (VVC). Therefore, comprehensive optimization of different types of PV inverters can be considered.
[0034] Exemplarily, the droop control function is:
[0035]
[0036] in, Indicates the active distribution network The photovoltaic inverter of each node is Reactive power at the moment, It is Nodes in The bus voltage observed at any moment, It is Nodes in The lower limit of the linear droop region voltage of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the linear droop region of the photovoltaic inverter at the moment, It is Nodes in The lower limit of the dead zone voltage of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the dead zone of the photovoltaic inverter at the moment, and It is Nodes in The slopes of the two linear droop regions of the photovoltaic inverter at the moment, It is Nodes in The basic reactive power of the photovoltaic inverter at the moment, It is Nodes in The maximum output reactive power of the PV inverter at the moment.
[0037] Figure 2 FIG. 4 is a schematic diagram of a droop control function according to an embodiment of the present invention.
[0038] like Figure 2 As shown in Figure 1, the droop control function can be composed of five piecewise functions with different conditions. Including piecewise functions corresponding to two maximum reactive output areas, two linear droop areas and dead zone areas. Figure 2 The D-length region in the figure is the dead zone, and the areas on both sides of the dead zone are linear droop regions. The other ends of the two linear droop regions are connected to the maximum reactive power output region. In other words, it is a logic-based function that selects the corresponding piecewise function based on the bus voltage to calculate the output reactive power. In the function model of this example, in order to fully utilize the capacity, when the bus voltage is lower than the lower voltage limit of the linear droop region, Or higher than the upper voltage limit of the linear droop region When the inverter generates or absorbs reactive power at its maximum capacity, if the bus voltage falls in the dead band, the inverter maintains its output as the set value, i.e. the basic reactive power. If the bus voltage is in the two slope regions, i.e. the linear droop region, the inverter reactive output is calculated based on the set point and the droop slope. and , the degree of change in the inverter's output reactive power when responding to a certain voltage change should be optimized and adjusted.
[0039] in, , ,in, It indicates the output reactive power change of the photovoltaic inverter per unit time in the first linear droop region. It represents the bus voltage change observed per unit time for the first linear droop region. It indicates the output reactive power change of the photovoltaic inverter per unit time in the second linear droop region. It represents the bus voltage change observed per unit time in the second linear droop region.
[0040] For example, to ensure the safe and stable operation of the photovoltaic inverter, the safe operation constraints of the photovoltaic inverter are as follows:
[0041] ;
[0042] in, yes Time Node The maximum reactive output power corresponding to the droop control function of the photovoltaic inverter on is a node The rated apparent power of the PV inverter on is a node The active power of the PV inverter on the PV panel.
[0043] For example, a mathematical model for voltage-reactive power control in a distribution network based on a complete droop control strategy is established, with the primary constraint of ensuring bus voltage is within the specified safety range and the goal of minimizing network losses. Specifically, the optimal power flow of an active distribution network is considered, and the parameters of the droop control function are used as decision variables to establish a mathematical model for voltage-reactive power control in an active distribution network that takes network losses into account.
[0044] For example, the objective function corresponding to the data model may be as follows:
[0045] ;
[0046] in, is the objective function, yes Timeline Active power loss on the droop control function , , , and These parameters are decision variables.
[0047] For example, the active power loss of the active power distribution network is as follows:
[0048] ;
[0049] in, It's a line The resistance, It's a line exist The current at the moment, To optimize the cycle, It is a collection of lines in the active distribution network.
[0050] Exemplarily, the constraints of the objective function include power flow constraints of the active power distribution network and safe operation constraints of the photovoltaic inverter.
[0051] For example, the power flow constraint condition of the active power distribution network may be:
[0052] ;
[0053] ;
[0054] in, yes Always installed on the node The actual active power output by the photovoltaic inverter, yes Time Node The active power required by the load, yes Time Node The actual component of the upper voltage phasor, yes Time Node The imaginary component of the upper voltage phasor, yes Timeline The actual component of the upper line admittance, yes Timeline The imaginary component of the upper line admittance, yes Always installed on the node The actual reactive power output by the photovoltaic inverter based on the droop control function, yes Time Node The reactive power required by the load, yes Time Node The real component of the upper voltage phasor, yes Time Node The imaginary component of the upper voltage phasor.
[0055] In addition, the node voltage also needs to be set within a safe range to ensure the safe operation of the active distribution network and photovoltaic inverters.
[0056] For example, the full operation constraint of the photovoltaic inverter is that the output reactive power of the photovoltaic inverter is within the maximum output value range of the photovoltaic inverter, and .
[0057] In one embodiment, based on the objective function and the constraints of the objective function, the state space, action space and reward function of the Markov decision process are defined, including: determining the state space of the Markov decision process based on the load active power, load reactive power, node voltage and photovoltaic output active power of the node where the photovoltaic inverter is located in the active distribution network; determining the action space of the Markov decision process based on the decision variables in the objective function; and determining the reward function of the Markov decision process based on the constraints of the objective function and the dependent variables in the objective function.
[0058] In this example, considering the uncertainty of real-time photovoltaic output, the difficulty in obtaining the parameter structure of the active distribution network system, and the difficulty in accurately modeling the power flow model, the state space, action space, and appropriate reward function of the Markov decision process are defined based on the mathematical model of the voltage-reactive power optimization of the active distribution network. The voltage-reactive power control problem of the active distribution network can be reformulated as a Markov decision process to overcome the above problems.
[0059] For example, the state space of the Markov decision process is defined, and the information that can reflect the operation status of the distribution network in the active distribution network flow model is expressed as the state space ,as follows:
[0060] ;
[0061] in, yes Time Node The active power required by the load, that is, the active power of the load mentioned above, yes Time Node The reactive power required by the load, that is, the reactive power of the load mentioned above, yes Nodes observed at all times The voltage on the node, that is, the voltage of the node mentioned above, yes Always installed on the node The actual output active power of the photovoltaic inverter is the above photovoltaic output active power.
[0062] For example, the action space in the Markov decision process is the variable in the droop control function. As shown in the following formula: .
[0063] For example, the reward function in the Markov decision process is composed of the objective function and constraints in the mathematical model, as follows:
[0064] ;
[0065] ;
[0066] ;
[0067] in, yes The single-step reward value of the agent at the moment, , , is the reward or penalty coefficient for each target in the reward function, is the total reward value, yes The degree of voltage exceeding the limit of nodes in the active distribution network at any time, The photovoltaic inverter A penalty function that always runs safely, yes Timeline Active power loss on It is the set of nodes of the active distribution network.
[0068] in, ,in, The value is and , Indicates the lower voltage limit, Indicates the upper voltage limit, Representation node exist Voltage at the moment.
[0069] In one embodiment, based on a reward function, observational learning is performed on the state space for each action in the action space, and the action with the maximum reward value is output as the target droop control strategy for a photovoltaic inverter. The method includes: simulating the environment of an active power distribution network based on the state space; determining an initial population based on the action space; and iterating the population from the initial population to perform the following operations: selecting the top K elite individuals in the population based on the fitness of each individual in the population, adding noise to the top K elite individuals, sampling individuals from the top K elite individuals after adding noise, and adding the sampled results to the population. Using an intelligent agent, each individual in the population after adding the sampled results interacts with the active power distribution network environment, and calculating the learning results of the intelligent agent using a reward function to obtain the fitness of each individual in the population after adding the sampled results. When the number of iterations of the population reaches a preset threshold, the iterative operation on the population is stopped, and the individual with the highest fitness in the population is output as the target droop control strategy for the photovoltaic inverter. Where K is an integer greater than 1.
[0070] It can be understood that the population after adding the sampling results in each iteration is used as the population for the next iteration. The fitness of each individual in the population after adding the sampling results is the fitness of the corresponding individual in the population for the next iteration.
[0071] Figure 3 Schematic diagram of a population iteration process according to an embodiment of the present invention.
[0072] like Figure 3As shown in Figure 1, this example uses the Evolutionary Strategy (ES) Embedded Soft Actor-Critic (SAC) algorithm (ES-SAC) to learn the Markov decision process and obtain the final droop control strategy.
[0073] In this example, an Evolutionary Strategy (ES) is integrated for policy search, enabling the Soft Actor-Critic (SAC) agent to learn from additional experience and improve sample efficiency. Furthermore, a multivariate normal distribution is updated based on an elite population, shifting the behavior distribution toward regions with higher fitness.
[0074] For example, the SAC algorithm is a model-free, action-criteria, off-policy algorithm. This example leverages its powerful deep neural networks (DNNs) to approximate the proposed droop control function with dead zone, and then leverages system information to obtain the optimal droop control function for each inverter. Compared to policy-based deep reinforcement learning (DRL) algorithms, the policy-based SAC algorithm can sample from the replay buffer to reduce the variance of experience. Furthermore, SAC possesses powerful exploration capabilities, making it more suitable for high-dimensional and complex VVC environments.
[0075] SAC considers a broader maximum entropy objective, which promotes the selection of random policies by increasing the expected entropy of the policy above a specified value. To further explain the SAC algorithm, we first introduce the entropy regularized reinforcement learning algorithm. SAC maximizes the trade-off between the entropy of the policy and the expected reward given. The specific principle is as follows:
[0076] ;
[0077] ;
[0078] in, is the entropy function, which is used to measure the randomness of the strategy; Indicates that the status Take action The probability of , which can also be called strategy or strategy probability; is the action variable, express The action of the moment, express The state of the moment, express Thus, since the entropy function can quantify the amount of information about taking different actions in a specified state, that is, the expected entropy, by taking all possible actions The sum of the information of the strategy is obtained in the state The entropy under , which can reflect the determinism or stochasticity of the policy.
[0079] Among them, the hyperparameters Is the trade-off coefficient, which is used to adjust the randomness of the optimal strategy. Its value range is ; represents the learned policy network, represents the policy network before learning, is the reward function, is Adaptability at all times, represents the soft Bellman equation, the result of which can be considered as an expected value.
[0080] in, Indicates the current state Execute the current action The next state obtained The reward value obtained when .
[0081] For example, SAC simultaneously learns two Function: First function and second function , and learn a policy network The parameters of the policy network are , the parameters of the evaluation network are , The value is or Understandably, first function The corresponding evaluation network parameters are ,second function The corresponding evaluation network parameters are .
[0082] Entropy regularization function Aims to minimize the soft Bellman equation, as follows:
[0083] ;
[0084] in, is the fitness function, Indicates that the status Select Action hour The function value of a function, , Indicates that the status Select Action hour The function value of a function, Indicates that the status Select Action The probability of express Moment of action.
[0085] Among them, the two after entropy regularization The function is used for the SAC algorithm. The SAC algorithm is based on two Get the minimum value in the function Therefore, in SAC The loss function corresponding to the function is as follows:
[0086] .
[0087] in, . and the new state Afterwards, the trajectory parameters are combined Stored in experience pool Among them, the experience pool Includes multiple experiences, each of which corresponds to a trajectory parameter combination.
[0088] in, It can be expressed as follows:
[0089] ;
[0090] in, , which means that in the policy network Searchable status in Next action , yes The reward value at that moment.
[0091] This example uses the reparameterization technique to optimize the policy. During the policy search process, Gaussian noise is added to the SAC and action samples are obtained. ,as follows:
[0092] ;
[0093] in, Indicates that the parameter is In the policy network of state and noise The action samples obtained by searching under the condition of represents the hyperbolic tangent function; It is to use DNNs method in the state The following parameters are The mean value generated by the policy network; It is to use DNNs method in the state The following parameters are The covariance generated by the policy network; is from the standard normal distribution Random noise sampled in .
[0094] As shown below, by using a single-step gradient descent method to update function:
[0095] .
[0096] in, express Loss function of function gradient.
[0097] As shown below, the strategy is updated using a single-step gradient ascent method:
[0098]
[0099] ;
[0100] in, Indicates that from the experience pool A small batch of experiences randomly drawn from Indicates that the parameter is In the policy network of state The action samples obtained by searching under the condition , Indicates that the status Select Action Sample hour The function value of a function, Indicates that the status Select Action Sample The probability of Representation Policy Network The loss function gradient.
[0101] Finally, the SAC objective evaluation network parameters Updates as follows:
[0102] ;
[0103] in, is a tiny constant used to update the weights in the evaluation network. Represents the evaluation network parameters target value.
[0104] For example, for the ES algorithm, in each iteration, the ES algorithm adds multivariate normal noise to the current population, selects the best performing individual, updates the parameters of the multivariate normal noise, and repeats this process until it finds the best strategy that best suits its problem.
[0105] For the policy network in the above example It can also be understood as an individual action, and can be iteratively updated using the following example.
[0106] This example converts the multivariate Gaussian distribution Applied to sampling from action groups, where is the mean vector, , is the covariance matrix, is the variance, is the identity matrix. Before training, the multivariate Gaussian distribution should be initialized as follows:
[0107] .
[0108] Then, since the optimal covariance matrix is usually unknown, ES starts from the current Generation overall population distribution Sampling action individuals , and gradually update it. In other words, this distribution is obtained by New candidate action individuals are sampled in each generation to guide the policy search. After all candidate action individuals are evaluated, the population distribution is updated in each generation. The sampling process can be described as follows:
[0109] ;
[0110] in, Represents the index of the ES-based action individual, Indicates the ES-based action individuals; Represents the number of ES-based action individuals in the action group.
[0111] In this example, for each iteration, the action behavior subject based on SAC Will be attached to the ES-based action behavior body Therefore, all action individuals are integrated into a population with a total number of All agents in the population then interact with the VVC simulated environment, and their respective objective function values F (fitness) are evaluated by the agent's critic network, which represents the stochastic reward at each time step in the agent's interaction. Simultaneously, the ES can provide information flow to the SAC agent by sending its trajectory to the replay buffer of the SAC agent. This replay buffer provides the mini-batch experience needed to update the SAC agent's action network and critic network.
[0112] Then, ES directly selects the elite action individual with the highest fitness to obtain the elite population, where the number of action individuals in the elite group is K. In addition, the K fittest individuals in the elite group are used to update the multivariate Gaussian distribution , from which the next generation is sampled after adding updated noise to prevent premature convergence. Among them, the multivariate Gaussian distribution The update is as follows:
[0113] ;
[0114] ;
[0115] in, Indicates updated , for the next iteration, Indicates updated , for the next iteration. Indicates adding noise The number of actions under the condition Indicates in Daidi The fitness of an individual action.
[0116] The SAC agent then interacts with the VVC simulated environment. The evaluation network randomly samples a mini-batch of experiences from the replay buffer and updates its network parameters using the single-step gradient descent method described in the SAC algorithm. The evaluation network and the mini-batch of experiences are then used to update the action network using the policy gradient calculated using the single-step gradient ascent method described in the SAC algorithm.
[0117] Finally, the updated evaluation network, action network, and multivariate Gaussian distribution will be used for the next generation until the optimal strategy is explored.
[0118] In this example, the main idea behind the ES-SAC algorithm is to combine the heuristic method of ES to generate various experiences. At the same time, SAC uses its powerful gradient-based method to learn from them. The SAC method after integrating ES enables pure DRL to learn various strategies and improve exploration capabilities.
[0119] Based on the principle of the above-mentioned improved algorithm, in order to optimize the optimal droop control curve for each photovoltaic, the present invention uses an improved deep reinforcement learning algorithm to solve the Markov decision process. Solving the Markov decision process can be regarded as an interaction process between the DRL agent and the environment, which can be specifically expressed as follows: the state space in the Markov decision process corresponds to the observation value that the DRL agent can obtain from the environment; the action space in the Markov decision process corresponds to the action output by the DRL agent based on the reward feedback from the environment; the reward function in the Markov decision process corresponds to the goal that the DRL agent needs to learn, and in the process of interacting with the environment, the reward value is maximized. The present invention is based on the continuous interaction between the improved ES-SAC agent and the active distribution network environment, so that the ES-SAC agent learns the optimal strategy, which corresponds to the optimal droop control parameters of each photovoltaic inverter, and then obtains the optimal droop control function based on the optimal parameters. The specific steps are as follows:
[0120] In step 1, the active power generation power of photovoltaics, the active power and reactive power of loads, and the node voltage in the active distribution network system are used as the state space of the ES-SAC algorithm reinforcement learning; the four voltage inflection points and the basic reactive power setting value in the droop control strategy of the photovoltaic inverter are used as the action space of the ES-SAC algorithm reinforcement learning; and the objective function and operation constraints of the system are used as the reward function of the ES-SAC algorithm reinforcement learning.
[0121] Step 2: Design the neural network parameters inside the ES-SAC agent to obtain the improved ES-SAC agent.
[0122] Step 3 simulates the interactive learning process between the active distribution network and the intelligent agent, and interactively generates experience data and puts it into the experience replay pool. At the same time, the improved ES-SAC intelligent agent can extract experience data from the experience pool and update the internal neural network parameters based on this, while generating new strategies to act on the distribution network environment.
[0123] Step 4: Repeat step 3 until the maximum number of training rounds is reached, and ensure that the reward shows an increasing convergence trend during the training process, thereby obtaining a trained improved ES-SAC intelligent agent. Based on the trained ES-SAC, the optimal active distribution network control strategy can be generated in real time.
[0124] This example is based on the continuous interaction between the improved ES-SAC agent and the active distribution network environment, which enables the ES-SAC agent to learn the optimal strategy. This strategy corresponds to the optimal droop control parameters for each PV inverter, and then the optimal droop control function is obtained based on the optimal parameters.
[0125] Compared with the prior art, the beneficial effects of the embodiments of the present invention are: the embodiments of the present invention construct a complete droop control function, study the active distribution network voltage-reactive power control model based on the complete droop control function, and use an improved deep reinforcement learning algorithm to optimize the droop control strategy, explore the superiority of the improved deep reinforcement learning algorithm in solving the voltage-reactive power control problem, and can enable photovoltaic inverters to participate in active distribution network voltage regulation.
[0126] Figure 4 2 is a schematic structural diagram of an active power distribution network according to an embodiment of the present invention.
[0127] Take the IEEE-33 node system in the power field as an example. Figure 4 As shown, one inverter-based PV is installed at each of nodes 4, 10, 13, 18, 21, 25, 27, 29, and 33. Figure 4 The PV locations and network topology consisting of 33 nodes are shown. The capacity of each PV inverter does not exceed 110% of its rated power. The base voltage of the IEEE-33-node system is 12.66 kV, and the set base power is 10 MVA. The node voltage per unit of the slack bus is 1 p.u., and the safe operating range of the node voltage is 0.95 pu-1.05 pu. Applying the methods provided in this example to this active distribution network allows for relevant verification experiments.
[0128] Figure 5 The figure is a structural block diagram of a droop control device for a photovoltaic inverter in an active power distribution network according to an embodiment of the present invention.
[0129] like Figure 5 As shown, the droop control device of the photovoltaic inverter in the active power distribution network may include:
[0130] The droop control function determination module 510 is configured to construct a droop control function for the photovoltaic inverter based on the following requirements: when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at maximum capacity; when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power;
[0131] an objective function construction module 520 for constructing an objective function using the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit, and the basic reactive power in the droop control function as decision variables and minimizing the active power loss of the active distribution network as a goal;
[0132] A redefinition module 530 is configured to define a state space, an action space, and a reward function of a Markov decision process based on the objective function and the constraints of the objective function;
[0133] The control strategy determination module 540 is configured to perform observation learning on the state space for each action in the action space based on the reward function and output the action with the maximum reward value as the target droop control strategy of the photovoltaic inverter.
[0134] In one embodiment, the redefinition module includes:
[0135] a state space determining unit, configured to determine a state space of a Markov decision process based on load active power, load reactive power, node voltage, and photovoltaic output active power of a node where the photovoltaic inverter is located in the active power distribution network;
[0136] an action space determining unit, configured to determine an action space of the Markov decision process based on decision variables in the objective function;
[0137] A reward function determination unit is used to determine the reward function of the Markov decision process based on the constraints of the objective function and the dependent variable in the objective function.
[0138] In one embodiment, the control strategy determination module includes:
[0139] an environment determining unit, configured to simulate an environment of the active power distribution network based on the state space;
[0140] A population initialization unit, configured to determine an initial population based on the action space;
[0141] a population iteration unit, configured to perform the following iterative operations on the population starting from the initial population: selecting the top K elite individuals in fitness ranking from the population based on the fitness of each individual in the population, adding noise to the top K elite individuals, sampling individuals from the top K elite individuals after the noise addition, and adding the sampling results to the population; using an intelligent agent to interactively learn between each individual in the population after the sampling results are added and the environment of the active power distribution network; and using the reward function to calculate the learning results of the intelligent agent to obtain the fitness of each individual in the population after the sampling results are added;
[0142] A strategy determination unit is configured to stop executing the iterative operation on the population when the number of iterations of the population reaches a preset threshold, and output the individual with the highest fitness in the population as the target droop control strategy of the photovoltaic inverter.
[0143] For the description of specific functions and examples of each module and submodule of the system in the embodiment of the present invention, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0144] In the technical solution of the present invention, the acquisition, storage and application of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0145] According to an embodiment of the present invention, the present invention further provides a system and a readable storage medium.
[0146] Figure 6 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0147] like Figure 6 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of electronic device 800. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0148] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0149] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the droop control method for a photovoltaic inverter in an active power distribution network. For example, in some embodiments, the droop control method for a photovoltaic inverter in an active power distribution network can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the droop control method for a photovoltaic inverter in an active power distribution network described above can be performed. Alternatively, in other embodiments, the calculation unit 801 may be configured to execute the droop control method for a photovoltaic inverter in an active power distribution network in any other appropriate manner (for example, by means of firmware).
[0150] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0151] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0152] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0154] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0155] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0156] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present invention can be achieved. This is not limited herein.
[0157] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A droop control method for a photovoltaic inverter in an active power distribution network, characterized in that: include: A droop control function for the photovoltaic inverter is constructed based on the following requirements: when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at maximum capacity; when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power; An objective function is constructed by taking the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit, and basic reactive power in the droop control function as decision variables and minimizing the active power loss of the active distribution network as the goal; Based on the objective function and the constraints of the objective function, defining a state space, an action space, and a reward function of a Markov decision process; Based on the reward function, observing and learning the state space for each action in the action space and outputting an action with the maximum reward value as the target droop control strategy of the photovoltaic inverter specifically includes: Based on the state space, simulating the environment of the active power distribution network; Based on the action space, determining an initial population; Starting from the initial population, the following iterative operation is performed on the population: based on the fitness of each individual in the population, the top K elite individuals ranked by fitness are selected from the population, noise is added to the top K elite individuals, individual sampling is performed from the top K elite individuals after the noise is added, and the sampling results are added to the population, and an intelligent agent is used to interactively learn each individual in the population after the sampling results are added with the environment of the active power distribution network, and the learning results of the intelligent agent are calculated using the reward function to obtain the fitness of each individual in the population after the sampling results are added; wherein, the population after the sampling results are added is used as the population for the next iteration, and the sampling results are: , Indicates that the parameter is In the policy network of state and noise The individual samples obtained by searching under the conditions of represents the hyperbolic tangent function; It is to use DNNs method in the state The following parameters are The mean value generated by the policy network; It is to use DNNs method in the state The following parameters are The covariance generated by the policy network; is from the standard normal distribution The random noise sampled in is parameterized as The policy network is the first K elite individuals after adding noise; When the number of iterations of the population reaches a preset threshold, the iterative operation on the population is stopped, and the individual with the highest fitness in the population is output as the target droop control strategy of the photovoltaic inverter.
2. The method according to claim 1, characterized in that The droop control function is: in, Indicates the active distribution network The photovoltaic inverter of each node is Reactive power at the moment, It is Nodes in The bus voltage observed at any moment, It is Nodes in The lower limit of the linear droop region voltage of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the linear droop region of the photovoltaic inverter at the moment, It is Nodes in The lower limit of the dead zone voltage of the photovoltaic inverter at the moment, It is Nodes in The upper voltage limit of the dead zone of the photovoltaic inverter at the moment, and It is Nodes in The slopes of the two linear droop regions of the photovoltaic inverter at the moment, It is Nodes in The basic reactive power of the photovoltaic inverter at the moment, It is Nodes in The maximum output reactive power of the PV inverter at the moment.
3. The method according to claim 1, characterized in that Defining the state space, action space, and reward function of the Markov decision process based on the objective function and the constraints of the objective function includes: Determining a state space of a Markov decision process based on load active power, load reactive power, node voltage, and photovoltaic output active power of a node where the photovoltaic inverter is located in the active power distribution network; determining an action space of the Markov decision process based on the decision variables in the objective function; A reward function of the Markov decision process is determined based on the constraints of the objective function and the dependent variable in the objective function.
4. The method according to claim 1, wherein The constraint conditions of the objective function include the power flow constraint conditions of the active power distribution network and the safe operation constraint conditions of the photovoltaic inverter.
5. A droop control device for a photovoltaic inverter in an active power distribution network, characterized in that: include: a droop control function determination module, configured to construct a droop control function for the photovoltaic inverter based on the following requirements: when the bus voltage in the active distribution network is lower than the lower voltage limit of the linear droop region, the photovoltaic inverter in the active distribution network generates reactive power at maximum capacity; when the bus voltage is higher than the upper voltage limit of the linear droop region, the photovoltaic inverter absorbs reactive power at maximum capacity; and when the bus voltage is between the lower voltage limit of the dead zone region and the upper voltage limit of the dead zone region, the photovoltaic inverter outputs basic reactive power; an objective function construction module, configured to construct an objective function using the linear droop region voltage lower limit, the linear droop region voltage upper limit, the dead zone region voltage upper limit, the dead zone region voltage lower limit, and basic reactive power in the droop control function as decision variables and minimizing the active power loss of the active distribution network as a goal; A redefinition module, configured to define a state space, an action space, and a reward function of a Markov decision process based on the objective function and the constraints of the objective function; a control strategy determination module, configured to perform observation learning on the state space for each action in the action space based on the reward function and output an action with the maximum reward value as a target droop control strategy for the photovoltaic inverter; Wherein, the control strategy determination module includes: an environment determining unit, configured to simulate an environment of the active power distribution network based on the state space; A population initialization unit, configured to determine an initial population based on the action space; A population iteration unit is configured to perform the following iterative operations on the population starting from the initial population: based on the fitness of each individual in the population, select the top K elite individuals in fitness ranking from the population, add noise to the top K elite individuals, sample individuals from the top K elite individuals after noise addition, and add the sampling results to the population; use an intelligent agent to interactively learn between each individual in the population after the sampling results are added and the environment of the active power distribution network, and use the reward function to calculate the learning results of the intelligent agent to obtain the fitness of each individual in the population after the sampling results are added; wherein the population after the sampling results are added is used as the population for the next iteration, and the sampling results are: , Indicates that the parameter is In the policy network of state and noise The individual samples obtained by searching under the conditions of represents the hyperbolic tangent function; It is to use DNNs method in the state The following parameters are The mean value generated by the policy network; It is to use DNNs method in the state The following parameters are The covariance generated by the policy network; is from the standard normal distribution The random noise sampled in is parameterized as The policy network is the first K elite individuals after adding noise; A strategy determination unit is configured to stop executing the iterative operation on the population when the number of iterations of the population reaches a preset threshold, and output the individual with the highest fitness in the population as the target droop control strategy of the photovoltaic inverter.
6. The device according to claim 5, characterized in that The redefinition module includes: a state space determining unit, configured to determine a state space of a Markov decision process based on load active power, load reactive power, node voltage, and photovoltaic output active power of a node where the photovoltaic inverter is located in the active power distribution network; an action space determining unit, configured to determine an action space of the Markov decision process based on decision variables in the objective function; A reward function determination unit is used to determine the reward function of the Markov decision process based on the constraints of the objective function and the dependent variable in the objective function.
7. A droop control system for a photovoltaic inverter in an active power distribution network, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the processor, and the processor is used to obtain the instructions from the memory and execute the instructions, so that the processor can execute the droop control method of the photovoltaic inverter in the active power distribution network according to any one of claims 1 to 4.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to be provided to a computer to instruct the computer to execute the droop control method for a photovoltaic inverter in an active power distribution network according to any one of claims 1 to 4.
Citation Information
Patent Citations
Photovoltaic inverter coordinated voltage reactive power control method and system and storage medium
CN118508537A
Distributed photovoltaic reactive voltage control method and system
CN118630776A