Environment perception based inverter adaptive power distribution method and system
By employing an environment-aware inverter adaptive power distribution method, which utilizes an intelligent agent to optimize dead zone and droop control parameters, the problem of voltage stability and reactive power distribution under load changes and light fluctuations in traditional inverter control methods is solved, achieving more efficient and stable grid operation.
Patent Information
- Application Number
- CN202511462918.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Traditional inverter reactive voltage control methods are difficult to adapt to changes in distribution network load and fluctuations in sunlight, resulting in poor voltage stability, unreasonable reactive power distribution and oscillations, which cannot meet the real-time, dynamic and intelligent control requirements of highly penetrated distributed energy distribution networks.
An environment-aware inverter adaptive power allocation method is adopted. By collecting distribution network data in real time, a pre-trained agent is used to adaptively optimize dead zone and droop control parameters, coordinate the reactive power allocation of multiple inverters, and avoid homogeneous response and oscillation.
It achieves adaptive control of the inverter, improves the stability and reliability of voltage at distribution network nodes, enhances the efficiency of reactive power distribution and the intelligence level of the system, and adapts to the access of distributed energy with high penetration rate.
Smart Images

Figure CN120933981B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reactive power regulation technology for inverters, specifically to an adaptive power allocation method and system for inverters based on environmental awareness. Background Technology
[0002] As the penetration rate of distributed photovoltaic and other renewable energy sources in distribution networks continues to increase, while promoting cleaner energy, it also brings severe challenges to the safe and stable operation of distribution networks. Among these challenges, voltage stability is particularly prominent. In distribution networks, reactive power balance is a key factor in maintaining node voltage stability. Traditional distribution networks rely mainly on the upstream grid and centralized compensation equipment for reactive power voltage support. However, this model is difficult to cope with the rapid and frequent voltage fluctuations caused by the random and intermittent output of distributed power sources.
[0003] To fully leverage the inherent regulation capabilities of distributed resources, inverter-based reactive power and voltage control technology has emerged. Among these, droop control, which draws inspiration from synchronous generator characteristics, is widely used. By establishing a linear relationship between inverter output voltage and reactive power, localized autonomous regulation is achieved. However, traditional droop control has gradually revealed inherent limitations in practical applications. First, its control parameters are typically fixed, making it difficult to adapt to complex operating conditions such as changes in distribution network load and fluctuations in sunlight, resulting in low regulation efficiency and, in extreme cases, even causing node voltage exceedances. Second, when multiple inverters distributed in different locations all employ fixed droop characteristics, they will produce homogeneous responses to the same voltage changes, lacking a coordination mechanism. This can easily lead to reactive power oscillations between inverters, affecting not only the stability of control performance but also potentially accelerating equipment wear. Furthermore, while a fixed deadband setting can prevent the inverter from being overly sensitive to small voltage fluctuations, it may result in a slow response and lack of flexibility when necessary. In summary, existing methods are inadequate in terms of voltage stability, reactive power allocation rationality, and equipment response efficiency, making it difficult to meet the demands of modern high-penetration distributed energy distribution networks for real-time, dynamic, and intelligent reactive power control. Summary of the Invention
[0004] The purpose of this invention is to solve the problems mentioned in the background art, such as the difficulty of adapting fixed control parameters to complex operating conditions and the unreasonable or even oscillating reactive power distribution caused by the lack of coordination among multiple inverters. Therefore, this invention proposes an inverter adaptive power distribution method and system based on environmental awareness.
[0005] A first aspect of this invention provides an environment-aware inverter adaptive power allocation method, the method comprising:
[0006] Real-time acquisition of operational status data of the target distribution network; the operational status data includes the node voltage of each inverter connected to the bus;
[0007] The operational status data is preprocessed to obtain the first environmental feature data;
[0008] The first environmental feature data is input into the pre-trained first agent to obtain the dead zone parameters of each inverter; the dead zone parameters include the upper and lower limits of the dead zone voltage.
[0009] The first environmental feature data and the dead zone parameter are combined to obtain the second environmental feature data;
[0010] The second environmental feature data is input into the pre-trained second agent to obtain the droop control parameters for each inverter; the droop control parameters include reference reactive power, undervoltage regulation coefficient and overvoltage regulation coefficient.
[0011] The reactive power of each inverter is determined based on the node voltage connected to the bus and its corresponding dead zone parameters.
[0012] Optionally, the first agent and the second agent are jointly trained using reinforcement learning. The training process includes:
[0013] Establish a simulation environment for the target distribution network to provide state feedback during the training process;
[0014] During each training cycle of the second agent, fixed dead zone voltage upper and lower limits are randomly generated as environmental parameters;
[0015] The second agent outputs the droop control parameters of each inverter at the first preset time step, and controls the reactive power based on the droop control parameters; it obtains the state feedback of the simulation environment and calculates the reward according to the preset reward function, stores the interaction data in the experience buffer; and updates the policy network of the second agent based on the soft actor critic algorithm.
[0016] After completing the individual training of the second agent:
[0017] The first agent outputs the dead-zone parameters of each inverter at the second preset time step and uses these dead-zone parameters as environmental parameters. At each time step, the second agent is invoked to interact with the simulation environment, and the multi-step rewards obtained by the second agent within that time step are summed as the reward of the first agent. The interaction data is stored in the experience buffer. The policy network of the first agent is updated based on the soft actor critic algorithm.
[0018] Optionally, the reward function for the second agent is:
[0019] ;
[0020] in, It is the reward for time step t; PL t It refers to the power loss of the power system; VE t It is a penalty for exceeding the node voltage limit; , These are weighting coefficients; It is the voltage over-limit penalty coefficient; It is the node voltage at which inverter i is connected to the bus at time step t+1; and These are the minimum and maximum allowable voltages when inverter i is connected to the bus, respectively. It is the ReLU activation function; It is a collection of inverters.
[0021] Optionally, determining the reactive power of each inverter based on the node voltage connected to the bus and its corresponding dead-time parameters includes:
[0022] If the node voltage of inverter i connected to the bus is If it is located within its dead zone, then the reactive power of inverter i is determined to be its reference reactive power. ;
[0023] If the node voltage of inverter i connected to the bus is Less than its dead zone voltage lower limit Then its reactive power is determined as: ;
[0024] If the node voltage of inverter i connected to the bus is Greater than its dead zone voltage limit Then its reactive power is determined as: ;
[0025] in, This is the reactive power of inverter i. This is the maximum reactive power capacity of inverter i; It is the undervoltage regulation coefficient, which is a negative number; It is the overpressure regulation coefficient, which is a negative number.
[0026] Optionally, the method further includes:
[0027] The first intelligent agent is invoked at a first preset frequency to output the dead zone parameters of each inverter during the target time period;
[0028] During the target time period, the second intelligent agent is invoked at a second preset frequency to output the droop control parameters of each inverter, and the reactive power of each inverter is dynamically adjusted under the constraint of the dead zone parameter.
[0029] A second aspect of this invention provides an environment-aware inverter adaptive power distribution system, the system comprising:
[0030] An environmental sensing module is used to collect real-time operating status data of the target power distribution network; the operating status data includes the node voltage of each inverter connected to the bus.
[0031] The preprocessing module is used to preprocess the running status data to obtain the first environmental feature data;
[0032] The long-term allocation module is used to input the first environmental feature data into the pre-trained first agent to obtain the dead-zone parameters of each inverter; the dead-zone parameters include the upper and lower limits of the dead-zone voltage.
[0033] The parameter integration module is used to combine the first environmental feature data and the dead zone parameter to obtain the second environmental feature data.
[0034] The short-term allocation module is used to input the second environmental feature data into the pre-trained second agent to obtain the droop control parameters of each inverter; the droop control parameters include reference reactive power, undervoltage regulation coefficient and overvoltage regulation coefficient.
[0035] The dynamic adjustment module is used to determine the reactive power of each inverter based on the node voltage of each inverter connected to the bus and its corresponding dead zone parameters.
[0036] Optionally, the system further includes a joint training module for jointly training the first agent and the second agent through reinforcement learning; the joint training module includes:
[0037] The environment simulation module is used to establish a simulation environment for the target power distribution network and to provide state feedback during the training process.
[0038] The missing parameter completion module is used to randomly generate fixed dead zone voltage upper and lower limits as environmental parameters in each training cycle of the second agent;
[0039] The second agent interactive learning module is used to call the second agent to output the droop control parameters of each inverter at the first preset time step, and control the reactive power based on the droop control parameters; obtain the state feedback of the simulation environment and calculate the reward according to the preset reward function, store the interactive data in the experience buffer; and update the policy network of the second agent based on the soft actor critic algorithm.
[0040] The first agent interaction learning module is used to: after completing the individual training of the second agent, call the first agent to output the dead zone parameters of each inverter at a second preset time step, and use the dead zone parameters as environmental parameters; at each time step, call the second agent to interact with the simulation environment, and sum the multi-step rewards obtained by the second agent in that time step as the reward of the first agent, and store the interaction data in the experience buffer; update the policy network of the first agent based on the soft actor critic algorithm.
[0041] Optionally, the reward function of the second agent is:
[0042] ;
[0043] in, It is the reward for time step t; PL t It refers to the power loss of the power system; VE t It is a penalty for exceeding the node voltage limit; , These are weighting coefficients; It is the voltage over-limit penalty coefficient; It is the node voltage at which inverter i is connected to the bus at time step t+1; and These are the minimum and maximum allowable voltages when inverter i is connected to the bus, respectively. It is the ReLU activation function; It is a collection of inverters.
[0044] Optionally, the dynamic adjustment module includes:
[0045] The voltage regulation module is used to adjust the node voltage of inverter i connected to the bus. If it is located within its dead zone, then the reactive power of inverter i is determined to be its reference reactive power. ;
[0046] The undervoltage regulation module is used to adjust the node voltage of inverter i connected to the bus. Less than its dead zone voltage lower limit Then its reactive power is determined as:
[0047] ;
[0048] The overvoltage regulation module is used to adjust the node voltage of inverter i connected to the bus. Greater than its dead zone voltage limit Then its reactive power is determined as:
[0049] ;
[0050] in, This is the reactive power of inverter i. This is the maximum reactive power capacity of inverter i; It is the undervoltage regulation coefficient, which is a negative number; It is the overpressure regulation coefficient, which is a negative number.
[0051] Optionally, the long-term allocation module is used to call the first intelligent agent at a first preset frequency and output the dead zone parameters of each inverter in the target time period; the short-term allocation module is used to call the second intelligent agent at a second preset frequency in the target time period and output the droop control parameters of each inverter.
[0052] The beneficial effects of this invention are:
[0053] By introducing environmental perception and intelligent agent decision-making mechanisms, the upper and lower limits of inverter dead-zone voltage and droop control parameters can be dynamically adjusted according to the real-time operating status of the distribution network. This achieves adaptive optimization of control parameters, overcoming the problem that traditional fixed parameters cannot cope with complex operating conditions such as load and light fluctuations. Simultaneously, it coordinates the reactive power distribution of multiple inverters, avoiding homogeneous responses and reactive power oscillations, making system regulation more efficient and stable. This significantly improves the stability, reliability, and real-time intelligence level of distribution network node voltages, meeting the needs of high-penetration distributed energy access. Attached Figure Description
[0054] Figure 1 A flowchart of an inverter adaptive power allocation method based on environment awareness provided for an embodiment of the present invention;
[0055] Figure 2 This is a comparison chart of photovoltaic voltage simulation on a sunny day provided in an embodiment of the present invention;
[0056] Figure 3 This is a simulation comparison chart of photovoltaic voltage on cloudy days provided in an embodiment of the present invention;
[0057] Figure 4 This is an architecture diagram of an inverter adaptive power distribution system based on environment awareness, provided for an embodiment of the present invention. Detailed Implementation
[0058] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0059] This invention provides an environment-aware adaptive power allocation method for inverters. See also... Figure 1 , Figure 1 A flowchart illustrating an environment-aware adaptive power allocation method for inverters, provided as an embodiment of the present invention, is shown. The method includes the following steps:
[0060] S101, real-time acquisition of operating status data of the target distribution network;
[0061] S102, Preprocess the operating status data to obtain the first environmental characteristic data;
[0062] S103, input the first environmental feature data into the pre-trained first agent to obtain the dead zone parameters of each inverter;
[0063] S104, combine the first environmental feature data and the dead zone parameter to obtain the second environmental feature data;
[0064] S105, input the second environmental feature data into the pre-trained second agent to obtain the droop control parameters of each inverter;
[0065] S106, determine the reactive power of each inverter based on the node voltage connected to the bus and its corresponding dead zone voltage upper and lower limits.
[0066] The operating status data includes the node voltage of each inverter connected to the bus, the active power of each inverter, the active and reactive loads of each bus, and the power loss of each branch. Dead zone parameters include the upper and lower limits of the dead zone voltage. Droop control parameters include the reference reactive power, undervoltage regulation coefficient, and overvoltage regulation coefficient.
[0067] This invention provides an inverter adaptive power allocation method based on environmental perception. By introducing environmental perception and intelligent agent decision-making mechanisms, it can dynamically adjust the upper and lower limits of the inverter dead zone voltage and the droop control parameters according to the real-time operating status of the distribution network. This achieves adaptive optimization of control parameters, overcoming the problem that traditional fixed parameters cannot cope with complex operating conditions such as load and light fluctuations. Simultaneously, it coordinates the reactive power allocation of multiple inverters, avoiding homogeneous responses and reactive power oscillations, making system regulation more efficient and stable. This significantly improves the stability, reliability, and real-time intelligence level of distribution network node voltages, meeting the needs of high-penetration distributed energy access.
[0068] In one implementation, step S106, based on the node voltage of each inverter connected to the bus and its corresponding dead-zone voltage upper and lower limits, determines the reactive power of each inverter, including:
[0069] Scenario 1: Dead-zone regulation: If the node voltage of inverter i connected to the bus... If it is located within its dead zone, then the reactive power of inverter i is determined. For its reference reactive power . A positive value indicates that reactive power is injected into the power grid, while a negative value indicates that reactive power is absorbed from the power grid.
[0070] Scenario 2: Undervoltage regulation: If the node voltage of inverter i connected to the bus is... Less than its dead zone voltage lower limit Then its reactive power for: .in, It is the maximum reactive power capacity of the target inverter i. It is the undervoltage regulation coefficient, which is a negative number.
[0071] Scenario 3: Overvoltage regulation: If the node voltage of inverter i connected to the bus... Greater than its dead zone voltage limit Then its reactive power for: ;in, It is the overpressure regulation coefficient, which is a negative number.
[0072] This implementation method controls the inverter's reactive power in segments under three conditions: dead zone, undervoltage, and overvoltage. Within the normal voltage fluctuation range (dead zone), a reference reactive power is maintained to avoid frequent tripping, reduce inverter losses, and extend inverter lifespan. When the voltage is below or above the dead zone limits, the reactive power output is dynamically adjusted according to undervoltage or overvoltage regulation coefficients and is constrained by maximum capacity. This allows for automatic and flexible support or absorption of reactive power under different operating conditions, stabilizing node voltage and reducing losses. It also prevents inverter overload and controls oscillations, improving the safety, economy, and power quality of the power grid.
[0073] In one implementation, a first intelligent agent is invoked at a first preset frequency to output the dead-time parameters of each inverter during the target time period. For example, the first intelligent agent is invoked once every hour, providing the dead-time parameters at the current moment, which remain unchanged for the next hour. The hour following this current moment is the target time period.
[0074] Within the target time period, the second intelligent agent is invoked at a second preset frequency to output the droop control parameters for each inverter. Under the constraint of the dead-zone parameter, the reactive power of each inverter is dynamically adjusted. For example, the second intelligent agent is invoked once every minute to provide the droop control parameters, which remain unchanged for the next minute.
[0075] In power distribution networks, photovoltaic (PV) output exhibits minute-level abrupt changes (e.g., cloud cover), while user load shows a long-term intraday trend (e.g., low load during the day and high load at night), with significant differences in the time scales of these two types of changes. By dividing the work between hourly-level decision dead zones and minute-level decision control coefficients, the first intelligent agent can optimize the dead zone based on long-term voltage distribution and system trends to adapt to slow-changing characteristics; the second intelligent agent can quickly respond to minute-level PV fluctuations by adjusting the droop control coefficient to suppress voltage deviations in real time, avoiding the problems of insufficient adaptation to slow-changing trends and delayed response to fast-changing fluctuations caused by fixed-time-scale control.
[0076] In one embodiment, the first agent and the second agent are jointly trained through reinforcement learning, and the training process includes:
[0077] Phase 1: Establish a simulation environment for the target distribution network to provide state feedback during the training process.
[0078] Phase Two: In this phase, only the second agent is trained; the first agent does not participate in parameter learning. A fixed environment is provided for the second agent by randomly generating dead zones, ensuring that the second agent can stably learn the droop control parameter optimization logic based on the dead zones. Specifically, this includes:
[0079] Initialize the parameters of the neural network (Actor / Critic network) of the second agent, and simultaneously initialize the training steps and time steps. An experience replay buffer is set up to store the four-tuple of state, action, reward, and next state during the interaction process. The initial dead zone range of each inverter is generated by uniform random distribution and fixed as the environmental parameters of the current training cycle.
[0080] Perform the following operations, using minutes as the time step:
[0081] Step 1, Status Input: Obtain the operating status of the distribution network. This includes the node voltage of each inverter connected to the bus, the active power of each inverter, the active and reactive loads of each bus, the power loss of each branch, and the dead zone parameters of each inverter.
[0082] Step 2, Action Output: Output the droop control parameters of each inverter through the Actor network. The reactive power is controlled based on the droop control parameters to perform simulation and obtain the next state. .
[0083] Step 3, Reward Calculation: Calculate the reward based on the preset reward function. The reward function is:
[0084] ;
[0085] in, It is the reward for time step t; PL t It refers to the power loss of the power system; VE t It is a penalty for exceeding the node voltage limit; , These are weighting coefficients. Set to 1, Adjust during training; It is the voltage over-limit penalty coefficient; The node voltage of inverter i connected to the bus at time step t+1; and These are the minimum and maximum allowable voltages when inverter i is connected to the bus, respectively. It is a ReLU activation function (outputs the original value only when the input is positive, otherwise outputs 0, that is, it only penalizes the case where the voltage is below the lower limit or above the upper limit). It is a collection of inverters.
[0086] Step 4, Experience Storage: ( , , , Store it in the experience replay buffer.
[0087] Step 5, Network Update: When the amount of experience meets the batch size, update the Critic network and Actor network based on the Soft Actor Critic (SAC) algorithm.
[0088] Repeat the above process until the reward curve stabilizes, indicating that the second agent has converged.
[0089] The third stage: After the second agent converges, the first agent begins to participate. The two learn together based on the SAC algorithm. The goal is to enable the first agent to learn to adapt to the dead zone optimization of the droop control parameters, while the second agent further optimizes the droop control parameters based on the dynamic dead zone.
[0090] Initialize the neural network (Actor / Critic network) parameters of the first agent (same as the network structure of the second agent), set its experience replay buffer, and use the converged network parameters of the second agent as initial values.
[0091] Using hours as the time step, perform the following operations:
[0092] Step 1, Status Input: Obtain the operating status of the distribution network. Compared to the input of the second agent, the only difference is the absence of a dead zone parameter.
[0093] Step 2, Action Output: The first agent outputs actions based on its running state. The dead-time parameters of each inverter are output through its Actor network. .
[0094] Step 3, Secondary Learning: Within multiple minute steps of the current hour step, the second agent ( , Given a state of ), the process of "outputting droop control parameters - calculating rewards - storing experience" is repeated, with the network updated in each round. The running status after the simulation at that hour step is also given. .
[0095] Step 4, Reward Calculation: The sum of the rewards for the second agent across all minute steps within the current hour step is defined as the reward for the first agent in the current hour step. This ensures that the effects of the two agents are bound together.
[0096] Step 5: Experience storage and network updates.
[0097] In this embodiment of the invention, the reward function simultaneously considers power loss and node voltage over-limit penalties, enabling the trained agent strategy to maintain voltage stability while reducing losses, thereby improving grid operating efficiency and reliability. The second agent first learns a droop control strategy to avoid training instability caused by simultaneously optimizing two sets of high-dimensional nonlinear parameters, ensuring smooth convergence of the reward curve. The first agent participates in training after the second agent converges, with the first agent updating dead-zone parameters (slow variables) hourly and the second agent updating droop parameters (fast variables) minutely. This ensures both long-term optimization and rapid response to short-term voltage fluctuations, allowing the dead-zone parameters and droop control parameters to be mutually adapted, forming a joint strategy of "dead-zone optimization + droop control optimization," thus improving the overall reactive power allocation effect.
[0098] In one embodiment, see Figure 2 , Figure 2 This is a comparison chart of photovoltaic voltage simulation on a sunny day, provided as an embodiment of the present invention. See also... Figure 3 , Figure 3 The above is a simulation comparison chart of photovoltaic voltage on cloudy days provided in an embodiment of the present invention.
[0099] The test system model adopts the IEEE 33-node distribution system, with photovoltaic (PV) access locations: a total of 9 inverter-type PV systems, located at nodes 4, 10, 13, 18, 21, 25, 27, 29, and 33. Voltage reference: The relaxation node voltage is set to 1 p.u., with an allowable voltage range of 0.95 p.u. to 1.05 p.u.
[0100] The load data required for this simulation uses residential load curves from the "Typical Load Curve Dataset" released by the State Grid Corporation of China. Photovoltaic power generation data is sourced from historical meteorological data released by the National Meteorological Science Data Center of the China Meteorological Administration. Power curves are generated based on hourly "horizontal surface irradiance" and "ambient temperature" data for one year, calculated using a standard photovoltaic power generation model. To simulate rapid fluctuations in sunlight under cloudy conditions, higher temporal resolution (15-minute intervals) irradiance data is used to generate photovoltaic output in the corresponding scenario.
[0101] Figure 2Part (a) shows the voltage distribution of each node under the condition of fixed droop control parameters for each photovoltaic inverter on a sunny day, while part (b) shows the voltage distribution of each node under the condition of adjusting the droop control parameters of the photovoltaic inverter using the inverter adaptive power allocation method of the present invention on a sunny day. The comparison shows that, compared to the case with fixed parameters, the simulation results of the present invention show a more concentrated and stable voltage, with all node voltages effectively controlled within a safe range. Therefore, the method of the present invention significantly improves voltage stability and eliminates voltage exceedance problems on sunny days by coordinating and optimizing the droop control functions of each photovoltaic inverter.
[0102] Figure 3 The blue curve represents the voltage variation of bus 33 under the condition of fixed droop control parameters for each photovoltaic inverter on a cloudy day, while the red curve represents the voltage variation of bus 33 under the condition of adjusting the droop control parameters of the photovoltaic inverters using the inverter adaptive power allocation method of this invention on a cloudy day. Compared to the simulation results with drastic voltage fluctuations and obvious limit exceedances when the parameters are fixed, the simulation results of this invention show that the voltage is effectively stabilized within a safe range. Therefore, even on cloudy days with drastic fluctuations in photovoltaic output, this invention can quickly respond to changes in system state, maintain voltage safety, and ensure the robustness and adaptability of the photovoltaic power generation system.
[0103] This invention provides an environment-aware inverter adaptive power distribution system. See also... Figure 4 , Figure 4 This is an architecture diagram of an environment-aware inverter adaptive power distribution system provided in an embodiment of the present invention. The system includes:
[0104] The environmental sensing module is used to collect real-time operational status data of the target power distribution network.
[0105] The preprocessing module is used to preprocess the running status data to obtain the first environmental feature data.
[0106] The long-term allocation module is used to input the first environmental feature data into the pre-trained first agent to obtain the dead-zone parameters of each inverter.
[0107] The parameter integration module is used to combine the first environmental feature data and the dead zone parameters to obtain the second environmental feature data.
[0108] The short-term allocation module is used to input the second environmental feature data into the pre-trained second agent to obtain the droop control parameters of each inverter.
[0109] The dynamic adjustment module is used to determine the reactive power of each inverter based on the node voltage of each inverter connected to the bus and its corresponding dead zone parameters.
[0110] The operating status data includes the node voltage of each inverter connected to the bus, the active power of each inverter, the active and reactive loads of each bus, and the power loss of each branch. Dead zone parameters include the upper and lower limits of the dead zone voltage. Droop control parameters include the reference reactive power, undervoltage regulation coefficient, and overvoltage regulation coefficient.
[0111] This invention provides an inverter adaptive power distribution system based on environmental perception. By introducing environmental perception and intelligent agent decision-making mechanisms, it can dynamically adjust the upper and lower limits of the inverter dead-zone voltage and the droop control parameters according to the real-time operating status of the distribution network. This achieves adaptive optimization of control parameters, overcoming the problem that traditional fixed parameters cannot cope with complex operating conditions such as load and light fluctuations. Simultaneously, it coordinates the reactive power distribution of multiple inverters, avoiding homogeneous responses and reactive power oscillations, making the system regulation more efficient and stable. This significantly improves the stability, reliability, and real-time intelligence level of distribution network node voltages, meeting the needs of high-penetration distributed energy access.
[0112] In one implementation, the long-term allocation module calls the first intelligent agent at a first preset frequency to output the dead-zone parameters of each inverter during the target time period; the short-term allocation module calls the second intelligent agent at a second preset frequency during the target time period to output the droop control parameters of each inverter.
[0113] In one embodiment, the system further includes a joint training module for jointly training the first agent and the second agent through reinforcement learning. The joint training module includes:
[0114] The environment simulation module is used to establish a simulation environment for the target power distribution network and to provide state feedback during the training process.
[0115] The missing parameter completion module is used to randomly generate fixed dead zone voltage upper and lower limits as environmental parameters in each training cycle of the second agent.
[0116] The second agent interactive learning module is used to call the second agent to output the droop control parameters of each inverter at a first preset time step, and control the reactive power based on the droop control parameters; obtain the state feedback of the simulation environment and calculate the reward according to the preset reward function, store the interactive data in the experience buffer; and update the policy network of the second agent based on the soft actor critic algorithm.
[0117] The first agent interaction learning module is used to: after completing the individual training of the second agent, call the first agent to output the dead zone parameters of each inverter at a second preset time step, and use the dead zone parameters as environmental parameters; at each time step, call the second agent to interact with the simulation environment, and sum the multi-step rewards obtained by the second agent in that time step as the reward of the first agent, and store the interaction data in the experience buffer; update the policy network of the first agent based on the soft actor critic algorithm.
[0118] In one embodiment, the dynamic adjustment module includes:
[0119] The voltage regulation module is used to adjust the node voltage of inverter i connected to the bus. If it is located within its dead zone, then the reactive power of inverter i is determined to be its reference reactive power. .
[0120] The undervoltage regulation module is used to adjust the node voltage of inverter i connected to the bus. Less than its dead zone voltage lower limit Then its reactive power is determined as:
[0121] .
[0122] The overvoltage regulation module is used to adjust the node voltage of inverter i connected to the bus. Greater than its dead zone voltage limit Then its reactive power is determined as:
[0123] .
[0124] in, This is the reactive power of inverter i. This is the maximum reactive power capacity of inverter i; It is the undervoltage regulation coefficient, which is a negative number; It is the overpressure regulation coefficient, which is a negative number.
[0125] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall fall within the scope of the present invention.
Claims
1. An inverter adaptive power allocation method based on environment perception, characterized in that, The method includes: Real-time acquisition of operational status data of the target distribution network; the operational status data includes the node voltage of each inverter connected to the bus; The operational status data is preprocessed to obtain the first environmental feature data; The first environmental feature data is input into the pre-trained first agent to obtain the dead zone parameters of each inverter; the dead zone parameters include the upper and lower limits of the dead zone voltage. The first environmental feature data and the dead zone parameter are combined to obtain the second environmental feature data; The second environmental feature data is input into the pre-trained second agent to obtain the droop control parameters for each inverter; the droop control parameters include reference reactive power, undervoltage regulation coefficient and overvoltage regulation coefficient. The reactive power of each inverter is determined based on the node voltage connected to the bus and its corresponding dead zone parameters and droop control parameters. The first and second agents are jointly trained using reinforcement learning. The training process includes: Establish a simulation environment for the target distribution network to provide state feedback during the training process; During each training cycle of the second agent, fixed dead zone voltage upper and lower limits are randomly generated as environmental parameters; The second agent outputs the droop control parameters of each inverter at the first preset time step, and controls the reactive power based on the droop control parameters; it obtains the state feedback of the simulation environment and calculates the reward according to the preset reward function, stores the interaction data in the experience buffer; and updates the policy network of the second agent based on the soft actor critic algorithm. After completing the individual training of the second agent: The first agent outputs the dead-zone parameters of each inverter at the second preset time step and uses these dead-zone parameters as environmental parameters. At each time step, the second agent is invoked to interact with the simulation environment, and the multi-step rewards obtained by the second agent within that time step are summed as the reward of the first agent. The interaction data is stored in the experience buffer. The policy network of the first agent is updated based on the soft actor critic algorithm.
2. The inverter adaptive power allocation method based on environment perception according to claim 1, characterized in that, The reward function for the second agent is: ; in, It is the reward for time step t; PL t It refers to the power loss of the power system; VE t It is a penalty for exceeding the node voltage limit; , These are weighting coefficients; It is the voltage over-limit penalty coefficient; It is the node voltage at which inverter i is connected to the bus at time step t+1; and These are the minimum and maximum allowable voltages when inverter i is connected to the bus, respectively. It is the ReLU activation function; It is a collection of inverters.
3. The inverter adaptive power allocation method based on environment perception according to claim 1, characterized in that, The determination of the reactive power of each inverter based on the node voltage connected to the bus and its corresponding dead-time parameters includes: If the node voltage of inverter i connected to the bus is If it is located within its dead zone, then the reactive power of inverter i is determined to be its reference reactive power. ; If the node voltage of inverter i connected to the bus is Less than its dead zone voltage lower limit Then its reactive power is determined as: ; If the node voltage of inverter i connected to the bus is Greater than its dead zone voltage limit Then its reactive power is determined as: ; in, This is the reactive power of inverter i. This is the maximum reactive power capacity of inverter i; It is the undervoltage regulation coefficient, which is a negative number; It is the overpressure regulation coefficient, which is a negative number.
4. The inverter adaptive power allocation method based on environment perception according to claim 1, characterized in that, The method further includes: The first intelligent agent is invoked at a first preset frequency to output the dead zone parameters of each inverter during the target time period; During the target time period, the second intelligent agent is invoked at a second preset frequency to output the droop control parameters of each inverter, and the reactive power of each inverter is dynamically adjusted under the constraint of the dead zone parameter.
5. An inverter adaptive power distribution system based on environment perception, characterized in that, The system includes: An environmental sensing module is used to collect real-time operating status data of the target power distribution network; the operating status data includes the node voltage of each inverter connected to the bus. The preprocessing module is used to preprocess the running status data to obtain the first environmental feature data; The long-term allocation module is used to input the first environmental feature data into the pre-trained first agent to obtain the dead-zone parameters of each inverter; the dead-zone parameters include the upper and lower limits of the dead-zone voltage. The parameter integration module is used to combine the first environmental feature data and the dead zone parameter to obtain the second environmental feature data. The short-term allocation module is used to input the second environmental feature data into the pre-trained second agent to obtain the droop control parameters of each inverter; the droop control parameters include reference reactive power, undervoltage regulation coefficient and overvoltage regulation coefficient. The dynamic adjustment module is used to determine the reactive power of each inverter based on the node voltage of each inverter connected to the bus and its corresponding dead zone parameters and droop control parameters. A joint training module is used to jointly train the first agent and the second agent through reinforcement learning; the joint training module includes: The environment simulation module is used to establish a simulation environment for the target power distribution network and to provide state feedback during the training process. The missing parameter completion module is used to randomly generate fixed dead zone voltage upper and lower limits as environmental parameters in each training cycle of the second agent; The second agent interactive learning module is used to call the second agent to output the droop control parameters of each inverter at the first preset time step, and control the reactive power based on the droop control parameters; obtain the state feedback of the simulation environment and calculate the reward according to the preset reward function, store the interactive data in the experience buffer; and update the policy network of the second agent based on the soft actor critic algorithm. The first agent interactive learning module is used to: after completing the individual training of the second agent, call the first agent to output the dead zone parameters of each inverter at a second preset time step, and use the dead zone parameters as environmental parameters; at each time step, call the second agent to interact with the simulation environment, and sum the multi-step rewards obtained by the second agent within that time step as the reward of the first agent, and store the interaction data in the experience buffer; update the policy network of the first agent based on the soft actor critic algorithm.
6. The inverter adaptive power distribution system based on environment perception according to claim 5, characterized in that, The reward function for the second agent is: ; in, It is the reward for time step t; PL t It refers to the power loss of the power system; VE t It is a penalty for exceeding the node voltage limit; , These are weighting coefficients; It is the voltage over-limit penalty coefficient; It is the node voltage at which inverter i is connected to the bus at time step t+1; and These are the minimum and maximum allowable voltages when inverter i is connected to the bus, respectively. It is the ReLU activation function; It is a collection of inverters.
7. The inverter adaptive power distribution system based on environment perception according to claim 5, characterized in that, The dynamic adjustment module includes: The voltage regulation module is used to adjust the node voltage of inverter i connected to the bus. If it is located within its dead zone, then the reactive power of inverter i is determined to be its reference reactive power. ; The undervoltage regulation module is used to adjust the node voltage of inverter i connected to the bus. Less than its dead zone voltage lower limit Then its reactive power is determined as: ; The overvoltage regulation module is used to adjust the node voltage of inverter i connected to the bus. Greater than its dead zone voltage limit Then its reactive power is determined as: ; in, This is the reactive power of inverter i. This is the maximum reactive power capacity of inverter i; It is the undervoltage regulation coefficient, which is a negative number; It is the overpressure regulation coefficient, which is a negative number.
8. The inverter adaptive power distribution system based on environment perception according to claim 5, characterized in that, The long-term allocation module is used to call the first intelligent agent at a first preset frequency and output the dead zone parameters of each inverter in the target time period; the short-term allocation module is used to call the second intelligent agent at a second preset frequency in the target time period and output the droop control parameters of each inverter.
Citation Information
Patent Citations
Reactive voltage control method based on multi-time-scale multi-agent deep reinforcement learning
CN113363997A
Improved droop curve dynamic voltage regulation method and voltage regulation system for closed-loop verification
CN115693770A
Voltage regulation method and device, equipment and storage medium
CN115733147A
Droop control method and system for photovoltaic inverter in active power distribution network
CN119966008A
Multi-end collaborative voltage treatment method and system for power distribution network with high-penetration-rate photovoltaic access, and storage medium
WO2023093537A1