Reactive power compensation adaptive control method based on distribution network

By collecting distribution network parameters in real time and combining the DDPG algorithm and OU noise mechanism, the load status is identified and reactive power is adjusted. This solves the problem of response hysteresis and robustness of existing reactive power compensation methods in complex environments, and realizes high-precision, adaptive reactive power regulation and scheduling optimization.

CN120955695BActive Publication Date: 2026-04-03LIANYUNGANG ZHITUO ENERGY SAVING ELECTRIC CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing reactive power compensation methods suffer from slow response, low adjustment accuracy, and poor robustness in scenarios with frequent load fluctuations, diverse load types, or uncertainties in distributed power sources. They are unable to meet the needs of modern power distribution systems for fine-grained reactive power adjustment and lack the ability to recognize and dynamically learn the system status in real time.

Method used

By collecting power distribution network parameters in real time, identifying load status and calculating reactive power demand, and combining the DDPG algorithm, OU noise mechanism and curiosity mechanism, reactive power is adjusted. Control commands are issued using the IEC 61850 protocol, the feedback effect of equipment is monitored and reactive power is optimized, and an adjustment target table is constructed to achieve adaptive control.

Benefits of technology

It significantly improves the accuracy, response efficiency and system stability of reactive power compensation control, enhances the ability to perceive and adaptively regulate multi-source heterogeneous information, and optimizes the standardization of reactive power dispatch.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120955695B_ABST
    Figure CN120955695B_ABST
Patent Text Reader

Abstract

This invention discloses an adaptive control method for reactive power compensation based on a distribution network, relating to the field of power system technology. The method includes real-time acquisition and preprocessing of distribution network operating parameters; identifying load status and calculating reactive power demand through data processing; acquiring the status and parameter configuration of reactive power compensation equipment; obtaining a regulation target table by calculating reactive power deviation; and scheduling and allocating equipment. Based on reactive power demand, reactive power deviation, and equipment matching data, the method uses the DDPG algorithm combined with an OU noise mechanism and a curiosity mechanism to regulate and control reactive power, and then distributes the regulated reactive power to the compensation equipment for reactive power compensation. This invention achieves the adaptive control objective for reactive power compensation in complex power grid environments by constructing a reinforcement learning framework that integrates perception, decision-making, learning, and control, combined with multi-source state modeling, reward mechanism-driven approaches, refined Q-value evaluation, continuous strategy optimization, and physical constraint fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system technology, and in particular to an adaptive control method for reactive power compensation based on distribution networks. Background Technology

[0002] With the development of smart grids and the widespread integration of distributed power sources, the operating environment of modern distribution networks is becoming increasingly complex, and the system's requirements for power quality are constantly increasing. Reactive power, as a crucial parameter in grid operation, not only affects voltage stability but also directly relates to energy transmission efficiency and equipment lifespan. Traditional reactive power compensation methods often employ setpoint control or regulation strategies based on preset rules, such as capacitor bank group switching and static var generator (SVG) setpoint operation. These methods often exhibit problems such as slow response, low regulation accuracy, and poor robustness when facing scenarios with frequent load fluctuations, diverse load types, or uncertainties in distributed power sources, making it difficult to meet the increasingly complex needs of distribution systems for refined reactive power regulation.

[0003] While existing research has explored various reactive power compensation methods, such as employing intelligent control algorithms based on fuzzy control, neural networks, or particle swarm optimization for dynamic management of compensation equipment, these methods generally suffer from limitations. First, the control algorithms lack real-time awareness and dynamic learning capabilities regarding system states, making them unable to adequately adapt to the interference effects of multi-source information (such as harmonic pollution and load type changes) in the distribution network. Second, the reactive power regulation execution process lacks real-time integration of equipment adjustable capacity, health status, and execution feedback, leading to problems such as unbalanced resource scheduling and shortened equipment lifespan. Furthermore, current technologies lack unified and refined standards for reactive power demand identification and regulation priority judgment, making it difficult to balance network-wide coordinated control with local regulation accuracy, resulting in limited overall system optimization. Therefore, current technologies have significant shortcomings in sensing capabilities for multi-source heterogeneous information in the distribution network, adaptive reactive power regulation, and standardized scheduling. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a reactive power compensation adaptive control method based on distribution networks, which solves the problems of sensing capability, adaptive reactive power regulation and scheduling standardization in distribution networks with multi-source heterogeneous information.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] In a first aspect, the present invention provides a reactive power compensation adaptive control method based on a power distribution network, comprising,

[0008] Real-time acquisition and preprocessing of power distribution network operating parameters; identification of load status and calculation of reactive power demand through data processing.

[0009] The status and parameter configuration of reactive power compensation equipment are collected, the adjustment target table is obtained by calculating the reactive power deviation, and the equipment is scheduled and allocated.

[0010] Based on reactive power demand, reactive power deviation and equipment matching data, the reactive power is adjusted and controlled by the DDPG algorithm combined with the OU noise mechanism and the curiosity mechanism, and the adjusted reactive power is sent to the compensation equipment to perform reactive power compensation.

[0011] The system collects feedback on the performance of the equipment, evaluates the effectiveness, and optimizes reactive power based on the evaluation results.

[0012] As a preferred embodiment of the reactive power compensation adaptive control method based on power distribution network described in this invention, the step of obtaining the adjustment target table by calculating the reactive power deviation and scheduling and allocating equipment refers to sampling the pre-processed three-phase currents at equal intervals to obtain discrete signals. And obtain after normalization. Based on the number of sampling points H and the time series index Constructing Nuttall window functions and time-domain mirror version And perform temporal convolution to obtain a symmetric weighted window function. ;

[0013] Through symmetric weighted window function Discrete signals after normalization Perform point-by-point weighting to obtain the final windowed signal used by FFT. Perform a full-phase FFT transform to obtain the spectral amplitude spectrum. Through the maximum spectral amplitude spectrum Extracting complex angles as signal phase The amplitudes of the left neighbor spectral line, the main spectral line, and the right neighbor spectral line are extracted. , , The frequency shift is corrected using three-spectral-line interpolation. and frequency component amplitude B;

[0014] The theoretical fundamental frequency is set as Calculate the ratio of each observed frequency to the fundamental frequency. Theoretical integer harmonic order Based on the corrected amplitude of each frequency component To distinguish between the fundamental frequency amplitude and the amplitudes of higher harmonics, when the frequency component amplitude... harmonic number When it equals 1, it is the fundamental amplitude. Amplitude of all frequency components harmonic number When it is greater than 1, it is the amplitude of a higher harmonic. Through the fundamental amplitude and higher harmonic amplitude Calculate the total harmonic distortion of the current. ;

[0015] The reactive power currently being output by the equipment Initial reactive power deviation calculated from theoretical reactive power requirements , set through statistical analysis threshold ,when Less than the threshold If the reactive power measurement value is normal, then the reactive power measurement value is normal; otherwise, the reactive power measurement value is significantly distorted and should be corrected by weighting. ;

[0016] The reactive power deviation is obtained by recalculating the corrected reactive power measurement values. The effect of this deviation on voltage is evaluated using a limiting sensitivity model. The sensitivity of voltage to reactive power is defined as... And calculate the compensation priority for each node. Perform descending sorting to determine the compensation order for each node, and then calculate the reactive power deviation. Node priority Total harmonic distortion of current and remaining adjustment capacity of equipment Establish a target adjustment table.

[0017] As a preferred embodiment of the reactive power compensation adaptive control method based on power distribution networks described in this invention, the reactive power factor is adjusted and controlled based on reactive power demand, reactive power deviation, and equipment matching data, using the DDPG algorithm combined with the OU noise mechanism and the curiosity mechanism. Remaining adjustment capacity of equipment reactive power deviation Unified encoding as state vector And design an external reward function. and internal reward function The combined reward mechanism obtains the comprehensive reward function. The Q-value function is modeled using the Dueling Q architecture, which decomposes the Q-value function into state-value functions. and action advantage function ;

[0018] Through the state value function and action advantage function The Q-value function is modeled and mean normalized to output the final Q-value. During policy training, the comprehensive reward function is used. Construct TD target Training the Q-network;

[0019] During the training of the Q-network, quadruplets are obtained through interaction with the environment and stored in the experience pool. Prioritized experience replay is used based on the TD error. The sampling probability is obtained by weighting the samples. And update the Q network parameter set through the RAdam optimizer. Simultaneously, a policy network structure Actor based on the DDPG algorithm is adopted to manage the state. Mapped to action output And use the OU noise mechanism to control the output action. Obtain by noise disturbance The policy network parameters are optimized using a advantage-weighted policy gradient method. Output adjustment action With the current output capacity of the device The superposition of these values ​​generates the reactive power target adjustment value. .

[0020] As a preferred embodiment of the reactive power compensation adaptive control method based on power distribution network described in this invention, wherein: the step of sending the adjusted reactive power to the compensation equipment to perform reactive power compensation refers to adjusting the target reactive power value... With control commands The instructions are encapsulated into device commands and transmitted to the device using the IEC 61850 protocol. The device then adjusts its settings based on the current operating status and the reactive power target value. The comparisons are made, and the output is adjusted accordingly.

[0021] As a preferred embodiment of the reactive power compensation adaptive control method based on power distribution networks described in this invention, the following steps are taken: the feedback effect of the data collection equipment is evaluated, and the reactive power index is optimized based on the evaluation results. The feedback information from the monitoring and compensation equipment is used to determine whether the equipment has reached the target adjustment value. Based on the execution status and feedback information, further strategy adjustments are made. If the target is met, the operational status before and after the adjustment is recorded. , The strategy reinforces the current adjustment action; if the target is not met, the system state after execution is collected for evaluation and learning, and subsequent states are used. The key operating parameters of the compensation equipment in the current cycle are updated as the state input for the adjustment strategy in the next cycle to optimize reactive power output.

[0022] As a preferred embodiment of the reactive power compensation adaptive control method based on power distribution networks described in this invention, the step of acquiring the status and parameter configuration of the reactive power compensation equipment refers to acquiring key parameters of the reactive power compensation equipment in real time through a monitoring system, and calculating the remaining regulation capacity of each reactive power compensation equipment based on the equipment's maximum output capacity and current actual output capacity. Based on the remaining adjustment capacity of the equipment Further monitor the health status of the equipment. By collecting data in real time and combining it with the remaining adjustment capacity and equipment health assessment results, generate an equipment capacity status table.

[0023] As a preferred embodiment of the reactive power compensation adaptive control method based on the power distribution network described in this invention, the step of identifying the load state and calculating the reactive power demand by processing data refers to calculating the load power factor by using the processed active power P. Set the load power factor threshold When the load power factor Greater than the load power factor threshold If the power demand is less than the maximum load, it is considered a light load. When the load power factor... Less than the load power factor threshold If the power demand exceeds the maximum load, it is considered a heavy load. When the load power factor... Large fluctuations indicate a fluctuating load. After identifying the load type, calculate the reactive power demand at each load point. .

[0024] As a preferred embodiment of the reactive power compensation adaptive control method based on power distribution network described in this invention, the real-time acquisition and preprocessing of power distribution network operating parameters refers to deploying three-phase multi-functional intelligent monitoring terminals at key locations to acquire data in real time and perform data filtering and standardization processing.

[0025] In a second aspect, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements any step of the reactive power compensation adaptive control method based on the distribution network as described in the first aspect of the present invention.

[0026] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the reactive power compensation adaptive control method based on the distribution network as described in the first aspect of the present invention.

[0027] The beneficial effects of this invention are as follows: This invention significantly improves the reactive power compensation control accuracy, response efficiency and system stability by using multiple data preprocessing mechanisms such as moving average and normalization, load classification method based on power factor identification, priority ranking mechanism based on harmonic sensing, and adaptive control strategy that integrates DDPG (Deep Deterministic Policy Gradient) algorithm with OU noise and curiosity mechanism. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a flowchart of the reactive power compensation adaptive control method based on the power distribution network in Example 1.

[0030] Figure 2 This is a diagram showing the device execution feedback evaluation and optimization in Example 1. Detailed Implementation

[0031] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0032] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0033] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0034] Example 1, referring to Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides a reactive power compensation adaptive control method based on a power distribution network, including the following steps:

[0035] S1. Real-time acquisition and preprocessing of power distribution network operating parameters; identification of load status and calculation of reactive power demand through data processing.

[0036] Specifically, real-time acquisition and preprocessing of power distribution network operating parameters refers to deploying three-phase multi-functional intelligent monitoring terminals at key locations (low-voltage side of distribution transformers, busbars of distribution lines, user access points, and grid connection points of distributed power sources (photovoltaics, energy storage)) to collect data (three-phase voltage, current, active power, and reactive power) in real time, and performing data filtering and standardization processing.

[0037] The data filtering described above uses a moving average method to filter the data in order to smooth out signal variation trends.

[0038]

[0039] in, It is the smoothed value at time t. is the original data values ​​(three-phase voltage, current, active power, reactive power) of the first i sampling points, and n is the sliding window length;

[0040] The data is standardized as follows:

[0041]

[0042] in, X is the normalized value, and X is the original data value (three-phase voltage, current, active power, reactive power). , These are the historical maximum and minimum values ​​of the original data value X, respectively.

[0043] By deploying intelligent sensing nodes at high density, the system achieved global perception and local response capabilities for its status. Through moving average filtering, dynamic stability and jitter suppression of monitoring data were achieved, significantly improving the reactive power control system's ability to identify dynamic trends and its robustness, and reducing the false judgment rate. Through normalization methods, the system achieved scale uniformity of data features and balance of model inputs, significantly improving the versatility and generalization ability of the control algorithm in different nodes and scenarios.

[0044] Furthermore, by processing data to identify load status and calculating reactive power demand, the load power factor is calculated using the processed active power P.

[0045]

[0046] in, P is the load power factor, where P is active power and U is reactive power.

[0047] Setting load power factor thresholds using a rule-based approach. When the load power factor Greater than the load power factor threshold If the power demand is less than the maximum load, it is considered a light load. When the load power factor... Less than the load power factor threshold If the power demand exceeds the maximum load, it is considered a heavy load. When the load power factor... Large fluctuations indicate a fluctuating load. After identifying the load type, calculate the reactive power demand at each load point:

[0048]

[0049] in, P is reactive power demand, and P is active power. It is the angle of load power factor. It is the angle of the target load power factor. It is the tangent of the power factor angle under the current load. It is the tangent of the target power factor angle;

[0050] The calculation formula is:

[0051]

[0052] in, It is the load power factor.

[0053] By calculating the load power factor, an objective measurement of load operating efficiency is achieved. By identifying load types based on power factor and load level rules, fine-grained modeling of load response characteristic differences is realized, enhancing the personalization and adaptability of subsequent control strategies. By calculating reactive power demand through triangular relationships, dynamic modeling based on the deviation between the actual load state and the target operating state is realized, making reactive power compensation strategies more targeted and accurate.

[0054] S2. Collect the status and parameter configuration of the reactive power compensation equipment, obtain the adjustment target table by calculating the reactive power deviation, and schedule and allocate the equipment.

[0055] Specifically, collecting the status and parameter configuration of reactive power compensation equipment refers to collecting key parameters (maximum output capacity, current output capacity, response time, and online status) of the reactive power compensation equipment in real time through the monitoring system, and calculating the remaining regulation capacity of each reactive power compensation device using the equipment's maximum output capacity and current actual output capacity.

[0056]

[0057] in, It refers to the remaining adjustment capacity of the equipment. This is the maximum output capacity of the device. It is the actual reactive power currently output by the equipment;

[0058] Based on equipment remaining adjustment capacity Further monitor the health status of the equipment to ensure it is not in a faulty or abnormal state. Using real-time collected data, combined with remaining regulation capacity and equipment health assessment results, generate an equipment capacity status table. The table should include the equipment name and maximum output capacity. Current output capacity Residual regulating capacity Response time, online status, and health status.

[0059] By collecting key operating parameters of the equipment in real time, high-frequency dynamic perception of the reactive power compensation capacity status is achieved, laying a data foundation for subsequent adjustment capacity assessment and control strategy allocation. By calculating the remaining adjustment capacity in real time, the remaining adjustment potential of the equipment is measured in a refined manner, improving the boundary accuracy and dynamism of the compensation control strategy formulation. By monitoring and analyzing the health status of the equipment, the authenticity of the adjustment capacity is verified, and the fault tolerance of the system under abnormal conditions is improved. By generating an equipment capacity status table, centralized management of the control capacity, response characteristics and health level of various reactive power compensation equipment is achieved.

[0060] Furthermore, by calculating the reactive power deviation, a regulation target table is obtained, and the equipment is scheduled and allocated. The pre-processed three-phase currents are sampled at equal intervals to obtain discrete signals.

[0061]

[0062] Where x(h) is a discrete signal, and W is the current amplitude. It refers to the signal frequency (fundamental frequency and harmonics). It is the sampling frequency. H is the initial phase, and H is the number of sampling points. It is a time series index;

[0063] The discrete signal x(h) is normalized to eliminate the DC offset effect:

[0064]

[0065] in, It is a discrete signal after normalization. It is the average value of the signal;

[0066] Based on the number of sampling points H and the time series index Constructing the Nuttall window function:

[0067]

[0068] in, It is the Nuttall window function. , , , These are the coefficients of the window function. , , It refers to the fundamental frequency, the second harmonic, and the third harmonic;

[0069] To improve temporal symmetry (i.e., the symmetry of the window about the central axis), a Nuttall window function is constructed. Time-domain mirror version:

[0070]

[0071] in, It is a time-domain mirror version;

[0072] Nuttall window function With time-domain mirror version Perform temporal convolution:

[0073]

[0074] in, It is a symmetric weighted window function. It is a temporal convolution operation;

[0075] Through symmetric weighted window function Discrete signals after normalization After point-by-point weighting, the final windowed signal used by the FFT is obtained:

[0076]

[0077] in, It is a signal weighted by a window function;

[0078] The signal weighted using a window function Perform a full-phase FFT transformation. The full-phase FFT expression is:

[0079]

[0080] in, Here, B is the frequency amplitude spectrum at frequency u, M(·) is the frequency response of the full-phase window function at frequency u, and H is the number of sampling points. Frequency offset is the initial phase, e is the exponential base, and l is the basic element for constructing complex numbers;

[0081] Through the maximum spectral amplitude spectrum Extracting complex angles as signal phase:

[0082]

[0083] in, It is the signal phase;

[0084] Through the amplitude spectrum Extract the amplitude of the left neighbor spectral line, the amplitude of the main spectral line (maximum), and the amplitude of the right neighbor spectral line:

[0085]

[0086] in, , , These are the amplitudes of the left adjacent spectral line, the amplitude of the main spectral line (maximum), and the amplitude of the right adjacent spectral line, respectively.

[0087] Based on the amplitude of the left neighbor spectral line, the amplitude of the main spectral line (maximum), and the amplitude of the right neighbor spectral line. , , The frequency shift is corrected using three-spectral-line interpolation. And the frequency component amplitude B, used to improve the estimation accuracy of frequency and amplitude;

[0088] The three-spectral-line interpolation correction frequency offset pass , , Calculate the spectral line shape factor :

[0089]

[0090] via spectral shape factor Correcting frequency offset :

[0091]

[0092] in, This is the corrected frequency offset. It is an inverse function relationship between the spectral shape factor and the frequency shift, belonging to the inverse sinc function. It is the spectral shape factor;

[0093] The amplitude of the frequency component B corrected by the three-spectral-line interpolation:

[0094]

[0095] in, It is the corrected frequency component amplitude. It is the frequency offset. Weighting function;

[0096] The theoretical fundamental frequency is set as Calculate the ratio of each observed frequency to the fundamental frequency. The closest integer multiple relationship:

[0097]

[0098] in, It is the fundamental frequency. With observation frequency The closest theoretical integer harmonic order, It is a floor function;

[0099] Based on the corrected frequency component amplitude corresponding to each frequency component To distinguish between the fundamental frequency amplitude and the amplitudes of higher harmonics, when the frequency component amplitude... harmonic number When it equals 1, it is the fundamental amplitude. Amplitude of all frequency components harmonic number When it is greater than 1, it is the amplitude of a higher harmonic. Through the fundamental amplitude and higher harmonic amplitude Calculate the total harmonic distortion of the current. :

[0100]

[0101] Where G is the number of times the harmonic frequency components are extracted;

[0102] The reactive power currently being output by the equipment The initial reactive power deviation is calculated from the theoretical reactive power requirement and used for subsequent reactive power correction.

[0103]

[0104] in, It is reactive power deviation. This is the theoretical reactive power requirement;

[0105] Setting through statistical analysis threshold ,when Less than the threshold If the reactive power measurement value is normal, then the reactive power measurement value is normal; otherwise, the reactive power measurement value is significantly distorted and should be weighted and corrected.

[0106]

[0107] in, This is the corrected reactive power measurement value. It is an empirical coefficient;

[0108] The reactive power deviation is obtained by recalculating the corrected reactive power measurement values. Based on the accurate acquisition of reactive power deviation at each node, the impact of this deviation on voltage is evaluated using a limiting sensitivity model, thereby calculating the compensation priority. The sensitivity of voltage to reactive power is defined as follows:

[0109]

[0110] in, It is the sensitivity of the voltage at node i (the compensation device) to reactive power. It is the voltage at node i. It is the reactive power measurement value after correction at node i. It is the reactive power disturbance;

[0111] Based on the sensitivity of voltage to reactive power Calculate the compensation priority for each node (compensation device):

[0112]

[0113] in, It is the compensation priority of node i;

[0114] The compensation priorities of all nodes are sorted in descending order to determine the compensation sequence for each node. This sequence is then used to calculate the reactive power deviation. Node priority Total harmonic distortion of current and remaining adjustment capacity of equipment Establish adjustment target table (add new items to the table) The field is used to mark the degree of harmonic influence so that the controller can use an anti-interference compensation mechanism when adjusting. The recommended equipment field combines node priority and the adjustment capacity of various compensation equipment (static var compensator (SVG), capacitor bank) to automatically recommend the most suitable compensation equipment.

[0115] By sampling and normalizing three-phase current at equal intervals, a reliable foundation is provided for high-precision harmonic extraction and reactive power assessment, improving the data perception quality before reactive power compensation control. By constructing and convolving a time-domain symmetric Nuttall window function, a high-resolution, low-leakage window function is obtained for FFT analysis, effectively improving the accuracy of FFT spectrum transformation and providing high-quality input for harmonic determination and frequency offset estimation, thus enhancing the sensing capability of the control system. By performing full-phase FFT after signal weighting, the accuracy of spectrum analysis is improved, enabling reactive power compensation control to more accurately identify nonlinear disturbance sources and fluctuation characteristics. By correcting frequency offset and amplitude using three-spectral-line interpolation, the accurate identification of frequency drift and higher harmonics is strengthened, improving the reactive power control strategy. The frequency domain sensing capability of the system enables a quantitative description of the harmonic pollution level of the current signal through comparison with the theoretical fundamental frequency and judgment of integer multiples, providing a basis for introducing a harmonic suppression mechanism for reactive power control. By calculating and correcting reactive power deviation, the robustness of reactive power assessment is improved, the system's self-recovery capability to measurement anomalies is enhanced, and more reliable compensation decisions are achieved. By evaluating the voltage response to reactive power through the limit sensitivity model, the system maximizes the voltage support effect with minimal resource input, optimizes the cost-effectiveness of reactive power compensation actions, and realizes an adaptive compensation strategy that takes into account multiple factors such as harmonic interference, equipment capacity, and adjustment priority by constructing an adjustment target table and introducing a harmonic influence degree and recommended equipment mechanism, thereby significantly improving the overall intelligence level of system control.

[0116] S3. Based on reactive power demand, reactive power deviation and equipment matching data, the reactive power is adjusted and controlled by the DDPG algorithm combined with the OU noise mechanism and the curiosity mechanism, and the adjusted reactive power is sent to the compensation equipment to perform reactive power compensation.

[0117] Specifically, based on reactive power demand, reactive power deviation, and equipment matching data, the DDPG algorithm, combined with the OU noise mechanism and the curiosity mechanism, is used to adjust and control the reactive power index and load power factor. Remaining adjustment capacity of equipment reactive power deviation Unified encoding as state vector Furthermore, a joint reward mechanism combining external and internal reward functions is designed to guide the learning behavior of the agent (which learns how to choose actions through interaction with the environment).

[0118] The external reward function rewards actions based on the deviation from the control target (reactive power deviation), and the calculation formula is as follows:

[0119]

[0120] in, It is an external reward function. , These are energy consumption and error weighting coefficients, set through fuzzy logic. It is the power consumption coefficient of reactive power regulation, which is set through experiments. This is the current action. It is a moment reactive power deviation;

[0121] The internal reward function employs a curiosity mechanism as a compensation mechanism, rewarding the agent for exploring new states. This encourages the agent to explore new and unknown load changes, preventing the system from getting trapped in local optima. The calculation formula is as follows:

[0122]

[0123] in, It is an internal reward function. It is a curiosity incentive coefficient, set through an adaptive mechanism. The current policy is related to the state. The predicted value, It is a state The actual value, It is the square of the Euclidean distance;

[0124] The external reward function and the internal reward function are combined to form the final comprehensive reward function:

[0125]

[0126] in, It is the final comprehensive reward function;

[0127] In terms of policy network structure, the Dueling Q architecture is used to model the Q-value function, which is then decomposed into a state value function and an action advantage function.

[0128] The state value function utilizes reactive power deviation. Current action In addition to environmental variables (power factor), state features are extracted and a state value function is output through a multilayer perceptron (MLP). This is used to identify whether the system is in a "high-quality, adjustable state" under different reactive power demand scenarios. It is the set of all trainable parameters of a multilayer perceptron;

[0129] The action advantage function utilizes the state and actions The action advantage function is output by extracting action features through a multilayer perceptron (MLP). ,in, It is the set of parameters for the action advantage function;

[0130] Through the state value function and action advantage function The Q-value function is modeled and mean normalization is applied to eliminate relative evaluation bias between actions, outputting the final Q-value:

[0131]

[0132] in, It is the final Q-value, representing the state. Take action below The expected long-term returns It is the size of the action set. It is the set of parameters for the entire Q network;

[0133] During policy training, a comprehensive reward function is used. Constructing a TD target to train a Q network:

[0134]

[0135] in, It is a TD target. It's a discount factor, set through machine learning. It is the maximum action value of the next state estimated by the target Q-network. It is the set of parameters for the target Q-network;

[0136] The parameter set of the target Q-network Perform soft updates to prevent target network parameters from being updated. With main network parameters To address the oscillation issues caused by synchronous updates and enhance training stability:

[0137]

[0138] in, It is the soft update rate;

[0139] During the training of the Q-network, quadruplets (state, action, reward, and next state) are obtained through interaction with the environment and stored in the experience pool. Prioritized Experience Replay (PER) is used to weight samples based on the TD error, and training samples are selected from these weighted samples so that the model prioritizes learning experiences with large prediction errors.

[0140]

[0141]

[0142] in, It is the TD error of sample c. It's an instant reward. It is a discount factor. It is the target Critic network. It is the next action generated by the target policy network. It is the sampling probability of sample c. It is a priority coefficient, set through a genetic algorithm. It is the TD error of sample j;

[0143] After sampling, the Q network parameter set is updated using the RAdam optimizer. To improve stability during the initial training phase, the formula has been updated as follows:

[0144]

[0145] in, It is the updated set of network parameters. It's the learning rate. , These are the unbiased correction values ​​for the first-order and second-order gradient momentum estimates, respectively. It is a smoothing term to prevent division by zero. It is a dynamic correction factor;

[0146] Meanwhile, to improve the continuity and robustness of action selection, a policy network structure Actor based on the DDPG algorithm is adopted to handle the state. Mapped to action output:

[0147]

[0148] in, It is in state The following actions were taken. These are policy network parameters;

[0149] The OU noise mechanism is used to analyze the movements during training. Incorporate noise disturbances to enhance strategy exploration capabilities:

[0150]

[0151] in, It is the action after noise disturbance. It is the change in noise at time t;

[0152] The update formula is:

[0153]

[0154] in, It is the change in noise at time t. It is the attenuation coefficient of OU noise, which is set through statistical analysis. It is the mean of the noise. It is the noise value at the time preceding time t. is the standard deviation of the noise, and N(0,1) is random noise with a standard normal distribution;

[0155] Standard deviation of noise It is a function of time t, which changes with different training phases and is adjusted in stages as follows:

[0156]

[0157] in, In the initial stage, agents are encouraged to conduct extensive exploration to avoid local optima. This is the intermediate stage, where the system begins to converge towards the target region, but still retains a certain degree of exploratory nature. This is the later stage, where the agent gradually focuses on efficient strategies. This is the convergence phase, where the agent fully enters the convergence phase to maximize the accuracy of the policy. , , It is a milestone threshold for the number of time steps during the training process of the agent, which is set through a callback mechanism;

[0158] Optimize policy network parameters using a advantage-weighted policy gradient method. The objective function is:

[0159]

[0160] in, It is the objective function. It is the expectation sampled from the current strategy. It is the logarithm of the probability of the current policy on the action. It is the action advantage function;

[0161] Optimize policy network parameters using objective function strategy Output adjustment action With the current output capacity of the device The superposition of these values ​​generates the reactive power target adjustment value:

[0162]

[0163] in, This is the reactive power target adjustment value. It is the adjustment action generated by the strategy.

[0164] By uniformly encoding load power factor, equipment residual regulation capacity, and reactive power deviation into state vectors, a foundation is laid for adaptive reactive power control based on reinforcement learning. This enables the agent to identify and respond to complex system state changes. Through the design of a joint reward mechanism combining external and internal reward functions, effective control of reactive power regulation and adaptive response to new load patterns are achieved, improving the robustness and generalization ability of the controller. This is further enhanced by the use of Dueling... The Q-architecture models the Q-value function, improving the policy network's ability to identify optimal actions under different states and enhancing the accuracy of reactive power regulation actions. Mean normalization of the Q-value function improves the stability and consistency of the control strategy's regulation behavior, reducing the risk of erroneous actions in reactive power regulation. A TD target is constructed using a comprehensive reward function, and a soft update mechanism is employed to train the target Q-network, accelerating training convergence and reducing power system fluctuation risks during training. Prioritized Experience Replay (PER) is used to sample training samples with TD error as weight, improving the control strategy's response under abnormal conditions and enhancing system robustness and fault adaptability. The RAdam optimizer optimizes network parameters, achieving uncertainty control and stable convergence in the early training phase. A DDPG-based Actor network maps states to actions, and an OU noise perturbation mechanism is introduced to enhance the policy's exploration capability and output continuity. A pluralistic weighted policy gradient method optimizes the Actor network, making the policy gradient more focused on actions that significantly improve control performance. Finally, by superimposing the regulation actions generated by the Actor network with the current output capacity of the equipment, a reactive power target regulation value is formed, achieving dynamic adaptive regulation of the compensation equipment.

[0165] Furthermore, the adjusted reactive power is sent to the compensation equipment to perform reactive power compensation, which means adjusting the target reactive power value. With control commands Encapsulated into device instructions, ensuring they can be sent to and executed by the device via industrial control protocols. The device instruction encapsulation includes the control node number, device number, control type, and target reactive power adjustment value. Regarding response time requirements, the IEC 61850 protocol is used to transmit encapsulated packets of equipment commands to the equipment. The equipment adjusts its current operating status (current reactive power) and the target reactive power value accordingly. The comparisons are made, and the output is adjusted accordingly.

[0166] By encapsulating reactive power target adjustment values ​​and control commands into device commands, the intelligent control strategy is transformed into standardized commands at the device end, enhancing the standardization and operability of system commands and improving the real-time performance and reliability of the control system. The encapsulated control commands include control node number, device number, control type, target adjustment value, and response time requirements, ensuring the accuracy of device control targets and the controllability of responses. By using the IEC 61850 protocol to transmit the encapsulated device command packets to the device, standardization, structuring, and interoperability of communication are guaranteed. After receiving the encapsulated commands, the device compares its current operating status with the target adjustment value and automatically adjusts its output, achieving closed-loop control execution of the adjustment behavior.

[0167] S4. Collect feedback from the equipment to evaluate its performance and optimize reactive power based on the evaluation results.

[0168] Specifically, the system collects and evaluates the feedback from the equipment's performance. Based on the evaluation results, it optimizes the reactive power compensation system by monitoring the feedback information of the equipment's adjustments (the adjustment range (the amount of adjustment the equipment has currently completed)) to determine whether the equipment has reached the target adjustment value. The system monitors the execution status (whether the device is performing adjustments normally) and makes further strategy adjustments based on feedback. If the target is met, the system records the operating status before and after the adjustment. , The strategy reinforces the current adjustment action. If the target is not met, the system status (node ​​voltage, current, presence or absence of harmonics, frequency) after the action is collected for evaluation and learning, and the subsequent status is used. Update the key operating parameters of the compensation equipment in the current cycle (reactive power, power factor, reactive power residual, and reactive power change) as the state input for the adjustment strategy in the next cycle to optimize reactive power output.

[0169] By monitoring the feedback information of the compensation equipment's regulation, dynamic perception and status recognition of the actual execution effect are achieved. By determining whether the equipment has reached the target regulation value and making strategy adjustments based on the execution status, adaptive correction and dynamic response of the control strategy are realized. By recording the operating data before and after regulation in the target state, positive feedback learning is realized, which strengthens high-quality regulation actions. By collecting the subsequent state of the system under non-target conditions for evaluation and learning, negative feedback analysis of failure experience is realized. By updating the key indicators in the subsequent state to the operating parameters of the compensation equipment, it is helpful to realize cross-cycle strategy evolution, enabling the regulation system to dynamically adapt to the non-stationary changes in the power grid state over time, thereby ensuring continuous optimization and long-term stability of the control effect.

[0170] This embodiment also provides a computer device applicable to the reactive power compensation adaptive control method based on power distribution network, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the reactive power compensation adaptive control method based on power distribution network proposed in the above embodiment.

[0171] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0172] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the reactive power compensation adaptive control method based on the power distribution network proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

Claims

1. A reactive power compensation adaptive control method based on power distribution networks, characterized in that: include, Real-time acquisition and preprocessing of power distribution network operating parameters; identification of load status and calculation of reactive power demand through data processing. The status and parameter configuration of reactive power compensation equipment are collected, the adjustment target table is obtained by calculating the reactive power deviation, and the equipment is scheduled and allocated. Based on reactive power demand, reactive power deviation and equipment matching data, the reactive power is adjusted and controlled by the DDPG algorithm combined with the OU noise mechanism and the curiosity mechanism, and the adjusted reactive power is sent to the compensation equipment to perform reactive power compensation. Collect feedback from the equipment to evaluate its performance and optimize reactive power based on the evaluation results; The process of obtaining the adjustment target table by calculating the reactive power deviation and scheduling and allocating equipment involves sampling the pre-processed three-phase current at equal intervals to obtain discrete signals. And obtain after normalization. Based on the number of sampling points H and the time series index Constructing Nuttall window functions and time-domain mirror version And perform temporal convolution to obtain a symmetric weighted window function. ; Using a symmetric weighted window function Discrete signals after normalization Perform point-by-point weighting to obtain the final windowed signal used by FFT. Perform a full-phase FFT transform to obtain the spectral amplitude spectrum. Through the maximum spectral amplitude spectrum Extracting complex angles as signal phase The amplitudes of the left neighbor spectral line, the main spectral line, and the right neighbor spectral line are extracted. , , The frequency shift is corrected using three-spectral-line interpolation. and frequency component amplitude B; The theoretical fundamental frequency is set as Calculate the ratio of each observed frequency to the fundamental frequency. Theoretical integer harmonic order Based on the corrected amplitude of each frequency component To distinguish between the fundamental frequency amplitude and the amplitudes of higher harmonics, when the frequency component amplitude... harmonic number When it equals 1, it is the fundamental amplitude. Amplitude of all frequency components harmonic number When it is greater than 1, it is the amplitude of a higher harmonic. Through the fundamental amplitude and higher harmonic amplitude Calculate the total harmonic distortion of the current. ; The reactive power currently being output by the equipment Calculate the initial reactive power deviation from the theoretical reactive power requirement. , set through statistical analysis threshold ,when Less than the threshold If the reactive power measurement value is normal, then the reactive power measurement value is normal; otherwise, the reactive power measurement value is significantly distorted and should be corrected by weighting. ; The reactive power deviation is obtained by recalculating the corrected reactive power measurement values. The effect of this deviation on voltage is evaluated using a limiting sensitivity model. The sensitivity of voltage to reactive power is defined as... And calculate the compensation priority for each node. Perform descending sorting to determine the compensation order for each node, and then calculate the reactive power deviation. Node priority Total harmonic distortion of current and remaining adjustment capacity of equipment Establish a target adjustment table.

2. The reactive power compensation adaptive control method based on distribution network as described in claim 1, characterized in that: Based on reactive power demand, reactive power deviation, and equipment matching data, the DDPG algorithm, combined with the OU noise mechanism and the curiosity mechanism, adjusts and controls the reactive power index to adjust the load power factor. Remaining adjustment capacity of equipment reactive power deviation Unified encoding as state vector And design an external reward function. and internal reward function The combined reward mechanism obtains the comprehensive reward function. The Q-value function is modeled using the Dueling Q architecture, which decomposes the Q-value function into state-value functions. and action advantage function ; Through the state value function and action advantage function The Q-value function is modeled and mean normalized to output the final Q-value. During policy training, the comprehensive reward function is used. Construct TD target Training the Q-network; During the training of the Q-network, quadruplets are obtained through interaction with the environment and stored in the experience pool. Prioritized experience replay is used based on the TD error. The sampling probability is obtained by weighting the samples. And update the Q network parameter set through the RAdam optimizer. Simultaneously, a policy network structure Actor based on the DDPG algorithm is adopted to manage the state. Mapped to action output And use the OU noise mechanism to control the output action. Obtain by noise disturbance The policy network parameters are optimized using a advantage-weighted policy gradient method. Output adjustment action With the current output capacity of the device The superposition of these values ​​generates the reactive power target adjustment value. .

3. The reactive power compensation adaptive control method based on distribution network as described in claim 2, characterized in that: The step of sending the adjusted reactive power to the compensation equipment to perform reactive power compensation refers to adjusting the target reactive power value. With control commands The instructions are encapsulated into device commands and transmitted to the device using the IEC 61850 protocol. The device then adjusts its settings based on the current operating status and the reactive power target value. The comparisons are made, and the output is adjusted accordingly.

4. The reactive power compensation adaptive control method based on distribution network as described in claim 3, characterized in that: The collection device performs feedback evaluation, and based on the evaluation results, optimizes the reactive power index monitoring and compensation device's adjustment feedback information to determine whether the device has reached the target adjustment value. Based on the execution status and feedback information, further strategy adjustments are made. If the target is met, the operational status before and after the adjustment is recorded. , The strategy reinforces the current adjustment action; if the target is not met, the system state after execution is collected for evaluation and learning, and subsequent states are used. The key operating parameters of the compensation equipment in the current cycle are updated as the state input for the adjustment strategy in the next cycle to optimize reactive power output.

5. The reactive power compensation adaptive control method based on distribution network as described in claim 4, characterized in that: The acquisition of the status and parameter configuration of the reactive power compensation equipment refers to the real-time acquisition of key parameters of the reactive power compensation equipment through the monitoring system, and the calculation of the remaining regulation capacity of each reactive power compensation device based on the maximum output capacity and the current actual output capacity. Based on the remaining adjustment capacity of the equipment Further monitor the health status of the equipment. By collecting data in real time and combining it with the remaining adjustment capacity and equipment health assessment results, generate an equipment capacity status table.

6. The reactive power compensation adaptive control method based on distribution network as described in claim 5, characterized in that: The process of identifying load status and calculating reactive power demand by processing data refers to calculating the load power factor using the processed active power P. Set the load power factor threshold When the load power factor Greater than the load power factor threshold If the power demand is less than the maximum load, it is considered a light load. When the load power factor... Less than the load power factor threshold If the power demand exceeds the maximum load, it is considered a heavy load. When the load power factor... Large fluctuations indicate a fluctuating load. After identifying the load type, calculate the reactive power demand at each load point. .

7. The reactive power compensation adaptive control method based on distribution network as described in claim 6, characterized in that: The real-time acquisition and preprocessing of power distribution network operating parameters refers to deploying three-phase multi-functional intelligent monitoring terminals at key locations to collect data in real time and perform data filtering and standardization processing.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the reactive power compensation adaptive control method based on the power distribution network as described in any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the reactive power compensation adaptive control method based on the power distribution network as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-source reactive compensation control method, system, program, medium and equipment

    CN119231559A

  • V2G-based alternating current charging power adjusting method and device

    CN119891330A