Gfm / gfl hybrid control cooperative optimization system and method for weak grid with high penetration of new energy
The GFM/GFL hybrid control collaborative optimization system solves the problem of insufficient control interpretability and auditability in the grid connection of high-penetration new energy sources in weak power grids, achieving a balance between economy, stability and power quality, and improving the reliability and interpretability of the control system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KUNMING UNIVERSITY
- Filing Date
- 2026-04-28
- Publication Date
- 2026-05-29
AI Technical Summary
Under the condition of high penetration of new energy in weak power grids, existing technologies are unable to coordinate minute-level scheduling, second-level setting and millisecond-level execution at the same time, resulting in insufficient control interpretability and auditability. Furthermore, traditional control methods lack strict online action and feasible constraint connection under extreme conditions, affecting system stability and power quality.
A GFM/GFL hybrid control collaborative optimization system is adopted, including state acquisition and system strength assessment, upper-level two-stage sub-Blu-ray scheduling, middle-level security reinforcement learning tuning, master-slave hybrid control execution, and security protection and rollback modules. Through multi-timescale collaborative optimization, the system achieves the unification of minute-level scheduling, second-level tuning, and millisecond-level execution. Combined with security protection mechanisms such as barrier functions and action projection, the system ensures the reliability and interpretability of control.
It achieves a balance between economy, stability and power quality under the condition of high penetration of new energy in weak grids, avoids phase angle changes and current surges in the traditional hard switching mode, improves controllability and reliability of engineering deployment under extreme conditions, and enhances convergence efficiency and online deployment interpretability in new scenarios.
Smart Images

Figure CN122118970A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of new energy grid connection control technology, and in particular to a GFM / GFL hybrid control collaborative optimization system and method for high-penetration new energy grid connection in weak power grids. Background Technology
[0002] With the large-scale connection of new energy power generation devices such as wind power and photovoltaics to the grid, the proportion of traditional synchronous machines continues to decline, the system inertia, damping and fault support capabilities are weakened, and voltage fluctuations, frequency fluctuations, phase-locked loop instability and oscillation problems under weak grid conditions are becoming increasingly prominent.
[0003] Existing technologies typically improve grid-connected control from the perspectives of pure GFL control, pure GFM control, threshold switching control, rule-based adaptive control, or single-layer optimization. Pure GFL control is susceptible to the effects of phase-locked loops, grid impedance, and current loop coupling under weak grid conditions; while pure GFM control has strong voltage source support capabilities, it may also cause oscillations or reduced economic efficiency under strong grid conditions or parameter mismatch.
[0004] Existing methods often employ single-time-scale optimization or empirical rule-based mode switching, which struggle to simultaneously coordinate minute-level scheduling, second-level tuning, and millisecond-level execution. Furthermore, most solutions still rely on post-event protection or soft penalty mechanisms for handling safety boundaries, resulting in a lack of strict coordination between online actions and feasible constraints. This leads to insufficient interpretability and auditability of controls under extreme conditions.
[0005] In the context of high-penetration renewable energy grid connection, engineering faces three constraints: First, insufficient inertia and damping under weak grid conditions lead to a contraction of transient stability margin; second, grid connection standards are continuously being strengthened, and controllers must simultaneously meet fault ride-through, frequency support, and power quality boundaries; third, on-site operation requires "explainability and auditability," and cannot rely on uncontrollable black-box switching logic.
[0006] Therefore, there is an urgent need for a GFM / GFL hybrid control collaborative optimization method and system that can achieve multi-time-scale collaboration among the scheduling, control and execution layers, while taking into account stability, economy, power quality and engineering safety, for high-penetration renewable energy scenarios in weak power grids. Summary of the Invention
[0007] To overcome the aforementioned problems in the existing technology, this invention proposes a GFM / GFL hybrid control collaborative optimization system and method for high-penetration renewable energy grid connection in weak power grids.
[0008] The technical solution adopted by this invention to solve its technical problem is: a GFM / GFL hybrid control collaborative optimization system for high-penetration renewable energy grid connection in weak power grids, including a status acquisition and system strength assessment module, an upper-level two-stage sub-Blu-ray scheduling module, a middle-level security reinforcement learning tuning module, a master-slave hybrid control execution module, and a security protection and rollback module. The status acquisition and system strength assessment module is used to collect real-time operating status data of the grid connection point, estimate the background impedance, and calculate the system strength. The upper-level two-stage sub-Brow bar scheduling module performs multi-objective scheduling according to the first preset scheduling cycle to obtain the first-stage decision quantity; The mid-level security reinforcement learning tuning module generates the allocation correction quantity and control parameter action based on the operating status data collected by the state acquisition and system strength assessment module and the first-stage decision quantity generated by the upper-level two-stage sub-Blu-ray scheduling module, thereby obtaining the real-time allocation coefficient. The master-slave hybrid control execution module is based on real-time ratio coefficient execution device-level master-slave GFM / GFL hybrid control, wherein the GFM channel provides a reference voltage, and the GFL channel generates a correction current and maps it to a correction voltage through virtual impedance to obtain a unified execution reference. The security protection and rollback module is used to switch the system control mode, which includes conservative rule control and reinforcement learning control.
[0009] The aforementioned GFM / GFL hybrid control collaborative optimization system for high-penetration renewable energy grid connection in weak power grids operates as follows: The safety protection and rollback module performs look-ahead checks on voltage deviation, frequency deviation, current amplitude, frequency change rate, and dominant oscillation amplitude; when an action exceeds the limits, feasible region projection is first executed; if continuous limit exceedances, sustained oscillation growth, excessive frequency change rate, or stability margin below the threshold are still detected after projection, conservative rule control is switched to within a preset time; when the minimum dwell time is met and continuous absence of limit exceedances and oscillation growth is maintained, and the system returns to the safe set, rollback is released and reinforcement learning control is restored.
[0010] The aforementioned GFM / GFL hybrid control and collaborative optimization system for high-penetration renewable energy grid integration in weak power grids, specifically the master-slave GFM / GFL hybrid control, involves the GFM channel outputting a reference voltage. The correction current is output from the GFL channel. The correction voltage is obtained through virtual impedance mapping. The unified execution reference satisfies: ; ; in, Limit the correction amplitude to ensure that a single PWM modulation link is always available.
[0011] A GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid integration in weak power grids, based on the aforementioned collaborative optimization system, includes the following steps: Step 1: Collect the operating status data of the PCC side of the grid connection point. The operating status data includes at least voltage, current, active power, reactive power, frequency deviation, frequency change rate, harmonic sensitivity index, new energy output, load information and background impedance estimate. Step 2: Perform upper-level two-stage distributed bar scheduling according to the first preset scheduling cycle to obtain the first-stage decision quantity. The first-stage decision quantity includes at least the active power reference P. g Reactive power reference Q g Reserve capacity R and GFM ratio target α target And obtain the robust executable interval or contingency plan set based on the backtracking decision set of uncertainties; Step 3: Perform mid-layer security reinforcement learning tuning according to the second preset control cycle, based on the operating status data and α. target Generate the ratio correction amount Δα RL and control parameter actions, and combine the ratio correction amount with the α target The real-time proportioning coefficient α is obtained by fusion. t ; Step 4, based on the real-time proportioning coefficient α t The execution device-level master-slave GFM / GFL hybrid control is implemented, where the GFM channel provides a reference voltage, and the GFL channel generates a correction current and maps it to a correction voltage through virtual impedance to obtain a unified execution reference. Step 5: For the unified execution reference current limiting, action projection, constraint look-ahead verification and backoff control, restore safety reinforcement learning control when the safety release condition is met, forming a closed-loop collaborative control with minute-level scheduling, second-level tuning and millisecond-level execution.
[0012] The aforementioned GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid integration in weak power grids obtains the background impedance estimate through low-amplitude active injection and synchronous detection, according to... Calculate the estimated background impedance, where To estimate the background impedance of the power grid at the k-th sampling time, This represents the increment of the fundamental voltage at the k-th sampling time. The increment of the fundamental current at the k-th sampling time; and according to The equivalent system strength is obtained, where K scr This is the base conversion factor from impedance to SCR. To prevent extremely small positive numbers from being divided by zero, the operating conditions are divided into extremely weak network, weak network, and strong network according to the equivalent system strength, so as to adjust the minimum GFM ratio, the upper limit of the rate of change of motion, and the damping parameters respectively.
[0013] In the aforementioned GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid integration in weak power grids, the first-stage decision quantity in the upper-level two-stage distributed bar scheduling is x(1)=[P g Q g ,R,α target The second-stage backtracking decision quantity is y(ξ)=[ΔP] g ,ΔQ g ,r ↑ ,r ↓ ,Δα,P curt ](ξ), where ξ represents the uncertainty, which includes at least the fluctuations in wind and solar power output, load fluctuations, and equivalent impedance disturbances, ΔP g ΔQ represents the active power adjustment. g Represents the reactive power adjustment amount, r ↑ This indicates an upward adjustment of the reserve adjustment amount, r ↓ This indicates a reduction in the reserve adjustment amount, Δα represents the adjustment amount for the GFM / GFL mixing ratio, and P curt This represents the power curtailment; the dispatching process simultaneously considers operating costs, stability risks, and power quality, and satisfies the BLU opportunity constraint that the stability margin is not lower than the threshold and the probability of default is not higher than the preset upper limit ε.
[0014] In the aforementioned GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid integration in weak power grids, the state vector in the mid-level security reinforcement learning tuning includes at least the following: ,in, For second-level observable harmonic sensitivity indicators, The magnitude of the estimated background impedance of the power grid. The phase angle is the estimated value of the background impedance of the power grid. For voltage deviation, RoCoF is the frequency deviation, α is the rate of change of frequency, and α is the frequency deviation. t-1 Let α be the GFM matching coefficient at time t-1. targe,t The target GFM ratio coefficient is issued by the upper-level scheduler at time t; the action vector includes at least [Δα]. RL H v D v ,k p ,k q ,ω pll ], where H v D represents the virtual inertia parameter. v k represents the virtual damping parameter. p k represents the active power-frequency droop control coefficient.q This represents the reactive power-voltage droop control coefficient. Define the barrier function as h(s) = TSI(s) - TSI min and satisfy control barrier constraints. , ,in, For the safety attenuation coefficient, s t Let a be the system state vector at time t. t Let be the action vector at time t. To take action a t Then, the predicted or estimated value of the state at the next moment; so that the policy prioritizes outputting actions that satisfy the safety boundary during training and online operation; The real-time proportioning coefficient is based on α. t =clip(α target ,t+Δα RL ,t,α min ,α max The calculation is performed, and feasible domain projection, rate of change constraint and minimum dwell time constraint are applied to the action to suppress mode chattering and control shocks.
[0015] The aforementioned GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid integration in weak power grids, specifically the device-level master-slave GFM / GFL hybrid control, involves the output of a reference voltage from the GFM channel. The correction current is output from the GFL channel. The correction voltage is obtained through virtual impedance mapping. The unified execution reference satisfies: ; ; in, Limit the correction amplitude to ensure that a single PWM modulation link can always be implemented; Before fusion, phase alignment, slope limiting, and admission criterion checks are performed on the GFL channels to avoid conflicts between parallel dual voltage sources.
[0016] In the aforementioned GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid connection in weak power grids, step 5 involves backoff control triggered based on a safety set Xsafe, which at least constrains voltage deviation, frequency deviation, and current amplitude. When the number of consecutive limit violations reaches a threshold, the dominant oscillation amplitude continues to increase, the frequency change rate exceeds a threshold, or the stability margin is lower than a threshold, conservative rule control is switched to within a preset time limit. After meeting the minimum dwell time and continuously maintaining no limit violations, no oscillation growth, and the state returning to the safety set, the backoff is released and safety reinforcement learning control is restored.
[0017] The beneficial effects of the present invention are: (1) By unifying and coordinating minute-level two-stage split-brush scheduling, second-level safe reinforcement learning tuning and millisecond-level safe execution, it is possible to simultaneously take into account economy, stability and power quality; (2) By adopting a master-slave hybrid control with GFM master control and GFL correction and a continuous ratio migration method, the phase angle change and current impact in the traditional hard switching mode are avoided; (3) By combining barrier function security constraints, action projection, current limiting and backoff control into a multi-layer security protection mechanism, the controllability and reliability of engineering deployment under extreme weak network conditions are improved; (4) By introducing digital twin knowledge base, agent model, meta-learning hot start and policy distillation, the convergence efficiency and online deployment interpretability in new scenarios are improved. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the multi-timescale collaborative framework of the present invention; Figure 2 This is a schematic diagram of the master-slave GFM / GFL hybrid control structure of the present invention; Figure 3 This is a flowchart of the method of the present invention; Figure 4 This is a flowchart illustrating the security protection and rollback process of this invention; Figure 5 The following is a performance comparison chart of the embodiments of the present invention and the comparative method, wherein (a) is a graph showing the change of grid connection point voltage over time; and (b) is a graph showing the change of grid connection point frequency over time. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] This embodiment discloses a GFM / GFL hybrid control collaborative optimization method for high-penetration renewable energy grid integration in weak power grids, such as... Figure 3 As shown, it includes the following steps: S1. Collect operating status data from the PCC side of the grid connection point. The operating status data includes at least three-phase voltage, three-phase current, active power P, reactive power Q, frequency deviation Δf, frequency change rate RoCoF, and harmonic sensitivity index bh. fast New energy output, load information, and background impedance estimation .
[0021] S2. Obtain the equivalent system strength based on the background impedance estimate. Preferably, the fundamental component is extracted through low-amplitude active injection and synchronous detection, according to... The background impedance estimate is obtained where To estimate the background impedance of the power grid at the k-th sampling time, This represents the increment of the fundamental voltage at the k-th sampling time. This represents the increment of the fundamental current at the k-th sampling time; then... The equivalent system strength is obtained; and the working conditions are divided into extremely weak network, weak network and strong network according to the equivalent system strength, so as to adjust the minimum GFM ratio, rate of change of motion and damping parameter boundary respectively.
[0022] S3. Execute the upper-level two-stage multi-objective scheduling using a minute-level scheduling cycle. The decision quantity for the first stage is x(1)=[P g Q g ,R,α target ], where P g For useful reference, Q g For reactive power reference, R is the reserve capacity, and α is the reactive power reference. target The target for GFM matching is y(ξ) = [ΔP]; the second-stage backtracking decision quantity is y(ξ) = [ΔP] g ,ΔQ g ,r ↑ ,r ↓ ,Δα,P curt ](ξ), where ξ represents the uncertainty, which includes at least the fluctuations in wind and solar power output, load fluctuations, and equivalent impedance disturbances, ΔP g ΔQ represents the active power adjustment. g Represents the reactive power adjustment amount, r ↑ This indicates an upward adjustment of the reserve adjustment amount, r ↓ This indicates a reduction in the reserve adjustment amount, Δα represents the adjustment amount for the GFM / GFL mixing ratio, and P curt This represents the amount of power abandoned. The scheduling objectives include at least operating cost objectives, stability risk objectives, and power quality objectives, and satisfy the Bruker chance constraint that the stability margin is not lower than a preset threshold and the probability of default is not higher than the upper limit ε.
[0023] S4. Perform mid-level security reinforcement learning tuning in seconds-level control cycles.
[0024] Constructing state vectors Construct action vector a t =[Δα RL H v D v ,k p ,k q ,ω pll ], where Δα RL H is the correction amount for the proportioning. v For virtual inertia parameters, D v k is a virtual damping parameter. p and k qThese are the active power-frequency droop coefficient and the reactive power-voltage droop coefficient, respectively. pll These are parameters related to the phase-locked loop.
[0025] S5. In the mid-layer security reinforcement learning tuning, a barrier function security constraint is introduced. The barrier function is defined as h(s) = TSI(s) - TSI. min Where TSI is the transient stability index; and it satisfies the control barrier constraint. , , The safety attenuation factor is used to limit the rate at which the safety margin decreases. The larger the value, the more stringent the requirements for safety in the next moment; The smaller the value, the slower the safety margin decreases; t Let a be the system state vector at time t. t Let be the action vector at time t. To take action a t Then, the predicted or estimated value of the state at the next time step is obtained. Preferably, a Lagrange multiplication strategy is used to update the reinforcement learning policy so that it prioritizes outputting actions that satisfy the safety boundary.
[0026] The significance of control barrier constraints lies in the fact that in the current state s t Take action a t Then, the safety margin for predicting the next moment should not decrease too quickly, and the system should try to always stay within the safety boundary.
[0027] S6. Merge the upper-level scheduling output with the middle-level actions to obtain the real-time allocation coefficient α. t =clip(α target ,t+Δα RL ,t,α min ,α max It also applies feasible domain projection, rate of change constraints, and minimum dwell time constraints to the action to suppress mode chattering and control shocks.
[0028] S7. Actuator-level master-slave GFM / GFL hybrid control. The GFM channel serves as the master control channel, outputting a reference voltage. The GFL channel outputs correction current as a correction channel. And it is mapped to a correction voltage through virtual impedance. ; Unified execution reference to meet ,in Preferably, phase alignment, admission verification, slope limiting, and minimum dwell time constraints are performed on the GFL channel before fusion to avoid conflicts between parallel dual voltage sources and reduce switching current surges.
[0029] S8. At the millisecond-level execution layer, the unified execution reference execution current is limited, the constraint look-ahead check is performed, the action projection is performed, and the rollback is controlled.
[0030] Building a security set When the number of consecutive limit violations reaches the threshold, the amplitude of the dominant oscillation continues to increase, the rate of change of frequency exceeds the threshold, or the stability margin is lower than the threshold, switch to conservative rule control within a preset time limit; after meeting the minimum dwell time and continuously maintaining no limit violations, no oscillation growth, and returning to the safe set, release the rollback and restore safe reinforcement learning control.
[0031] S9. Write the runtime data into the digital twin knowledge base and update the agent model or distillation strategy. Construct a digital twin knowledge base to record status, scheduling decisions, disturbances, actions, and performance labels. Train the agent model based on the knowledge base and perform meta-learning hot start and strategy distillation to accelerate parameter tuning under new operating conditions, reduce online deployment complexity, and improve control interpretability.
[0032] like Figure 1 As shown, this invention divides the control process into three time scales: upper-level scheduling, middle-level tuning, and lower-level execution. The upper-level scheduling cycle is 5 to 15 minutes, and Pg, Qg, R, and α are obtained based on load level, renewable energy output, system strength, and historical statistical information. target The middle layer uses a control cycle of 0.5 to 2 seconds to generate a ratio correction amount Δα based on the real-time status. RL And control parameter actions. The lower layer completes unified voltage / current inner loop control, current limiting and backoff protection with an execution cycle of 0.1 milliseconds to 1 millisecond.
[0033] In this embodiment, the active power P and reactive power Q of the grid connection point PCC are respectively expressed in the dq coordinate system through... , Calculate, where, Unified definition in synchronous rotating reference frame Below, v d For the d-axis voltage component, v q For the q-axis voltage component, i d For the d-axis current component, i q For the q-axis current component; where phase compensation is performed on the phase-locked loop estimated phase of the GFL channel to map the GFL channel and the GFM channel to a unified reference system, so as to reduce the control error caused by the aliasing of multiple reference systems.
[0034] In this embodiment, the upper-level two-stage distributed bar multi-objective scheduling targets operating cost, stability risk, and power quality. The operating cost target includes at least generation cost, reserve cost, and curtailment cost; the stability risk target includes at least system strength-related risk, ratio mismatch risk, and oscillation risk; and the power quality target includes at least voltage deviation, frequency deviation, and harmonic indices. Uncertainties related to wind and solar power output, load disturbances, and equivalent impedance disturbances are constrained using distributed fuzzy sets, and the default probability of the stability margin is limited by distributed bar chance constraints.
[0035] In this embodiment, the state vector of the mid-layer security reinforcement learning includes not only conventional voltage, frequency, and impedance, but also a second-level harmonic sensitivity index, bh. fast And the α issued by the upper level target This ensures that mid-level actions can be corrected within the constraints of the upper-level scheduling baseline. The action vector includes a ratio correction amount Δα. RL Virtual inertia H v Damping D v droop coefficient k p and k q and the phase-locked loop parameter ω pll .
[0036] In this embodiment, reinforcement learning employs a security policy update method with barrier function constraints. We define h(s) = TSI(s) - TSI min When predicting the state s_hat,t+1 at the next moment causes h(s_hat,t+1) to decrease too quickly or cross the safety boundary, the cost of the corresponding action is increased through barrier constraints. During online operation, action projection, slope limiting, flow limiting, and rollback protection are still retained as the last line of defense at the engineering level.
[0037] like Figure 2 As shown, the GFM channel is used to output a reference voltage, where the active power-frequency relationship can be expressed as ω*=ω0-k p (P-P0), the virtual synchronous machine dynamically satisfies 2H v *dω / dt=P*-PD v (ω-ω0), the reactive power-voltage relationship can be expressed as V*=V0-k q (Q-Q0). The GFL channel generates a correction current based on a phase-locked loop and a power outer loop. And mapped to virtual impedance The unified execution reference after integration satisfies... ,in .
[0038] In this embodiment, to avoid parallel dual-voltage source conflict between GFM and GFL, the phase difference δθ is gradually aligned before the GFL correction is injected: ; ; Where, ω gfm Indicates the angular frequency of the GFM channel. This represents the original current reference vector generated by the GFL channel in the dq coordinate system. This indicates the current reference that has undergone rotational compensation.
[0039] Rate of change of motion Set upper limit r α An admission condition is set for the mixing zone: entry into the mixing zone is only permitted when both the phase difference and voltage difference are below a preset threshold. This admission condition is used in conjunction with the minimum dwell time to reduce chattering and shocks during the post-fault recovery process.
[0040] like Figure 4 As shown, the safety protection and rollback process preferably adopts the following logic: First, the voltage deviation, frequency deviation, current amplitude, frequency change rate, and dominant oscillation amplitude are checked in advance; when the action exceeds the limit, feasible region projection is performed first; if continuous limit exceedance, continuous oscillation growth, frequency change rate exceeding the limit, or stability margin below the threshold are still detected after projection, conservative rule control is switched to within a preset time; when the minimum dwell time is met and there are no limit exceedances, no oscillation growth, and the system returns to the safe set, the rollback is released and reinforcement learning control is restored.
[0041] In an optional embodiment, when the equivalent system strength is in an extremely weak network, the minimum GFM ratio α is increased. min And tighten the boundary of the rate of change of motion, while raising the damping parameter D. v When the equivalent system strength is at the strong network level, k should be appropriately reduced. α And increase the weight of economic factors to reduce unnecessary support expenses.
[0042] In one alternative embodiment, a digital twin knowledge base Used to train the surrogate model Fφ, where z t It should include at least stability margin, harmonic parameters, recovery time, damping ratio, and the real part of key eigenvalues.
[0043] For the new working conditions, a meta-learning warm start is performed using scenario context vectors: ; ; Where c is the scene context vector (containing SCR interval, penetration rate, and fault label). After training is complete, policy distillation is performed: ; in, Represents the optimal student strategy parameters. This represents the weighting coefficient of the KL divergence term.
[0044] Complex strategies are compressed into lightweight student strategies for rapid online simulations and fault-tolerant takeover.
[0045] In an optional embodiment, to improve the deployability of the project, the upper-layer two-stage distributed bar scheduling module can be deployed on the scheduling master station or edge server, and the middle-layer security reinforcement learning tuning module and master-slave hybrid control execution module can be deployed in the field controller or grid-connected converter controller, and the slow variables and event quantities can be exchanged through the industrial communication network.
[0046] This embodiment discloses a GFM / GFL hybrid control and collaborative optimization system for high-penetration renewable energy grid connection in weak power grids. It includes a state acquisition and system strength assessment module, an upper-level two-stage sub-Browser scheduling module, a middle-level security reinforcement learning tuning module, a master-slave hybrid control execution module, and a security protection and rollback module. The state acquisition and system strength assessment module is used to collect real-time operating status data of the grid connection point, estimate background impedance, and calculate system strength. The upper-level two-stage sub-Browser scheduling module performs multi-objective scheduling according to a first preset scheduling cycle to obtain the first-stage decision quantity. The middle-level security reinforcement learning tuning module... The system generates a ratio correction quantity and control parameter action based on the operating status data collected by the status acquisition and system strength assessment module and the first-stage decision quantity generated by the upper-level two-stage sub-Blu-ray bar scheduling module, thereby obtaining the real-time ratio coefficient. The master-slave hybrid control execution module executes device-level master-slave GFM / GFL hybrid control based on the real-time ratio coefficient. The GFM channel provides a reference voltage, and the GFL channel generates a correction current and maps it to a correction voltage through virtual impedance to obtain a unified execution reference. The safety protection and rollback module is used to switch the system control mode, which includes conservative rule control and reinforcement learning control.
[0047] To verify the effectiveness of the method of this invention, this embodiment selects a typical weak power grid with high penetration of new energy grid connection scenario for simulation verification. An IEEE 30-node retrofit system is constructed as the test object, with a system baseline capacity of 100 MVA, and a new energy grid connection device using GFM / GFL hybrid control is configured at the grid connection point. The upper-level scheduling layer adopts a minute-level update mechanism, the middle-level security reinforcement learning tuning layer adopts a second-level update mechanism, and the lower-level execution layer adopts a millisecond-level control and protection mechanism. The simulation adopts a dual-time-domain configuration, where the EMT simulation step size is 20 μs, the RMS simulation step size is 10 ms, and the transient analysis window is 10 s.
[0048] Before the experiment, the system was stabilized near its rated operating point, and the steady-state grid connection point voltage, frequency, active power, reactive power, current, and estimated background impedance were recorded. The background impedance estimation results were used to determine that the current operating condition belonged to a weak network operating zone. Two sets of comparative schemes were then configured: the first set was the comparative method, using the original cooperative control method without the two-stage sub-Blu-ray bar scheduling and barrier function security reinforcement learning cooperative mechanism of this invention; the second set was the method of this invention, namely, the cooperative optimization method of "two-stage sub-Blu-ray bar scheduling + barrier function security reinforcement learning + master-slave GFM / GFL hybrid control + security protection and backoff". Except for the control algorithm, the network parameters, fault conditions, initial operating conditions, and measurement conditions were kept consistent between the two sets of schemes.
[0049] During the fault injection phase, a voltage dip disturbance is applied at the grid connection point at t=1.0, causing the grid connection point voltage to drop to 0.2 pu and remain there for 150 ms. The fault is cleared at t=1.15 s. During the fault and after fault clearing, the grid connection point voltage recovery curve, grid connection point frequency recovery curve, peak current, frequency change rate, over-limit duration, and total harmonic distortion are continuously collected. Simultaneously, the upper-level scheduling output, as well as the matching correction and real-time matching coefficient output from the middle-level controller, are recorded to analyze the closed-loop coordination process of the system from scheduling to control to execution before and after the fault.
[0050] During the data processing phase, the following indicators were calculated for both schemes: Total Harmonic Distortion (THD), Fault Recovery Time (RT), Stability Margin (SM), and Normalized Total Operating Cost. The recovery time is defined as the shortest time it takes for the grid-connected voltage and frequency to return to the allowable error band and remain stable. The stability margin is used to comprehensively characterize the recovery time, overshoot, and harmonic performance. The normalized total operating cost is used to characterize the economic efficiency of the control strategy while ensuring stable operation. To ensure the representativeness of the results, this embodiment uses a typical weak grid fault scenario as the final demonstration result, as shown in Table 1. The grid-connected voltage recovery curve and grid-connected frequency recovery curve are plotted as Chinese figures, as shown below. Figure 5 As shown.
[0051] Experimental results show that, under the same fault conditions, the method of the present invention has better overall performance than the comparative method. Specifically, the total harmonic distortion of the method of the present invention is reduced from 2.40% to 1.90%, a reduction of 20.8%; the fault recovery time is shortened from 0.218s to 0.111s, a reduction of 49.1%; the stability margin is increased from 0.4417 to 0.6117, an increase of 38.5%; and the normalized total operating cost is reduced from 1.000 to 0.9315, a reduction of 6.85%. In terms of safety, the maximum voltage deviation is reduced from 0.3521pu to 0.1845pu; the peak frequency change rate is reduced from 2.3456Hz / s to 1.2345Hz / s; and the total duration of over-limit is shortened from 456.78ms to 123.45ms. This demonstrates that the method of the present invention can achieve faster voltage and frequency recovery under fault disturbances in the context of high-penetration renewable energy grid connection in weak power grids, while taking into account power quality, stability and operational economy, thus verifying the effectiveness of the method of the present invention.
[0052] Table 1
[0053] The above embodiments are merely exemplary embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art can make various modifications or equivalent substitutions to the present invention within its scope and spirit, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of the present invention.
Claims
1. A GFM / GFL hybrid control and collaborative optimization system for high-penetration renewable energy grid integration in weak power grids, characterized in that, It includes a status acquisition and system strength assessment module, an upper-level two-stage sub-Blu-ray bar scheduling module, a middle-level security reinforcement learning tuning module, a master-slave hybrid control execution module, and a security protection and rollback module. The status acquisition and system strength assessment module is used to collect grid connection point operation status data in real time, estimate background impedance, and calculate system strength. The upper-level two-stage sub-Brow bar scheduling module performs multi-objective scheduling according to the first preset scheduling cycle to obtain the first-stage decision quantity; The mid-level security reinforcement learning tuning module generates the allocation correction quantity and control parameter action based on the operating status data collected by the state acquisition and system strength assessment module and the first-stage decision quantity generated by the upper-level two-stage sub-Blu-ray scheduling module, thereby obtaining the real-time allocation coefficient. The master-slave hybrid control execution module is based on real-time ratio coefficient execution device-level master-slave GFM / GFL hybrid control, wherein the GFM channel provides a reference voltage, and the GFL channel generates a correction current and maps it to a correction voltage through virtual impedance to obtain a unified execution reference. The security protection and rollback module is used to switch the system control mode, which includes conservative rule control and reinforcement learning control.
2. The GFM / GFL hybrid control and collaborative optimization system for high-penetration renewable energy grid integration in weak power grids according to claim 1, characterized in that, The operating logic of the safety protection and rollback module is as follows: it performs look-ahead checks on voltage deviation, frequency deviation, current amplitude, frequency change rate, and dominant oscillation amplitude; when an action exceeds the limit, it first performs feasible region projection; if continuous limit exceedance, continuous oscillation growth, frequency change rate exceeding the limit, or stability margin below the threshold are still detected after projection, it switches to conservative rule control within a preset time; when the minimum dwell time is met and there are no limit exceedances, no oscillation growth, and the system returns to the safe set, it releases the rollback and resumes reinforcement learning control.
3. The GFM / GFL hybrid control and collaborative optimization system for high-penetration renewable energy grid integration in weak power grids according to claim 1, characterized in that, The master-slave GFM / GFL hybrid control specifically involves: the GFM channel outputting a reference voltage. The correction current is output from the GFL channel. The correction voltage is obtained through virtual impedance mapping. The unified execution reference satisfies: ; ; in, Limit the correction amplitude to ensure that a single PWM modulation link is always available.
4. A GFM / GFL hybrid control and collaborative optimization method for high-penetration renewable energy grid integration in weak power grids, characterized in that, The collaborative optimization system according to any one of claims 1-3 includes the following steps: Step 1: Collect the operating status data of the PCC side of the grid connection point. The operating status data includes at least voltage, current, active power, reactive power, frequency deviation, frequency change rate, harmonic sensitivity index, new energy output, load information and background impedance estimate. Step 2: Perform upper-level two-stage distributed bar scheduling according to the first preset scheduling cycle to obtain the first-stage decision quantity. The first-stage decision quantity includes at least the active power reference P. g Reactive power reference Q g Reserve capacity R and GFM ratio target α target And obtain the robust executable interval or contingency plan set based on the backtracking decision set of uncertainties; Step 3: Perform mid-layer security reinforcement learning tuning according to the second preset control cycle, based on the operating status data and α. target Generate the ratio correction amount Δα RL and control parameter actions, and combine the ratio correction amount with the α target The real-time proportioning coefficient α is obtained by fusion. t ; Step 4, based on the real-time proportioning coefficient α t The execution device-level master-slave GFM / GFL hybrid control is implemented, where the GFM channel provides a reference voltage, and the GFL channel generates a correction current and maps it to a correction voltage through virtual impedance to obtain a unified execution reference. Step 5: For the unified execution reference current limiting, action projection, constraint look-ahead verification and backoff control, restore safety reinforcement learning control when the safety release condition is met, forming a closed-loop collaborative control with minute-level scheduling, second-level tuning and millisecond-level execution.
5. The GFM / GFL hybrid control and collaborative optimization method for high-penetration renewable energy grid connection in weak power grids according to claim 4, characterized in that, The background impedance estimate is obtained through low-amplitude active injection and synchronous detection, according to... Calculate the estimated background impedance, where To estimate the background impedance of the power grid at the k-th sampling time, This represents the increment of the fundamental voltage at the k-th sampling time. The increment of the fundamental current at the k-th sampling time; and according to The equivalent system strength is obtained, where K scr This is the base conversion factor from impedance to SCR. To prevent extremely small positive numbers from being divided by zero, the operating conditions are divided into extremely weak network, weak network, and strong network according to the equivalent system strength, so as to adjust the minimum GFM ratio, the upper limit of the rate of change of motion, and the damping parameters respectively.
6. The GFM / GFL hybrid control and collaborative optimization method for high-penetration renewable energy grid connection in weak power grids according to claim 4, characterized in that, In the aforementioned upper-level two-stage distributed bar scheduling, the decision quantity for the first stage is x(1)=[P g Q g ,R,α target The second-stage backtracking decision quantity is y(ξ)=[ΔP] g ,ΔQ g ,r ↑ ,r ↓ ,Δα,P curt ](ξ), where ξ represents the uncertainty, which includes at least the fluctuations in wind and solar power output, load fluctuations, and equivalent impedance disturbances, ΔP g ΔQ represents the active power adjustment. g Represents the reactive power adjustment amount, r ↑ This indicates an upward adjustment of the reserve adjustment amount, r ↓ This indicates a reduction in the reserve adjustment amount, Δα represents the adjustment amount for the GFM / GFL mixing ratio, and P curt This represents the power curtailment; the dispatching process simultaneously considers operating costs, stability risks, and power quality, and satisfies the BLU opportunity constraint that the stability margin is not lower than the threshold and the probability of default is not higher than the preset upper limit ε.
7. The GFM / GFL hybrid control and collaborative optimization method for high-penetration renewable energy grid connection in weak power grids according to claim 4, characterized in that, In the mid-layer security reinforcement learning tuning, the state vector includes at least the following: ,in, For second-level observable harmonic sensitivity indicators, The magnitude of the estimated background impedance of the power grid. The phase angle is the estimated value of the background impedance of the power grid. For voltage deviation, RoCoF is the frequency deviation, α is the rate of change of frequency, and α is the frequency deviation. t-1 Let α be the GFM matching coefficient at time t-1. targe,t The target GFM ratio coefficient is issued by the upper-level scheduler at time t; the action vector includes at least [Δα]. RL H v D v ,k p ,k q ,ω pll ], where H v D represents the virtual inertia parameter. v k represents the virtual damping parameter. p k represents the active power-frequency droop control coefficient. q This represents the reactive power-voltage droop control coefficient. Define the barrier function as h(s) = TSI(s) - TSI min and satisfy control barrier constraints. , ,in, For the safety attenuation coefficient, s t Let a be the system state vector at time t. t Let be the action vector at time t. To take action a t Then, the predicted or estimated value of the state at the next moment; so that the policy prioritizes outputting actions that satisfy the safety boundary during training and online operation; The real-time proportioning coefficient is based on α. t =clip(α target ,t+Δα RL ,t,α min ,α max The calculation is performed, and feasible domain projection, rate of change constraint and minimum dwell time constraint are applied to the action to suppress mode chattering and control shocks.
8. The GFM / GFL hybrid control and collaborative optimization method for high-penetration renewable energy grid connection in weak power grids according to claim 4, characterized in that, The device-level master-slave GFM / GFL hybrid control specifically involves: outputting a reference voltage from the GFM channel. The correction current is output from the GFL channel. The correction voltage is obtained through virtual impedance mapping. The unified execution reference satisfies: ; ; in, Limit the correction amplitude to ensure that a single PWM modulation link can always be implemented; Before fusion, phase alignment, slope limiting, and admission criterion checks are performed on the GFL channels to avoid conflicts between parallel dual voltage sources.
9. The GFM / GFL hybrid control and collaborative optimization method for high-penetration renewable energy grid integration in weak power grids according to claim 4, characterized in that, In step 5, the backoff control is triggered based on the safety set Xsafe, which at least constrains the voltage deviation, frequency deviation, and current amplitude. When the number of consecutive limit violations reaches the threshold, the amplitude of the dominant oscillation continues to increase, the rate of change of frequency exceeds the threshold, or the stability margin is lower than the threshold, the system switches to conservative rule control within a preset time limit. After the minimum dwell time is met and the system continuously maintains no limit violations, no oscillation growth, and the state returns to the safe set, the system releases the rollback and resumes safe reinforcement learning control.