Grid-connected inverter and off-grid inverter dual-mode adaptive switching control based on dqn

CN122801388APending Publication Date: 2026-09-22HEFEI UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610592829.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0012]本发明所要解决的技术问题是:并网逆变器切换中,在实际工况快速变化或阻抗未知的场景下,辨识过程往往存在实时性差、可靠性低以及引入额外扰动等问题

Benefits of technology

1.本发明无需依赖电网阻抗模型或SCR辨识,通过DQN算法直接挖掘逆变器输出的有功、THD等状态特征与电网强弱环境的内在关联,实现了在物理参数完全未知条件下的自适应决策,规避了建模误差对控制性能的影响。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122801388A_ABST
    Figure CN122801388A_ABST
Patent Text Reader

Abstract

The application discloses a grid-connected inverter follow-constructing network dual-mode adaptive switching control method based on DQN and belongs to the electrical engineering field. The method constructs a state space, extracts key characteristic parameters, designs a reward function, uses a DQN algorithm for training to obtain an optimal control mode, and uses the trained agent to realize stable operation of the grid-connected inverter in the optimal control mode. The method not only does not need to change the original control loop, but also can optimize the state space and the reward function according to actual requirements, and improve the stable operation boundary of the grid-connected inverter under the condition of a dynamic power grid. Compared with the prior art, the application effectively overcomes the condition of relying on power grid impedance identification or SCR estimation, realizes adaptive control mode selection of the grid-connected inverter under different power grid conditions, and realizes an intelligent mapping relationship between the control mode and the system operation state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of electrical engineering and relates to a dual-mode adaptive switching control method for grid-connected inverters based on DQN. Background Technology

[0002] With the large-scale grid connection of new energy sources such as wind power and photovoltaics, the power system operating environment exhibits characteristics such as an expanded range of grid impedance variations and grid strength uncertainty, posing new challenges to the stable operation of grid-connected inverters. Grid-following control and grid-connecting control have complementary stability characteristics under different grid conditions: GFL control has better dynamic performance in strong grids, while GFM control can maintain system stability under weak grid conditions. Therefore, achieving adaptive operation of grid-connected inverters under different grid conditions through dual-mode switching has become an important research direction. Existing dual-mode switching strategies typically select the control mode based on grid impedance identification or short-circuit ratio estimation. However, in actual operation, grid parameters are often difficult to obtain accurately, and the impedance identification process has limitations in real-time performance and reliability, making it unsuitable for scenarios where grid impedance is unknown or changes rapidly.

[0003] 1) The paper "Seamless switching method between grid-following and grid-forming control for renewable energy conversion systems" published in IEEE Transactions on Industry Applications, Volume 61, Issue 1 in 2024 (also published in IEEE Transactions on Industry Applications, Volume 61, Issue 1 in 2024) proposed a seamless switching method between grid-following and grid-forming control under grid-connected conditions. However, it can only adapt to a single grid-connected operating condition, lacks an adaptive switching mechanism based on dynamic changes in grid strength, has insufficient transient disturbance suppression capability, and does not take into account the stability and versatility of grid-forming control under strong grid conditions. It is difficult to meet the stable operation requirements of new energy conversion systems under all operating conditions and with high robustness in complex grid environments.

[0004] 2) Chinese patent document CN110021959A, authorized on April 2, 2019, entitled "Dual-mode control method for grid-connected inverters based on short-circuit ratio under weak power grids," discloses a dual-mode control method for grid-connected inverters based on short-circuit ratio under weak power grids. It uses the equivalent short-circuit ratio of the system to characterize the strength of the power grid, providing a basis for the dual-mode switching of the current source and voltage source of the grid-connected inverter, thereby improving the grid stability under weak power grids. However, the detection system usually relies on grid impedance identification or short-circuit ratio estimation to determine the grid strength. In actual new energy power plants, grid impedance is often difficult to obtain accurately and exhibits obvious fluctuation characteristics with changes in operating conditions. Although the impedance identification method based on signal injection can obtain grid parameters, its identification process may introduce additional disturbances. At the same time, there are certain difficulties in achieving online identification in multi-inverter systems.

[0005] 3) In his dissertation, "Research on Impedance Adaptive Dual-Mode Control of High-Penetration New Energy Power Generation Grid-Connected Inverters [D]" (Doctoral Dissertation, Hefei University of Technology, 2020. DOI:10.27101-d.cnki.ghfgu.2020. 000048), Li Ming used the D-segmentation method to plot the parameter stability domains of current-source mode grid-connected inverters and voltage-source mode grid-connected inverters as a function of SCR, and found the dual-mode switching boundary under multi-parameter constraints. However, in actual new energy power plant applications, harmonic disturbances introduced during the identification process may induce system instability, and there is a problem of difficulty in impedance identification. This makes it difficult for the adaptive switching control to identify the magnitude of the grid impedance in time when faced with sudden changes in grid impedance, leading to misjudgment of switching.

[0006] 4) The paper "The dual-mode combined control strategy for centralized photovoltaic grid-connected inverters based on double-split transformers," published in IEEE Transactions on Industrial Electronics, Volume 68, Issue 12, 2021, proposes three control schemes: a full current source, a hybrid mode, and a full voltage source. It uses the D-segmentation method to define the stability domain of different modes under grid impedance fluctuations, selecting the switching of grid-connected inverter control and effectively broadening the stable operating range of the inverter. However, in practical applications, the switching determination of this scheme highly depends on accurate grid impedance identification, making it difficult to obtain real-time and accurate impedance parameters.

[0007] 5) The paper "A Dual-Mode Grid-Connected Stability Control Strategy Based on Adaptive Grid Impedance in Weak Grids," published in the *Acta Energiae Solaris Sinica*, Vol. 42, No. 7, 2021, proposed an adaptive switching control scheme for current source-voltage source modes based on grid impedance identification. Starting from the relationship between the monotonicity and damping characteristics of power transmission and the magnitude of grid impedance, it derived the grid impedance stability operating boundaries for both grid-following control and grid-connecting control modes, achieving adaptive switching based on the grid impedance. However, in practical applications, the triggering of this scheme's switching action is highly dependent on the real-time identification results of the grid impedance.

[0008] 6) The paper "Data-driven optimal control strategy for virtual synchronous generator via deep reinforcement learning approach," published in Volume 9, Issue 4 of the *Journal of Modern Power Systems and Clean Energy* in 2021, transforms the problem of adaptive adjustment of virtual inertia and damping coefficients into a reinforcement learning decision-making task. It achieves adaptive optimization of control parameters based on the real-time operating state of the system, thus expanding the operating range. However, although this scheme uses reinforcement learning algorithms to achieve adaptive parameter optimization, it is only applicable to virtual synchronous generator control, and traditional single control cannot operate stably under a wide range of grid impedance fluctuations compared to dual-mode switching control.

[0009] In summary, the existing technology has the following problems: 1) In weak grid environments, traditional dual-mode switching control mostly relies on grid impedance identification or short-circuit ratio estimation. Grid parameters are difficult to obtain accurately, and the identification process has poor real-time performance and reliability, making it unsuitable for scenarios where grid impedance is unknown or fluctuates rapidly.

[0010] 2) Single control mode or traditional switching scheme cannot take into account both the dynamic performance of strong power grid and the stability support of weak power grid, and the scope of application is limited, making it difficult to meet the stable operation requirements under a wide range of power grid intensity fluctuations.

[0011] 3) Under the condition that both network-following control and network-building control can operate stably, the existing dual-mode switching lacks an adaptive decision based on the control objective. It only relies on preset thresholds or fixed rules to select a single control mode and cannot find the optimal control in this range. Summary of the Invention

[0012] The technical problem this invention aims to solve is that during grid-connected inverter switching, in scenarios with rapidly changing actual operating conditions or unknown impedance, the identification process often suffers from poor real-time performance, low reliability, and the introduction of additional disturbances. This invention introduces a Deep Q-Network (DQN) reinforcement learning mechanism to achieve adaptive decision-making for control modes without impedance identification. Through a reward function primarily based on stability, it effectively ensures that the system can automatically find the optimal control mode under unknown impedance changes, realizing intelligent mapping and collaborative optimization of control modes and complex power grid environments.

[0013] The technical solution of the present invention is as follows: A dual-mode adaptive switching control method for grid-connected inverters based on DQN is disclosed. The method involves a circuit comprising a grid-connected inverter, line impedance, and a three-phase power grid connected in sequence. An intelligent agent adjusts the operating state of the grid-connected inverter in real time to maintain its operational stability. The method includes the following steps: Step 1: Collect key characteristic parameters of the grid-connected inverter. S : Including the total harmonic distortion of the current of the grid-connected inverter THD and its rate of change ΔTHD Inverter output active power P and its rate of change ΔP ; Step 2: Construct the action space A of discrete actions, divide the control of the grid-connected inverter into grid-following control and grid-connecting control, and record the signal that changes the control mode of the inverter as the action signal. The action signal of grid-following control is 0, and the action signal of grid-connecting control is 1. Step 3, based on the key feature parameters extracted in Step 1 S A reward function R is constructed from two dimensions: adaptability and optimization ability. This reward function prioritizes the system's adaptability and then considers its optimization ability. An adaptability evaluation index is obtained through the reward function R. And the ability to strive for excellence assessment indicators ; Step 4, based on key feature parameters S The action space A in step 2 and the reward function R in step 3 are trained using the DQN algorithm to obtain the optimal control mode Ψ(s). t ), s t To match the optimal control mode Ψ(s) t The corresponding key feature parameters; Step 5: The reward function R is used to evaluate the performance of the agent in the dual-mode adaptive switching control of the grid-connected inverter, obtaining the optimal agent. Then, the optimal agent is used to perform dual-mode adaptive switching control to achieve the optimal control mode Ψ(s). t Stable operation of grid-connected inverters.

[0014] Preferably, the reward function R in step 3 includes a control adaptation reward. m 1. Adaptability Rewards m 2. And the optimization capability reward, wherein the optimization capability reward includes a first power quality optimization reward. n 1. Second power quality optimization reward n 2 and third power quality optimization rewards n 3; The expression for the reward function R is:

[0015] in, P ref This is the reference value for the active power output of the inverter; The adaptability assessment indicators And the ability to strive for excellence assessment indicators The expression is: m= m 1+ m 2, = n 1+ n 2+ n 3.

[0016] Preferably, the DQN algorithm in step 4 comprises three neural networks: a target Q-network, an online Q-network, and an action evaluation network, wherein the neural network parameters of the target Q-network are denoted as θ. Q The neural network parameters of the online Q-network are denoted as θ. Q ’ The neural network parameters of the action evaluation network are denoted as θ; The optimal control mode Ψ(s) t The expression for ) is as follows:

[0017] In the formula, s t =( THD t , ΔTHD t , P t , ΔP t ), THD t , ΔTHD t , P t , ΔP t These represent the total harmonic distortion (THD) of the current, the rate of change of the THD, the inverter output active power, and the rate of change of the inverter output active power, respectively, corresponding to the optimal control mode; at It is the optimal control mode Ψ(s) t The output action signal, a t = (0, 1).

[0018] Preferably, the implementation process of step 5 is as follows: The expression for the optimal agent is given as follows:

[0019] in, C For comprehensive evaluation indicators, x Let i be the number of action signals emitted by the agent in one evaluation, and let i be the sequence number of the action signals emitted by the agent in one training session. Then, the following judgment is made: like C ≤ xm 1. In this case, the grid-connected inverter operates in an unstable mode. like C > xm 1 but C ≠ x ( m If 1+n), then the grid-connected inverter operates in a non-optimal control mode, but the grid-connected inverter maintains stable operation; like C > xm 1 and C = x ( m If 1+n), then the optimal intelligent agent controls the grid-connected inverter to operate stably in the optimal control mode Ψ(s). t )Down.

[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention does not rely on grid impedance models or SCR identification. It directly mines the intrinsic correlation between the active power and THD state characteristics of the inverter output and the strong and weak grid environment through the DQN algorithm, realizing adaptive decision-making under the condition of completely unknown physical parameters and avoiding the impact of modeling errors on control performance.

[0021] 2. This invention constructs a hierarchical reward function based on stability. By using the order-of-magnitude weight difference between THD and power deviation, it ensures that the stability index is the first prerequisite when the agent searches for a strategy, and pursues the optimization of power tracking performance while ensuring the safe operation of the power grid.

[0022] 3. Unlike static methods that only focus on the current sampling point, this invention introduces electrical quantities with time-series characteristics such as power change rate, THD change rate, and delay signal into the state space. This enables the agent to capture the evolution trend of grid impedance changes, and can not only cope with static steady-state conditions, but also quickly find the optimal control in the transient process of rapid fluctuations in grid strength.

[0023] 4. This invention fully utilizes the large-scale parallel computing capabilities in an offline simulation environment to find the optimal intelligent agent and the optimal control law through a neural network. In the online application stage, the optimal control can be automatically found under different working conditions simply by using the intelligent agent.

[0024] 5. This invention utilizes the nonlinear function approximation capability of deep neural networks to learn complex nonlinear laws that are difficult to quantify in traditional control theory. When facing extreme uncertain environments caused by the high penetration rate of new energy sources, the intelligent agent can switch control modes through adaptive strategies, thereby expanding the stability margin of traditional single control. Attached Figure Description

[0025] Figure 1 This is a flowchart of the grid-connected inverter and grid-connected dual-mode adaptive switching control based on DQN as described in this invention.

[0026] Figure 2 This is a control block diagram of the control method of the present invention.

[0027] Figure 3 This is a diagram illustrating the training process of the intelligent agent described in this invention.

[0028] Figure 4 This refers to the action signal given by the intelligent agent of this invention.

[0029] Figure 5 The simulation results show the output active and reactive power when the optimal intelligent agent control of this invention is stable. Detailed Implementation

[0030] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0031] Figure 2 This is a control diagram of the control method of the present invention. As can be seen from the upper part of the diagram, the circuit involved in the control method includes a grid-connected inverter, a line impedance, and a three-phase power grid connected in sequence.

[0032] In addition, from Figure 2 It can be seen that a DC capacitor C is also included in the grid-connected inverter. dc And LC filters, LC filters include filter inductors L f Filter capacitor C f and resistance R f . Figure 2 R ong and Lg These are the line impedance and line reactance, respectively, in the line impedance.

[0033] Figure 1 This is a flowchart of the control method of the present invention. Figure 1 and Figure 2 As can be seen, this invention provides a grid-connected inverter-grid dual-mode adaptive switching control method based on DQN, which adjusts the operating state of the grid-connected inverter in real time through an intelligent agent to maintain the stability of the grid-connected inverter during operation, including the following steps: Step 1: Collect key characteristic parameters of the grid-connected inverter. S : Including the total harmonic distortion of the current of the grid-connected inverter THD and its rate of change ΔTHD Inverter output active power P and its rate of change ΔP .

[0034] Step 2: Construct the action space A of discrete actions, divide the control of the grid-connected inverter into grid-following control and grid-connecting control, and record the signal that changes the control mode of the inverter as the action signal, where the action signal of grid-following control is 0 and the action signal of grid-connecting control is 1.

[0035] Step 3, based on the key feature parameters extracted in Step 1 S A reward function R is constructed from two dimensions: adaptability and optimization ability. This reward function prioritizes the system's adaptability and then considers its optimization ability. An adaptability evaluation index is obtained through the reward function R. And the ability to strive for excellence assessment indicators .

[0036] In this embodiment, the reward function R includes a control adaptation reward. m 1. Adaptability Rewards m 2. And the optimization capability reward, wherein the optimization capability reward includes a first power quality optimization reward. n 1. Second power quality optimization reward n 2 and third power quality optimization rewards n 3; The expression for the reward function R is:

[0037] in, P ref This is the reference value for the active power output of the inverter.

[0038] The adaptability assessment indicators And the ability to strive for excellence assessment indicators The expression is: m=m 1+ m 2, = n 1+ n 2+ n 3.

[0039] Step 4, based on key feature parameters S The action space A in step 2 and the reward function R in step 3 are trained using the DQN algorithm to obtain the optimal control mode Ψ(s). t ), s t To match the optimal control mode Ψ(s) t The key feature parameters corresponding to ).

[0040] In this embodiment, the DQN algorithm comprises three neural networks: a target Q-network, an online Q-network, and an action evaluation network. The neural network parameters of the target Q-network are denoted as θ. Q The neural network parameters of the online Q-network are denoted as θ. Q ’ The neural network parameters of the action evaluation network are denoted as θ; The optimal control mode Ψ(s) t The expression for ) is as follows:

[0041] In the formula, s t =( THD t , ΔTHD t , P t , ΔP t ), THD t , ΔTHD t , P t , ΔP t They are respectively with Total harmonic distortion of current in grid-connected inverters under optimal control mode 、 The rate of change of total harmonic distortion of current, inverter output active power, and the rate of change of inverter output active power; a t It is the optimal control mode Ψ(s) t The output action signal, a t = (0, 1).

[0042] Step 5: The reward function R is used to evaluate the performance of the agent in the dual-mode adaptive switching control of the grid-connected inverter, obtaining the optimal agent. Then, the optimal agent is used to perform dual-mode adaptive switching control to achieve the optimal control mode Ψ(s). t Stable operation of grid-connected inverters.

[0043] In this embodiment, step 5 is implemented as follows: The expression for the optimal agent is given as follows:

[0044] in, C For comprehensive evaluation indicators, x Let i be the number of action signals emitted by the agent in one evaluation, and let i be the sequence number of the action signals emitted by the agent in one training session. Then, the following judgment is made: like C ≤ xm 1. In this case, the grid-connected inverter operates in an unstable mode. like C > xm 1 but C ≠ x ( m If 1+n), then the grid-connected inverter operates in a non-optimal control mode, but the grid-connected inverter maintains stable operation; like C > xm 1 and C = x ( m If 1+n), then the optimal intelligent agent controls the grid-connected inverter to operate stably in the optimal control mode Ψ(s). t )Down.

[0045] Simulations were performed to demonstrate the beneficial effects of the present invention.

[0046] In this embodiment, the simulation results of a test condition where the grid impedance varies over a wide range from strong grid to weak grid back to strong grid are taken as an example. , , , , .

[0047] Figure 3 This is a diagram illustrating the training process of the intelligent agent described in this invention. Specifically, taking... , , , , The reward value initially decreased because the agent was randomly exploring in the unknown environment and had not yet learned an effective control strategy. After continuous training, the reward value gradually converged to a maximum of over 500, indicating that the agent autonomously learned and discovered the optimal grid-connected control strategy. It's worth noting that in the later stages of training, even after converging to the maximum reward value, the agent might still explore, leading to a decrease in the reward value. This is due to the greedy algorithm causing the agent to continuously explore.

[0048] Figure 4 This refers to the action signal provided by the intelligent agent of this invention. Specifically, when evaluating using this invention, the intelligent agent outputs discrete control mode action signals based on the key characteristic parameters of the inverter: an action value of 0 corresponds to the grid-following control mode, and an action value of 1 corresponds to the grid-building control mode. This figure visually demonstrates that under conditions of dynamically changing grid strength and unknown line impedance, the intelligent agent can autonomously and in real-time adaptively switch between the grid-following and grid-building control modes, verifying the effectiveness of the DQN algorithm in switching control under unknown grid conditions.

[0049] Figure 5 The simulation results show the output active power P and reactive power Q when the optimal intelligent agent control of this invention is stable. Specifically, during the adaptive switching of control modes, the active power P can quickly track the reference value without significant overshoot or oscillation; the reactive power Q output is stable, effectively supporting the grid voltage. This waveform verifies that the proposed control strategy can simultaneously meet the requirements of system stability and power point tracking performance, achieving stable operation under unknown conditions.

[0050] In summary, Figure 3 , Figure 4 , Figure 5 The simulation results shown are consistent with the results of this invention, and the adaptive dual-mode switching control for unknown impedance changes is effectively realized.

Claims

1. A grid-connected inverter-grid dual-mode adaptive switching control method based on DQN, wherein the control method involves a circuit comprising a grid-connected inverter, line impedance, and a three-phase power grid connected in sequence; characterized in that, To maintain the stability of grid-connected inverters by adjusting their operating status in real time through an intelligent agent, the following steps are included: Step 1: Collect key characteristic parameters of the grid-connected inverter. S : Including the total harmonic distortion of the current of the grid-connected inverter THD and its rate of change ΔTHD Inverter output active power P and its rate of change ΔP ; Step 2: Construct the action space A of discrete actions, divide the control of the grid-connected inverter into grid-following control and grid-connecting control, and record the signal that changes the control mode of the inverter as the action signal. The action signal of grid-following control is 0, and the action signal of grid-connecting control is 1. Step 3, based on the key feature parameters extracted in Step 1 S A reward function R is constructed from two dimensions: adaptability and optimization ability. This reward function prioritizes the system's adaptability and then considers its optimization ability. An adaptability evaluation index is obtained through the reward function R. And the ability to strive for excellence assessment indicators ; Step 4, based on key feature parameters S The action space A in step 2 and the reward function R in step 3 are trained using the DQN algorithm to obtain the optimal control mode Ψ(s). t ), s t To match the optimal control mode Ψ(s) t The corresponding key feature parameters; Step 5: The reward function R is used to evaluate the performance of the agent in the dual-mode adaptive switching control of the grid-connected inverter, obtaining the optimal agent. Then, the optimal agent is used to perform dual-mode adaptive switching control to achieve the optimal control mode Ψ(s). t Stable operation of grid-connected inverters.

2. The grid-connected inverter-grid dual-mode adaptive switching control method based on DQN according to claim 1, characterized in that, The reward function R described in step 3 includes a reward for controlling adaptive ability. m 1. Adaptability Rewards m 2. And the optimization capability reward, wherein the optimization capability reward includes a first power quality optimization reward. n 1. Second power quality optimization reward n 2 and third power quality optimization rewards n 3; The expression for the reward function R is: in, P ref This is the reference value for the active power output of the inverter; The adaptability assessment indicators And the ability to strive for excellence assessment indicators The expression is: m= m 1+ m 2, = n 1+ n 2+ n 3.

3. The grid-connected inverter-grid dual-mode adaptive switching control method based on DQN according to claim 1, characterized in that, The DQN algorithm described in step 4 comprises three neural networks: a target Q-network, an online Q-network, and an action evaluation network. The neural network parameters of the target Q-network are denoted as θ. Q The neural network parameters of the online Q-network are denoted as θ. Q ’ The neural network parameters of the action evaluation network are denoted as θ; The optimal control mode Ψ(s) t The expression for ) is as follows: In the formula, s t =( THD t , ΔTHD t , P t , ΔP t ), THD t , ΔTHD t , P t , ΔP t These represent the total harmonic distortion (THD) of the current, the rate of change of the THD, the inverter output active power, and the rate of change of the inverter output active power, respectively, corresponding to the optimal control mode; a t It is the optimal control mode Ψ(s) t The output action signal, a t = (0, 1).

4. The grid-connected inverter-grid dual-mode adaptive switching control method based on DQN according to claim 1, characterized in that, The implementation process of step 5 is as follows: The expression for the optimal agent is given as follows: in, C For comprehensive evaluation indicators, x Let i be the number of action signals emitted by the agent in one evaluation, and let i be the sequence number of the action signals emitted by the agent in one training session. Then, the following judgment is made: like C ≤ xm 1. In this case, the grid-connected inverter operates in an unstable mode. like C > xm 1 but C ≠ x ( m If 1+n), then the grid-connected inverter operates in a non-optimal control mode, but the grid-connected inverter maintains stable operation; like C > xm 1 and C = x ( m If 1+n), then the optimal intelligent agent controls the grid-connected inverter to operate stably in the optimal control mode Ψ(s). t )Down.

Citation Information

Patent Citations

  • Double-mode control method for grid-connected inverter based on short-circuit ratio under weak power grid

    CN110021959A