A control method and control device of a grid-connected converter and an electric energy router
By dynamically adjusting the adaptive coefficients of the intelligent model to coordinate the control mode of the grid-connected converter, the problem of unstable operation under different power grid conditions in the existing technology is solved, and optimal coordination and stability improvement are achieved under all operating conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHANGSHA UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2026-03-17
- Publication Date
- 2026-05-12
AI Technical Summary
The existing control strategies for grid-connected converters exhibit synchronization performance contradictions under both strong and weak grid conditions, making it difficult to achieve optimal coordination in complex and ever-changing grid environments, leading to operational instability.
By introducing an intelligent model, real-time power grid operation status information is acquired, and adaptive coefficients are dynamically output using a decision model. This coordinates the integration of grid-following control mode and grid-building control mode, generates control commands, and achieves adaptive control.
Achieving the optimal balance between tracking accuracy and support capability under all operating conditions enhances the system's dynamic response and stability, ensuring a smooth transition of control strategies and the system's robustness.
Smart Images

Figure CN121863529B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of power electronics technology, and in particular relates to a control method, control device and power router for a grid-connected converter. Background Technology
[0002] With the large-scale integration of intermittent renewable energy sources such as wind power and photovoltaics into the power system, the strength and inertia of the power grid are showing a downward trend, which puts forward higher requirements for the operational stability and control performance of grid-connected converters. As a key interface between renewable energy power generation units and the power grid, the control strategy of grid-connected converters needs to meet multiple objectives such as fast synchronization tracking, dynamic stability support, and fault ride-through in a complex and ever-changing power grid environment, which has become the research focus of the current power electronics and system control field.
[0003] In existing technologies, grid-connected converters mainly adopt a single grid-following control or grid-building control strategy. Grid-following control relies on phase-locked loops to track the grid voltage phase in real time, which has excellent synchronization performance under strong grid conditions. However, its control is essentially "grid following" and cannot provide active support for the system in terms of voltage and frequency. It is prone to instability and oscillation under weak grid or fault conditions. Grid-building control can autonomously build and stabilize AC voltage by simulating the external characteristics of a synchronous generator, providing the grid with the necessary inertia and damping. However, its dynamic response is relatively slow, and there is a risk of synchronization instability when the grid frequency fluctuates rapidly or the phase changes abruptly. These two single control modes have inherent contradictions in performance and are difficult to meet the needs of different operating conditions. Summary of the Invention
[0004] The purpose of this application is to provide a control method, control device, and power router for a grid-connected converter. The control method, control device, and power router provided in this application introduce an intelligent model that can make autonomous decisions based on real-time grid operating status information, and fuse two control signals based on the adaptive coefficients dynamically output by the model. This fundamentally overcomes the limitation of the fixed performance of a single control mode and effectively solves the problem in the prior art that it is impossible to achieve optimal coordination according to changes in operating conditions.
[0005] This application provides a control method for a grid-connected converter, the method comprising:
[0006] Real-time acquisition of power grid operating status information;
[0007] The operational status information is input into a trained decision model, which then dynamically outputs at least one adaptive coefficient to coordinate the network-following control mode and the network-building control mode based on the operational status information.
[0008] Based on the at least one adaptive coefficient, the first control signal from the network control link and the second control signal from the network construction control link are fused to generate a control command.
[0009] The operation of the grid-connected converter is controlled according to the control command.
[0010] Optionally, the operating status information includes: the short-circuit ratio characterizing the grid strength, the deviation between the output power and the reference power of the grid-connected converter, and the deviation between the common coupling point voltage and the reference voltage.
[0011] Optionally, the method further includes:
[0012] The decision model is trained using a dual-delay deep deterministic policy gradient algorithm.
[0013] Optionally, the decision model includes: an online Actor network and two online Critic networks, as well as a corresponding Actor target network and two Critic target networks.
[0014] Optionally, the decision model trained based on the dual-delay deep deterministic policy gradient algorithm includes:
[0015] Initialize the parameters of the Actor online network, the two Critic online networks, the Actor target network, and the two Critic target networks, and create an experience replay pool;
[0016] At each time step, the acquired current training state information is input into the Actor online network to obtain the current state action. Noise is added to the current state action to obtain a noisy action. Based on the current training state information and the preset reward function, the reward value is obtained. The next training state information is obtained, and the current training state information, the current state action, the noisy action, the reward value, and the next training state information are stored as experience data in the experience replay pool.
[0017] Experience data is sampled from the experience replay pool. Based on the sampled data, the Actor target network and the two Critic target networks are used to obtain the target action value. The parameters of the two Critic online networks are updated with the goal of minimizing the error between the action value output by the two Critic online networks based on the sampled data and the target action value.
[0018] Every preset number of update steps, the Actor online network is updated using a deterministic strategy gradient update, and the parameters of the Actor target network and the two Critic target networks are updated using a soft update method.
[0019] Repeat the above steps until the preset training termination condition is met.
[0020] Optionally, the current training state information includes the deviation between the training output power and the reference power, and the deviation between the training common coupling point voltage and the reference voltage;
[0021] The reward value is negatively correlated with the deviation between the training output power and the reference power, and with the deviation between the training common coupling point voltage and the reference voltage.
[0022] Optionally, the at least one adaptive coefficient includes: a network-following adaptive coefficient and a network-building adaptive coefficient, wherein the sum of the network-following adaptive coefficient and the network-building adaptive coefficient is 1.
[0023] Optionally, fusing the first control signal from the network control link and the second control signal from the network construction control link according to the at least one adaptive coefficient includes:
[0024] The first control signal from the network tracking control link and the second control signal from the network construction control link are weighted and summed according to the weights determined by the network tracking adaptive coefficient and the network construction adaptive coefficient.
[0025] This application also provides a control device for a grid-connected converter, the control device comprising:
[0026] The acquisition module is used to acquire real-time operating status information of the power grid;
[0027] The decision module is used to input the operating status information into the trained decision model, and the decision model dynamically outputs at least one adaptive coefficient for coordinating the network-following control mode and the network-building control mode based on the operating status information.
[0028] The fusion module is used to fuse the first control signal from the network control link and the second control signal from the network construction control link according to the at least one adaptive coefficient, and generate control commands.
[0029] The control module is used to control the operation of the grid-connected converter according to the control commands.
[0030] This application also provides an electric power router, including the control device, grid-connected converter, converter and hybrid energy storage module as described above;
[0031] The converter is connected between the grid-connected converter and the hybrid energy storage module to achieve power transmission and electrical isolation.
[0032] The hybrid energy storage module includes a lithium battery and a supercapacitor connected in parallel.
[0033] Compared with existing technologies, the control method, control device, and power router for a grid-connected converter provided in this application acquire real-time grid operating status information and input the operating status information into a trained decision model. The decision model dynamically outputs at least one adaptive coefficient to coordinate the grid-following control mode and the grid-connecting control mode based on the operating status information. Based on the at least one adaptive coefficient, the first control signal from the grid-following control link and the second control signal from the grid-connecting control link are fused to generate control commands. The grid-connected converter is then controlled according to these control commands. This application introduces an intelligent model capable of autonomously making decisions based on real-time grid operating status information and fuses the two control signals based on the adaptive coefficients dynamically output by the model. This fundamentally overcomes the limitation of the fixed performance of a single control mode and effectively solves the problem in existing technologies that cannot achieve optimal coordination according to changes in operating conditions. Specifically, this application achieves the following technical effects:
[0034] 1. Achieve optimal adaptive coordination under all operating conditions: The decision model autonomously and dynamically outputs coordinated control coefficients based on real-time power grid operation status information, enabling the grid-connected converter to adaptively adjust its control characteristics. Under operating conditions requiring high-precision tracking, the grid-following control component is enhanced, and under operating conditions requiring active support, the grid-building control component is enhanced, thereby achieving the optimal balance between tracking accuracy and support capability across all operating conditions.
[0035] 2. Improve system dynamic response and stability: The fusion coefficient, which is dynamically adjusted based on real-time status information, enables the control system to respond quickly to changes in grid power and voltage. When disturbances occur, the smooth and rapid adjustment of the control strategy focus not only effectively avoids the instability risk under a single control mode, but also significantly improves the overall transient stability and dynamic quality of the system.
[0036] 3. Ensure smooth and disturbance-free transition of control strategy: The continuous fusion of two control signals is achieved through continuously changing adaptive coefficients, which completely avoids the output jump and system oscillation caused by sudden changes in control mode or mismatch of fixed weights in traditional schemes, and ensures the stability and reliability of the system during various transition processes.
[0037] 4. Improving the control system with intelligent adaptability: By adopting a trained decision-making model, the control system is able to make online decisions based on real-time data, thereby proactively adapting to the complex and ever-changing power grid operating environment and improving the overall robustness and adaptability of the system under unknown or time-varying operating conditions. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart of a control method for a grid-connected converter disclosed in an embodiment of this application;
[0040] Figure 2 This is a schematic diagram of the structure of a control device for a grid-connected converter disclosed in an embodiment of this application;
[0041] Figure 3 This is a schematic diagram of the structure of a power router disclosed in an embodiment of this application;
[0042] Figure 4 This is a circuit diagram of a power router disclosed in an embodiment of this application;
[0043] Figure 5 This is a circuit diagram of a power submodule disclosed in an embodiment of this application;
[0044] Figure 6 This is a circuit diagram of another power submodule disclosed in an embodiment of this application. Detailed Implementation
[0045] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0046] It should be noted that when a component is referred to as being "fixed to" or "set on" another component, it can be directly on or indirectly set on the other component; when a component is referred to as being "connected to" another component, it can be directly connected to or indirectly connected to the other component.
[0047] It should be noted that the structures, proportions, sizes, etc., shown in the accompanying drawings of this specification are only for the purpose of assisting those skilled in the art in understanding and reading the content disclosed in the specification, and are not intended to limit the conditions under which this application can be implemented. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in the proportions, or adjustments to the size should still fall within the scope of the technical content disclosed in this application, provided that they do not affect the effects and purposes that this application can produce.
[0048] like Figure 1 As shown, this application provides a control method for a grid-connected converter, the method comprising:
[0049] S11. Obtain real-time power grid operation status information;
[0050] In this embodiment, the operating status information refers to key electrical quantities that reflect the current operating conditions of the power grid and the output status of the grid-connected converter. This information is collected and calculated in real time by sensors and measurement units, providing accurate and timely data input for subsequent intelligent decision-making. This step ensures that the control strategy is formulated based on an accurate perception of the actual state of the power grid, which is a prerequisite for achieving adaptive control.
[0051] S12. Input the operating status information into the trained decision model, and the decision model dynamically outputs at least one adaptive coefficient for coordinating the network-following control mode and the network-building control mode based on the operating status information.
[0052] In this embodiment, the trained decision model has the ability to map the optimal decision in real time based on the input information. After receiving the real-time operating status information, the model calculates and outputs one or more adaptive coefficients for coordinating the two control modes online through the complex mapping relationship learned internally. The coefficients are continuously changing, and their values directly determine the contribution ratio of network-following control and network-building control in the final control command under the current operating conditions, thereby realizing an intelligent decision-making process from sensing the state to generating a strategy.
[0053] S13. Based on at least one adaptive coefficient, fuse the first control signal from the network control link and the second control signal from the network construction control link to generate a control command.
[0054] In this embodiment, the fusion process is manifested as a synthesis based on dynamic weights. Specifically, the core of the fusion operation is to dynamically construct the phase of the virtual internal potential of the grid-connected converter. The phase of the virtual internal potential consists of two parts: one part is the phase tracking signal generated by the grid-following control link (such as the phase-locked loop), and the other part is the phase construction signal generated by the grid-building control link (such as power synchronization control). The adaptive coefficients output by the decision model are used as weights to perform weighted calculations on the two phase signals, thereby synthesizing the final virtual internal potential phase command. This process realizes the smooth and disturbance-free fusion of the two control modes. The generated final control command has both the fast tracking capability of grid-following control and the active support characteristics of grid-building control, and its hybrid characteristics can be dynamically adjusted according to the coefficients.
[0055] S14. Control the operation of the grid-connected converter according to the control command.
[0056] In this embodiment, the fused control commands (usually expressed as reference values for voltage or current) are applied to the pulse width modulation (PWM) or corresponding control unit of the grid-connected converter, thereby driving the operation of the power switching devices and ultimately controlling the output of the AC side of the converter. In this way, the actual operating characteristics of the grid-connected converter can be adjusted in real time according to the control commands, so that it exhibits stronger grid following when needed and stronger grid support capability at other times, thereby dynamically adapting to the current grid demand.
[0057] Compared with existing technologies, the control method, control device, and power router for a grid-connected converter provided in this application acquire real-time grid operating status information and input the operating status information into a trained decision model. The decision model dynamically outputs at least one adaptive coefficient to coordinate the grid-following control mode and the grid-connecting control mode based on the operating status information. Based on the at least one adaptive coefficient, the first control signal from the grid-following control link and the second control signal from the grid-connecting control link are fused to generate control commands. The grid-connected converter is then controlled according to these control commands. This application introduces an intelligent model capable of autonomously making decisions based on real-time grid operating status information and fuses the two control signals based on the adaptive coefficients dynamically output by the model. This fundamentally overcomes the limitation of the fixed performance of a single control mode and effectively solves the problem in existing technologies that cannot achieve optimal coordination according to changes in operating conditions. Specifically, this application achieves the following technical effects:
[0058] 1. Achieve optimal adaptive coordination under all operating conditions: The decision model autonomously and dynamically outputs coordinated control coefficients based on real-time power grid operation status information, enabling the grid-connected converter to adaptively adjust its control characteristics. Under operating conditions requiring high-precision tracking, the grid-following control component is enhanced, and under operating conditions requiring active support, the grid-building control component is enhanced, thereby achieving the optimal balance between tracking accuracy and support capability across all operating conditions.
[0059] 2. Improve system dynamic response and stability: The fusion coefficient, which is dynamically adjusted based on real-time status information, enables the control system to respond quickly to changes in grid power and voltage. When disturbances occur, the smooth and rapid adjustment of the control strategy focus not only effectively avoids the instability risk under a single control mode, but also significantly improves the overall transient stability and dynamic quality of the system.
[0060] 3. Ensure smooth and disturbance-free transition of control strategy: The continuous fusion of two control signals is achieved through continuously changing adaptive coefficients, which completely avoids the output jump and system oscillation caused by sudden changes in control mode or mismatch of fixed weights in traditional schemes, and ensures the stability and reliability of the system during various transition processes.
[0061] 4. Improving the control system with intelligent adaptability: By adopting a trained decision-making model, the control system is able to make online decisions based on real-time data, thereby proactively adapting to the complex and ever-changing power grid operating environment and improving the overall robustness and adaptability of the system under unknown or time-varying operating conditions.
[0062] As one implementation method, in this application embodiment, the operating status information includes: the short-circuit ratio characterizing the grid strength, the deviation between the output power and the reference power of the grid-connected converter, and the deviation between the voltage at the common coupling point and the reference voltage.
[0063] In this embodiment, the short-circuit ratio, output power deviation, and common coupling point voltage deviation are selected as core operating status information. This aims to provide the decision-making model with key inputs that can comprehensively and accurately characterize the current grid operating conditions and system operating requirements. Specifically, the short-circuit ratio directly quantifies the grid's connection strength and is the fundamental basis for determining whether the control strategy should favor following the grid (strong grid) or building a new grid (weak grid). The output power deviation reflects the tracking accuracy of the grid-connected converter to the power command and is a basic requirement for ensuring power quality and dispatch response. The common coupling point voltage deviation characterizes the system's demand for voltage support and is a key signal that triggers the active adjustment role of grid-building control. These three together constitute a multi-dimensional and complementary observation set, enabling the decision-making model to comprehensively evaluate the grid's stability requirements, power balance status, and voltage support requirements, thereby providing sufficient and necessary judgment basis for generating adaptive coordination coefficients.
[0064] As one implementation method, in this embodiment of the application, the method further includes:
[0065] S21. The decision model is obtained by training based on the dual-delay deep deterministic policy gradient algorithm.
[0066] In this embodiment, the decision model is trained based on the dual-delay deep deterministic policy gradient (TD3) algorithm, and its input state space is:
[0067] S=(SCR,ΔP,ΔU);
[0068] Where SCR is the short-circuit ratio at the common coupling point, ΔP is the deviation between the output power and the reference power, and ΔU is the deviation between the voltage at the common coupling point and the reference voltage.
[0069] The formula for calculating the short-circuit ratio (SCR) is:
[0070] ;
[0071] Among them, S SC For short-circuit capacity, S N Z is the rated capacity of the generator. This is the per-unit value of the equivalent impedance of the power grid.
[0072] Short-circuit capacity S SC The voltage U at the common coupling point can be used as a reference. o(abc) With short-circuit current I abc The calculation yields the following result:
[0073] .
[0074] The decision model is trained using a deep reinforcement learning framework that involves trial and error between the agent and the environment: the agent takes the operating state information as the observed state and the adaptive coefficient as the output action. By trying different actions and receiving the immediate reward calculated by the preset reward function, the agent learns to make the optimal decision in a complex power grid environment and finally learns the strategy that maps the optimal adaptive coefficient according to the real-time state.
[0075] As one implementation method, in this embodiment of the application, the decision model includes: an online Actor network and two online Critic networks, as well as a corresponding Actor target network and two Critic target networks.
[0076] The decision model employs the double-delay deep deterministic policy gradient (TD3) algorithm, which is based on the Actor-Critic framework and includes a... Parameterized online Actor network and two separate entities... θ 1. θ 2. A parameterized online Critic network. In addition, a corresponding [system / mechanism] is also provided. 'Parameterized Actor target network and two by , The parameterized Critic target network, with its dual Critic structure, aims to alleviate the overestimation problem of the action-value function and improve the stability and robustness of policy learning.
[0077] As one implementation method, in this embodiment of the application, step S21 includes:
[0078] S211. Initialize the parameters of the Actor online network, two Critic online networks, the Actor target network, and the two Critic target networks, and create an experience replay pool;
[0079] In this embodiment, the Actor online network parameters are first initialized. and two Critic online network parameters θ 1. θ 2. Simultaneously initialize the corresponding target network parameters: , ;
[0080] Next, set the maximum number of training rounds N, the maximum number of steps per round T, and the policy update frequency d (i.e., the Actor online network updates once for every d updates of the Critic online network).
[0081] Then create an experience replay pool R to store experience tuples, i.e., experience data (s, a , a' (r, s').
[0082] S212. At each time step, the obtained current training state information is input into the Actor online network to obtain the current state action. Noise is added to the current state action to obtain the noisy action. The reward value is obtained according to the current training state information and the preset reward function. The next training state information is obtained. The current training state information, the current state action, the noisy action, the reward value and the next training state information are stored as experience data in the experience replay pool.
[0083] In this embodiment, firstly, at each time step t, the agent observes the current training state information s, inputs the current training state information s into the Actor online network, and obtains the current state action output by the Actor online network. ;
[0084] Then add Gaussian noise to the motion. ε To obtain the current state action with noise. a' = a + ε ,in, Noise standard deviation It can decay during the training process and execute the noisy current state action;
[0085] Then, based on the current training state information and the preset reward function, the reward value r is obtained, and the next training state information is observed. s' ;
[0086] Then, combine the current training state information s, the current state action a, and the noisy current state action. a' The reward value r and the next training state information s' are used as empirical data (s, a , a' ,r,s') are stored in the experience replay pool R.
[0087] S213. Sample experience data from the experience replay pool. Based on the sampled data, obtain the target action value of the Actor target network and two Critic target networks. With the goal of minimizing the error between the action value output by the two Critic online networks based on the sampled data and the target action value, update the parameters of the two Critic online networks.
[0088] In this embodiment, M samples (s) are first randomly sampled from the experience replay pool R. a , a' ,r,s');
[0089] For each sample, the next training state information s' is input into the Actor target network to obtain the next state action output by the Actor target network. Add Gaussian noise to the next state action. ε To obtain the next state action with noise. ;
[0090] Then combine the next training state information s' with the noisy next state action. The input is fed into two Critic target networks, yielding two action values output by each network. The target action value is then calculated based on the reward value and the two action values, using the following formula:
[0091] ;
[0092] Where y represents the target action value; r represents the reward value; and γ is the discount factor. The action value output by the Critic target network; To obtain the minimum value of the two action values output by the two Critic target networks;
[0093] The parameters of the Critic network are then updated by minimizing the mean square error between the action values output by the two Critic online networks and the target action value. The specific calculation formula is as follows:
[0094] ;
[0095] in, represents the parameters of two online Critic networks; argmin is an abbreviation for argument of the minimum, which is the input variable that minimizes the objective function; M is the number of samples sampled from the experience replay pool; y is the target action value. The action value output by Critic online network.
[0096] S214. Every preset number of update steps, the Actor online network is updated using a deterministic strategy gradient update, and the parameters of the Actor target network and the two Critic target networks are updated using a soft update method.
[0097] In this embodiment, every preset update step d (i.e., every d updates of the Critic network parameters), the parameters of the Actor online network are updated using a deterministic policy gradient. The formula for calculating the policy gradient is:
[0098] ;
[0099] in, For expected returns; Indicates J to Differentiate; M is the number of samples taken from the empirical replay pool; The action value output by Critic's online network; It represents the current state action output by the Actor online network.
[0100] Simultaneously, the parameters of the Actor target network and the two Critic target networks are soft-updated, as follows:
[0101] ;
[0102] Where τ is the soft update factor; These are the parameters for two Critic online networks; Here are the parameters for the two Critic target networks. For the parameters of the Actor online network, These are the parameters of the Actor target network.
[0103] S215. Repeat the above steps until the preset training termination condition is met.
[0104] In this embodiment, steps S212 to S214 are repeated until the preset maximum number of training rounds N is reached, or other preset performance requirements (such as reward convergence) are met. After training is completed, the trained decision model is obtained.
[0105] In one implementation, the current training state information in this application includes the deviation between the training output power and the reference power, and the deviation between the training common coupling point voltage and the reference voltage; the reward value is negatively correlated with the deviation between the training output power and the reference power, and the deviation between the training common coupling point voltage and the reference voltage.
[0106] In this embodiment, the reward function is designed to negatively correlate the instantaneous reward value with the absolute values of power deviation and voltage deviation, providing a clear and stable optimization guide for the training of the decision model. Its core principle is that when the actual output power of the grid-connected converter or the voltage at the point of common coupling deviates from its reference value, system performance degrades. At this time, the reward function returns a lower reward value (or a higher penalty value), thus informing the agent that the current control strategy (i.e., the adaptive coefficient of the output) is inadequate, prompting it to adjust the strategy under similar future conditions. Conversely, when the deviation decreases, the reward value increases, thereby incentivizing and reinforcing the current strategy. This mechanism guides the agent to spontaneously explore and ultimately learn the optimal control strategy that simultaneously minimizes power tracking error and voltage support deviation during long-term learning.
[0107] The specific expression for the reward function is:
[0108] ;
[0109] in, r This is the reward value; λ 1 represents the weighting coefficient for the deviation between the training output power and the reference power. λ 2 represents the weighting coefficients for the deviation between the common coupling point voltage and the reference voltage during training. λ 1 and λ Both 2 are positive numbers and are adjusted during training; ΔP is the deviation between the training output power and the reference power; ΔU is the deviation between the training common coupling point voltage and the reference voltage.
[0110] In one implementation, in this application embodiment, at least one adaptive coefficient includes: a network-following adaptive coefficient and a network-building adaptive coefficient, the sum of which is 1.
[0111] In this embodiment, the constraint that the sum of the following network adaptive coefficient and the network construction adaptive coefficient is always 1 essentially constructs a normalized and complete weight allocation mechanism. This mechanism mathematically establishes a clear, continuous, and non-redundant decision space, ensuring that the total weight resources allocated to the following network control and the network construction control are fixed at 100%. This design brings multiple technical advantages: First, it makes the representation of the control mode extremely clear and intuitive. When the coefficient is (1,0), it corresponds to pure following network control; when it is (0,1), it corresponds to pure network construction control; and any point between the two corresponds to a mixed state with different proportions, which facilitates the understanding and debugging of the strategy; Second, This fundamentally ensures the continuity of the control signal. The fusion process of the two control signals will not produce unnecessary gain changes or distortions due to fluctuations in the sum of weights. This makes the characteristic changes of the final control command entirely due to the ebb and flow between the two modes, resulting in a natural and smooth transition. Finally, it provides a well-structured and clearly defined output action space for the decision model (or other optimization algorithms). The model only needs to focus on outputting a continuously varying adaptive coefficient in the [0,1] interval to unambiguously determine the complete control structure. This greatly simplifies the complexity of strategy learning and optimization and ensures that the control system has a definite physical meaning under any decision.
[0112] As one implementation method, in this embodiment of the application, the first control signal from the network-following control link and the second control signal from the network-building control link are fused according to at least one adaptive coefficient, including: weighting and summing the first control signal from the network-following control link and the second control signal from the network-building control link according to the weights determined by the network-following adaptive coefficient and the network-building adaptive coefficient.
[0113] In this embodiment, the phase signal specifically targeting the virtual internal potential of the grid-connected converter is fused. The two control signals are weighted and summed using weights determined by adaptive coefficients. This is the core mathematical operation for achieving control command fusion. This method has significant advantages in both engineering and physical aspects: In engineering implementation, the weighted summation calculation is simple and deterministic, involving only one multiplication and one addition operation, resulting in a low computational burden. It is easy to implement efficiently and reliably in digital signal processors or microcontrollers, meeting the speed requirements of real-time control systems. In terms of physics and control, this linear fusion method allows for a clear and intuitive quantitative expression of the contribution ratio of grid-connected control signals (such as the phase deviation of the phase-locked loop output) and grid-connected control signals (such as the phase deviation of the power synchronization control output) in the final virtual internal potential phase. The weighting coefficients directly and linearly determine the characteristic bias of the output phase.
[0114] More importantly, by continuously changing the coefficients, this method can achieve a smooth and stepless transition between all intermediate states from pure follow-net control to pure network control, completely avoiding output jumps, torque pulsations or system oscillations that may be caused by sudden changes or switching of control modes. This ensures the continuity and stability of the control process and the dynamic response of the system. This weight-based linear fusion provides a solid foundation for building a cooperative control architecture that has both clear physical meaning and is conducive to stable implementation.
[0115] Specifically, the phase θ of the virtual internal potential is determined by the following equation:
[0116] ;
[0117] Where θ0 is the power grid reference phase angle, r PLL r is the adaptive coefficient for the network; PSC For the adaptive coefficients of the network construction; Δθ PLL The phase angle deviation generated by the phase-locked loop (PLL) element; Δθ PSC The phase angle deviation generated by the power synchronization control (PSC) loop is obtained by integrating the angular velocity deviation of the PSC active control loop output.
[0118] As one implementation, the phase-locked loop (PLL) circuit tracks the voltage U at the point of common coupling. o(abc) The phase, output angular velocity deviation Δω PLL and the corresponding phase angle deviation Δθ PLL =∫Δω PLL The control law of the phase-locked loop is expressed as:
[0119] ;
[0120] in, u oq The q-axis component of the voltage at the common coupling point; The proportional gain of the phase-locked loop PI controller; The integral coefficients of the phase-locked loop PI controller are denoted as .
[0121] In one implementation, the power synchronization control (PSC) circuit includes an active power control loop and a reactive power control loop to simulate the external characteristics of a synchronous generator. The active power control loop is defined by an active power reference value P. ref With actual output active power P e The difference-driven mechanism is achieved by introducing virtual inertia J, damping coefficient D, and active-frequency droop coefficient k. p Generate angular velocity deviation Δω PSC This process simulates the rotor motion characteristics of a synchronous generator, and its dynamic equation can be described as follows:
[0122] JdΔωPSC / dt=(P ref -P e ) / ω0-(D+k p / ω0)Δω PSC ;
[0123] Where ω0 is the rated angular velocity of the power grid.
[0124] The angular velocity deviation Δω PSC The phase angle deviation Δθ of the network control element is obtained after integration. PSC The specific formula is: Δθ PSC =∫Δω PSC dt.
[0125] In one implementation method, the reactive power control loop of the power synchronization control stage achieves voltage support by adjusting the virtual electromotive force E. The control law of power synchronization control is expressed as:
[0126] E=E0+k i ∫[(Q ref Q e )+k q (U n U o(abc) )]dt;
[0127] Where E0 is the no-load electromotive force; Q ref and Q e These are the reference and actual values for reactive power, respectively; U n and U o(abc) These are the effective values of the rated voltage and the voltage at the point of common coupling, respectively; k i and k q This is the control factor.
[0128] As one implementation method, to simulate the stator impedance of a synchronous generator, a virtual impedance element is used to adjust the virtual electromotive force E in order to generate a reference current i. abc :
[0129] ;
[0130] Among them, L f This is the virtual inductance value; u o(abc) R is the voltage at the common coupling point; f This is a virtual resistance value.
[0131] As one implementation, the system further includes a current inner-loop control loop for fast and accurate tracking of the reference current. This loop is implemented in a dq rotating coordinate system and generates the converter's voltage reference command through a PI regulator, grid voltage feedforward, and cross-decoupling compensation. Its control law is expressed in the frequency domain as follows:
[0132] ;
[0133] in, u d , u q The d-axis and q-axis reference voltage commands output by the controller; k p1 and k p2 These are the proportional parameters of the current loop PI controller; k i1 and k i2 The integral coefficient of the inner current loop PI controller; s For the Laplace operator; i * d , i * q The d-axis and q-axis reference currents given to the outer loop; i d , i q These are the actual measured d-axis and q-axis currents; u gd , u gq The d-axis and q-axis components of the measured grid voltage are used as feedforward signals to improve dynamic response; ω0 is the grid's rated angular velocity. Using the rated angular velocity ω0 instead of the real-time angular velocity in the decoupling term helps maintain the stability and robustness of the control system when the grid frequency fluctuates slightly; L0=(L s +L arm / 2) is the total equivalent inductance of the system from the controller reference point to the grid point, L s L is the grid-side inductance, i.e., the equivalent inductance of the line between the AC output of the grid-connected converter and the grid connection point. arm For each bridge arm of the grid-connected converter, there is the bridge arm inductance.
[0134] As one implementation method, the grid-connected converter is a modular multilevel converter, and the method further includes:
[0135] S31. Detect the circulating current component of each phase arm of the modular multilevel converter;
[0136] In this embodiment, the modular multilevel converter (MMC) can specifically be a three-phase six-arm structure. The circulating current component specifically refers to the unwanted circulating current flowing between the three-phase arms of the MMC. Due to the difficulty in maintaining absolute balance of the capacitor voltages of each phase submodule and the imperfect symmetry of the circuit parameters, circulating currents mainly composed of DC and second harmonic components will be generated between the upper and lower arms. Detecting this circulating current component is a prerequisite for effective suppression, and it is usually obtained by measuring the current of each phase's upper and lower arms and separating it through specific coordinate transformations or calculations.
[0137] S32. Input the circulating current component to the proportional-integral resonant controller, wherein the resonant frequency of the proportional-integral resonant controller is set to twice the frequency of the grid.
[0138] In this embodiment, a proportional-integral (PIR) resonant controller is used for closed-loop regulation of the circulating current component. The controller's resonant frequency is precisely set to twice the power grid frequency, aiming to provide extremely high open-loop gain for the second harmonic component, which has the highest amplitude and is the most harmful in the circulating current. This achieves precise and efficient suppression of the second harmonic circulating current. At the same time, the proportional-integral (PI) part of the controller is used to eliminate the DC bias component in the circulating current. This composite controller structure achieves targeted compensation for key components of the circulating current.
[0139] S33. Adjust the reference value of the bridge arm voltage according to the output of the proportional-integral resonant controller to suppress the circulating current component of each phase bridge arm.
[0140] In this embodiment, the output of the PIR controller is used as a compensation term and superimposed on the original voltage reference value of each phase arm. This compensation term is equivalent to injecting a cancellation voltage command into the control loop that is opposite in phase and has the same amplitude as the detected circulating current component. When this adjusted arm voltage reference value is modulated to generate a switching signal and drive the submodule to operate, a voltage that can actively cancel the original circulating current will be generated in the arm, thereby forming closed-loop suppression. This method cancels the circulating current path electrically through control means, effectively attenuating the circulating current amplitude, reducing additional losses and current harmonic distortion in the arm, improving the overall system operating efficiency and waveform quality, and without increasing any hardware costs.
[0141] Preferably, this control method also combines the nearest level approximation modulation (NLM) strategy with the capacitor voltage sorting algorithm to further suppress circulating current and achieve equalization of the capacitor voltage of the submodule from the modulation level.
[0142] like Figure 2 As shown in the embodiment of this application, a control device for a grid-connected converter is also provided. The control device 100 includes:
[0143] The acquisition module 110 is used to acquire real-time power grid operating status information;
[0144] The decision module 120 is used to input the operating status information into the trained decision model, and the decision model dynamically outputs at least one adaptive coefficient for coordinating the network-following control mode and the network-building control mode based on the operating status information.
[0145] The fusion module 130 is used to fuse the first control signal from the network control link and the second control signal from the network construction control link according to at least one adaptive coefficient to generate control commands.
[0146] The control module 140 is used to control the operation of the grid-connected converter according to control commands.
[0147] like Figure 3 and Figure 4 As shown, this application embodiment also provides an electric power router, including the control device 100, grid-connected converter 200, converter 300 and hybrid energy storage module 400 as described above; the converter 300 is connected between the grid-connected converter 200 and the hybrid energy storage module 400 to realize power transmission and electrical isolation; the hybrid energy storage module 400 includes a lithium battery C1 and a supercapacitor C2 connected in parallel.
[0148] In this embodiment, the control device 100 is signal-connected to the grid-connected converter 200, and is used to collect voltage and current operating status information of the grid-connected converter 200 and output control commands to it. The grid-connected converter 200 is preferably a modular multilevel converter (MMC) with a three-phase six-arm structure. Each arm 210 includes N cascaded power submodules 211, where N is a positive integer greater than 1.
[0149] Specifically, such as Figure 4 As shown, taking phase A as an example, each bridge arm has a bridge arm inductor (Larm) connected in series at its end. Specifically, the last power submodule of the upper bridge arm of phase A is connected to the first terminal of the first inductor (Larm). The second terminal of the first inductor (Larm) is simultaneously connected to the first terminal of the second inductor (Larm) of the lower bridge arm of phase A and the first terminal of the line resistor R1 of phase A. The second terminal of the second inductor (Larm) is connected to the first power submodule of the lower bridge arm of phase A. The second terminal of the line resistor R1 of phase A is connected to the first terminal of the line inductor Ls of phase A, and the second terminal of the line inductor Ls of phase A is connected to phase A of the AC subgrid. The connection method for phases B and C is the same as that for phase A. The AC side port of the grid-connected converter 200 is the point of common coupling (PCC).
[0150] The converter 300 includes multiple conversion sub-modules 310, and the hybrid energy storage module 400 includes multiple hybrid energy storage sub-modules 410. The conversion sub-modules 310 are connected between the power sub-module 211 and the hybrid energy storage sub-modules 410. The hybrid energy storage sub-modules 410 include a lithium battery C1 and a supercapacitor C2 connected in parallel.
[0151] The converter 300 is preferably a three-bridge active power converter (TAB), which employs a high-frequency isolation transformer and is equipped with an anti-saturation core design. This not only enables bidirectional power transmission and electrical isolation between the grid-connected converter 200 and the hybrid energy storage module 400, but also effectively suppresses core saturation, avoiding the risk of inrush current and system instability caused by it. In the hybrid energy storage module 400, the lithium battery C1 has high energy density characteristics, suitable for energy throughput and storage; the supercapacitor C2 has high power density and fast charge and discharge characteristics, suitable for instantaneous response to frequent large power fluctuations in the microgrid. The two work together to improve the dynamic performance and operational flexibility of the power router.
[0152] Furthermore, the power submodule 211 can be a full-bridge submodule (FBSM) or a half-bridge submodule (HBSM). The full-bridge submodule has DC fault current blocking capability, which helps to improve system reliability; the half-bridge submodule has a simple structure and lower cost and operating losses. In actual design, it can be flexibly selected or combined according to different requirements such as cost, efficiency, and fault ride-through capability. For example, a full-bridge submodule can be used in the first power unit of each bridge arm to enhance reliability, while half-bridge submodules can be used in the rest to optimize economy.
[0153] Specifically, one circuit implementation of the full-bridge submodule FBSM is as follows: Figure 5 As shown, it includes four controllable switching devices (such as IGBTs or MOSFETs) and a DC support capacitor. By controlling the on and off states of the four switching devices, it can output positive, negative, or zero levels and has DC fault current blocking capability. One circuit implementation of the half-bridge submodule HBSM is shown below. Figure 6 As shown, it includes two controllable switching devices and a DC support capacitor, and its structure is simpler, making it the most commonly used unit for constructing multi-level voltages. Figure 5 and Figure 6 The circuit shown is merely an example. In other embodiments of this application, the power submodule 211 may also employ other circuit topologies that achieve the same function.
[0154] In the control of the grid-connected converter 200, the virtual inductance Lf and virtual resistance Rf in the aforementioned virtual impedance module are used to simulate the stator impedance of the synchronous generator; while the total equivalent inductance L0 in the current inner loop module is composed of the line inductance Ls and the bridge arm inductance Larm, i.e., L0 = (Ls + Larm / 2). The line resistance R1 serves as the equivalent resistance between the PCC and the AC subgrid, participating in the power transmission and dynamic response calculations of the system.
[0155] The above description is merely an embodiment of the present invention. It should be noted that those skilled in the art can make improvements without departing from the inventive concept of the present invention, but these improvements all fall within the protection scope of the present invention.
Claims
1. A control method for a grid-connected converter, characterized in that, The method includes: Real-time acquisition of power grid operating status information; The operational status information is input into a trained decision model, which then dynamically outputs at least one adaptive coefficient to coordinate the network-following control mode and the network-building control mode based on the operational status information. Based on the at least one adaptive coefficient, the first control signal from the network control link and the second control signal from the network construction control link are fused to generate a control command. The operation of the grid-connected converter is controlled according to the control command. The method further includes: The decision model is trained using a dual-delay deep deterministic policy gradient algorithm. The decision-making model includes: an online Actor network and two online Critic networks, as well as a corresponding Actor target network and two Critic target networks; The decision model trained using the dual-delay deep deterministic policy gradient algorithm includes: Initialize the parameters of the Actor online network, the two Critic online networks, the Actor target network, and the two Critic target networks, and create an experience replay pool; At each time step, the acquired current training state information is input into the Actor online network to obtain the current state action. Noise is added to the current state action to obtain a noisy action. Based on the current training state information and the preset reward function, the reward value is obtained. The next training state information is obtained, and the current training state information, the current state action, the noisy action, the reward value, and the next training state information are stored as experience data in the experience replay pool. Experience data is sampled from the experience replay pool. Based on the sampled data, the Actor target network and the two Critic target networks are used to obtain the target action value. The parameters of the two Critic online networks are updated with the goal of minimizing the error between the action value output by the two Critic online networks based on the sampled data and the target action value. Every preset number of update steps, the Actor online network is updated using a deterministic strategy gradient update, and the parameters of the Actor target network and the two Critic target networks are updated using a soft update method. Repeat the above steps until the preset training termination condition is met; The at least one adaptive coefficient includes: a network-following adaptive coefficient and a network-building adaptive coefficient, wherein the sum of the network-following adaptive coefficient and the network-building adaptive coefficient is 1.
2. The method according to claim 1, characterized in that, The operating status information includes: the short-circuit ratio characterizing the grid strength, the deviation between the output power and the reference power of the grid-connected converter, and the deviation between the voltage at the common coupling point and the reference voltage.
3. The method according to claim 1, characterized in that, The current training state information includes the deviation between the training output power and the reference power, and the deviation between the training common coupling point voltage and the reference voltage. The reward value is negatively correlated with the deviation between the training output power and the reference power, and with the deviation between the training common coupling point voltage and the reference voltage.
4. The method according to claim 1, characterized in that, The step of fusing the first control signal from the network-following control link and the second control signal from the network-building control link according to the at least one adaptive coefficient includes: The first control signal from the network tracking control link and the second control signal from the network construction control link are weighted and summed according to the weights determined by the network tracking adaptive coefficient and the network construction adaptive coefficient.
5. A control device for a grid-connected converter, used to implement the method according to any one of claims 1 to 4, characterized in that, The control device includes: The acquisition module is used to acquire real-time operating status information of the power grid; The decision module is used to input the operating status information into the trained decision model, and the decision model dynamically outputs at least one adaptive coefficient for coordinating the network-following control mode and the network-building control mode based on the operating status information. The fusion module is used to fuse the first control signal from the network control link and the second control signal from the network construction control link according to the at least one adaptive coefficient, and generate control commands. The control module is used to control the operation of the grid-connected converter according to the control commands.
6. A power router, characterized in that, Includes the control device, grid-connected converter, converter, and hybrid energy storage module as described in claim 5; The converter is connected between the grid-connected converter and the hybrid energy storage module to achieve power transmission and electrical isolation. The hybrid energy storage module includes a lithium battery and a supercapacitor connected in parallel.