Reinforcement Learning-Based Adaptive Stability Controller for Power Electronic Transformer Feed Networks
By integrating a reinforcement learning-based adaptive stability controller on the low-voltage side of the power electronic transformer, the wide-frequency domain instability problem of the power electronic transformer feeding the grid under a high proportion of renewable energy access is solved, and the adaptive stability and robustness of the system are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-30
- Publication Date
- 2026-04-03
AI Technical Summary
With a high proportion of renewable energy and power electronic equipment connected to the grid, traditional feeder grids face broadband oscillations and instability problems. In particular, when the number of converters connected to the low-voltage side of power electronic transformers increases, the robustness of stability control strategies and parameters is difficult to cope with the stability challenges of system parameter changes and multi-timescale coupling.
An adaptive stability controller based on reinforcement learning is integrated into the low-voltage side converter of the power electronic transformer. Through state observation, reinforcement learning algorithm and action strategy selection module, the matching degree between the complex frequency domain impedance of the power electronic transformer and the equivalent impedance of the power grid is optimized, and the stability controller parameters are adaptively adjusted to improve system stability.
It improves the stability and robustness of power electronic transformer power grid systems, and can automatically adjust the control strategy when system parameters change to maintain system stability. It is suitable for various scenarios and does not require changes to the control structure of new energy power generation converters.
Smart Images

Figure CN115296341B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power system control technology, specifically relating to an adaptive stability controller for power electronic transformer feeder networks based on reinforcement learning. Background Technology
[0002] Currently, with the high penetration rate of distributed renewable energy into the grid, traditional feeder networks face numerous challenges, especially the broadband oscillation problem caused by the high proportion of renewable energy and high proportion of power electronic equipment (referred to as "dual high"). The harmonics generated by instability can propagate to other grids through traditional transformers and transmission lines, affecting the operation of the entire power system. Power electronic transformers can control output voltage, improve system power quality, and enhance the reliability and carrying capacity of the grid. Introducing power electronic transformers into traditional grids to construct new feeder network systems is an effective measure to solve the oscillation problem.
[0003] However, with a high proportion of distributed renewable energy, the power electronic transformer feeder grid still faces the risk of instability, especially when the number of converters connected to the low-voltage side of the power electronic transformer increases and the system component parameters change. Modifying the power electronic transformer control strategy can effectively improve the stability of the feeder grid system, but the selection of the stability control strategy and the adjustment of parameters have a great impact on stability. Considering the uncertainty of the power and component parameters of the power electronic transformer feeder grid system, as well as the characteristics of the system's wideband oscillation, the robustness of fixed stability control strategies and control parameters is greatly challenged in the face of various stability problems caused by the multi-timescale coupling of the "high-voltage and high-renewable energy" power system. Summary of the Invention
[0004] In view of the shortcomings of the prior art, the purpose of this invention is to provide an adaptive stability controller for power electronic transformer feeder networks based on reinforcement learning, so as to solve the problems mentioned in the background art.
[0005] The objective of this invention can be achieved through the following technical solutions:
[0006] The power electronic transformer feeder system based on reinforcement learning is used to solve the wide frequency domain instability phenomenon in the small signal stability range of the power electronic transformer feeder system, including filter resonance and controller multi-time scale coupling instability. It optimizes the matching degree between the complex frequency domain impedance of the power electronic transformer and the equivalent impedance of the grid, thereby improving the stability of the power electronic transformer feeder system. The adaptive stability controller is integrated into the low-voltage side converter of the power electronic transformer. The adaptive stability controller includes a stability controller and an adaptive controller. The adaptive controller adjusts the parameters of the stability controller.
[0007] The adaptive controller is trained using a reinforcement learning algorithm. The training structure includes a state observation module, a reinforcement learning algorithm module, and an action policy selection module. The state observation module, reinforcement learning algorithm module, and action policy selection module together form an agent that interacts with the power electronic transformer feed grid for training.
[0008] Preferably, the stability controller is based on the principle of complex frequency domain impedance reshaping, and the stability controller reshapes the complex frequency domain impedance of the low-voltage side converter of the power electronic transformer.
[0009] Preferably, the stability controller control strategy includes passive control, active damping, virtual impedance, digital filter, and lead-lag compensation control strategy.
[0010] Preferably, the state observation module extracts the state information of the power electronic transformer feeder system and assigns action rewards to the intelligent agent based on the state information;
[0011] The reinforcement learning algorithm module is trained based on the state information of the power electronic transformer feeder system, the actions taken by the agent, and the rewards obtained after taking the actions, to optimize the action strategy.
[0012] The action strategy selection module selects an action strategy based on the state information of the power electronic transformer feeder system, and then changes the parameters of the stability controller to stabilize the power electronic transformer feeder system.
[0013] Preferably, the state information extracted by the state observation module includes the output voltage, output current, filter inductor current, and power of the low-voltage side converter of the power electronic transformer, including their effective value, average value, and harmonic components.
[0014] Preferably, the algorithms used in the reinforcement learning algorithm module include Q-learning, deep reinforcement learning, DDPG, and Actor-Critic algorithms.
[0015] A reinforcement learning method for an adaptive stability controller of a power electronic transformer feeder grid is proposed, the reinforcement learning method being as follows:
[0016] First, initialize the quantities related to the reinforcement learning algorithm, including the learning rate, discount factor, neural network parameters, and training convergence flag;
[0017] Then, training is performed for each scenario. In each scenario, the parameters of the power electronic transformer feeder system are initialized, and the state observation module acquires the current state information S. t The action strategy selection module selects action a based on the current status information. tThe parameters of the stability control strategy are updated. After the parameters change, the system state will change accordingly. The current state S of the system is observed using the state observation module. t+1 And obtain the state S t Take action a t Reward R t ;
[0018] Next, the reinforcement learning algorithm updates its action selection strategy based on the above information. After updating the action strategy, it determines whether the scene can end by checking whether the state has reached the target state or whether the number of action steps has reached the set maximum value. If the scene has not reached the end flag, the updated action selection strategy module then selects the action based on the state S. t+1 Choose action a t+1 Repeat the previous process until the training for this scene is completed. After the training for each scene is completed, determine the convergence state of the agent and whether the training can be terminated.
[0019] Finally, if the number of training sessions reaches the set value or the agent reaches the convergence condition, the training can be terminated. Otherwise, the power electronic transformer feed grid parameters are reinitialized, and the training process for each session is repeated until the termination condition is met. After the training ends, the final optimized action strategy is obtained.
[0020] Preferably, the feature is that after training is completed, the state observation module and the action selection strategy module are ported to the adaptive controller. When the power electronic transformer feed grid system becomes unstable, the state observation module identifies the state of the power electronic transformer feed grid system, and the action selection module adjusts the stability controller parameters according to the current system state to stabilize the power electronic transformer feed grid system.
[0021] The beneficial effects of this invention are:
[0022] 1. The adaptive stability controller proposed in this invention can not only improve the stability of the power grid system, but also, after training, even if the parameters of the power electronic transformer power grid system change and the system instability phenomenon is different from before, the adaptive stability controller can automatically adjust the parameters of the stability control strategy to improve the system stability and the robustness of the stability control strategy. Since the adaptive stability controller is set in the voltage controller of the low-voltage side converter of the power electronic transformer, there is no need to change the control structure of the many new energy power generation converters connected to it, and it can be flexibly applied in a variety of occasions. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is an application structure diagram of the present invention;
[0025] Figure 2 This is a control block diagram of the low-voltage side converter of the power electronic transformer in this invention;
[0026] Figure 3 This is a block diagram of the adaptive stabilization controller training in this invention;
[0027] Figure 4 This is a flowchart of the adaptive stabilization controller training process in this invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] This invention addresses wide-frequency domain instability (including but not limited to filter resonance and controller multi-time-scale coupling instability) within the small-signal stability range of power electronic transformer feeder systems. It adds a stability controller to the low-voltage side converter control of the power electronic transformer to improve the stability of the feeder system. A reinforcement learning algorithm is used to achieve adaptive adjustment of the stability controller parameters. After training, the adaptive controller can automatically adjust the stability controller parameters based on external input state information to maintain the stability of the feeder system under different operating conditions and scenarios.
[0030] A controller other than stability control typically includes voltage control and current control components. Voltage and phase angle reference values can be set directly or generated by a power-controlled phase-locked loop (PLL) mechanism. The voltage reference value is compared with the actual voltage value, then passed through the voltage controller and finally the stability controller to obtain the current reference value. The current reference value is compared with the actual value, then passed through the current controller to output a modulation voltage, which regulates the output voltage of the power electronic transformer.
[0031] The power electronic transformer feeder network adaptive stability controller based on reinforcement learning adds an adaptive stability controller to the low-voltage side converter control of the power electronic transformer to reshape the equivalent complex frequency domain impedance of the converter, optimize its matching degree with the low-voltage grid impedance, and improve the stability of the feeder network system. The adaptive stability controller is integrated into the voltage controller of the low-voltage side converter of the power electronic transformer. In order to improve the robustness of the stability control strategy, the adaptive stability controller parameter adjustment is implemented using a reinforcement learning algorithm. After training, the adaptive controller 9 can automatically adjust the parameters of the stability controller 7 according to the changes in external input state information to maintain the stability of the feeder network system under different operating conditions and scenarios.
[0032] Adaptive stability control is achieved through reinforcement learning training. The training framework includes a state observation module, a reinforcement learning algorithm module, and an action policy selection module. These three modules together form a reinforcement learning agent, which interacts with the power electronic transformer feed grid during training. After training, the agent ultimately achieves the function of selecting the optimal action based on the system state. The state observation module extracts the current system state information and assigns action rewards based on this information. The reinforcement learning algorithm module is trained based on the system state, the actions taken by the agent, and the rewards obtained after taking those actions, updating the action policy selection module. The action policy selection module then selects an appropriate action based on the current state, changing the stability controller parameters.
[0033] Figure 1 This is an application structure diagram of the present invention, including a medium-voltage power grid 1 on the medium-voltage side of the power electronic transformer, a power electronic transformer 2, a filter 3, a low-voltage AC voltage 4, and a new energy power generation converter group 5 connected to the low-voltage side of the power electronic transformer;
[0034] The power electronic transformer 2 is composed of a cascaded AC / DC rectifier 2-1, an isolated DC / DC converter 2-2, and a DC / AC inverter 2-3. The isolated DC / DC converter 2-2 decouples the medium-voltage and low-voltage sides of the power electronic transformer. A capacitor connected to the DC side of the power electronic transformer reduces voltage ripple. An LC filter 3 is connected to the low-voltage side of the power electronic transformer to filter out harmonic components.
[0035] The new energy power generation converter group 5 is connected after the filter 3. Its control can adopt current control, power control, voltage control and droop control. The low-voltage side converter of the power electronic transformer coupled with the new energy power generation converter group 5 can be centrally controlled by adopting a stable control strategy. It is not necessary to add a stable control module to all new energy converters. The functions related to artificial intelligence algorithms involved in the new energy converter can also be centrally implemented by the low-voltage side of the power electronic transformer in most cases.
[0036] Control strategies for low-voltage side converters of power electronic transformers, such as Figure 2 As shown, this control is a classic dual-loop control strategy implemented in a rotating coordinate system (dq axis), including voltage controller 6 (VC), current controller 8 (CC), etc.
[0037] Without a stable control strategy, the actual voltage value obtained by coordinate transformation is compared with the reference value and input to the voltage controller 6. The current reference value is directly output and compared with the actual dq axis current value and input to the current controller 8. The output quantity is adjusted by PWM to regulate the output voltage. The voltage controller 6 and the current controller 8 can adopt PI control to ensure zero steady-state error tracking. The adaptive stable controller includes a stable controller 7 and an adaptive controller 9. The stable controller based on the impedance reshaping principle can adopt control strategies such as active damping, virtual impedance, digital filter, and lead-lag compensation.
[0038] Since wide-frequency domain stability issues are frequently encountered in "high-voltage and high-efficiency" power systems, a stability controller 7 is connected in series after the voltage controller 6 to reshape the complex frequency domain impedance of the converter and improve the stability of the power grid system. The parameters of the stability controller 7 are adjusted by the adaptive controller 9, and its input is a certain state variable of the system. Figure 2 The selected current is the filter inductor current, and the output is the parameter of the stabilizing controller. Through the adaptive controller 9, the system can adaptively adjust the parameters of the stabilizing controller 7 according to the system state, thereby improving the robustness of the stabilizing control strategy. The above shows and describes an application example. In practical applications, the stabilizing controller can also be integrated into the original control system in parallel or in a hybrid manner, and the connection location is not limited to the control forward channel. These changes and improvements all fall within the scope of the present invention as claimed.
[0039] The control of the low-voltage side converter of the power electronic transformer can also be realized in a two-phase stationary coordinate system (αβ axis). Correspondingly, in order to achieve zero steady-state error control, the voltage controller 6 and the current controller 8 adopt PR control.
[0040] Figure 3 This is a training block diagram of the adaptive stability controller of the present invention. The training adopts a reinforcement learning algorithm, and the application environment is a power electronic transformer feeder system that uses an adaptive stability controller for stability control.
[0041] The training framework includes a state observation module, a reinforcement learning algorithm module, and an action policy selection module. These three modules together form a reinforcement learning agent that continuously interacts with the power electronic transformer feed grid for training, ensuring that the system can select the optimal action based on the state.
[0042] The status observation module extracts current system status information, including but not limited to harmonic order and harmonic amplitude of current, voltage or power waveforms, and assigns action rewards based on the current status information;
[0043] Reinforcement learning algorithms are trained based on the system state, the actions taken by the agent, and the rewards obtained after taking actions. They continuously update the policy for choosing actions. Training algorithms can employ Q-learning, deep reinforcement learning, DDPG, or Actor-Critic algorithms, etc.
[0044] The action selection strategy module will choose an appropriate action based on the current state and change the parameters of the stability controller.
[0045] Figure 4 The flowchart shows the training control method for an adaptive stability controller for power electronic transformer feeder networks based on reinforcement learning. The training control method is as follows:
[0046] First, initialize the relevant quantities for the reinforcement learning algorithm, including the learning rate, discount factor, neural network parameters, and training convergence flag. Then, train each scene. Considering that the system component parameters may change during actual operation, it may be necessary to randomly assign values to the component parameters according to the actual situation during training to increase the robustness of the adaptive control algorithm. In the training of each scene, initialize the power electronic transformer feeder system parameters, and the state observation module acquires the current state information S. t The action strategy selection module selects action a based on the current status information. t The parameters of the stability control strategy are updated. After the parameters change, the system state will change accordingly. The current state S of the system is observed using the state observation module. t+1 And obtain the state S t Take action a t Reward R t Then, the reinforcement learning algorithm updates its action selection policy based on this information. After updating the action policy, it determines whether the scene can end by checking if the state has reached the target state or if the number of action steps has reached the set maximum value. If the scene has not reached the end flag, the updated action selection policy module then selects the action based on the state S. t+1 Choose action a t+1 Repeat the previous process until the training of this scene ends. After the training of each scene ends, determine the convergence state of the agent and whether the training can be terminated. If the number of training scenes reaches the set value or the agent reaches the convergence condition, the training can be terminated. Otherwise, reinitialize the power electronic transformer feed grid parameters and repeat the training process of each scene until the termination condition is met, end the training, and obtain the final optimized action strategy.
[0047] After training, the state observation module and the action selection strategy module are ported to the adaptive controller. When the power electronic transformer feed grid system becomes unstable, the state observation module identifies the state of the power electronic transformer feed grid system, and the action selection module adjusts the stability controller parameters according to the current system state to stabilize the power electronic transformer feed grid system.
[0048] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0049] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.
Claims
1. A reinforcement learning-based adaptive stability controller for power electronic transformer feeder grids, used to address wide-frequency domain instability within the small-signal stability range of power electronic transformer feeder grid systems, including filter resonance and controller multi-time-scale coupling instability, optimizes the matching degree between the complex frequency domain impedance of the power electronic transformer and the equivalent impedance of the grid, and improves the stability of the power electronic transformer feeder grid system, characterized in that... The adaptive stability controller is integrated into the low-voltage side converter of the power electronic transformer. The adaptive stability controller includes a stability controller and an adaptive controller, and the adaptive controller adjusts the parameters of the stability controller. The adaptive controller is trained using a reinforcement learning algorithm. The training structure includes a state observation module, a reinforcement learning algorithm module, and an action policy selection module. The state observation module, reinforcement learning algorithm module, and action policy selection module together form an intelligent agent that interacts with the power electronic transformer feed grid for training. The stability controller is based on the principle of complex frequency domain impedance reshaping, and reshapes the complex frequency domain impedance of the low-voltage side converter of the power electronic transformer.
2. The power electronic transformer feeder adaptive stability controller based on reinforcement learning according to claim 1, characterized in that, The control strategies of the stability controller include passive control, active damping, virtual impedance, digital filtering, and lead-lag compensation control strategies.
3. The power electronic transformer feeder adaptive stability controller based on reinforcement learning according to claim 1, characterized in that, The state observation module extracts the state information of the power electronic transformer feeder system and assigns action rewards to the intelligent agent based on the state information. The reinforcement learning algorithm module is trained based on the state information of the power electronic transformer feeder system, the actions taken by the agent, and the rewards obtained after taking the actions, to optimize the action strategy. The action strategy selection module selects an action strategy based on the state information of the power electronic transformer feeder system, and then changes the parameters of the stability controller to stabilize the power electronic transformer feeder system.
4. The power electronic transformer feeder adaptive stability controller based on reinforcement learning as described in claim 3, characterized in that, The state information extracted by the state observation module includes the output voltage, output current, filter inductor current, and power of the low-voltage side converter of the power electronic transformer, including their effective value, average value, and harmonic components.
5. The power electronic transformer feeder adaptive stability controller based on reinforcement learning as described in claim 3, characterized in that, The algorithms used in the reinforcement learning algorithm module include Q-learning, deep reinforcement learning, DDPG, and Actor-Critic algorithms.
6. The reinforcement learning method for an adaptive stability controller for a power electronic transformer feeder grid based on reinforcement learning according to any one of claims 1-5, characterized in that, The reinforcement learning method is as follows: First, initialize the quantities related to the reinforcement learning algorithm, including the learning rate, discount factor, neural network parameters, and training convergence flag; Then, training is performed for each scenario. In each scenario, the parameters of the power electronic transformer feeder system are initialized, and the state observation module acquires the current state information. S t The action strategy selection module selects an action based on the current status information. a t The parameters of the stability control strategy are updated. After the parameters change, the system state will change accordingly. The current state of the system is observed using the state observation module. S t+1 And obtain the state. S t Take action a t Rewards R t ; Next, the reinforcement learning algorithm updates its action selection strategy based on the above information. After updating the action strategy, it determines whether the scene can end by checking whether the state has reached the target state or whether the number of action steps has reached the set maximum value. If the scene has not reached the end flag, the updated action selection strategy module then determines the end time based on the state. S t+1 Choose Action a t+1 Repeat the previous process until the training for this scene is completed. After the training for each scene is completed, determine the convergence state of the agent and whether the training can be terminated. Finally, if the number of training sessions reaches the set value or the agent reaches the convergence condition, the training can be terminated. Otherwise, the power electronic transformer feed grid parameters are reinitialized, and the training process for each session is repeated until the termination condition is met. After the training ends, the final optimized action strategy is obtained.
7. The reinforcement learning method for the adaptive stability controller of power electronic transformer feeder grid based on reinforcement learning according to claim 6, characterized in that, After training, the state observation module and the action selection strategy module are ported to the adaptive controller. When the power electronic transformer feed grid system becomes unstable, the state observation module identifies the state of the power electronic transformer feed grid system, and the action selection module adjusts the stability controller parameters according to the current system state to stabilize the power electronic transformer feed grid system.
Citation Information
Patent Citations
Adaptive adjustment inverter controller
CN110880774A