Power converter control method based on TD3 algorithm with multiple experience replay pools

By using a power converter control method based on the multi-experience replay pool TD3 algorithm, a dynamic perception layer, a strategy optimization layer and a lightweight execution layer are constructed, which solves the dynamic response and parameter setting problems of traditional PID control in complex power converter systems, and achieves efficient and stable control and lightweight deployment.

CN120110126BActive Publication Date: 2025-09-26GUANGDONG DIANBANG NEW ENERGY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510257750.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-09-26
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Traditional PID control algorithms in power converters have difficulty in meeting the dynamic response speed and multi-objective coordination requirements under complex topologies and high power density systems. Especially when faced with nonlinear characteristics and multi-physical field coupling effects, parameter tuning relies on empirical trial and error methods and it is difficult to achieve dynamic characteristic identification and parameter time-varying characteristic tracking.

Method used

A power converter control method based on the multi-experience replay pool TD3 algorithm is adopted to construct a power converter control system, including a dynamic perception layer, a strategy optimization layer and a lightweight execution layer. Through the improved TD3 algorithm and LSTM-RBF network, combined with a multi-experience replay buffer pool and a reward function, PID parameter adaptive tuning and lightweight deployment are achieved.

Benefits of technology

It significantly improves the control stability and efficiency of the power converter under complex working conditions, reduces the dependence on precise mathematical models, realizes intelligent optimization of control parameters, solves the resource bottleneck problem of LSTM in embedded deployment, and improves the interpretability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110126B_ABST
    Figure CN120110126B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of power converter management technology, and more specifically, to a power converter control method based on a multi-experience replay pool TD3 algorithm. The method comprises the following steps: S1, constructing a power converter control system: the power converter control system is composed of a power converter, a dynamic perception layer, a strategy optimization layer, a lightweight execution layer, and a PID controller; S2, implementing an improved TD3 algorithm: improving the TD3 framework based on the control characteristics of the power converter, adopting an innovative architecture of a multi-experience replay buffer pool, taking the stability, transient penalty, and safety of the power converter as an innovative reward function of the comprehensive reward value, and distilling the Actor online LSTM network into an RBF network. The design of the present invention adopts a TD3-PID hierarchical control structure to achieve optimal control under complex working conditions; it reduces dependence on precise mathematical models, and at the same time realizes intelligent optimization of control parameters through reinforcement learning, thereby improving control stability, significantly reducing computational complexity, and improving the interpretability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power converter management, and in particular to a power converter control method based on a multi-experience replay pool TD3 algorithm. Background Art

[0002] In the field of power converter control, the PID algorithm has long dominated industrial control due to its architectural simplicity, dynamic robustness, and engineering universality. However, the increasing penetration of renewable energy generation, the dynamic reconfiguration of smart grids, and the evolution of ultra-fast charging technologies (typically exemplified by the 650kW liquid-cooled supercharging system) have driven two major technological evolutions in power conversion systems: increasing topological complexity (with the prevalence of three-level and multi-level architectures) and an order of magnitude increase in power density. This has placed new demands on control algorithms for dynamic response speed and multi-objective coordination capabilities.

[0003] Traditional PID control faces dual performance bottlenecks in engineering applications: in the modeling dimension, its parameter adjustment has long relied on empirical trial and error due to the limitations of the system's nonlinear characteristics and multi-physical field coupling effects; in the control dimension, faced with multi-dimensional challenges in new application scenarios, including the wide-range grid impedance disturbances of grid-connected inverters, the nonlinear frequency domain characteristics of LLC resonant converters, and the electromechanical-electromagnetic multi-time-scale coupling dynamics of energy storage converters, its fixed parameter architecture makes it difficult to achieve dynamic characteristic identification and parameter time-varying characteristic tracking.

[0004] To address the aforementioned control efficiency bottlenecks, intelligently enhanced control strategies have become a research focus in recent years. Deep reinforcement learning, by constructing a self-evolutionary mechanism combining environment interaction, strategy evaluation, and parameter iteration, has demonstrated significant advantages in the field of dynamic system control. For example, while the invention patent "DC / DC Converter Control Method Based on TD3 Reinforcement Learning Algorithm" introduces a reinforcement learning framework, it still faces technical limitations in handling the multi-timescale coupled dynamics of power converters and lightweight deployment of control strategies. In light of this, we propose a power converter control method based on the TD3 algorithm with multiple experience replay pools. Summary of the Invention

[0005] The object of the present invention is to provide a power converter control method based on a multi-experience replay pool TD3 algorithm to solve the problems raised in the above background technology.

[0006] To solve the above technical problems, the present invention provides a power converter control method based on a multi-experience replay pool TD3 algorithm, comprising the following steps:

[0007] S1. Build a power converter control system: The power converter control system consists of a power converter, a dynamic perception layer, a strategy optimization layer, a lightweight execution layer, and a PID controller;

[0008] S2. Implementation of the Improved TD3 Algorithm: The TD3 framework is improved to address the control characteristics of power converters. An innovative architecture with multiple experience replay buffer pools is adopted. The stability, transient penalty, and safety of the power converter are used as an innovative reward function for the comprehensive reward value. The Actor online LSTM network is distilled into an RBF network to achieve network lightweighting. Specifically, the following are included:

[0009] S2.1. Establishing a Markov model for the power converter PID control task , the power converter output voltage , output voltage With reference voltage Error and first-order difference and second-order differences As an observation variable, the online Actor network For the amount of action, use Calculate the PID control law, take the stability, transient penalty and safety of the power converter as the comprehensive reward value, and Store in multiple experience replay buffer pool;

[0010] S2.2. The TD3 (Twin Delayed Deep Deterministic Policy Gradient) framework is used, consisting of one actor network and two critic networks. Both the actor and critic networks are equipped with online networks and target networks. Both networks are designed based on the LSTM architecture. The LSTM architecture enhances the network's ability to model historical state-action sequences, making it suitable for continuous control tasks in dynamic systems.

[0011] S2.3, the current state quantity Input Actor online network and generate actions To enhance the exploration of the strategy, Gaussian noise is superimposed on the output of the Actor online network. , generating the final action ;Will Converts to power converter control signals PWM / PFM / PSM, adjusts the switching state of power electronic devices, and collects real-time stability bonus values , safety bonus value and the state at the next moment ;

[0012] S2.4, based on the training environment, generate a variety of abnormal scenarios in the training environment to complete a training task of TD3 reinforcement learning; during the training process of TD3 network, the state quantity of abnormal samples will be , amount of movement , reward value and the next moment state , packaged into sample entries and stored in the abnormal situation pool; the experience with large TD error is stored in the high priority pool; the experience with small TD error is stored in the ordinary replay pool, focusing on collecting experiences when the system is in a stable operating state to ensure that the learned strategy can maintain the stability of the system;

[0013] S2.5. Collect samples from the high-priority pool and the normal playback pool according to the set ratio, and extract a certain number of samples from the abnormal pool; and amount of movement As the input of the two evaluation networks-online networks, the output is the cumulative return value , Numbers for the two Critic online networks;

[0014] S2.6. As the input of the Actor target network, it generates the output action volume , two Q-value estimates are generated through two Critic target networks ,in Numbers of the two Critic target networks , Smooth the noise for the target strategy; then calculate the target Q value according to the Bellman equation ,in, is the discount factor;

[0015] S2.7. Calculate the loss function of two online critic networks , which is calculated as follows:

[0016] , minimize the loss function of each online critic network through the optimizer and update the parameters of each online critic network;

[0017] S2.8. When the online critic network training is completed After a cycle, ; Calculate the loss function of the online Actor network , which is calculated as follows: , minimize the loss function of the online Actor network through the optimizer and update the parameters of the online Actor network;

[0018] S2.9. Repeat steps S2.5 to S2.8. When the online Actor network converges, distill the online Actor network LSTM into an RBF network.

[0019] S2.10. Deploy the RBF network as a PID controller to the power converter to achieve control of the power converter.

[0020] As a further improvement of this technical solution, in S1, the core of constructing the power converter control system includes constructing a dynamic perception layer, a strategy optimization layer, and a lightweight execution layer; wherein the construction of the dynamic perception layer includes the construction of the state space, the action space, and the reward function, and the ultimate goal is to construct an experience replay pool; specifically, the steps include:

[0021] A1. State space modeling: defining a state space based on a multidimensional observation vector;

[0022] A2. Action space design: constructing a continuous action space;

[0023] A3. Composite Reward Function Design: The reward function is designed to comprehensively consider stability and safety, minimizing voltage errors and guiding the agent to make optimal decisions while ensuring that the power converter remains in optimal health.

[0024] A4. Build multiple experience replay buffer pools: By introducing multiple specially designed experience pools, different types of replay pools should be optimized according to their specific purposes.

[0025] As a further improvement of this technical solution, in A1, the state space based on the multidimensional observation vector is defined as:

[0026] ;

[0027] in, The real-time output voltage of the power converter; is the voltage tracking error, is the reference voltage of the power converter; is the first-order difference of the error, which is used to reflect the dynamic response rate; is the second-order difference of the error, which is used to characterize the nonlinearity of the system.

[0028] As a further improvement of this technical solution, in A2, the continuous action space is constructed as follows:

[0029] ;

[0030] in, They correspond to the PID parameter increments respectively, and the time-varying parameter adjustment is achieved through the Actor network output.

[0031] As a further improvement of this technical solution, in A3, the reward function includes:

[0032] First, stability reward: introducing an exponential error penalty term to enhance steady-state accuracy;

[0033] Design a reward function that can effectively reflect stability. You can use the absolute value of the difference between the reference voltage and the current voltage and multiply it by a coefficient to calculate the reward. This reward function should encourage the agent to make the output parameter close to the set value, and give smaller rewards or penalties when the deviation is large:

[0034] ;

[0035] in, is the stability bonus value, is the weight coefficient used to adjust the intensity of the reward, is the current output voltage, is the reference voltage, is an exponential factor, usually 1 or 2; choose different The shape of the reward curve can be changed: represents a linear relationship, and It emphasizes greater deviation penalties; It is a small threshold used to define the range of "close"; is an additional bonus coefficient used to provide additional incentives when the voltage is close to the set value;

[0036] Second, transient penalty: adding a differential constraint term to suppress overshoot;

[0037] Penalize overshoot or undershoot to reduce extreme behavior during transients;

[0038] ;

[0039] in, is the transient penalty value, is the weight coefficient, is the change in voltage during the transient process;

[0040] Third, security boundaries: ;

[0041] Ensure that all operations are within safe limits and impose severe penalties for any behavior that exceeds safe limits;

[0042] ;

[0043] in, is the security bonus value, is the weight coefficient, and are the maximum allowable output voltage and current respectively;

[0044] Fourth, comprehensive reward function: combine the above sub-parts into a comprehensive reward function, which can be done by weighted summation;

[0045] ;

[0046] in, For comprehensive rewards, are the weights of each sub-reward, and these weights can be adjusted according to task requirements to balance the importance of different goals.

[0047] As a further improvement of this technical solution, in A4, building a multi-experience replay buffer pool specifically includes the following:

[0048] A total of three experience pools are designed, including an abnormal situation pool, a high-priority pool, and a normal replay pool. The abnormal situation pool is used to store experiences of dealing with sudden disturbances or abnormal situations, helping the agent learn how to recover or adapt quickly. The high-priority pool is used to store experiences with large TD errors. The normal replay pool is used to store experiences with small TD errors.

[0049] Regarding the capacity of the multi-experience replay buffer pool, the normal replay pool should be set to a larger capacity to receive experiences with smaller TD errors, provide diverse training samples, and use a FIFO or random replacement strategy to delete old data when the upper limit is reached. The high-priority pool stores important experiences with larger TD errors and maintains a small but sufficient capacity. The abnormal situation pool is used to store experiences that respond to sudden disturbances, and the capacity is flexibly adjusted according to actual needs. By properly configuring the capacity of each pool, the training data can be maximized, ensuring the model's learning efficiency, generalization ability, and robustness while avoiding the storage of excessive data that is no longer relevant.

[0050] To efficiently sample data from the multi-experience replay buffer pool, a weighted comprehensive sampling strategy is used in each training iteration: the majority of samples are randomly sampled from the common experience pool to ensure the latest environmental interaction results; a small number of important samples are sampled from the high-priority pool according to priority to quickly correct learning deviations; a small number of samples are occasionally sampled from the abnormal pool to ensure that the system learns to cope with extreme conditions; to balance the importance of each pool and ensure the overall learning effect, a weighted sampling method can be used:

[0051] ;

[0052] in, is the total sample set, Indicates that from The samples drawn from the pool, are the corresponding weights; these weights can be adjusted according to task requirements and technical feasibility to maximize the overall performance.

[0053] As a further improvement of this technical solution, in S2.2, the Actor online network does not rely on predefined labels or target values, but predicts the state-action pair Q value calculated by the Critic network based on the scalar reward value from the real-time feedback of the system and the long-term benefit that characterizes the current control strategy. , the input is the state space , covering the system output voltage, error and its differential signal, the output is the PID controller parameter increment , used to dynamically adjust the control strategy, optimize parameters through the gradient ascent method, and update network parameters based on the action value evaluation (Q value) and environmental reward signal provided by the critic network without predefined labels;

[0054] The Critic network is trained by minimizing the difference between the predicted Q value and the actual immediate reward plus the maximum Q value of the next state, namely the TD error. It does not use pre-specified labels, but relies on environmental feedback and its own estimation of future rewards. The input of the Critic network is the joint state-action pair , the output of the Critic network is a state-action pair The Q-value estimation is used to quantify the quality of the action. The optimization goal of the Critic is to minimize the temporal difference error, also known as the TD error. Through the LSTM neural network, it can better consider the influence of previous states and actions.

[0055] The collaborative optimization mechanism of the actor and critic network: the critic network guides the actor strategy optimization through Q-value evaluation, and the actor adjusts its actions based on the critic feedback to maximize the long-term cumulative reward.

[0056] As a further improvement of the present technical solution, in S2.4, the training environment of the power converter includes a deep learning intelligent agent Actor online LSTM network and a power converter, and the interaction between the two, wherein the intelligent agent Actor online LSTM network generates actions according to the state, and the intelligent agent Actor online network optimizes the control strategy through continuous interaction with the power converter, ultimately achieving efficient and stable operation under dynamic working conditions.

[0057] As a further improvement of the present technical solution, in S2.4, the abnormal scenarios that may occur in the training environment include at least load mutation, input voltage disturbance, overvoltage, overcurrent, etc.

[0058] As a further improvement of this technical solution, in S2.9, the process of distilling the online Actor network LSTM into an RBF network includes:

[0059] In a reinforcement learning environment, we train an LSTM-based actor online network to converge its control strategy, using it as a teacher model and constructing a lightweight RBF network as a student model.

[0060] Extract from the teacher model Input-output pairs, build a distillation training dataset, that is, run the converged Actor LSTM online network in the training environment, and collect state observation vector sequences , record the control action sequence corresponding to the Actor LSTM online network , divide the data into training set and validation set;

[0061] The distillation training process is to make the RBF network output close to the action of the Actor LSTM online network through supervised learning. The RBF student model is trained using Mini-batch SGD with MSE as the loss function and the Adam optimizer with a low learning rate. The loss function is:

[0062] ;

[0063] in, is the loss function, is the time period, is the action of the Actor LSTM online network, The action output by the RBF network.

[0064] Compared with the prior art, the present invention has the following beneficial effects:

[0065] 1. In this power converter control method based on the multi-experience replay pool TD3 algorithm, a TD3-PID hierarchical control structure is used to achieve optimal control under complex working conditions. The architecture consists of three core modules: the intelligent decision-making layer constructs a dynamic adjustment mechanism through deep reinforcement learning, collects the output voltage in real time and calculates its deviation from the reference value and key state parameters such as the first-order and second-order differences, and uses a multi-experience replay pool mechanism to enhance the training stability of the TD3 algorithm; the knowledge distillation layer innovatively introduces the LSTM-RBF network conversion module, and uses knowledge distillation technology to transfer the temporal feature extraction capability in the Actor network to the radial basis function network. , achieving lightweight deployment of control strategies; the parameter self-correction execution layer adopts a PID controller with online adjustment capability. Through the nonlinear mapping relationship between the multi-dimensional action space output by the Actor-Critic network and the PID parameter increment, it effectively copes with the complex operating conditions faced by power converters, including multi-time scale coupling problems such as variability in operating states, strong nonlinear characteristics, and electromechanical-electromagnetic transient interactions; compared with traditional control methods, this solution significantly reduces the dependence on precise mathematical models. At the same time, through the self-learning characteristics of reinforcement learning, it realizes intelligent optimization of control parameters, avoiding the subjective limitations of manual parameter tuning;

[0066] 2. In this power converter control method based on the multi-experience replay pool TD3 algorithm, the reward function design fully considers voltage errors and out-of-bounds penalties, as well as dynamic scenarios (such as transient overshoot), improving control stability.

[0067] 3. This power converter control method based on the multi-experience replay pool (TD3) algorithm distills the online actor network LSTM into an RBF network. This distillation allows the RBF network to inherit the control strategy of the actor LSTM online network while significantly reducing computational complexity, meeting the requirements of real-time power converter control. This method, while retaining the dynamic performance advantages of the reinforcement learning controller, addresses the resource bottleneck of the LSTM in embedded deployments.

[0068] 4. In this power converter control method based on the multi-experience replay pool TD3 algorithm, the deep learning LSTM network is distilled into an RBF network, which improves the interpretability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 FIG. 1 is a diagram of an exemplary overall control scheme of a power converter in the present invention;

[0070] Figure 2 This is an exemplary multi-experience replay pool process architecture diagram of the present invention;

[0071] Figure 3 This is a diagram of an exemplary TD3 network innovation architecture in the present invention;

[0072] Figure 4 This is an exemplary Actor-Critic network architecture diagram in the present invention;

[0073] Figure 5 This is an exemplary online actor network LSTM distilled into RBF network architecture diagram in the present invention;

[0074] Figure 6 This is an exemplary RBF network deployment architecture diagram in the present invention. DETAILED DESCRIPTION

[0075] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0076] Example 1

[0077] like Figures 1-6 As shown, this embodiment provides a power converter control method based on a multi-experience replay pool TD3 algorithm, comprising the following steps:

[0078] S1. Build a power converter control system: The power converter control system consists of a power converter, a dynamic perception layer, a strategy optimization layer, a lightweight execution layer, and a PID controller;

[0079] Among them, the core of the power converter control system is a three-layer control architecture that integrates deep reinforcement learning and knowledge distillation, such as Figure 1 As shown, including:

[0080] Dynamic perception layer: Integrates multimodal state observation and priority experience replay mechanism;

[0081] Strategy optimization layer: TD3 algorithm is used to realize PID parameter adaptive tuning;

[0082] Lightweight execution layer: Embedded deployment is achieved through LSTM-RBF distillation.

[0083] Specifically, the dynamic perception layer includes the construction of the experience replay buffer pool, which involves the design of state, action and reward functions, and is intended to store the experience generated when the agent interacts with the environment. , each experience is represented by the current state , actions taken , rewards received and the next state The state space includes the output voltage of the power converter , output voltage With reference voltage Error and first-order difference and second-order differences ; The action space covers the increment of PID control parameters , reflecting the dynamic process of system adjustment; the reward function comprehensively considers stability, overshoot or undershoot, and safety to guide the agent to make optimal decisions. Both the state space and the action space are continuous spaces, ensuring that the model can handle complex nonlinear relationships. Furthermore, by introducing multiple experience replay buffer pools, including a high-priority pool, a normal experience pool, and an abnormal situation pool, it not only more effectively manages and utilizes training data, but also significantly improves learning efficiency, generalization ability, and robustness, enabling the agent to achieve more efficient and stable control in complex and changing task environments.

[0084] The policy optimization layer uses the TD3 algorithm to achieve adaptive tuning of PID parameters. The TD3 algorithm is based on an actor-critic architecture, where the actor selects actions and the critic evaluates the quality of the actions. Both use LSTM neural networks as the training body. The policy network actor receives data input from the experience replay buffer pool and its output action is the incremental value of the PID controller parameters. In other words, the TD3 algorithm is used to learn and self-tune the PID controller parameters. The PID controller uses the action parameters generated by the actor. To adjust the control signal, the control signal is mapped to PWM, PFM and PSM modulation signals to control the stability of the output voltage of the power converter.

[0085] The lightweight execution layer achieves embedded deployment through LSTM-RBF distillation. Before applying the Actor LSTM online network to the PID controller, the Actor LSTM online network is distilled into an RBF network, which can significantly reduce the computational complexity and meet the requirements of real-time control of the power converter.

[0086] In addition, power converters are power converters, which include but are not limited to Buck, Boost, Buck-Boost, CUK and resonant converters, as well as flyback converters, forward converters, push-pull converters, half-bridge converters and full-bridge converters, and parallel topology architectures of these types of converters, and power converters that use ZVS and ZCS technologies to reduce switching losses of power devices.

[0087] Furthermore, in S1, the core of building a power converter control system includes constructing a dynamic perception layer, a strategy optimization layer, and a lightweight execution layer. The construction of the dynamic perception layer includes the construction of the state space, action space, and reward function. The ultimate goal is to build an experience replay pool. Traditional experience replay pools have shortcomings such as low sample utilization and insufficient coverage of abnormal working conditions. This solution improves training efficiency by constructing multiple experience replay pools and a dynamic sampling strategy. The specific steps include the following:

[0088] A1. State space modeling: defining a state space based on a multidimensional observation vector;

[0089] ;

[0090] in, The real-time output voltage of the power converter; is the voltage tracking error, is the reference voltage of the power converter; is the first-order difference of the error (used to reflect the dynamic response rate); is the second-order difference of the error (used to characterize the nonlinearity of the system);

[0091] A2. Action space design: constructing a continuous action space;

[0092] ;

[0093] in, They correspond to PID parameter increments respectively, and the time-varying parameter adjustment is achieved through the Actor network output

[0094] A3. Composite Reward Function Design: The reward function is designed to comprehensively consider stability and safety, aiming to minimize voltage errors and guide the agent to make optimal decisions while ensuring that the power converter maintains optimal health.

[0095] Among them, the reward function includes:

[0096] First, stability reward: introducing an exponential error penalty term to enhance steady-state accuracy;

[0097] Design a reward function that can effectively reflect stability. You can use the absolute value of the difference between the reference voltage and the current voltage and multiply it by a coefficient to calculate the reward. This reward function should encourage the agent to make the output parameter (such as voltage) close to the set value, and give smaller rewards or penalties when the deviation is large:

[0098] ;

[0099] in, is the stability bonus value, is the weight coefficient used to adjust the intensity of the reward, is the current output voltage, is the reference voltage (set value), is an exponential factor, usually 1 or 2; choose different The shape of the reward curve can be changed: represents a linear relationship, and It emphasizes greater deviation penalties; It is a small threshold used to define the range of "close"; is an additional bonus coefficient used to provide additional incentives when the voltage is close to the set value;

[0100] Second, transient penalty: adding a differential constraint term to suppress overshoot;

[0101] Penalize overshoot or undershoot to reduce extreme behavior during transients;

[0102] ;

[0103] in, is the transient penalty value, is the weight coefficient, is the change in voltage during the transient process;

[0104] Third, security boundaries: ;

[0105] Ensure that all operations are within safe limits and impose severe penalties for any behavior that exceeds safe limits;

[0106] ;

[0107] in, is the security bonus value, is the weight coefficient, and are the maximum allowable output voltage and current respectively;

[0108] Fourth, comprehensive reward function: combine the above sub-parts into a comprehensive reward function, which can be done by weighted summation;

[0109] ;

[0110] in, For comprehensive rewards, are the weights of each sub-reward, and these weights can be adjusted according to task requirements to balance the importance of different goals.

[0111] A4. Build multiple experience replay buffer pools: By introducing multiple specially designed experience pools, different types of replay pools should be optimized according to their specific purposes, such as storing important experiences, recording abnormal situations, etc. And by balancing the sampling method of the importance of each pool, combined with TD error calculation and network update, maximize the use of diverse training data, and improve the learning efficiency, generalization ability and robustness of the model. Figure 2 As shown, specifically including the following:

[0112] A total of three experience pools are designed, including an abnormal situation pool, a high-priority pool, and a normal replay pool. The abnormal situation pool is used to store experiences of responding to sudden disturbances or abnormal situations, helping the agent learn how to recover or adapt quickly. The high-priority pool is used to store experiences with large TD errors (e.g., the first 10%). The normal replay pool is used to store experiences with small TD errors (e.g., the last 90%). It focuses on collecting experiences when the system is in a stable operating state to ensure that the learned strategy can maintain system stability.

[0113] Regarding the capacity of the multi-experience replay buffer pool, the normal replay pool should be set to a larger capacity (e.g., 100,000 experiences) to receive experiences with smaller TD errors and provide diverse training samples. When the upper limit is reached, a FIFO or random replacement strategy should be used to delete old data. The high-priority pool stores important experiences with larger TD errors and maintains a smaller but sufficient capacity (e.g., 5,000 experiences). The abnormal situation pool is used to store experiences that respond to sudden disturbances, and its capacity can be flexibly adjusted based on actual needs. By properly configuring the capacity of each pool, the training data can be maximized, ensuring model learning efficiency, generalization ability, and robustness while avoiding the storage of excessive irrelevant data.

[0114] To efficiently sample data from the multi-experience replay buffer pool, a weighted comprehensive sampling strategy is used in each training iteration: a large number of samples (e.g., 80%) are randomly sampled from the normal experience pool to ensure the latest environmental interaction results; a small number of important samples (e.g., 15%) are sampled from the high-priority pool according to priority to quickly correct learning deviations; a small number of samples (e.g., 5) are occasionally sampled from the abnormal pool to ensure that the system learns to cope with extreme conditions; to balance the importance of each pool and ensure the overall learning effect, a weighted sampling method can be used:

[0115] ;

[0116] in, is the total sample set, Indicates that from The samples drawn from the pool, are the corresponding weights. These weights can be adjusted according to task requirements and technical feasibility to maximize the overall performance.

[0117] S2. Implementation of improved TD3 algorithm: Figure 3 As shown in the figure, the TD3 framework is improved to target the control characteristics of power converters. An innovative architecture of multiple experience replay buffer pools is adopted. The stability, transient penalty, and safety of the power converter are used as the innovative reward function for the comprehensive reward value. The Actor online LSTM network is distilled into an RBF network to achieve network lightweighting. Specifically, the following are included:

[0118] S2.1. Establishing a Markov model for the power converter PID control task , the power converter output voltage , output voltage With reference voltage Error and first-order difference and second-order differences As an observation variable, the online Actor network For the amount of action, use Calculate the PID control law, take the stability, transient penalty and safety of the power converter as the comprehensive reward value, and Store in multiple experience replay buffer pool;

[0119] S2.2, adopt the TD3 (Twin Delayed Deep Deterministic Policy Gradient) framework consisting of 1 Actor network and 2 Critic networks, such as Figure 4 As shown in the figure, both the Actor network and the Critic network are equipped with an online network and a target network. Both networks are designed based on the LSTM architecture. The LSTM architecture enhances the network's ability to model historical state-action sequences and is suitable for continuous control tasks in dynamic systems.

[0120] The Actor online network does not rely on predefined labels or target values, but predicts the state-action pair Q value calculated by the Critic network based on the scalar reward value from the real-time feedback of the system and the long-term benefit that characterizes the current control strategy. , the input is the state space , covering the system output voltage, error and its differential signal, the output is the PID controller parameter increment , used to dynamically adjust the control strategy, optimize parameters through the gradient ascent method, and update network parameters based on the action value evaluation (Q value) and environmental reward signal provided by the critic network without predefined labels;

[0121] The Critic network is trained by minimizing the difference between the predicted Q value and the actual immediate reward plus the maximum Q value of the next state (i.e., TD error). It does not use pre-specified labels, but relies on environmental feedback and its own estimation of future rewards. The input of the Critic network is the joint state-action pair , the output of the Critic network is a state-action pair The Q-value estimation is used to quantify the quality of the action. The optimization goal of Critic is to minimize the temporal difference error (TD error), and through the LSTM neural network, it can better consider the influence of previous states and actions;

[0122] The collaborative optimization mechanism of the actor and critic network: the critic network guides the actor strategy optimization through Q-value evaluation, and the actor adjusts its actions based on the critic feedback to maximize the long-term cumulative reward.

[0123] S2.3, the current state quantity (voltage, voltage error, differential value, etc.) input Actor online network to generate action To enhance the exploration of the strategy, Gaussian noise is superimposed on the output of the Actor online network. , generating the final action ;Will Converts to power converter control signals (PWM / PFM / PSM), adjusts the switching state of power electronic devices, and collects real-time stability bonus values (such as voltage fluctuation suppression), safety bonus value (such as overcurrent / overvoltage protection) and the status at the next moment The power converter training environment consists of a deep learning agent, an online LSTM network, and the interaction between the two. This is a virtualized concept. The online LSTM network generates actions (such as adjusting the PWM signal) based on states (such as voltage, voltage error, first-order difference, and second-order difference). The online network optimizes the control strategy through continuous interaction with the power converter (state perception → action execution → reward feedback), ultimately achieving efficient and stable operation under dynamic conditions (such as sudden load changes and input voltage disturbances).

[0124] S2.4. Based on the training environment, a variety of abnormal scenarios such as load mutation, input voltage disturbance, overvoltage, overcurrent, etc. are generated in the training environment to complete a training task of TD3 reinforcement learning; during the training process of TD3 network, the state quantity of abnormal samples such as overcurrent and overvoltage will be generated. , amount of movement , reward value and the next moment state , packaged into sample entries and stored in the abnormal situation pool; the experience with large TD error (for example, the first 10%) is stored in the high priority pool; the experience with small TD error (for example, the last 90%) is stored in the normal replay pool. Focus on collecting experiences when the system is in a stable operating state to ensure that the learned strategy can maintain the stability of the system;

[0125] S2.5. Collect samples from the high-priority pool and the normal playback pool according to the set ratio, and extract a certain number of samples from the abnormal pool; and amount of movement As the input of the two evaluation networks-online networks, the output is the cumulative return value , Numbers for the two Critic online networks;

[0126] S2.6. As the input of the Actor target network, it generates the output action volume , two Q-value estimates are generated through two Critic target networks ,in Numbers of the two Critic target networks , Smooth the noise for the target strategy; then calculate the target Q value according to the Bellman equation ,in, is the discount factor (e.g. );

[0127] S2.7. Calculate the loss function of two online critic networks , which is calculated as follows:

[0128] , minimize the loss function of each online critic network through the optimizer and update the parameters of each online critic network;

[0129] S2.8. When the online critic network training is completed After a cycle, ; Calculate the loss function of the online Actor network , which is calculated as follows: , minimize the loss function of the online Actor network through the optimizer and update the parameters of the online Actor network;

[0130] S2.9. Repeat steps S2.5 to S2.8. When the online Actor network converges, distill the online Actor network LSTM into an RBF network.

[0131] Among them, Figure 5 As shown in Figure 2, the process of distilling the online Actor network LSTM into an RBF network includes:

[0132] In a reinforcement learning environment, we train an LSTM-based actor online network to converge its control strategy, using it as a teacher model and constructing a lightweight RBF network as a student model.

[0133] Extract from the teacher model Input-output pairs, build a distillation training dataset, that is, run the converged Actor LSTM online network in the training environment, and collect state observation vector sequences , record the control action sequence corresponding to the Actor LSTM online network , divide the data into training set and validation set;

[0134] The distillation training process is to make the RBF network output close to the action of the Actor LSTM online network through supervised learning. The RBF student model is trained using Mini-batch SGD with MSE as the loss function and the Adam optimizer with a low learning rate. The loss function is:

[0135] ;

[0136] in, is the loss function, is the time period, is the action of the Actor LSTM online network, The action output by the RBF network.

[0137] Through this distillation process, the RBF network inherits the control strategy of the Actor LSTM online network while significantly reducing computational complexity, meeting the real-time control requirements of power converters. This approach addresses the resource bottleneck of LSTM in embedded deployments while retaining the dynamic performance advantages of reinforcement learning controllers.

[0138] S2.10, deploy the RBF network as a PID controller to the power converter to achieve the control of the power converter, such as Figure 6 shown.

[0139] In summary, this technical solution proposes multiple experience replay buffer pools, including a high-priority pool, a general experience pool, and an abnormal situation pool. This not only more effectively manages and utilizes training data, but also significantly improves learning efficiency, generalization ability, and robustness, enabling the intelligent agent to achieve more efficient and stable control in complex and changing task environments.

[0140] At the same time, the output voltage of the power converter is proposed , output voltage With reference voltage Error and first-order difference and second-order differences The state space as the observed variables;

[0141] It also proposes that the action space is defined as the increment of PID parameters, namely the proportional gain increment, the integral gain increment and the differential gain increment.

[0142] A comprehensive reward function of stability reward and security reward is proposed; three experience pools are designed: an abnormal situation pool, a high priority pool, and a normal replay pool; and the online actor network LSTM is distilled into an RBF network.

[0143] Those skilled in the art will appreciate that the process of implementing all or part of the steps of the above embodiments may be accomplished by hardware, or by instructing related hardware through a program.

[0144] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.

Claims

1. A power converter control method based on a multi-experience playback pool TD3 algorithm, characterized in that: The steps include: S1. Build a power converter control system: The power converter control system consists of a power converter, a dynamic perception layer, a strategy optimization layer, a lightweight execution layer, and a PID controller; S2. Implementation of the Improved TD3 Algorithm: The TD3 framework is improved to address the control characteristics of power converters. An innovative architecture with multiple experience replay buffer pools is adopted. The stability, transient penalty, and safety of the power converter are used as an innovative reward function for the comprehensive reward value. The Actor online LSTM network is distilled into an RBF network to achieve network lightweighting. Specifically, the following are included: S2.

1. Establishing a Markov model for the power converter PID control task , the power converter output voltage , output voltage With reference voltage Error and first-order difference and second-order differences As an observation variable, the online Actor network For the amount of action, use Calculate the PID control law, take the stability, transient penalty and safety of the power converter as the comprehensive reward value, and Store in multiple experience replay buffer pool; S2.

2. Use the TD3 framework, which consists of one actor network and two critic networks. Both the actor network and the critic network are equipped with an online network and a target network. Both networks are designed based on the LSTM architecture. S2.3, the current state quantity Input Actor online network and generate actions , superimpose Gaussian noise on the output of the Actor online network , generating the final action ;Will Converts to power converter control signals PWM / PFM / PSM, adjusts the switching state of power electronic devices, and collects real-time stability bonus values , safety bonus value and the state at the next moment ; S2.4, based on the training environment, generate a variety of abnormal scenarios in the training environment to complete a training task of TD3 reinforcement learning; during the training process of TD3 network, the state quantity of abnormal samples will be , amount of movement , reward value and the next moment state , packaged into sample entries and stored in the abnormal situation pool; the experience with larger TD error is stored in the high priority pool; the experience with smaller TD error is stored in the normal replay pool; S2.

5. Collect samples from the high-priority pool and the normal playback pool according to the set ratio, and extract a certain number of samples from the abnormal pool; and amount of movement As the input of two evaluation networks - online networks, namely the Critic online network, the output is the cumulative return value , Numbers for the two Critic online networks; S2.

6. As the input of the Actor target network, it generates the output action volume , two Q-value estimates are generated through two Critic target networks ,in Numbers of the two Critic target networks , Smooth the noise for the target strategy; then calculate the target Q value according to the Bellman equation ,in, is the discount factor; S2.

7. Calculate the loss function of two online critic networks , which is calculated as follows: , minimize the loss function of each online critic network through the optimizer and update the parameters of each online critic network; S2.

8. When the online critic network training is completed After a cycle, ; Calculate the loss function of the online Actor network , which is calculated as follows: , minimize the loss function of the online Actor network through the optimizer and update the parameters of the online Actor network; S2.

9. Repeat steps S2.5 to S2.

8. When the online Actor network converges, distill the online Actor network LSTM into an RBF network. S2.

10. Deploy the RBF network as a PID controller to the power converter to achieve control of the power converter.

2. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 1, characterized in that: In S1, the core of constructing the power converter control system includes constructing a dynamic perception layer, a strategy optimization layer, and a lightweight execution layer. The construction of the dynamic perception layer includes the construction of the state space, the action space, and the reward function, with the ultimate goal of constructing an experience replay pool. The specific steps include: A1. State space modeling: defining a state space based on a multidimensional observation vector; A2. Action space design: constructing a continuous action space; A3. Composite Reward Function Design: The reward function is designed to comprehensively consider stability and safety, minimizing voltage errors and guiding the agent to make optimal decisions while ensuring that the power converter remains in optimal health. A4. Build multiple experience replay buffer pools: By introducing multiple specially designed experience pools, different types of replay pools should be optimized according to their specific purposes.

3. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 2, characterized in that: In A1, the state space based on the multidimensional observation vector is defined as: ; in, The real-time output voltage of the power converter; is the voltage tracking error, is the reference voltage of the power converter; is the first-order difference of error; is the second-order difference of the error.

4. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 3, characterized in that: In A2, the continuous action space is constructed as: ; in, They correspond to the PID parameter increments respectively, and the time-varying parameter adjustment is achieved through the Actor network output.

5. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 4 is characterized in that: In A3, the reward function includes: First, stability reward: introducing an exponential error penalty term to enhance steady-state accuracy; Design a reward function that can effectively reflect stability. Use the absolute value of the difference between the reference voltage and the current voltage and multiply it by a coefficient to calculate the reward. This reward function makes the output parameter close to the set value and gives a smaller reward or penalty when the deviation is large: ; in, is the stability bonus value, is the weight coefficient used to adjust the intensity of the reward, is the current output voltage, is the reference voltage, is an exponential factor, usually taking the value of 1 or 2; Is a small threshold used to define the range of "close"; is an additional bonus coefficient used to provide additional incentives when the voltage is close to the set value; Second, transient penalty: adding a differential constraint term to suppress overshoot; Penalize overshoot or undershoot to reduce extreme behavior during transients; ; in, is the transient penalty value, is the weight coefficient, is the change in voltage during the transient process; Third, security boundaries: ; Ensure that all operations are within safe limits and impose severe penalties for any behavior that exceeds safe limits; ; in, is the security bonus value, is the weight coefficient, and are the maximum allowable output voltage and current respectively; Fourth, comprehensive reward function: combine the above sub-parts into a comprehensive reward function using a weighted summation method; ; in, For comprehensive rewards, is the weight of each sub-part reward.

6. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 5, characterized in that: In A4, building a multi-experience replay buffer pool specifically includes the following: A total of three experience pools are designed, including an abnormal situation pool, a high-priority pool, and a normal replay pool. The abnormal situation pool is used to store experiences in dealing with sudden disturbances or abnormal situations; the high-priority pool is used to store experiences with large TD errors; and the normal replay pool is used to store experiences with small TD errors. Regarding the capacity of the multi-experience replay buffer pool, the normal replay pool should be set to a larger capacity to receive experiences with smaller TD errors, provide diverse training samples, and use a FIFO or random replacement strategy to delete old data when the upper limit is reached. The high-priority pool stores important experiences with larger TD errors and maintains a smaller but sufficient capacity. The abnormal situation pool is used to store experiences that respond to sudden disturbances, and the capacity can be flexibly adjusted according to actual needs. To efficiently sample data from the multi-experience replay buffer pool, a weighted comprehensive sampling strategy is used in each training iteration: the majority of samples are randomly sampled from the common experience pool to ensure the latest environmental interaction results; a small number of important samples are sampled from the high-priority pool according to priority to quickly correct learning deviations; a small number of samples are occasionally sampled from the abnormal pool to ensure that the system learns to cope with extreme conditions; to balance the importance of each pool and ensure the overall learning effect, a weighted sampling method can be used: ; in, is the total sample set, Indicates that from The samples drawn from the pool, is the corresponding weight.

7. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 6, characterized in that: In S2.2, the Actor online network does not rely on predefined labels or target values, but predicts the state-action pair Q value calculated by the Critic network based on the scalar reward value from the real-time feedback of the system and the long-term benefit that characterizes the current control strategy. , the input is the state space , covering the system output voltage, error and its differential signal, the output is the PID controller parameter increment , used to dynamically adjust the control strategy, optimize parameters through the gradient ascent method, and update network parameters based on the action value evaluation (Q value) and environmental reward signal provided by the Critic network without predefined labels; The Critic network is trained by minimizing the difference between the predicted Q value and the actual immediate reward plus the maximum Q value of the next state, namely the TD error. It does not use pre-specified labels, but relies on environmental feedback and its own estimation of future rewards. The input of the Critic network is the joint state-action pair , the output of the Critic network is a state-action pair The Q value estimation is used to quantify the quality of the action. The optimization goal of the Critic is to minimize the temporal difference error, also known as the TD error. The collaborative optimization mechanism of the actor and critic network: the critic network guides the actor strategy optimization through Q-value evaluation, and the actor adjusts its actions based on the critic feedback to maximize the long-term cumulative reward.

8. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 7, characterized in that: In S2.4, the training environment of the power converter includes the deep learning agent Actor online LSTM network and the power converter, and the interaction between the two, wherein the agent Actor online LSTM network generates actions according to the state, and the agent Actor online network optimizes the control strategy through continuous interaction with the power converter, ultimately achieving efficient and stable operation under dynamic working conditions.

9. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 8, characterized in that: In S2.4, the abnormal scenarios in the training environment include at least load mutation, input voltage disturbance, overvoltage, and overcurrent.

10. The power converter control method based on the multi-experience replay pool TD3 algorithm according to claim 9, characterized in that: In S2.9, the process of distilling the online Actor network LSTM into an RBF network includes: In a reinforcement learning environment, we train an LSTM-based actor online network to converge its control strategy, using it as a teacher model and constructing a lightweight RBF network as a student model. Extract from the teacher model Input-output pairs, build a distillation training dataset, that is, run the converged Actor LSTM online network in the training environment, and collect state observation vector sequences , record the control action sequence corresponding to the Actor LSTM online network , divide the data into training set and validation set; The distillation training process is to make the RBF network output close to the action of the Actor LSTM online network through supervised learning. The RBF student model is trained using MSE as the loss function, small block gradient descent, Adam optimizer, and a low learning rate. The loss function is: ; in, is the loss function, is the time period, is the action of the Actor LSTM online network, The action output by the RBF network.

Citation Information

Patent Citations

  • Aircraft decoupling-free attitude control method based on TD3 multi-experience pool reinforcement learning

    CN115857530A

  • SA-TD3 algorithm based on TD error priority sampling and Adam adaptive learning rate

    CN117787383A