Digital switching power supply neural network controller training method based on reinforcement learning

By using reinforcement learning training methods in the simulation environment of digital switching power supply and combining Gaussian reward function for multiple training, the problem of random control results of digital switching power supply is solved, and the robustness and adaptability of the system is improved.

CN120010272AActive Publication Date: 2025-05-16NORTHWESTERN POLYTECHNICAL UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510482196.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-16
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing reinforcement learning training methods have relatively random control results for digital switching power supplies, which cannot effectively ensure optimal performance, and are difficult to adapt to load sudden changes and input voltage fluctuations.

Method used

Using the digital switching power supply neural network controller training method based on reinforcement learning, by establishing a power-level circuit model and reinforcement learning agent of the digital switching power supply in a simulation environment, multiple reinforcement learning training are performed using symmetric Gaussian reward functions and asymmetric Gaussian reward functions, and the reward function parameters are adjusted to optimize the control strategy.

Benefits of technology

Automatic exploration of optimal control strategies in complex nonlinear systems is realized, which significantly improves the robustness and adaptability of the system, and can maintain the stability of the output voltage under load transients and input voltage fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010272A_ABST
    Figure CN120010272A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of reinforcement learning. The invention provides a digital switching power supply neural network controller training method based on reinforcement learning. According to the embodiment of the invention, the neural network controller of the digital switching power supply is designed, mathematical modeling is not needed, and an excellent control effect can be realized only by using the design and control of the training process. Quantitative adjustment of control performance and voltage overshoot can be realized by using a Gaussian reward function based on adjustable parameters, and a controller conforming to performance indexes is designed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the technical field of reinforcement learning, and in particular to a training method for a neural network controller of a digital switching power supply based on reinforcement learning. Background Art

[0002] With the rapid development of integrated circuits and information technology, the functions of electronic systems are becoming increasingly powerful, and the scale of circuits is growing continuously, resulting in a sharp increase in current consumption and frequent load mutations, which puts higher requirements on the performance of power supplies. Compared with traditional analog control switching power supplies, digital control switching power supplies (or digital control DC-DC switching converters, referred to as digital power supplies) have the advantages of being programmable, reconfigurable, flexible in design, able to execute various advanced and complex control algorithms, able to achieve multi-phase or multi-channel simultaneous output, good system control robustness, and able to communicate with the host in real time. It has obvious advantages in improving the steady-state and transient performance of power supplies, improving power density and conversion efficiency, reducing manufacturing costs, integrating components, and shortening product development cycles. Its application scope is gradually expanding, and it has broad application prospects in consumer electronics, industrial electronics, cloud computing and big data centers, as well as military and aerospace electronics.

[0003] At present, digital power products generally use traditional linear (such as PID control) or nonlinear (such as fuzzy logic control, sliding mode control) control methods. Since these traditional control algorithms are based on the linear steady-state small signal model of the system or the optimal control of "expert experience", it is difficult to further improve the performance of the power supply by using traditional linear or nonlinear control methods for complex nonlinear time-varying systems such as switching power supplies. Since neural networks have the advantages of highly parallel structure, strong learning ability, approximation ability of arbitrary functions, and super fault tolerance, they are particularly suitable for complex nonlinear time-varying systems such as switching power supplies. Through the training of neural network models, the optimal control of switching power supply systems can be achieved. Therefore, digital power supply intelligent control based on neural networks is an inevitable development trend of digital power supplies. In recent years, it has received extensive attention from academia and industry, and has become a research and development hotspot in the field of power supplies today.

[0004] In the related art, there are currently two main methods for training neural network control in the field of digital switching power supplies, namely supervised learning and reinforcement learning. The training process of supervised learning requires a large amount of data to be collected and labeled, and the performance of the controller is optimized by continuously reducing the difference between the output of the neural network controller and the calibration data. However, how to collect high-quality data that is consistent with the current digital power supply parameters and working conditions is a difficulty. Reinforcement learning as a method for training digital switching power supply neural network controllers has attracted attention because its training process does not require small signal modeling of the switching converter or labeled data. Reinforcement learning allows the neural network controller to be trained in the process of interacting with the controlled circuit in a simulated working environment. However, there are also problems with the existing reinforcement learning training process for digital switching power supplies. The training results of reinforcement learning for switching converters are relatively random, the final performance of the controller cannot be well controlled, and whether its performance is optimal cannot be guaranteed.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that this section is intended to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art by virtue of being included in this section. Summary of the invention

[0007] The purpose of the embodiments of the present disclosure is to provide a digital switching power supply neural network controller training method based on reinforcement learning, thereby overcoming one or more problems caused by the limitations and defects of related technologies at least to a certain extent.

[0008] According to an embodiment of the present disclosure, a method for training a neural network controller of a digital switching power supply based on reinforcement learning is provided, the method comprising: Step S1, establishing a power level circuit model and a reinforcement learning agent of a digital switching power supply in a simulation environment; wherein the power level circuit model includes an input power supply, a switch tube, a filter circuit and an output end, and the reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network, wherein the actor network is used to receive a state vector output by an observer, and the critic network is used to evaluate a control strategy output by the actor network; Step S2, based on the reinforcement learning agent and the power level circuit model, a symmetric Gaussian reward function SGRF is used to initialize the value of the SGRF parameter, and the executor network and the critic network are trained for the first reinforcement learning by the DDPG algorithm, and the reward function parameters are gradually reduced until the training does not converge; Step S3, based on the network parameters of the actor-critic network trained by the first reinforcement learning, an asymmetric Gaussian reward function AGRF is used to initialize the AGRF parameters, and the actor-critic network is trained by the second reinforcement learning through the DDPG algorithm, and the values ​​of the AGRF parameters are adjusted to eliminate the voltage overshoot in the startup phase; Step S4, based on the network parameters of the actor-critic network trained by the second reinforcement learning, introduce a load mutation event in the power level circuit model, use an asymmetric Gaussian reward function AGRF, and fix the SGRF parameters and AGRF parameters, and perform a third reinforcement learning training on the actor-critic network through the DDPG algorithm until the training is completed; Step S5, extracting the network parameters of the actor network in the actor-critic network to construct a neural network controller of the digital switching power supply to control the stability of the output voltage in real time.

[0009] Furthermore, the state vector includes an output voltage error, an error integral and an error differential value, and outputs a duty cycle signal to the power stage circuit model.

[0010] Furthermore, step S2 specifically includes: Step S21, initialize the actor-critic network, use the symmetric Gaussian reward function SGRF, and initialize the SGRF parameters The value of is 1; Step S22, executing a digital switch power supply startup event; Step S23, using the DDPG algorithm to perform the first reinforcement learning training on the actor-critic network; Step S24, if the output voltage of the digital switching power supply under the control of the controller during the training of the executor network and the critic network Stable , then the DDPG algorithm converges; Step S25, set the SGRF parameters Reduce by 0.05, repeat step S23 and step S24, and determine whether the DDPG algorithm converges; Step S26: If the DDPG algorithm does not converge, take the last SGRF parameter that can converge The value of is taken as the SGRF parameter with the best performance The value of , and the first reinforcement learning training ends; if the DDPG algorithm converges, step S25 is repeated until the DDPG algorithm does not converge.

[0011] Furthermore, step S3 specifically includes: Step S31, import the network parameters of the actor-critic network trained by the first reinforcement learning into the new training environment, and use the asymmetric Gaussian reward function AGRF, SGRF parameters The value of continues the SGRF parameter with the best performance in step S26 The value of AGRF parameter The value of is initially set to 0; Step S32, executing a digital switch power supply startup event; Step S33, using the DDPG algorithm to perform a second reinforcement learning training on the actor-critic network; Step S34, determining the output voltage of the digital switching power supply under the control of the controller during the training of the performer network and the critic network Is there a voltage overshoot? If there is a voltage overshoot, set the SGRF parameter The value of is increased by 0.5, and steps S33 and S34 are repeated; if there is no voltage overshoot, the voltage output rises steadily during the startup phase, and the next step is entered.

[0012] Furthermore, step S4 specifically includes: Step S41, import the network parameters of the actor-critic network trained by the second reinforcement learning into the new training environment, and use the asymmetric Gaussian reward function AGRF, SGRF parameters The values ​​and AGRF parameters The value of takes the final value of the second reinforcement learning; Step S42, executing a digital switch power supply startup event and a load mutation event; Step S43, using the DDPG algorithm to perform a third reinforcement learning training on the actor-critic network until the training is completed.

[0013] Furthermore, the expression of the symmetric Gaussian reward function SGRF is:

[0014] in, is the AGRF parameter, is the output voltage, is the ideal output reference value of the output voltage; The expression of asymmetric Gaussian reward function AGRF is:

[0015] in, is the AGRF parameter.

[0016] Furthermore, the DDPG algorithm includes the following steps: Initialize the network parameters of the actor-critic network; Storing training data through the experience replay pool; Iteratively update the network parameters of the actor-critic network to maximize the cumulative reward value, and synchronously update the network parameters of the actor-critic network.

[0017] Furthermore, a digitally controlled DC-DC switching power supply is formed according to the power stage circuit model, the analog-to-digital converter in the feedback control loop, the trained digital switching power supply neural network controller and the digital pulse width modulator; wherein, The power stage circuit model turns on and off the power switch tube according to the switch signal of the previous duty cycle to adjust the input voltage Perform chopping; the chopped input voltage The filtering by capacitor C and inductor L forms a voltage lower than the input voltage. Output voltage , to complete the voltage from the input voltage To output voltage of blood pressure reduction; The analog-to-digital converter outputs a voltage Convert to value; and Value and reference voltage value Subtract, output voltage error ; The voltage error The value of The points and The differential value of is used as the output vector of the observer , which is used as the input vector of the trained digital switching power supply neural network controller, and the reward function by As input variable; input vector After calculation by the executor network, the duty cycle signal is output That is, the proportion of the power switch tube conduction time to the switching cycle; Digital pulse width modulator for duty cycle signal Modulation is performed to form a switching signal, which is then output to the power switch tube in the power stage circuit model to adjust the ratio of the on-time and off-time of the power switch tube so that the output voltage at the current moment is is adjusted.

[0018] Furthermore, load mutation events include: After the output voltage stabilizes, the load changes from no-load to heavy-load and maintains for a preset time, and then changes back to no-load.

[0019] Furthermore, the switching power supply startup event includes: The startup process refers to the transition of the switching power supply from the non-working state 0V output to the working state.

[0020] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects: In the embodiments of the present disclosure, through the above-mentioned reinforcement learning-based training method for the neural network controller of the digital switching power supply, on the one hand, for the design of the neural network controller of the digital switching power supply, there is no need to perform mathematical modeling, and excellent control effects can be achieved only by designing and controlling the training process. The Gaussian reward function based on adjustable parameters can be used to quantitatively adjust the control performance and voltage overshoot, and design a controller that meets the performance indicators. On the other hand, the method automatically explores the optimal control strategy through a trial and error mechanism, without manually designing complex control laws or adjusting PID parameters. It is possible to find the global optimal solution in a complex nonlinear system, avoiding the local optimal problem that may exist in the traditional design method. DC-DC converters are usually faced with dynamic changes such as load transients and input voltage fluctuations. Reinforcement learning can adjust the control strategy in real time to adapt to these changes, significantly improving the robustness and adaptability of the system. It can also show better fault tolerance for parameter perturbations. The method is based on data-driven and can work effectively even when the model is not completely clear. It reduces the cumbersome design and debugging process and shortens the development cycle. The method can quickly adapt to new working scenarios or performance requirements by updating the policy network. Further improvements can be made on the basis of this method, and other performance of the controller can be increased by adding training events during the training process. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0022] Figure 1 A diagram showing the steps of a training method for a digital switching power supply neural network controller based on reinforcement learning in an exemplary embodiment of the present disclosure; Figure 2 A schematic diagram showing a training environment for a digital power neural network controller in an exemplary embodiment of the present disclosure; Figure 3 A schematic diagram showing the structure of a conventional digitally controlled DC-DC switching power supply in an exemplary embodiment of the present disclosure is shown; Figure 4 A schematic diagram showing a basic topological structure of a DC-DC switching power supply power stage circuit in an exemplary embodiment of the present disclosure; Figure 5A schematic diagram of the structure of a DC-DC switching power supply with a neural network controller in an exemplary embodiment of the present disclosure is shown; Figure 6 A schematic diagram showing the network structure of an implementer network and a critic network in an exemplary embodiment of the present disclosure; Figure 7 A schematic diagram of a training process in an exemplary embodiment of the present disclosure is shown; Figure 8 A schematic diagram showing the construction of a reinforcement learning training environment in simulink with a Buck-type switching power supply as a controlled object in an exemplary embodiment of the present disclosure is shown; Fig. 9 A schematic diagram showing the construction of a Buck-type digital switching power supply in simulink in an exemplary embodiment of the present disclosure is shown; Fig.10 The different exemplary embodiments of the present disclosure are shown The reward value distribution of the reward function under the value; Fig.11 The different exemplary embodiments of the present disclosure are shown The number of training rounds under the value and the changing trend of the rewards obtained from training; Fig.12 The different exemplary embodiments of the present disclosure are shown The output voltage startup situation of the switching power supply under the value; Fig.13 The different exemplary embodiments of the present disclosure are shown The reward value distribution of the reward function under the value; Fig.14 The same exemplary embodiment of the present disclosure is shown Value, different The output voltage startup situation of the DC-DC switching power supply under the value; Fig.15 The output voltage waveform of the DC-DC switching power supply at the startup stage after three trainings in the exemplary embodiment of the present disclosure is shown; Fig.16 A schematic diagram showing that a trained actor network in an exemplary embodiment of the present disclosure is taken out separately and implemented in a real circuit control environment; Fig.17 A schematic diagram showing a test system in which an executor network is taken out and implemented in an FPGA and constituted in an exemplary embodiment of the present disclosure; Fig.18 A diagram showing steady-state performance results of a DC-DC switching power supply in an exemplary embodiment of the present disclosure is shown; Fig.19 A diagram showing transient performance results of a DC-DC switching power supply under input voltage fluctuation in an exemplary embodiment of the present disclosure; Fig. 20 The transient performance of the DC-DC switching power supply at startup in the exemplary embodiment of the present disclosure is shown; Fig.21 The load change transient performance of the DC-DC switching power supply under different capacitor and inductor combinations in the exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0023] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0024] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated descriptions will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0025] This example embodiment provides a method for training a digital switching power supply neural network controller based on reinforcement learning. Figure 1 As shown in , the digital switching power supply neural network controller training method based on reinforcement learning may include: Step S1, establishing a power level circuit model and a reinforcement learning agent of a digital switching power supply in a simulation environment; wherein the power level circuit model includes an input power supply, a switch tube, a filter circuit and an output end, and the reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network, wherein the actor network is used to receive a state vector output by an observer, and the critic network is used to evaluate a control strategy output by the actor network; Step S2, based on the reinforcement learning agent and the power level circuit model, a symmetric Gaussian reward function SGRF is used to initialize the value of the SGRF parameter, and the executor network and the critic network are trained for the first reinforcement learning by the DDPG algorithm, and the reward function parameters are gradually reduced until the training does not converge; Step S3, based on the network parameters of the actor-critic network trained by the first reinforcement learning, an asymmetric Gaussian reward function AGRF is used to initialize the AGRF parameters, and the actor-critic network is trained by the second reinforcement learning through the DDPG algorithm, and the values ​​of the AGRF parameters are adjusted to eliminate the voltage overshoot in the startup phase; Step S4, based on the network parameters of the actor-critic network trained by the second reinforcement learning, introduce a load mutation event in the power level circuit model, use an asymmetric Gaussian reward function AGRF, and fix the SGRF parameters and AGRF parameters, and perform a third reinforcement learning training on the actor-critic network through the DDPG algorithm until the training is completed; Step S5, extracting the network parameters of the actor network in the actor-critic network to construct a neural network controller of the digital switching power supply to control the stability of the output voltage in real time.

[0026] Through the above-mentioned reinforcement learning-based training method for the neural network controller of the digital switching power supply, on the one hand, for the design of the neural network controller of the digital switching power supply, there is no need for mathematical modeling, and excellent control effects can be achieved only by designing and controlling the training process. The Gaussian reward function based on adjustable parameters can be used to quantitatively adjust the control performance and voltage overshoot, and design a controller that meets the performance indicators. On the other hand, the method automatically explores the optimal control strategy through a trial-and-error mechanism, without manually designing complex control laws or adjusting PID parameters. It can find the global optimal solution in a complex nonlinear system, avoiding the local optimal problem that may exist in traditional design methods. DC-DC converters are usually faced with dynamic changes such as load transients and input voltage fluctuations. Reinforcement learning can adjust the control strategy in real time to adapt to these changes, significantly improving the robustness and adaptability of the system. It can also show better fault tolerance for parameter perturbations (such as component aging). The method is data-driven and can work effectively even when the model is not completely clear. It reduces the tedious design and debugging process and shortens the development cycle. The method can quickly adapt to new working scenarios or performance requirements by updating the strategy network. Further improvements can be made based on this method by increasing other controller performance (such as predictive control and self-healing capabilities) by adding training events during the training process.

[0027] Next, we will refer to Figures 1 to 21 The various steps of the above-mentioned reinforcement learning-based digital switching power supply neural network controller training method in this example implementation are described in more detail.

[0028] In step S1, a power level circuit model and a reinforcement learning agent of a digital switching power supply are established in a simulation environment; wherein the power level circuit model includes an input power supply, a switch tube, a filter circuit and an output terminal. The reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network, wherein the actor network is used to receive a state vector output by an observer, and the critic network is used to evaluate a control strategy output by the actor network.

[0029] Specifically, the training environment of this application is constructed in the Simulink component of the Matlab software and is divided into two parts, such as Figure 2 As shown, one part is the power stage part of the digital switching power supply, and the other part is the reinforcement learning agent including the neural network controller.

[0030] In a specific embodiment, the structure of the digitally controlled DC-DC switching power supply is as follows: Figure 3 The figure shows a typical feedback tracking control system, which consists of a power stage circuit (i.e., a power stage circuit model) and an analog-to-digital converter (ADC), a digital compensator, and a digital pulse width modulator (DPWM) in the feedback control loop. The ADC, digital compensator, and DPWM constitute the digital controller of the DC-DC switching power supply, which is usually integrated in a single chip in a digital control chip. , and Respectively represent the input voltage, output voltage and reference voltage of the DC-DC switching power supply. is the duty cycle output of the digital controller, representing the duty cycle signal, After passing through DPWM, a switching signal for controlling the switch is generated.

[0031] According to the relative size relationship of the voltage value before and after the DC-DC conversion, the power stage circuit of the DC-DC switching power supply has three basic topologies: Buck (step-down type), Boost (step-up type) and Buck-Boost (step-up and step-down type). Figure 4 The basic components of the power stage circuit are capacitors. C ,inductance L ,diode and power switch tube S , and the difference between the three topological structures is L , C , and S The positions of the three DC-DC switching power supplies are different. All of these three DC-DC switching power supplies adjust the output voltage by controlling the on-time (duty cycle) or on-cycle of the power switch tube S. Figure 4 a in the figure is a schematic diagram of Buck topology. Figure 4 b in the figure is a schematic diagram of the Boost topology. Figure 4 The c in FIG. is a schematic diagram of a Buck-Boost topology structure.

[0032] Taking the Buck type DC-DC switching power supply topology as an example, its working principle is described as follows: Figure 5As shown. For Buck type DC-DC switching power supply, is an output cycle of The duty cycle signal, in one switching cycle, the duration of the high level is The duration of low level is The duty cycle signal of period, S Conductivity, Cut off, at this time the input voltage passes L right C Charging is performed while part of the energy is transferred to the load . In the switching signal period, S Disconnect, at this time L The voltage across the two ends is reversed, so In forward bias and conduction, L and C The energy stored in the At a certain switching frequency, the above two working states are repeated continuously, thus providing power to the load. Providing an approximately constant supply voltage makes the output voltage approximately equal to the product of the steady-state duty cycle and the input voltage, so the Buck type DC-DC switching power supply has a step-down function. Defining the switching cycle of a DC-DC switching power supply and duty cycle for:

[0033] In the no-load stable stage, the control signal The duty cycle of the output tends to a constant value, that is, and The output voltage of the Buck DC-DC switching power supply is approximately equal to the input voltage multiplied by the duty cycle. Therefore, by changing the duty cycle The output voltage of the power supply can be set. During the operation of the DC-DC switching power supply, when the output voltage deviates from the set value due to load fluctuation or other reasons, the output voltage May deviate from the set At this time, it is necessary to dynamically adjust the duty cycle through the negative feedback system, so that the output voltage is finally stabilized at its set value, thereby providing a stable output voltage for the load.

[0034] by Figure 2 Taking the training of the neural network controller for the Buck circuit as an example, the voltage of the input power supply is a constant value. , the series of frequencies output by the neural network is fThe duty cycle of the switching signal can be modulated The power switch is turned on and off to complete the voltage change from arrive In order to maintain the output voltage under different loads and working conditions The output should be as stable as possible. ADC will The voltage is collected to form a digital control loop, and the neural network is used as its controller. Value and Subtract to get the current voltage error .

[0035] Will Figure 3 The digital compensator in the negative feedback loop is replaced by a trained digital switching power supply neural network controller to output a current duty cycle , and then the digital pulse width modulator generates a duty cycle according to the output of the current neural network. The switching signal is based on the Buck type structure. Figure 5 shown.

[0036] The voltage error The value of The integral and differential values ​​of are used as the output vector of the observer , which is used as the input vector of the neural network controller - the executor network. At the same time, the reward function by As an input variable, it quantifies the stability of the output voltage of the digital switching power supply at time t and feeds back a specific value. After calculation by the executor network, the output duty cycle That is, the proportion of the power switch conduction time to the switching cycle. Duty cycle signal After modulation by the digital pulse width modulator DPWM, a switching signal is formed and output to the power switch tube to adjust the ratio of the on-time and off-time of the power switch tube so that the output voltage is adjusted.

[0037] In a specific embodiment, the specific training objectives are: 1. The controller can ensure that the digital switching power supply will input power Accurately convert to The reference standard voltage.

[0038] 2. During the startup phase of the digital switching power supply, the voltage rises steadily from 0V to the target voltage , no large voltage overshoot occurs.

[0039] 3. When the load of the digital power supply changes within the rated output range, the output voltage can be quickly restored to .

[0040] Figure 2 In the reinforcement learning training simulation environment shown in the figure, reinforcement learning training is implemented, in which the control of the digital switching power supply requires a series of duty cycle signals To continuously adjust the output voltage, maintain stable voltage output, suppress and restore output voltage fluctuations caused by external factors such as load changes or input voltage, T is a specified reinforcement learning training cycle. T In the time period, this process can be regarded as a global optimal sequential decision problem. t to t + T During the training process, the overall voltage output stability of the digital switching power supply under the control of the neural network controller, i.e., the executor network It can be described as follows, where the discount factor ∈[0,1]:

[0041] like Figure 2 As shown in the figure, the actor network in the reinforcement learning agent needs to be trained to eventually serve as the digital switching power supply neural network controller. , where the input vector As the input of the network, the output behavior vector As the output of this network, the network represents the current control strategy The training objective of reinforcement learning can be described as T ,for The network updates its parameters so that it T A series of duty cycle signals output from the Performance values ​​of digital switching power supplies under control Increase, in the process The increase in the control strategy Evolution. The value can be defined as an expectation function, which describes the control strategy Next, in the state Lower Output The overall expected return after:

[0042] This function is composed of another neural network - the critic network To approximate the representation, during the training process, the critic network Through multiple training cyclesT Constantly self-update parameters to make it more reliable for the executor network The ability to judge is becoming more and more accurate, and its network of executors is helping Evaluate its output Is it optimal, thereby improving the training efficiency of the executor network. Figure 6 Shown by and The structure consisting of two neural networks is called the actor-critic network structure.

[0043] In the reinforcement learning process, the executor network serves as the main training object in the reinforcement learning agent. After the training, the parameters are finalized and it will be used alone as a neural network controller in the digital switching power supply. The critic network assists the executor network in quickly updating the network parameters during the training interaction. Its mission is completed after the training.

[0044] During the training process, the deep deterministic policy gradient (DDPG) algorithm is used to train the actor-critic network structure. The pseudo code is as follows:

[0045] In steps S2 to S5, step S2: based on the reinforcement learning agent and the power-level circuit model, a symmetric Gaussian reward function SGRF is used to initialize the value of the SGRF parameter, and the executor network and the critic network are subjected to the first reinforcement learning training through the DDPG algorithm, and the reward function parameters are gradually reduced until convergence; step S3: based on the network parameters of the executor-critic network trained by the first reinforcement learning, an asymmetric Gaussian reward function AGRF is used to initialize the AGRF parameters, and the executor-critic network is subjected to the second reinforcement learning training through the DDPG algorithm, and the value of the AGRF parameter is adjusted to eliminate the voltage overshoot in the startup phase; step S4: based on the network parameters of the executor-critic network trained by the second reinforcement learning, a load mutation event is introduced into the power-level circuit model, an asymmetric Gaussian reward function AGRF is used, and the SGRF parameters and the AGRF parameters are fixed, and the executor-critic network is subjected to the third reinforcement learning training through the DDPG algorithm until the training is completed; step S5: the network parameters of the executor network in the executor-critic network are extracted as the neural network controller of the digital switching power supply, which is used to control the stability of the output voltage in real time.

[0046] Specifically, Figure 7 As shown, the training process is divided into three stages, corresponding to the method proposed in this application: in the first stage, after multiple trials, the minimum value that can make the digital switching power supply output smoothly is found. In the second stage, after many trials, we found the value that can eliminate the overshoot voltage at the startup stage of the digital switching power supply. value; in the third stage, the sudden change of load is added to the training event set, so that the transient performance of the controller for sudden change of load is enhanced. The specific training process is as follows: The first stage of training (i.e. the first reinforcement learning training): [1] Critic-executor network initialization, reward function is SGRF, initial =1; [2] Execute the digital switching power supply startup event during simulation; [3] Use the DDPG algorithm to perform reinforcement learning training in the simulation process; [4] Observe the output voltage of the digital switching power supply under the control of the neural network training process. , after multiple training cycles, can it be stabilized at , and use this to determine whether the training has converged. During the training process, the controller controls Finally stabilized at It is considered convergent, otherwise it is considered non-convergent; [5] The value is reduced by 0.05, and steps [3] and [4] are repeated to observe whether the DDPG algorithm converges. If it does not converge, take the last convergent The value is the best performance value, proceed directly to step [6]; if convergence, it is considered There is room for the value to continue to decrease, and step [5] is repeated until it fails to converge.

[0047] The second stage of reinforcement learning training (i.e. the second reinforcement learning training): [6] Import the critic-executor network parameters trained in the previous stage into the new training environment, and use the reward function as AGRF. The value continues the first stage and the final selection has the best performance value, The value is initially set to 0; [7] Execute the digital switching power supply startup event during simulation; [8] Use the DDPG algorithm to perform reinforcement learning training in simulation; [9] Observe the output voltage of the digital switching power supply under the control of the neural network training process. Is there any voltage overshoot during the startup phase?

[10] If the voltage overshoot still exists, The value is increased by 0.5, and steps [8] and [9] are repeated. If the overshoot no longer exists and the voltage output rises steadily during the startup phase, proceed to step

[11] . The third stage of reinforcement learning training (i.e. the third reinforcement learning training):

[11] The critic-executor network parameters trained in the previous stage are imported into the new training environment, and the reward function is AGRF. Value and The value is the final value of the previous stage of training;

[12] During the simulation process, the digital switching power supply startup event and load jump event are executed;

[13] used the DDPG algorithm to perform reinforcement learning training in the simulation process;

[14] After the training is completed, the performer network is extracted separately and used as a neural network controller as the digital controller of the digital switching power supply.

[0048] Furthermore, some terms in the above training process description are explained: 1. Gaussian reward function: exist Figure 2 In the training environment shown, the reward function is used to evaluate the output performance of the current digital switching power supply, the output voltage and The smaller the gap, the better the current output voltage performance. The evaluation of the reward function provides the reinforcement learning training process with an evaluation standard and a direction for evolution through training. This application proposes to use a Gaussian function with parameters as a reward function, which is divided into a symmetric Gaussian reward function SGRF (Symmetric Gaussian reward function) and an asymmetric Gaussian reward function AGRF (Asymmetric Gaussian reward function), and its expression is as follows:

[0049]

[0050] SGRF and AGRF The smaller the parameter is, the stricter the reward function is and the better the performance of the trained controller is. The larger the parameter, the more severe the penalty imposed by the reward function on the overshoot of the output voltage of the digital switching power supply, and the trained controller will suppress the voltage overshoot more strongly.

[0051] 2. Training event set: Since reinforcement learning is completed in the process of interaction between the neural network controller and the controlled object, that is, the power stage circuit of the digital switching power supply, the designer of the neural network will set the events to occur in the power stage circuit according to the training objectives during the training process. The training events involved in this application include two, one is "startup" and the other is "load jump".

[0052] a) Startup: The startup process refers to the transition of the switching power supply from the non-working state to the working state. In the non-working state, the input voltage of the switching power supply is maintained at , the output voltage is 0V, the neural network controller has no output initially, and the digital pulse width modulator DPWM and the analog-to-digital converter ADC do not work. From the start time 0, the neural network controller, DPWM and ADC all start to work according to their own inherent output frequency. Stable to When there is no obvious fluctuation nearby, the startup event ends.

[0053] b) Load jump: After the output voltage stabilizes, the load jumps from no-load to heavy-load, maintains the heavy-load for a fixed period of time, and then jumps from heavy-load to no-load.

[0054] The digital switching power supply power stage circuit in the above training basic environment takes the Buck type as an example. Figure 4 To train other types of switching power supply power stage circuits, only the positions of the rectifier diode, power switch tube, inductor and capacitor need to be changed, and the neural network controller, its control loop and neural network training environment remain unchanged.

[0055] In a specific embodiment, the specific environment settings used in this example are: 1. Build a training environment: like Figure 2 The figure shows a reinforcement learning training environment with Buck as the controlled object. The entire environment is built in the Simulink simulation environment. The intelligent agent Agent continuously evolves itself in the interaction with the Buck switching power supply power stage circuit to improve the critic network and the executor network. The agent is implemented using the reinforcement learning toolbox in Simulink.

[0056] The reinforcement learning training environment built in Simulink with Buck type switching power supply as the controlled object is as follows: Figure 8 shown.

[0057] Buck type digital switching power supply, such as Fig. 9 shown. Fig. 9 The specific parameters of the Buck switching power supply are shown in Table 1. The training goal is to reduce the 5V input voltage to 2.5V output.

[0058] Table 1 Parameters of Buck switch circuit

[0059] 2. Set the training random variables required in the environment In order to enhance the robustness and generalization ability of the trained neural network controller, random variables must be added during training. The range of variables in the two stages of training is shown in Table 2 and Table 3.

[0060] Table 2 The range of changes in random variables required during training

[0061] Table 3. Hyperparameters and adjustment ranges required during training

[0062] 3. Neural network size and training parameters: The two neural networks used in the training are the executor network and the critic network. The network structure and number of nodes are as follows: Figure 6 shown.

[0063] IV. Interim Results of the Training Process 1. The first training and its results The reward function SGRF used in the first training is:

[0064] Depend on =1 to start training, gradually reduce Value, final <0.1, the training will no longer converge. Fig.10 As shown, different The more concentrated the reward value distribution is in the middle, the better the training effect is. However, correspondingly, Fig.11 As shown, different The number of training rounds and the changing trend of the training rewards under the value. The more training rounds, the more difficult it is to increase the total reward during training. More rounds are needed to achieve stable convergence. When the total reward is stable above 8.5K, it is considered convergence. Fig.12 As shown, it demonstrates the use of different The output voltage waveform of the controller trained by the reward function of the value during the startup process of the Buck switching power supply shows The smaller it is, the higher the control quality, the shorter the stabilization time, and the smaller the steady-state error. According to the training situation, the best result selected in the first training is Value is =0.1.

[0065] Use save("save path / file name.mat", "agent") in Matlab to save the results of the previous training, and then use the command load("save path / file name.mat", "agent") to load it at the beginning of the second training, so that the second training can start from the results of the first training.

[0066] 2. Second training and its results Reward function used for training: AGRF (7) Fig.13 Shows the difference The reward value distribution of the reward function under the parameter value. The higher the asymmetry of the reward value, the stronger the overshoot suppression for voltage startup. Fig.14 Demonstrated the use of the same Value, different The output voltage waveform of the controller trained by the reward function under the value of during the startup of the Buck type switching power supply shows The larger the value, the smaller the startup overshoot, the smoother the transient response during startup, and the safer the startup. When the value is >7, there is no significant difference in the training effect.

[0067] Use save("save path / file name.mat", "agent") in Matlab to save the results of the previous training, and then use the command load("save path / file name.mat", "agent") to load it at the beginning of the second training, so that the second training can start from the results of the first training.

[0068] 3. The third training and its results In the third training, the reward function is the same as the second, but the load change is added in the simulation. The load used in this case is a constant power load. The load mutation is added in the third training. The load change value is random, and the range is shown in Table 3 to improve the transient response to load changes. Similarly, use save ("save path / file name.mat", "agent") in matlab; save the results of the second training, and then use the command load ("save path / file name.mat", "agent") at the beginning of the third training to load it, so that the third training can start with the results of the second training.

[0069] After three training sessions, the output effect of the switching power supply during the startup phase changes as follows: Fig.15 shown.

[0070] V. Deployment of Neural Network Controller After training, the network weights and bias values ​​of the executor network in the reinforcement learning agent in the simulation environment are extracted separately. The network structure is rebuilt in the FPGA unit or digital IC as a digital controller for controlling the switching power supply in reality. At this point, the digital controller design is completed. Fig.16 Shown is a schematic diagram of the trained actor network being taken out separately and implemented in a real circuit control environment.

[0071] 6. Test verification results: like Fig.17 As shown, the executor network is taken out and implemented in FPGA to form a test and verification system.

[0072] Fig.18 Demonstrates the switching power supply in Equal to 5V, The output voltage of the load is 2.7V, 2.2V and 1.9V respectively. Output status. The blue waveform is the load current at this time , unified as 1A, the purple waveform is the current switching signal, and its frequency is 2MHz. According to the test results, When the output voltages are 2.7V, 2.2V and 1.9V respectively, the output voltages represented by the green waveforms are 2.65V, 2.14V and 1.83V respectively, the steady-state errors are 50mV, 60mV and 70mV respectively, and the output ripples are 250mV, 220mV and 200mV respectively. The test results show that That is, when the reference voltage changes, the output voltage can be adjusted to the corresponding voltage level.

[0073] Fig.19 Shows the switching power supply input voltage When the output voltage changes during operation The purple waveform is the input voltage The green waveform is The blue waveform is the load current at this time. The results show that When a jump occurs, There were no major fluctuations and we were fully able to adapt.

[0074] Fig. 20 Shows the switching power supply input voltage When starting from 0V, the output voltage The purple waveform is the input voltage The green waveform is The results show that during the startup phase of the switching power supply, Follow The rise is smooth and there is no large overshoot voltage, which is in line with expectations.

[0075] Fig.21 The output voltage is shown when the switching power supply load changes from 0.5W to 3W, and from 3W to 0.5W. Dynamic waveform to test the load change for output voltage This test is conducted under three different circuit component parameter combinations to verify the impact of load jump on the output voltage when the key components in the DC-DC switching power supply power stage circuit, capacitors and inductors drift. The blue waveform is the load current at this time. The green waveform is The results show that under standard capacitance and inductance, C =2.2μF, L =2.2μH, when the load current jumps by 1A, the overshoot of the output voltage does not exceed 70mV, and the stabilization time is less than 16μs. C =1μF, L =2.2μH, that is, when the capacitance value decreases, the load current jumps by 1A, the overshoot of the output voltage does not exceed 220mV, and the stabilization time is less than 110μs. C =2.2μF, L =4.7μH, that is, when the inductance value increases, the load current jumps by 1A, the overshoot of the output voltage does not exceed 130mV, and the stabilization time is less than 50μs. C =4.7μF, L =4.7μH, that is, when both the capacitance and inductance become larger, the load current jumps by 1A, the overshoot and overshoot of the output voltage do not exceed 240mV, and the stabilization time is less than 40μs.

[0076] The above test results show that the neural network controller trained by the training method proposed in this application can make the feedback loop of the DC-DC switching power supply more adaptive and robust.

[0077] Through the above-mentioned reinforcement learning-based training method for the neural network controller of the digital switching power supply, on the one hand, for the design of the neural network controller of the digital switching power supply, there is no need for mathematical modeling, and excellent control effects can be achieved only by designing and controlling the training process. The Gaussian reward function based on adjustable parameters can be used to quantitatively adjust the control performance and voltage overshoot, and design a controller that meets the performance indicators. On the other hand, the method automatically explores the optimal control strategy through a trial-and-error mechanism, without manually designing complex control laws or adjusting PID parameters. It can find the global optimal solution in a complex nonlinear system, avoiding the local optimal problem that may exist in traditional design methods. DC-DC converters are usually faced with dynamic changes such as load transients and input voltage fluctuations. Reinforcement learning can adjust the control strategy in real time to adapt to these changes, significantly improving the robustness and adaptability of the system. It can also show better fault tolerance for parameter perturbations (such as component aging). The method is data-driven and can work effectively even when the model is not completely clear. It reduces the tedious design and debugging process and shortens the development cycle. The method can quickly adapt to new working scenarios or performance requirements by updating the strategy network. Further improvements can be made based on this method by increasing other controller performance (such as predictive control and self-healing capabilities) by adding training events during the training process.

[0078] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.

[0079] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.

Claims

1. A training method for a digital switching power supply neural network controller based on reinforcement learning, characterized in that: The method includes: Step S1, establishing a power level circuit model and a reinforcement learning agent of a digital switching power supply in a simulation environment; wherein the power level circuit model includes an input power supply, a switch tube, a filter circuit and an output end, and the reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network, wherein the actor network is used to receive a state vector output by an observer, and the critic network is used to evaluate a control strategy output by the actor network; Step S2, based on the reinforcement learning agent and the power level circuit model, a symmetric Gaussian reward function SGRF is used to initialize the value of the SGRF parameter, and the executor network and the critic network are trained for the first reinforcement learning by the DDPG algorithm, and the reward function parameters are gradually reduced until the training does not converge; Step S3, based on the network parameters of the actor-critic network trained by the first reinforcement learning, an asymmetric Gaussian reward function AGRF is used to initialize the AGRF parameters, and the actor-critic network is trained by the second reinforcement learning through the DDPG algorithm, and the values ​​of the AGRF parameters are adjusted to eliminate the voltage overshoot in the startup phase; Step S4, based on the network parameters of the actor-critic network trained by the second reinforcement learning, introduce a load mutation event in the power level circuit model, use an asymmetric Gaussian reward function AGRF, and fix the SGRF parameters and AGRF parameters, and perform a third reinforcement learning training on the actor-critic network through the DDPG algorithm until the training is completed; Step S5, extracting the network parameters of the actor network in the actor-critic network to construct a neural network controller of the digital switching power supply to control the stability of the output voltage in real time.

2. According to the reinforcement learning-based digital switching power supply neural network controller training method of claim 1, it is characterized in that: The state vector includes the output voltage error, the integral error and the differential error, and outputs the duty cycle signal to the power stage circuit model.

3. According to the reinforcement learning-based digital switching power supply neural network controller training method of claim 1, it is characterized in that: Step S2 specifically includes: Step S21, initialize the actor-critic network, use the symmetric Gaussian reward function SGRF, and initialize the SGRF parameters The value of is 1; Step S22, executing a digital switch power supply startup event; Step S23, using the DDPG algorithm to perform the first reinforcement learning training on the actor-critic network; Step S24, if the output voltage of the digital switching power supply under the control of the controller during the training of the executor network and the critic network Stable , then the DDPG algorithm converges; Step S25, set the SGRF parameters Reduce by 0.05, repeat step S23 and step S24, and determine whether the DDPG algorithm converges; Step S26: If the DDPG algorithm does not converge, take the last SGRF parameter that can converge The value of is taken as the SGRF parameter with the best performance The value of , and the first reinforcement learning training ends; if the DDPG algorithm converges, step S25 is repeated until the DDPG algorithm does not converge.

4. According to claim 3, the training method of a digital switching power supply neural network controller based on reinforcement learning is characterized in that: Step S3 specifically includes: Step S31, import the network parameters of the actor-critic network trained by the first reinforcement learning into the new training environment, and use the asymmetric Gaussian reward function AGRF, SGRF parameters The value of continues the SGRF parameter with the best performance in step S26 The value of AGRF parameter The value of is initially set to 0; Step S32, executing a digital switch power supply startup event; Step S33, using the DDPG algorithm to perform a second reinforcement learning training on the actor-critic network; Step S34, determining the output voltage of the digital switching power supply under the control of the controller during the training of the performer network and the critic network Is there a voltage overshoot? If there is a voltage overshoot, set the SGRF parameter The value of is increased by 0.5, and steps S33 and S34 are repeated; if there is no voltage overshoot, the voltage output rises steadily during the startup phase, and the next step is entered.

5. According to claim 4, the training method for a digital switching power supply neural network controller based on reinforcement learning is characterized in that: Step S4 specifically includes: Step S41, import the network parameters of the actor-critic network trained by the second reinforcement learning into the new training environment, and use the asymmetric Gaussian reward function AGRF, SGRF parameters The values ​​and AGRF parameters The value of takes the final value of the second reinforcement learning; Step S42, executing a digital switch power supply startup event and a load mutation event; Step S43, using the DDPG algorithm to perform a third reinforcement learning training on the actor-critic network until the training is completed.

6. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 5 is characterized in that: The expression of the symmetric Gaussian reward function SGRF is: in, is the SGRF parameter, is the output voltage, is the ideal output reference value of the output voltage; The expression of asymmetric Gaussian reward function AGRF is: in, is the AGRF parameter.

7. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 6 is characterized in that: The DDPG algorithm consists of the following steps: Initialize the network parameters of the actor-critic network; Storing training data through the experience replay pool; Iteratively update the network parameters of the actor-critic network to maximize the cumulative reward value, and synchronously update the network parameters of the actor-critic network.

8. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 7 is characterized in that: According to the power stage circuit model, the analog-to-digital converter in the feedback control loop, the trained digital switching power supply neural network controller and the digital pulse width modulator, a digital controlled DC-DC switching power supply is formed; wherein, The power stage circuit model turns on and off the power switch tube according to the switch signal of the previous duty cycle to adjust the input voltage Perform chopping; the chopped input voltage The filtering by capacitor C and inductor L forms a voltage lower than the input voltage. Output voltage , to complete the voltage from the input voltage To output voltage of blood pressure reduction; The analog-to-digital converter outputs a voltage Convert to value; and Value and reference voltage value Subtract, output voltage error ; The voltage error The value of The points and The differential value of is used as the output vector of the observer , which is used as the input vector of the trained digital switching power supply neural network controller, and the reward function by As input variable; input vector After calculation by the executor network, the duty cycle signal is output That is, the proportion of the power switch tube conduction time to the switching cycle; Digital pulse width modulator for duty cycle signal Modulation is performed to form a switching signal, which is then output to the power switch tube in the power stage circuit model to adjust the ratio of the on-time and off-time of the power switch tube so that the output voltage at the current moment is is adjusted.

9. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 8 is characterized in that: Load mutation events include: After the output voltage stabilizes, the load changes from no-load to heavy-load and maintains for a preset time, and then changes back to no-load.

10. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 9 is characterized in that: Switching power supply startup events include: The startup process refers to the transition of the switching power supply from the non-working state 0V output to the working state.

Citation Information

Patent Citations

  • Secondary frequency modulation control method and system for distributed energy storage system

    CN111224433A

  • Digital DC-DC converter loop control method, system, device and medium

    CN116345898A

  • Modulation strategy design method and system of DC-DC converter based on reinforcement learning

    CN117318480A

  • Dynamic response optimization method and device, computer equipment, storage medium and product

    CN118554761A

  • Model-free micro-grid control method and system based on transfer learning and medium

    CN119275824A