Training Method for Neural Network Controller of Digital Switching Power Supply Based on Reinforcement Learning

By using reinforcement learning training methods in the simulation environment of digital switching power supply and using Gaussian reward function for multiple training, the problem of random training results of digital switching power supply is solved, and excellent control effect and system robustness are achieved.

CN120010272BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510482196.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-06-27
Estimated Expiration
2045-04-17

AI Technical Summary

Technical Problem

The existing reinforcement learning training process has relatively random training results for digital switching power supplies. The final performance of the controller cannot be well controlled, and whether the performance is optimal cannot be guaranteed.

Method used

Using the digital switching power supply neural network controller training method based on reinforcement learning, the power-level circuit model and reinforcement learning agent of the digital switching power supply are established in the simulation environment, and multiple reinforcement learning training are performed using symmetric Gaussian reward functions and asymmetric Gaussian reward functions, and the reward function parameters are gradually adjusted to optimize the control strategy.

Benefits of technology

It realizes excellent control effect on digital switching power supplies, and does not require mathematical modeling, and can find global optimal solutions in complex nonlinear systems, improving the robustness and adaptability of the system, reducing the cumbersome process of design and debugging, and shortening the development cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010272B_ABST
    Figure CN120010272B_ABST
Patent Text Reader

Abstract

This application belongs to the technical field of reinforcement learning. This application provides a method for training a neural network controller for a digital switching power supply based on reinforcement learning. In the embodiments of the present disclosure, for the design of a neural network controller for a digital switching power supply, mathematical modeling is not required, and excellent control effects can be achieved only by designing and controlling the training process. By using a Gaussian reward function based on adjustable parameters, the control performance and voltage overshoot can be quantitatively adjusted to design a controller that meets the performance indicators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the technical field of reinforcement learning, and particularly to a method for training a neural network controller of a digital switching power supply based on reinforcement learning. Background Art

[0002] With the rapid development of integrated circuits and information technology, the functions of electronic systems are becoming increasingly powerful and the circuit scale is continuously growing, resulting in a sharp increase in the consumed current and frequent load mutations, which poses higher requirements for the performance of the power supply. Compared with traditional analog-controlled switching power supplies, digital-controlled switching power supplies (or digital-controlled DC-DC switching converters, simply referred to as digital power supplies) have the advantages of being programmable, reconfigurable, flexible in design, capable of executing various advanced and complex control algorithms, capable of realizing multi-phase or multi-channel simultaneous output, good system control robustness, and being able to communicate with the host in real time. They have obvious advantages in improving the steady-state and transient performance of the power supply, increasing the power density and conversion efficiency, reducing the manufacturing cost, integrating components, and shortening the product development cycle. Their application scope is gradually expanding and they have broad application prospects in fields such as consumer electronics, industrial electronics, cloud computing and big data centers, and military aerospace electronics.

[0003] Currently, digital power products generally adopt traditional linear (such as PID control) or nonlinear (such as fuzzy logic control, sliding mode control) control methods. Since these traditional control algorithms are all optimization controls based on the linear time-invariant small-signal model or "expert experience" of the system, for complex nonlinear time-varying systems such as switching power supplies, it is difficult to further improve the performance of the power supply by using traditional linear or nonlinear control methods. Since neural networks have the advantages of a highly parallel structure, powerful learning ability, approximation ability for any function, and strong fault tolerance, they are particularly suitable for complex nonlinear time-varying systems such as switching power supplies. By training the neural network model, optimal control of the switching power supply system can be achieved. Therefore, intelligent control of digital power based on neural networks is an inevitable development trend of digital power and has received extensive attention from the academic and industrial circles in recent years and has become a research and development hotspot in the current power field.

[0004] In the related art, there are mainly two methods for training neural network control in the field of digital switching power supplies at present, namely supervised learning and reinforcement learning. The training process of supervised learning requires a large amount of data collection and labeling, and relies on continuously reducing the difference between the output of the neural network controller and the calibrated data to optimize the performance of the controller. However, how to collect high-quality data consistent with the current digital power supply parameters and working conditions is a difficult point. Reinforcement learning, as a method for training neural network controllers of digital switching power supplies, has received attention because its training process neither requires small-signal modeling of the switching converter nor data labeling. Reinforcement learning allows the neural network controller to be trained during the interaction with the controlled circuit in a simulation working environment. However, there are also problems with the existing reinforcement learning training process for digital switching power supplies. The training results of reinforcement learning for switching converters are relatively random, the final performance of the controller cannot be well controlled, and it cannot be guaranteed whether its performance reaches the optimal state.

[0005] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.

[0006] It should be noted that this part aims to provide background or context for the technical solutions of the present disclosure stated in the claims. The description herein is not admitted to be prior art merely because it is included in this part. Summary of the Invention

[0007] The purpose of the embodiments of the present disclosure is to provide a method for training a neural network controller of a digital switching power supply based on reinforcement learning, thereby at least to a certain extent overcoming one or more problems caused by the limitations and defects of the related art.

[0008] According to the embodiments of the present disclosure, there is provided a method for training a neural network controller of a digital switching power supply based on reinforcement learning, the method comprising:

[0009] Step S1, establishing a power stage circuit model of a digital switching power supply and a reinforcement learning agent in a simulation environment; wherein, the power stage circuit model includes an input power supply, a switching tube, a filter circuit and an output terminal, and the reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network. The actor network is used to receive the state vector output by the observer, and the critic network is used to evaluate the control strategy output by the actor network;

[0010] Step S2, based on the reinforcement learning agent and the power stage circuit model, adopting a symmetric Gaussian reward function SGRF, initializing the value of the SGRF parameter, and performing the first reinforcement learning training on the actor network and the critic network through the DDPG algorithm, gradually reducing the reward function parameter until the training does not converge;

[0011] Step S3: Based on the network parameters of the actor-critic network trained by the first reinforcement learning, use the Asymmetric Gaussian Reward Function (AGRF) to initialize the AGRF parameters. Then, through the Deep Deterministic Policy Gradient (DDPG) algorithm, perform the second reinforcement learning training on the actor-critic network, and adjust the value of the AGRF parameters to eliminate the voltage overshoot in the startup phase.

[0012] Step S4: Based on the network parameters of the actor-critic network trained by the second reinforcement learning, introduce a load mutation event into the power stage circuit model. Use the Asymmetric Gaussian Reward Function (AGRF), and fix the parameters of the Symmetric Gaussian Reward Function (SGRF) and the AGRF. Then, through the DDPG algorithm, perform the third reinforcement learning training on the actor-critic network until the training is completed.

[0013] Step S5: Extract the network parameters of the actor network in the actor-critic network to construct the neural network controller of the digital switching power supply, so as to control the stability of the output voltage in real time.

[0014] Furthermore, the state vector includes the output voltage error, error integral, and error differential value, and outputs a duty cycle signal to the power stage circuit model.

[0015] Furthermore, Step S2 specifically includes:

[0016] Step S21: Initialize the actor-critic network, use the Symmetric Gaussian Reward Function (SGRF), and initialize the value of the SGRF parameter to 1. The value is 1.

[0017] Step S22: Execute the startup event of the digital switching power supply.

[0018] Step S23: Use the DDPG algorithm to perform the first reinforcement learning training on the actor-critic network.

[0019] Step S24: If the output voltage controlled by the controller during the training of the actor network and the critic network of the digital switching power supply stabilizes at, then the DDPG algorithm converges. Stabilizes at , then the DDPG algorithm converges.

[0020] Step S25: Decrease the SGRF parameter by 0.05, and repeat Step S23 and Step S24, and determine whether the DDPG algorithm converges. Decrease the SGRF parameter by 0.05, re-perform Step S23 and Step S24, and determine whether the DDPG algorithm converges.

[0021] Step S26: If the DDPG algorithm does not converge, then take the value of the previous SGRF parameter that can converge as the SGRF parameter with the optimal performance, and end the first reinforcement learning training; if the DDPG algorithm converges, then repeat Step S25 until the DDPG algorithm does not converge. Take the value of the previous SGRF parameter that can converge as the SGRF parameter with the optimal performance And end the first reinforcement learning training; if the DDPG algorithm converges, then repeat Step S25 until the DDPG algorithm does not converge.

[0022] Further, step S3 specifically includes:

[0023] Step S31: Import the network parameters of the actor-critic network trained by the first reinforcement learning into a new training environment. Adopt the asymmetric Gaussian reward function AGRF, and the value of the SGRF parameter continues the value of the SGRF parameter with the optimal performance in step S26, and the value of the AGRF parameter is initially set to 0; is initially set to 0;

[0024] Step S32: Execute the digital switched-mode power supply startup event;

[0025] Step S33: Use the DDPG algorithm to perform the second reinforcement learning training on the actor-critic network;

[0026] Step S34: Determine whether there is voltage overshoot in the output voltage controlled by the controller during the training of the actor network and the critic network of the digital switched-mode power supply ;

[0027] If there is voltage overshoot, increase the value of the SGRF parameter by 0.5, and re-perform step S33 and step S34; if there is no voltage overshoot, the voltage output during the startup phase rises steadily, and proceed to the next step.

[0028] Further, step S4 specifically includes:

[0029] Step S41: Import the network parameters of the actor-critic network trained by the second reinforcement learning into a new training environment. Adopt the asymmetric Gaussian reward function AGRF, and the value of the SGRF parameter and the value of the AGRF parameter take the final values of the second reinforcement learning;

[0030] Step S42: Execute the digital switched-mode power supply startup event and the load mutation event;

[0031] Step S43: Use the DDPG algorithm to perform the third reinforcement learning training on the actor-critic network until the training is completed.

[0032] Further, the expression of the symmetric Gaussian reward function SGRF is:

[0033]

[0034] where is the AGRF parameter, is the output voltage, is the ideal output reference value of the output voltage;

[0035] The expression of the Asymmetric Gaussian Reward Function (AGRF) is as follows:

[0036]

[0037] where, are the AGRF parameters.

[0038] Furthermore, the DDPG algorithm includes the following steps:

[0039] Initialize the network parameters of the actor-critic network;

[0040] Store the training data through the experience replay pool;

[0041] Iteratively update the network parameters of the actor-critic network to maximize the cumulative reward value, and synchronously update the network parameters of the actor-critic network.

[0042] Furthermore, a digital controlled DC-DC switching power supply is composed of a power stage circuit model, an analog-to-digital converter in the feedback control loop, a trained digital switching power supply neural network controller, and a digital pulse width modulator; where,

[0043] The power stage circuit model turns on and off the power switch tube according to the switching signal power of the duty cycle at the previous moment, so as to chop the input voltage The chopped input voltage forms an output voltage lower than the input voltage through the filtering of the capacitor C and the inductor L, so as to complete the step-down of the voltage from the input voltage to the output voltage ;

[0044] The analog-to-digital converter converts the output voltage into value; and subtracts the value from the reference voltage value to output the voltage error ;

[0045] The value of the voltage error , the integral of , and the differential value of are used as the output vector of the observer, and this vector is used as the input vector of the trained digital switching power supply neural network controller. The reward function uses as the input variable; the input vector is calculated by the actor network to output the duty cycle signal That is, the time ratio of the conduction time of the power switch tube to the switching period;

[0046] The digital pulse width modulator modulates the duty cycle signal to form a switching signal and output it to the power switch tube in the power stage circuit model to adjust the ratio of the conduction time and the turn-off time of the power switch tube, so that the output voltage at the current moment is adjusted.

[0047] Furthermore, the load mutation event includes:

[0048] After the output voltage is stable, the load jumps from no-load to heavy-load and maintains a preset time, and then jumps back to no-load.

[0049] Furthermore, the switching power supply startup event includes:

[0050] The startup process refers to the transition of the switching power supply from the non-working state with 0V output to the working state.

[0051] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:

[0052] In the embodiments of the present disclosure, through the above-mentioned digital switching power supply neural network controller training method based on reinforcement learning, on the one hand, for the design of the neural network controller of the digital switching power supply, there is no need to perform mathematical modeling, and excellent control effects can be achieved only by using the design and control of the training process. Using the Gaussian reward function based on adjustable parameters can realize the quantitative adjustment of control performance and voltage overshoot, and design a controller that meets the performance indicators. On the other hand, this method automatically explores the optimal control strategy through a trial-and-error mechanism, without manually designing complex control laws or adjusting PID parameters. It can find the global optimal solution in complex non-linear systems, avoiding the possible local optimal problems of traditional design methods. DC-DC converters usually face dynamic changes such as load transients and input voltage fluctuations. Reinforcement learning can adjust the control strategy in real time to adapt to these changes, significantly improving the robustness and adaptability of the system. It can also show better fault tolerance for parameter perturbations. This method is data-driven and can work effectively even when the model is not completely clear. It reduces the cumbersome design and debugging process and shortens the development cycle. This method can quickly adapt to new working scenarios or performance requirements by updating the policy network. Further improvements can be made on the basis of this method, and other performances of the controller can be increased by adding training events during the training process. Description of the Drawings

[0053] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0054] Figure 1 A step diagram showing the method for training a neural network controller of a digital switching power supply based on reinforcement learning in an exemplary embodiment of the present disclosure;

[0055] Figure 2 A schematic diagram showing the training environment of a digital power neural network controller in an exemplary embodiment of the present disclosure;

[0056] Figure 3 A schematic diagram showing the structure of a traditional digital control DC-DC switching power supply in an exemplary embodiment of the present disclosure;

[0057] Figure 4 A schematic diagram showing the basic topological structure of the power stage circuit of a DC-DC switching power supply in an exemplary embodiment of the present disclosure;

[0058] Figure 5 A schematic diagram showing the structure of a DC-DC switching power supply with a neural network controller in an exemplary embodiment of the present disclosure;

[0059] Figure 6 A schematic diagram showing the network structures of the actor network and the critic network in an exemplary embodiment of the present disclosure;

[0060] Figure 7 A schematic diagram showing the training process in an exemplary embodiment of the present disclosure;

[0061] Figure 8 A schematic diagram showing the construction of a reinforcement learning training environment with a Buck-type switching power supply as the controlled object in Simulink in an exemplary embodiment of the present disclosure;

[0062] Figure 9 A schematic diagram showing the construction of a Buck-type digital switching power supply in Simulink in an exemplary embodiment of the present disclosure;

[0063] Figure 10 Showing in an exemplary embodiment of the present disclosure different The distribution of the reward values of the reward function at different values;

[0064] Figure 11 Showing in an exemplary embodiment of the present disclosure different The number of training episodes and the changing trend of the rewards obtained from training at different values;

[0065] Figure 12 Shows the startup situation of the output voltage of the switching power supply at different values in the exemplary embodiments of the present disclosure;

[0066] Figure 13 Shows the distribution of the reward values of the reward function at different values in the exemplary embodiments of the present disclosure;

[0067] Figure 14 Shows the same value and different values of the startup situation of the output voltage of the DC-DC switching power supply in the exemplary embodiments of the present disclosure;

[0068] Figure 15 Shows the output voltage waveform of the DC-DC switching power supply during the startup phase after three trainings in the exemplary embodiments of the present disclosure;

[0069] Figure 16 Shows a schematic diagram of the trained executor network being separately taken out and implemented in a real circuit control environment in the exemplary embodiments of the present disclosure;

[0070] Figure 17 Shows a schematic diagram of the test system formed by implementing and composing the executor network in the FPGA after it is taken out in the exemplary embodiments of the present disclosure;

[0071] Figure 18 Shows the steady-state performance result graph of the DC-DC switching power supply in the exemplary embodiments of the present disclosure;

[0072] Figure 19 Shows the transient performance result graph of the DC-DC switching power supply under the condition of input voltage fluctuation in the exemplary embodiments of the present disclosure;

[0073] Figure 20 Shows the transient performance of the DC-DC switching power supply during startup in the exemplary embodiments of the present disclosure;

[0074] Figure 21 Shows the transient performance of the load change of the DC-DC switching power supply under different combinations of capacitance and inductance in the exemplary embodiments of the present disclosure. Detailed implementation manners

[0075] Now, the exemplary embodiments will be described more comprehensively with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0076] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.

[0077] In this exemplary embodiment, a method for training a neural network controller of a digital switching power supply based on reinforcement learning is provided. Referring to Figure 1 as shown, the method for training a neural network controller of a digital switching power supply based on reinforcement learning may include:

[0078] Step S1, establishing a power stage circuit model of a digital switching power supply and a reinforcement learning agent in a simulation environment; wherein, the power stage circuit model includes an input power supply, a switching tube, a filter circuit and an output terminal, and the reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network. The actor network is used to receive the state vector output by an observer, and the critic network is used to evaluate the control strategy output by the actor network;

[0079] Step S2, based on the reinforcement learning agent and the power stage circuit model, using a symmetric Gaussian reward function SGRF, initializing the value of the SGRF parameter, and performing the first reinforcement learning training on the actor network and the critic network through the DDPG algorithm, gradually reducing the reward function parameter until the training does not converge;

[0080] Step S3, based on the network parameters of the actor-critic network trained in the first reinforcement learning, using an asymmetric Gaussian reward function AGRF, initializing the AGRF parameter, and performing the second reinforcement learning training on the actor-critic network through the DDPG algorithm, adjusting the value of the AGRF parameter to eliminate the voltage overshoot in the startup phase;

[0081] Step S4, based on the network parameters of the actor-critic network trained in the second reinforcement learning, introducing a load mutation event into the power stage circuit model, using the asymmetric Gaussian reward function AGRF, and fixing the SGRF parameter and the AGRF parameter, and performing the third reinforcement learning training on the actor-critic network through the DDPG algorithm until the training is completed;

[0082] Step S5, extracting the network parameters of the actor network in the actor-critic network to construct a neural network controller of the digital switching power supply to real-time control the stability of the output voltage.

[0083] Through the above-mentioned training method of the neural network controller for digital switching power supplies based on reinforcement learning, on the one hand, for the design of the neural network controller for digital switching power supplies, mathematical modeling is not required, and excellent control effects can be achieved only by designing and controlling the training process. By using a Gaussian reward function based on adjustable parameters, the control performance and voltage overshoot can be quantitatively adjusted to design a controller that meets the performance indicators. On the other hand, this method automatically explores the optimal control strategy through a trial-and-error mechanism, without manually designing complex control laws or adjusting PID parameters. It can find the global optimal solution in complex non-linear systems, avoiding the local optimal problems that may exist in traditional design methods. DC-DC converters usually face dynamic changes such as load transients and input voltage fluctuations. Reinforcement learning can adjust the control strategy in real time to adapt to these changes, significantly improving the robustness and adaptability of the system. It can also show better fault tolerance for parameter perturbations (such as component aging). This method is data-driven and can work effectively even when the model is not fully clear. It reduces the cumbersome design and debugging process and shortens the development cycle. This method can quickly adapt to new working scenarios or performance requirements by updating the policy network. Based on this method, further improvements can be made to increase other performance of the controller (such as predictive control and self-healing ability) by adding training events during the training process.

[0084] Next, with reference to Figures 1 to 21 each step of the above-mentioned training method of the neural network controller for digital switching power supplies based on reinforcement learning in the present exemplary embodiment will be described in more detail.

[0085] In step S1, a power stage circuit model of a digital switching power supply and a reinforcement learning agent are established in a simulation environment; wherein, the power stage circuit model includes an input power supply, a switching transistor, a filter circuit, and an output terminal. The reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network. The actor network is used to receive the state vector output by the observer, and the critic network is used to evaluate the control strategy output by the actor network.

[0086] Specifically, the training environment of the present application is constructed in the Simulink component of the Matlab software and is divided into two parts, as Figure 2 shown, one part is the power stage part of the digital switching power supply, and the other part is the reinforcement learning agent including the neural network controller.

[0087] In a specific embodiment, the structure of the digital control DC-DC switching power supply is as Figure 3As shown, it is a typical feedback tracking control system, which consists of a power stage circuit (i.e., a power stage circuit model) and an analog-to-digital converter (ADC), a digital compensator, and a digital pulse width modulator (DPWM) in the feedback control loop. The ADC, digital compensator, and DPWM constitute the digital controller of the DC-DC switching power supply, which is usually monolithically integrated in a digital control chip. Among them , and respectively represent the input voltage, output voltage, and reference voltage of the DC-DC switching power supply. is the duty cycle output of the digital controller, representing the duty cycle signal. After passing through the DPWM, a switching signal for controlling the switch is generated.

[0088] According to the relative magnitude relationship of the voltage values before and after the DC-DC conversion, the power stage circuit of the DC-DC switching power supply has three basic topologies: Buck (step-down), Boost (step-up), and Buck-Boost (step-up / step-down), as Figure 4 shown. The basic components of the power stage circuit are capacitors C , inductors L , diodes , and power switch transistors S . The difference between the three topologies lies in the L , C , , and S 's different positions. All three of these DC-DC switching power supplies regulate the output voltage by controlling the conduction time (duty cycle) or conduction period of the power switch transistor S. Among them, Figure 4 a in is a schematic diagram of the Buck-type topology, Figure 4 b in is a schematic diagram of the Boost-type topology, Figure 4 c in is a schematic diagram of the Buck-Boost-type topology.

[0089] Taking the Buck-type DC-DC switching power supply topology as an example, its working principle is described as Figure 5 shown. For the Buck-type DC-DC switching power supply, is a duty cycle signal with an output period of . Within one switching cycle, the duration of the high level is , and the duration of the low level is . During the of the duty cycle signal , S conducts. Up to this point, the input voltage passes through L to C for charging, and at the same time, part of the energy is transferred to the load . During the of the switching signal, S is disconnected. At this time, the voltage across L reverses, so is forward-biased and conducts. The energy stored in L and C continues to supply power to the load . At a certain switching frequency, the above two operating states are continuously repeated, thereby providing an approximately constant supply voltage for the load , making the output voltage approximately equal to the product of the steady-state duty cycle and the input voltage. Therefore, the Buck-type DC-DC switching power supply has a step-down function.

[0090] Define the switching period and the duty cycle of the DC-DC switching power supply as:

[0091]

[0092] In the no-load stable stage, at this time, the duty cycle output by the control signal tends to a constant value approximately, that is, the ratio of to . Since the output voltage of the Buck-type DC-DC switching power supply is approximately equal to the input voltage multiplied by the duty cycle, therefore, by changing the duty cycle , the magnitude of the output voltage of the power supply can be set. During the operation of the DC-DC switching power supply, when the output voltage deviates from the set value due to load fluctuations or other reasons, the output voltage may deviate from the set . At this time, it is necessary to dynamically adjust the duty cycle through the negative feedback system, and finally make the output voltage stable at its set value, so as to provide a stable output voltage for the load.

[0093] Taking the training of the neural network controller for the Buck-type circuit in Figure 2 as an example, the voltage of the input power supply is a constant value . A series of switching signals with a modifiable duty cycle and a frequency of f are output through the neural network to turn on and off the power switch tube, thereby completing the step-down of the voltage from to . In order to maintain the output voltage as stable as possible at under different loads and operating conditions, the ADC will to Voltage acquisition is performed to form a digital control loop, with a neural network as its controller. After ADC acquisition, the value is subtracted from to obtain the current voltage error .

[0094] The digital compensator in the negative feedback loop in Figure 3 is replaced by a trained digital switching power supply neural network controller, which is used to output a current duty cycle . Then, the digital pulse width modulator generates a switching signal with a duty cycle of based on the output of the current neural network. Its structure is based on the Buck type as shown in Figure 5 .

[0095] The value of the voltage error , the integral of , and the differential value are used as the output vector of the observer. This vector serves as the input vector for the neural network controller - the executor network. At the same time, the reward function uses as the input variable, which quantifies the stability of the output voltage of the digital switching power supply at time t and is fed back as a specific value. The input vector is calculated by the executor network to output the duty cycle , that is, the time ratio of the conduction time of the power switch tube to the switching period. The duty cycle signal is modulated by the digital pulse width modulator DPWM to form a switching signal, which is output to the power switch tube to adjust the ratio of the conduction time and the turn-off time of the power switch tube, so that the output voltage is adjusted.

[0096] In a specific embodiment, the specific training objectives are as follows:

[0097] 1. The controller can ensure that the digital switching power supply accurately converts the input power supply to the reference standard voltage.

[0098] 2. During the start-up phase of the digital switching power supply, it rises smoothly from 0V to the target voltage , without significant voltage overshoot.

[0099] 3. When the load of the digital power supply changes within the output rated range, the output voltage can be quickly adjusted and restored to .

[0100] Figure 2 In the reinforcement learning training simulation environment shown in Continuously adjust the output voltage to maintain a stable voltage output, suppress and recover the output voltage fluctuations caused by external factors such as load changes or input voltage. T is a specified reinforcement learning training cycle. During T this time period, this process can be regarded as a global optimal sequential decision-making problem. At time t to t + T During the training process, the overall voltage output stability of the digital switching power supply under the control of the neural network controller, i.e., the executor network, can be described by the following formula, where the discount factor ∈[0,1]:

[0101]

[0102] As Figure 2 shown, the executor network in the reinforcement learning agent that needs to be trained ultimately as the neural network controller of the digital switching power supply , where the input vector is used as the input of this network, and the output behavior vector is used as the output of this network. This network represents the current control policy . The training objective of reinforcement learning can be described as through multiple training cycles T , for the network to update its parameters, so that the performance value T of the digital switching power supply controlled by a series of duty cycle signals output within the cycle increases. During this process, the increase represents the evolution of the control policy . The expected value can be defined as an expected function, which describes the overall expected benefit after outputting under the control policy in the state :

[0103]

[0104] This function is approximately represented by another neural network - the critic network . During the training process, the critic network constantly updates its parameters through multiple training cycles T , making its judgment ability on the executor network more and more accurate. It assists the executor network to evaluate whether the it outputs is optimal, thereby improving the training efficiency of the executor network. Figure 6The structure shown by and composed of two neural networks is called an actor-critic network structure.

[0105] During the reinforcement learning process, the actor network is the main training object in the reinforcement learning agent. After the training is completed, the parameters are finalized and it will be used alone as a neural network controller in the digital switching power supply. The critic network assists the actor network in quickly updating the network parameters during the training interaction process, and its mission is completed after the training is over.

[0106] During the training process, the DDPG (deep deterministic policy gradient) algorithm, that is, the deep deterministic policy gradient algorithm, is specifically used for the training of the actor-critic network structure. Its pseudo-code execution is as follows:

[0107]

[0108] In steps S2 to S5, step S2: Based on the reinforcement learning agent and the power stage circuit model, the symmetric Gaussian reward function SGRF is adopted to initialize the value of the SGRF parameter. The actor network and the critic network are trained by the DDPG algorithm for the first time of reinforcement learning training, and the reward function parameter is gradually reduced until convergence; step S3: Based on the network parameters of the actor-critic network trained in the first reinforcement learning, the asymmetric Gaussian reward function AGRF is adopted to initialize the AGRF parameter. The actor-critic network is trained by the DDPG algorithm for the second time of reinforcement learning training, and the value of the AGRF parameter is adjusted to eliminate the voltage overshoot in the startup stage; step S4: Based on the network parameters of the actor-critic network trained in the second reinforcement learning, a load mutation event is introduced into the power stage circuit model, the asymmetric Gaussian reward function AGRF is used, and the SGRF parameter and the AGRF parameter are fixed. The actor-critic network is trained by the DDPG algorithm for the third time of reinforcement learning training until the training is completed; step S5: The network parameters of the actor network in the actor-critic network are extracted as the neural network controller of the digital switching power supply to be used for real-time control of the stability of the output voltage.

[0109] Specifically, as Figure 7 shown, the training process is divided into three stages, corresponding to the method proposed in this application: In the first stage, through multiple trial trainings, the smallest value that can make the digital switching power supply output stably is found; in the second stage, through multiple trial trainings, the value that can just eliminate the overshoot voltage in the startup stage of the digital switching power supply is found; in the third stage, by adding load mutations to the training event set, the transient performance of the controller for load mutations is enhanced. The specific training process is as follows:

[0110] The first - stage training (i.e., the first reinforcement learning training):

[0111] [1] Initialize the critic - actor network, with the reward function being SGRF, and the initial = 1;

[0112] [2] Execute the digital switching power supply startup event during the simulation process;

[0113] [3] Use the DDPG algorithm to perform reinforcement learning training during the simulation process;

[0114] [4] Observe the output voltage under the control of the digital switching power supply during the neural network training process. After multiple training cycles, whether it can stabilize at , and judge whether the training converges based on this. The finally stabilized under the control of the controller during the training process at is regarded as convergence, otherwise it is regarded as non - convergence;

[0115] [5] Decrease the value by 0.05, and repeat steps [3] and [4]. Observe whether the DDPG algorithm converges. If it does not converge, then take the previous convergent value as the optimal performance value, and directly proceed to step [6]; if it converges, then it is considered that there is room for the value to continue to decrease, and repeat step [5] until it does not converge.

[0116] The second - stage reinforcement learning training (i.e., the second reinforcement learning training):

[0117] [6] Import the parameters of the trained critic - actor network in the previous stage into the new training environment. The reward function is AGRF, and the value continues with the optimal performance value finally selected in the first stage, and the

[0118] value is initially set to 0;[7] Execute the digital switching power supply startup event during the simulation process;

[0119] [8] Use the DDPG algorithm to perform reinforcement learning training during the simulation process;

[0120] [9] Observe whether there is still voltage overshoot in the output voltage under the control of the digital switching power supply during the neural network training process during the startup stage; Whether there is still voltage overshoot during the startup stage;

[0121]

[10] If there is still voltage overshoot, then Increase the value by 0.5 and repeat steps [8] and [9]. If the overshoot no longer exists and the output voltage during the startup phase rises steadily, proceed to step

[11] .

[0122] The third stage of reinforcement learning training (i.e., the third reinforcement learning training):

[0123]

[11] Import the parameters of the trained critic - actor network from the previous stage into the new training environment. The reward function is AGRF. The value and The value takes the final value of the previous stage of training.

[0124]

[12] Execute the digital switching power supply startup event and the load jump event during the simulation process.

[0125]

[13] Use the DDPG algorithm to perform reinforcement learning training during the simulation process.

[0126]

[14] After the training is completed, extract the actor network separately as the neural network controller, which serves as the digital controller of the digital switching power supply.

[0127] Furthermore, some terms in the above - described training process are explained as follows:

[0128] 1. Gaussian reward function:

[0129] In the Figure 2 shown training environment, the reward function is used to evaluate the output performance of the current digital switching power supply. The smaller the gap between the output voltage and , the better the performance of the current output voltage. The evaluation of the reward function provides a criterion for the reinforcement learning training process and a direction for it to evolve through training. This application proposes to use a Gaussian function with parameters as the reward function, which is divided into the symmetric Gaussian reward function SGRF (Symmetric Gaussian reward function) and the asymmetric Gaussian reward function AGRF (Asymmetric Gaussian reward function), and their expressions are as follows:

[0130]

[0131]

[0132] In SGRF and AGRF, the The smaller the parameter of, the stricter the reward function and the better the performance of the trained controller. In AGRF, the The larger the parameter is, the more severe the penalty of the reward function for the overshoot part of the output voltage of the digital switching power supply is, and the stronger the suppression of the voltage overshoot by the trained controller is.

[0133] 2. Training event set:

[0134] Since reinforcement learning is completed during the interaction between the neural network controller and the controlled object, i.e., the power stage circuit of the digital switching power supply, during the training process, the designer of the neural network will set the events that the power stage circuit should occur according to the training objectives. The training events involved in this application include two, one is "start-up", and the other is "load jump".

[0135] a) Start-up: The start-up process refers to the transition of the switching power supply from a non-operating state to an operating state. In the non-operating state, the input voltage of the switching power supply remains , the output voltage is 0V, the neural network controller has no initial output, and both the digital pulse width modulator DPWM and the analog-to-digital converter ADC do not work. Starting from the start-up time 0, both the neural network controller and the DPWM and ADC start to work according to their own inherent output frequencies. When the output voltage stabilizes to nearby and there is no obvious fluctuation anymore, the start-up event ends.

[0136] b) Load jump: After the output voltage stabilizes, the load jumps from no-load to heavy-load, maintains the heavy-load for a fixed period of time, and then jumps from heavy-load to no-load.

[0137] In the above basic training environment, the power stage circuit of the digital switching power supply takes the Buck type as an example. If it is necessary to train other types of power stage circuits of the switching power supply in Figure 4 , only the positions of the rectifier diode, power switch tube, inductor, and capacitor need to be changed, and the neural network controller, its control loop, and the neural network training environment remain unchanged.

[0138] In a specific embodiment, the specific environment settings used in this example:

[0139] 1. Build the training environment:

[0140] As Figure 2 shown is the reinforcement learning training environment with the Buck type as the controlled object. The entire environment is built in the Simulink simulation environment. During the interaction between the intelligent agent Agent and the power stage circuit of the Buck type switching power supply, it continuously evolves itself to improve the critic network and the executor network. Agent is implemented using the reinforcement learning toolbox in Simulink.

[0141] The reinforcement learning training environment built in Simulink with the Buck type switching power supply as the controlled object, asFigure 8 as shown

[0142] The Buck-type digital switching power supply, such as Figure 9 as shown Figure 9 The specific parameters of the Buck-type switching power supply in it are shown in Table 1, and the training goal is to reduce the 5V input voltage to 2.5V output.

[0143] Table 1 Parameters of the Buck-type switching circuit

[0144]

[0145] 2. Set the training random variables required in the environment

[0146] In order to enhance the robustness and generalization ability of the neural network controller after training, random variables must be added during training. The variation ranges of the variables in the two stages of training are shown in Table 2 and Table 3.

[0147] Table 2 Variation ranges of the random variables required during training

[0148]

[0149] Table 3 Hyperparameters required during training and their adjustment ranges

[0150]

[0151] 3. Neural network scale and training parameters:

[0152] The two neural networks used in training are the executor network and the critic network respectively. Their network structures and the number of nodes are as Figure 6 shown

[0153] IV. Intermediate results during training

[0154] 1. The first training and its results

[0155] The reward function SGRF used in the first training:

[0156]

[0157] Starting from = 1 for training, gradually reduce the value. Finally, when < 0.1, the training no longer converges. As Figure 10 shown, it is the reward value distribution of the reward function under different values. The more concentrated the reward value distribution is in the middle, the better the training effect. However, correspondingly, as Figure 11 shown, it is for different The number of training rounds under a certain value and the changing trend of the rewards obtained from training. The more training rounds there are, the more difficult it is for the total rewards in training to increase. More rounds are needed to converge stably. When the total rewards are stable above 8.5K, it is considered to have converged. As Figure 12 shown, it shows the waveforms of the output voltage during the startup process of the Buck-type switching power supply for the controllers trained using reward functions with different values, showing that the smaller it is, the higher the control quality, the shorter the settling time, and the smaller the steady-state error. According to the training situation, the best value finally selected for the first training is = 0.1.

[0158] In Matlab, use save("save path / file name.mat", "agent"); to save the results of the previous training, and then use the command load("save path / file name.mat", "agent") at the start of the second training to load, which can make the second training start from the results of the first training.

[0159] 2. The second training and its results

[0160] Reward function used in training: AGRF

[0161] (7)

[0162] Figure 13 It shows the distribution of the reward values of the reward function under different parameter values. The higher the asymmetry of the reward value, the stronger the overshoot suppression for voltage startup. Figure 14 It shows the waveforms of the output voltage during the startup process of the Buck-type switching power supply for the controllers trained using reward functions with the same value and different values, showing that the larger it is, the smaller the startup overshoot, the smoother the transient response during the startup process, and the safer the startup. In this case, when > 7, there is no obvious difference in the training effect.

[0163] In Matlab, use save("save path / file name.mat", "agent"); to save the results of the previous training, and then use the command load("save path / file name.mat", "agent") at the start of the second training to load, which can make the second training start from the results of the first training.

[0164] 3. The third training and its results

[0165] In the third training, the reward function is the same as the second, but the load change is added in the simulation. The load used in this case is a constant power load. The load mutation is added in the third training. The load change value is random, and the range is shown in Table 3 to improve the transient response to load changes. Similarly, use save ("save path / file name.mat", "agent") in matlab; save the results of the second training, and then use the command load ("save path / file name.mat", "agent") at the beginning of the third training to load it, so that the third training can start with the results of the second training.

[0166] After three training sessions, the output effect of the switching power supply during the startup phase changes as follows: Figure 15 shown.

[0167] V. Deployment of Neural Network Controller

[0168] After training, the network weights and bias values ​​of the executor network in the reinforcement learning agent in the simulation environment are extracted separately. The network structure is rebuilt in the FPGA unit or digital IC as a digital controller for controlling the switching power supply in reality. At this point, the digital controller design is completed. Figure 16 Shown is a schematic diagram of the trained actor network being taken out separately and implemented in a real circuit control environment.

[0169] 6. Test verification results:

[0170] like Figure 17 As shown, the executor network is taken out and implemented in FPGA to form a test and verification system.

[0171] Figure 18 Demonstrates the switching power supply in Equal to 5V, The output voltage under heavy load is 2.7V, 2.2V and 1.9V respectively. Output status. The blue waveform is the load current at this time , unified as 1A, the purple waveform is the current switching signal, and its frequency is 2MHz. According to the test results, When the output voltages are 2.7V, 2.2V and 1.9V respectively, the output voltages represented by the green waveforms are 2.65V, 2.14V and 1.83V respectively, the steady-state errors are 50mV, 60mV and 70mV respectively, and the output ripples are 250mV, 220mV and 200mV respectively. The test results show that That is, when the reference voltage changes, the output voltage can be adjusted to the corresponding voltage level.

[0172] Figure 19Shows the input voltage of the switching power supply when it changes during operation, the output voltage of the dynamic waveform. The purple waveform is the input voltage waveform, the green waveform is waveform, and the blue waveform is the load current at this time . The results show that when jumps, does not fluctuate greatly and can fully adapt.

[0173] Figure 20 Shows the input voltage of the switching power supply when starting from 0V first, the output voltage of the dynamic waveform. The purple waveform is the input voltage waveform, and the green waveform is waveform. The results show that during the start-up stage of the switching power supply, rises smoothly with and there is no large overshoot voltage, meeting the expectations.

[0174] Figure 21 Shows the dynamic waveform of the output voltage when the load of the switching power supply jumps from 0.5W to 3W and from 3W to 0.5W, to test the impact of load changes on the output voltage . This test is carried out under three different combinations of circuit element parameters to verify the impact of load jumps on the output voltage when the key components in the DC-DC switching power supply power stage circuit, capacitors and inductors, drift. The blue waveform is the change waveform of the load current at this time, and the green waveform is waveform. The results show that under standard capacitors and inductors, that is C =2.2μF, L =2.2μH, when the load current jumps by 1A amplitude, the overshoot and undershoot of the output voltage do not exceed 70mV, and the settling time is less than 16μs. When the capacitor and inductor are C =1μF, L =2.2μH, that is, when the capacitance value of the capacitor becomes smaller, when the load current jumps by 1A amplitude, the overshoot and undershoot of the output voltage do not exceed 220mV, and the settling time is less than 110μs. When the capacitor and inductor are C =2.2μF, L =4.7μH, that is, when the inductance value of the inductor becomes larger, when the load current jumps by 1A amplitude, the overshoot and undershoot of the output voltage do not exceed 130mV, and the settling time is less than 50μs. When the capacitor and inductor are C =4.7μF, LWhen L = 4.7 μH, that is, when both the capacitance and the inductance increase, the load current jumps by 1 A, the overshoot and undershoot of the output voltage do not exceed 240 mV, and the settling time is less than 40 μs.

[0175] The above test results show that the neural network controller trained by the training method proposed in this application can make the feedback loop of the DC-DC switching power supply more adaptive and robust.

[0176] Through the above-mentioned training method of the neural network controller for digital switching power supplies based on reinforcement learning, on the one hand, for the design of the neural network controller of the digital switching power supply, there is no need to carry out mathematical modeling, and excellent control effects can be achieved only by using the design and control of the training process. By using the Gaussian reward function based on adjustable parameters, the control performance and voltage overshoot can be quantitatively adjusted to design a controller that meets the performance indicators. On the other hand, this method automatically explores the optimal control strategy through a trial-and-error mechanism, without manually designing complex control laws or adjusting PID parameters. It can find the global optimal solution in complex non-linear systems, avoiding the possible local optimal problems of traditional design methods. DC-DC converters usually face dynamic changes such as load transients and input voltage fluctuations. Reinforcement learning can adjust the control strategy in real time to adapt to these changes, significantly improving the robustness and adaptability of the system. It can also show better fault tolerance for parameter perturbations (such as component aging). This method is data-driven and can work effectively even when the model is not fully clear. It reduces the cumbersome design and debugging process and shortens the development cycle. This method can quickly adapt to new working scenarios or performance requirements by updating the policy network. Further improvements can be made on the basis of this method, and other performances of the controller (such as predictive control and self-healing ability) can be increased by adding training events during the training process.

[0177] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0178] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A training method for a digital switching power supply neural network controller based on reinforcement learning, characterized in that: The method includes: Step S1, establishing a power level circuit model and a reinforcement learning agent of a digital switching power supply in a simulation environment; wherein the power level circuit model includes an input power supply, a switch tube, a filter circuit and an output end, and the reinforcement learning agent includes an actor-critic network composed of an actor network and a critic network, wherein the actor network is used to receive a state vector output by an observer, and the critic network is used to evaluate a control strategy output by the actor network; Step S2, based on the reinforcement learning agent and the power level circuit model, a symmetric Gaussian reward function SGRF is used to initialize the value of the SGRF parameter, and the executor network and the critic network are trained for the first reinforcement learning by the DDPG algorithm, and the reward function parameters are gradually reduced until the training does not converge; Step S3, based on the network parameters of the actor-critic network trained by the first reinforcement learning, an asymmetric Gaussian reward function AGRF is used to initialize the AGRF parameters, and the actor-critic network is trained by the second reinforcement learning through the DDPG algorithm, and the values ​​of the AGRF parameters are adjusted to eliminate the voltage overshoot in the startup phase; Step S4, based on the network parameters of the actor-critic network trained by the second reinforcement learning, introduce a load mutation event in the power level circuit model, use an asymmetric Gaussian reward function AGRF, and fix the SGRF parameters and AGRF parameters, and perform a third reinforcement learning training on the actor-critic network through the DDPG algorithm until the training is completed; Step S5, extracting the network parameters of the actor network in the actor-critic network to construct a neural network controller of the digital switching power supply to control the stability of the output voltage in real time.

2. According to the reinforcement learning-based digital switching power supply neural network controller training method of claim 1, it is characterized in that: The state vector includes the output voltage error, the integral error and the differential error, and outputs the duty cycle signal to the power stage circuit model.

3. According to claim 1, the training method for a digital switching power supply neural network controller based on reinforcement learning is characterized in that: Step S2 specifically includes: Step S21, initialize the actor-critic network, use the symmetric Gaussian reward function SGRF, and initialize the SGRF parameters The value of is 1; Step S22, executing a digital switch power supply startup event; Step S23, using the DDPG algorithm to perform the first reinforcement learning training on the actor-critic network; Step S24, if the output voltage of the digital switching power supply under the control of the controller during the training of the executor network and the critic network Stable , then the DDPG algorithm converges; Step S25, set the SGRF parameters Reduce by 0.05, repeat step S23 and step S24, and determine whether the DDPG algorithm converges; Step S26: If the DDPG algorithm does not converge, take the last SGRF parameter that can converge The value of is taken as the SGRF parameter with the best performance The value of , and the first reinforcement learning training ends; if the DDPG algorithm converges, step S25 is repeated until the DDPG algorithm does not converge.

4. According to claim 3, the training method of a digital switching power supply neural network controller based on reinforcement learning is characterized in that: Step S3 specifically includes: Step S31, import the network parameters of the actor-critic network trained by the first reinforcement learning into the new training environment, and use the asymmetric Gaussian reward function AGRF, SGRF parameters The value of continues the SGRF parameter with the best performance in step S26 The value of AGRF parameter The value of is initially set to 0; Step S32, executing a digital switch power supply startup event; Step S33, using the DDPG algorithm to perform a second reinforcement learning training on the actor-critic network; Step S34, determining the output voltage of the digital switching power supply under the control of the controller during the training of the performer network and the critic network Is there a voltage overshoot? If there is a voltage overshoot, set the SGRF parameter The value of is increased by 0.5, and steps S33 and S34 are repeated; if there is no voltage overshoot, the voltage output rises steadily during the startup phase, and the next step is entered.

5. According to claim 4, the training method for a digital switching power supply neural network controller based on reinforcement learning is characterized in that: Step S4 specifically includes: Step S41, import the network parameters of the actor-critic network trained by the second reinforcement learning into the new training environment, and use the asymmetric Gaussian reward function AGRF, SGRF parameters The values ​​and AGRF parameters The value of takes the final value of the second reinforcement learning; Step S42, executing a digital switch power supply startup event and a load mutation event; Step S43, using the DDPG algorithm to perform a third reinforcement learning training on the actor-critic network until the training is completed.

6. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 5 is characterized in that: The expression of the symmetric Gaussian reward function SGRF is: in, is the SGRF parameter, is the output voltage, is the ideal output reference value of the output voltage; The expression of asymmetric Gaussian reward function AGRF is: in, is the AGRF parameter.

7. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 6 is characterized in that: The DDPG algorithm consists of the following steps: Initialize the network parameters of the actor-critic network; Storing training data through the experience replay pool; Iteratively update the network parameters of the actor-critic network to maximize the cumulative reward value, and synchronously update the network parameters of the actor-critic network.

8. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 7 is characterized in that: According to the power stage circuit model, the analog-to-digital converter in the feedback control loop, the trained digital switching power supply neural network controller and the digital pulse width modulator, a digital controlled DC-DC switching power supply is formed; wherein, The power stage circuit model turns on and off the power switch tube according to the switch signal of the previous duty cycle to adjust the input voltage Perform chopping; the chopped input voltage The filtering by capacitor C and inductor L forms a voltage lower than the input voltage. Output voltage , to complete the voltage from the input voltage To output voltage of blood pressure reduction; The analog-to-digital converter outputs a voltage Convert to value; and Value and reference voltage value Subtract, output voltage error ; The voltage error The value of The points and The differential value of is used as the output vector of the observer , which is used as the input vector of the trained digital switching power supply neural network controller, and the reward function by As input variable; input vector After calculation by the executor network, the duty cycle signal is output That is, the proportion of the power switch tube conduction time to the switching cycle; Digital pulse width modulator for duty cycle signal Modulation is performed to form a switching signal, which is then output to the power switch tube in the power stage circuit model to adjust the ratio of the on-time and off-time of the power switch tube so that the output voltage at the current moment is is adjusted.

9. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 8 is characterized in that: Load mutation events include: After the output voltage stabilizes, the load changes from no-load to heavy-load and maintains for a preset time, and then changes back to no-load.

10. The digital switching power supply neural network controller training method based on reinforcement learning according to claim 9 is characterized in that: Switching power supply startup events include: The startup process refers to the transition of the switching power supply from the non-working state 0V output to the working state.

Citation Information

Patent Citations

  • Secondary frequency modulation control method and system for distributed energy storage system

    CN111224433A

  • Digital DC-DC converter loop control method, system, device and medium

    CN116345898A