Asymmetric Regulation Method of Bilateral Switches in Dual Active Bridge Converters
Through deep reinforcement learning and random forest model combined with PI-controlled bilateral switch asymmetric adjustment method, the problem of high computational complexity of existing DAB converters is solved, efficient output voltage and power regulation is achieved, and the system's operating efficiency and the transferability of the adjustment strategy is improved.
Patent Information
- Application Number
- CN202510751800.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-06-06
AI Technical Summary
The bilateral switch asymmetric adjustment strategy of existing DAB converters relies on complex multimodal derivation and convex optimization theory, and is highly computationally cost-effective and difficult to achieve optimal transmission efficiency.
A bilateral switch asymmetric adjustment method based on deep reinforcement learning algorithm is adopted, combined with a random forest model, and through numerical modeling of circuit physical laws and PI control, the optimal adjustment strategy is obtained in real time, reducing the computational complexity and improving adaptive adjustment capabilities.
It realizes efficient and accurate output voltage and power regulation of DAB converters, reduces calculation costs, improves the operating efficiency of the system and the transferability of the adjustment strategy.
Smart Images

Figure CN120301168B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power electronic converter control, and more particularly to a bilateral switch asymmetric regulation method for a dual active bridge converter. Background Art
[0002] In recent years, the rapid penetration of renewable energy has driven the rapid development of DC microgrids and energy storage technologies. Compared to traditional power architectures, this new power system, which incorporates a large amount of renewable energy, places higher demands on the flexibility and stability of DC converters. To this end, Dual Active Bridge (DAB) converters, with their bidirectional energy transmission characteristics, high-frequency isolation, easy soft switching, and flexible load access capabilities, have become widely used.
[0003] To improve the efficiency of DAB converters, a large number of control strategies have emerged. Phase Shift Modulation (PSM) is a classic control method for DAB converters, including single phase shift (SPS), extended phase shift (EPS), dual phase shift (DPS), and triple phase shift (TPS). SPS is the easiest to implement, but it struggles to achieve soft switching when input and output voltages are mismatched or when the load is light, leading to problems such as soft switching difficulties and low efficiency. To address this, EPS and DPS add a degree of freedom to mitigate these issues. To further improve DAB efficiency, the three-degree-of-freedom (TPS) control scheme has been proposed. Compared to the above three control methods, it offers a wider soft switching range and achieves higher efficiency.
[0004] Because the trigger pulses in the aforementioned PSM strategies all have a fixed 50% duty cycle, achieving a wide soft-switching range and relatively low current stress is difficult. To address this, the asymmetric duty modulation (ADM) method has been proposed. Under this modulation strategy, the switches in the same bridge arm can have different on-times, thereby widening the soft-switching range. Based on PSM and ADM, the three-degree-of-freedom dual-side asymmetric duty modulation (DSAM) and dual-side asymmetric phase-shift modulation (DAPM) methods have emerged. They offer significant advantages under light loads and wide voltage ratios.
[0005] The current mainstream approach to optimizing DAB steady-state regulation relies on piecewise linear models or harmonic analysis to derive analytical expressions for maximum inductor current, transmitted power, and soft switching conditions. Optimization targets include variables such as maximum inductor current, RMS current, soft switching operation, and transmitted reactive power. The optimal regulation angle is then determined using the Karush-Kuhn-Tucker (KKT) condition or traditional optimization algorithms. However, these methods rely on complex theoretical derivations or require continuous real-time iteration to determine the optimal regulation result, resulting in lengthy processing times and high computational costs.
[0006] While artificial intelligence (AI) has garnered widespread attention in power electronics regulation in recent years for its exceptional nonlinear fitting and adaptive learning capabilities, its current application in DAB converters has primarily focused on symmetrical phase-shift regulation methods. For example, for PSM regulation, online optimizers such as reinforcement learning, neural networks, fuzzy neural networks, and neural network model predictive control are being used to improve DAB converter efficiency. However, due to a lack of theoretical knowledge and high control complexity, the application of AI technology in bilateral asymmetric regulation is significantly limited compared to PSM.
[0007] From the above content, it can be seen that in practical applications, the existing bilateral switch asymmetric regulation strategies of DAB converters all rely on complex multi-modal derivation and convex optimization theory solutions. A new regulation strategy that is universal for all modes, does not require theoretical derivation, and can achieve optimal transmission efficiency is urgently needed. Summary of the Invention
[0008] In response to the defects of existing DAB converter analysis theories and adjustment methods, the present invention provides a bilateral switch asymmetric adjustment method for a dual active bridge converter. This method is based on the circuit transient behavior and numerical simulation method of the DAB converter, accurately simulates the dynamic behavior of each state variable of the circuit, and uses advanced deep reinforcement learning algorithms for real-time solution, significantly improving the system's adaptive adjustment capability and operating efficiency.
[0009] The technical solution adopted by the present invention is as follows: a bilateral switch asymmetric regulation method of a dual active bridge converter, comprising:
[0010] Obtain the system parameters and operating parameters of the dual active bridge DC / DC converter. The system parameters include switching period, switching frequency, circuit output capacitance, and transformation ratio; the operating parameters include input voltage, output voltage, and output power.
[0011] Input the reference values of input voltage, output power and output voltage at the current moment into the random forest model trained by the optimal adjustment angle dataset, and obtain the reference value of the optimal static adjustment angle at the next moment in real time. ;
[0012] Based on the obtained system parameters and operating parameters, the reference value of the transient phase shift angle is obtained by using the PI control principle. ;
[0013] Based on the obtained reference value of the static adjustment angle and the reference value of the transient phase shift angle, precise adjustment of the output voltage and output power of the dual active bridge converter is achieved.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for asymmetric regulation of bilateral switches of a dual active bridge converter is implemented.
[0015] The present invention also provides a storage medium containing computer executable instructions, which, when executed by a computer processor, implement the method for asymmetric regulation of bilateral switches of a dual active bridge converter.
[0016] The present invention also provides a computer program product, comprising a computer program, which implements the above-mentioned bilateral switch asymmetric regulation method of a dual active bridge converter when executed by a processor.
[0017] The beneficial effects of the present invention are:
[0018] (1) The method of the present invention obtains performance evaluation data of different switching regulation strategies through a numerical modeling method based on circuit physics laws, without the need for complex modal theory derivation. The method is accurate and concise, and is more convenient than the existing theoretical analysis model of DAB converters.
[0019] (2) By combining an advanced bilateral asymmetric regulation strategy with PI control, the DAB converter not only achieves precise regulation of output voltage and output power, but also has superior steady-state performance, such as minimum current stress, maximum transmission efficiency, and soft switching range.
[0020] (3) Without the need for complex derivation or repeated iteration of the optimization algorithm, the method of the present invention uses a deep reinforcement learning method to accurately capture the nonlinear mapping relationship between the expected operating conditions and the optimal adjustment strategy. The deployment time cost is low, and the adjustment strategy has good transferability and generalization under different operating conditions and converter prototype parameters.
[0021] (4) When deploying the underlying controller, a simplified random forest model is used to fit the optimal regulation strategy of deep reinforcement learning, thereby reducing the computational cost of the optimal regulation strategy and achieving real-time optimization of the operating efficiency of the DAB converter under complex working conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0023] Figure 1 1 is a topological diagram of a dual active bridge DAB converter according to an embodiment of the present invention;
[0024] Figure 2 This is a schematic diagram of the DAPM adjustment method for the DAB converter in Example 1 of the present invention;
[0025] Figure 3 Schematic diagram of the deep reinforcement learning optimization algorithm in an embodiment of the present invention;
[0026] Figure 4 Schematic diagram of a random forest model according to an embodiment of the present invention;
[0027] Figure 5 is a schematic diagram of a DSAM adjustment method for a DAB converter in Example 2 of the present invention;
[0028] Figure 6 This is a performance comparison chart of the deep reinforcement learning algorithm in Example 1 of the present invention and other algorithms;
[0029] Figure 7 is the optimal adjustment strategy obtained by the deep reinforcement learning algorithm in Example 2 of the present invention;
[0030] Figure 8 This is a performance comparison chart of the deep reinforcement learning algorithm in Example 2 of the present invention and other algorithms;
[0031] Figure 9 This is a comparison chart of the fitting accuracy of the random forest model in Example 2 of the present invention and other methods;
[0032] Figure 10 This is a comparison chart of the pre-training time and online fitting time of the random forest model in Example 2 of the present invention and other methods;
[0033] Figure 11 This is an experimental waveform diagram of the DAB prototype connected to a constant power load in Example 2 of the present invention;
[0034] Figure 12 This is a comparison chart of the operating efficiency of the DAB prototype in Example 2 of the present invention and other adjustment methods. DETAILED DESCRIPTION
[0035] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.
[0036] The present invention provides a bilateral switch asymmetric regulation method for a dual active bridge converter, comprising the following steps:
[0037] Step 1: Obtain the system parameters and operating parameters of the dual active bridge DC / DC converter. The system parameters include the switching period, switching frequency, circuit output capacitance, and transformation ratio; the operating parameters include the current input voltage, output voltage, and output power.
[0038] Step 2: Set the current input voltage , reference value of output power , output voltage reference value Input into the random forest model trained based on the optimal adjustment angle dataset to predict the reference value of the optimal static adjustment angle at the next moment ;
[0039] Step 3: Based on the obtained system parameters and operating parameters, use the PI control transient phase shift angle related expression to solve and obtain the reference value of the transient phase shift angle ;
[0040] Step 4: Based on the obtained reference value of the static adjustment angle and the reference value of the transient phase shift angle, a corresponding PWM drive signal is generated, thereby achieving precise adjustment of the output voltage and output power of the dual active bridge converter while significantly improving the operating efficiency of the DAB converter.
[0041] In the present invention, the transient phase shift angle related expression is obtained by obtaining the next moment PI control strategy output by the tracking error of the output voltage control.
[0042] The optimal adjustment angle dataset is obtained by solving the efficiency optimization adjustment problem through deep reinforcement learning. The efficiency optimization adjustment problem takes the minimum effective value of the inductor current as the optimization target, and the transmission power is bias-free and fully soft-switching operation as the constraints. Then, the above three objectives are integrated into the reward function of reinforcement learning, and a DAB numerical simulation system is constructed as an interactive environment for reinforcement learning to train a deep reinforcement learning model that can accurately capture the nonlinear mapping relationship between the expected operating conditions and the optimal adjustment angle. Then, all possible operating conditions are considered. Input into the trained deep reinforcement model to obtain the optimal adjustment angle dataset under any working conditions.
[0043] The random forest model is trained based on an optimal adjustment angle dataset using input voltage, reference output power, and reference output voltage under specified operating conditions as input, and a reference static adjustment angle as output. The random forest model structure includes multiple different least squares boosting (LSBoost) trees.
[0044] Example 1:
[0045] First, a bilateral switch asymmetric phase shift adjustment method for a dual active bridge converter of the present invention is formulated, comprising:
[0046] Based on the principle of bilateral asymmetric phase shift regulation method (DAPM) and the physical laws of DAB converter circuit, a numerical simulation model of the DAB converter circuit is built in proportion to obtain the changes in state variables of each circuit when adjusting different phase shift angles under specific working conditions, providing a rapid strategy evaluation method for subsequent regulation strategies.
[0047] In order to ensure the transient control margin, the present invention adjusts the duty cycle of the H bridge of the primary side of the converter to and the duty cycle of the secondary H-bridge As a static adjustment angle, only the external phase shift angle As a dynamic adjustment variable in real-time control, the real-time optimization of steady-state efficiency is achieved. The present invention takes the minimum effective value of inductor current, no transmission power deviation and full soft switching (Zero Voltage Switching, ZVS) operation as the optimization goal, and applies the deep reinforcement learning algorithm to solve it, and obtains the static adjustment angle reference value that ensures the optimal steady-state efficiency of the system. , thereby obtaining the optimal adjustment angle data set; then combined with random forest fitting technology, based on the optimal adjustment angle data set, with the input voltage, output power reference value, output voltage reference value under specified working conditions as input, and the reference value of static adjustment angle as output, training is performed to obtain a trained optimal adjustment random forest model. The trained optimal adjustment random forest model is a universal expression of the optimal static phase shift angle under any working condition, which is convenient for the deployment of underlying control strategies.
[0048] Finally, with the minimum output voltage Tracking deviation is the control target, that is, ,in is the reference value of DAB output voltage, Based on the current output voltage, a transient regulation angle was initially designed using the PI control principle. Based on the obtained reference values for the static regulation angle and transient phase shift angle, PWM drive signals corresponding to the eight switching transistors were generated. This not only achieves precise regulation of the dual active bridge converter's output voltage and output power, but also significantly improves the operating efficiency of the DAB converter.
[0049] Specifically, the following steps are included:
[0050] Step 1: Construct a general numerical simulation model of the DAB converter under full modal regulation, including the following sub-steps:
[0051] The circuit topology of the DAB converter is as follows: Figure 1 As shown, the primary side H-bridge inverter and secondary side H-bridge rectifier Controllable switches and The switching cycle of the controllable switch is , the switching frequency is ; is the primary H-bridge output voltage; the circuit output capacitance Used to reduce output voltage ripple, The input (primary H-bridge) and output (secondary H-bridge) voltages are connected by inductors. Connected to an isolation transformer T with a transformation ratio of n:1, the primary voltage of the transformer is Size , v cd is the transformer secondary H-bridge input voltage; The initial opening moment is defined as the starting point 0, and the DAPM adjustment waveform is as follows Figure 2 Phase shift control is the mainstream control method for DAB steady-state operation. for Hysteresis The duty cycle is used to control the direction and size of power flow, and the duty cycle and internal phase shift angle Used to reduce cycle power and expand soft switching ZVS range, is half a switching cycle.
[0052] According to Kirchhoff's law, the inductor current , the average value of transmission power in a single DAB cycle , output voltage The following dynamic equations are satisfied:
[0053] (1)
[0054] Where, is the load current of DAB. is the numerical simulation cycle.
[0055] Further, if Figure 2 As shown in the figure, when the working condition and adjustment angle are given, the DAPM adjustment characteristics can be used to obtain the time at each moment in a single switching cycle. and The waveform of the inductor current and the transmission power value under specific working conditions can be obtained by numerical simulation. The numerical value can also be used to obtain the ZVS situation of each switching device, which can be used as a quick evaluation scheme for the performance of subsequent regulation strategies.
[0056] Step 2: Build a deep reinforcement learning algorithm to solve the optimal regulation problem, which includes the following sub-steps:
[0057] Step 2.1: Construct the DAB converter efficiency optimization model under the DAPM regulation method:
[0058] In the DAB converter phase shift adjustment method, current stress and soft switching are often used as important performance indicators. According to the adjustment margin of each phase shift angle on system performance, the external phase shift angle is selected. is the transient adaptive control quantity, Set as static adjustment angle. When the inductor current is the minimum effective value, , no power deviation and full soft switching operation As the optimization goal, in order to achieve the best steady-state operating efficiency, the optimization problem of the bilateral asymmetric regulation strategy can be expressed as follows:
[0059] (2)
[0060] Therefore, the regulation strategy optimization problem contains a minimization objective, an equality constraint and six inequality constraints, so the formulation of the above optimal DAPM regulation strategy is a strongly nonlinear non-convex optimization problem.
[0061] Step 2.2: Build a regulation strategy optimization algorithm based on deep reinforcement learning:
[0062] Before implementing the deep reinforcement learning algorithm, a Markov decision process (MDP) for the bilateral asymmetric regulation subproblem is first established. The decision maker for each regulation angle can be considered an intelligent agent, which accumulates experience in state transitions through interaction with the numerical simulation environment for use in deep reinforcement learning network training.
[0063] In MDP, each agent that formulates the adjustment angle selects the adjustment strategy action value according to the current DAB converter system state. After receiving the adjustment action of the agent, the numerical simulation environment feeds back its strategy reward value, and the agent adjusts the next action value according to the reward value. The Markov decision process of the adjustment strategy optimization problem contains three elements: the state space , action space And the reward function The definition is as follows:
[0064] In the DAB real-time efficiency optimal control, the controller is based on the system state observation value To determine the best adjustment angle action , therefore, the state space of the agent can be expressed as .in, Respectively represent The instantaneous value of input voltage, reference value of output voltage and reference value of output power at each moment. Similarly, For the Three adjustment strategies at each moment Since deep reinforcement learning is difficult to solve optimization problems with complex constraints, complex constraints such as full soft switching operation and no transmission power deviation are imposed on the objective function by introducing a penalty function. The original problem is then transformed into an unconstrained single-objective optimization problem, and the reward function is The settings are as follows:
[0065] (3)
[0066] Where, is the active power penalty factor, Optimize the inductor current effective value index, is the soft switching penalty term.
[0067] (4)
[0068] (5)
[0069] (6)
[0070] Where, is the reference value of output power, To adjust the action The instantaneous value of DAB transmission power; 、 、 、 、 is the adjustment coefficient of each optimization index, and they are all positive numbers; is the power deviation less than the threshold The regulation strategy imposes a smaller penalty, so the active power penalty factor The design is a piecewise function, and ; When all switches meet the soft switching ZVS constraint under a specific regulation strategy, Set it to 1, otherwise to 150 to ensure the effective implementation of ZVS.
[0071] The Deep Deterministic Policy Gradient (DDPG) algorithm uses a dual neural network architecture, namely the online policy network and target policy network in the Actor part, and the online Q network and target Q network in the Critic part, to accelerate the convergence speed and improve the stability of the algorithm. The neural network is mainly used to simulate the value of the Q function. Represents the value function of the action to adjust the policy, used to estimate the given system state Take specific regulatory actions The expected cumulative reward obtained when is expressed as follows:
[0072] (7)
[0073] Where, represents the discount factor, Indicates the time step symbol.
[0074] The specific training process of DDPG is as follows Figure 3 As shown, first the online strategy network in Actor is based on the current state Select the adjustment action , the numerical simulation environment generates and , and then the action reward value Feedback to the agent is used to update the policy network parameters The target strategy network in the Actor is based on the next state sampled from the experience pool Select the next adjustment action , network parameters Regularly from the parameters In which, the mean value of state and action is used as the loss function of the policy network .
[0075] (8)
[0076] Similarly, in the Critic part, the online Q network is based on the current state and action value Calculate the current action value , target Q value , and update the online Q network parameters The Target Q network is based on the state and actions Calculate the actual Q value of the next state , network parameters Regularly pass Update, its loss function The definition is as follows:
[0077] (9)
[0078] Using performance objective functions To measure regulatory action Pros and cons:
[0079] (10)
[0080] Where, the DAB environment state is randomly generated according to the actual operating range, and the action of the adjustment strategy is Satisfy the distribution function , so When the state satisfy Distribution expected value.
[0081] The algorithm training goal is set to minimize the network loss, so the policy value gradient is used To soft update the Actor network parameters.
[0082] (11)
[0083] Where, Indicates that the online policy network parameters are The strategy value gradient when Represents the online Q network parameters as Time Action Strategy The value gradient, This means that the online policy network parameters are The action policy gradient at .
[0084] The target network parameters are updated periodically using the soft update formula as follows:
[0085] (12)
[0086] (13)
[0087] Where, Indicates the soft update factor.
[0088] After the above deep reinforcement learning algorithm is trained, all possible operating conditions of the DAB converter are input into the algorithm to quickly obtain the corresponding optimal bilateral asymmetric regulation strategy data set.
[0089] Step 3: Random forest model based on the best adjustment angle dataset, such as Figure 4shown.
[0090] Since deep reinforcement learning has a complex network architecture and is difficult to deploy directly in the underlying controller, a simplified random forest model is used to output the optimal steady-state adjustment angle in real time. First, Bootstrap sampling is performed on the best bilateral asymmetric adjustment strategy dataset obtained by deep reinforcement learning to introduce randomness in the training process and improve the robustness of the model. The out-of-bag OOB data that is not sampled is used for subsequent model performance verification. Therefore, the optimal adjustment dataset is split into training samples , then use multiple LSBoost weak fitters for fitting, and use voting to aggregate and obtain a unique strong fitter, so as to achieve excellent fitting effect with a streamlined network architecture and deploy it using the underlying controller.
[0091] Step 4: Adaptive transient regulation strategy
[0092] In order to ensure the adaptive adjustment capability of the system output voltage and output power, select is the static adjustment angle, which is directly generated by the random forest model, and As the dynamic adjustment angle, the most commonly used PI control strategy is selected. In DAB transient control, the output voltage is usually Tracking is the goal, then the output voltage tracking error of the regulation strategy is e rr for:
[0093] (14)
[0094] Therefore, the output of the dynamic adjustment angle is:
[0095] (15)
[0096] Where, and Represents the proportional and integral coefficients of the PI controller respectively; due to the dynamic adjustment angle in actual regulation The value range is , so the output of the PI controller is limited.
[0097] Therefore, while accurately tracking the DAB output power and voltage, the steady-state efficiency of the converter DAB can be improved by reasonably adjusting the static and transient adjustment angles.
[0098] After formulating the bilateral switch asymmetric phase shift adjustment strategy of the dual active bridge converter of the present invention, the bilateral switch asymmetric phase shift adjustment method of the dual active bridge converter of the present invention can be performed, which is specifically as follows:
[0099] Acquire the operating parameters of the dual active bridge converter in real time, including the current input voltage, output voltage, output power, and reference value of the output voltage;
[0100] Input the reference values of the current input voltage, output voltage, and output power into the trained random forest model deployed in the controller in step 2 to obtain the reference value of the static adjustment angle at the next moment. ;
[0101] Based on the obtained operating parameters and system parameters, the transient phase shift angle related expression (Formula 15) deployed in the controller is used to solve the reference value of the transient phase shift angle. ; Among them, the system parameters include the switching period, switching frequency, transformation ratio, circuit output capacitance, etc. of the dual active bridge converter.
[0102] Based on the obtained reference values of the static adjustment angle and the transient phase shift angle, PWM drive signals corresponding to the eight switching tubes are generated, thereby achieving precise adjustment of the output voltage and output power of the dual active bridge converter while significantly improving the operating efficiency of the DAB converter.
[0103] Example 2:
[0104] This embodiment is the same as Example 1, except that in step 1 of the process of formulating a bilateral switch asymmetric phase shift adjustment method for a dual active bridge converter of the present invention, the bilateral asymmetric phase shift adjustment DAPM is replaced by the bilateral asymmetric duty cycle adjustment DSAM, and the remaining steps are the same as Example 1.
[0105] Different from DAPM, the primary full bridge of DSAM also adopts asymmetric duty cycle regulation instead of phase shift regulation with a fixed duty cycle of 50%. Represents the primary full-bridge switch tube The conduction time of DSAM is as follows: Figure 5 Similarly, the conditions for achieving full soft switching under DSAM regulation can be obtained:
[0106] (16)
[0107] Furthermore, after formulating the DAB converter bilateral switch asymmetric duty cycle adjustment strategy of the present invention, the DAB converter bilateral switch asymmetric duty cycle adjustment method of the present invention can be performed, which is specifically as follows:
[0108] Acquire the operating parameters of the dual active bridge converter in real time, including the current input voltage, output voltage, output power, and output voltage reference value;
[0109] Input the reference values of input voltage, output power, and output voltage at the current moment into the trained random forest model deployed in the controller in step 2 to obtain the reference value of the static adjustment angle at the next moment. ;
[0110] Based on the obtained operating parameters and system parameters, the transient phase shift angle related expression (Formula 15) deployed in the controller is used to solve the reference value of the transient phase shift angle. ; Among them, the system parameters include the switching period, switching frequency, transformation ratio, circuit output capacitance, etc. of the dual active bridge converter.
[0111] Based on the obtained reference values of the static adjustment angle and the transient phase shift angle, PWM drive signals corresponding to the eight switching tubes are generated, thereby achieving precise adjustment of the output voltage and output power of the dual active bridge converter while significantly improving the operating efficiency of the DAB converter.
[0112] In order to further verify the effectiveness of the control strategy proposed in this invention, an inductor The output capacitance is 54mH The numerical simulation model and physical prototype of the DAB converter with a transformer ratio of 5:3, switching frequency 100kHz, control cycle The time taken is 0.1ms, and the underlying microcontroller is a DSP-TMS320F28335. The adjustment method of the present invention (Examples 1 and 2) was compared with the Soft Actor-Critic (SAC) algorithm, the Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm, the Particle Swarm Optimization (PSO), and the nested heuristic algorithm (PSO&GA) combining PSO and the Genetic Algorithm (GA).
[0113] Figure 6 This is a performance comparison chart of the DAPM regulation strategy of the DDPG method proposed in Example 1 and other algorithms. The performance includes the average value of the inductor current RMS, the average power deviation, and the success rate of full soft switching. It can be seen that the present invention is superior to deep reinforcement learning algorithms such as SAC and TD3 in terms of power deviation, minimizing the RMS value of the inductor current, and achieving soft switching. Although both the PSO&GA algorithm and the present invention method have good global optimization performance, traditional heuristic algorithms require multiple iterations to optimize for different operating conditions. This type of optimization algorithm is difficult to meet the real-time requirements of the controller and has a high deployment time cost. The present invention method can accurately capture the nonlinear relationship between the optimal regulation strategy and the operating conditions, and can quickly obtain the real-time optimal regulation strategy without the need for online optimization.
[0114] Figure 7 The different output powers obtained by the DDPG method of the present invention under the DSAM regulation of Example 2 are and voltage gain The best adjustment strategy when ,in, . Figure 8 The performance comparison chart of DDPG and other algorithms under DSAM regulation is as similar to DAPM. Compared with reinforcement learning algorithms such as TD3 and SAC, DDPG algorithm has obvious advantages, and its performance is basically on par with PSO and GA algorithms.
[0115] Figure 9 The following is a comparison chart of the fitting performance of the random forest model under DSAM regulation in Example 2 and other fitting methods. The control group includes a neural network model, an XGBoost extreme gradient boosting model, and a Support Vector Regression (SVR) support vector regression model. It can be found that the goodness of fit of the random forest model proposed in the present invention is As high as 0.991, the mean squared error (MSE) is also the smallest among the four methods. Figure 10 The figure below compares the pre-training time and online fitting time of the random forest model with other methods. This shows that the random forest model not only achieves high fitting accuracy but also reduces both training and online fitting times. Furthermore, the actual execution time of each model was tested on a physical prototype device, revealing that the random forest model only takes 2.6 μs, while the neural network model takes 33 μs. The controller is a DSP-TMS320F28335 with a main frequency of 150 MHz. Therefore, the proposed random forest model has a lower deployment cost in the converter's underlying regulator.
[0116] Figure 11 The following is a dynamic waveform diagram of each state variable of the DAB converter under DSAM regulation in Example 2. The load is a constant power load CPL with poor stability to fully verify the effectiveness of the proposed regulation method. The output voltage reference value is 200 V. The voltage is 100 V, and the constant power load power switches periodically between 50 and 150 W. It can be found that when the load power fluctuates periodically, the proposed regulation strategy can achieve fast and accurate control of the output voltage and output power with almost no deviation.
[0117] Figure 12This figure compares the operating efficiency of the DAB converter under DSAM regulation in Example 2 with other regulation methods. It can be seen that Example 2 has a significant efficiency advantage under light load compared to the three-phase shift and single-phase shift regulation strategies. Therefore, the proposed lightweight regulation method not only ensures precise output control but also enables the DAB converter to achieve good operating efficiency under light load.
[0118] Corresponding to the above-mentioned embodiment of the method for adjusting the asymmetric switches on the bilateral sides of a DAB converter, the present invention further provides an electronic device comprising one or more processors for implementing the above-mentioned method for adjusting the asymmetric switches on the bilateral sides of a DAB converter.
[0119] The device embodiments can be implemented through software, hardware, or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running them. From the hardware level, it mainly includes a processor, memory, network interface, and non-volatile memory. In addition, any device with data processing capabilities in which the device in the embodiment is located may also include other hardware according to the actual function of the device with data processing capabilities, which will not be described in detail.
[0120] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0121] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present invention. A person of ordinary skill in the art can understand and implement the present invention without inventive work.
[0122] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the method for adjusting the asymmetric switching of a DAB converter bilaterally in the above embodiment is implemented.
[0123] The present invention also provides a computer program product, comprising a computer program, which implements the above-mentioned bilateral switch asymmetric regulation method of a dual active bridge converter when executed by a processor.
[0124] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0125] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0126] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0127] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0129] Those skilled in the art may make various modifications or additions to the described embodiments or replace them with similar methods without departing from the spirit of the present invention or exceeding the scope defined by the appended claims.
[0130] The above embodiments are intended only to illustrate the design concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. The scope of protection of the present invention is not limited to the above embodiments. Therefore, any equivalent changes or modifications made based on the principles and design concepts disclosed in the present invention are within the scope of protection of the present invention.
Claims
1. A bilateral switch asymmetric regulation method for a dual active bridge converter, characterized in that: include: Obtain the system parameters and operating parameters of the dual active bridge DC / DC converter. The system parameters include switching period, switching frequency, circuit output capacitance, and transformation ratio; the operating parameters include input voltage, output voltage, and output power. Input the reference values of input voltage, output power and output voltage at the current moment into the random forest model trained by the optimal adjustment angle dataset, and obtain the reference value of the optimal static adjustment angle at the next moment in real time. ; Based on the obtained system parameters and operating parameters, the reference value of the transient phase shift angle is obtained by using the PI control principle. ; Based on the obtained reference value of the static adjustment angle and the reference value of the transient phase shift angle, precise adjustment of the output voltage and output power of the dual active bridge converter is achieved.
2. The method according to claim 1, characterized in that The random forest model is trained based on the optimal adjustment angle dataset obtained by deep reinforcement learning, with the input voltage under specified working conditions. , output voltage reference value and output power reference values The reference value of the static adjustment angle is obtained by training as the input and the output.
3. The method according to claim 1, characterized in that The random forest model is composed of multiple LSBoost weak learners.
4. The method according to claim 1, wherein The related expression of the transient phase shift angle is an analytical expression of the transient phase shift angle at the next moment obtained by controlling the tracking error of the output voltage, and the control strategy is PI control.
5. The method according to claim 1, wherein The optimal adjustment angle data set is constructed and obtained by the following method: taking the minimum effective value of the inductor current as , no power deviation and full soft switching operation are the optimization objectives to design the reward function, and the deep reinforcement learning algorithm is used for global optimization to obtain the given Optimal static phase shift angle adjustment strategy under As each sample of the optimal adjustment angle dataset; is the input voltage, is the output voltage reference value and is the reference value of output power.
6. The method according to claim 5, characterized in that The deep reinforcement learning algorithm uses a deep deterministic policy gradient algorithm and a dual neural network architecture, while also introducing an experience replay mechanism to improve the convergence of the algorithm.
7. The method according to claim 5, characterized in that The optimization problem is based on the overall goal of minimizing the effective value of the inductor current, with no power deviation and full soft switching operation as constraints. The full soft switching operation and no transmission power deviation constraints are imposed on the reward function by introducing a piecewise penalty function. The original problem is simplified to an unconstrained single-objective optimization problem.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for adjusting bilateral switches asymmetricly of a dual active bridge converter according to any one of claims 1 to 7 is implemented.
9. A storage medium containing computer-executable instructions, characterized in that: When executed by a computer processor, the computer executable instructions implement a bilateral switch asymmetry regulation method for a dual active bridge converter according to any one of claims 1 to 7.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for adjusting bilateral switches asymmetric in a dual active bridge converter according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
LADRC control method and system based on DAB converter
CN114499209A
Dual-active-bridge optimal unilateral asymmetric three-degree-of-freedom phase shift modulation method
CN115622412A