An optimization control method of a dual-active full-bridge converter fused with an AI optimization algorithm

By optimizing the phase-shift control of a dual active full-bridge converter using the TD3 algorithm of deep reinforcement learning, the problem of finding the optimal current performance in traditional control methods is solved, and efficient current control and real-time optimization of the converter are achieved.

CN119853398BActive Publication Date: 2025-11-25UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411606977.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-12
Publication Date
2025-11-25
Estimated Expiration
2044-11-12

AI Technical Summary

Technical Problem

The multi-phase-shift control method of traditional dual active full-bridge converters has multiple phase-shift control variables, which makes it difficult to solve for the optimal current performance and increases the control complexity. Furthermore, it is difficult to achieve the optimal combination of control variables in the design of the closed-loop controller.

Method used

The phase-shifting control variables of the dual active full-bridge converter are trained using the deep reinforcement learning TD3 algorithm. By constructing an actor-critic framework and a reward function, the control variables are optimized to achieve control of the lowest root mean square current and current stress.

Benefits of technology

It enables easy finding of the optimal phase-shift control variable under various operating conditions, reduces root mean square current and current stress, improves converter efficiency, and achieves real-time optimization control in closed-loop control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119853398B_ABST
    Figure CN119853398B_ABST
Patent Text Reader

Abstract

This invention discloses an optimized control method for a dual active full-bridge converter that integrates AI optimization algorithms, while simultaneously controlling the operating states (V1, V2, P) of the dual active full-bridge converter. o ) and phase shift control quantities (D1, D2, D φ Offline training was conducted to obtain multiple sets of triple phase-shift control variables for the dual active full-bridge converter under the lowest RMS current and current stress, and these variables were integrated into a deep reinforcement learning model. Finally, in practical applications, the V1, V2, and P values ​​of the dual active full-bridge DC-DC converter were analyzed. o Sampling is performed, and a deep reinforcement learning model is invoked based on the actual magnitude of the sampled values ​​to map out the phase shift angle (D1, D2, D) that matches the optimal current performance. φ The current characteristics of the dual active full-bridge converter are optimized based on the final triple phase-shift control variables.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of DC-DC converter control technology, and more specifically, relates to an optimized control method for a dual active full-bridge converter that integrates AI optimization algorithms. Background Technology

[0002] The dual-active-bridge (DAB) DC-DC converter was first proposed in the early 1990s, such as... Figure 1 As shown, it includes a high-frequency power electronic transformer T. r A series inductor L r It consists of one input-side full-bridge and one output-side full-bridge. As one of the most popular bidirectional topologies, the dual active full-bridge converter offers advantages such as electrical isolation, high power density, wide voltage transmission range, and ease of soft switching, making it widely used in electric vehicles, smart grids, and renewable energy systems.

[0003] In traditional dual active full-bridge converters, multiple phase-shift control methods are used, such as Figure 2 As shown, the two switching devices in each bridge arm employ complementary switching modes, with each switching device having a 180° conduction phase (ignoring dead time). The transmission power is controlled by adjusting the phase difference between the four bridge arms. This control method has multiple phase-shift control variables. By combining these variables, the root-mean-square current and current stress of the dual active full-bridge converter can be reduced under a given transmission power condition. However, the presence of multiple phase-shift control variables in this modulation method makes finding the optimal current performance and the control complexity extremely high.

[0004] Taking general three-phase shift control as an example, given the input voltage V1 and the output voltage V2, with the switching frequency remaining constant, there are up to three control variables, such as... Figure 2 As shown, control variable D1 is the inner phase shift of full-bridge FB1, control variable D2 is the inner phase shift of full-bridge FB2, and control variable D... φ For v p and v s 'Phase shifting between center points. In traditional multi-phase-shift control methods, finding the optimal set of phase-shift control variables to optimize the current performance of a dual active full-bridge converter is extremely difficult. Furthermore, designing a closed-loop controller that makes the control variables approximate the optimal combination of control variables is also very challenging.' Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide an optimization control method for dual active full-bridge converters that integrates AI optimization algorithms. The method trains the phase-shifting control variables that control the dual active full-bridge converters using the advanced deep reinforcement learning TD3 algorithm, thereby achieving optimized control of the dual active full-bridge DC-DC converters.

[0006] To achieve the above-mentioned objectives, this invention discloses an optimization control method for a dual active full-bridge converter that integrates AI optimization algorithms, characterized by comprising the following steps:

[0007] (1) First, set the operating range of the dual active full-bridge converter, and then construct the mathematical expressions for the current characteristics and transmission power of the dual active full-bridge converter under general three-phase shift control.

[0008] (2) Build a deep reinforcement learning model based on the actor-critic framework, and construct the reward function of deep reinforcement learning by combining the mathematical expressions of the current characteristics and transmission power of the dual active full-bridge converter. Then, use the TD3 algorithm in the AI ​​optimization algorithm to train the deep reinforcement learning model until the deep reinforcement learning model converges.

[0009] (3) Real-time acquisition of input voltage V1, output voltage V2, and transmission power P of the dual active full-bridge DC converter. o Then, the state variables s = (V1, V2, P) are formed. o Then, the state variable s is input into the trained deep reinforcement learning model, and the action a = (D1, D2, D...) is fitted through the actor network π. φ This yields the phase shift angles (D1, D2, D) under optimal current performance. φ Finally, based on the phase shift angles (D1, D2, D...), φ )Optimize the control of the dual active full-bridge converter.

[0010] The objective of this invention is achieved as follows:

[0011] This invention discloses an optimized control method for a dual active full-bridge converter that integrates AI optimization algorithms, while simultaneously controlling the operating states (V1, V2, P) of the dual active full-bridge converter. o ) and phase shift control quantities (D1, D2, D φ Offline training was conducted to obtain multiple sets of triple phase-shift control variables for the dual active full-bridge converter under the lowest RMS current and current stress, and these variables were integrated into a deep reinforcement learning model. Finally, in practical applications, the V1, V2, and P values ​​of the dual active full-bridge DC-DC converter were analyzed. o Sampling is performed, and a deep reinforcement learning model is invoked based on the actual magnitude of the sampled values ​​to map out the phase shift angle (D1, D2, D) that matches the optimal current performance. φThe current characteristics of the dual active full-bridge converter are optimized based on the final triple phase-shift control variables.

[0012] Meanwhile, the dual active full-bridge converter optimization control method integrating AI optimization algorithm of the present invention also has the following beneficial effects:

[0013] (1) This invention uses deep reinforcement learning to analyze the operating states (V1, V2, P) of a dual active full-bridge converter. o ) and phase shift control quantities (D1, D2, D φ By training various values ​​of the phase-shift control variables offline, it is easy to find the optimal set of phase-shift control variables, thereby reducing the root mean square current and current stress, and thus improving the efficiency of the dual active full-bridge converter.

[0014] (2) In closed-loop control, by sampling V1, V2 and P... o The corresponding values ​​are searched in the trained deep learning model to map out the phase shift angles (D1, D2, D) that match the optimal current performance. φ Then, the dual active full-bridge DC-DC converter is optimized and controlled based on this set of phase-shift control variables;

[0015] (3) When real-time acquisition of V1, V2 and P of the dual active full-bridge DC converter o When the corresponding values ​​are not within a predefined range, the phase-shifting control variables (D1, D2, D...) can be trained online through deep reinforcement learning. φ It can achieve real-time control of the dual active full-bridge converter. Attached Figure Description

[0016] Figure 1 This is a topology diagram of a dual active full-bridge converter;

[0017] Figure 2 This is a partial voltage and current waveform diagram of a dual active converter;

[0018] Figure 3 This invention relates to a current optimization control structure for a dual active full-bridge converter based on deep reinforcement learning.

[0019] Figure 4 These are the results of the root mean square current experiment;

[0020] Figure 5 These are the results of a current stress experiment;

[0021] Figure 6 These are the results of transmission efficiency experiments. Detailed Implementation

[0022] The specific embodiments of the present invention will now be described with reference to the accompanying drawings to enable those skilled in the art to better understand the invention. It should be particularly noted that in the following description, detailed descriptions of known functions and designs that might obscure the main content of the invention will be omitted here.

[0023] Example

[0024] For ease of description, the relevant technical terms appearing in the specific implementation method will be explained first:

[0025] In this embodiment, as Figure 1 As shown, the dual active full-bridge converter includes a high-frequency power electronic transformer T. r A series inductor L r The system consists of one input-side full-bridge and one output-side full-bridge. The input-side full-bridge comprises two arms, arm 1 and arm 2; arm 1 contains two switching devices, S1 and S2; arm 2 contains two switching devices, S3 and S4. The output-side full-bridge comprises two arms, arm 3 and arm 4; arm 3 contains two switching devices, S5 and S6; arm 4 contains two switching devices, S7 and S8. The two switching devices in each arm employ complementary switching modes, and the conduction phase of each switching device is 180° (ignoring dead time).

[0026] like Figure 2 As shown, control variable D1 is the inner phase shift of full-bridge FB1, control variable D2 is the inner phase shift of full-bridge FB2, and control variable D... φ For v p and v s The range of the three control variables for the outer phase shift between the center points is limited to the phase shift angle within half a cycle, that is, (D1, D2, D...). φ The magnitude of v lies within the range of (0 to π). p To input the voltage difference between the midpoints of the two arms of the full bridge, v s To output the voltage difference between the midpoints of the two arms of the full-bridge circuit, the transformer turns ratio is n:1, V s 'for v s The equivalent voltage to the primary side of the transformer, v s The amplitude is V1, v s The amplitude of ' is equal to nV2,i Lk This represents the current flowing through the inductor.

[0027] The following is a detailed description of the optimization control method for a dual active full-bridge converter that integrates AI optimization algorithms, which includes the following steps:

[0028] S1. Set the operating range of the dual active full-bridge converter and construct mathematical expressions for the current characteristics and transmission power of the dual active full-bridge converter under general three-phase-shift control.

[0029] S1.1 Set the range of input voltage V1, the range of output voltage V2, and the desired transmission power P of the dual active full-bridge DC-DC converter. o The range; in this embodiment, the range of the input voltage V1 is set to 100V to 140V, the range of the output voltage V2 is set to 40V to 50V, and the transmission power P is set to [missing information]. o The range is 80W to 200W;

[0030] S1.2 Construct mathematical expressions for the current characteristics and transmission power of a dual active full-bridge converter under general three-phase-shift control;

[0031]

[0032] Where V1 is the input voltage of the dual active full-bridge converter, V2 is the output voltage of the dual active full-bridge converter, control variable D1 is the inner phase shift of full-bridge FB1, control variable D2 is the inner phase shift of full-bridge FB2, and control variable D... φ For v p and v s The phase shift between the center points is limited to a range of three control variables within a half-cycle phase shift angle, i.e., (D1, D2, D...). φ The magnitude of P lies within the range of (0 to π). o The transmission power of the dual active full-bridge converter is given by coefficients n = 1, 3, 5, 7, ..., where L is the inductance value, A and B are intermediate variables, and I... rms f is the root mean square current of the dual active full-bridge converter. s w0 is the switching frequency, i is the angular frequency, and w0 is the angular frequency. L This represents the inductor current.

[0033] S2. Train a deep reinforcement learning model using the TD3 algorithm;

[0034] In this embodiment, we train a deep reinforcement learning model using the TD3 algorithm in the AI ​​optimization algorithm. During the training process, we target the root mean square current and current stress of the dual active full-bridge converter, and perform offline training on a certain range of input voltage V1, output voltage V2, and desired transmission power Po to obtain the triple phase-shift control variables (D1, D2, D...) corresponding to the optimal current performance. φ The specific process is as follows:

[0035] S2.1, Set the reward function r = -[α(P) for the deep reinforcement learning model. o -P o ') 2 +β|Irms |], where P o ' represents the transfer power in the reinforcement learning process, and α and β are weight coefficients; through simulation experiments, the value of α was selected as the baseline value of 1, and the value of β was selected as 10;

[0036] S2.2 Build a deep reinforcement learning model based on the actor-critic framework, including: an actor network π and two critic networks θ1 and θ2, as well as a target actor network π' and two target critic networks θ1' and θ2';

[0037] S2.3, Within the operating range of the dual active full-bridge converter, V1, V2, P o Random sampling yields the current state s. t =(V 1,t V 2,t ,P o,t ), and then state s t The input is fed into the actor network π, and action a is fitted. t =(D 1,t D 2,t D φ,t ),like Figure 3 As shown;

[0038] S2.4, Action a t Substitute into formula (1) to calculate the inductor current i L Then determine the inductor current i L Does it meet the following conditions:

[0039]

[0040] If the inductor current i L If formula (2) is satisfied, then (s) t ,a t Substituting the values ​​into the reward function of the deep reinforcement learning model, we can calculate the reward value r(s) at the current time step. t ,a t Otherwise, increase the weighting coefficient β. Through simulation experiments, the optimal β value is 50. Then, (s) t ,a t Substituting the values ​​into the reward function of the deep reinforcement learning model, we can calculate the reward value r(s) at the current time step. t ,a t );

[0041] S2.5, Repeat step S2.3 to obtain state s t+1 =(V 1,t+1 V 2,t+1 ,P o,t+1 ) and action a t+1 =(D1,t+1 D 2,t+1 D φ,t+1 );

[0042] S2.6, will (s t+1 ,a t+1 Substitute them into the target critic network θ respectively i ', get the output value

[0043] S2.7 During training, the critic network adjusts its behavior based on the current action a. t and state s t Output Q value The Q-value must closely resemble the corresponding value generated by the target critic network; therefore, we select... Using this as a baseline, the two critic networks θ are then updated using the mean square Bellman error function. i ;

[0044]

[0045] Where N is

[0046] S2.8, with The maximum value is the target update actor network π;

[0047]

[0048] in, Let J(π) represent the gradient, E represent the expectation, and s ~ R represent the expectation. Indicates (s) t ,a t Substitute into the critic network θ i The Q value output later;

[0049] S2.9, Update the two target critic networks θ i ';

[0050] θ i '←τθ i +(1-τ)θ i ',i=1,2

[0051] S2.10, Update the target actor network π';

[0052] π'←τπ+(1-τ)π'

[0053] Where τ represents the soft update factor;

[0054] S2.11. Repeat steps S2.3) to S2.10 until the deep reinforcement learning model converges;

[0055] S3. Control the dual active full-bridge converter;

[0056] Real-time acquisition of V1, V2, and P of the dual active full-bridge DC-DC converter o Then, the current state s = (V1, V2, P) is formed. o Then, the state s is input into the trained actor network π to fit the action a = (D1, D2, D...). φ This yields the phase shift angles (D1, D2, D) under optimal current performance. φ Finally, based on the phase shift angles (D1, D2, D...), φ )Optimize the control of the dual active full-bridge converter.

[0057] Experimental simulation

[0058] In this embodiment, when the input voltage V1 is 100V, the output voltage V2 is 50V, and the transmission power P is... o The corresponding root mean square current experimental diagram is as follows: Figure 4 As shown, SPS represents the RMS current experimental diagram corresponding to the existing single-phase-shift control strategy, DPS represents the RMS current experimental diagram corresponding to the existing dual-phase-shift control strategy, EPS represents the RMS current experimental diagram corresponding to the existing extended phase-shift control strategy, QATPS represents the RMS current experimental diagram corresponding to the three-phase-shift control strategy optimized by Q-learning algorithm combined with ANN neural network, and DUTPS represents the RMS current experimental diagram corresponding to the current performance optimization control strategy of the present invention that integrates AI optimization algorithm. When the input voltage V1 is 100V, the output voltage V2 is 50V, and the transmission power P... o The corresponding current stress experimental diagram is as follows: Figure 5 As shown, SPS represents the current stress experiment diagram corresponding to the existing single-phase-shift control strategy, DPS represents the current stress experiment diagram corresponding to the existing dual-phase-shift control strategy, EPS represents the current stress experiment diagram corresponding to the existing extended phase-shift control strategy, QATPS represents the current stress experiment diagram corresponding to the three-phase-shift control strategy optimized by Q-learning algorithm combined with ANN neural network, and DUTPS represents the current stress experiment diagram corresponding to the current performance optimization control strategy of the present invention that integrates AI optimization algorithm. When the input voltage V1 is 100V, the output voltage V2 is 50V, and the transmission power P... o The corresponding efficiency experiment graph is as follows Figure 5As shown, SPS represents the efficiency experiment diagram corresponding to the existing single-phase-shift control strategy, DPS represents the efficiency experiment diagram corresponding to the existing dual-phase-shift control strategy, EPS represents the efficiency experiment diagram corresponding to the existing extended phase-shift control strategy, QATPS represents the efficiency experiment diagram corresponding to the three-phase-shift control strategy optimized by Q-learning algorithm combined with ANN neural network, and DUTPS represents the efficiency experiment diagram corresponding to the current performance optimization control strategy of the present invention that integrates AI optimization algorithm. Figure 4 , Figure 5 and Figure 6 It can be seen that the current characteristic optimization control method for dual active full-bridge converters based on deep reinforcement learning provided by this invention has relatively low root mean square current and current stress, and can improve the efficiency of dual active full-bridge DC converters.

[0059] Although the illustrative specific embodiments of the present invention have been described above to enable those skilled in the art to understand the invention, it should be understood that the invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the invention as defined and determined by the appended claims, and all inventions utilizing the concept of the present invention are protected.

Claims

1. A dual active full-bridge converter optimization control method integrating AI optimization algorithms, characterized in that, Includes the following steps: (1) Set the operating range of the dual active full-bridge converter and construct mathematical expressions for the current characteristics and transmission power of the dual active full-bridge converter under general three-phase shift control. (1.1) Set the input voltage V1 range, output voltage V2 range, and desired transmission power P for the dual active full-bridge converter application. o Scope; (1.2) Construct mathematical expressions for the current characteristics and transmission power of a dual active full-bridge converter under general three-phase-shift control; Where V1 is the input voltage of the dual active full-bridge converter, V2 is the output voltage of the dual active full-bridge converter, control variable D1 is the inner phase shift of full-bridge FB1, control variable D2 is the inner phase shift of full-bridge FB2, and control variable D... φ For v p and v s The phase shift between the center points is limited to a range of three control variables within a half-cycle phase shift angle, i.e., (D1, D2, D...). φ The magnitude of P lies within the range of (0 to π). o I represents the transmission power of the dual active full-bridge converter. rms f is the root mean square current of the dual active full-bridge converter. s w0 is the switching frequency, i is the angular frequency, and w0 is the angular frequency. L This is the expression for the inductor current; (2) Build a deep reinforcement learning model; (2.1) Set the reward function r = -[α(P) for the deep reinforcement learning model. o -P o ') 2 +β|I rms |], where P o ' represents the transfer power in the reinforcement learning process, and α and β are weight coefficients; (2.2) Construct a deep reinforcement learning model based on the actor-critic framework, including: an actor network φ and two critic networks θ1 and θ2, as well as a target actor network φ' and two target critic networks θ1' and θ2'; (2.3) Within the operating range of the dual active full-bridge converter, V1, V2, P o Random sampling yields the current state s. t =(V 1,t V 2,t ,P o,t ), and then state s t The input is fed into the actor network φ, and the action a is fitted. t =(D 1,t D 2,t D φ,t ); (2.4) Action a t Substitute into formula (1) to calculate the inductor current i L Then determine the inductor current i L Does it meet the following conditions: If the inductor current i L If formula (2) is satisfied, then (s) t ,a t Substituting the values ​​into the reward function of the deep reinforcement learning model, we can calculate the reward value r(s) at the current time step. t ,a t Otherwise, increase the weighting coefficient β in increments, and then (s) t ,a t Substituting the values ​​into the reward function of the deep reinforcement learning model, we can calculate the reward value r(s) at the current time step. t ,a t ); (2.5) Repeat step (2.3) to obtain state s t+1 =(V 1,t+1 V 2,t+1 ,P o,t+1 ) and action a t+1 =(D 1,t+1 D 2,t+1 D φ,t+1 ); (2.6) will (s) t+1 ,a t+1 Substitute them into the target critic network θ′ respectively i , obtain the output value (2.7) Select Then, the two critic networks θ are updated using the mean square Bellman error function. i ; (2.8) with The maximum value is the target update actor network φ; (2.9) Update the two target critic networks θ i '; i i '←tθ i +(1-τ)θ i ',i=1,2 (2.10) Update the target actor network φ'; φ'←τφ+(1-τ)φ' (2.11) Repeat steps (2.3) to (2.10) until the deep reinforcement learning model converges; (3) Control the dual active full-bridge converter; Real-time acquisition of V1, V2, and P of the dual active full-bridge DC-DC converter o Then determine the sampled V1, V2, and P. o The corresponding values ​​are searched in the deep reinforcement learning model in step (2.4) to map out the phase shift angles (D1, D2, D...) that match the optimal current performance. φ Then, the dual active full-bridge converter is optimized and controlled based on the phase-shift control variable.

Citation Information

Patent Citations

  • Optimal control method of dual-active full-bridge direct-current converter

    CN110707935A

  • Energy consumption optimization method for dual-active half-bridge direct-current converter

    CN112685951A