Hydrogen production converter output current control and ripple suppression method based on deep learning
Through deep learning, the phase shift time of the hydrogen-making converter is optimized, and the output current fluctuation and ripple suppression of the hydrogen-making converter is solved, achieving stable output and efficient hydrogen-making.
Patent Information
- Application Number
- CN202510656000.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art has failed to effectively analyze and control the impact of t1-1 and t2-2 on the circuit in the hydrogen-making converter, resulting in fluctuations in the output current and fail to achieve effective current ripple suppression.
Using deep learning method, by sampling voltage and load-side current at the three ports of the hydrogen-making converter, the agent is trained using the depth deterministic strategy gradient algorithm (DDPG) to optimize the phase shift times t1_1, t2_2, t1_3, t3_3 to achieve stable control of output current and ripple suppression.
The stable control of the output current of the hydrogen-making converter and low ripple output are realized, improving the hydrogen-making efficiency and equipment safety.
Smart Images

Figure CN120498266A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of power electronics technology, and in particular to a method for output current control and ripple suppression of a hydrogen production converter based on deep learning. Background Art
[0002] As a green energy source, hydrogen boasts widespread availability, flexibility, and high efficiency. Water electrolysis is considered the most promising method for hydrogen production. Electrolyzer operation places special demands on energy supply, requiring the accompanying DC-DC converter to possess a high step-down ratio, low output voltage, high current output, and excellent efficiency and stability. Low current ripple is also a key requirement. Excessive fluctuations not only reduce hydrogen production efficiency but also impact the safety and lifespan of hydrogen storage equipment.
[0003] In recent years, the advantages of deep learning in complex system prediction and modeling have promoted its application in the intelligent control of hydrogen energy systems. For example:
[0004] A Current Ripple Suppression Strategy of TAB for PhotovoltaicHydrogen Production, a paper on t 3-3 The article points out that the phase shift control scheme of t can be simply calculated. 1-3 , t 3-3 Control is performed to achieve output current ripple suppression. However, this solution has no control freedom and t is not analyzed. 1-1 , t 2-2 The impact on the circuit is not analyzed, and the problem of output current fluctuation of the hydrogen production converter under actual operation is not analyzed.
[0005] The Chinese invention patent application number 202510196085.3 provides a method for suppressing the current ripple of a hydrogen production converter. The method adds a delay t to the original H-bridge circuit of port 1 and port 3 of the hydrogen production converter. 1_1 , t 1_3 , t 3_3 By performing differential evolution algorithm optimization, the optimal delay time corresponding to the minimum of the three current peaks at the load port is obtained, thereby achieving the suppression of the input electrolytic cell current ripple. However, this solution does not analyze t 2-2 The impact on the circuit is not analyzed, and the problem of output current fluctuation of the hydrogen production converter under actual operation is not analyzed. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a method for output current control and ripple suppression of a hydrogen production converter based on deep learning. By increasing the control freedom, the t 1-1 , t 2-2In order to improve the impact on the circuit, a deep learning method is added to achieve output current stability control and ripple suppression, so that the hydrogen production converter has the ability to maintain output current stability and low ripple output.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0008] A method for controlling output current and suppressing ripple of a hydrogen production converter based on deep learning, comprising the following steps:
[0009] Step 1: sampling the voltages of the three ports of the hydrogen production converter and the current at the load end;
[0010] Step 2: Based on the given time segments, the output current and the output current peak value of the third port are calculated using the average current method;
[0011] Step 3: Using the deep learning control strategy, initialize the parameters, and receive the state S of the hydrogen converter environment. t (V1 * ,V2 * ,V3 * ,I out * ) and action a t (t 1_1 ,t 2_2 ,t 1_3 ,t 3_3 ), build a data set, find the phase shift time when the actual output current is consistent with the set value, and find I peak The optimal solution of the third port and the optimal delay time t within the third port 1_1 , t 2_2 , t 1_3 , t 3_3 ;
[0012] Step 4: Use the deep deterministic policy gradient algorithm to calculate the state S t 、S t+1 and action value a t To train the agent so that it can track the output current I out Reference value;
[0013] Step 5: Determine whether the difference between the set value and the algorithm output value is less than 0.001. If not, return to step 4 and perform deep learning again;
[0014] Step 6: Phase shift time t based on training completion 1_1 , t 2_2 , t 1_3 , t 3_3 Calculate the output current peak value, compare different peak values, and select the phase shift combination with the smallest output current peak value t1_1 , t 2_2 , t 1_3 , t 3_3 , observe the actual output current I o size;
[0015] Step 7: Determine whether the difference between the output current setting value and the actual output current value is less than 0.5. If not, return to step 4 and perform deep learning again;
[0016] Step 8: If satisfied, get the optimal phase shift combination t 1_1 , t 2_2 , t 1_3 , t 3_3 , that is, the hydrogen production converter can operate with the lowest output current ripple while maintaining the output current stability.
[0017] A further improvement of the technical solution of the present invention is that: in step 2, the output current I of the third port is calculated out and the output current peak I peak is calculated as follows:
[0018] The time segment relationship after the superposition of the three voltage waveforms at the port is:
[0019] t1=t0+t 1_1
[0020] t2=t0+t 1_2
[0021] t3=t0+t 1_2 +t 2_2
[0022] t4=t0+t 1_3
[0023] t5=t0+t 1_3 +t 3_3
[0024] t6=t0+0.5*T
[0025] Among them, the initial time is t0, which corresponds to the low estimate of the load current, t1 to t5 are the time points corresponding to the change of the load current waveform, t6 is the time corresponding to the load current reaching the peak value, t 1_2 Set to a fixed value;
[0026] According to the sampling results in step 1, the expression of the load current in each time period is analyzed:
[0027] i0=-i6
[0028]
[0029] Wherein, U1 is the input voltage of the first port, U2 is the input voltage of the second port, and U3 is the voltage of the third port; L 13 is the equivalent inductance between the first port and the third port; L 23 is the equivalent inductance between the second port and the third port;
[0030] The equivalent inductance is calculated as follows:
[0031]
[0032] The average current method is used to solve the problem. Combining the above time segment relationship, the output current I is obtained. out The peak current I corresponding to the load terminal at time t6 peak :
[0033]
[0034] The peak current of the load terminal corresponding to time t6 is I peak , then I peak for:
[0035]
[0036] The further improvement of the technical solution of the present invention is that: in step 3, the specific optimization process of the deep learning algorithm is as follows: a control strategy based on deep reinforcement learning is adopted to achieve stable control of the output current of the hydrogen production converter, and the suppression of the output current ripple is considered; the hydrogen production converter environment receiving state S t (V1 * ,V2 * ,V3 * ,I out * ) and action a t (t 1_1 , t 2_2 , t 1_3 , t 3_3 ) to generate the next state S t+1 ; In each round, an action is performed through the action function, which is composed of the sampled state value S t Decision; the action function is defined as follows:
[0037] a t =μ(s t |θ μ )+N
[0038] Where N is random noise; μ is the deterministic policy function, θ μ As a parameter.
[0039] A further improvement of the technical solution of the present invention is that in step 4, the agent training process is specifically as follows:
[0040] The agent uses the state S through the DDPG algorithm t 、S t+1 and action value a t To track the output current reference and calculate the output current peak value corresponding to each phase shift time combination; In the DDPG algorithm, there are two agents, including the Actor network and the Critic network; each network contains two hidden layers, each with 100 neurons;
[0041] Critic Network Q(s t ,a t ∣θ Q ) is determined by the parameter θ Q Indicates that Q is the dynamic-value function, and the input state s t and action value a t , generating an estimated Q(s t ,a t ∣θ Q )value;
[0042] Actor network μ(s t ∣θ μ ) is determined by the parameter θ μ Indicates that according to the input state s t Execution strategy; the strategy is implemented by using the estimated Q(s t ,a t ∣θ Q ) value to update and generate action value a t ;
[0043] Set the action value a t After being placed in the environment of the hydrogen converter, the reward value r t and the next state s t+1 Able to perform calculations; during the entire training process, all variables {s t ,a t ,r t ,s t+1} are stored in the experience replay buffer; then, a small batch n is randomly sampled from the experience replay buffer to update the Actor network and Critic network of each round;
[0044] Critic Network Q(s t ,a t ∣θ Q ) is updated by minimizing the loss function as follows:
[0045]
[0046] The actor network is updated via a gradient strategy as follows:
[0047]
[0048] In order to ensure the stability and convergence of the learning process, the target critic network Q′(s t ,a t ∣θ Q ) and the target Actor network μ′(s t ∣θ μ ) Perform soft updates on the Critic network and the Actor network;
[0049] Update target network weights θ Q′ and θ μ′ The formula is as follows, where τ is the learning rate;
[0050] θ Q′ ←τθ Q +(1-τ)θ Q′
[0051] θ μ′ ←τθ μ +(1-τ)θ μ′
[0052] Target Critic Network Q′(s t ,a t ∣θ Q )’s output y i As shown below, γ is the discount factor:
[0053] y i =r(s,a)+γQ′(s′,μ′(s′)|θ Q′ |s t+1 )
[0054] To ensure that the calculated output current matches the set value, a reward function is established based on the difference between the calculated value and the set value. The reward function is as follows:
[0055]
[0056] A further improvement of the technical solution of the present invention is that: in step 5, it specifically includes: judging whether the absolute value of the difference between the output current value given by the intelligent agent and the set value is less than 0.001. If it is satisfied, the intelligent agent can follow the output current set value. At this time, the intelligent agent training is completed and actual experiments can be carried out; if it is not satisfied, return to step 4 for a new cycle.
[0057] A further improvement of the technical solution of the present invention is that: in step 6, specifically including: according to the combination t of multiple groups of phase shift times given by the intelligent agent 1_1 , t 2_2 , t1_3 , t 3_3 , substitute into the peak current expression to solve the output current peak value; compare the peak currents, select the lowest value, record it as the minimum peak current value and record this set of phase shift time t 1_1 , t 2_2 , t 1_3 , t 3_3 , input the signal into the hydrogen production converter circuit and observe the actual output current value I o .
[0058] A further improvement of the technical solution of the present invention is that: in step 7, it specifically includes: judging whether the absolute value of the difference between the actual output current value and the set current value is less than 0.5, if it is satisfied, then the intelligent agent not only can the algorithm calculation meet the set value requirement, but also the actual output can meet the requirement; if not, then according to the actual output current I o and set value I out * Update the absolute value of the difference I out * , set the updated output current setting value I out * Return to step 4 for a new cycle; output current setting value I out * The update process satisfies the following formula:
[0059]
[0060] A further improvement of the technical solution of the present invention is that: in step 8, specifically including: according to the phase shift combination t with the minimum output current peak value obtained by optimization 1_1 , t 2_2 , t 1_3 , t 3_3 , recorded as the phase shift time combination that satisfies the actual output current matching the set value and the output current ripple is minimized; at this time, the hydrogen production converter can stably operate in a working state where the output current meets the set value and the output current ripple is minimized.
[0061] Due to the adoption of the above technical solution, the technical advancements achieved by the present invention are:
[0062] The present invention uses a deep learning method to control the four phase shift control schemes of the hydrogen production converter. 1-1 , t 2-2 , t 1-3 , t 3-3 These four degrees of freedom of phase shift time achieve stable output current control and ripple suppression of the hydrogen production converter, so that the hydrogen production converter can produce more hydrogen while maintaining stable output current control. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive efforts.
[0064] Figure 1 is a circuit diagram of the hydrogen production converter in the present invention;
[0065] Figure 2 This is a flow chart of a method for output current control and ripple suppression of a hydrogen production converter based on deep learning provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0066] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or apparatuses.
[0067] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside" and the like to indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0068] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "several" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0069] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:
[0070] like Figure 1As shown, the existing hydrogen production converter (also known as a triple-active bridge converter) includes a first H-bridge module, a second H-bridge module, and a third H-bridge module. The energy flow relationship of the hydrogen production converter is two-terminal input and single-terminal output. The first port of the first H-bridge module is a photovoltaic DC power supply module, the second port of the second H-bridge module is a battery module, and the third port of the third H-bridge module is an electrolyzer module, which also serves as the hydrogen production port. The first port serves as the main power source, supplying power to the second and third ports. The second port receives input from the first port and supplies power to the third port. The third port receives energy input from the first and second ports and consumes energy as a load port.
[0071] The first H-bridge module consists of the first switch tube S a1 , the second switch tube S a2 , the third switch tube S a3 , the fourth switch tube S a4 , a first filter capacitor C1, a first excitation inductor L1 and a first three-winding transformer with a voltage ratio of N1; a first switch tube S a1 and the fourth switch tube S a4 Diagonal arrangement to achieve coordinated switch control; the second switch tube S a2 and the third switch S a3 Diagonal setting to achieve coordinated switch control, the diagonal switch combination is (S a1 , S a4 ) and (S a2 , S a3 ), realizing switch complementary control;
[0072] The second H-bridge module consists of the fifth switch tube S b1 , the sixth switch tube S b2 , the seventh switch tube S b3 , the eighth switch tube S b4 , a second filter capacitor C2, a second excitation inductor L2 and a second three-winding transformer with a voltage ratio of N2; a fifth switch tube S b1 and the eighth switch tube S b4 Diagonal arrangement to achieve coordinated switch control; the sixth switch tube S b2 and the seventh switch tube S b3 Diagonal setting to achieve coordinated switch control, the diagonal switch combination is (S b1 , S b4 ) and (S b2 , S b3 ), realizing switch complementary control;
[0073] The third H-bridge module consists of the ninth switch tube S c1 , the tenth switch tube S c2 , the eleventh switch tube S c3 , the twelfth switch tube S c4, a third filter capacitor C3, a third excitation inductor L3 and a third three-winding transformer with a voltage ratio of N3; a ninth switch tube S c1 and the twelfth switch S c4 Diagonal arrangement to achieve coordinated switch control; the tenth switch tube S c2 and the eleventh switch S c3 Diagonal setting to achieve coordinated switch control, the diagonal switch combination is (S c1 , S c4 ) and (S c2 , S c3 ) to achieve complementary switch control.
[0074] like Figure 2 As shown in the figure, a method for output current control and ripple suppression of a hydrogen production converter based on deep learning is presented. A delay t is added to the original H-bridge circuit of the first, second and third ports of the hydrogen production converter. 1_1 , t 2_2 , t 1_3 , t 3_3 By performing deep learning algorithm optimization, the optimal delay time combination corresponding to the minimum current peak of the hydrogen production port is obtained while ensuring that the actual output current remains stable, thereby achieving the suppression of the current ripple of the hydrogen production port output current, and can provide the electrolyzer with input current more accurately, effectively and stably, with good practical value; specifically, the following steps are included:
[0075] Step 1: The input voltage U1 of the first port, the input voltage U2 of the second port, the voltage U3 of the third port, and the current I of the third port (i.e., the load end) are measured. o sampling;
[0076] Step 2: Based on the given time segment, use the average current method to calculate the output current I of the third port. out and the output current peak I peak ;
[0077] The peak output current of the third port is calculated as follows:
[0078] The time segment relationship after superposition of the third port voltage waveform is:
[0079] t1=t0+t 1_1
[0080] t2=t0+t 1_2
[0081] t3=t0+t 1_2 +t 2_2
[0082] t4=t0+t 1_3
[0083] t5=t0+t 1_3 +t 3_3
[0084] t6=t0+0.5*T
[0085] The initial time is t0, which corresponds to an underestimate of the load current; t1 to t5 are time points corresponding to changes in the load current waveform; and t6 is the time corresponding to the load current reaching a peak value.
[0086] According to the voltage sampling results of the three ports and the load current sampling results in step 1, the expression of the load current in each time period is analyzed:
[0087] i0=-i6
[0088]
[0089] Wherein, U1 is the input voltage of the first port, U2 is the input voltage of the second port, and U3 is the voltage of the third port (load end); L 13 is the equivalent inductance between the first port and the third port; L 23 is the equivalent inductance between the second and third ports, which is calculated as follows:
[0090]
[0091] The average current method is used to solve the problem. Combining the above time segment relationship, the output current I is obtained. out And the peak output current of the load end corresponding to time t6:
[0092]
[0093] The peak current of the load terminal corresponding to time t6 is I peak , then I peak for:
[0094]
[0095] Step 3: Using the deep learning (DRL) control strategy, initializing the parameters, the hydrogen converter environment receiving state S t (V1 * ,V2 * ,V3 * ,I out * ) and action a t (t 1_1 ,t 2_2 ,t 1_3 ,t 3_3), build a data set, find the phase shift time when the actual output current is consistent with the set value, and find I peak The optimal solution of the third port and the optimal delay time t within the third port 1_1 , t 2_2 , t 1_3 , t 3_3 .
[0096] The specific optimization process of the deep learning algorithm is as follows:
[0097] A control strategy based on deep reinforcement learning (DRL) is adopted to achieve stable control of the output current of the hydrogen production converter, while taking into account the suppression of the output current ripple.
[0098] Hydrogen converter environmental receiving state S t (V1 * ,V2 * ,V3 * ,I out * ) and action a t (t 1_1 , t 2_2 , t 1_3 , t 3_3 ) to generate the next state S t+1 Environmental reception: refers to the hydrogen converter receiving V1*, V2*, V3*, I out * and t 1-1 , t 2-2 , t 1-3 , t 3-3 These two sets of data are then optimized through deep learning methods.
[0099] In each round, an action is executed by the action function, which is composed of the sampled state value S t Decision. The action function is defined as follows:
[0100] a t =μ(s t |θ μ )+N
[0101] Here N is random noise, which can enable the agent to explore. μ is the deterministic policy function, θ μ All the values that can be learned in the network are called parameters.
[0102] Step 4: Use the Deep Deterministic Policy Gradient (DDPG) algorithm to calculate the state S t 、S t+1 and action a t To train the agent to track the output current I out Reference value.
[0103] The agent training process is as follows:
[0104] The agent uses the state S through the DDPG algorithm t , S t+1 and action value a t This is used to track the output current reference and calculate the peak output current for each phase shift time combination. In the DDPG algorithm, there are two agents: an actor network and a critic network. Each network consists of two hidden layers, each with 100 neurons.
[0105] Critic Network Q(s t ,a t ∣θ Q ) is determined by the parameter θ Q Indicates that Q is the dynamic-value function, and the input state s t and action value a t , generating an estimated Q(s t ,a t ∣θ Q )value.
[0106] Actor network μ(s t ∣θ μ ) is determined by the parameter θ μ Indicates that according to the input state s t Execute the strategy. The strategy is implemented by using the estimated Q(s t ,a t ∣θ Q ) value to update and generate action value a t .
[0107] Set the action value a t After being placed in the environment of the hydrogen converter, the reward value r t and the next state s t+1 Can be calculated. During the entire training process, all variables {s t ,a t ,r t ,s t+1} are stored in the experience replay buffer. Then, a mini-batch n is randomly sampled from the experience replay buffer to update the Actor network and Critic network in each round.
[0108] Critic Network Q(s t ,a t ∣θ Q ) is updated by minimizing the loss function as follows:
[0109]
[0110] The actor network is updated via a gradient strategy as follows:
[0111]
[0112] In order to ensure the stability and convergence of the learning process, the target critic network Q′(s t ,a t ∣θ Q ) and the target Actor network μ′(s t ∣θ μ ) Perform soft updates on the Critic network and the Actor network. Update the target network weight θ Q′ and θ μ′ The formula is as follows:
[0113] θ Q′ ←τθ Q +(1-τ)θ Q′
[0114] θ μ′ ←τθ μ +(1-τ)θ μ′
[0115] Where τ is the learning rate.
[0116] Target Critic Network Q′(s t ,a t ∣θ Q )’s output y i As shown below:
[0117] y i =r(s,a)+γQ′(s′,μ′(s′)|θ Q′ |s t+1 )
[0118] Where γ is the discount factor.
[0119] To ensure that the calculated output current matches the set value, a reward function is established based on the difference between the calculated value and the set value. The reward function is as follows:
[0120]
[0121] Step 5: Determine whether the difference between the set value and the algorithm output value is less than 0.001. If not, return to step 4 and perform deep learning again.
[0122] The process of judging whether the final training result is the set result is as follows:
[0123] Determine whether the absolute value of the difference between the output current value given by the agent and the set value is less than 0.001. If not, return to step 4 for a new cycle. If satisfied, then according to the combination of multiple phase shift times t given by the agent 1_1 , t 2_2 , t 1_3 , t 3_3 , substitute into the peak current expression to solve the output current peak value. Compare the peak currents, select the lowest value, record it as the minimum peak current value and record this set of phase shift time t 1_1 , t 2_2 , t 1_3 , t 3_3 , input the signal into the three active bridge converter circuit and observe the actual output current value I o .
[0124] Step 6: Phase shift time t based on training completion 1_1 , t 2_2 , t 1_3 , t 3_3 Calculate the output current peak value, compare different peak values, and select the phase shift combination with the smallest output current peak value t 1_1 , t 2_2 , t 1_3 , t 3_3 , observe the actual output current I o size.
[0125] Step 7: Determine whether the difference between the output current setting value and the actual output current value is less than 0.5. If not, return to step 4 and perform deep learning again;
[0126] The process of judging whether the actual output result is consistent with the set value is as follows:
[0127] Determine whether the absolute value of the difference between the actual output current value and the set current value is less than 0.5. If not, the actual output current I o and set value I out * Update the absolute value of the difference I out * , set the updated output current setting value I out * Return to step 4 for a new cycle. Output current setting value I out * The update process satisfies the following formula:
[0128]
[0129] If it satisfies the requirement, the phase shift combination t with the minimum output current peak value is obtained. 1_1 , t 2_2 , t1_3 , t 3_3 , recorded as the phase shift time combination that satisfies the actual output current matching the set value and minimizes the output current ripple. At this point, the hydrogen converter can stably operate in a state where the output current meets the set value and the output current ripple is minimized.
[0130] Step 8: If satisfied, get the optimal phase shift combination t 1_1 , t 2_2 , t 1_3 , t 3_3 , that is, the hydrogen production converter can operate with the lowest output current ripple while maintaining the output current stability.
[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for output current control and ripple suppression of a hydrogen production converter based on deep learning, characterized in that: The following steps are involved: Step 1: sampling the voltages of the three ports of the hydrogen production converter and the current at the load end; Step 2: Based on the given time segments, the output current and the output current peak value of the third port are calculated using the average current method; Step 3: Using the deep learning control strategy, initialize the parameters, and receive the state S of the hydrogen converter environment. t (V1 * ,V2 * ,V3 * ,I out * ) and action a t (t 1_1 ,t 2_2 ,t 1_3 ,t 3_3 ), build a data set, find the phase shift time when the actual output current is consistent with the set value, and find I peak The optimal solution of the third port and the optimal delay time t within the third port 1_1 , t 2_2 , t 1_3 , t 3_3 ; Step 4: Use the deep deterministic policy gradient algorithm to calculate the state S t 、S t+1 and action value a t To train the agent so that it can track the output current I out Reference value; Step 5: Determine whether the difference between the set value and the algorithm output value is less than 0.
001. If not, return to step 4 and perform deep learning again; Step 6: Phase shift time t based on training completion 1_1 , t 2_2 , t 1_3 , t 3_3 Calculate the output current peak value, compare different peak values, and select the phase shift combination with the smallest output current peak value t 1_1 , t 2_2 , t 1_3 , t 3_3 , observe the actual output current I o size; Step 7: Determine whether the difference between the output current setting value and the actual output current value is less than 0.
5. If not, return to step 4 and perform deep learning again; Step 8: If satisfied, get the optimal phase shift combination t 1_1 , t 2_2 , t 1_3 , t 3_3 , that is, the hydrogen production converter can operate with the lowest output current ripple while maintaining the output current stability.
2. The method for controlling output current and suppressing ripple of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 2, the output current I of the third port is calculated. out and the output current peak I peak is calculated as follows: The time segment relationship after the superposition of the three voltage waveforms at the port is: t1=t0+t 1_1 t2=t0+t 1_2 t3=t0+t 1_2 +t 2_2 t4=t0+t 1_3 t5=t0+t 1_3 +t 3_3 t6=t0+0.5*T Among them, the initial time is t0, which corresponds to the low estimate of the load current, t1 to t5 are the time points corresponding to the change of the load current waveform, t6 is the time corresponding to the load current reaching the peak value, t 1_2 Set to a fixed value; According to the sampling results in step 1, the expression of the load current in each time period is analyzed: i0=-i6 Wherein, U1 is the input voltage of the first port, U2 is the input voltage of the second port, and U3 is the voltage of the third port; L 13 is the equivalent inductance between the first port and the third port; L 23 is the equivalent inductance between the second port and the third port; The equivalent inductance is calculated as follows: The average current method is used to solve the problem. Combining the above time segment relationship, the output current I is obtained. out The peak current I corresponding to the load terminal at time t6 peak : The peak current of the load terminal corresponding to time t6 is I peak , then I peak for:
3. The method for output current control and ripple suppression of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 3, the specific optimization process of the deep learning algorithm is as follows: A control strategy based on deep reinforcement learning is used to achieve stable control of the output current of the hydrogen converter and consider the suppression of the output current ripple; the hydrogen converter environment receiving state S t (V1 * ,V2 * ,V3 * ,I out * ) and action a t (t 1_1 , t 2_2 , t 1_3 , t 3_3 ) to generate the next state S t+1 ; In each round, an action is performed through the action function, which is composed of the sampled state value S t Decision; the action function is defined as follows: a t =μ(s t |θ μ )+N Where N is random noise; μ is the deterministic policy function, θ μ As a parameter.
4. The method for output current control and ripple suppression of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 4, the agent training process is as follows: The agent uses the state S through the DDPG algorithm t 、S t+1 and action value a t To track the output current reference and calculate the output current peak value corresponding to each phase shift time combination; In the DDPG algorithm, there are two agents, including the Actor network and the Critic network; each network contains two hidden layers, each with 100 neurons; Critic Network Q(s t ,a t ∣θ Q ) is determined by the parameter θ Q Indicates that Q is the dynamic-value function, and the input state s t and action value a t , generating an estimated Q(s t ,a t ∣θ Q )value; Actor network μ(s t ∣θ μ ) is determined by the parameter θ μ Indicates that according to the input state s t Execution strategy; the strategy is implemented by using the estimated Q(s t ,a t ∣θ Q ) value to update and generate action value a t ; Set the action value a t After being placed in the environment of the hydrogen converter, the reward value r t and the next state s t+1 Able to perform calculations; during the entire training process, all variables {s t ,a t ,r t ,s t+1 } are stored in the experience replay buffer; then, a small batch n is randomly sampled from the experience replay buffer to update the Actor network and Critic network of each round; Critic Network Q(s t ,a t ∣θ Q ) is updated by minimizing the loss function as follows: The actor network is updated via a gradient strategy as follows: In order to ensure the stability and convergence of the learning process, the target critic network Q′(s t ,a t ∣θ Q ) and the target Actor network μ′(s t ∣θ μ ) Perform soft updates on the Critic network and the Actor network; Update target network weights θ Q′ and θ μ′ The formula is as follows, where τ is the learning rate; i Q′ ←tth Q +(1-τ)θ Q′ i μ′ ←tth μ +(1-τ)θ μ′ Target Critic Network Q′(s t ,a t ∣θ Q )’s output y i As shown below, γ is the discount factor: y i =r(s,a)+γQ′(s′,μ′(s′)|θ Q′ |s t+1 ) To ensure that the calculated output current matches the set value, a reward function is established based on the difference between the calculated value and the set value. The reward function is as follows:
5. The method for output current control and ripple suppression of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 5, it specifically includes: judging whether the absolute value of the difference between the output current value given by the intelligent agent and the set value is less than 0.
001. If it is satisfied, the intelligent agent can follow the output current set value. At this time, the intelligent agent training is completed and actual experiments can be carried out; if it is not satisfied, return to step 4 for a new cycle.
6. The method for output current control and ripple suppression of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 6, specifically including: according to the combination t of multiple phase shift times given by the intelligent agent 1_1 , t 2_2 , t 1_3 , t 3_3 , substitute into the peak current expression to solve the output current peak value; compare the peak currents, select the lowest value, record it as the minimum peak current value and record this set of phase shift time t 1_1 , t 2_2 , t 1_3 , t 3_3 , input the signal into the hydrogen production converter circuit and observe the actual output current value I o .
7. The method for controlling output current and suppressing ripple of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 7, it specifically includes: judging whether the absolute value of the difference between the actual output current value and the set current value is less than 0.
5. If it is satisfied, the intelligent agent not only calculates the value by the algorithm but also the actual output can meet the requirement; if not, the intelligent agent calculates the value by the algorithm by the actual output current I o and set value I out * Update the absolute value of the difference I out * , set the updated output current setting value I out * Return to step 4 for a new cycle; output current setting value I out * The update process satisfies the following formula:
8. The method for controlling output current and suppressing ripple of a hydrogen production converter based on deep learning according to claim 1, characterized in that: In step 8, the following steps are specifically performed: according to the phase shift combination t with the minimum output current peak value obtained by searching for the optimal value, 1_1 , t 2_2 , t 1_3 , t 3_3 , recorded as the phase shift time combination that satisfies the matching of the actual output current with the set value and the minimum output current ripple; at this time, the hydrogen production converter can stably operate in a working state where the output current meets the set value and the output current ripple is minimum.
Citation Information
Patent Citations
Current ripple suppression method for hydrogen production converter
CN120016805A