Learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning

By combining imitation learning and deep reinforcement learning in fuel cell hybrid vehicles, constructing the global optimal trajectory and initializing the neural network, the problems of long training cycle and unstable results of traditional deep reinforcement learning are solved, an efficient and stable energy management strategy is achieved, and the total driving cost is reduced.

CN117993293BActive Publication Date: 2025-09-12SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410130880.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-09-12
Estimated Expiration
2044-01-31

AI Technical Summary

Technical Problem

Traditional deep reinforcement learning algorithms have long training cycles, difficulty escaping local optimality, and unstable results in fuel cell hybrid vehicle energy management, resulting in high training costs and poor results.

Method used

An imitation learning algorithm is used to construct the global optimal trajectory. The neural network is initialized through dynamic programming and behavior cloning algorithm. Combined with deep reinforcement learning, the energy management strategy is optimized. The global optimal trajectory obtained by imitation learning is used as the initial policy network of deep reinforcement learning, and training is performed until the reward function converges.

Benefits of technology

It improves training efficiency and optimization effects, reduces total driving costs, enhances the stability and adaptability of energy management, and achieves performance similar to dynamic programming.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117993293B_ABST
    Figure CN117993293B_ABST
Patent Text Reader

Abstract

This invention discloses a learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning. The method includes: constructing a simulation environment, training conditions, and test conditions; extracting the global optimal trajectory of the training condition based on dynamic programming; using an imitation learning algorithm to imitate the global optimal trajectory to obtain inheritable neural network parameters; and using the neural network obtained by the imitation learning algorithm as the initialization policy network for a deep reinforcement learning algorithm, initiating reinforcement learning training until the deep reinforcement learning algorithm converges. This method fully combines the advantages of optimization-based methods and deep reinforcement learning methods, overcoming the shortcomings of traditional deep reinforcement learning algorithms and improving training efficiency and optimization effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an automobile energy management method, and in particular to an energy management method for a learning-type fuel cell hybrid automobile embedded with imitation learning. Background Art

[0002] Fuel cell electric vehicles (FCVs) utilize fuel cell systems (FCSs) as their primary power source, offering excellent range while maintaining zero emissions, thus holding the potential for widespread application. However, FCSs have limited dynamic output characteristics and lack the rapid charge and discharge capabilities of batteries. Fuel cell hybrid electric vehicles (FCHEVs) circumvent this FCV shortcoming by employing an auxiliary power source. This additional power source not only improves driving range but also facilitates energy recovery during braking. With the introduction of diverse power sources, energy management strategies (EMSs) are needed to optimize the distribution of multiple energy flows, thereby coordinating these diverse power sources and achieving energy savings.

[0003] In recent years, reinforcement learning (RL) has been applied to training EMS. During training, the intelligent agent continuously interacts with the environment, seeking the optimal strategy through trial and error. A trained EMS exhibits fast inference speed, enabling real-time application on an onboard controller. Deep reinforcement learning (DRL), a derivative of RL, is characterized by the use of deep neural networks (DNNs) to approximate value functions. Benefiting from the powerful fitting capabilities of DNNs, DRL exhibits stronger learning capabilities compared to RL, can scale to high-dimensional state spaces, and can achieve near-global optimal performance. DRL-based EMS has become a new research direction. However, DRL algorithms also have certain limitations, such as long training cycles, difficulty escaping local optima, and unstable final results. Prolonged training incurs significant time and computational overhead, while unstable training results hinder practical application. Summary of the Invention

[0004] Purpose of the invention: The purpose of the present invention is to provide a learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning, which utilizes imitation learning (IL) to maximize the advantages of optimization-based methods such as global optimality stability and short computation time, to make up for the shortcomings of traditional DRL algorithms and improve training efficiency and optimization effects.

[0005] Technical Solution: The energy management method for a learning-based fuel cell hybrid vehicle embedded with imitation learning described in the present invention includes:

[0006] (1) Construct a simulation environment and build an FCHEV model, including the FCHEV power system structure, fuel cell hydrogen consumption model and life model, and the power battery electric-thermal-life coupling model. The power battery electric-thermal-life coupling model includes a second-order RC electric model, a two-state thermal model, and an energy throughput aging model; construct training data, including training conditions and test conditions;

[0007] (2) Establish a reinforcement learning environment based on the FCHEV model and FCHEV energy management strategy, set the state space, action space and reward function; extract the global optimal trajectory of the training condition based on dynamic programming;

[0008] (3) Using imitation learning algorithms, we imitate the global optimal trajectory and obtain inheritable neural network parameters;

[0009] (4) The neural network that fully imitates the global optimal trajectory is used as the initialization policy network of the deep reinforcement learning algorithm, and reinforcement learning training is started until the reward function converges;

[0010] (5) The trained parameterized neural network strategy is loaded into the FCHEV vehicle controller to realize real-time online application; the target domain FCHEV executes the trained energy management strategy.

[0011] Furthermore, in step (1), the training conditions adopt the China Light Vehicle Test Conditions - Passenger Cars and West Virginia University Interstate Highway Conditions; and the test conditions adopt the standard China City Conditions.

[0012] Furthermore, in step (1), when building the FCHEV model, the efficiency diagram of the quasi-steady-state motor model and the fuel cell output characteristic curve are used as prior knowledge;

[0013] The efficiency diagram of the quasi-steady-state motor model is used to construct the relationship between motor efficiency and wheel speed and torque. The corresponding motor efficiency is obtained by interpolation, thereby solving the required power of the vehicle at any time.

[0014] The fuel cell output characteristic curve is used to construct the relationship between the fuel cell power, hydrogen consumption rate and fuel cell stack efficiency, so as to solve the hydrogen consumption rate at any time.

[0015] Furthermore, step (1) includes:

[0016] (1-1) Construct the power system structure of FCHEV:

[0017] At time step t, the longitudinal traction force F of the vehicle t The calculation is as follows:

[0018]

[0019] Where m is the total mass of the vehicle; g is the acceleration due to gravity; f is the rolling resistance coefficient; θ is the road slope; A is the area in front of the vehicle; C D is the air resistance coefficient; v h is the speed of the vehicle; δ is the rotational mass coefficient; a h is the vehicle’s acceleration;

[0020] Wheel speed W w and drive shaft torque T w It is expressed as follows:

[0021]

[0022] Among them, r w is the wheel radius;

[0023] Motor speed W m and torque T m The calculation is as follows:

[0024]

[0025] Among them, R fd is the final drive gear ratio; η fd is the efficiency of the drive shaft;

[0026] The power P required by the vehicle req The calculation is as follows:

[0027]

[0028] Among them, η m is the motor efficiency, obtained by interpolation of the efficiency diagram of the quasi-steady-state motor model; correspondingly, P req can be represented as follows:

[0029] P req =P DC / DC +P bat (5)

[0030] Among them, P DC / DC is the output power of the DC / DC converter, P bat It is the power of lithium-ion battery pack, including charging and discharging process;

[0031] (1-2) Constructing FCHEV fuel cell hydrogen consumption model and life model:

[0032] Hydrogen consumption rate of fuel cell stack The calculation is as follows:

[0033]

[0034] Among them, L v Indicates the lower calorific value of hydrogen; η fcs Indicates the efficiency of the fuel cell stack; power P fcs and hydrogen combustion rate and efficiency η fcs The relationship between them is represented by the fuel cell output characteristic curve;

[0035] Overall performance degradation of the fuel cell system fcs It can be expressed by discrete expressions for four different types of adverse driving conditions with load variation:

[0036]

[0037] Where n is the number of time steps, d ss (t), d low (t), d high (t), d cha (t) is the performance degradation caused by the start-stop condition, low power condition, high power load and load change condition at time t.

[0038] Furthermore, step (1) further includes:

[0039] (1-3) In the second-order RC electrical model, two RC branches are used to simulate the polarization effect, and its governing equation is as follows:

[0040]

[0041]

[0042]

[0043] V t (t) = V oc (SoC)+V p1 (t)+V p2 (t)+R S I(t) (11)

[0044] Where SoC(t) represents the state of charge at time step t; I(t) is the load current at time step t; C n Indicates the capacitance of the fuel cell; V t (t) is the terminal voltage at time step t; V p1 and V p2 is the polarization voltage across the RC branch, which is determined by the capacitor C p1 、C p2 and resistor R p1 、R P2 parameterization; R SRepresents the surface resistance of the battery; V oc (SoC) represents the open circuit voltage, which is a function of SoC.

[0045] Furthermore, steps (1-3) further include:

[0046] In the two-state thermal model, according to the principle of conservation of thermal energy, the following equation is given:

[0047]

[0048]

[0049]

[0050] Among them, C c is the equivalent thermal capacitance of the battery cell; C s is the equivalent thermal capacitance of the battery surface; T s (t) is the battery surface temperature; T c (t) is the battery core temperature; T a (t) is the average temperature inside the battery; T f (t) is the ambient temperature; R c is the thermal resistance caused by heat conduction inside the battery; R u is the thermal resistance caused by convection on the battery surface;

[0051] The heat generation rate, which is a combination of ohmic heat, polarization heat, and irreversible entropy heat, is represented by H(t):

[0052] H(t)=I(t)[V p1 (t)+V p2 (t)+R s I(t)]+I(t)[T a (t)+273]E n (SoC,t) (15)

[0053] Among them, E n It represents the entropy change during the electrochemical reaction.

[0054] Furthermore, steps (1-3) further include:

[0055] The energy throughput aging model is used to assess battery degradation. It assumes that the battery can withstand a certain amount of cumulative charge flow before it is scrapped. Therefore, the dynamic calculation of the battery health SOH is as follows:

[0056]

[0057] Among them, ΔSOH t Indicates the change in battery health; Δt indicates the current duration, N(c,T a) is the equivalent number of cycles until the battery system reaches the end of its life; the capacity loss empirical model based on the Arrhenius equation takes into account the effects of discharge rate C-rate (c) and internal temperature, and the equation is as follows:

[0058]

[0059] Where, ΔC n is the percentage of capacity loss; B(c) represents the pre-exponential factor; E a represents activation energy; R is the ideal gas constant; Ah represents ampere-hour throughput; z is the power law factor equal to 0.55;

[0060] E a (c)=31700-370.3·c (18)

[0061] c represents the battery charge and discharge rate;

[0062] When C n When it drops by 20%, the battery will reach the end of its life; the derivation of Ah and N is as follows:

[0063]

[0064] N(c,T a )=3600·Ah(c,T a ) / C n (20)

[0065] Finally, the change in battery health SOH can be calculated based on the given current, temperature and battery dynamics using Equation (16) to understand the aging of the battery pack.

[0066] Furthermore, in step (2), the state space is defined as follows:

[0067] s=[SOC,SOH bat ,SOH fcs ,P bat ,P fcs ,v h ,a h ] (twenty one)

[0068] Among them, SOC is the state of charge of the battery; SOH bat Is the health status of the power battery; SOH fcs is the health status of the fuel cell stack; P bat Power battery power; P fcs is the power of the fuel cell stack; v h is the vehicle speed; a h is the vehicle acceleration;

[0069] Define the action space as the output power of the fuel cell system:

[0070] a=P fcs ∈[0,60]kW (22)

[0071] The reward function is defined as follows:

[0072]

[0073] Among them, ρ1, ρ2, and ρ3 are the hydrogen price, fuel cell system replacement price, and power battery pack replacement price respectively; the weight coefficient ω determines the relative importance of capital cost to the battery SOC value; SOC ref is the reference value of SOC;

[0074] After defining the state space, action space and reward function, the global optimal trajectory Tr of the energy management strategy under training conditions can be solved through dynamic programming. * , the global optimal trajectory Tr * The state-action pair (s t ,a t ) composition, s t represents the state at time t, a t Represents the action at time t.

[0075] Furthermore, in step (3), the imitation learning algorithm adopts a behavioral cloning algorithm, which specifically includes:

[0076] (3-1) Initialize the Actor network of the behavior cloning algorithm;

[0077] (3-2) The goal of the behavior cloning algorithm is to recover the expert strategy from one or more expert trajectories. If a complete trajectory is represented by Tr * , then the dataset consisting of m trajectories can be expressed as The Tr obtained in step (2) * As

[0078] (3-3) Use Gaussian distribution to represent the strategy of continuous action space. For each state s, the strategy It can be expressed as a Gaussian policy distribution:

[0079]

[0080] in, represents the probability distribution of selecting all actions in state s, represents Gaussian distribution, μ θ (s) represents the mean of the Gaussian distribution, represents the variance of the Gaussian distribution;

[0081] (3-4) Establish an optimization problem to make the Gaussian strategy generated by the Actor network closest to the expert strategy:

[0082]

[0083] in, represents the mean of the Gaussian strategy, represents the variance of the Gaussian strategy;

[0084] (3-5) Calculate the gradient using the gradient descent method Train the Actor network, trying to approach the optimal solution to the optimization problem until the pre-set maximum number of iterations is reached, the training ends, and then the neural network is saved.

[0085] Furthermore, in step (4), the deep reinforcement learning algorithm uses the proximal strategy to optimize PPO, specifically including:

[0086] (4-1) The Actor network fully simulated by the behavioral cloning algorithm is used as the initialization strategy network of PPO, and the value network is randomly initialized at the same time;

[0087] (4-2) Let the strategy before each strategy update be π old ,π old Interact with the environment for a fixed number of steps to obtain multiple state-action pairs (s, a);

[0088] (4-3) Calculate for all state-action pairs (s, a), the specific action a in state s and the policy π old The relative improvement in total reward obtained compared to a randomly selected action in (·|s)

[0089] (4-4) Optimize the agent objective function of the policy network. The objective function L is expressed as:

[0090]

[0091] Among them, π θ represents the new policy network parameterized by θ obtained by updating the current policy, π old represents the old policy network obtained after the last policy update. In formula (26), a clipping threshold ∈ ≥ 0 is used to control the size of each policy update and calculate the gradient And continuously update the parameters θ of the policy network by solving the following optimization problem through gradient descent method:

[0092]

[0093] (4-5) Optimize the objective function of the value network, and its optimization goal is:

[0094]

[0095] Among them, V Φ Represents a PPO value network parameterized by Φ, calculating the gradient The value network parameters Φ are also updated using the gradient descent method:

[0096]

[0097] (4-6) Repeat steps (4-2) to (4-5) until the preset maximum number of iterations is reached, the training ends, and then the neural network parameters are saved and downloaded.

[0098] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:

[0099] This paper designs an imitation learning (IL)-embedded deep reinforcement learning (DRL) framework for the energy management strategy (EMS) of fuel cell hybrid electric vehicles (FCHEVs). It first uses imitation learning (IL) to imitate the global optimal trajectory obtained by dynamic programming (DP). Then, the learned neural network is used as the initial policy network for deep reinforcement learning (DRL), and DRL training is started and continued until convergence. The present invention fully combines the advantages of optimization-based methods and deep reinforcement learning methods, perfectly compensating for the shortcomings of the DRL method, can improve training efficiency and optimization effects, and has strong potential for guiding training. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 This is a block diagram of a learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning provided by an embodiment of the present application;

[0101] Figure 2 is a diagram of motor efficiency in an embodiment of the present application;

[0102] Figure 3 is the fuel cell output characteristic curve in the embodiment of the present application;

[0103] Figure 4 This is a diagram of a power battery model in an embodiment of the present application;

[0104] Figure 5 (a) is the training condition, Figure 5 (b) in the figure is the test condition;

[0105] Figure 6 This is a structural diagram of the fuel cell hybrid vehicle power system in an embodiment of the present application. DETAILED DESCRIPTION

[0106] The present invention will be further described below with reference to the accompanying drawings.

[0107] like Figure 1 As shown, the embodiment of the present application provides a learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning, which specifically includes the following steps:

[0108] (1) Using Python, a simulation environment for a fuel cell hybrid electric vehicle (FCHEV) model and energy management strategy (EMS) was built. The FCHEV model includes the FCHEV power system structure, the fuel cell hydrogen consumption model and life model, and the power battery electric-thermal-life coupling model. Training data, including training conditions and test conditions, was constructed to serve as FCHEV driving data.

[0109] When building the FCHEV model, the efficiency diagram of the preloaded quasi-steady-state motor model and the fuel cell output characteristic curve are used as prior knowledge; the efficiency diagram of the quasi-steady-state motor model is used to construct the relationship between the motor efficiency and the wheel speed and torque, such as Figure 2 As shown, the corresponding motor efficiency is obtained by interpolation, thereby solving the required power of the vehicle at any time; the fuel cell output characteristic curve is used to construct the relationship between the fuel cell power and hydrogen consumption rate and the fuel cell stack efficiency, as shown Figure 3 As shown, the hydrogen consumption rate at any time is solved.

[0110] The lithium-ion battery pack is simulated by a power battery electric-thermal-life coupling model consisting of a second-order RC electric model, a two-state thermal model, and an energy throughput aging model, so that the battery health (SOH) of the battery at any time can be solved. The battery model is as follows: Figure 4 shown.

[0111] The training dataset consists of highway conditions and urban road conditions. Figure 5 The training conditions shown in (a) of FIG. 1 cover low-speed to high-speed conditions, so that the training results of the present invention can be used on various roads. The training conditions include the China Light Vehicle Test Conditions - Passenger Cars (CLTC-P) and the West Virginia University Interstate Highway (WVU-INTER) conditions. In addition, the following conditions are also used: Figure 5 The standard China City condition shown in (b) is used as a test condition to test the robustness of the obtained strategy.

[0112] Step (1) specifically includes the following steps:

[0113] (1-1) Construct the power system structure of FCHEV:

[0114] The research object of the embodiment of the present application is a fuel cell hybrid electric bus, whose power system structure is as follows: Figure 6 shown.

[0115] At time step t, the longitudinal traction force F of the vehicle t The calculation is as follows:

[0116]

[0117] Where m is the total mass of the vehicle; g is the acceleration due to gravity; f is the rolling resistance coefficient; θ is the road slope; A is the area in front of the vehicle; C D is the air resistance coefficient; v h is the speed of the vehicle; δ is the rotational mass coefficient; a h is the acceleration of the vehicle; the speed v of the vehicle in the simulation scene h and acceleration a h Obtained through the interactive interface;

[0118] Wheel speed W w and drive shaft torque T w It is expressed as follows:

[0119]

[0120] Among them, r w is the wheel radius;

[0121] Motor speed W m and torque T m The calculation is as follows:

[0122]

[0123] Among them, R fd is the final drive gear ratio; η fd is the efficiency of the drive shaft;

[0124] The power P required by the vehicle req The calculation is as follows:

[0125]

[0126] Among them, η m is the motor efficiency, obtained by interpolation of the efficiency diagram of the quasi-steady-state motor model; correspondingly, P req can be represented as follows:

[0127] P req =P DC / DC +P bat (5)

[0128] Among them, P DC / DC It is a DC / DC converter (i.e. Figure 6 The output power of the DC-DC transformer shown in the figure, P bat It is the power of lithium-ion battery pack, including charging and discharging process;

[0129] (1-2) Constructing FCHEV fuel cell hydrogen consumption model and life model:

[0130] The fuel cell, as the main power source of the FCHEV, converts the chemical energy of hydrogen and oxygen into electrical energy. The fuel cell hydrogen consumption model and life model are used to construct the fuel cell stack.

[0131] Hydrogen consumption rate of fuel cell stack The calculation is as follows:

[0132]

[0133] Among them, L v Indicates the lower calorific value of hydrogen, equal to 120kJ / g; η fcs Indicates the efficiency of the fuel cell stack; power P fcs and hydrogen combustion rate and efficiency η fcs The relationship between them is represented by the fuel cell output characteristic curve.

[0134] Overall performance degradation of the fuel cell system fcs It can be expressed by discrete expressions for four different types of adverse driving conditions with load variation:

[0135]

[0136] Where n is the number of time steps, d ss (t), d low (t), d high (t), d cha (t) is the performance degradation caused by the start-stop condition, low power condition, high power load and load change condition at time t.

[0137] (1-3) The power battery pack, as the FCHEV's second energy storage device, provides peak power for the vehicle and smooths the fuel cell system's output. A coupled electrical-thermal-lifetime model is used to construct the power battery system, consisting of three sub-models: a second-order RC electrical model, a two-state thermal model, and an energy throughput aging model.

[0138] In the second-order RC electrical model, two RC branches are used to simulate the polarization effect, and its governing equation is as follows:

[0139]

[0140]

[0141]

[0142] V t (t) = V oc (SoC)+V p1 (t)+V p2 (t)+R S I(t) (11)

[0143] Where SoC(t) represents the state of charge at time step t; I(t) is the load current at time step t; C n Indicates the capacitance of the fuel cell; V t (t) is the terminal voltage at time step t; V p1 and V p2 is the polarization voltage across the RC branch, which is determined by the capacitor C p1 、C p2 and resistor R p1 、R P2 parameterization; R S Represents the surface resistance of the battery; V oc (SoC) represents the open circuit voltage, which is a function of SoC;

[0144] In the two-state thermal model, according to the principle of conservation of thermal energy, the following equation is given:

[0145]

[0146]

[0147]

[0148] Among them, C c is the equivalent thermal capacitance of the battery cell; C s is the equivalent thermal capacitance of the battery surface; T s (t) is the battery surface temperature, T c (t) is the battery core temperature, T a (t) is the average temperature inside the battery, T f (t) is the ambient temperature, in °C; R c is the thermal resistance caused by heat conduction inside the battery; R u is the thermal resistance caused by convection on the battery surface;

[0149] The heat generation rate, which is a combination of ohmic heat, polarization heat, and irreversible entropy heat, is represented by H(t):

[0150] H(t)=I(t)[V p1 (t)+V p2 (t)+R sI(t)]+I(t)[T a (t)+273]E n (SoC,t) (15)

[0151] Among them, E n It represents the entropy change during the electrochemical reaction.

[0152] The energy throughput aging model is used to assess battery degradation. It assumes that the battery can withstand a certain amount of cumulative charge flow before it fails. Therefore, the dynamic calculation of the battery health (SOH) is as follows:

[0153]

[0154] Among them, ΔSOH t Indicates the change in battery health; Δt indicates the current duration, N(c,T a ) is the equivalent number of cycles until the battery system reaches the end of its life; the capacity loss empirical model based on the Arrhenius equation takes into account the effects of discharge rate C-rate (c) and internal temperature, and the equation is as follows:

[0155]

[0156] Where, ΔC n is the percentage of capacity loss; B(c) represents the pre-exponential factor; E a represents the activation energy in J / mol; R is the ideal gas constant, which is equal to 8.314 J / (mol·K); Ah represents the ampere-hour throughput; z is the power law factor equal to 0.55;

[0157] E a (c)=31700-370.3·c (18)

[0158] c represents the battery charge and discharge rate;

[0159] When C n When it drops by 20%, the battery will reach the end of its life; the derivation of Ah and N is as follows:

[0160]

[0161] N(c,T a )=3600·Ah(c,T a ) / C n (20)

[0162] Finally, the change in battery health (SOH) can be calculated based on the given current, temperature, and battery dynamics using Equation (16) to understand the aging of the battery pack.

[0163] (2) Establish a reinforcement learning environment based on the FCHEV model and the FCHEV (fuel cell hybrid vehicle) energy management strategy (EMS), set the state space, action space and reward function; based on dynamic programming (DP), extract the global optimal trajectory of the training condition;

[0164] Combining the speed, acceleration, and battery SOC information in the FCHEV model and energy management strategy, the state space is defined as follows:

[0165] s=[SOC,SOH bat ,SOH fcs ,P bat ,P fcs ,v h ,a h ] (twenty one)

[0166] Among them, SOC is the state of charge of the battery; SOH bat Is the health status of the power battery; SOH fcs is the health status of the fuel cell stack; P bat Power battery power; P fcs is the power of the fuel cell stack; v h is the vehicle speed; a h is the vehicle acceleration;

[0167] Define the action space as the output power of the fuel cell system:

[0168] a=P fcs ∈[0,60]kW (22)

[0169] The main optimization objectives of the learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning provided in the embodiments of this application include: (1) reducing the total driving cost, which is composed of hydrogen consumption, fuel cell degradation, and power battery aging. (2) maintaining the SOC near the initial SOC. Therefore, the reward function is defined as follows:

[0170]

[0171] Among them, ρ1, ρ2, and ρ3 are the hydrogen price, fuel cell system replacement price, and power battery pack replacement price, respectively, which means that the first two objectives can be normalized by capital cost; the weight coefficient ω determines the relative importance of capital cost to the battery SOC value; SOC ref is the reference value of SOC, which is taken as 0.5;

[0172] After defining the state space, action space and reward function, the global optimal trajectory Tr of the energy management strategy (EMS) under the training condition can be solved by dynamic programming (DP) * , the global optimal trajectory Tr* The state-action pair (s t ,a t ) composition, s t represents the state at time t, a t Represents the action at time t.

[0173] (3) Using the imitation learning algorithm (IL), the global optimal trajectory is imitated to obtain inheritable neural network parameters;

[0174] In the embodiment of the present application, the behavior cloning algorithm (BC) is used as the algorithm of the imitation learning (IL) part of the framework, which specifically includes:

[0175] (3-1) Initialize the Actor network of the Behavior Cloning Algorithm (BC);

[0176] (3-2) The goal of the behavior cloning algorithm (BC) is to recover the expert strategy from one or more expert trajectories. If a complete trajectory is represented by Tr * , then the dataset consisting of m trajectories can be expressed as The Tr obtained in step (2) * As

[0177] (3-3) Use Gaussian distribution to represent the strategy of continuous action space. Specifically, for each state s, the strategy It can be expressed as a Gaussian policy distribution:

[0178]

[0179] in, represents the probability distribution of selecting all actions in state s, represents Gaussian distribution, μ θ (s) represents the mean of the Gaussian distribution, represents the variance of the Gaussian distribution;

[0180] (3-4) Establish an optimization problem to make the Gaussian strategy generated by the Actor network closest to the expert strategy:

[0181]

[0182] in, represents the mean of the Gaussian strategy, represents the variance of the Gaussian strategy;

[0183] (3-5) Calculate the gradient using the gradient descent method Train the Actor network, trying to approach the optimal solution to the optimization problem until the pre-set maximum number of iterations is reached, the training ends, and then the neural network is saved.

[0184] (4) The neural network that fully imitates the global optimal trajectory is used as the initialization policy network of the deep reinforcement learning algorithm (DRL) and reinforcement learning training is started until the reward function converges;

[0185] By fully imitating the expert trajectory of IL, a trained and inheritable neural network can be obtained, which is used as the initialization policy network of DRL to start reinforcement learning training. In the embodiment of this application, the proximal policy optimization PPO is used as the algorithm of the deep reinforcement learning (DRL) part of the framework, specifically including:

[0186] (4-1) The Actor network fully simulated by the behavioral cloning algorithm (BC) is used as the initialization strategy network of PPO, and the value network is randomly initialized;

[0187] (4-2) Let the strategy before each strategy update be π old ,π old Interact with the environment for a fixed number of steps to obtain multiple state-action pairs (s, a);

[0188] (4-3) Calculate for all state-action pairs (s, a), the specific action a in state s and the policy π old The relative improvement in total reward obtained compared to a randomly selected action in (·|s)

[0189] (4-4) Optimize the agent objective function of the policy network. The objective function L is expressed as:

[0190]

[0191] Among them, π θ represents the new policy network parameterized by θ obtained by updating the current policy, π old represents the old policy network obtained after the last policy update. In formula (26), a clipping threshold ∈ ≥ 0 is used to control the size of each policy update and calculate the gradient And continuously update the parameters θ of the policy network by solving the following optimization problem through gradient descent method:

[0192]

[0193] (4-5) Optimize the objective function of the value network, and its optimization goal is:

[0194]

[0195] Among them, V Φ Represents a PPO value network parameterized by Φ, calculating the gradient The value network parameters Φ are also updated using the gradient descent method:

[0196]

[0197] (4-6) Repeat steps (4-2) to (4-5) until the preset maximum number of iterations is reached, the training ends, and then the neural network parameters are saved and downloaded.

[0198] (5) The trained parameterized neural network strategy is loaded into the FCHEV vehicle controller to realize real-time online application; the target domain FCHEV executes the trained energy management strategy (EMS).

[0199] In the energy management method of a learning-based fuel cell hybrid vehicle embedded with imitation learning provided in the embodiment of the present application, the EMS takes into account the degradation of the fuel cell and the aging of the power battery, with the ultimate goal of reducing the total driving cost; Figure 1 The EMS of the IL-embedded DRL framework shown in the figure outperforms the traditional DRL method in terms of training efficiency and optimization effect. It outperforms the separate PPO-based EMS in terms of training efficiency, fuel economy, and total driving cost, achieving performance very close to the DP benchmark. The simulation results in different driving conditions demonstrate its good adaptability.

Claims

1. A learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning, characterized in that: include: (1) Construct a simulation environment and build an FCHEV model, including the FCHEV power system structure, fuel cell hydrogen consumption model and life model, and the power battery electric-thermal-life coupling model. The power battery electric-thermal-life coupling model includes a second-order RC electric model, a two-state thermal model, and an energy throughput aging model; construct training data, including training conditions and test conditions; (2) Establish a reinforcement learning environment based on the FCHEV model and FCHEV energy management strategy, set the state space, action space and reward function; extract the global optimal trajectory of the training condition based on dynamic programming; (3) Using imitation learning algorithms, we imitate the global optimal trajectory and obtain inheritable neural network parameters; (4) The neural network that fully imitates the global optimal trajectory is used as the initialization policy network of the deep reinforcement learning algorithm, and reinforcement learning training is started until the reward function converges; In step (4), the deep reinforcement learning algorithm uses the proximal strategy to optimize PPO, which specifically includes: (4-1) The Actor network fully simulated by the behavioral cloning algorithm is used as the initialization strategy network of PPO, and the value network is randomly initialized at the same time; (4-2) Let the strategy before each strategy update be π old ,π old Interact with the environment for a fixed number of steps to obtain multiple state-action pairs (s, a); (4-3) Calculate for all state-action pairs (s, a), the specific action a in state s and the policy π old The relative improvement in total reward obtained compared to a randomly selected action in (·|s) (4-4) Optimize the agent objective function of the policy network. The objective function L is expressed as: Among them, π θ represents the new policy network parameterized by θ obtained by updating the current policy, π old represents the old policy network obtained after the last policy update. In formula (26), a clipping threshold ∈ ≥ 0 is used to control the size of each policy update and calculate the gradient And continuously update the parameters θ of the policy network by solving the following optimization problem through gradient descent method: (4-5) Optimize the objective function of the value network, and its optimization goal is: Among them, V Φ Represents a PPO value network parameterized by φ, calculating the gradient The value network parameter φ is also updated by gradient descent: (4-6) Repeat steps (4-2) to (4-5) until the preset maximum number of iterations is reached, the training ends, and then save and download the neural network parameters; (5) The trained parameterized neural network strategy is loaded into the FCHEV vehicle controller to realize real-time online application; the target domain FCHEV executes the trained energy management strategy.

2. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 1 is characterized in that: In step (1), the training conditions adopt the China Light Vehicle Test Conditions - Passenger Cars and West Virginia University Interstate Highway Conditions; the test conditions adopt the standard China City Conditions.

3. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 1 is characterized in that: In step (1), when building the FCHEV model, the efficiency diagram of the quasi-steady-state motor model and the fuel cell output characteristic curve are used as prior knowledge; The efficiency diagram of the quasi-steady-state motor model is used to construct the relationship between motor efficiency and wheel speed and torque. The corresponding motor efficiency is obtained by interpolation, thereby solving the required power of the vehicle at any time. The fuel cell output characteristic curve is used to construct the relationship between the fuel cell power, hydrogen consumption rate and fuel cell stack efficiency, so as to solve the hydrogen consumption rate at any time.

4. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 3 is characterized in that: Step (1) includes: (1-1) Construct the power system structure of FCHEV: At time step t, the longitudinal traction force F of the vehicle t The calculation is as follows: Where m is the total mass of the vehicle; g is the acceleration due to gravity; f is the rolling resistance coefficient; θ is the road slope; A is the area in front of the vehicle; C D is the air resistance coefficient; v h is the speed of the vehicle; δ is the rotational mass coefficient; a h is the vehicle’s acceleration; Wheel speed W w and drive shaft torque T w It is expressed as follows: Among them, r w is the wheel radius; Motor speed W m and torque T m The calculation is as follows: Among them, R fd is the final drive gear ratio; η fd is the efficiency of the drive shaft; The power P required by the vehicle req The calculation is as follows: Among them, η m is the motor efficiency, obtained by interpolation of the efficiency diagram of the quasi-steady-state motor model; correspondingly, P req can be represented as follows: P req =P DC / DC +P bat (5) Among them, P DC / DC is the output power of the DC / DC converter, P bat It is the power of lithium-ion battery pack, including charging and discharging process; (1-2) Constructing FCHEV fuel cell hydrogen consumption model and life model: Hydrogen consumption rate of fuel cell stack The calculation is as follows: Among them, L v Indicates the lower calorific value of hydrogen; η fcs Indicates the efficiency of the fuel cell stack; power P fcs and hydrogen combustion rate and efficiency η fcs The relationship between them is represented by the fuel cell output characteristic curve; Overall performance degradation of the fuel cell system fcs It can be expressed by discrete expressions for four different types of adverse driving conditions with load variation: Where n is the number of time steps, d ss (t), d low (t), d high (t), d cha (t) is the performance degradation caused by the start-stop condition, low power condition, high power load and load change condition at time t.

5. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 4 is characterized in that: Step (1) further includes: (1-3) In the second-order RC electrical model, two RC branches are used to simulate the polarization effect, and its governing equation is as follows: V t (t)=V oc (SoC)+V p1 (t)+V p2 (t)+R S I(t) (11) Where SoC(t) represents the state of charge at time step t; I(t) is the load current at time step t; C n Indicates the capacitance of the fuel cell; V t (t) is the terminal voltage at time step t; V p1 and V p2 is the polarization voltage across the RC branch, which is determined by the capacitor C p1 、C p2 and resistor R p1 、R P2 parameterization; R S Represents the surface resistance of the battery; V oc (SoC) represents the open circuit voltage, which is a function of SoC.

6. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 5 is characterized in that: Steps (1-3) also include: In the two-state thermal model, according to the principle of conservation of thermal energy, the following equation is given: Among them, C c is the equivalent thermal capacitance of the battery cell; C s is the equivalent thermal capacitance of the battery surface; T s (t) is the battery surface temperature; T c (t) is the battery core temperature; T a (t) is the average temperature inside the battery; T f (t) is the ambient temperature; R c is the thermal resistance caused by heat conduction inside the battery; R u is the thermal resistance caused by convection on the battery surface; The heat generation rate, which is a combination of ohmic heat, polarization heat, and irreversible entropy heat, is represented by H(t): H(t)=I(t)[V p1 (t)+V p2 (t)+R s I(t)]+I(t)[T a (t)+273]E n (SoC,t) (15) Among them, E n It represents the entropy change during the electrochemical reaction.

7. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 6 is characterized in that: Steps (1-3) also include: The energy throughput aging model is used to assess battery degradation. It assumes that the battery can withstand a certain amount of cumulative charge flow before it is scrapped. Therefore, the dynamic calculation of the battery health SOH is as follows: Among them, ΔSOH t Indicates the change in battery health; Δt indicates the current duration, N(c,T a ) is the equivalent number of cycles until the battery system reaches the end of its life; the capacity loss empirical model based on the Arrhenius equation takes into account the effects of discharge rate C-rate (c) and internal temperature, and the equation is as follows: Where, ΔC n is the percentage of capacity loss; B(c) represents the pre-exponential factor; E a represents activation energy; R is the ideal gas constant; Ah represents ampere-hour throughput; z is the power law factor equal to 0.55; E a (c)=31700-370.3·c (18) c represents the battery charge and discharge rate; When C n When it drops by 20%, the battery will reach the end of its life; the derivation of Ah and N is as follows: N(c,T a )=3600·Ah(c,T a ) / C n (20) Finally, the change in battery health SOH can be calculated based on the given current, temperature and battery dynamics using Equation (16) to understand the aging of the battery pack.

8. The method for energy management of a fuel cell hybrid vehicle embedded with imitation learning according to any one of claims 4 to 7, characterized in that: In step (2), the state space is defined as follows: s=[SOC,SOH bat ,SOH fcs ,P bat ,P fcs ,v h ,a h ] (21) Among them, SOC is the state of charge of the battery; SOH bat Is the health status of the power battery; SOH fcs is the health status of the fuel cell stack; P bat Power battery power; P fcs is the power of the fuel cell stack; v h is the vehicle speed; a h is the vehicle acceleration; Define the action space as the output power of the fuel cell system: a=P fcs ∈[0,60]kW (22) The reward function is defined as follows: Among them, ρ1, ρ2, and ρ3 are the hydrogen price, fuel cell system replacement price, and power battery pack replacement price respectively; the weight coefficient ω determines the relative importance of capital cost to the battery SOC value; SOC ref is the reference value of SOC; After defining the state space, action space and reward function, the global optimal trajectory Tr of the energy management strategy under training conditions can be solved through dynamic programming. * , the global optimal trajectory Tr * The state-action pair (s t ,a t ) composition, s t represents the state at time t, a t Represents the action at time t.

9. The learning-based fuel cell hybrid vehicle energy management method embedded with imitation learning according to claim 8, characterized in that: In step (3), the imitation learning algorithm adopts the behavioral cloning algorithm, which specifically includes: (3-1) Initialize the Actor network of the behavior cloning algorithm; (3-2) The goal of the behavior cloning algorithm is to recover the expert strategy from one or more expert trajectories. If a complete trajectory is represented by Tr * , then the dataset consisting of m trajectories can be expressed as The Tr obtained in step (2) * As (3-3) Use Gaussian distribution to represent the strategy of continuous action space. For each state s, the strategy It can be expressed as a Gaussian policy distribution: in, represents the probability distribution of selecting all actions in state s, represents Gaussian distribution, μ θ (s) represents the mean of the Gaussian distribution, represents the variance of the Gaussian distribution; (3-4) Establish an optimization problem to make the Gaussian strategy generated by the Actor network closest to the expert strategy: in, represents the mean of the Gaussian strategy, represents the variance of the Gaussian strategy; (3-5) Calculate the gradient using the gradient descent method Train the Actor network, trying to approach the optimal solution to the optimization problem until the pre-set maximum number of iterations is reached, the training ends, and then the neural network is saved.

Citation Information

Patent Citations

  • Hybrid electric vehicle hierarchical prediction energy management method fused with deep reinforcement learning

    CN113525396A

  • Energy system control

    US20220161687A1