Eco-driving cooperative control method for fuel cell vehicles based on SAC algorithm

By optimizing the power distribution of fuel cell vehicles using a neural network based on the SAC algorithm, the response lag problem of fuel cell vehicles when the load changes is solved, coordinated control of ecological driving and vehicle following is achieved, and the following performance and comfort of fuel cell vehicles are improved.

CN118991823BActive Publication Date: 2025-09-19CHONGQING JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411175997.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-09-19
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

When the load of fuel cell vehicles changes dramatically, the reaction gas response lags and regenerative braking energy cannot be recovered. In addition, the existing energy management strategies have limited effectiveness in complex driving environments, resulting in increased system complexity.

Method used

A neural network based on the SAC algorithm is used for eco-driving collaborative control of fuel cell vehicles. By acquiring the target vehicle state information and following vehicle information, the state space, action space and reward function are constructed to optimize the power distribution of the fuel cell vehicle and achieve collaborative control of the vehicle acceleration and power system output power.

Benefits of technology

It achieves good following performance and comfort of fuel cell vehicles in complex driving environments, reduces the life degradation of the fuel cell stack, and improves adaptability to working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118991823B_ABST
    Figure CN118991823B_ABST
Patent Text Reader

Abstract

The present invention provides a fuel cell vehicle eco-driving collaborative control method based on the SAC algorithm, comprising the following steps: obtaining target vehicle status information, and determining control requirement information of the target vehicle based on the vehicle status information; constructing a SAC-based neural network in a controller, including a state space and an action space for eco-driving, and determining a reward function based on the control requirement information; inputting the target vehicle status information into the SAC-based neural network for offline training, and obtaining the real-time status information of the target vehicle and inputting it into the trained SAC neural network to determine an optimal control strategy. The method can determine the power distribution state of the fuel cell vehicle based on the actual status information of the target fuel cell vehicle and the following vehicle information, thereby deciding the acceleration of the entire vehicle and the output power of the power system, thereby effectively realizing the collaborative control of eco-driving and vehicle following.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vehicle driving control method, and in particular to a fuel cell vehicle eco-driving collaborative control method based on a SAC algorithm. Background Art

[0002] Hydrogen has the characteristics of high calorific value, cleanness and carbon-free, and is regarded as a key energy source for solving global energy and environmental problems. Moreover, hydrogen refueling is fast and the conversion efficiency is high, making it the main power source for fuel cell vehicles (FCVs). This makes FCVs have the advantages of both internal combustion engine vehicles and electric vehicles, making them the most promising new energy vehicle model in the future.

[0003] However, when the load of FCVs changes dramatically, the response of the reaction gas will lag behind the power demand. At the same time, the one-way power output makes it impossible to recover regenerative braking energy. To solve this technical problem, FCVs usually use supercapacitors or power batteries as auxiliary power sources to form a hybrid system.

[0004] Because multiple energy sources increase system complexity, a stable, efficient, and reliable energy management strategy is crucial for improving fuel cell vehicle performance. Typically, energy management optimizes powertrain energy distribution at a specific vehicle speed, resulting in limited energy-saving effects. Furthermore, the complexity and uncertainty of the driving environment significantly impact energy management performance.

[0005] Therefore, in order to solve the above technical problems, it is urgent to propose a new technical means. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a fuel cell vehicle eco-driving collaborative control method based on the SAC algorithm, which can determine the power distribution state of the fuel cell vehicle based on the actual state information of the target fuel cell vehicle and the following vehicle information, thereby deciding the acceleration of the entire vehicle and the output power of the power system, thereby effectively realizing the collaborative control of eco-driving and vehicle following.

[0007] The present invention provides a fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm, comprising the following steps:

[0008] Acquiring target vehicle status information, and determining control requirement information of the target vehicle based on the vehicle status information;

[0009] A SAC-based neural network is constructed in the controller, including the state space and action space of eco-driving, and the reward function is determined based on the control demand information;

[0010] The state information of the target vehicle is input into the SAC-based neural network for offline training, and the real-time state information of the target vehicle is obtained and input into the trained SAC neural network to determine the optimal control strategy.

[0011] Furthermore, the target vehicle status information includes the following distance between the target vehicle and the vehicle preceding the target vehicle, the speed of the target vehicle, the speed of the vehicle preceding the target vehicle, the SOC of the fuel cell stack of the target vehicle, and the output power of the fuel cell stack of the target vehicle.

[0012] Furthermore, the control requirement information includes a following distance model of the target vehicle relative to the preceding vehicle, a drive motor model of the target vehicle, a power battery model, a fuel cell stack electrochemical model, and a fuel cell stack life model.

[0013] Furthermore, the following distance model of the target vehicle relative to the preceding vehicle is:

[0014] D(t+1)=D(t)+v rel (t)Δt;

[0015] v rel (t) = v p (t)-v h (t);

[0016]

[0017] Where: D(t+1) represents the following distance between the target vehicle and the preceding vehicle at time t+1, D(t) represents the following distance between the target vehicle and the preceding vehicle at time t, d1 represents the minimum following distance after the target vehicle brakes, d2 represents the maximum following distance after the target vehicle brakes, τ1 and τ2 represent the response time of the braking system, a1 and a2 represent the braking acceleration of the target vehicle, and v h represents the speed of the target vehicle, v rel (t) represents the relative speed between the target vehicle and the preceding vehicle, and They represent the upper and lower limits of the target vehicle’s expected following distance area, v p (t) is the speed of the vehicle in front of the target vehicle, D max and D min represents the maximum following distance and the minimum following distance, and α represents the adjustment coefficient.

[0018] Furthermore, the drive motor model of the target vehicle is:

[0019]

[0020] Where: P reqis the target vehicle's required power, m, g, and f are the vehicle's mass, gravitational acceleration, and rolling resistance coefficient, respectively. C d is the air resistance coefficient, A and δ are the conversion coefficients of frontal area and rotating mass respectively, μ is the road slope; P mot is the motor power requirement, P fc_net and P bat are the fuel cell stack net power and battery output power, η mot is the motor efficiency.

[0021] Furthermore, the power battery model is:

[0022]

[0023] Where: V bat Indicates the voltage of the power battery, V oc , I bat and R bat Represent the open circuit voltage, battery current and internal resistance of the power battery respectively, P bat is the battery power, SOC in Represents the initial value of the power battery SOC, Q bat Represents the nominal capacity of the power battery.

[0024] Furthermore, the electrochemical model of the fuel cell stack is:

[0025]

[0026] Where: T fc , R and F represent the temperature, gas constant and Faraday constant of the fuel cell stack respectively, P H2 and P O2 are the anode hydrogen partial pressure and cathode oxygen partial pressure of the fuel cell stack respectively; v0 represents the voltage at zero current density, i is the current density, i=I / A cell , I and A cell Represent the fuel cell stack current and the effective area of ​​the battery, v a and c2 are constants, c1 and c3 are constants, R ohm is the internal resistance.

[0027] Furthermore, the fuel cell stack life model is:

[0028] ΔD fc =k p [d low (t)+d high (t)+d load-change (t)+d start-stop (t)];

[0029]

[0030] Where: ΔD fc is the performance degradation rate of the fuel cell in a single time step, P low and P high represents the low power threshold and high power threshold of the fuel cell, |ΔP| represents the absolute value of the difference between the output power of the fuel cell at the previous time step and the output power at the current moment; k p is the road correction coefficient, k i (i=1,2,3,4) represents the decay coefficient of each type of working condition, and k1 is 1.26×10 -3 , k2 is 1.47×10 -3 , k3 takes 9.88×10 -7 , k4 is 1.96×10 -3 .

[0031] Furthermore, building a neural network based on SAC specifically includes:

[0032] The SAC-based neural network is a pruned double Q-learning network, which includes four parameterized soft Q-functions Q θ (s t ,a t ) and a policy function π φ (a t |s t );

[0033] Construct state space s t :s t =[D(t),v h (t),v rel (t),SOC(t),P fc (t)];

[0034] Constructing the action space:

[0035] And the acceleration and following distance constraints are:

[0036]

[0037] Constructing the reward function:

[0038]

[0039] in: is the weight coefficient, which represents the vehicle following cost C ACC and energy management cost C EMS The relative importance between ACC Based on the following distance cost C safe and jerk cost C jerk and the collision penalty term C collisionComposition, δ1, δ2 are C jerk and C collision The weight coefficient, C price C represents the price cost equivalent to the reduction of fuel cell stack SOH and hydrogen consumption. fc,tem Indicates the fuel cell stack overtemperature penalty, C SOC represents the power battery SOC maintenance cost, λ1 and λ2 are C fc,tem and C SOC The weight coefficient, p H2 and p fc are the unit prices of hydrogen and fuel cell stack, T fc,ref and SOC ref Indicates the fuel cell stack target temperature reference value and the lithium battery SOC maintenance reference value.

[0040] Furthermore, the state information of the target vehicle is input into the SAC-based neural network for offline training, specifically including:

[0041] Initialize the algorithm’s hyperparameters, the soft Q network parameters θ1, θ2 and θ′1, θ′2, and the policy network parameters φ;

[0042] Obtain the target vehicle's status information and input it into the SAC-based neural network;

[0043] The SAC-based neural network is based on the state s t Give the corresponding action a t , the power system performs corresponding power output;

[0044] Return the next moment state s t+1 and reward r t , the obtained tuple (s t ,a t ,r t ,s t+1 ) into the experience replay pool;

[0045] The optimal strategy function J(π * ) can be expressed by the following formula:

[0046]

[0047] Where: T represents the total duration of data interaction between the controller and the environment; E represents the mathematical expectation, s t and a t Represent the state and action at time t, ρ π is the optimal strategy π * The trajectory distribution under γ represents the discount factor of the reward, r(s t ,a t ) refers to the reward obtained by the state-action pair, πφ represents the policy function with parameter φ, -logπ φ (a t |s t ) indicates action a t Entropy at the current state; α is the temperature coefficient;

[0048] The update objective of the soft Q-function is to minimize the soft Bellman residual:

[0049]

[0050] Among them (s t ,a t ,r t ,s t+1 ) is a tuple sampled from the experience replay pool D, θ′ represents the network parameters of the target soft Q-function; in order to avoid overestimation of the Q value, the loss function for evaluating the soft Q network is expressed as follows:

[0051]

[0052] The evaluation network and the target network are composed of two groups of networks, whose parameters can be expressed as θ1, θ2 and θ′1, θ′2 respectively. During the training process, the target network with a lower Q value is selected to update the evaluation network. At the same time, the parameters of the target soft-Q network are replaced by soft updates:

[0053] θ′ i | i=1,2 =δθ i +(1-δ)θ′ i

[0054] Among them, δ is the soft update factor;

[0055] The loss function of the policy network is expressed as follows:

[0056]

[0057] Beneficial effects of the present invention: Through the present invention, the power distribution state of the fuel cell vehicle can be determined based on the actual state information and following vehicle information of the target fuel cell vehicle, thereby deciding the acceleration of the entire vehicle and the output power of the power system, thereby effectively realizing the coordinated control of ecological driving and vehicle following, so that the vehicle has good following performance and comfort during driving, and reduces the SOH degradation of the fuel cell stack, so that the fuel cell vehicle has good adaptability to working conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The present invention will be further described below with reference to the accompanying drawings and embodiments:

[0059] Figure 1Flowchart of the present invention. DETAILED DESCRIPTION

[0060] The present invention is further described in detail below:

[0061] The present invention provides a fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm, comprising the following steps:

[0062] Acquiring target vehicle status information, and determining control requirement information of the target vehicle based on the vehicle status information;

[0063] A SAC-based neural network is constructed in the controller, encompassing the state space and action space for eco-driving, and a reward function is determined based on control demand information. The controller is the target vehicle's vehicle controller, responsible for controlling the target vehicle and acquiring response information from the vehicle's sensors. It also interacts with the preceding vehicle to obtain its speed information.

[0064] The target vehicle's state information is input into a SAC-based neural network for offline training. The target vehicle's real-time state information is then acquired and input into the trained SAC neural network to determine the optimal control strategy. This invention determines the fuel cell vehicle's power distribution based on the target vehicle's actual state information and vehicle-following information, thereby determining the vehicle's acceleration and power system output, effectively achieving coordinated control of eco-driving and vehicle-following.

[0065] In this embodiment, the target vehicle status information includes the following distance between the target vehicle and the vehicle in front of the target vehicle, the speed of the target vehicle, the speed of the vehicle in front of the target vehicle, the SOC of the fuel cell stack of the target vehicle, and the output power of the fuel cell stack of the target vehicle.

[0066] In this embodiment, the control requirement information includes a following distance model of the target vehicle relative to the preceding vehicle, a drive motor model of the target vehicle, a power battery model, a fuel cell stack electrochemical model, and a fuel cell stack life model.

[0067] Specifically: The following distance model of the target vehicle relative to the preceding vehicle is:

[0068] D(t+1)=D(t)+v rel (t)Δt;

[0069] v rel (t) = v p (t)-v h (t);

[0070]

[0071] Where: D(t+1) represents the following distance between the target vehicle and the preceding vehicle at time t+1, D(t) represents the following distance between the target vehicle and the preceding vehicle at time t, d1 represents the minimum following distance after the target vehicle brakes, d2 represents the maximum following distance after the target vehicle brakes, τ1 and τ2 represent the response time of the braking system, a1 and a2 represent the braking acceleration of the target vehicle, and v h represents the speed of the target vehicle, v rel (t) represents the relative speed between the target vehicle and the preceding vehicle, and They represent the upper and lower limits of the target vehicle’s expected following distance area, v p (t) is the speed of the vehicle in front of the target vehicle, D max and D min represents the maximum following distance and the minimum following distance, and α represents the adjustment coefficient.

[0072] The drive motor model of the target vehicle is:

[0073]

[0074] Where: P req is the target vehicle's required power, m, g, and f are the vehicle's mass, gravitational acceleration, and rolling resistance coefficient, respectively. C d is the air resistance coefficient, A and δ are the conversion coefficients of frontal area and rotating mass respectively, μ is the road slope; P mot is the motor power requirement, P fc_net and P bat are the fuel cell stack net power and battery output power, η mot is the motor efficiency.

[0075] The power battery model is:

[0076]

[0077] Where: V bat Indicates the voltage of the power battery, V oc , I bat and R bat Represent the open circuit voltage, battery current and internal resistance of the power battery respectively, P bat is the battery power, SOC in Represents the initial value of the power battery SOC, Q bat Represents the nominal capacity of the power battery.

[0078] The electrochemical model of the fuel cell stack is:

[0079]

[0080] Where: Tfc , R and F represent the temperature, gas constant and Faraday constant of the fuel cell stack respectively, P H2 and P O2 are the anode hydrogen partial pressure and cathode oxygen partial pressure of the fuel cell stack respectively; v0 represents the voltage at zero current density, i is the current density, i=I / A cell , I and A cell Represent the fuel cell stack current and the effective area of ​​the battery, v a and c2 constants, which can be determined by the existing empirical formula of fuel cell stack temperature and oxygen partial pressure, c1 and c3 are constants, R ohm is the internal resistance, i max Indicates the current density that causes a sharp drop in voltage.

[0081] The fuel cell stack life model is:

[0082] ΔD fc =k p [d low (t)+d high (t)+d load-change (t)+d start-stop (t)];

[0083]

[0084] Where: ΔD fc is the performance degradation rate of the fuel cell in a single time step, P low and P high represents the low power threshold and high power threshold of the fuel cell, |ΔP| represents the absolute value of the difference between the output power of the fuel cell at the previous time step and the output power at the current moment; k p is the road correction coefficient, k i (i=1,2,3,4) represents the decay coefficient of each type of working condition, k1 is 1.26×10 -3 , k2 is 1.47×10 -3 , k3 takes 9.88×10 -7 , k4 is 1.96×10 -3 .

[0085] In this embodiment, constructing a neural network based on SAC specifically includes:

[0086] SAC neural network is the abbreviation of Soft Actor-Critic in English, and its Chinese name is Soft Actor-Critic Neural Network. The neural network based on SAC is a pruned double Q-learning network, which includes four parameterized soft Q-functions Q θ (s t ,a t ) and a policy function πφ (a t |s t ); and the control quantity of the entire neural network is determined by the acceleration a of the target vehicle h and the output power P of the fuel cell stack fc composition,

[0087] Construct state space s t :s t =[D(t),v h (t),v rel (t),SOC(t),P fc (t)];

[0088] Constructing the action space:

[0089] a h,bound Indicates the defined acceleration limit value, which is 3m / s 2 ; and the acceleration and following distance constraints are:

[0090]

[0091] In the early stages of training, the second acceleration constraint prevents most collisions. However, due to the discrete time step and the agent's exploration, when the leading vehicle's speed is low, the lead vehicle's rapid acceleration can lead to a collision. To ensure proper training, after a collision, the lead vehicle's following distance is set to the minimum following distance, as per the second following distance constraint. The lead vehicle's speed is also set to the leading vehicle's speed, but a penalty is applied for collisions.

[0092] Constructing the reward function:

[0093]

[0094] in: is the weight coefficient, which represents the vehicle following cost C ACC and energy management cost C EMS The relative importance between ACC Based on the following distance cost C safe and jerk cost C jerk and the collision penalty term C collision Composition, δ1, δ2 are C jerk and C collision The weight coefficient, C price C represents the price cost equivalent to the reduction of fuel cell stack SOH and hydrogen consumption. fc,tem Indicates the fuel cell stack overtemperature penalty, C SOC represents the power battery SOC maintenance cost, λ1 and λ2 are C fc,tem and C SOCThe weight coefficient, p H2 and p fc are the unit prices of hydrogen and fuel cell stack, T fc,ref and SOC ref represents the target temperature reference value for the fuel cell stack and the reference value for maintaining the lithium battery SOC. Based on the aforementioned reward function, the vehicle following system meets the following safety, traffic efficiency, and driving comfort requirements. Safety and traffic efficiency are represented by the following distance, while driving comfort is represented by the rate of change of acceleration. The primary energy management goal is to minimize hydrogen consumption, while the reduction of the fuel cell's SOH is equivalent to a price and included in the total cost. To ensure that the fuel cell stack operates at an appropriate temperature, an overtemperature penalty is set. Finally, the lithium battery SOC should be set to a reference value and maintained within a certain range.

[0095] In this embodiment, inputting the target vehicle's state information into the SAC-based neural network for offline training specifically includes:

[0096] Initialize the algorithm’s hyperparameters, the soft Q network parameters θ1, θ2 and θ′1, θ′2, and the policy network parameters φ;

[0097] Obtain the target vehicle's status information and input it into the SAC-based neural network;

[0098] The SAC-based neural network is based on the state s t Give the corresponding action a t , the power system performs corresponding power output;

[0099] Return the next moment state s t+1 and reward r t , the obtained tuple (s t ,a t ,r t ,s t+1 ) into the experience replay pool;

[0100] The optimal strategy function J(π * ) can be expressed by the following formula:

[0101]

[0102] Where: T represents the total duration of data interaction between the controller and the environment, that is, the duration of data interaction between the controller and the preceding vehicle, because the speed state of the preceding vehicle is to be obtained; E represents the mathematical expectation, s t and a t Represent the state and action at time t, ρ π is the optimal strategy π * The trajectory distribution under γ represents the discount factor of the reward, r(s t ,at ) refers to the reward obtained by the state-action pair, π φ represents the policy function with parameter φ, -logπ φ (a t |s t ) indicates action a t Entropy at the current state; α is the temperature coefficient;

[0103] Using the standard actor-critic framework and the maximum entropy principle, the SAC algorithm constructs a parameterized soft Q-function Q θ (s t ,a t ) and the policy function π φ (a t |s t ), the value function can be modeled as an expressive neural network to evaluate the quality of the output action, and the policy function is expressed by a Gaussian function whose mean and variance are the output of the neural network.

[0104] The update objective of the soft Q-function is to minimize the soft Bellman residual:

[0105]

[0106] Among them (s t ,a t ,r t ,s t+1 ) is a tuple sampled from the experience replay pool D, and θ′ represents the network parameters of the target soft Q-function. To avoid overestimation of the Q value, the pruned double Q-learning technique is introduced into the SAC algorithm. The loss function for evaluating the soft Q network is expressed as follows:

[0107]

[0108] The evaluation network and the target network are composed of two groups of networks, whose parameters can be expressed as θ1, θ2 and θ′1, θ′2 respectively. During the training process, the target network with a lower Q value is selected to update the evaluation network. At the same time, the parameters of the target soft-Q network are replaced by soft updates:

[0109] θ′ i | i=1,2 =δθ i +(1-δ)θ′ i

[0110] Among them, δ is the soft update factor;

[0111] The loss function of the policy network is expressed as follows:

[0112]

[0113] After the SAC God General network training is completed, the current speed of the preceding vehicle and the target vehicle's own state are input into the trained neural network to predict the target vehicle's action under the optimal strategy at the next moment. Then, the acceleration of the target vehicle and the output power of the power system are determined based on the action (this determination process can be carried out using existing technology), thereby achieving coordinated control of the target vehicle's following and ecological driving.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm, characterized by: The following steps are involved: Acquiring target vehicle status information, and determining control requirement information of the target vehicle based on the vehicle status information; A SAC-based neural network is constructed in the controller, including the state space and action space of eco-driving, and the reward function is determined based on the control demand information; The target vehicle's status information is input into the SAC-based neural network for offline training, and the real-time status information of the target vehicle is obtained and input into the trained SAC neural network to determine the optimal control strategy; The target vehicle state information includes the following distance between the target vehicle and the vehicle preceding the target vehicle, the speed of the target vehicle, the speed of the vehicle preceding the target vehicle, the SOC of the fuel cell stack of the target vehicle, and the output power of the fuel cell stack of the target vehicle; The control requirement information includes a following distance model of the target vehicle relative to the preceding vehicle, a drive motor model of the target vehicle, a power battery model, a fuel cell stack electrochemical model, and a fuel cell stack life model; The following distance model of the target vehicle relative to the preceding vehicle is: D(t+1)=D(t)+v rel (t)Δt; v rel (t)=v p (t)-v h (t); Where: D(t+1) represents the following distance between the target vehicle and the preceding vehicle at time t+1, D(t) represents the following distance between the target vehicle and the preceding vehicle at time t, d1 represents the minimum following distance after the target vehicle brakes, d2 represents the maximum following distance after the target vehicle brakes, τ1 and τ2 represent the response time of the braking system, a1 and a2 represent the braking acceleration of the target vehicle, and v h represents the speed of the target vehicle, v rel (t) represents the relative speed between the target vehicle and the preceding vehicle, and They represent the upper and lower limits of the target vehicle’s expected following distance area, v p (t) is the speed of the vehicle in front of the target vehicle, D max and D min represents the maximum following distance and the minimum following distance, and α represents the adjustment coefficient.

2. The fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm according to claim 1 is characterized by: The drive motor model of the target vehicle is: Where: P req is the required power of the target vehicle, m, g and f are the vehicle mass, gravity acceleration and rolling resistance coefficient respectively, C d is the air resistance coefficient, A and δ are the conversion coefficients of frontal area and rotating mass respectively, μ is the road slope; P mot is the motor power requirement, P fc_net and P bat are the fuel cell stack net power and battery output power, η mot is the motor efficiency.

3. The fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm according to claim 1 is characterized by: The power battery model is: Where: V bat Indicates the voltage of the power battery, V oc , I bat and R bat Represent the open circuit voltage, battery current and internal resistance of the power battery respectively, P bat is the battery power, SOC in Represents the initial value of the power battery SOC, Q bat Represents the nominal capacity of the power battery.

4. The fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm according to claim 1 is characterized by: The electrochemical model of the fuel cell stack is: Where: T fc , R and F represent the temperature, gas constant and Faraday constant of the fuel cell stack respectively, P H2 and P O2 are the anode hydrogen partial pressure and cathode oxygen partial pressure of the fuel cell stack respectively; v0 represents the voltage at zero current density, i is the current density, i=I / A cell , I and A cell Represent the fuel cell stack current and the effective area of ​​the battery, v a and c2 are constants, c1 and c3 are constants, R ohm is the internal resistance.

5. The fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm according to claim 1 is characterized by: The fuel cell stack life model is: ΔD fc =k p [d low (t)+d high (t)+d load-change (t)+d start-stop (t)]; Where: ΔD fc is the performance degradation rate of the fuel cell in a single time step, P low and P high represents the low power threshold and high power threshold of the fuel cell, |ΔP| represents the absolute value of the difference between the output power of the fuel cell at the previous time step and the output power at the current moment; k p is the road correction coefficient, k i (i=1,2,3,4) represents the decay coefficient of each type of working condition, and k1 is 1.26×10 -3 , k2 is 1.47×10 -3 , k3 takes 9.88×10 -7 , k4 is 1.96×10 -3 .

6. The fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm according to claim 1 is characterized by: Building a neural network based on SAC specifically includes: The SAC-based neural network is a pruned double Q-learning network, which includes four parameterized soft Q-functions Q θ (s t ,a t ) and a policy function π φ (a t |s t ); Construct state space s t :s t =[D(t),v h (t),v rel (t),SOC(t),P fc (t)]; Constructing the action space: And the acceleration and following distance constraints are: Constructing the reward function: in: is the weight coefficient, which represents the vehicle following cost C ACC and energy management cost C EMS The relative importance between ACC Based on the following distance cost C safe and jerk cost C jerk and the collision penalty term C collision Composition, δ1, δ2 are C jerk and C collision The weight coefficient, C price C represents the price cost equivalent to the reduction of fuel cell stack SOH and hydrogen consumption. fc,tem Indicates the fuel cell stack overtemperature penalty, C SOC represents the power battery SOC maintenance cost, λ1 and λ2 are C fc,tem and C SOC The weight coefficient, p H2 and p fc are the unit prices of hydrogen and fuel cell stack, T fc,ref and SOC ref Indicates the fuel cell stack target temperature reference value and the lithium battery SOC maintenance reference value.

7. The fuel cell vehicle eco-driving cooperative control method based on the SAC algorithm according to claim 6 is characterized by: Inputting the target vehicle's state information into the SAC-based neural network for offline training specifically includes: Initialize the algorithm’s hyperparameters, the soft Q network parameters θ1, θ2 and θ′1, θ′2, and the policy network parameters φ; Obtain the target vehicle's status information and input it into the SAC-based neural network; The SAC-based neural network is based on the state s t Give the corresponding action a t , the power system performs corresponding power output; Return the next moment state s t+1 and reward r t , the obtained tuple (s t ,a t ,r t ,s t+1 ) into the experience replay pool; The optimal strategy function J(π * ) can be expressed by the following formula: Where: T represents the total duration of data interaction between the controller and the environment; E represents the mathematical expectation, s t and a t Represent the state and action at time t, ρ π is the optimal strategy π * The trajectory distribution under γ represents the discount factor of the reward, r(s t ,a t ) refers to the reward obtained by the state-action pair, π φ represents the policy function with parameter φ, -logπ φ (a t |s t ) indicates action a t Entropy at the current state; α is the temperature coefficient; The update objective of the soft Q-function is to minimize the soft Bellman residual: Among them (s t ,a t ,r t ,s t+1 ) is a tuple sampled from the experience replay pool D, θ′ represents the network parameters of the target soft Q-function; in order to avoid overestimation of the Q value, the loss function for evaluating the soft Q network is expressed as follows: The evaluation network and the target network are composed of two groups of networks, whose parameters can be expressed as θ1, θ2 and θ′1, θ′2 respectively. During the training process, the target network with a lower Q value is selected to update the evaluation network. At the same time, the parameters of the target soft-Q network are replaced by soft updates: I will i | i=1,2 =sth i +(1-δ)θ' i Among them, δ is the soft update factor; The loss function of the policy network is expressed as follows:

Citation Information

Patent Citations

  • Bend road car following control method based on safe control domain

    CN109969183A

  • Multi-model reasoning acceleration system and method for automatic driving full-scene perception

    CN116306938A