An energy-saving driving method for a networked fuel cell bus with comprehensive comfort enhancement
By integrating EMS and ACC and combining deep reinforcement learning algorithms to optimize the energy-saving driving strategy of fuel cell hybrid buses, the impact of road surface smoothness on ride comfort is resolved, achieving the effect of improving ride comfort while ensuring energy saving and safety.
Patent Information
- Application Number
- CN202411026157.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-07-30
AI Technical Summary
Existing fuel cell hybrid bus energy-saving driving strategies fail to fully consider the impact of road surface smoothness on ride comfort, resulting in passengers feeling uncomfortable on bumpy roads. Traditional energy-saving strategies mainly focus on energy consumption and power system status, ignoring vertical comfort.
It adopts deep reinforcement learning algorithms, integrates energy management system (EMS) and adaptive cruise control (ACC), optimizes energy-saving driving strategies through deep reinforcement learning algorithms, combines road surface smoothness and vertical comfort, establishes a quantitative relationship model, uses V2X technology to obtain real-time road condition information, optimizes vehicle motion status, and enhances comprehensive comfort.
The ride comfort of fuel cell hybrid buses was improved while ensuring energy efficiency and safety. The improved deep reinforcement learning algorithm improved the training performance and generalization of the strategy, thereby enhancing the overall comfort.
Smart Images

Figure CN118963124B_ABST
Abstract
Description
Technical Field
[0001] This paper proposes an energy-saving driving strategy for a comfort-enhanced fuel cell hybrid electric bus. This strategy achieves energy-saving driving for fuel cell hybrid vehicles by integrating energy management system (EMS) and adaptive cruise control (ACC) optimization. Furthermore, this method utilizes deep reinforcement learning techniques, such as the Differential Flexible Actor-3-Critic (SSA3C) algorithm, to optimize the energy-saving driving strategy. While meeting multiple performance criteria such as safety, durability, and efficiency, this strategy leverages enhanced road surface information to optimize energy consumption. Background Art
[0002] Global challenges such as carbon emissions and the energy crisis are driving governments and automakers towards electrified and environmentally friendly transportation. Fuel cell hybrid buses (FCHEBs) have attracted significant attention due to their high energy density, fast refueling, and zero CO2 emissions. However, while pursuing efficient energy utilization and environmental protection, the impact of road bumps on bus comfort has become increasingly prominent. To meet high power demands, especially for heavy vehicles, FCHEBs utilize lithium-ion battery packs to support the fuel cell system (FCS), addressing issues such as low power-to-weight ratio and long charging times. However, under complex road conditions, especially those with poor road surface roughness, the vertical comfort of the vehicle, or the passenger experience, can be severely impacted. Advances in intelligent control technology have enabled the effective control and optimization of complex FCHEB powertrains. However, current energy-saving driving strategies mostly aim to minimize energy consumption and improve driving economy. These strategies primarily focus on two aspects: optimizing fuel consumption and powertrain status, and controlling driving behavior in traffic, while neglecting road surface information and the crucial factor of vertical comfort. This means that when the road surface is bumpy, the vehicle's vibration damping and cushioning systems may not function fully, leading to increased passenger discomfort. The former can be modeled as an energy management system (EMS) task, while the latter encompasses various vehicle motion control tasks depending on the traffic scenario. Among these, the car-following scenario, closely related to the adaptive cruise control (ACC) system, is the most energy-efficient and popular. To enhance the ride comfort of FCHEBs, energy-saving driving strategies must fully consider the impact of road surface information and vertical comfort. Leveraging V2X (vehicle-to-everything) technology to connect vehicles, roadside units, and cloud platforms enhances multiple driving capabilities and facilitates real-time access to road condition information, enabling early prediction and adjustment of the vehicle's motion state. Therefore, integrating EMS and ACC optimization into research on energy-saving driving in connected FCHEBs is of great significance. Summary of the Invention
[0003] Purpose of the Invention: To address the aforementioned technical issues, the present invention provides an energy-saving driving method for a connected fuel cell bus with enhanced comprehensive comfort. This method integrates energy management system (EMS) and adaptive cruise control (ACC) optimization to achieve energy-saving driving for fuel cell hybrid buses. By considering road surface smoothness and vertical comfort, and incorporating these into the energy-saving driving strategy along with other indicators such as longitudinal comfort, energy efficiency, safety, and durability, a quantitative relationship model between road surface smoothness, vertical comfort, and vehicle speed is established. This method actively controls and improves the overall driving comfort of the vehicle, independent of the vehicle's suspension type and parameters.
[0004] Technical solution: The present invention adopts the following technical solution: a comprehensive comfort-enhanced networked fuel cell bus energy-saving driving method, comprising the following steps:
[0005] Using a connected fuel cell hybrid bus (FCHEB) as the research object, a collaborative optimization goal for energy-saving driving was established. While considering safety, energy efficiency, and durability, comprehensive driving comfort was also improved. This comfort included vertical comfort, which used weighted root mean square acceleration (WRMSA) as an indicator for assessing human exposure to whole-body vibration. The relationship between WRMSA, the International Roughness Index (IRI), and driving speed was expressed using a linear function.
[0006] The collaborative optimization objectives are quantified into reward coefficients. The state space, action space, and reward function are set, and a comprehensive comfort-enhanced energy-saving driving strategy is derived based on a deep reinforcement learning algorithm. The state variables in the state space include the state of charge of the power battery, the health status of the power battery, fuel cell, and motor, the speed and acceleration of the main vehicle, the speed and acceleration of the preceding vehicle, the distance between the two pairs of vehicles, and the IRI. The action space includes the speed and power distribution of the main vehicle. The reward function integrates the reward coefficients of each objective.
[0007] The energy-saving driving strategy is trained offline to obtain an inheritable parameterized neural network strategy, which is then downloaded to the vehicle controller to achieve real-time online application of the energy-saving driving strategy.
[0008] Furthermore, the safety bonus coefficient is expressed as:
[0009]
[0010] Among them D h,l (t), D safe (t) and D max (t) represents the relative distance, safe distance and maximum distance between the two vehicles at time step t respectively; Indicates the maximum speed, v h Indicates the speed of the host vehicle;
[0011] The comprehensive comfort bonus coefficient is expressed as:
[0012]
[0013] Where, jerk(t) is the rate of change of acceleration, a r is the acceleration a h The instantaneous difference, WRMSAa w The instantaneous difference
[0014] Energy-saving bonus coefficient includes R m (t) and R soc (t), expressed as:
[0015]
[0016] in, is the hydrogen consumption rate, SOC(t) is the battery state of charge at time step t, SOC tar is the target value of SOC;
[0017] Durability bonus coefficients include R bat (t), R FCS (t) and R mot (t), expressed as:
[0018]
[0019] Among them, ΔSOH Ba (t), D FCS (t) and CLR(t) are the battery health, FCS degradation rate and cumulative loss rate of the motor at time step t, respectively.
[0020] Furthermore, the state defined in the state space is expressed as:
[0021] s=[SOC,SOH s ,a h ,v h ,a l ,v l ,D h,l ,IRI]
[0022] Among them, SOC represents the state of charge of the power battery, SOH s =(SOH bat ,SOH FCS ,SOH mot ) represents the health status of the power battery, fuel cell and motor, a h ,v h represents the acceleration and speed of the host vehicle, a l ,vl Indicates the acceleration and speed of the preceding vehicle;
[0023] The action is represented as:
[0024] a=[a h ,P FCS ]
[0025] Among them, a h ∈[-1.5,1]m / s 2 , P FCS Represents the output power of the fuel cell, P FCS ∈[0,60]kW.
[0026] Furthermore, the collaborative optimization goal of energy-saving driving also considers driving efficiency. Driving efficiency refers to maintaining a short headway time TH while ensuring driving safety. The probability density function of the lognormal distribution of TH within the safety range is:
[0027]
[0028] Where x(t) represents the TH at time step t, μ and σ represent the mean and standard deviation of TH.
[0029] Furthermore, the total reward function is expressed as:
[0030]
[0031] Among them, s, a represent the state and action, R m (t), R soc (t), R FCS (t), R bat (t), R not (t), R comfort (t), R safe (t), R eff (t) represents the fuel energy saving bonus coefficient, electricity energy saving bonus coefficient, FCS durability bonus coefficient, power battery pack durability bonus coefficient, drive motor durability bonus coefficient, comfort bonus coefficient, safety bonus coefficient and efficiency bonus coefficient respectively; β1, β2, β3, β4, β5, ω and σ are weight coefficients.
[0032] Furthermore, WRMSAa w The relationship between the International Roughness Index IRI and driving speed v is expressed as:
[0033] a w =k1IRI+k2v+k3·IRI·v+b
[0034] Among them, k1, k2, k3, and b are coefficients. The measured IRI at different speeds is used to evaluate WRMSA, and the values of each coefficient are determined by multiple linear regression analysis.
[0035] Furthermore, each state variable of the networked fuel cell hybrid bus is obtained through sensing devices and networked devices, or each state variable is obtained from a simulation environment; the simulation environment includes a fuel cell vehicle model, a car-following scenario model, and an electrified component health model;
[0036] The fuel cell vehicle model uses a fuel cell system FCS and a group of lithium-ion power batteries as the main energy sources; in the vehicle dynamics model, the required driving power P req and driving force F req , the corresponding resistance includes the inertial force F a , rolling resistance F r , road slope resistance F i and aerodynamic drag F w , described as follows:
[0037]
[0038] Where v represents the bus speed. The FCHEB drive motor model is based on a quasi-steady-state method. The quasi-steady-state model represents the motor as:
[0039]
[0040] Among them, P mot , T mot , n mot and η mot (T mot ,n mot ) represent the motor’s power, torque, speed and efficiency at a certain torque and speed respectively;
[0041] FCS uses an equivalent circuit model consisting of an ideal voltage source and three series resistors; the H2 consumption rate can then be calculated:
[0042]
[0043] Among them, P FCS represents the power of FCS, η FCS Indicates the efficiency of FCS, L v Indicates the lower heating value of hydrogen;
[0044] Use equivalent circuit to simulate the power battery:
[0045]
[0046] Among them, VOC is the open circuit voltage, I is the load current, R bat is the internal resistance of the battery, SOC is the state of charge, SOC0 is the initial value of the state of charge, C n is the capacity of the battery; Indicates the total amount of electricity consumed or charged by the battery during this period;
[0047] Power battery output power P bat and FCSP FCS Maintain the following relationship:
[0048] P req =T mot ·n mot =P bat +P FCS ·η DC / DC
[0049] Among them, η DC / DC Represents the efficiency of the DC / DC converter;
[0050] The following scene model is composed of the main vehicle speed v h (t) and acceleration a h (t), speed of the preceding vehicle v l (t) and acceleration a l (t), mileage of the preceding vehicle L l (t), the main vehicle's mileage L h (t) and the relative distance D between the two vehicles h,l (t) Description: The relationship between speed and distance is as follows:
[0051]
[0052] Among them, L veh is the length of the vehicle;
[0053] In the following process, define the maximum relative distance D between the two vehicles. max and safety distance D safe ;
[0054] The electric component health model consists of three parts: power battery, FCS and motor;
[0055] The degradation D of FCS at time step t FCS (t) and state of health SOH FCS (t) is calculated by the following formula:
[0056]
[0057] Where n is the total time step, d ss (i),d low(i),d high (i),d cha (i) start-stop, low power load, high power load and load change at time i;
[0058] Degradation ΔSOH(t) of power batteries and healthy state SOH bat (t) is shown in the following formula:
[0059]
[0060] Where I(t) is the load current at time step t, Δt is the current duration, and N(c,T a ) is the equivalent number of cycles when the power battery reaches the end of its life, T a is the average temperature inside the battery, c is the discharge rate;
[0061] The motor's cumulative loss ratio (CLR) and motor health status (SOH mot ) is defined as follows:
[0062]
[0063] Among them, W losses is the total energy loss during the life of the motor; η mot , P mot , η rated and P rated They are the efficiency, power, rated efficiency and rated power of the motor respectively; t life is the manufacturer's recommended motor life.
[0064] Furthermore, the deep reinforcement learning algorithm adopts a flexible actor-critic (SAC) algorithm. In the algorithm framework, the intelligent agent interacts with the real environment or the simulated environment. The intelligent agent observes the current environmental state information, selects and executes actions according to the strategy, enters a new environmental state, and obtains rewards from environmental feedback. The sampling efficiency is improved through experience replay technology. At the same time, the state, action, and reward information are stored. This cycle is repeated to achieve training of energy-saving driving strategies.
[0065] Furthermore, the deep reinforcement learning algorithm adopts the differential flexible actor-3-critic (SSA3C) algorithm. The intelligent agent in the algorithm framework interacts with the simulation environment. The intelligent agent observes the current environmental state information, selects and executes actions according to the strategy, enters the new environmental state, and obtains rewards from the environmental feedback. The sampling efficiency is improved through experience replay technology. At the same time, the state, action, and reward information are stored. This cycle is repeated to achieve the training of energy-saving driving strategies. The training steps are as follows:
[0066] Step 1. Initialize the actor network and critic network for approximating the policy and value function: Initialize the target critic network, which is used to assist in estimating the Q value and reduce the estimation error; set the hyperparameters of the algorithm, set the step size parameter ∈ and kernel function h of the differential gradient descent SVGD; define a storage space D as the experience replay pool and initialize it;
[0067] Step 2: The Actor network transformed based on SVGD interacts with the established training environment: a set of particles {a 0}, update these particles through multiple iterative steps l∈[1,L] to adapt them to the target distribution:
[0068] {a l+1}←{a l}+∈h({a l},s t )
[0069] Among them, ∈ is used to control the size of the particle update in each iteration; h is used to measure the similarity between two particles; by deriving the closed-form expression of the particle distribution at the lth iteration and the particle part of the final step L, the strategy is generated:
[0070]
[0071] Where q represents the distribution; π φ (a t ∣s t ) is a policy π with φ as parameter given state s t Output action a t The probability distribution of ; det represents the determinant; I is an identity matrix; is the gradient used to adjust the particle position; then, the reward r is obtained by calculating the reward function and observing t and the next state s t+1 ;Experience(s t ,a t ,r t ,s t+1 )Save to the experience replay pool And update the state matrix s←s t+1 ;
[0072] Step 3: From the Experience Replay Pool Randomly sample subsets Get N(s t ,a t ,r t ,s t+1 ) for subsequent training;
[0073] Step 4. Calculate the Q-value function of the Critic network, which is used to evaluate the value of actions under different states. Use truncated triple Q learning to calculate the minimum two Q-values and take their average as the final Q-value:
[0074]
[0075] i∈{1,2,3},j∈{1,2}.
[0076] Step 5: Use Kullback-Leibler divergence D KL , for the policy function π(a t |s t ) to improve:
[0077]
[0078] Among them, Z(s t ) represents the logarithmic partition function, α is the temperature parameter; at the same time, the Q-value function is learned by minimizing the soft Bellman residual to update the parameters θ of the Critic network:
[0079]
[0080] Among them, (s t ,a t ,r t ,s t+1 ) represents from The mini-batch sampling in ; Q′ represents the target Q value function of the target Critic network; γ is the discount factor; then, the parameter θ′ of the target evaluation network Q′ is updated:
[0081] θ′←(1-τ)θ′+τθ
[0082] Where, τ is the soft update factor;
[0083] Step 6. The entropy H is estimated by taking the logarithm over a set of particles:
[0084]
[0085] Among them, Tr represents the trace of the matrix, a 0 is the initial particle swarm, and are the particle swarms after the lth iteration and the last iteration L respectively;
[0086] Step 7: Minimize the strategy induced by the sampler particles The expected KL divergence between the Q value is used to calculate the parameter φ * :
[0087]
[0088] Among them, t sd To limit the standard deviation of particle updates, σ φ is the mean of the initial distribution;
[0089] Step 8. Maximize the following objective function to make the entropy of the policy close to the target entropy value, and finally update the temperature coefficient α:
[0090]
[0091] in, is the target entropy, which is the opposite of the action dimension.
[0092] The present invention also provides a computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort.
[0093] Beneficial effects: Compared with the existing technology, the present invention has the following beneficial effects: The present invention proposes an energy-saving driving method for networked fuel cell vehicles based on a deep reinforcement learning algorithm, taking into account road smoothness and vertical comfort, and integrating it with longitudinal comfort, energy saving, safety, durability, traffic efficiency and other factors into the energy-saving driving strategy of networked fuel cell buses, thereby improving multiple performance while enhancing comprehensive comfort. In addition, the present invention further addresses the problem of insufficient action expression and Q-value estimation in traditional SAC algorithms, and uses differential gradient descent (SVGD) and truncated triple Q learning (3C) to improve the actor-critic network, thereby improving training performance and adaptability to multiple tasks. In addition, the present invention enhances the generalization of the evaluation strategy by constructing a combined driving cycle of a standard driving cycle and an IRI distribution. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 Schematic diagram of the overall framework of the energy-saving driving method according to an embodiment of the present invention.
[0095] Figure 2 It is a schematic structural diagram of a fuel cell hybrid vehicle in an embodiment of the present invention.
[0096] Figure 3 It is a schematic diagram of the connected car-following scenario model in an embodiment of the present invention.
[0097] Figure 4 Schematic diagram of the interaction among IRI, vehicle speed and WRMSA in an embodiment of the present invention.
[0098] Figure 5 Schematic diagram of a composite driving cycle in an embodiment of the present invention. DETAILED DESCRIPTION
[0099] In order to make the purpose, technical solution and advantages of the present invention more clear, the technical solution of the present application is further described in detail below with reference to the accompanying drawings. The described embodiment is only a part of the embodiments involved in this application. All non-innovative embodiments based on this embodiment by other researchers in this field fall within the scope of protection of this application.
[0100] The embodiment of the present invention discloses an energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort. First, a networked fuel cell hybrid bus FCHEB is taken as a research object, and a collaborative optimization goal of energy-saving driving is established. While considering safety, energy saving and durability, the comprehensive driving comfort is increased. The comfort includes vertical comfort. The vertical comfort uses weighted root mean square acceleration WRMSA as an indicator for evaluating human exposure to whole-body vibration. The relationship between WRMSA, international roughness index IRI and driving speed is expressed by a linear function. Then, the collaborative optimization goal is quantified into a reward coefficient, and the state space, The action space and reward function are used to obtain a comprehensive comfort-enhanced energy-saving driving strategy based on a deep reinforcement learning algorithm. The state variables in the state space include the state of charge of the power battery, the health status of the power battery, fuel cell and motor, the speed and acceleration of the main vehicle, the speed and acceleration of the preceding vehicle, the distance between the two pairs of vehicles and IRI. The action space includes the main vehicle speed and power distribution. The reward function integrates the reward coefficients of each target. The energy-saving driving strategy is then trained offline to obtain an inheritable parameterized neural network strategy. The parameterized neural network strategy obtained through offline training is downloaded to the vehicle controller to realize the real-time online application of the energy-saving driving strategy.
[0101] The following combination Figure 1 Taking a simulation scenario as an example, the specific steps of the energy-saving driving method according to an embodiment of the present invention are described in detail:
[0102] Step 1: Build a simulation environment for training, including a fuel cell vehicle model, a car-following scenario model, an electrified component health model, and a vertical comfort model. Use vehicle-to-cloud communication (V2C) to obtain road information, quantify its relationship with vertical comfort, and enhance the coordinated optimization of the energy management system (EMS) and adaptive cruise control (ACC) in energy-saving driving.
[0103] Step 2: Obtain a comprehensive comfort-enhanced energy-saving driving strategy based on a deep reinforcement learning algorithm, such as the Flexible Actor-Critic (SAC) algorithm. This embodiment further improves the traditional SAC algorithm and proposes a comprehensive comfort-enhanced energy-saving driving strategy based on the Differential Flexible Actor-3-Critic (SSA3C) algorithm. Actor networks, critic networks, and target critic networks are created to construct a training network for energy-saving driving. An optimization goal for energy-saving driving is established, and the optimization goal is quantified into a reward coefficient. The state space, action space, and reward function required for training are set.
[0104] Step 3: The intelligent agent interacts with the environment. The C-WTVC-IRI combined cycle constructed by the China-World Heavy Commercial Vehicle Transient Cycle (C-WTVC) and the Long-Term Pavement Performance (LTPP) program is used as training data. The energy-saving driving strategy is trained offline through the SSA3C algorithm to obtain an inheritable parameterized neural network strategy.
[0105] Step 4: Download the parameterized neural network strategy obtained through offline training to the vehicle controller of the fuel cell hybrid bus to realize the real-time online application of the energy-saving driving strategy.
[0106] In a preferred embodiment of the present invention, the step 1 of constructing a simulation environment for training specifically includes the following steps:
[0107] 1. Fuel cell vehicle model, with a fuel cell system (FCS) and a group of lithium-ion power batteries as the main energy source, the vehicle structure is as follows Figure 2 As shown. Since the energy of both power components is converted into motion, in the vehicle dynamics model, the required driving power P req and driving force F req , the corresponding resistance includes the inertial force F a , rolling resistance F r , road slope resistance F i and aerodynamic drag F w , which can be described as follows:
[0108]
[0109] Where v represents the bus speed, m represents the bus mass, a represents the bus lateral acceleration, δ represents the rotational mass conversion coefficient, μ represents the rolling resistance coefficient, and A w C is the frontal area of the vehicle. d is the air resistance coefficient, θ is the road slope, g is the acceleration due to gravity, 9.8m / s 2 .
[0110] Secondly, the FCHEB drive motor model is based on a quasi-steady-state method. The quasi-steady-state model represents the motor as:
[0111]
[0112] Among them, P mot , T mot , n mot and η mot (T mot ,n mot ) represent the motor's power, torque, speed and efficiency at a certain torque and speed respectively.
[0113] Secondly, FCS uses an equivalent circuit model consisting of an ideal voltage source and three series resistors. Then, the H2 consumption rate can be calculated:
[0114]
[0115] Among them, P FCS represents the power of FCS, η FCS Indicates the efficiency of FCS, L v Indicates the lower calorific value of hydrogen (120kJ / h).
[0116] Then, the power battery is simulated using an equivalent circuit:
[0117]
[0118] Where V OC is the open circuit voltage, I is the load current, R bat is the internal resistance of the battery, SOC is the state of charge, and SOC0 is the initial value of the state of charge, C n is the capacity of the battery, Indicates the total amount of power consumed or charged by the battery during this period.
[0119] In addition, due to the application of DC / DC converter, its current matches the DC bus and the power battery output power P bat and FCSP FC Keep as follows:
[0120] P req =T mot ·n mot =P bat +P FCS ·η DC / DC #(5)
[0121] Among them, η DC / DC Represents the efficiency of the DC / DC converter.
[0122] 2. Car-following scenario model simulation environment, such as Figure 3 As shown, the main vehicle speed v h (t) and acceleration a h (t), speed of the preceding vehicle v l (t) and acceleration a l (t), mileage of the preceding vehicle L l (t), the main vehicle's mileage L h (t) and the relative distance D between the two vehicles h,l (t) Description. The host vehicle can obtain the speed and acceleration of the preceding vehicle at each time step, as well as the distance between the two vehicles, uploaded to the cloud via V2C. It can also obtain the road surface roughness information of the preceding road section from the cloud. The speed and distance can be derived as follows:
[0123]
[0124] Among them, L veh is the length of the vehicle.
[0125] During the following process, while paying attention to the passengers' comfort, the vehicle in front and the main vehicle should maintain a safe and appropriate distance. Therefore, the maximum relative distance between the two vehicles is D max and safety distance D safe The safety distance is defined as:
[0126]
[0127] Among them, D safe Considered as the distance between two vehicles D h,l The minimum value of t d is the sum of the braking delay and reaction time, which is 1.5s; d0 is the safe distance between the main vehicle and the preceding vehicle after stopping, which is set to 3m; a max It is the maximum acceleration in an emergency situation, and its value is 6.68m / s 2 .
[0128] 3. The durability model of the electrified components consists of three parts: power battery, FCS and motor:
[0129] (1) Load change, start-stop, low power load and high power load are the four operating conditions that lead to FCS degradation. Correspondingly, the degradation of FCS at time step t is D FCS (t) and state of health SOH FCS (t) can be calculated by the following formula:
[0130]
[0131] Where n is the total time step. ss (i),d low (i),d hig (i),dcha (i) are start-stop, low power load, high power load and load change at time i.
[0132] (2) The degradation of power batteries is based on a coupled electrothermal aging model. The degradation and health state (SOH) of power batteries are shown in the following equation:
[0133]
[0134] Where I(t) is the load current at time step t, Δt is the current duration, and N(c,T a ) is the equivalent number of cycles when the power battery reaches the end of life (EOL), T a is the average temperature inside the battery, and c is the discharge rate.
[0135] (3) Regarding motor degradation, the state of health (SOH) is defined based on the energy loss accumulated over time. If the motor’s cumulative loss ratio (CLR) is exceeded, the motor’s state of health (SOH) mot The motor's cumulative loss ratio (CLR) and motor health status (SOH) mot ) is defined as follows:
[0136]
[0137] Among them, W losses is the total energy loss during the life of the motor; η mot , P mot , η rated and P rated They are the efficiency, power, rated efficiency and rated power of the motor respectively; t life is the manufacturer's recommended lifespan for the motor, in this case 30,000 hours.
[0138] 4. Vertical comfort model Figure 4 As shown in Figure 3, the quantitative correlation between vertical comfort index and road roughness and driving speed is used to simplify the vertical comfort evaluation.
[0139] For vertical comfort index, weighted root mean square acceleration (WRMSA) is used as an indicator to evaluate human exposure to whole body vibration. Since human sensation and acceleration differ significantly between 0.5 and 80 Hz, WRMSA can be calculated as:
[0140]
[0141] Among them, ω i ,u i , l iare the coefficient, upper limit frequency and lower limit frequency of the i-th one-third octave band respectively; Sa(f) is the power spectral density of vibration acceleration at frequency f.
[0142] Quantifying road roughness involves a large amount of raw road profile data, so the International Roughness Index (IRI) is used to represent road characteristics. The IRI is a macro-parameter related to vertical comfort, which is described as the cumulative vertical displacement of a specific quarter-size car model suspension at a speed of 80 km / h:
[0143]
[0144] Where l is the evaluation length of IRI, x is the driving distance of the vehicle; z s (x) is the vertical displacement of the sprung mass of the suspension when the vehicle travels a distance x, and z u is the vertical displacement of the unsprung mass of the suspension when the vehicle travels a distance x.
[0145] Regression models are often used to study the interaction between WRMSA, IRI, and driving speed. At different driving speeds, the relationship between WRMSA and IRI can be expressed by a linear function. Therefore, the linear function is designed as:
[0146] a w =k1IRI+k2v+k3·IRI·v+b#(13)
[0147] Here, k1, k2, k3, and b are coefficients. Driving behavior is determined by control commands and is unrelated to external factors and passenger comfort, so the independent variables are uncorrelated. WRMSA was evaluated using the measured IRI at different speeds. A normal distribution was chosen for multiple linear regression analysis, resulting in k1 set to 0.4428, k2 to 0.0157, k3 to 0.0004, and b to -0.901.
[0148] In a preferred embodiment of the present invention, the step 2 specifically includes the following steps:
[0149] First, if Figure 1 As shown in , an Actor network is constructed, denoted as π(s,a|φ), where φ is the network parameter, the input of the Actor network is the current state s, and the output is the action a; a Critic network is constructed, denoted as Q(s,a|θ), where θ is the network parameter, the input of the Critic network is the current state s and the action a output by the Actor network, and the output is the value function Q and gradient information; a target Critic network Q′(s,a|θ′) is constructed, and the network structure and parameter structure of the target network are the same as those of the Critic network, and θ′ is the parameter of the target Critic network.
[0150] Next, we establish the optimization goals for energy-saving driving: five performance optimization goals: (1) safety, (2) comfort, (3) energy saving, (4) durability, and (5) efficiency. Correspondingly, we quantify the energy-saving driving optimization goals into a reward coefficient R:
[0151] For the optimization goal (1), that is, to maintain driving safety, the relative distance D between the two vehicles h,l Used as a basis for judging the safety of following a vehicle:
[0152]
[0153] When D h,l (t)≤0 means that when the main vehicle collides with the vehicle in front, the main vehicle should be severely punished. At this time, the reward coefficient is the maximum speed When D h,l (t) is less than the safety distance D safe At (t), the reward coefficient is the speed v of the main vehicle. h , that is, the slower the speed, the smaller the risk. h,l (t) is greater than the maximum following distance D max (t), the reward coefficient is the difference between the two D h,l (t)-D max (t).
[0154] For optimization goal (2), i.e. maintaining driving comfort, longitudinal comfort (vehicle forward direction) also needs to be considered on the basis of vertical comfort. The rate of change of acceleration (jerk) is introduced as an indicator of longitudinal comfort:
[0155]
[0156] Where, jerk(t) is the rate of change of acceleration, a r is the acceleration a h The instantaneous difference, WRMSAa w The instantaneous difference.
[0157] For the optimization goal (3), i.e. reducing hydrogen consumption and keeping SOC within a reasonable range, the corresponding reward coefficient R m (t) and R soc (t) are as follows:
[0158]
[0159] in, is the hydrogen consumption rate in g / s, SOC(t) is the battery state of charge (SOC) at time step t, SOC tar is the target value of SOC;
[0160] Regarding the optimization goal (4), which is to reduce the degradation of multiple electrified components, including the power battery pack, FCS, and drive motor, the problem can be defined as follows:
[0161]
[0162] Among them, ΔSOH Ba (t), D FCS (t) and CLR(t) are the battery health, FCS degradation rate and cumulative loss rate of the motor at time step t, respectively.
[0163] For optimization objective (5), driving efficiency refers to maintaining a short headway (TH) while ensuring driving safety. A short TH within a safe range can improve traffic flow efficiency. Based on NGSIM data, the appropriate TH during vehicle following can be determined to be 1.3s. The probability density function of the log-normal distribution of TH is:
[0164]
[0165] Where x(t) represents the TH at time step t, μ and σ represent the mean and standard deviation of TH, which are 0.4226 and 0.4365 respectively.
[0166] Furthermore, in a preferred embodiment of the present invention, in step 2, the state space, action space, and reward function required for training are set based on the optimization goal:
[0167] The characteristics of the V2C following environment should be represented in the state space as: the main vehicle HV power system characteristics, the main vehicle HV and the front vehicle LV motion characteristics and the IRI distribution of the road ahead. Therefore, the state variables defined in the state space include SOC, SOH s (SOH bat ,SOH FCS ,SOH mot ),a h ,v h ,a l ,v l ,D h,l and IRI:
[0168] s=[SOC,SOH s ,a h ,v h ,a l ,v l ,D h,l ,IRI]#(19)
[0169] The above state variables can be obtained through sensing devices and networked devices, or from the simulation environment constructed in step one.
[0170] In addition, the input actions of the energy-saving driving strategy should reflect the EMS power distribution inside the vehicle and the ACC motion control outside the vehicle. Therefore, the action space is represented by a h and P FCS Features:
[0171] a=[a h ,P FCS ]#(20)
[0172] Among them, a h ∈[-1.5,1]m / s 2 ;P FCS Represents the output power of the fuel cell, P FCS ∈[0,60]kW.
[0173] Ultimately, the reward function directly determines the training results of the agent. In order to achieve the five goals of energy-saving driving, the total reward function is configured as follows:
[0174]
[0175] Among them, β1, β2, β3, β4, β5, ω, and σ are weight coefficients. In this embodiment, they can be set as follows: β1 is the price of hydrogen, β2 is the total electricity consumption, β3, β4, and β5 are the replacement prices of the FCS, power battery pack, and drive motor, respectively. σ represents the weight coefficient for balancing driving safety, comfort, and efficiency, and is fixed at 1 / 10. ω is the relative weight coefficient.
[0176] In a preferred embodiment of the present invention, the training data in step 3 is: a C-WTVC-IRI combined cycle constructed using the C-WTVC (China-World Heavy Duty Commercial Vehicle Transient Cycle) standard driving cycle and LTPP (Long Term Pavement Performance) road condition data as training data. The standard driving cycle and road roughness are representative standards for evaluating vehicle driving and road characteristics, respectively. Figure 5 As shown in the figure, the C-WTVC standard driving cycle consists of urban, suburban, and highway scenarios. Based on this, the IRI ranges corresponding to different roads and IRI data from the LTPP are referenced, with an evaluation interval of 100 meters and a normal distribution of IRIs set, forming a C-WTVC-IRI combined cycle. This combination, used in training, enables the intelligent agent to gain a more comprehensive understanding of the interplay between road conditions, vehicle maneuvering, and driving comfort, significantly enhancing the generalization of energy-saving driving training strategies.
[0177] In a preferred embodiment of the present invention, step three is specifically as follows: the intelligent agent in the SSA3C algorithm framework interacts with the simulation environment, observes the current environmental state information, selects and executes actions according to the strategy, enters a new environmental state, and obtains rewards from environmental feedback. The sampling efficiency is improved through experience replay technology, and at the same time, information such as state, action, and reward is stored. This cycle is repeated to achieve training of energy-saving driving strategies. The training steps are as follows:
[0178] Step 1: First, initialize the Actor network and Critic network for approximating the policy and value function. Second, initialize the target Critic network, which is used to assist in estimating the Q value and reduce the estimation error. Next, set the hyperparameters of the algorithm, such as the learning rate, batch size, discount factor, and temperature parameter α (used to adjust the trade-off between policy entropy and value function). Then, set the step size parameter ∈ and kernel function h of the differential gradient descent (SVGD). Finally, define a storage space. As an experience replay pool, and initialize it.
[0179] Step 2: The Actor network transformed based on SVGD interacts with the established training environment: a set of particles {a 0}, update these particles through multiple iterative steps l∈[1,L] to adapt them to the target distribution:
[0180] {a l+1}←{a l}+∈h({a l},s t )#(twenty two)
[0181] Where ∈ is a small step size parameter used to control the size of the particle update in each iteration; h is the kernel function. In this embodiment, the radial basis function is used to measure the similarity between two particles and guide their update. Finally, by deriving the closed-form expression of the particle distribution at the lth iteration and the particle part of the final step L, the strategy is generated:
[0182]
[0183] Where q represents the distribution; π φ (a t ∣s t ) is a policy π with φ as parameter given state s t Output action a t The probability distribution of ; det represents the determinant; I is an identity matrix; is the gradient used to adjust the particle position; then, the reward r is obtained by calculating the reward function and observing t and the next state st+1 . Will experience (s t ,a t ,r t ,s t+1 )Save to the experience replay pool And update the state matrix s←s t+1 .
[0184] Step 3: From the Experience Replay Pool Randomly sample subsets Get N(s t ,a t ,r t ,s t+1 ) for subsequent training.
[0185] Step 4. Calculate the Q-value function of the Critic network, which is used to evaluate the value of actions under different states. Use truncated triple Q learning to calculate the minimum two Q-values and take their average as the final Q-value:
[0186]
[0187] Step 5. Then, using the Kullback-Leibler (KL) divergence D KL , for the policy function π(a t |s t ) has been improved:
[0188]
[0189] Among them, Z(s t ) represents the logarithmic partition function. At the same time, the Q-value function is learned by minimizing the soft Bellman residual to update the parameters θ of the Critic network:
[0190]
[0191] in, For experience playback; (s t ,a t ,r t ,s t+1 ) represents from The mini-batch sampling in ; Q′ represents the target Q-value function of the target critic network; γ is a discount factor that determines the importance of future rewards in the current decision, with a value between [0,1]. The closer γ is to 1, the greater the weight of future rewards, and the closer γ is to 0, the greater the weight of current rewards. Subsequently, the parameter θ′ of the target evaluation network Q′ is updated:
[0192] θ′←(1-τ)θ′+τθ#(27)
[0193] Where τ is the soft update factor.
[0194] Step 6. Subsequently, the entropy H can be calculated by taking the logarithm of a set of particles, that is, logq L (a L |s t ) are averaged to estimate:
[0195]
[0196]
[0197] Among them, Tr represents the trace of the matrix, a 0 is the initial particle swarm, and are the particle swarms after the lth iteration and the last iteration L respectively.
[0198] Step 7: Minimize the strategy induced by the sampler particles The expected KL divergence between the Q value is used to calculate the parameter φ * :
[0199]
[0200] Among them, t sd To limit the standard deviation of particle updates, σ φ is the mean of the initial distribution.
[0201] Step 8. Maximize the following objective function to make the entropy of the policy close to the target entropy value, and finally update the temperature coefficient α:
[0202]
[0203] in, is the target entropy, which is the opposite of the action dimension.
[0204] In a preferred embodiment of the present invention, step four is specifically: downloading the parameterized neural network strategy obtained through offline training to the vehicle controller of the fuel cell bus to realize real-time online application: the fuel cell bus executes the trained energy management strategy and adaptive cruise control.
[0205] An embodiment of the present invention also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the energy-saving driving method of a networked fuel cell bus with enhanced comprehensive comfort.
[0206] The above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A comprehensive comfort-enhanced networked fuel cell bus energy-saving driving method, characterized in that: The steps include: Using a connected fuel cell hybrid bus (FCHEB) as the research object, a collaborative optimization goal for energy-saving driving was established. While considering safety, energy efficiency, and durability, comprehensive driving comfort was also improved. This comfort included vertical comfort, which used weighted root mean square acceleration (WRMSA) as an indicator for assessing human exposure to whole-body vibration. The relationship between WRMSA, the International Roughness Index (IRI), and driving speed was expressed using a linear function. The collaborative optimization goal is quantified into a reward coefficient, and the state space, action space, and reward function are set. Based on the deep reinforcement learning algorithm, a comprehensive comfort-enhanced energy-saving driving strategy is obtained. The state variables in the state space include the state of charge of the power battery, the health status of the power battery, fuel cell, and motor, the speed and acceleration of the main vehicle, the speed and acceleration of the preceding vehicle, the relative distance between the two vehicles, and the IRI. The action space includes the speed of the main vehicle and the power distribution. The reward function integrates the reward coefficients of each goal. The deep reinforcement learning algorithm adopts the differential flexible actor-3-critic (SSA3C) algorithm. The intelligent agent in the algorithm framework interacts with the simulation environment. The intelligent agent observes the current environmental state information, selects and executes actions according to the strategy, enters the new environmental state, and obtains rewards from the environmental feedback. The sampling efficiency is improved by the experience replay technology. At the same time, the state, action, and reward information are stored. This cycle is repeated to achieve the training of the energy-saving driving strategy. The energy-saving driving strategy is trained offline to obtain an inheritable parameterized neural network strategy, which is then downloaded to the vehicle controller to achieve real-time online application of the energy-saving driving strategy.
2. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: The security bonus coefficient is expressed as: Among them D h,l (t), D safe (t) and D max (t) represents the relative distance, safe distance and maximum distance between the two vehicles at time step t respectively; Indicates the maximum speed, v h Indicates the speed of the host vehicle; The comprehensive comfort bonus coefficient is expressed as: Where, jerk(t) is the rate of change of acceleration, a r is the acceleration a h The instantaneous difference, WRMSAa w The instantaneous difference Energy-saving bonus coefficient includes R m (t) and R soc (t), expressed as: in, is the hydrogen consumption rate, SOC(t) is the battery state of charge at time step t, SOC tar is the target value of SOC; Durability bonus coefficients include R bat (t), R FCS (t) and R mot (t), expressed as: Among them, ΔSOH Bat (t), D FCS (t) and CLR(t) are the battery health, FCS degradation rate and cumulative loss rate of the motor at time step t, respectively.
3. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: The states defined in the state space are represented as: s=[SOC,SOH s ,a h ,v h ,a l ,v l ,D h,l ,IRI] Among them, SOH represents the state of charge of the power battery. s =(SOH bat ,SOH FCS ,SOH mot ) represents the health status of the power battery, fuel cell and motor, a h ,v h represents the acceleration and speed of the host vehicle, a l ,v l Indicates the acceleration and speed of the preceding vehicle; The action is represented as: a=[a h ,P FCS ] Among them, a h ∈[-1.5,1]m / s 2 , P FCS Represents the output power of the fuel cell, P FCS ∈[0,60]kW.
4. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: The collaborative optimization goal of energy-saving driving also considers driving efficiency. Driving efficiency refers to maintaining a short headway time TH while ensuring driving safety. The probability density function of the log-normal distribution of TH within the safety range is: Where x(t) represents the TH at time step t, μ and σ represent the mean and standard deviation of TH.
5. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: The total reward function is expressed as: Among them, s, a represent the state and action, R m (t), R soc (t), R FCS (t), R bat (t), R mot (t), R comfort (t), R safe (t), R eff (t) represents the fuel energy saving bonus coefficient, electricity energy saving bonus coefficient, FCS durability bonus coefficient, power battery pack durability bonus coefficient, drive motor durability bonus coefficient, comfort bonus coefficient, safety bonus coefficient and efficiency bonus coefficient respectively; β1, β2, β3, β4, β5, ω and σ are weight coefficients.
6. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: WRMSA w The relationship between the International Roughness Index IRI and driving speed v is expressed as: a w =k1IRI+k2v+k3 IRI v+b Among them, k1, k2, k3, and b are coefficients. The measured IRI at different speeds is used to evaluate WRMSA, and the values of each coefficient are determined by multiple linear regression analysis.
7. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: Acquiring various state variables of a connected fuel cell hybrid bus through sensing devices and connected devices, or acquiring various state variables from a simulation environment; the simulation environment includes a fuel cell vehicle model, a car-following scenario model, and an electrified component health model; The fuel cell vehicle model uses a fuel cell system FCS and a group of lithium-ion power batteries as the main energy sources; in the vehicle dynamics model, the required driving power P req and driving force F req , the corresponding resistance includes the inertial force F a , rolling resistance F r , road slope resistance F i and aerodynamic drag F w , described as follows: Where v represents the bus speed. The FCHEB drive motor model is based on a quasi-steady-state method. The quasi-steady-state model represents the motor as: Among them, P mot , T mot , n mot and η mot (T mot ,n mot ) represent the motor’s power, torque, speed and efficiency at a certain torque and speed respectively; FCS uses an equivalent circuit model consisting of an ideal voltage source and three series resistors; the H2 consumption rate can then be calculated: Among them, P FCS represents the power of FCS, η FCS Indicates the efficiency of FCS, L v Indicates the lower heating value of hydrogen; Use equivalent circuit to simulate the power battery: Among them, V OC is the open circuit voltage, I is the load current, R bat is the internal resistance of the battery, SOC is the state of charge, SOC0 is the initial value of the state of charge, C n is the capacity of the battery; Indicates the total amount of electricity consumed or charged by the battery during this period; Power battery output power P bat and FCSP FCS Maintain the following relationship: P req =T mot ·n mot =P bat +P FCS ·η DC / DC Among them, η DC / DC Represents the efficiency of the DC / DC converter; The following scene model is composed of the main vehicle speed v h (t) and acceleration a h (t), speed of the preceding vehicle v l (t) and acceleration a l (t), mileage of the preceding vehicle L l (t), the main vehicle's mileage L h (t) and the relative distance D between the two vehicles h,l (t) Description: The relationship between speed and distance is as follows: Among them, L veh is the length of the vehicle; In the following process, define the maximum relative distance D between the two vehicles. max and safety distance D safe ; The electric component health model consists of three parts: power battery, FCS and motor; The degradation D of FCS at time step t FCS (t) and state of health SOH FCS (t) is calculated by the following formula: Where n is the total time step, d ss (i),d low (i),d high (i),d cha (i) start-stop, low power load, high power load and load change at time i; Degradation ΔSOH(t) of power batteries and healthy state SOH bat (t) is shown in the following formula: Where I(t) is the load current at time step t, Δt is the current duration, and N(c,T a ) is the equivalent number of cycles when the power battery reaches the end of its life, T a is the average temperature inside the battery, c is the discharge rate; The motor's cumulative loss ratio (CLR) and motor health status (SOH mot ) is defined as follows: Among them, W losses is the total energy loss during the life of the motor; η mot , P mot , η rated and P rated They are the efficiency, power, rated efficiency and rated power of the motor respectively; t life is the manufacturer's recommended motor life.
8. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: The deep reinforcement learning algorithm adopts the flexible actor-critic (SAC) algorithm. In the algorithm framework, the intelligent agent interacts with the real environment or the simulated environment. The intelligent agent observes the current environmental state information, selects and executes actions according to the strategy, enters the new environmental state, and obtains rewards in the form of environmental feedback. The sampling efficiency is improved through experience replay technology. At the same time, the state, action, and reward information are stored. This cycle is repeated to achieve the training of energy-saving driving strategies.
9. The energy-saving driving method for a networked fuel cell bus with enhanced comprehensive comfort according to claim 1, characterized in that: The training steps of the deep reinforcement learning algorithm are as follows: Step 1. Initialize the actor network and critic network for approximating the policy and value function: Initialize the target critic network, which is used to assist in estimating the Q value and reduce the estimation error; set the hyperparameters of the algorithm, set the step size parameter ∈ and kernel function h of the differential gradient descent SVGD; define a storage space As an experience replay pool, and initialize it; Step 2: The Actor network transformed based on SVGD interacts with the established training environment: a set of particles {a 0 }, update these particles through multiple iterative steps l∈[1,L] to adapt them to the target distribution: {a l+1 }←{a l }+∈h({a l },s t ) Among them, ∈ is used to control the size of the particle update in each iteration; h is used to measure the similarity between two particles; by deriving the closed-form expression of the particle distribution at the lth iteration and the particle part of the final step L, the strategy is generated: Where q represents the distribution; π φ (a t ∣s t ) is a policy π with φ as parameter given state s t Output action a t The probability distribution of ; det represents the determinant; I is an identity matrix; is the gradient used to adjust the particle position; then, the reward r is obtained by calculating the reward function and observing t and the next state s t+1 ;Experience(s t ,a t ,r t ,s t+1 )Save to the experience replay pool And update the state matrix s←s t+1 ; Step 3: From the Experience Replay Pool Randomly sample subsets Get N(s t ,a t ,r t ,s t+1 ) for subsequent training; Step 4. Calculate the Q-value function of the Critic network, which is used to evaluate the value of actions under different states. Use truncated triple Q learning to calculate the minimum two Q-values and take their average as the final Q-value: Step 5: Use Kullback-Leibler divergence D KL , for the policy function π(a t |s t ) to improve: Among them, Z(s t ) represents the logarithmic partition function, α is the temperature parameter; at the same time, the Q-value function is learned by minimizing the soft Bellman residual to update the parameters θ of the Critic network: Among them, (s t ,a t ,r t ,s t+1 ) represents from The mini-batch sampling in ; Q′ represents the target Q value function of the target Critic network; γ is the discount factor; then, the parameter θ′ of the target evaluation network Q′ is updated: θ′←(1-τ)θ′+τθ Where, τ is the soft update factor; Step 6. The entropy H is estimated by taking the logarithm over a set of particles: Among them, Tr represents the trace of the matrix, a 0 is the initial particle swarm, and are the particle swarms after the lth iteration and the last iteration L respectively; Step 7: Minimize the strategy induced by the sampler particles The expected KL divergence between the Q value is used to calculate the parameter φ * : Among them, t sd To limit the standard deviation of particle updates, σ φ is the mean of the initial distribution; Step 8. Maximize the following objective function to make the entropy of the policy close to the target entropy value, and finally update the temperature coefficient α: in, is the target entropy, which is the opposite of the action dimension.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by the processor, the steps of the energy-saving driving method of a networked fuel cell bus with enhanced comprehensive comfort are implemented according to any one of claims 1 to 9.
Citation Information
Patent Citations
Speed control multi-target optimized car following algorithm of automatic driving vehicle
CN109709956A
Fuel cell automobile energy-saving driving optimization method and device
CN114103971A