New energy automobile following speed planning and energy distribution method

By constructing a collaborative architecture between upper and lower intelligent agents and combining TD3 and PER algorithms to optimize speed planning and energy management, the problems of dynamic traffic adaptability and energy source aging in traditional methods are solved, and the safety, comfort, energy efficiency and durability are improved.

CN121947523APending Publication Date: 2026-05-01SHANGHAI HANAO NEW ENERGY TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI HANAO NEW ENERGY TECH CO LTD
Filing Date
2025-12-11
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional car-following control methods are not adaptable enough to dynamic traffic environments, and energy management strategies ignore the impact of energy source aging, resulting in poor safety, comfort and energy efficiency, and the optimization effect of conventional algorithms is limited.

Method used

A collaborative architecture between upper and lower layer intelligent agents is constructed. The dual-delay deep deterministic policy gradient algorithm TD3 and the priority experience replay mechanism algorithm PER are combined to optimize velocity planning and energy management. The optimal policy is generated through deep reinforcement learning training.

Benefits of technology

It improves the safety, comfort, energy efficiency, and system durability of hydrogen fuel cell hybrid vehicles in intelligent driving scenarios, adapts to complex dynamic traffic environments, optimizes energy distribution, and extends device lifespan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121947523A_ABST
    Figure CN121947523A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent control of new energy vehicles, in particular to a new energy vehicle following speed planning and energy distribution method, which realizes deep coordination of following speed planning and energy distribution by constructing a layered intelligent agent architecture, adapts to high requirements on dynamic decision in an intelligent driving scene, and improves the driving efficiency. The problem of performance loss caused by separation of the two in a traditional method is solved. And reinforcement learning training of a TD3 algorithm and a PER mechanism is combined, so that the adaptability to a complex dynamic traffic environment is enhanced, and variable car-following scenes in intelligent driving can be better coped with. The power distribution is optimized by incorporating an energy source durability model, the device loss is reduced, and the energy efficiency and the system life are both considered. The layered architecture also improves the decision-making efficiency, meets the real-time requirement of intelligent driving, achieves the comprehensive improvement of safety, comfort, energy efficiency and system durability, and provides a better solution for the application of the hydrogen fuel cell hybrid electric vehicle in the field of intelligent driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for new energy vehicles, specifically a method for planning the following speed and distributing energy in a new energy vehicle. Background Technology

[0002] With the escalating global energy crisis and environmental pollution, traditional gasoline-powered vehicles, due to their inherent high fuel consumption and emissions, are no longer able to meet the demands of sustainable development. Hydrogen fuel cell hybrid vehicles, as a new generation of clean energy vehicles, combine the high energy density of hydrogen fuel cells with the rapid response characteristics of power batteries. While achieving zero carbon emissions, they effectively alleviate the range anxiety of pure electric vehicles, becoming an important development direction in the future of intelligent transportation.

[0003] However, this technology still faces many challenges in practical applications, especially in the coordinated optimization of vehicle dynamics control and its corresponding energy management in intelligent driving scenarios, where mature solutions have not yet been developed. Specifically: As one of the most basic driving modes, car-following places extremely high demands on safety, comfort, and energy efficiency. Traditional car-following control methods mostly rely on fixed safety distance models, whose static decision-making logic is difficult to adapt to complex and dynamic traffic environments. Existing energy management strategies often ignore the impact of frequent acceleration and deceleration during car-following on the lifespan of fuel cells and lithium batteries, leading to decreased system durability and ultimately significantly increasing the long-term operating costs of the vehicle. In addition, single-sensor state perception schemes lack reliability in adverse weather conditions, and conventional control algorithms have limitations such as weak adaptability and difficulty in approximating the global optimum, further restricting the performance of hydrogen fuel cell hybrid vehicles in intelligent driving and transportation systems.

[0004] It is worth noting that among learning-based algorithms, deep reinforcement learning has shown great potential in the fields of speed planning and energy management. Its core idea is derived from dynamic programming. After sufficient training and convergence, the agent can achieve near-global optimal control effect and has excellent optimization performance. It provides a feasible path to solve the technical bottlenecks of vehicle following and energy management mentioned above, but there is no specific technical solution in the existing technology.

[0005] Based on the above reasons, this invention designs a new energy vehicle following speed planning and energy allocation method. By constructing an upper and lower layer intelligent agent collaborative architecture, it solves the problems of weak adaptability of traditional following control to dynamic traffic, neglect of the impact of energy source aging in energy management strategies, and limited optimization effect of conventional algorithms, and achieves synergistic improvement of safety, comfort, energy efficiency and durability in following scenarios. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method for planning the speed of following a vehicle and its energy allocation. By constructing a collaborative architecture of upper and lower intelligent agents, this invention solves the problems of weak adaptability of traditional following control to dynamic traffic, neglect of the impact of energy source aging in energy management strategies, and limited optimization effect of conventional algorithms. This invention achieves a synergistic improvement in safety, comfort, energy efficiency and durability in following scenarios.

[0007] To achieve the above objectives, this invention provides a method for planning the following speed and allocating energy for a new energy vehicle, comprising the following steps: S1. Establish environmental models: Establish a following environment model and a hydrogen fuel cell hybrid electric vehicle powertrain model. S2 is optimized by combining the dual-delay deep deterministic policy gradient algorithm TD3 with the priority experience replay mechanism algorithm PER to construct the hierarchy of the intelligent agent: the hierarchical structure of the upper-level speed planning intelligent agent is constructed according to the car-following environment model, and the hierarchical structure of the lower-level energy management intelligent agent is constructed according to the hydrogen fuel cell hybrid electric vehicle transmission system model. The target speed parameter output by the upper-level speed planning intelligent agent is used as the core constraint condition for the lower-level energy management intelligent agent to perform dynamic power allocation. S3 defines the state space and action space of the upper-level velocity planning agent and designs the reward function, and generates the optimal car-following strategy through deep reinforcement learning training. S4 defines the state space and action space of the lower-level energy management agent and designs the reward function, and generates the optimal energy management strategy through deep reinforcement learning training. S5 enables collaborative control of hierarchical intelligent agents in fuel cell hybrid electric vehicles in car-following scenarios.

[0008] The hydrogen fuel cell hybrid electric vehicle powertrain model in S1 includes a hydrogen fuel cell model considering durability, a lithium battery model considering durability, and a powertrain model.

[0009] The hydrogen fuel cell models that consider durability specifically include: Single cell voltage of hydrogen fuel cells for With activation polarization loss voltage Ohmic polarization loss voltage Concentration polarization loss voltage The difference is calculated using the following formula: Formula 1: ; in, For the temperature of the fuel cell stack, The gas constant is... It is Faraday's constant. and These are the partial pressures of hydrogen and oxygen, respectively. Nernst voltage; Net power output of hydrogen fuel cells The calculation formula is as follows: Formula 2: ; in, For fuel cell stack power, Power consumption for auxiliary equipment, For fuel cell stack current, This refers to the number of individual cells; The fuel cell degradation model is calculated using the following formula: Formula 3: ; in, This represents the total performance degradation of the fuel cell. and These correspond to the durations of high-load conditions (greater than 80% of rated power) and idling conditions (less than 5% of rated power), respectively. This indicates the number of times the fuel cell has been started and stopped. This represents the change in the output power of the fuel cell.

[0010] Lithium-ion battery models that consider durability specifically include: The charging and discharging behavior of the battery is characterized by an equivalent internal resistance model. Indicates open-circuit voltage. This indicates the total internal resistance of the battery pack. Indicates the battery output power. This indicates the battery current. With battery power demand The formula for calculating the relationship between them is: Formula 4: ; The SOC of a lithium battery is obtained by the ampere-hour integration method, and the calculation formula is as follows: Formula 5: ; in, This represents the initial state of charge (SOC) of the lithium battery. This refers to the nominal capacity of the lithium battery. The capacity reduction caused by changes in the physical and chemical properties of lithium-ion batteries during energy transmission and recycling is simulated and predicted using an aging model based on energy throughput. The calculation formula is as follows: Formula Six: ; in, The percentage of capacity loss. As the pre-exponential factor, This refers to the charge / discharge rate. Let be the ideal gas constant. This refers to the absolute temperature of the battery. To accumulate throughput per ampere-hour.

[0011] The powertrain model specifically includes: The formula for calculating vehicle traction force is as follows: Formula 7: ; in, For rolling friction resistance, For air resistance, To increase resistance, For slope resistance; For the quality of the vehicle, It is the acceleration due to gravity. The rolling resistance coefficient, For road slope, air density, The air drag coefficient, The vehicle's frontal area. The moment of inertia coefficient of the vehicle; By the speed of the motor and torque Derive the electric power of the motor. : Formula 8: ; Among them, motor efficiency It is torque and rotational speed The function can be determined by the efficiency diagram of the drive motor; The power required by the powertrain is provided by both lithium-ion batteries and fuel cells, and the formula for calculating its power balance is as follows: Formula Nine: ; in, For the power required by the powertrain, and The output power of fuel cells and lithium-ion batteries, respectively. and The efficiency of the DC / AC converter and the DC / DC converter are respectively.

[0012] The car-following environment model specifically includes: Dynamic safety clearances are used for constraints, and a minimum distance is defined. Maximum distance from the ideal This makes the actual following distance The formula for calculating a value between the two is as follows: Formula 10: ; Formula 11: ; Formula 12: ; in, The distance traveled by the vehicle in front. The initial distance between the two vehicles. The distance traveled by the main vehicle; Main vehicle speed, It is the sum of braking delay and driver reaction time. This represents an emergency deceleration.

[0013] The specific content of S2 includes: S2-1, After initializing the network parameters, the agent first explores the environment by relying on random processes to obtain the initial state of the system parameters; S2-2, Entering the training loop: The agent interacts with the environment to generate experience samples, and manages the experience replay pool based on the PER mechanism. In the Priority Experience Replay Mechanism (PER) algorithm, the sampling probability is defined as: Formula Thirteen: ; in, Priority factor, priority , It is the TD error of the j-th empirical rule; Use small positive numbers to avoid ignoring the experience when the TD error is zero; S2-3, samples are drawn from the playback pool with probability S2-2. To correct for bias introduced by preferential sampling, an importance sampling weight is used, calculated using the following formula: Formula Fourteen: ; in, For the total number of experiences, This is a compensation coefficient used to control the adjustment of importance sampling weights; S2-4, During network updates, in the dual-delay deep deterministic policy gradient algorithm TD3, the Critic network is optimized using gradient descent, with the loss function being: Formula 15: ; S2-5, the Actor network is based on the Critic network for value evaluation and is optimized using the gradient ascent method. Its parameter update formula is: Formula Sixteen: ; The target network parameters are synchronized via a soft update mechanism, using the following formula: Formula 17: ; Formula 18: ; After each preset number of training steps, the current network parameters are slowly synchronized to the target network according to the soft update formula to maintain the stability of the training process. Through multiple iterations, the agent converges to the optimal strategy that adapts to the vehicle following and energy management scenarios, and outputs precise control decisions.

[0014] The specific content of S3 includes: S3-1, the state space of the upper-level velocity planning agent is defined as follows: Formula 19: ; in, This represents the input state of the upper-level intelligent agent. These represent the speed and acceleration of the vehicle in front, respectively. This is the distance between the two workshops; S3-2, the action space of the upper-level velocity planning agent determines its feasible operations in a specific scenario, defined as the vehicle's acceleration, and its formula is: Formula 20: ; S3-3, the upper-level velocity planning agent learns the mapping from the state space to the action space by optimizing the reward function, in order to balance safety, comfort, and energy efficiency. Its calculation formula is as follows: Formula 21: ; Formula 22: ; Formula 23: ; Formula 24: ; in, , and These correspond to the comfort, safety, and economy bonuses, respectively. , and These are weighting coefficients used to adjust the priority of each objective; , and These are energy consumption characteristic parameters, determined by the specific vehicle model. The nominal kinetic energy is used; a constant is introduced. This ensures that each training iteration reaches the maximum number of rounds.

[0015] The specific content of S4 includes: S4-1: Based on the driving speed provided by the upper-layer speed planning agent, the lower-layer energy management agent performs energy management functions. S4-2, the state space definition formula for the lower-level energy management agent is as follows: Formula 25: ; S4-3, the power of the fuel cell system, after normalization, is used as the action state output of the lower-level energy management agent, and its formula is: Formula 26: ; S4-4, the optimization objective of the lower-level energy management agent is to generate the optimal power allocation strategy to minimize overall operating costs and maintain the battery SOC within a reasonable fluctuation range. This includes hydrogen consumption costs, fuel cell degradation costs, and lithium-ion battery aging costs; to effectively maintain SOC within the target range, a penalty is introduced. ; S4-5, based on the requirements of S4-4, sets the reward function for the lower-level energy management agent, with the following formula: Formula 27: ; Formula 28: ; Formula 29: ; in, It is a positive weighting factor for overall operating costs. It is a positive weighting factor for maintaining battery SOC; and The selection of factors is based on comprehensive analysis and multiple experiments, taking into account the sensitivity of each value to the optimization objective, in order to ensure a balance between different objectives within the same dimension; and These are the unit prices of fuel cells and lithium-ion batteries, respectively. This indicates the mass of hydrogen gas consumed. Indicates the degradation rate of the fuel cell. Indicates the aging rate of lithium-ion batteries; This indicates the maximum power of the fuel cell system. This indicates the total energy storage of a lithium-ion battery pack.

[0016] The specific content of S5 is as follows: Under the given speed condition of the preceding vehicle, the master vehicle first completes speed planning through the upper-level speed planning agent. After determining the speed condition of the master vehicle, the lower-level energy management agent performs energy management. Through the cooperation of the two agents, the vehicle completes the driving task under the predetermined condition.

[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention achieves deep collaboration between speed planning and energy allocation by constructing a hierarchical intelligent agent architecture, which is adapted to the high requirements of dynamic decision-making in intelligent driving scenarios and solves the performance loss problem caused by the separation of the two in traditional methods.

[0018] This invention combines the TD3 algorithm with the PER mechanism for reinforcement learning training, enhancing adaptability to complex and dynamic traffic environments and better handling the ever-changing car-following scenarios in intelligent driving. Simultaneously, by incorporating an energy source durability model to optimize power allocation, it reduces device losses while balancing energy efficiency and system lifespan. The layered architecture also improves decision-making efficiency, meeting the real-time requirements of intelligent driving and achieving a comprehensive improvement in safety, comfort, energy efficiency, and system durability, providing a superior solution for the application of hydrogen fuel cell hybrid vehicles in the field of intelligent driving. Attached Figure Description

[0019] Figure 1 This is a topology diagram of a hydrogen fuel cell hybrid electric vehicle in an embodiment of the present invention; Figure 2 This is a schematic diagram of a vehicle following model in an embodiment of the present invention; Figure 3 This is a control block diagram of the hierarchical intelligent agent in an embodiment of the present invention; Figure 4 This is a graph showing the reward change of the speed planning agent in an embodiment of the present invention; Figure 5 This is a graph showing the reward change of the energy management agent in an embodiment of the present invention; Figure 6 This is a distance result diagram of the following speed planning under a given preceding vehicle condition in an embodiment of the present invention; Figure 7 This is a power allocation diagram of the energy management method in this embodiment of the invention at a given car-following speed; Figure 8 This is a comparison chart of fuel cell degradation between the energy management method in this embodiment of the invention and other comparative methods at a given race speed. Detailed Implementation

[0020] The present invention will now be further described with reference to the accompanying drawings.

[0021] See Figures 1-8 This invention provides a method for planning the following speed and distributing energy in a new energy vehicle, comprising the following steps: S1. Establish environmental models: Establish a following environment model and a hydrogen fuel cell hybrid electric vehicle powertrain model. S2 is optimized by combining the dual-delay deep deterministic policy gradient algorithm TD3 with the priority experience replay mechanism algorithm PER to construct the hierarchy of the intelligent agent: the hierarchical structure of the upper-level speed planning intelligent agent is constructed according to the car-following environment model, and the hierarchical structure of the lower-level energy management intelligent agent is constructed according to the hydrogen fuel cell hybrid electric vehicle transmission system model. The target speed parameter output by the upper-level speed planning intelligent agent is used as the core constraint condition for the lower-level energy management intelligent agent to perform dynamic power allocation. S3 defines the state space and action space of the upper-level velocity planning agent and designs the reward function, and generates the optimal car-following strategy through deep reinforcement learning training. S4 defines the state space and action space of the lower-level energy management agent and designs the reward function, and generates the optimal energy management strategy through deep reinforcement learning training. S5 enables collaborative control of hierarchical intelligent agents in fuel cell hybrid electric vehicles in car-following scenarios.

[0022] like Figure 1 The topological structure diagram of the fuel cell hybrid electric vehicle of the present invention is shown, wherein the hydrogen fuel cell is connected to the DC bus through a DC / DC converter; the lithium-ion battery is directly connected to the DC bus and can be charged or discharged according to the vehicle's operating status; the electrical energy drives the motor through a DC / AC inverter to ensure the normal operation of the vehicle.

[0023] The hydrogen fuel cell hybrid electric vehicle powertrain model in S1 includes a hydrogen fuel cell model considering durability, a lithium battery model considering durability, and a powertrain model.

[0024] The hydrogen fuel cell models that consider durability specifically include: Single cell voltage of hydrogen fuel cells for With activation polarization loss voltage Ohmic polarization loss voltage Concentration polarization loss voltage The difference is calculated using the following formula: ; in, For the temperature of the fuel cell stack, The gas constant is... It is Faraday's constant. and These are the partial pressures of hydrogen and oxygen, respectively. Nernst voltage; ; in, For the temperature of the fuel cell stack, The gas constant is... It is Faraday's constant. and These are the partial pressures of hydrogen and oxygen, respectively.

[0025] Activation polarization loss voltage, also known as activation overvoltage, originates from the electrochemical reaction process in a fuel cell: electron transfer and the breaking or formation of chemical bonds both consume energy. Therefore, a portion of the energy output by the battery is used to overcome the activation energy, and the resulting energy loss forms the activation overvoltage. The expression for activation overvoltage is: ; ; in, For flow transport coefficient, The heat transfer coefficient, The thermal conductivity coefficient, The coefficient for the electrochemical dynamic reaction, For fuel cell current, This represents the oxygen concentration on the cathode catalyst layer.

[0026] Ohmic polarization loss voltage refers to the energy loss caused by the internal resistance of a conductor when current flows through it. According to Ohm's law, its calculation formula is: ; ; ; in, The ionic impedance of the electrolyte. It includes the resistance of the bipolar plates, battery interconnects, contacts, and other battery components through which electrons flow. It is the resistivity of the membrane. It is the thickness of the membrane. It is the activated area of ​​the battery. It refers to the water content.

[0027] Concentration polarization loss voltage is mainly due to the inability of reactants to reach the electrode surface in a timely manner during the electrochemical reaction, leading to a decrease in reactant concentration and thus affecting the battery voltage. For proton exchange membrane fuel cells, concentration polarization voltage is primarily related to current, and its expression is: ; in, For coefficients, For actual current density, This represents the maximum current density of the fuel cell stack.

[0028] Net power output of hydrogen fuel cells The calculation formula is as follows: ; in, For fuel cell stack power, Power consumption for auxiliary equipment, For fuel cell stack current, This refers to the number of individual cells; The fuel cell degradation model is calculated using the following formula: ; in, This represents the total performance degradation of the fuel cell. and These correspond to the durations of high-load conditions (greater than 80% of rated power) and idling conditions (less than 5% of rated power), respectively. This indicates the number of times the fuel cell has been started and stopped. This represents the change in the output power of the fuel cell.

[0029] Lithium-ion battery models that consider durability specifically include: The charging and discharging behavior of the battery is characterized by an equivalent internal resistance model. Indicates open-circuit voltage. This indicates the total internal resistance of the battery pack. Indicates the battery output power. This indicates the battery current. With battery power demand The formula for calculating the relationship between them is: ; The SOC of a lithium battery is obtained by the ampere-hour integration method, and the calculation formula is as follows: ; in, This represents the initial state of charge (SOC) of the lithium battery. This refers to the nominal capacity of the lithium battery. The capacity reduction caused by changes in the physical and chemical properties of lithium-ion batteries during energy transmission and recycling is simulated and predicted using an aging model based on energy throughput. The calculation formula is as follows: ; in, The percentage of capacity loss. As the pre-exponential factor, This refers to the charge / discharge rate. The ideal gas constant This refers to the absolute temperature of the battery. To accumulate throughput per ampere-hour.

[0030] The powertrain model specifically includes: The formula for calculating vehicle traction force is as follows: ; in, For rolling friction resistance, For air resistance, To increase resistance, For slope resistance; For the quality of the vehicle, It is the acceleration due to gravity. The rolling resistance coefficient, For road slope, air density, The air drag coefficient, The vehicle's frontal area. The moment of inertia coefficient of the vehicle; By the speed of the motor and torque Derive the electric power of the motor. : ; Among them, motor efficiency It is torque and rotational speed The function can be determined by the efficiency diagram of the drive motor; The power required by the powertrain is provided by both lithium-ion batteries and fuel cells, and the formula for calculating its power balance is as follows: ; in, For the power required by the powertrain, and The output power of fuel cells and lithium-ion batteries, respectively. and The efficiency of the DC / AC converter and the DC / DC converter are respectively.

[0031] The car-following environment model specifically includes: like Figure 2 As shown, dynamic safety distances are used for constraints, and a minimum distance is defined. Maximum distance from the ideal This makes the actual following distance The formula for calculating a value between the two is as follows: ; ; ; in, The distance traveled by the vehicle in front. The initial distance between the two vehicles. The distance traveled by the main vehicle; Main vehicle speed, It is the sum of braking delay and driver reaction time. This represents an emergency deceleration.

[0032] The specific content of S2 includes: S2-1, After initializing the network parameters, the agent first explores the environment by relying on random processes to obtain the initial state of the system parameters; like Figure 3As shown in S2-2, the training loop begins: the agent interacts with the environment to generate experience samples, and manages the experience replay pool based on the PER mechanism. In the Priority Experience Replay Mechanism (PER) algorithm, the sampling probability of experience is defined as: ; in, Priority factor, priority , It is the TD error of the j-th empirical rule; Use small positive numbers to avoid ignoring the experience when the TD error is zero; S2-3, samples are drawn from the playback pool with probability S2-2. To correct for bias introduced by preferential sampling, an importance sampling weight is used, calculated using the following formula: ; in, For the total number of experiences, This is a compensation coefficient used to control the adjustment of importance sampling weights; S2-4, During network updates, in the dual-delay deep deterministic policy gradient algorithm TD3, the Critic network is optimized using gradient descent, with the loss function being: ; S2-5, the Actor network is based on the Critic network for value evaluation and is optimized using the gradient ascent method. Its parameter update formula is: ; The target network parameters are synchronized via a soft update mechanism, using the following formula: ; ; After each preset number of training steps, the current network parameters are slowly synchronized to the target network according to the soft update formula to maintain the stability of the training process. Through multiple iterations, the agent converges to the optimal strategy that adapts to the vehicle following and energy management scenarios, and outputs precise control decisions.

[0033] The specific content of S3 includes: S3-1, the state space of the upper-level velocity planning agent is defined as follows: ; in, This represents the input state of the upper-level intelligent agent. These represent the speed and acceleration of the vehicle in front, respectively. This is the distance between the two workshops; S3-2, the action space of the upper-level velocity planning agent determines its feasible operations in a specific scenario, defined as the vehicle's acceleration, and its formula is: ; S3-3, the upper-level velocity planning agent learns the mapping from the state space to the action space by optimizing the reward function, in order to balance safety, comfort, and energy efficiency. Its calculation formula is as follows: ; ; ; ; in, , and These correspond to the comfort, safety, and economy bonuses, respectively. , and These are weighting coefficients used to adjust the priority of each objective; , and These are energy consumption characteristic parameters, determined by the specific vehicle model. The nominal kinetic energy is used; a constant is introduced. This ensures that each training iteration reaches the maximum number of rounds.

[0034] After completing the above settings, the agent training phase begins: The environment module and the agent algorithm module are connected to build an interactive learning framework between the environment and the agent; agent hyperparameters and experience pool capacity are configured, and the training iteration process is initiated. When the total reward value obtained during the iteration process tends to stabilize and converge, and the learning effect reaches the expected level, training is terminated and the model is saved. See the appendix for the reward change curve of the velocity planning agent. Figure 4 .

[0035] The specific content of S4 includes: S4-1: Based on the driving speed provided by the upper-layer speed planning agent, the lower-layer energy management agent performs energy management functions. The energy management agent comprehensively considers the operating costs and degradation characteristics of the two energy sources, as their output power is interdependent: when the output power of one energy source changes, its impact on lifespan or cost will have a reciprocal effect on the lifespan or cost of the other energy source. Therefore, the core of this step lies in how to efficiently allocate the power of the two energy sources to achieve the dual goals of extending lifespan and reducing costs while meeting vehicle performance requirements.

[0036] S4-2, the state space definition formula for the lower-level energy management agent is as follows: ; S4-3, the power of the fuel cell system, after normalization, is used as the action state output of the lower-level energy management agent, and its formula is: ; S4-4, the optimization objective of the lower-level energy management agent is to generate the optimal power allocation strategy to minimize overall operating costs and maintain the battery SOC within a reasonable fluctuation range. This includes hydrogen consumption costs, fuel cell degradation costs, and lithium-ion battery aging costs; to effectively maintain SOC within the target range, a penalty is introduced. ; S4-5, based on the requirements of S4-4, sets the reward function for the lower-level energy management agent, with the following formula: ; ; ; in, It is a positive weighting factor for overall operating costs. It is a positive weighting factor for maintaining battery SOC; and The selection of factors is based on comprehensive analysis and multiple experiments, taking into account the sensitivity of each value to the optimization objective, in order to ensure a balance between different objectives within the same dimension; and These are the unit prices of fuel cells and lithium-ion batteries, respectively. This indicates the mass of hydrogen gas consumed. Indicates the degradation rate of the fuel cell. Indicates the aging rate of lithium-ion batteries; This indicates the maximum power of the fuel cell system. This indicates the total energy storage of a lithium-ion battery pack.

[0037] The specific content of S5 is as follows: Under the given speed condition of the preceding vehicle, the master vehicle first completes speed planning through the upper-level speed planning agent. After determining the speed condition of the master vehicle, the lower-level energy management agent performs energy management. Through the cooperation of the two agents, the vehicle completes the driving task under the predetermined condition.

[0038] The above are merely preferred embodiments of the present invention, intended only to aid in understanding the method and core ideas of this application. The scope of protection of the present invention is not limited to the above embodiments; all technical solutions falling within the scope of the present invention's concept are within its protection. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

[0039] This invention comprehensively addresses the shortcomings of existing technologies in the coordinated optimization of vehicle dynamics control and corresponding energy management for hydrogen fuel cell vehicles in intelligent driving scenarios. It also addresses the need for further research and improvement in the combined application of deep reinforcement learning in speed planning and energy management. By combining the dual-delay deep deterministic policy gradient (TD3) with the priority experience replay (PER) algorithm, the decision-making stability and efficiency of the intelligent agent are improved. A hierarchical intelligent agent architecture is established. The upper-layer speed planning agent uses 5G-CV2X vehicle-to-vehicle cooperative communication to obtain the speed curve actively uploaded by the preceding vehicle in real time. This curve is then verified and fused with relative distance information collected by millimeter-wave radar to generate a safe and comfortable following speed. The lower-layer energy management agent, based on the target vehicle speed output from the upper layer, introduces a dual-energy-source aging model to dynamically allocate the output power of the fuel cell and lithium battery, balancing energy efficiency and device aging costs while meeting the vehicle's power requirements.

Claims

1. A method for planning the following speed and allocating energy for a new energy vehicle, characterized in that, Includes the following steps: S1. Establish environmental models: Establish a following environment model and a hydrogen fuel cell hybrid electric vehicle powertrain model. S2, the hierarchy of the intelligent agent is constructed by optimizing the algorithm based on the dual-delay deep deterministic strategy gradient algorithm TD3 combined with the priority experience replay mechanism algorithm PER: the hierarchical structure of the upper-level speed planning intelligent agent is constructed according to the car-following environment model, and the hierarchical structure of the lower-level energy management intelligent agent is constructed according to the hydrogen fuel cell hybrid electric vehicle transmission system model. The target speed parameter output by the upper-level speed planning intelligent agent is used as the core constraint condition for the lower-level energy management intelligent agent to perform dynamic power allocation. S3 defines the upper-level velocity planning agent. state space and Action space A reward function was designed, and the optimal following strategy was generated through deep reinforcement learning training. S4 defines the state space and action space of the lower-level energy management agent and designs the reward function, and generates the optimal energy management strategy through deep reinforcement learning training. S5 enables collaborative control of hierarchical intelligent agents in fuel cell hybrid electric vehicles in car-following scenarios.

2. The new energy vehicle following speed planning and energy distribution method according to claim 1, characterized in that, The hydrogen fuel cell hybrid electric vehicle powertrain model in S1 includes a hydrogen fuel cell model considering durability, a lithium battery model considering durability, and a powertrain model.

3. The new energy vehicle following speed planning and energy distribution method according to claim 2, characterized in that, The hydrogen fuel cell model considering durability specifically includes: Single cell voltage of hydrogen fuel cells for With activation polarization loss voltage Ohmic polarization loss voltage Concentration polarization loss voltage The difference is calculated using the following formula: Formula 1: ; in, For the temperature of the fuel cell stack, The gas constant is... It is Faraday's constant. and These are the partial pressures of hydrogen and oxygen, respectively. Nernst voltage; Net power output of hydrogen fuel cells The calculation formula is as follows: Formula 2: ; in, For fuel cell stack power, Power consumption for auxiliary equipment, For fuel cell stack current, This refers to the number of individual cells; The fuel cell degradation model is calculated using the following formula: Formula 3: ; in, This represents the total performance degradation of the fuel cell. and These correspond to the durations of high-load conditions (greater than 80% of rated power) and idling conditions (less than 5% of rated power), respectively. This indicates the number of times the fuel cell has been started and stopped. This represents the change in the output power of the fuel cell.

4. The new energy vehicle following speed planning and energy distribution method according to claim 2, characterized in that, The lithium battery model considering durability specifically includes: The charging and discharging behavior of the battery is characterized by an equivalent internal resistance model. Indicates open-circuit voltage. This indicates the total internal resistance of the battery pack. Indicates the battery output power. This indicates the battery current. With battery power demand The formula for calculating the relationship between them is: Formula 4: ; The SOC of a lithium battery is obtained by the ampere-hour integration method, and the calculation formula is as follows: Formula 5: ; in, This represents the initial state of charge (SOC) of the lithium battery. This refers to the nominal capacity of the lithium battery. The capacity reduction caused by changes in the physical and chemical properties of lithium-ion batteries during energy transmission and recycling is simulated and predicted using an aging model based on energy throughput. The calculation formula is as follows: Formula Six: ; in, The percentage of capacity loss. As the pre-exponential factor, This refers to the charge / discharge rate. Let be the ideal gas constant. This refers to the absolute temperature of the battery. To accumulate throughput per ampere-hour.

5. The new energy vehicle following speed planning and energy distribution method according to claim 2, characterized in that, The powertrain model specifically includes: The formula for calculating vehicle traction force is as follows: Formula 7: ; in, For rolling friction resistance, For air resistance, To increase resistance, For slope resistance; For the quality of the vehicle, It is the acceleration due to gravity. The rolling resistance coefficient, For road slope, air density, The air drag coefficient, The vehicle's frontal area. The moment of inertia coefficient of the vehicle; By the speed of the motor and torque Derive the electric power of the motor. : Formula 8: ; Among them, motor efficiency It is torque and rotational speed The function can be determined by the efficiency diagram of the drive motor; The power required by the powertrain is provided by both lithium-ion batteries and fuel cells, and the formula for calculating its power balance is as follows: Formula Nine: ; in, For the power required by the powertrain, and The output power of fuel cells and lithium-ion batteries, respectively. and The efficiency of the DC / AC converter and the DC / DC converter are respectively.

6. The new energy vehicle following speed planning and energy distribution method according to claim 1, characterized in that, The car-following environment model specifically includes: Dynamic safety clearances are used for constraints, and a minimum distance is defined. Maximum distance from the ideal This makes the actual following distance The formula for calculating a value between the two is as follows: Formula 10: ; Formula 11: ; Official Twelve: ; in, The distance traveled by the vehicle in front. The initial distance between the two vehicles. The distance traveled by the main vehicle; Main vehicle speed, It is the sum of braking delay and driver reaction time. This represents an emergency deceleration.

7. The new energy vehicle following speed planning and energy distribution method according to claim 6, characterized in that, The specific content of S2 includes: S2-1, After initializing the network parameters, the agent first explores the environment by relying on a random process to obtain the initial state of the system parameters; S2-2, Entering the training loop: The agent interacts with the environment to generate experience samples, and manages the experience replay pool based on the PER mechanism. In the Priority Experience Replay Mechanism (PER) algorithm, the sampling probability of experience is defined as: Formula Thirteen: ; in, Priority factor, priority , It is the TD error of the j-th empirical rule; The value is a small positive number to avoid the empirical value being ignored when the TD error is zero; S2-3, Sampling is performed from the playback pool according to the probability stated in S2-2. Simultaneously, to correct for bias introduced by preferential sampling, an importance sampling weight is used, calculated using the following formula: Formula Fourteen: ; in, For the total number of experiences, This is a compensation coefficient used to control the adjustment of importance sampling weights; S2-4, During network updates, in the dual-delay deep deterministic policy gradient algorithm TD3, the Critic network is optimized using gradient descent, with the loss function being: Formula 15: ; S2-5, the Actor network is based on the Critic network for value evaluation and is optimized using the gradient ascent method. Its parameter update formula is: Formula Sixteen: ; The target network parameters are synchronized via a soft update mechanism, using the following formula: Formula 17: ; Formula 18: ; After each preset number of training steps are completed, the current network parameters are slowly synchronized to the target network according to the soft update formula to maintain the stability of the training process. Through multiple iterations, the agent converges to the optimal strategy that adapts to the vehicle following and energy management scenarios, and outputs precise control decisions.

8. The new energy vehicle following speed planning and energy distribution method according to claim 6, characterized in that, The specific content of S3 includes: S3-1, the state space of the upper-layer velocity planning agent is defined by the following formula: Formula 19: ; in, This represents the input state of the upper-level intelligent agent. These represent the speed and acceleration of the vehicle in front, respectively. This is the distance between the two workshops; S3-2, the action space of the upper-layer velocity planning agent determines its feasible operations in a specific scenario, defined as the vehicle acceleration, and its formula is: Formula 20: ; S3-3, the upper-layer velocity planning agent learns the mapping from the state space to the action space by optimizing the reward function, in order to balance safety, comfort, and energy efficiency. The calculation formula is as follows: Formula 21: ; Formula 22: ; Formula 23: ; Formula 24: ; in, , and These correspond to the comfort, safety, and economy bonuses, respectively. , and These are weighting coefficients used to adjust the priority of each objective; , and These are energy consumption characteristic parameters, determined by the specific vehicle model. The nominal kinetic energy is used; a constant is introduced. This ensures that each training iteration reaches the maximum number of rounds.

9. The new energy vehicle following speed planning and energy distribution method according to claim 6, characterized in that, The specific content of S4 includes: S4-1, Based on the driving speed provided by the upper-layer speed planning agent, the lower-layer energy management agent performs energy management functions; S4-2, the state space definition formula for the lower-level energy management agent is as follows: Formula 25: ; S4-3, the power of the fuel cell system, after normalization, is used as the action state output of the lower-level energy management agent, and its formula is: Formula 26: ; S4-4, the optimization objective of the lower-level energy management agent is to generate an optimal power allocation strategy to minimize overall operating costs and maintain battery SOC within a reasonable fluctuation range. The overall operating cost... This includes hydrogen consumption costs, fuel cell degradation costs, and lithium-ion battery aging costs; to effectively maintain SOC within the target range, a penalty is introduced. ; S4-5, according to the requirements of S4-4, set the reward function of the lower-level energy management agent, the formula of which is: Formula 27: ; Formula 28: ; Formula 29: ; in, It is a positive weighting factor for overall operating costs. It is a positive weighting factor for maintaining battery SOC; the aforementioned and The selection of factors is based on comprehensive analysis and multiple experiments, taking into account the sensitivity of each value to the optimization objective, in order to ensure a balance between different objectives within the same dimension; and These are the unit prices of fuel cells and lithium-ion batteries, respectively. This indicates the mass of hydrogen gas consumed. Indicates the degradation rate of the fuel cell. Indicates the aging rate of lithium-ion batteries; This indicates the maximum power of the fuel cell system. This indicates the total energy storage of a lithium-ion battery pack.

10. The new energy vehicle following speed planning and energy distribution method according to claim 1, characterized in that, The specific content of S5 is as follows: Under the given speed condition of the preceding vehicle, the master vehicle first completes speed planning through the upper-level speed planning agent to determine the master vehicle speed condition, and then the lower-level energy management agent performs energy management. Through the cooperation of the two agents, the vehicle completes the driving task under the predetermined condition.

Citation Information

Cited By

  • An adaptive health state hierarchical fuel cell vehicle energy management method

    CN122211256A