Intelligent network connection HEV vehicle-road cooperation hierarchical ecological driving control method and system

By using intelligent connected HEVs for speed planning and power split control, and by optimizing vehicle speed trajectory and power source distribution using DRL algorithm and A-ECMS, the problems of fuel consumption optimization and real-time performance in existing technologies are solved, and multi-objective optimization of safety, trafficability and fuel economy is achieved.

CN115955712BActive Publication Date: 2026-04-17SHANGHAI JIAOTONG UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI JIAOTONG UNIV
Filing Date
2022-11-21
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing intelligent connected HEV control strategies fail to fully utilize the characteristics of multiple power sources, making it difficult to optimize fuel consumption in complex traffic scenarios. Furthermore, their high computational complexity makes it difficult to meet real-time requirements.

Method used

The system employs reinforcement learning algorithms combined with V2X information and vehicle status. Through vehicle speed planning and power distribution control modules, it optimizes vehicle speed trajectory and power allocation. The DRL algorithm is used to integrate multi-objective reward functions, and A-ECMS is used for real-time fuel economy management.

Benefits of technology

It achieves multi-objective optimization of safety, traffic flow, and fuel economy in complex traffic scenarios, reducing computational complexity and improving real-time performance and fuel economy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115955712B_ABST
    Figure CN115955712B_ABST
Patent Text Reader

Abstract

This invention relates to a hierarchical ecological driving control method and system for intelligent connected HEVs (Hybrid Electric Vehicles). The method includes: a speed planning step: receiving V2X information and combining it with the vehicle's state information to solve for the optimal speed trajectory; and a power distribution control step: receiving the optimal speed trajectory information, using it as a reference for following, and allocating power output from different power sources based on real-time power demand to maximize vehicle fuel economy. Compared with existing technologies, this invention achieves higher fuel economy while ensuring safety and traffic flow, and can still complete multi-objective optimal speed planning under real-time conditions in complex traffic scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle control technology, and in particular to a method and system for intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control. Background Technology

[0002] In recent years, while automobiles have made travel more convenient, they have also brought a series of challenges, including energy crises, environmental problems, and traffic safety. Therefore, the research and development of new energy and connected vehicles has received increasing attention. Intelligent connected HEVs represent the future form of automobiles and an effective solution to existing energy, environmental, and traffic issues.

[0003] To respond to this trend, exploring safe, efficient, and environmentally friendly control strategies for intelligent connected HEVs is imperative. HEV control strategies need to dynamically allocate engine and motor output power during driving to optimize engine operating curves and maintain appropriate battery state of charge (SOC), thereby improving fuel economy. In the application scenarios of intelligent connected HEVs, control strategies not only need to consider the power distribution between the engine and motor at the powertrain level but also involve speed planning at the vehicle level, comprehensively considering complex traffic information such as the speed of the vehicle ahead, traffic lights, and road speed limits. Currently, solutions for intelligent connected HEV eco-driving technology often extend from the traditional HEV EMS (Electric Vehicle Management System), modeling eco-driving strategies as optimal control problems, where fuel consumption is the optimization objective, and vehicle status, traffic flow information, and traffic light phases serve as constraints.

[0004] The invention disclosed in CN113561793A presents a dynamically constrained intelligent fuel cell vehicle EMS, proposing a fuel cell vehicle control strategy based on dynamic constraints, including a vehicle speed planning module and an energy management module. The vehicle speed planning module uses the fast projected gradient method to obtain the optimal speed that minimizes energy consumption at the current moment, and uses a safety distance as a dynamic constraint for the speed planning problem to ensure safety. The energy management module, based on model predictive control, establishes a control function, controls the current, voltage, and battery SOC of the target vehicle according to the current planned vehicle speed to achieve power distribution, and outputs vehicle driving force in the form of torque.

[0005] The invention disclosed in CN110264757A presents a hierarchical speed planning method for intelligent connected vehicles based on continuous traffic light information, aiming to enable vehicles to pass through all intersections with minimal stops. This invention calculates the phase and distance information of traffic lights at intersections ahead of the vehicle and transforms it into the ideal time and speed for the vehicle to reach each intersection. It then divides the entire driving process according to the intersection location, constructing the speed planning for each segment between two intersections as an optimal control problem. Using the obtained ideal time and speed for reaching the intersection as constraints, the optimal speed for each segment is solved.

[0006] The invention disclosed in CN112498334A presents a robust energy management method for intelligent connected hybrid electric vehicles (HEVs), proposing a reinforcement learning-based energy management method that incorporates traffic information. This method first obtains the driver's driving intention, determines the driving mode based on this intention, and then obtains the speed and torque requirements from the driving mode. Subsequently, it acquires traffic information through intelligent transportation systems to predict global operating conditions and short-term real-time operating conditions (i.e., whether there is congestion ahead). Based on the global operating conditions, dynamic programming is used to solve for the optimal State of Charge (SOC) reference trajectory, and this trajectory is updated periodically using a rolling time domain based on the short-term real-time operating conditions. Finally, energy management is achieved through reinforcement learning, where the state variables are vehicle speed, the difference between the current SOC and the reference trajectory, and the deviation of the engine operating point from the high-efficiency region; the control variable is the engine output power.

[0007] The invention disclosed in CN112526883A presents a vehicle energy management method based on intelligent connected vehicle information, proposing a hierarchical speed planning and energy management method for pure electric vehicles. In terms of speed planning, this method collects intelligent connected vehicle information such as traffic lights and the position of the vehicle ahead, as well as vehicle state information such as power limits and acceleration limits. Using this information as constraints, it solves for the optimal vehicle acceleration based on dynamic programming. In terms of energy management, based on the obtained optimal vehicle speed trajectory, this method employs a fuzzy control-based approach to allocate and manage the electricity consumption of the electric vehicle's power system and non-power systems.

[0008] The aforementioned prior art has the following drawbacks:

[0009] 1. Current CAV control strategies have limited definitions of traffic scenario information. V2X information is often defined as simple traffic light information, average vehicle speed information, or binary congestion status information, which differs greatly from the actual application scenarios of CAV.

[0010] 2. Most existing intelligent connected HEV control strategies are limited to optimal speed planning based on V2X information, without taking into account the characteristics of HEV's multiple power sources and combining EMS based on power split control to further optimize fuel consumption.

[0011] 3. Current control strategies for intelligent connected HEVs often extend from traditional HEV EMS research. For example, the existing patents mentioned in the technical background almost all employ traditional optimal control methods such as dynamic programming, model predictive control, and fuzzy strategy control, using fuel consumption as the optimization objective and vehicle state, traffic flow information, and traffic light phases as constraints. When applied to complex traffic scenarios, the complex and diverse traffic information significantly increases the computational complexity of such optimal control-based driving strategies. Therefore, these methods, including dynamic programming, fuzzy control, and model predictive control, struggle to meet real-time requirements. Summary of the Invention

[0012] The purpose of this invention is to overcome the shortcomings of the existing technology, which does not take into account the characteristics of multiple power sources in HEVs and further optimize fuel consumption by combining EMS based on power split control, and to provide an intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control method and system.

[0013] The objective of this invention can be achieved through the following technical solutions:

[0014] A hierarchical ecological driving control method for intelligent connected HEVs (Hybrid Electric Vehicles) includes:

[0015] Vehicle speed planning steps: Receive V2X information and combine it with the vehicle's status information to solve for the optimal vehicle speed trajectory;

[0016] Power split control steps: Receive the optimal vehicle speed trajectory information, follow the optimal vehicle speed trajectory information as a reference, and allocate the power output of different power sources according to the real-time power demand, with the aim of maximizing vehicle fuel economy.

[0017] Furthermore, in the vehicle speed planning step, the V2X information and the vehicle state information are fused and processed as input, and then the optimal vehicle speed trajectory is solved by reinforcement learning algorithm with multiple optimization objectives, including safety, trafficability and fuel economy.

[0018] Furthermore, the reinforcement learning algorithm is the DRL algorithm, which determines the observation state s based on the environmental input at the current control time n. n According to the observed state s n Use strategy π to select action a n According to the system, action a is executed. n Then, the next state s is obtained. n+1 Action a is calculated using the reward function. n Reward r n Let Q be the expected total cumulative reward for the action, starting from the current state.π (s n ,a n The optimal policy π that maximizes the Q-value is obtained through training. * .

[0019] Furthermore, the state s n The expression is:

[0020] s n =[V ego D ego V a V pre ,a pre D head ,i,SOC,V light_min V light_max ] T

[0021] In the formula, V ego D ego V a V pre ,a pre D head ,i,SOC,D V2S ,t r_rem ,t g_rem V light_min V light_max These represent the minimum and maximum values ​​of the following parameters: the vehicle's reference speed, the distance traveled, the vehicle's actual speed, the speed of the vehicle in front, the acceleration of the vehicle in front, the headway, the road gradient, the battery's state of charge (SOC), and the speed range within which the vehicle can pass through the traffic light with a green light window. n The observed state at time n;

[0022] The reinforcement learning algorithm employs strategy π. * According to the observed state s n Output the action a n As the vehicle's acceleration a ego (n), thus calculating the desired speed of the vehicle;

[0023] The expression for calculating the desired speed of the vehicle is as follows:

[0024] v(n+1)=v(n)+a ego (n)·Δt

[0025] In the formula, v(n+1) is the expected speed of the vehicle at time n+1, v(n) is the expected speed of the vehicle at time n, and a ego (n) represents the vehicle's acceleration at time n, and Δt represents the control step size.

[0026] Furthermore, the strategy π satisfies road speed limit constraints, traffic light rule constraints, and preceding vehicle constraints, and the objective function of the strategy π optimization process is:

[0027] J=α1fuel+α2light+α3safety

[0028] In the formula, J is the objective function, fuel is the cumulative fuel consumption, light is the cumulative red light violation penalty, safety is the cumulative collision or speeding penalty, and α1, α2 and α3 are all coefficients.

[0029] Furthermore, the reward function includes the sum of the traffic light reward function, the safety reward function, and the fuel consumption reward function, and the calculation expressions for the traffic light reward function, the safety reward function, and the fuel consumption reward function are as follows:

[0030] r light =f light (V light_min V light_max V ego )

[0031] r safety =f safety (min(V safety V limit ),V ego )

[0032]

[0033] In the formula, [V light_min V light_max [V] is the speed range that allows the vehicle to pass through the traffic light within the green light window, calculated based on the traffic light time, phase, and distance information between the vehicle and the traffic light. safety V represents the maximum safe speed at which a following vehicle can avoid a collision if the vehicle in front brakes suddenly. ego For reference speed of the vehicle, For fuel consumption rate, r light Let r be the traffic light reward function. safety For the safety reward function, r fuel Let f be the fuel consumption reward function. light (V light_min V light_max V ego ) is based on V light_min V light_max V ego Solve for the traffic light reward function, f safety (min(V safety V limit ),V ego ) is based on V safety Vlimit The minimum value and V ego Solve for the safety reward function.

[0034] Furthermore, the reward function also includes a shaping function F that guides the controlled vehicle through the traffic light. light The shaping function F that guides the controlled vehicle to follow safely. safety The shaping function F guides the controlled vehicles to improve traffic efficiency. time The expression for the reward function is:

[0035] R = r fuel +r light +r safety +F light +F safety +F time

[0036] In the formula, R is the reward function.

[0037] Furthermore, in the power split control step, the driving power demand is calculated through the driver model, so that the vehicle speed follows the optimal vehicle speed trajectory information, and the driving power demand is distributed to different power sources through EMS. The EMS is A-ECMS, and the EMS performs energy management with the minimum equivalent fuel consumption rate.

[0038] The formula for calculating the equivalent fuel consumption rate is as follows:

[0039]

[0040] In the formula, Let s be the system's total equivalent fuel consumption rate, engine fuel consumption rate, and battery equivalent fuel consumption rate at time t, respectively, and s be the equivalent factor, P demand (t) represents the required driving power at time t, and the engine power P. e (t) and battery power P b The sum of (t);

[0041] The EMS discretizes the engine power from minimum to maximum power with a certain step size to obtain an engine power sequence. Given the required drive power at any given time, it calculates the engine power P based on a preset engine fuel consumption map and motor efficiency map. e (t) Equivalent fuel consumption values ​​corresponding to all possible values and the smallest The corresponding engine power and motor power serve as the optimal power allocation strategy at that moment.

[0042] Furthermore, the EMS solution process also includes dynamically adjusting the equivalent factor s in real time using PID control based on the error between the current battery SOC and the expected SOC, so as to ensure that the SOC is maintained at an appropriate level.

[0043] The present invention also provides a control system based on the above-described intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control method, comprising:

[0044] The vehicle speed planning module is configured to execute the vehicle speed planning steps;

[0045] The power shunt control module is configured to execute the power shunt control steps.

[0046] The vehicle powertrain system is configured to receive power output commands from different power sources output by the power split control module and respond accordingly.

[0047] Compared with the prior art, the present invention has the following advantages:

[0048] (1) The intelligent connected HEV ecological driving strategy proposed in this invention provides a richer definition of V2X information in the driving environment, including information such as traffic light phase and position, speed and position of the vehicle in front, road speed limit and gradient. Therefore, the ecological driving strategy proposed in this invention is more suitable for the application scenarios of intelligent connected HEVs and has stronger practical significance.

[0049] (2) The intelligent connected HEV ecological driving strategy proposed in this invention combines a vehicle speed planning module based on V2X information and a power split control module based on HEV power system information—that is, collecting V2X information and having both fuel and battery power sources. By combining upper-level vehicle speed planning and lower-level energy management, it fully leverages the two characteristics of intelligent connected HEV, thus achieving higher fuel economy while ensuring safety and trafficability.

[0050] (3) The DRL-based vehicle speed planning algorithm proposed in this invention can use deep learning to fuse complex V2X information, thereby avoiding the curse of dimensionality. Therefore, it can still complete the optimal vehicle speed planning for multiple objectives under the condition of meeting real-time requirements in complex traffic scenarios.

[0051] (4) The data-driven DRL-based method adopted in this invention can fuse complex traffic information and perform unsupervised learning based on the multi-objective reward function designed in this invention, which takes into account safety, trafficability, and fuel economy. Compared with traditional strategies based on rules and optimal control methods, the proposed method greatly improves the processing speed of multi-dimensional information and can find the global optimal solution of multi-objective optimization problems with lower time complexity. Attached Figure Description

[0052] Figure 1 This is a schematic diagram illustrating the working principle of a hierarchical ecological driving control method for intelligent connected HEV vehicles and road collaboration provided in an embodiment of the present invention.

[0053] Figure 2 This is a schematic block diagram of a hierarchical ecological driving control method for intelligent connected HEV vehicles and road cooperation provided in an embodiment of the present invention;

[0054] Figure 3 This is a schematic block diagram of a reinforcement learning algorithm provided in an embodiment of the present invention;

[0055] Figure 4 This is a schematic block diagram of an eco-driving strategy algorithm provided in an embodiment of the present invention;

[0056] Figure 5 This is a schematic block diagram of an A-ECMS algorithm provided in an embodiment of the present invention. Detailed Implementation

[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0058] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.

[0059] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0060] It should be noted that some terms used in this invention are explained as follows:

[0061] V2X: Vehicle-to-Everything, vehicle-to-everything wireless communication technology

[0062] CAV: Connected Autonomous Vehicle

[0063] HEV: Hybrid Electric Vehicle

[0064] EMS: Energy Management Strategy

[0065] SOC: State of Charge

[0066] DRL: Deep Reinforcement Learning

[0067] MDP: Markov Decision Process

[0068] TD3: Twin-delayed Deep deterministic policy gradient.

[0069] PMP: Pontryagin's Minimum Principle

[0070] A-ECMS: Adaptive equivalent fuel consumption minimization strategy.

[0071] ECMS: Equivalent Consumption Minimization Strategy

[0072] PID stands for Proportional-Integral-Derivative, meaning proportional-integral-derivative control.

[0073] Example 1

[0074] Addressing the current development needs for intelligent connectivity and low carbon emissions in vehicles, this invention aims to design a multi-objective ecological driving control method suitable for intelligent connected HEVs. CAVs can obtain rich traffic scene information through V2X communication, including road information, traffic light phase information, and time information. Ecological driving technology can then construct a multi-objective optimal control problem that balances safety, traffic flow, comfort, and economy based on this external information and the current internal vehicle state. An optimization algorithm is used to solve for the optimal vehicle speed in real time, which serves as a reference speed during driving. The vehicle is then controlled to follow this reference speed while distributing power demand among different power sources to achieve higher fuel economy. The working principle of this intelligent connected HEV ecological driving strategy is as follows: Figure 1 As shown.

[0075] This embodiment provides a hierarchical ecological driving control method for intelligent connected HEVs with vehicle-road cooperation, including:

[0076] Vehicle speed planning steps: Receive V2X information and combine it with the vehicle's status information to solve for the optimal vehicle speed trajectory;

[0077] Power split control steps: Receive optimal vehicle speed trajectory information, follow the optimal vehicle speed trajectory information as a reference, and allocate the power output of different power sources according to real-time power demand, with the aim of maximizing vehicle fuel economy.

[0078] In the vehicle speed planning step, V2X information and vehicle status information are fused and processed as input. Then, the optimal vehicle speed trajectory is solved through reinforcement learning algorithm with multiple optimization objectives, including safety, traffic flow and fuel economy.

[0079] The reinforcement learning algorithm used can be the TD3 algorithm, deep Q network, deep deterministic policy gradient, or other DRL algorithms.

[0080] This embodiment also provides a corresponding system embodiment of the above method, specifically including:

[0081] A control system based on the above-described intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control method includes:

[0082] The vehicle speed planning module is configured to execute vehicle speed planning steps;

[0083] The power shunt control module is configured to execute power shunt control steps.

[0084] The vehicle powertrain is configured to receive power output commands from different power sources from the power split control module and respond accordingly.

[0085] This solution mainly consists of a vehicle speed planning module and a power split control module. V2X information is obtained from the vehicle communication system of the intelligent connected HEV. Vehicle speed planning calculates the optimal speed trajectory based on the V2X information and the vehicle's status information. The power split control module uses this optimal speed as a reference and follows it, allocating power output from different power sources according to real-time power demand, with the aim of maximizing vehicle fuel economy. The vehicle powertrain system receives control commands from the power split control module and responds accordingly. The technical block diagram of this system is shown below. Figure 2 As shown.

[0086] The focus of this invention is Figure 2 The ecological driving strategy system shown in the diagram will be explained in detail in the following sections. The main technical solutions are divided into the development of the vehicle speed planning module and the power split control module.

[0087] As a preferred implementation, the reinforcement learning algorithm is the DRL algorithm, which determines the observation state s based on the environmental input. n According to the observed state s n Use strategy π to select action a n According to the system, action a is executed. n The next state s n+1 Action a is calculated using the reward function. n Reward r n Let Q be the expected total accumulated reward from the current state in the future. π (s n ,a n The optimal policy π that maximizes the Q-value is obtained through training. * .

[0088] In other words, the development process of the vehicle speed planning module includes:

[0089] The vehicle speed planning module is mainly developed based on the DRL algorithm. This module integrates V2X information (including traffic light phase and position, speed and position of the vehicle in front, road gradient and speed limit, etc.) with the vehicle's state information (including SOC, instantaneous fuel consumption, actual vehicle speed, etc.) and learns the optimal speed that satisfies trafficability, safety and fuel economy through a designed multi-objective reward function and shaping function.

[0090] 1.1 Overview of DRL Algorithm Background

[0091] In reinforcement learning, there are two interacting entities: the agent and the environment. The agent perceives the state *s* of the environment and selects an appropriate action *a* based on its learned policy *π* to maximize long-term gains. Reinforcement learning algorithms simplify the real-world environment into a Multiplicative Probabilistic (MDP) problem, where changes in the environment's state depend only on the current state *s* and the agent's action *a*. The agent's actions are determined by the policy *π*, which is the probability distribution of actions within the current state.

[0092] In reinforcement learning, the environment calculates the agent's reward at each step based on the agent's actions using a reward function. The action value function (or Q-value) calculates the agent's expected total reward from the current state. The action value function allows the agent to choose strategies that are more advantageous in the long run. A flowchart of the reinforcement learning algorithm is shown below. Figure 3 As shown.

[0093] The goal of reinforcement learning algorithms is to find the optimal policy π that maximizes the Q-value. *Based on this, this embodiment selects the state-of-the-art DRL algorithm TD3 to solve this eco-driving strategy problem.

[0094] 1.2 Intelligent Connected HEV Ecosystem Driving Strategy Based on DRL

[0095] The traffic information obtainable through V2X communication of intelligent vehicles is extremely complex. The input information of the intelligent agent should be simplified as much as possible while fully reflecting the current traffic state of the controlled vehicle, so as to reduce the learning difficulty of the DRL algorithm. Therefore, the state variables selected in this embodiment are as follows:

[0096] s n =[V ego D ego V a V pre ,a pre D head ,i,SOC,V light_min V light_max ] T (1)

[0097] In the formula, V ego D ego V a V pre ,a pre D head ,i,SOC,D V2S ,t r_rem ,t g_rem V light_min V light_max These represent the minimum and maximum values ​​of the following parameters: the vehicle's reference speed, the distance traveled, the vehicle's actual speed, the speed of the vehicle in front, the acceleration of the vehicle in front, the headway, the road gradient, the battery's state of charge (SOC), and the speed range within which the vehicle can pass through the traffic light with a green light window. t The state at time n; the output of the TD3 agent is the vehicle's acceleration a. ego The desired speed of the vehicle can be expressed as v(n+1) = v(n) + a ego (n)·Δt;

[0098] In the formula, v(n+1) is the expected speed of the vehicle at time n+1, v(n) is the expected speed of the vehicle at time n, and a ego (n) represents the vehicle's acceleration at time n, and Δt represents the control step size.

[0099] Efficient eco-driving control strategies can unlock the energy-saving potential of intelligent connected HEVs by optimizing vehicle speed. Simultaneously, the control strategy must also meet constraints such as road speed limits, traffic light rules, and the presence of the vehicle ahead. Therefore, the objective function can be expressed as:

[0100]

[0101] In the formula, J is the objective function, fuel is the cumulative fuel consumption, light is the cumulative red light violation penalty, safety is the cumulative collision or speeding penalty, and α1, α2 and α3 are all coefficients.

[0102] As a preferred implementation method,

[0103] The traffic light reward function and safety reward function, which are directly derived from the objective function (2), are too sparse and delayed. Feedback is only generated when a vehicle runs a red light or a collision occurs. This makes it difficult for the agent to understand the possible impact of the current action on the future state, which is not conducive to the convergence of the training process. Therefore, considering the characteristics of the DRL algorithm, this implementation redesigns the traffic light reward function and safety reward function.

[0104] For the traffic light reward function, this implementation calculates the speed range [V] that allows the vehicle to pass through the traffic light during the green light window based on the traffic light time, phase, and distance information between the controlled vehicle and the traffic light. light_min V light_max When the planned vehicle speed is not within the window speed range, the reward function (3) is calculated based on the difference between the planned vehicle speed and the range. For the safety reward function, this implementation calculates the maximum safe speed V that the following vehicle will not collide with under emergency braking conditions of the preceding vehicle based on the Krauss following model. safety The reward function is calculated based on the safe speed, road speed limit, and planned vehicle speed (4). The fuel consumption reward function does not need to be redesigned, that is, it is the opposite of the fuel consumption rate.

[0105] r light =f light (V light_min V light_max V ego (3)

[0106] r safety =f safety (min(V safety V limit ),V ego (4)

[0107]

[0108] In the formula, [V light_min V light_max [V] is the speed range that allows the vehicle to pass through the traffic light within the green light window, calculated based on the traffic light time, phase, and distance information between the vehicle and the traffic light. safety V represents the maximum safe speed at which a following vehicle can avoid a collision if the vehicle in front brakes suddenly.ego For reference speed of the vehicle, For fuel consumption rate, r light Let r be the traffic light reward function. safety For the safety reward function, r fuel Let f be the fuel consumption reward function. light (V light_min V light_max V ego ) is based on V light_min V light_max V ego Solve for the traffic light reward function, f safety (min(V safety V limit ),V ego ) is based on V safety V limit The minimum value and V ego Solve for the safety reward function.

[0109] Furthermore, as a preferred implementation, the basic reward functions (3) and (4) are insufficient to encourage the TD3 agent to learn safe and efficient ecological driving behavior. Therefore, this implementation also designs a shaping function F to guide the controlled vehicle through traffic lights. light The shaping function F that guides the controlled vehicle to follow safely. safety The shaping function F guides the controlled vehicles to improve traffic efficiency. time During training, prior knowledge is provided to the TD3 agent to compensate for the deficiencies of the basic reward functions (3) and (4). To ensure that the added shaping function does not alter the final learning goal of the TD3 agent, this embodiment designs the aforementioned shaping reward function based on the potential energy-based reward shaping theory, which guarantees policy invariance. Therefore, our designed total reward function is:

[0110] R = r fuel +r light +r safety +F light +F safety +F time (6)

[0111] In the formula, R is the total reward function.

[0112] This total reward function provides the agent with fused V2X traffic information and vehicle state information, and promotes the reinforcement learning algorithm to learn the correct policy more efficiently. The flowchart of the DRL-based eco-driving policy algorithm is shown below. Figure 4As shown. The reinforcement learning environment here includes both an intelligent connected traffic environment and a power shunt control module. The intelligent connected traffic environment provides V2X information, while the power shunt control module is responsible for outputting vehicle powertrain control commands based on the optimal vehicle speed and providing the reinforcement learning agent with vehicle status information such as battery SOC, instantaneous fuel consumption, and vehicle speed.

[0113] After completing the theoretical design of the algorithm, this embodiment uses the open-source machine learning library PyTorch to implement the proposed reference speed planning module based on TD3. This speed planning module can plan the optimal speed in real time based on multiple optimization objectives (safety, traffic flow, fuel economy).

[0114] As a preferred implementation, in the power split control step, the driving power demand is calculated through the driver model, so that the vehicle speed follows the optimal vehicle speed trajectory information, and the driving power demand is distributed to different power sources through the EMS.

[0115] The EMS in the power shunt control step can be ECMS, model predictive control, dynamic programming, or other EMS types.

[0116] Essentially, the power split control module consists of a driver model and an EMS (Electronic Power Management System). Given a reference speed provided by the vehicle speed planning module, the driver model determines the appropriate drive power demand, ensuring the vehicle speed follows the reference speed trajectory. The EMS then distributes this drive power demand to different power sources, optimizing the engine operating range and maintaining battery SOC (State of Charge). Finally, it outputs control commands (including throttle, brake, engine power, and motor power) corresponding to the power distribution strategy to the intelligent connected HEV powertrain system. This embodiment uses a PID-based driver model for speed following and employs an Adaptive Equivalent Fuel Minimum Strategy (A-ECMS), an improvement on ECMS, based on PMP (Power Management System) as the EMS. Since PID control is widely used in various fields, this section will focus on the A-ECMS algorithm.

[0117] 2.1 Overview of ECMS Algorithm

[0118] ECMS is an optimization method for HEV energy management. It converts the electrical energy consumed by the motor into equivalent fuel consumption using an equivalent factor, aiming to minimize instantaneous equivalent fuel consumption to obtain a locally optimal control variable (power allocation). Based on this idea of ​​ECMS, the optimization objective function of the system at each moment is the equivalent fuel consumption value:

[0119]

[0120] in, These are the system's total equivalent fuel consumption rate, engine fuel consumption rate, and battery equivalent fuel consumption rate, respectively. The battery equivalent fuel consumption rate is determined by the battery's electrical power P. b (t) and equivalent factor s are calculated. Due to the engine power P... e (t) and battery power P b The sum of (t) is the demand-driven power P. demand (t), therefore It can be expressed as the following formula:

[0121]

[0122] In the formula, Let P be the total equivalent fuel consumption rate of the system and the engine fuel consumption rate at time t, respectively, and s be the equivalent factor. demand (t) represents the required driving power at time t, and the engine power P. e (t) and battery power P b The sum of (t);

[0123] The goal of the ECMS algorithm is to determine the variable P in equation (8). e (t) is used to minimize the optimization objective at each time step. The value of P. In specific execution, the ECMS algorithm discretizes the engine power from minimum to maximum power with a certain step size to obtain the engine power sequence, which is P. e (t) Possible values. Given the required driving power at any given time, the algorithm will calculate P based on the engine specific fuel consumption Map and the motor efficiency Map. e (t) all possible values ​​corresponding to and the smallest The corresponding engine power and motor power (calculated from the vehicle's required driving power) serve as the optimal power allocation strategy at that moment.

[0124] 2.2 Development of a Power Shunt Control Module Based on A-ECMS

[0125] The equivalent factor s in the ECMS algorithm is crucial to the algorithm's performance. However, the s that enables the algorithm to find the global optimal solution can only be obtained under known operating conditions. Therefore, the ECMS algorithm has difficulty in ensuring optimal fuel economy while meeting real-time requirements.

[0126] As a preferred implementation, this embodiment employs another method for determining the equivalent factor, namely the A-ECMS algorithm. This algorithm dynamically adjusts the equivalent factor s in real time using PID control based on the error between the current battery SOC and the desired SOC, ensuring that the SOC is maintained at an appropriate level.

[0127] Next, we will provide an overview of the specific working process of the power split control module based on the A-ECMS algorithm. First, the driver model uses a PID control algorithm to calculate the driving power required to follow the reference vehicle speed trajectory at the current moment. Then, the A-ECMS algorithm consults the engine specific fuel consumption map and the motor efficiency map to calculate the fuel consumption corresponding to each control strategy, selecting the optimal power split strategy. Finally, the power split control module outputs a power distribution control command to the vehicle powertrain system. Based on this command, the vehicle powertrain system determines the operating states of the battery, motor, and engine to achieve tracking of the reference vehicle speed. The A-ECMS algorithm block diagram is shown below. Figure 5 As shown.

[0128] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A hierarchical ecological driving control method for intelligent connected HEVs (Hybrid Electric Vehicles), characterized in that, include: Vehicle speed planning steps: Receive V2X information and combine it with the vehicle's status information to solve for the optimal vehicle speed trajectory; Power split control steps: Receive the optimal vehicle speed trajectory information, follow the optimal vehicle speed trajectory information as a reference, and allocate the power output of different power sources according to the real-time power demand in order to maximize the vehicle's fuel economy. In the vehicle speed planning step, the V2X information and the vehicle status information are fused and processed as input. Then, the optimal vehicle speed trajectory is solved by reinforcement learning algorithm with multiple optimization objectives, including safety, traffic flow and fuel economy. The reinforcement learning algorithm is the DRL algorithm, which determines the observation state s based on the environmental input at the current control time n. n According to the observed state s n Use strategy π to select action a n According to the system, action a is executed. n Then, the next state s is obtained. n+1 Action a is calculated using the reward function. n Reward r n Let Q be the expected total cumulative reward for the action, starting from the current state. π (s n ,a n The optimal policy π that maximizes the Q-value is obtained through training. * ; The reward function includes the sum of the traffic light reward function, the safety reward function, and the fuel consumption reward function. The calculation expressions for the traffic light reward function, the safety reward function, and the fuel consumption reward function are as follows: r light =f light (V light_min ,V light_max ,V ego ) r safety =f safety (min(V safety ,V limit ),V ego ) In the formula, [V light_min V light_max [V] is the speed range that allows the vehicle to pass through the traffic light within the green light window, calculated based on the traffic light time, phase, and distance information between the vehicle and the traffic light. safety V represents the maximum safe speed at which a following vehicle can avoid a collision if the vehicle in front brakes suddenly. ego For reference speed of the vehicle, For fuel consumption rate, r light Let r be the traffic light reward function. safety For the safety reward function, r fuel Let f be the fuel consumption reward function. light (V light_min V light_max V ego ) is based on V light_min V light_max V ego Solve for the traffic light reward function, f safety (min(V safety V limit ),V ego ) is based on V safety V limit The minimum value and V ego Solve for the safety reward function; The reward function also includes a shaping function F that guides the controlled vehicle through the traffic light. light The shaping function F that guides the controlled vehicle to follow safely. safety The shaping function F guides the controlled vehicles to improve traffic efficiency. time The expression for the reward function is: R=r fuel +r light +r safety +F light +F safety +F time In the formula, R is the reward function; In the power split control step, the driving power demand is calculated by the driver model, so that the vehicle speed follows the optimal vehicle speed trajectory information. The driving power demand is distributed to different power sources by the EMS, which is an A-ECMS. The EMS manages energy with the minimum equivalent fuel consumption rate. The formula for calculating the equivalent fuel consumption rate is as follows: In the formula, Let s be the system's total equivalent fuel consumption rate, engine fuel consumption rate, and battery equivalent fuel consumption rate at time t, respectively, and s be the equivalent factor, P demand (t) represents the required driving power at time t, and the engine power P. e (t) and battery power P b The sum of (t); The EMS discretizes the engine power from minimum to maximum power with a preset step size to obtain an engine power sequence. Given the required drive power at any given time, it calculates the engine power P based on preset engine fuel consumption Map and motor efficiency Map. e (t) Equivalent fuel consumption values ​​corresponding to all possible values and the smallest The corresponding engine power and motor power serve as the optimal power allocation strategy at that moment; The EMS solution process also includes dynamically adjusting the equivalent factor s in real time using PID control based on the error between the current battery SOC and the expected SOC, so as to ensure that the SOC is maintained at an appropriate level.

2. The intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control method according to claim 1, characterized in that, The state s n The expression is: s n =[V ego ,D ego ,V a ,V pre ,a pre ,D head ,i,SOC,V light_min ,V light_max ] T In the formula, V ego D ego V a V pre ,a pre D head ,i,SOC,D V2S ,t r_rem ,t g_rem V light_min V light_max These represent the minimum and maximum values ​​of the following parameters: the vehicle's reference speed, the distance traveled, the vehicle's actual speed, the speed of the vehicle in front, the acceleration of the vehicle in front, the headway, the road gradient, the battery's state of charge (SOC), and the speed range within which the vehicle can pass through the traffic light with a green light window. n The observed state at time n; The reinforcement learning algorithm employs strategy π. * According to the observed state s n Output the action a n As the vehicle's acceleration a ego (n), thus calculating the desired speed of the vehicle; The expression for calculating the desired speed of the vehicle is as follows: v(n+1)=v(n)+a ego (n)·△t In the formula, v(n+1) is the expected speed of the vehicle at time n+1, v(n) is the expected speed of the vehicle at time n, and a ego (n) represents the vehicle's acceleration at time n, and Δt represents the control step size.

3. The intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control method according to claim 1, characterized in that, The strategy π satisfies road speed limit constraints, traffic light rule constraints, and preceding vehicle constraints. The objective function for the optimization process of strategy π is: J=α1fuel+α2light+α3safety In the formula, J is the objective function, fuel is the cumulative fuel consumption, light is the cumulative red light violation penalty, safety is the cumulative collision or speeding penalty, and α1, α2 and α3 are all coefficients.

4. A control system based on the intelligent connected HEV vehicle-road cooperative hierarchical ecological driving control method as described in any one of claims 1-3, characterized in that, include: The vehicle speed planning module is configured to execute the vehicle speed planning steps; The power shunt control module is configured to execute the power shunt control steps. The vehicle powertrain system is configured to receive power output commands from different power sources output by the power split control module and respond accordingly.

Citation Information

Patent Citations

  • Intelligent networked vehicle layered speed planning method based on continuous signal light information

    CN110264757A

  • Robust energy management method and system for intelligent networked hybrid electric vehicle

    CN112498334A

  • Vehicle energy management method based on intelligent network connection information

    CN112526883A

  • Dynamically constrained intelligent fuel cell vehicle energy management strategy

    CN113561793A

  • HEV energy management hierarchical control method considering lane changing behaviors in network connection environment

    CN111959492A