Ecological driving-oriented networked motorcade intelligent control method

By constructing a deep reinforcement learning simulation environment and a two-layer strategy ecological driving framework, the problem of insufficient energy consumption models in existing fleet control methods is solved, a comprehensive balance between energy efficiency, driving comfort and safety is achieved, and the ecological driving performance and operational sustainability of the connected fleet are improved.

CN120673603APending Publication Date: 2025-09-19SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511019360.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing fleet control methods do not fully consider the microscopic energy consumption model when defining the eco-driving speed, ignore the energy consumption changes under different driving conditions and the degradation performance of the energy source, resulting in the failure to fully balance performance such as energy efficiency, driving comfort and safety in multi-objective optimization.

Method used

A deep reinforcement learning simulation environment was constructed, including fleet spacing and vehicle dynamics models, a composite energy source model, and a traffic light speed planning model. Based on a two-layer policy ecological driving framework, a double-delay deep deterministic policy gradient algorithm (TD3) was used for training. The trained two-layer policy ecological driving framework was obtained for intelligent control of connected fleets.

Benefits of technology

By introducing the degradation characteristics of fuel cells and batteries and the microscopic energy consumption model, the accuracy and effectiveness of the control strategy are improved, the practical application value of the control strategy for fleet operation and the energy efficiency control accuracy are realized, and the system's perception and adaptability to the health status of energy sources are enhanced to adapt to complex and changing traffic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673603A_ABST
    Figure CN120673603A_ABST
Patent Text Reader

Abstract

The invention discloses a networked motorcade intelligent control method for ecological driving, and belongs to the field of networked motorcade control, and the method comprises the steps: building a double-layer strategy ecological driving framework; the upper layer strategy takes passing efficiency and stability as indexes to obtain a queue dynamic management strategy; according to the lower-layer strategy, the driving comfort, the energy-saving performance and the energy source durability of the motorcade serve as performance indexes, and the expected ecological driving speed and the power distribution scheme of the energy source are optimized; and a control scheme of the to-be-processed motorcade is obtained by using the trained double-layer strategy ecological driving framework, and intelligent control of the networked motorcade is completed. The problems that an existing motorcade control method focuses on traffic efficiency, a microscopic energy consumption model is not fully considered when the ecological driving speed is defined, and energy consumption changes and energy source degradation performance under different driving conditions are ignored, so that performance such as energy efficiency, driving comfort and safety cannot be comprehensively balanced in multi-objective optimization are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of connected vehicle fleet control, and in particular relates to an intelligent control method for a connected vehicle fleet oriented to ecological driving. Background Art

[0002] In recent years, rapid urbanization has exacerbated traffic congestion, accidents, and high energy consumption. Connected fleet eco-driving technology offers a potential solution. Through information sharing and coordinated control, fleet members can not only make real-time decisions based on their own status, but also share traffic information with the preceding and following vehicles, as well as with surrounding vehicles, thereby fully optimizing traffic system efficiency. Fuel cell hybrid vehicles, with their zero-pollution, high-efficiency, and low-noise advantages, are considered an ideal platform for connected fleet eco-driving technology.

[0003] Eco-driving technology enables fleets to travel uninterrupted through traffic lights, effectively reducing energy losses from stopping and starting. Considering the challenges posed by the coordinated optimization of the hybrid energy source system of fuel cell hybrid vehicles, particularly the impact of the degradation effects of fuel cells and batteries on the overall energy efficiency of the fleet, eco-driving becomes a complex problem with multi-objective optimization characteristics that requires dynamic adjustment during fleet operation. However, existing fleet control methods focus on traffic efficiency and fail to fully consider microscopic energy consumption models when defining eco-driving speeds. They also ignore energy consumption variations under different driving conditions and the degradation performance of energy sources. This results in a failure to fully balance multiple factors such as energy efficiency, driving comfort, and safety in multi-objective optimization. Summary of the Invention

[0004] To address the above-mentioned shortcomings in the existing technology, the present invention provides an intelligent control method for a connected fleet for eco-driving. This method solves the problem that existing fleet control methods focus on traffic efficiency, fail to fully consider microscopic energy consumption models when defining eco-driving speeds, and ignore energy consumption changes under different driving conditions and energy source degradation performance. As a result, the method fails to fully balance performance issues such as energy efficiency, driving comfort, and safety in multi-objective optimization.

[0005] In order to achieve the above-mentioned purpose, the technical solution adopted by the present invention is: an intelligent control method for a connected fleet for ecological driving, comprising: Constructing a deep reinforcement learning simulation environment, which includes a fleet spacing and vehicle dynamics model, a composite energy source model, a vehicle demand power model, and a traffic light speed planning model; Based on the deep reinforcement learning simulation environment, a two-layer strategy ecological driving framework is constructed; The two-layer policy ecological driving framework is trained based on the double-delay deep deterministic policy gradient algorithm to obtain the trained two-layer policy ecological driving framework; The trained two-layer strategy ecological driving framework is used to obtain the control plan for the fleet to be processed, completing the intelligent control of the connected fleet.

[0006] The beneficial effects of the present invention are as follows: by constructing a two-layer strategy framework, the upper-layer strategy is used to optimize traffic efficiency and stability, and the lower-layer strategy is used for energy management of the composite energy source system, thereby collaboratively achieving multi-objective optimization. By introducing the degradation characteristics of fuel cells and batteries and microscopic energy consumption models into the ecological driving control process, the optimization deviation caused by insufficient energy consumption modeling in existing methods is effectively avoided, thereby improving the practical application value of the control strategy and the accuracy of energy efficiency control. Finally, by designing a dual-delay deep deterministic policy gradient algorithm (TD3) to train the control strategy, the control system has the ability to learn and adaptively adjust online, can adapt to complex and changing traffic environments, and realize intelligent management and control of the fleet's operating status.

[0007] Furthermore, the fleet spacing and vehicle dynamics model is:

[0008] in, is the total length of the convoy; is the total number of team members; is the distance between vehicles; The vehicle captain; is the current speed of the fleet; is the reaction time; is the current acceleration of the team; is the mass of the vehicle; the traction provided to the vehicle; is the air resistance; is the rolling resistance; is the slope resistance; is the air density; is the air resistance coefficient; is the vehicle's front windshield area; is the rolling resistance coefficient; is the acceleration due to gravity; is the road slope; is a sine function; is the cosine function.

[0009] The beneficial effect of the above further solution is that by introducing an accurate speed estimation model that takes into account vehicle spacing, dynamic response and vehicle force factors, it provides theoretical support and modeling basis for achieving more refined ecological driving control.

[0010] Furthermore, the composite energy source model includes a hydrogen consumption sub-model, a fuel cell voltage degradation sub-model, and a lithium battery SOC sub-model considering capacity degradation:

[0011]

[0012]

[0013] in, is the actual hydrogen consumption of the fuel cell; is the current running time; is the output power of the fuel cell; for the efficiency of the fuel cell system; is the lower calorific value of hydrogen; is the output voltage of the fuel cell; is the power of the auxiliary system; is the fuel cell voltage degradation rate; is the correction factor for real driving; is the degradation rate of the fuel cell under start-stop operation; is the degradation rate of the fuel cell under low-power operation, where the low-power operation is an operation with a power lower than a power threshold; is the degradation rate of the fuel cell under high-power operation, where the high-power operation is an operation with a power not less than a power threshold; is the degradation rate of the fuel cell during power fluctuations; is the average number of starts and stops of the fuel cell; is the operating time of the fuel cell in low power state; is the operating time of the fuel cell in high power state; is the power change of the fuel cell; is the SOC value of the lithium battery at the current moment; is the initial SOC of the lithium battery; The charging and discharging efficiency of lithium batteries; is the lithium battery current; is the lithium battery capacity; is the nominal capacity of the lithium battery; is the degradation degree of lithium battery capacity; is the number of charge and discharge cycles.

[0014] The beneficial effects of the above further scheme are: providing a modeling basis for accurate evaluation of fuel cell hydrogen consumption; providing a theoretical basis and modeling basis for enhancing the control strategy's perception and adaptability to the health status of the fuel cell system; providing a theoretical basis and modeling basis for enhancing the control strategy's perception and adaptability to the health status of the lithium battery system.

[0015] Furthermore, the vehicle demand power model is:

[0016] in, The vehicle requires power; the traction provided to the vehicle; is the current speed of the fleet; is the motor efficiency; is the efficiency of the unidirectional DC / DC converter; is the efficiency of the bidirectional DC / DC converter; is the output power of the fuel cell; is the power of lithium battery.

[0017] The beneficial effect of the above further solution is to provide a modeling basis for accurately calculating the required power of the vehicle.

[0018] Furthermore, the traffic light speed planning model is:

[0019] in, The minimum speed at which a convoy can pass through an intersection without stopping under a green light; is the current speed of the fleet; The maximum speed at which a convoy can pass through an intersection without stopping under a green light; The maximum speed required by the current road; The signal light status; for the green light; For the team The distance between the vehicle and the intersection; The remaining time of the current signal light; for red lights; The duration of the green light.

[0020] The beneficial effect of the above further solution is: by introducing a green light speed optimization consultation model and dynamically adjusting the upper and lower speed limits in combination with the status of traffic lights, the team members can be guided to pass through the intersection reasonably.

[0021] Furthermore, the dual-strategy eco-driving framework includes an upper-strategy for obtaining a fleet dynamic management strategy and a lower-strategy for obtaining a desired eco-driving speed and power allocation scheme for the fleet.

[0022] The beneficial effects of the above-mentioned further scheme are as follows: This scheme improves the efficiency and safety of the fleet at intersections by constructing a state space that integrates information such as traffic signals, vehicle distances, and speeds, and combines diverse queue reorganization actions with refined reward functions. At the same time, it fully considers queue stability and speed adaptability to achieve a more efficient and coordinated fleet ecological driving strategy under signal light control. In the lower-level control, an energy efficiency evaluation function based on the equivalent consumption minimum strategy is constructed. By introducing an energy source degradation model, the control system's perception and adaptability to the health status of the energy source are enhanced. The energy source power boundary, battery SOC boundary, vehicle speed and acceleration boundary are integrated to construct a penalty term based on the Lagrangian function, thereby significantly improving the fleet's ecological driving performance and operational sustainability.

[0023] Furthermore, the action space, state space and reward function of the upper-level strategy are:

[0024]

[0025]

[0026] in, is the state space of the upper-level strategy L1; The signal light status; For the team The distance between the vehicle and the intersection; is the current speed of the fleet; The remaining time of the current signal light; How long the green light lasts; is the action space of the upper-level strategy L1; For the team Separation of vehicles; is the total number of convoy members; pass means the convoy passes through the intersection as a whole without separation; stop means the convoy stops before the intersection as a whole; is the reward function of the upper strategy L1; is the weight coefficient of the traffic efficiency item; is the weight coefficient of the fleet stability term; a reward for safe passage for the convoy; To encourage overall traffic rewards, Indicates separation, It means not separated; Reward for traffic efficiency; A reward for fleet stability; The maximum speed at which a convoy can pass through an intersection without stopping under a green light; The minimum speed at which a convoy can pass through an intersection without stopping under a green light; The maximum speed required by the current road; The minimum speed required for the current road; It is the comfortable acceleration of the controlled vehicles in the convoy.

[0027] The beneficial effects of this further solution are as follows: By constructing a state space that integrates information such as traffic signals, vehicle distances, and speed, and combining it with diverse platoon reorganization actions and a refined reward function, this solution improves platoon efficiency and safety at intersections. Furthermore, by fully considering platoon stability and speed adaptability, it achieves a more efficient and coordinated platoon ecological driving strategy under signal control.

[0028] Furthermore, the state space, action space and reward function of the lower-level strategy are:

[0029]

[0030]

[0031]

[0032] in, is the state space of the lower-level strategy L2; For the team The distance between the vehicle and the intersection; is the current speed of the fleet; is the SOC value of the lithium battery at the current moment; is the output power of the fuel cell; is the road slope; is the rolling resistance coefficient; is the action space of the lower-level strategy L2; is the acceleration of the vehicle; The maximum deceleration allowed for the fleet; The maximum acceleration allowed by the team; is the power change of the fuel cell; is the reward function of the lower layer strategy L2; Rewards for fuel consumption; A reward for driving comfort; For the Lagrange multipliers for term constraints; For the Item constraints; is the number of items; is the speed constraint; is the fuel cell power constraint; is the acceleration constraint term; is the lithium battery power constraint; is the SOC constraint item; The maximum speed required by the current road; is the fuel cell nominal power; is the current acceleration of the team; The maximum charging power of the lithium battery; is the maximum discharge power of the lithium battery; is the power of the lithium battery; is the actual hydrogen consumption of the fuel cell; is the adaptive equivalent coefficient; is the weight of energy performance degradation; is the virtual hydrogen consumption of the lithium battery; is the degradation degree of lithium battery capacity; is the fuel cell voltage degradation rate; is the discharge efficiency of the lithium battery; is the average efficiency of the fuel cell system; is the air density; The charging efficiency of lithium batteries; is the impact, whose value is the acceleration with sampling time The amount of change; is the weight coefficient used to adjust the acceptance of the shock amplitude; Adaptive equivalent coefficient A constant determined by the upper limit of ; is the weight coefficient of lithium battery SOC; is the reference SOC of lithium battery.

[0033] The beneficial effects of this further solution are as follows: This solution constructs an energy efficiency evaluation function based on an equivalent consumption minimization strategy within the lower-level control layer. By introducing an energy source degradation model, the control system's ability to perceive and adapt to the energy source's health status is enhanced. Furthermore, by integrating energy source power boundaries, battery SOC boundaries, and vehicle speed and acceleration boundaries, a penalty term based on a Lagrangian function is constructed, significantly improving the fleet's eco-driving performance and operational sustainability.

[0034] Furthermore, the update process of the critic network and the orator network in the upper-level strategy is:

[0035]

[0036]

[0037]

[0038] in, For the upper-level strategy The critic estimates the network loss function and the critic estimates the network parameters gradient; For the upper-level strategy A critic estimates the parameters of the network; is the capacity of the small batch sample; Index for mini-batch samples; is the target Q value of the upper-level strategy; For the upper-level strategy A critic estimates the Q value of the network for the current state-action pair; for Estimating network parameters for critics gradient; Estimating network numbers for critics; is the current state of the upper-level policy; is the current action of the upper-level strategy; is the reward value of the upper-level strategy; is the discount on future rewards of the upper-level strategy; Estimate the Q value of the next state-action pair for the first critic target network of the upper strategy; The Q value of the second critic target network of the upper strategy for the next state-action pair is estimated; is the state of the upper-level strategy at the next moment; The parameters of the orator target network of the upper-level strategy; are the parameters of the first critic target network of the upper strategy; are the parameters of the second critic target network of the upper strategy; Estimate the network loss function for the speaker of the upper strategy and estimate the network parameters for the speaker gradient; Estimate network parameters for the upper-level strategy orator; Estimate the network action output parameters for the upper-level policy speaker gradient; Estimate the actions output by the network for the upper-level policy speaker; Estimate the Q-value of the network output for the first critic of the upper strategy to the action gradient; is the soft update rate of the upper layer strategy; For the upper-level strategy Parameters of the critic-target network; The first Parameters of the critic-target network; The parameters of the orator target network after the soft update of the upper layer strategy; Prioritize experience samples is the total capacity of the experience pool; E is the number of all samples; Normalize the priorities of all samples; To adjust the priority hyperparameter; For the The sampling value of an experience sample; is the TD error corresponding to the sample; is a constant.

[0039] The beneficial effects of this further approach are as follows: By constructing a two-layer policy network update mechanism based on TD error, introducing gradient update rules for the target and evaluation networks, and combining prioritized experience replay with a weighted sampling strategy, this approach improves policy convergence speed and learning stability. This mechanism effectively focuses on key state-action pairs, enhancing the policy's ability to learn rare but critical examples in complex driving conditions, significantly improving the convergence efficiency and generalization of the eco-driving policy.

[0040] Furthermore, the updating process of the critic network and the orator network in the lower-level strategy is as follows:

[0041]

[0042]

[0043] in, The first The critic estimates the network loss function and the critic estimates the network parameters gradient; The first A critic estimates the parameters of the network; is the capacity of the small batch sample; Index for mini-batch samples; is the target Q value of the lower-level strategy; The first A critic estimates the Q value of the network for the current state-action pair; for Estimating network parameters for critics gradient; Estimating network numbers for critics; is the current state of the underlying policy; is the current action of the lower-level strategy; is the reward value of the lower-level strategy; is the discount on future rewards of the underlying strategy; Estimate the Q value of the next state-action pair for the first critic target network of the lower layer strategy; Estimate the Q value of the next state-action pair for the second critic target network of the lower layer strategy; is the state of the lower-level strategy at the next moment; Parameters of the orator target network for the lower-level strategy; are the parameters of the first critic target network of the lower layer strategy; are the parameters of the second critic target network of the lower layer strategy; Estimate the network loss function for the speaker of the lower policy and estimate the network parameters for the speaker gradient; Estimate network parameters for the lower-level strategic orator; Estimate network action output parameters for the underlying policy speaker gradient; Estimate the actions output by the network for the underlying policy speaker; Estimate the Q-value of the network output for the first critic of the lower layer strategy to the action gradient; is the soft update rate of the lower layer strategy; The first Parameters of the critic-target network; The first value after the soft update of the lower layer strategy Parameters of the critic-target network; Parameters of the orator target network after soft update of the underlying policy.

[0044] The beneficial effects of this further approach are as follows: By constructing a two-layer policy network update mechanism based on TD error, introducing gradient update rules for the target and evaluation networks, and combining prioritized experience replay with a weighted sampling strategy, this approach improves policy convergence speed and learning stability. This mechanism effectively focuses on key state-action pairs, enhancing the policy's ability to learn rare but critical examples in complex driving conditions, significantly improving the convergence efficiency and generalization of the eco-driving policy. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Flow chart of the method of the present invention.

[0046] Figure 2 This is a topological diagram of the power system of a fuel cell hybrid vehicle in an embodiment of the present invention.

[0047] Figure 3 This is a schematic diagram of the ecological driving control method for a connected fleet based on deep reinforcement learning in an embodiment of the present invention. DETAILED DESCRIPTION

[0048] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0049] like Figure 1 As shown, in one embodiment of the present invention, a method for intelligent control of a connected fleet for eco-driving includes: Constructing a deep reinforcement learning simulation environment, which includes a fleet spacing and vehicle dynamics model, a composite energy source model, a vehicle demand power model, and a traffic light speed planning model; Based on the deep reinforcement learning simulation environment, a two-layer strategy ecological driving framework is constructed; The two-layer policy ecological driving framework is trained based on the double-delay deep deterministic policy gradient algorithm to obtain the trained two-layer policy ecological driving framework; The trained two-layer strategy ecological driving framework is used to obtain the control plan for the fleet to be processed, completing the intelligent control of the connected fleet.

[0050] In this embodiment, Figure 1 The flowchart of the connected fleet ecological driving control method based on deep reinforcement learning provided by the present invention is shown. The method first establishes a deep reinforcement learning simulation environment for the spacing and dynamics model of the connected fleet, a composite energy source model composed of fuel cells / batteries, and a traffic light speed planning model; secondly, based on the established model, the action, state space and reward mechanism of the upper-level ecological driving strategy are constructed, and the dynamic management strategy of the fleet is defined as the optimization goal; then, based on the established model, the action, state space and reward mechanism of the lower-level ecological driving strategy are constructed, and the expected ecological driving speed and power distribution plan of the fleet are defined as the optimization goal; and a deep deterministic policy gradient algorithm based on prioritized experience replay is designed to complete the training of the above-mentioned ecological driving framework using data-driven methods.

[0051] A fleet spacing and vehicle dynamics model is established, as is a fuel cell / battery model and a demand power model that considers energy source performance degradation. This is done to obtain the fleet's demand power, speed, acceleration, energy source degradation, hydrogen consumption, and battery state of charge (SoC) at the sampling moment. A green light speed optimization advisory model is also established to provide speed recommendations for the fleet to pass through intersections uninterruptedly.

[0052] Based on the established green light speed optimization consulting model and traffic light information, an upper-level strategy based on the TD3 ecological driving framework is designed. With traffic efficiency and safety as performance indicators, decisions are made on whether the fleet should pass as a whole, stop as a whole at the intersection, or pass separately under the road speed limit and vehicle acceleration limit.

[0053] Based on the platoon spacing model, vehicle dynamics model, energy source model, and platoon dynamic management strategy, a lower-level strategy based on the TD3 eco-driving framework is designed. Taking platoon driving comfort, energy efficiency, and energy source durability as performance indicators, the desired eco-driving speed and energy source power distribution scheme are optimized within the constraints of platoon speed and energy source physical parameters. According to the established ecological driving framework, a dataset consisting of road fleet location information, traffic light information, road slope, road friction coefficient, and vehicle driving status information under typical driving cycles obtained by the road infrastructure unit is used to train the ecological driving framework.

[0054] The platoon spacing and vehicle dynamics model is:

[0055] in, is the total length of the convoy; is the total number of team members; is the distance between vehicles; The vehicle captain; is the current speed of the fleet; is the reaction time; is the current acceleration of the team; is the mass of the vehicle; the traction provided to the vehicle; is the air resistance; is the rolling resistance; is the slope resistance; is the air density; is the air resistance coefficient; is the vehicle's front windshield area; is the rolling resistance coefficient; is the acceleration due to gravity; is the road slope; is a sine function; is the cosine function.

[0056] The composite energy source model includes a hydrogen consumption sub-model, a fuel cell voltage degradation sub-model, and a lithium battery SOC sub-model considering capacity degradation:

[0057]

[0058]

[0059] in, is the actual hydrogen consumption of the fuel cell; is the current running time; is the output power of the fuel cell; for the efficiency of the fuel cell system; is the lower calorific value of hydrogen; is the output voltage of the fuel cell; is the power of the auxiliary system; is the fuel cell voltage degradation rate; is the correction factor for real driving; is the degradation rate of the fuel cell under start-stop operation; is the degradation rate of the fuel cell under low-power operation, where the low-power operation is an operation with a power lower than a power threshold; is the degradation rate of the fuel cell under high-power operation, where the high-power operation is an operation with a power not less than a power threshold; is the degradation rate of the fuel cell during power fluctuations; is the average number of starts and stops of the fuel cell; is the operating time of the fuel cell in low power state; is the operating time of the fuel cell in high power state; is the power change of the fuel cell; is the SOC value of the lithium battery at the current moment; is the initial SOC of the lithium battery; The charging and discharging efficiency of lithium batteries; is the lithium battery current; is the lithium battery capacity; is the nominal capacity of the lithium battery; is the degradation degree of lithium battery capacity; is the number of charge and discharge cycles.

[0060] In this embodiment, the composite energy source system includes a fuel cell and a battery. The topology of the fuel cell hybrid vehicle power system is as follows: Figure 2 As shown, it includes a fuel cell, a battery, a DC-DC converter, a DC bus, a DC / AC inverter and a motor. Both the battery and the fuel cell provide electricity to meet the power demand, and the energy management system is designed to coordinate the above-mentioned composite energy sources. This embodiment uses a proton exchange membrane fuel cell. The performance of the fuel cell will gradually degrade with the use time and changes in working conditions, especially under conditions of high power output, low power output, frequent power changes or frequent starts and stops. Therefore, an empirical degradation model is used to predict the degradation trend under different working modes to extend its service life.

[0061] In this embodiment, and represent the operating time of the fuel cell in the lower power and higher power states, respectively, represents the power change of the fuel cell, / Number of starts and stops, , and Represents the degradation rate of fuel cells under four conditions in the experimental environment, Indicates the correction factor for real driving. Generally speaking, a battery is considered to have reached the end of its life when the nominal capacity loss reaches 20%.

[0062] The vehicle demand power model is:

[0063] in, The vehicle requires power; the traction provided to the vehicle; is the current speed of the fleet; is the motor efficiency; is the efficiency of the unidirectional DC / DC converter; is the efficiency of the bidirectional DC / DC converter; is the output power of the fuel cell; is the power of lithium battery.

[0064] The traffic light speed planning model is:

[0065] in, The minimum speed at which a convoy can pass through an intersection without stopping under a green light; is the current speed of the fleet; The maximum speed at which a convoy can pass through an intersection without stopping under a green light; The maximum speed required by the current road; The signal light status; for the green light; For the team The distance between the vehicle and the intersection; The remaining time of the current signal light; for red lights; The duration of the green light.

[0066] In this embodiment, a green light speed optimization advisory algorithm based on traffic light information is used to provide real-time driving speed recommendations for vehicles, ensuring that the convoy can pass through the green light without interruption. If the speed cannot meet the above constraints, the convoy will stop at the traffic light intersection.

[0067] The dual-strategy eco-driving framework includes an upper-strategy for obtaining a fleet dynamic management strategy and a lower-strategy for obtaining a fleet's desired eco-driving speed and power allocation scheme.

[0068] The action space, state space and reward function of the upper-level strategy are:

[0069]

[0070]

[0071] in, is the state space of the upper-level strategy L1; The signal light status; For the team The distance between the vehicle and the intersection; is the current speed of the fleet; The remaining time of the current signal light; How long the green light lasts; is the action space of the upper-level strategy L1; For the team Separation of vehicles; is the total number of convoy members; pass means the convoy passes through the intersection as a whole without separation; stop means the convoy stops before the intersection as a whole; is the reward function of the upper strategy L1; is the weight coefficient of the traffic efficiency item; is the weight coefficient of the fleet stability term; a reward for safe passage for the convoy; To encourage overall traffic rewards, Indicates separation, It means not separated; Reward for traffic efficiency; A reward for fleet stability; The maximum speed at which a convoy can pass through an intersection without stopping under a green light; The minimum speed at which a convoy can pass through an intersection without stopping under a green light; The maximum speed required by the current road; The minimum speed required for the current road; It is the comfortable acceleration of the controlled vehicles in the convoy.

[0072] In this example, a top-level ecological driving strategy based on a connected fleet is constructed based on the established model. The fleet consists of a leader vehicle and multiple follower vehicles. The leader vehicle is responsible for communicating with the road infrastructure unit to obtain red light information and also coordinates the entire fleet. Maintaining the fleet's current speed through the intersection is the optimal decision that the top-level strategy can provide. However, due to the fleet's distance from the intersection, the signal time, and the fleet's speed, the optimal decision is often idealistic. Therefore, it is necessary to design an incentive mechanism for dynamic fleet management to improve the fleet's efficiency and stability.

[0073] Indicates the comfortable acceleration of the controlled vehicles in the convoy. The purpose of this item is to constrain the speed range of the convoy to match the speed range allowed by the traffic light window. When the intersection of the optimized speed set and the road speed limit set is empty, it means that the convoy cannot pass through the intersection continuously within the speed limit range in its current form. In addition, when the traffic light is green, there is In this case, the optimized speed set will not exist and the convoy will not be able to pass through the intersection uninterruptedly. The project aims to encourage vehicles to pass through intersections in a coordinated manner. It represents the penalty for the time loss of distance-speed coupling, thereby guiding vehicles to arrive at the intersection faster and complete the passage as much as possible in the scenario where the remaining green light time is limited. The first half of the item is the maximum speed that the vehicle can theoretically reach within the remaining green light time, and the second half is the median value of the speed range. By balancing the vehicle's acceleration capability and speed, the stability of the fleet can be improved.

[0074] The state space, action space and reward function of the lower-level strategy are:

[0075]

[0076]

[0077]

[0078] in, is the state space of the lower-level strategy L2; For the team The distance between the vehicle and the intersection; is the current speed of the fleet; is the SOC value of the lithium battery at the current moment; is the output power of the fuel cell; is the road slope; is the rolling resistance coefficient; is the action space of the lower-level strategy L2; is the acceleration of the vehicle; The maximum deceleration allowed for the fleet; The maximum acceleration allowed by the team; is the power change of the fuel cell; is the reward function of the lower layer strategy L2; Rewards for fuel consumption; A reward for driving comfort; For the Lagrange multipliers for term constraints; For the Item constraints; is the number of items; is the speed constraint; is the fuel cell power constraint; is the acceleration constraint term; is the lithium battery power constraint; is the SOC constraint item; The maximum speed required by the current road; is the fuel cell nominal power; is the current acceleration of the team; The maximum charging power of the lithium battery; is the maximum discharge power of the lithium battery; is the power of the lithium battery; is the actual hydrogen consumption of the fuel cell; is the adaptive equivalent coefficient; is the weight of energy performance degradation; is the virtual hydrogen consumption of the lithium battery; is the degradation degree of lithium battery capacity; is the fuel cell voltage degradation rate; is the discharge efficiency of the lithium battery; is the average efficiency of the fuel cell system; is the air density; The charging efficiency of lithium batteries; is the impact, whose value is the acceleration with sampling time The amount of change; is the weight coefficient used to adjust the acceptance of the shock amplitude; Adaptive equivalent coefficient A constant determined by the upper limit of ; is the weight coefficient of lithium battery SOC; is the reference SOC of lithium battery.

[0079] In this embodiment, based on the established model and the designed upper-level ecological driving strategy, a lower-level ecological driving strategy based on the connected fleet is constructed to obtain the desired ecological driving speed and power distribution scheme of the energy source. The overall driving framework composed of the upper-level strategy and the lower-level strategy is as follows: Figure 3 The optimization objectives of the lower-level strategy include driving comfort, energy saving, and energy source durability.

[0080] Represents the Lagrangian function consisting of speed constraint, acceleration constraint, battery SOC and fuel cell / battery power constraint terms, by dynamically adjusting the Lagrangian multiplier , to ensure the selection of the optimal strategy while satisfying the safety constraints. When a safety constraint is violated, the corresponding multiplier The violation will increase by 5% of the violation level. If the constraint is satisfied, the multiplier It will decrease at an exponential decay rate of 2%, thereby achieving dynamic adjustment of the risk level.

[0081] In this embodiment, Obtained by the adaptive equivalent consumption strategy. This strategy equates the instantaneous consumption of the battery to the virtual hydrogen consumption. , by converting the multi-objective optimization problem involving the consumption of composite energy sources into a single-objective optimization of hydrogen consumption to reduce the design complexity, and designing an adaptive equivalent factor to maintain SoC stability. and Respectively represent the performance degradation of fuel cells and batteries, obtained from their respective degradation models, represents the average efficiency of the fuel cell system, and Respectively represent the charge and discharge efficiency of the battery, and Represent the adaptive equivalent coefficient and the weight of energy performance degradation respectively. The specific design is as follows:

[0082] in, represents the weight coefficient of battery SoC, Indicates the reference SoC of the battery, is an adaptive equivalent coefficient The value determined by the upper limit.

[0083] The reward function of the underlying policy Includes bonus for driving comfort Vehicle driving comfort is related to shock (i.e., the amount of change in acceleration). Lower shock amplitude can reduce vehicle bumps. It is used to adjust the tolerance to impact amplitude. The smaller the value, the more comfortable the driving experience.

[0084] The update process of the critic network and the orator network in the upper strategy is:

[0085]

[0086]

[0087]

[0088] in, For the upper-level strategy The critic estimates the network loss function and the critic estimates the network parameters gradient; For the upper-level strategy A critic estimates the parameters of the network; is the capacity of the small batch sample; Index for mini-batch samples; is the target Q value of the upper-level strategy; For the upper-level strategy A critic estimates the Q value of the network for the current state-action pair; for Estimating network parameters for critics gradient; Estimating network numbers for critics; is the current state of the upper-level policy; is the current action of the upper-level strategy; is the reward value of the upper-level strategy; is the discount on future rewards of the upper-level strategy; Estimate the Q value of the next state-action pair for the first critic target network of the upper strategy; The Q value of the second critic target network of the upper strategy for the next state-action pair is estimated; is the state of the upper-level strategy at the next moment; The parameters of the orator target network of the upper-level strategy; are the parameters of the first critic target network of the upper strategy; are the parameters of the second critic target network of the upper strategy; Estimate the network loss function for the speaker of the upper strategy and estimate the network parameters for the speaker gradient; Estimate network parameters for the upper-level strategy orator; Estimate the network action output parameters for the upper-level policy speaker gradient; Estimate the actions output by the network for the upper-level policy speaker; Estimate the Q-value of the network output for the first critic of the upper strategy to the action gradient; is the soft update rate of the upper layer strategy; For the upper-level strategy Parameters of the critic-target network; The first Parameters of the critic-target network; The parameters of the orator target network after the soft update of the upper layer strategy; Prioritize experience samples is the total capacity of the experience pool; E is the number of all samples; Normalize the priorities of all samples; To adjust the priority hyperparameter; For the The sampling value of an experience sample; is the TD error corresponding to the sample; is a constant.

[0089] The updating process of the critic network and the orator network in the lower layer strategy is:

[0090]

[0091]

[0092] in, The first The critic estimates the network loss function and the critic estimates the network parameters gradient; The first A critic estimates the parameters of the network; is the capacity of the small batch sample; Index for mini-batch samples; is the target Q value of the lower-level strategy; The first A critic estimates the Q value of the network for the current state-action pair; for Estimating network parameters for critics gradient; Estimating network numbers for critics; is the current state of the underlying policy; is the current action of the lower-level strategy; is the reward value of the lower-level strategy; is the discount on future rewards of the underlying strategy; Estimate the Q value of the next state-action pair for the first critic target network of the lower layer strategy; Estimate the Q value of the next state-action pair for the second critic target network of the lower layer strategy; is the state of the lower-level strategy at the next moment; Parameters of the orator target network for the lower-level strategy; are the parameters of the first critic target network of the lower layer strategy; are the parameters of the second critic target network of the lower layer strategy; Estimate the network loss function for the speaker of the lower policy and estimate the network parameters for the speaker gradient; Estimate network parameters for the lower-level strategic orator; Estimate network action output parameters for the underlying policy speaker gradient; Estimate the actions output by the network for the underlying policy speaker; Estimate the Q-value of the network output for the first critic of the lower layer strategy to the action gradient; is the soft update rate of the lower layer strategy; The first Parameters of the critic-target network; The first value after the soft update of the lower layer strategy Parameters of the critic-target network; Parameters of the orator target network after soft update of the underlying policy.

[0093] In this embodiment, both upper and lower layer strategies adopt the same network structure and update process.

[0094] In this embodiment, according to the hierarchical ecological driving strategy, a hierarchical deep reinforcement learning method based on the speech-critic network is designed and the required hierarchical quadruple is defined: "Current state ", "The next moment state ", "Execute action ", "Rewards A dataset consisting of traffic light information, road information, and various classic driving cycles was used for training to solve the aforementioned hierarchical deep reinforcement learning algorithm. The hyperparameters involved in the ecological driving strategy based on hierarchical deep reinforcement learning include: the critic and orator network learning rates, the future reward discount factor, the number of hidden layers, the number of neurons, the activation function type, and the soft update rate. The specific values ​​of these hyperparameters are shown in Table 1.

[0095] Table 1

[0096] In this embodiment, when using the trained eco-driving framework for policy solving, the upper and lower layer policy networks are first deployed to the vehicle control system, with a fixed structure for real-time reasoning. The vehicle perception system continuously collects current state information, including speed, required power, vehicle distance, traffic light status, intersection information, etc., and inputs it into the policy network. The upper layer policy determines whether to accelerate, wait, or decouple traffic based on the traffic light phase and queue status, and outputs corresponding macro-behavior instructions. Based on this, the lower layer policy combines the vehicle's energy state, driving status, and instructions transmitted from the upper layer to output target acceleration and power distribution, taking into account both traffic demand and energy consumption optimization. The control signal output by the policy is converted into vehicle action by the execution layer, and the environmental state is then updated for the next decision, completing closed-loop control.

Claims

1. An intelligent control method for a connected vehicle fleet for eco-driving, characterized in that: include: Constructing a deep reinforcement learning simulation environment, which includes a fleet spacing and vehicle dynamics model, a composite energy source model, a vehicle demand power model, and a traffic light speed planning model; Based on the deep reinforcement learning simulation environment, a two-layer strategy ecological driving framework is constructed; The two-layer policy ecological driving framework is trained based on the double-delay deep deterministic policy gradient algorithm to obtain the trained two-layer policy ecological driving framework; The trained two-layer strategy ecological driving framework is used to obtain the control plan for the fleet to be processed, completing the intelligent control of the connected fleet.

2. The intelligent control method for a connected fleet for eco-driving according to claim 1, characterized in that: The platoon spacing and vehicle dynamics model is: in, is the total length of the convoy; is the total number of team members; The distance between vehicles; The vehicle captain; is the current speed of the fleet; is the reaction time; is the current acceleration of the team; is the mass of the vehicle; the traction provided to the vehicle; is the air resistance; is the rolling resistance; is the slope resistance; is the air density; is the air resistance coefficient; is the vehicle's front windshield area; is the rolling resistance coefficient; is the acceleration due to gravity; is the road slope; is a sine function; is the cosine function.

3. The intelligent control method for a connected fleet for eco-driving according to claim 1, characterized in that: The composite energy source model includes a hydrogen consumption sub-model, a fuel cell voltage degradation sub-model, and a lithium battery SOC sub-model considering capacity degradation: in, is the actual hydrogen consumption of the fuel cell; is the current running time; is the output power of the fuel cell; for the efficiency of the fuel cell system; is the lower calorific value of hydrogen; is the output voltage of the fuel cell; is the power of the auxiliary system; is the fuel cell voltage degradation rate; is the correction factor for real driving; is the degradation rate of the fuel cell under start-stop operation; is the degradation rate of the fuel cell under low-power operation, where the low-power operation is an operation with a power lower than a power threshold; is the degradation rate of the fuel cell under high-power operation, where the high-power operation is an operation with a power not less than a power threshold; is the degradation rate of the fuel cell during power fluctuations; is the average number of starts and stops of the fuel cell; is the operating time of the fuel cell in low power state; is the operating time of the fuel cell in high power state; is the power change of the fuel cell; is the SOC value of the lithium battery at the current moment; is the initial SOC of the lithium battery; The charging and discharging efficiency of lithium batteries; is the lithium battery current; is the lithium battery capacity; is the nominal capacity of the lithium battery; is the degradation degree of lithium battery capacity; is the number of charge and discharge cycles.

4. The intelligent control method for a connected fleet for eco-driving according to claim 1, characterized in that: The vehicle demand power model is: in, The vehicle requires power; the traction provided to the vehicle; is the current speed of the fleet; is the motor efficiency; is the efficiency of the unidirectional DC / DC converter; is the efficiency of the bidirectional DC / DC converter; is the output power of the fuel cell; is the power of lithium battery.

5. The intelligent control method for a connected fleet oriented to ecological driving according to claim 1 is characterized in that: The traffic light speed planning model is: in, The minimum speed at which a convoy can pass through an intersection without stopping under a green light; is the current speed of the fleet; The maximum speed at which a convoy can pass through an intersection without stopping under a green light; The maximum speed required by the current road; The signal light status; for the green light; For the team The distance between the vehicle and the intersection; The remaining time of the current signal light; for red lights; The duration of the green light.

6. The intelligent control method for a connected fleet for eco-driving according to claim 1 is characterized in that: The dual-strategy eco-driving framework includes an upper-strategy for obtaining a fleet dynamic management strategy and a lower-strategy for obtaining a fleet's desired eco-driving speed and power allocation scheme.

7. The intelligent control method for a connected fleet oriented to ecological driving according to claim 6 is characterized in that: The action space, state space and reward function of the upper-level strategy are: in, is the state space of the upper-level strategy L1; The signal light status; For the team The distance between the vehicle and the intersection; is the current speed of the fleet; The remaining time of the current signal light; How long the green light lasts; is the action space of the upper-level strategy L1; For the team Separation of vehicles; is the total number of convoy members; pass means the convoy passes through the intersection as a whole without separation; stop means the convoy stops before the intersection as a whole; is the reward function of the upper strategy L1; is the weight coefficient of the traffic efficiency item; is the weight coefficient of the fleet stability term; a reward for safe passage for the convoy; To encourage overall traffic rewards, Indicates separation, It means not separated; Reward for traffic efficiency; A reward for fleet stability; The maximum speed at which a convoy can pass through an intersection without stopping under a green light; The minimum speed at which a convoy can pass through an intersection without stopping under a green light; The maximum speed required by the current road; The minimum speed required for the current road; It is the comfortable acceleration of the controlled vehicles in the convoy.

8. The intelligent control method for a connected fleet oriented to ecological driving according to claim 6 is characterized in that: The state space, action space and reward function of the lower-level strategy are: in, is the state space of the lower-level strategy L2; For the team The distance between the vehicle and the intersection; is the current speed of the fleet; is the SOC value of the lithium battery at the current moment; is the output power of the fuel cell; is the road slope; is the rolling resistance coefficient; is the action space of the lower-level strategy L2; is the acceleration of the vehicle; The maximum deceleration allowed by the team; The maximum acceleration allowed by the team; is the power change of the fuel cell; is the reward function of the lower layer strategy L2; Rewards for fuel consumption; A reward for driving comfort; For the Lagrange multipliers for term constraints; For the Item constraints; is the number of items; is the speed constraint; is the fuel cell power constraint; is the acceleration constraint term; is the lithium battery power constraint; is the SOC constraint item; The maximum speed required by the current road; is the fuel cell nominal power; is the current acceleration of the team; The maximum charging power of the lithium battery; is the maximum discharge power of the lithium battery; is the power of the lithium battery; is the actual hydrogen consumption of the fuel cell; is the adaptive equivalent coefficient; is the weight of energy performance degradation; is the virtual hydrogen consumption of the lithium battery; is the degradation degree of lithium battery capacity; is the fuel cell voltage degradation rate; is the discharge efficiency of the lithium battery; is the average efficiency of the fuel cell system; is the air density; The charging efficiency of lithium batteries; is the impact, whose value is the acceleration with sampling time The amount of change; is the weight coefficient used to adjust the acceptance of the shock amplitude; Adaptive equivalent coefficient A constant determined by the upper limit of ; is the weight coefficient of lithium battery SOC; is the reference SOC of lithium battery.

9. The intelligent control method for a connected fleet oriented to ecological driving according to claim 6, characterized in that: The update process of the critic network and the orator network in the upper strategy is: in, For the upper-level strategy The critic estimates the network loss function and the critic estimates the network parameters gradient; For the upper-level strategy A critic estimates the parameters of the network; is the capacity of the small batch sample; Index for mini-batch samples; is the target Q value of the upper-level strategy; For the upper-level strategy A critic estimates the Q value of the network for the current state-action pair; for Estimating network parameters for critics gradient; Estimating network numbers for critics; is the current state of the upper-level policy; is the current action of the upper-level strategy; is the reward value of the upper-level strategy; is the discount on future rewards of the upper-level strategy; Estimate the Q value of the state-action pair at the next moment for the first critic target network of the upper strategy; The Q value of the second critic target network of the upper strategy for the next state-action pair is estimated; is the state of the upper-level strategy at the next moment; The parameters of the orator target network of the upper-level strategy; are the parameters of the first critic target network of the upper strategy; are the parameters of the second critic target network of the upper strategy; Estimate the network loss function for the speaker of the upper strategy and estimate the network parameters for the speaker gradient; Estimating network parameters for the upper-level strategy orator; Estimate the network action output parameters for the upper-level policy speaker gradient; Estimate the actions output by the network for the upper-level policy speaker; Estimate the Q-value of the network output for the first critic of the upper strategy to the action gradient; is the soft update rate of the upper layer strategy; For the upper-level strategy Parameters of the critic-target network; The first Parameters of the critic-target network; The parameters of the orator target network after the soft update of the upper layer strategy; Prioritize experience samples is the total capacity of the experience pool; E is the number of all samples; Normalize the priorities of all samples; To adjust the priority hyperparameters; For the The sampling value of an experience sample; is the TD error corresponding to the sample; is a constant.

10. The intelligent control method for a connected fleet oriented to ecological driving according to claim 6, characterized in that: The updating process of the critic network and the orator network in the lower layer strategy is: in, The first The critic estimates the network loss function and the critic estimates the network parameters gradient; The first A critic estimates the parameters of the network; is the capacity of the small batch sample; Index for mini-batch samples; is the target Q value of the lower-level strategy; The first A critic estimates the Q value of the network for the current state-action pair; for Estimating network parameters for critics gradient; Estimating network numbers for critics; is the current state of the underlying policy; is the current action of the lower-level strategy; is the reward value of the lower-level strategy; is the discount on future rewards of the underlying strategy; Estimate the Q value of the next state-action pair for the first critic target network of the lower layer strategy; Estimate the Q value of the next state-action pair for the second critic target network of the lower layer strategy; is the state of the lower-level strategy at the next moment; Parameters of the orator target network for the lower-level strategy; are the parameters of the first critic target network of the lower layer strategy; are the parameters of the second critic target network of the lower layer strategy; Estimate the network loss function for the speaker of the lower layer strategy and estimate the network parameters for the speaker gradient; Estimate network parameters for the lower-level strategic orator; Estimate the network action output parameters for the underlying policy speaker gradient; Estimate the actions output by the network for the underlying policy speaker; Estimate the Q-value of the network output for the first critic of the lower layer strategy to the action gradient; is the soft update rate of the underlying strategy; The first Parameters of the critic-target network; The first value after the soft update of the lower layer strategy Parameters of the critic-target network; Parameters of the orator target network after soft update of the underlying policy.

Citation Information

Cited By

  • Multi-task concurrency and route planning speed control method based on large model

    CN121393152A