A Cooperative Vehicle Platooning Decision-making Method Based on Nested Graph Reinforcement Learning
By building nested graph models and deep reinforcement learning algorithms to optimize vehicle fleet decisions, the safety, efficiency and comfort of vehicle fleets in highway environments are solved, and efficient formation control and energy consumption reduction are achieved.
Patent Information
- Application Number
- CN202410196471.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-02-22
AI Technical Summary
The prior art is difficult to effectively integrate and utilize vehicle network information in a complex and dynamic highway environment to realize the optimization strategy of vehicle formations, especially when dealing with safety, efficiency and comfort issues in emergencies such as emergency braking, sudden lane changes and traffic congestion.
Build a nested graph model to represent the vehicle team, use deep reinforcement learning algorithms to optimize the safety following capabilities, energy saving effects and comfort level in the formation, optimize vehicle decision-making through feature extraction networks and reward functions, including graph attention layer, full connection layer and multi-head attention mechanism, and design safe driving, efficiency improvement, energy saving and passenger comfort reward functions.
Achieve fully integrated formation decision control under efficient communication in a high-dynamic and complex highway environment, improving the robustness, flexibility and effectiveness of vehicle formations, improving road capacity, reducing energy consumption and improving driving safety and comfort.
Smart Images

Figure CN118082805B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle networking and autonomous driving, and particularly relates to a formation decision-making method for connected vehicles based on nested graph reinforcement learning. Background Art
[0002] In the field of autonomous driving, vehicle platooning is a key technology, especially in highway environments. Vehicle platooning technology enables multiple vehicles to travel efficiently in a certain format and maintain a specific distance from each other, thereby increasing road capacity, reducing energy consumption, and enhancing driving safety. Nevertheless, vehicle platooning in highway environments faces multiple challenges, including the safety of high-speed driving, real-time response in dynamic environments, effective communication between vehicles, and decision-making strategies in complex traffic situations.
[0003] With the development of vehicle-to-everything (V2X) technology, intelligent vehicles can receive and process data from other vehicles, traffic infrastructure, and even pedestrians in real time. This provides new solutions for vehicle platooning, especially in dealing with complex traffic situations and real-time decision-making. However, effectively integrating and utilizing this data to achieve an optimized platooning strategy remains a technical challenge.
[0004] In addition, deep learning, especially deep reinforcement learning (DRL), has become a hot topic in autonomous driving research due to its potential in dealing with complex and non-linear problems. DRL can learn optimal strategies through interaction with the environment, which is crucial for vehicles to make rapid decisions under changing road conditions. However, applying DRL to vehicle platooning requires solving problems such as state space design, reward mechanism design, and the stability and robustness of the algorithm.
[0005] Therefore, there is an urgent need to develop a new type of vehicle platooning technology that can fully utilize the rich information provided by vehicle networking and combine the efficient decision-making ability of deep reinforcement learning to address the challenges of vehicle platooning in highway environments. Summary of the Invention
[0006] The present invention aims to solve the problem that traditional methods in the prior art are difficult to handle emergencies, such as emergency braking, vehicles suddenly changing lanes, traffic congestion, etc., and proposes a vehicle platooning method in a complex and dynamically changing actual traffic environment, especially in a multi-lane, high-speed, and high-density highway environment. By constructing a nested graph model to represent the vehicle fleet and using a deep reinforcement learning algorithm to optimize the safe following ability, energy-saving effect, and comfort level in the platoon.
[0007] To achieve the above object, the present invention provides the following solution: A formation decision-making method for connected intelligent vehicle formations based on nested graph reinforcement learning, comprising the following steps:
[0008] S1. Collect the vehicle state information within the formation, and process the state information to obtain an inter-formation nested graph and an intra-vehicle nested graph;
[0009] S2. Use a feature extraction network to extract features from the inter-formation nested graph and the intra-vehicle nested graph, and obtain the actions of each intelligent vehicle based on the extracted features; the feature extraction network includes: a graph attention layer, a fully connected layer, and an activation layer, and a multi-head attention mechanism is introduced to improve the feature extraction ability;
[0010] S3. Optimize the actions using a reward function to obtain an intelligent vehicle decision-making method.
[0011] Further preferably, the inter-formation nested graph includes: an inter-formation sub-graph feature matrix and an inter-formation sub-graph adjacency matrix;
[0012] The inter-formation sub-graph feature matrix includes:
[0013]
[0014] In the formula, M represents the number of formations; represents the inter-formation sub-graph feature matrix; F f is the feature number of the formation; V Li represents the longitudinal speed of the leading vehicle of the i-th formation; Y ai represents the longitudinal position of the leading vehicle of the i-th formation; Y bi represents the longitudinal position of the leading vehicle of the i-th formation; represents the average speed within the i-th formation; σ vi represents the average acceleration within the i-th formation; T i represents the collision time between the leading vehicle of the i-th formation and the vehicle in front; I i represents vehicle classification.
[0015] Further preferably, the inter-formation sub-graph adjacency matrix is constructed based on the maximum and minimum distances between formations and the lane to construct a lateral dimension graph;
[0016] The weight function of the inter-formation sub-graph adjacency matrix is:
[0017]
[0018] In the formula, Δy ij represents the distance along the lane between the leading vehicle of formation i and the leading vehicle of formation j; Y is the inter-formation distance threshold.
[0019] Further preferably, the inter-vehicle nested graph includes: an inter-vehicle sub-graph feature matrix and an inter-vehicle sub-graph adjacency matrix;
[0020] The inter-vehicle sub-graph feature matrix includes:
[0021]
[0022] In the formula, N represents the number of vehicles; represents the inter-vehicle sub-graph feature matrix; F v represents the number of features of each vehicle; V i represents the longitudinal speed of the i-th vehicle; Y i represents the longitudinal position of the i-th vehicle, ΔV i represents the relative speed between the i-th vehicle and the vehicle in front; ΔY i represents the relative longitudinal distance between the i-th vehicle and the vehicle in front; a i represents the acceleration of the i-th vehicle; T vi represents the collision time between the i-th vehicle and the vehicle in front; I vi represents vehicle classification.
[0023] Further preferably, the inter-vehicle sub-graph connection matrix is constructed based on the inter-vehicle distance and the difference in inter-vehicle speeds;
[0024] The weight function of the inter-vehicle sub-graph connection matrix is:
[0025]
[0026] In the formula, Δv ij represents the speed difference; Δd ij represents the distance between vehicle i and vehicle j along the lane direction; D represents the inter-vehicle distance threshold; V represents the speed threshold.
[0027] Further preferably, the reward function includes: a safe driving reward function, an efficiency improvement reward function, an energy saving reward function, and a passenger comfort reward function.
[0028] Further preferably, the safe driving reward function includes: a safe distance sub-reward function and a collision time sub-reward function;
[0029] The safe distance sub-reward function includes:
[0030]
[0031] In the formula, S n-1,n represents the position difference between the vehicle n-1 in front and the vehicle n itself at time t; D s represents the safe distance;
[0032] Among them,
[0033]
[0034] In the formula, v0 represents the vehicle's own speed; v f represents the speed of the vehicle in front; d0 represents the minimum safe distance; a max represents the maximum deceleration; τ represents the reaction time;
[0035] The collision time sub-reward function includes:
[0036]
[0037] In the formula, TTC represents the time to collision;
[0038] Among them,
[0039]
[0040] In the formula, V n (t) represents the absolute speed of the host vehicle n at time t; V n-1 (t) represents the absolute speed of the vehicle in front n-1 at time t; V n-1,n (t) represents the relative speed between the vehicle in front n-1 and the host vehicle n at time t; S n-1,n represents the position difference between the vehicle in front n-1 and the host vehicle n at time t.
[0041] Further preferably, the efficiency improvement reward function includes: a following distance sub-reward function and a vehicle speed efficiency sub-reward function; S n-1,n (t) represents the position difference between the vehicle in front n-1 and the host vehicle n at time t, V n-1,n (t) represents the speed difference between the vehicle in front n-1 and the host vehicle n at time t.
[0042] The following distance sub-reward function includes:
[0043]
[0044] In the formula, represents the penalty weight; S n-1,n represents the position difference between the vehicle in front n-1 and the host vehicle n at time t; Δd des represents the desired following distance;
[0045] The vehicle speed efficiency sub-reward function includes:
[0046]
[0047] In the formula, V n-1,n (t) represents the relative speed between the vehicle in front n-1 and the host vehicle n at time t; Δv limit represents the limit of the speed difference between the two vehicles.
[0048] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0049] (1) The proposed connected intelligent vehicle formation architecture of the present invention is a significant improvement over the current autonomous vehicle formation technology. In a highway environment with high dynamics, high complexity, and strong randomness, it realizes fully integrated formation decision-making control under efficient communication.
[0050] (2) The present invention uses nested graphs between vehicles and formations to represent the spatio-temporal interactions between vehicles and formations. It constructs a sub-adjacency matrix between formations based on the relative positions of the leading vehicles between formations, and constructs a sub-graph adjacency matrix between vehicles based on the relative positions and speed differences between vehicles. The nested graph can introduce heterogeneous vehicle interactions, formation-vehicle hierarchical interactions, and communication interactions between formations into decision-making modeling, which can improve the robustness, flexibility, and effectiveness of vehicle formation decision-making in an uncertain traffic environment.
[0051] (3) The present invention fuses and trains the hierarchical graphs between vehicles and formations, and extracts features through a multi-layer attention network and a fully connected network for three channels, which can effectively extract the interaction information between formations, between vehicles, and from formations to vehicles in the scenario, enabling RL vehicles to more flexibly cope with changing road conditions and traffic scenarios.
[0052] (4) The present invention designs a reward function from four perspectives: driving safety, efficiency improvement, energy conservation, and passenger comfort. Vehicle formations can maintain a specific distance from each other and drive efficiently, thereby increasing road capacity, reducing energy consumption, and enhancing driving safety and comfort. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0054] Figure 1 It is a schematic flowchart of a method for connected intelligent vehicle formation decision-making based on nested graph reinforcement learning according to an embodiment of the present invention;
[0055] Figure 2 It is a schematic diagram of the feature extraction process according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] Embodiment 1:
[0059] As Figure 1 shown, this embodiment provides a formation decision-making method for connected intelligent vehicle formations based on nested graph reinforcement learning. The process includes: Vehicle-to-base station communication: All vehicles in the formation send their status information (such as position, speed, acceleration) to the nearest base station through a wireless communication network. Base station-to-cloud transmission: After collecting vehicle information, the base station transmits the data to the cloud server for processing. Cloud computing: After receiving the data, the cloud performs calculations to generate decision instructions. Instruction transmission back to the base station: The cloud sends the decision instructions back to the base station. Base station-to-RL vehicle instruction distribution: After receiving the decision instructions from the cloud, the base station distributes them to each RL vehicle for execution.
[0060] Specifically, the method includes the following steps:
[0061] S1. Collect the status information of the vehicles in the formation and process the status information to obtain an inter-formation nested graph and an inter-vehicle nested graph.
[0062] In the architecture of the connected intelligent vehicle formation, there are multiple vehicle formations. The leading vehicle of each formation is an intelligent vehicle (RL vehicle) based on reinforcement learning (RL), and the following vehicles are multiple vehicles that comply with the intelligent driving model (IDM).
[0063] Among them, the status information of the vehicle includes: position, speed, acceleration, etc. Through the above status information, an inter-formation nested graph and an inter-vehicle nested graph are constructed.
[0064] In the inter-vehicle and inter-formation nested graphs constructed in this embodiment, it is defined that there are N formations in the traffic environment, and each formation has M driverless vehicles, and multi-dimensional interactions between formations and between vehicles are carried out in the traffic environment. The global traffic environment is a nested graph G(G vi ,E), where G vi =G(V i ,E i ) represents the sub-graph between the vehicles in the i-th formation, and V i(G) is a set of graph nodes, where 1 ≤ i ≤ N. E(G) is the set of edges of the inter-formation graph; the information interaction between formations is represented as (G vi , G vj ) ∈ E. Each intelligent vehicle is regarded as a node u (u ∈ V i ); the information interaction between vehicles is represented as an edge (u, v) ∈ E i .
[0065] For the temporal representation of the nested graph, the nested graph between formations and the nested graph between vehicles at the t-th time step are respectively represented as the sub-graph feature matrix between formations, the sub-graph adjacency matrix between formations, the sub-graph feature matrix between vehicles, and the sub-graph adjacency matrix between vehicles.
[0066] Among them, the sub-graph feature matrix between formations includes:
[0067]
[0068] In the formula, M represents the number of formations; represents the sub-graph feature matrix between formations; F f is the number of features of the formation; V Li represents the longitudinal speed of the leading vehicle of the i-th formation; Y ai represents the longitudinal position of the leading vehicle of the i-th formation; Y bi represents the longitudinal position of the leading vehicle of the i-th formation; represents the average speed within the i-th formation; σ vi represents the average acceleration within the i-th formation; T i represents the collision time between the leading vehicle of the i-th formation and the vehicle in front; T i = Δs i / Δv i ; Δs i represents the relative displacement between two vehicles along the lane traveling direction, and Δv i represents the relative speed between two vehicles along the lane traveling direction; I i represents vehicle classification.
[0069] On the premise that formations can share information by themselves and all formations within the specified range can share information in the traffic scenario, a lateral dimension graph is constructed based on the maximum and minimum distances between formations and the lane to construct the sub-graph adjacency matrix between formations. Its weight function is a distance function, and the function value decreases as the distance increases, ranging from 0 to 1. The specific formula is as follows:
[0070]
[0071] In the formula, Δy ijdenotes the distance along the lane between the leading vehicle of formation i and the leading vehicle of formation j; Y is the inter-formation distance threshold, and in this embodiment, Y = 500 m.
[0072] The sub-graph feature matrix between vehicles includes:
[0073]
[0074] In the formula, denotes the sub-graph feature matrix between vehicles; F v denotes the number of features of each vehicle; V i denotes the longitudinal speed of the i-th vehicle; Y i denotes the longitudinal position of the i-th vehicle, ΔV i denotes the relative speed between the i-th vehicle and the vehicle in front; ΔY i denotes the relative longitudinal distance between the i-th vehicle and the vehicle in front; a i denotes the acceleration of the i-th vehicle; T vi denotes the collision time between the i-th vehicle and the vehicle in front; I vi denotes vehicle classification.
[0075] The calculation of the sub-graph adjacency matrix between vehicles includes four prerequisites: 1) All unmanned vehicles within the control range of the same base station can share information, denoted as a v-ij = 1; 2) Information cannot be shared between IDM vehicles; 3) All unmanned vehicles can share the information of IDM vehicles within the constraint range; 4) A vehicle can share information with itself, denoted as a v-ii = 1. The sub-graph connection matrix between vehicles is constructed based on the distance between vehicles and the difference in vehicle speeds, ranging from 0 to 1. The weight function of the sub-graph connection matrix between vehicles is:
[0076]
[0077] In the formula, Δv ij denotes the speed difference; Δd ij denotes the distance along the lane direction between vehicle i and vehicle j; D denotes the vehicle distance threshold, which is 100 m in this embodiment; V denotes the speed threshold, which is 3 m / s in this embodiment.
[0078] S2. Use the feature extraction network to extract features from the nested graph between formations and the nested graph between vehicles, and obtain the actions of each intelligent vehicle based on the extracted features; the feature extraction network includes: a graph attention layer, a fully connected layer, and an activation layer, and a multi-head attention mechanism is introduced to improve the feature extraction ability;
[0079] Such as Figure 2As shown, feature extraction includes three parts: nested graph feature extraction between vehicles, nested graph feature extraction between formations, and feature extraction of fused features. The feature extraction network includes: a graph attention layer, a fully connected layer, and an activation layer, and a multi-head attention mechanism is introduced to improve the feature extraction ability of the model.
[0080] First, a layer of graph attention network is used to extract features between formations and features between vehicles; then a layer of graph attention network is used to comprehensively extract nested graph features between formations and nested graph features between vehicles to obtain fused features, and attention weight training is performed on the above three types of features.
[0081] After that, the above three types of features are independently extracted through two layers of fully connected networks, and finally feature aggregation is performed on them. The aggregated features are uniformly extracted through a layer of multi-head graph attention network, and the actions that each RL vehicle needs to execute are output through a layer of fully connected layer.
[0082] S3. Optimize the actions using a reward function to obtain an intelligent vehicle decision-making method.
[0083] In this embodiment, the reward function is further designed from four perspectives: safe driving, efficiency improvement, energy conservation, and passenger comfort. Specifically, the reward function includes: a safe driving reward function, an efficiency improvement reward function, an energy conservation reward function, and a passenger comfort reward function.
[0084] Perform spatio-temporal coupling on driving safety. The safe driving reward function includes: a safe distance sub-reward function and a time-to-collision sub-reward function;
[0085] The safe distance sub-reward function includes:
[0086]
[0087] In the formula, S n-1,n represents the position difference (headway) between the vehicle n-1 in front and the vehicle n itself at time t; D s represents the safe distance;
[0088] Among them, when both the front and rear vehicles decelerate at the maximum deceleration, the safe distance between the two vehicles is:
[0089]
[0090] In the formula, v0 represents the vehicle's own speed; v f represents the speed of the vehicle in front; d0 represents the minimum safe distance; a max represents the maximum deceleration; τ represents the reaction time, which is 1.5 s in this embodiment.
[0091] When the gap distance is less than the safe distance, the intelligent vehicle should be punished.
[0092] The collision time sub - reward function includes:
[0093]
[0094] In the formula, TTC represents the time - to - collision;
[0095] For the collision time sub - reward function, the collision time is a parameter widely used in the transportation field to measure the safety of vehicle following, which represents the time from the current moment to the collision of the leading vehicle and the following vehicle. Numerically, it is equal to the ratio of the position difference to the speed difference between two consecutive vehicles (when the leading vehicle is slower than the following vehicle). Specifically,
[0096]
[0097] In the formula, V n (t) represents the absolute speed of the ego - vehicle n at time t; V n-1 (t) represents the absolute speed of the leading vehicle n - 1 at time t; V n-1,n (t) represents the relative speed between the leading vehicle n - 1 and the ego - vehicle n at time t; S n-1,n represents the position difference (headway) between the leading vehicle n - 1 and the ego - vehicle n at time t.
[0098] When TTC is greater than 0 s and less than 4 s, the intelligent vehicle is punished, and a very large punishment is given to too small TTC values.
[0099] The driving efficiency is coupled from the vehicle distance and the vehicle speed. The efficiency improvement reward function includes: the following - distance sub - reward function and the vehicle - speed efficiency sub - reward function;
[0100] For the following - distance sub - reward function, during the following process, the ego - vehicle's following of the leading vehicle is mainly reflected in that the following distance approaches the desired following distance Δd des = 50 m, and the ego - vehicle speed approaches the leading vehicle speed; the following - distance sub - reward function includes:
[0101]
[0102] In the formula, represents the punishment weight; S n-1,n represents the position difference (headway) between the leading vehicle n - 1 and the ego - vehicle n at time t; Δd des represents the desired following distance;
[0103] The vehicle - speed efficiency sub - reward function includes:
[0104]
[0105] In the formula, V n-1,n (t) represents the relative speed between the leading vehicle n - 1 and the ego - vehicle n at time t; Δvlimit It represents the speed difference limit between two vehicles. In this embodiment, it is taken as 20 km / h.
[0106] Regarding the driving comfort, decoupling is performed based on the order of the derivative. The passenger comfort reward function is divided into: a second-order acceleration comfort sub-reward function and a third-order jerk comfort sub-reward function. For the second-order acceleration comfort sub-reward function, in the longitudinal motion, the smaller the absolute value of the acceleration, the higher the longitudinal comfort. Its expression is:
[0107]
[0108] In the formula, a(t) represents the acceleration. In this embodiment, the maximum absolute value of the acceleration is 4.5 m / s 2 .
[0109] For the third-order jerk comfort sub-reward function, jerk is the rate of change of acceleration, and jerk is used to measure driving comfort. In the longitudinal motion, the smaller the absolute value of the jerk, the higher the comfort. Therefore, the jerk is less than 2.944.5 m / s 3 . For the case where the jerk exceeds 2.944.5 m / s 3 , a penalty weight is given; the third-order jerk comfort sub-reward function is:
[0110]
[0111] In the formula, represents the penalty weight; j(t) represents the third-order jerk of the vehicle at time t.
[0112] The energy saving reward function is: r Energy = -Energy; in the formula, Energy is the total energy consumed by the RL vehicle at each time step.
[0113] The continuous DRL algorithm is adopted: including DDPG and its variants, SAC, PPO algorithms and various graph neural networks: including network models such as GCN, GAT, etc. for training. The actions at each time step are transmitted to the controller of the intelligent vehicle for interaction with the environment, and the rewards and the next state from the environment are received. According to the collected sample data, the parameters are continuously optimized so that the intelligent vehicle can better execute tasks and obtain more rewards. When the performance of the intelligent vehicle reaches a satisfactory level, the training process ends.
[0114] Embodiment 2
[0115] This embodiment provides a connected intelligent vehicle formation decision-making system based on nested graph reinforcement learning. The system is used to implement the method proposed in Embodiment 1 and includes: a connected intelligent vehicle formation architecture module, an inter-vehicle and inter-formation nested graph representation module, a hierarchical graph fusion decision network module, and a reinforcement learning training module.
[0116] The connected intelligent vehicle formation architecture module is used to collect the vehicle state information within the formation.
[0117] In the architecture of the connected intelligent vehicle formation, there are multiple vehicle formations. The leading vehicle of each formation is an intelligent vehicle (RL vehicle) based on reinforcement learning (RL), and the following vehicles are multiple vehicles that comply with the intelligent driving model (IDM).
[0118] Among them, the state information of the vehicle includes: position, speed, acceleration, etc. Through the above state information, an inter-formation nested graph and an inter-vehicle nested graph are constructed.
[0119] The inter-vehicle and inter-formation nested graph representation module is used to process the state information to obtain an inter-formation nested graph and an inter-vehicle nested graph.
[0120] In the inter-vehicle and inter-formation nested graphs constructed in this embodiment, it is defined that there are N formations in the traffic environment, and each formation has M unmanned vehicles, and multi-dimensional interactions between formations and between vehicles are carried out in the traffic environment. The global traffic environment is a nested graph G(G vi ,E), where G vi = G(V i ,E i ) represents the sub-graph between the vehicles in the i-th formation, V i (G) is the set of graph nodes, 1 ≤ i ≤ N. E(G) is the edge set of the inter-formation graph; the information interaction between formations is expressed as (G vi , G vj ) ∈ E. Each intelligent vehicle is regarded as a node u (u ∈ V i ) of the graph; the information interaction between vehicles is expressed as the edge (u, v) ∈ E i .
[0121] For the temporal representation of the nested graph, the inter-formation nested graph and the inter-vehicle nested graph are respectively represented as the inter-formation sub-graph feature matrix and the inter-formation sub-graph adjacency matrix, the inter-vehicle sub-graph feature matrix and the inter-vehicle sub-graph adjacency matrix at the t-th time step.
[0122] Among them, the inter-formation sub-graph feature matrix includes:
[0123]
[0124] In the formula, M represents the number of formations; Denote the sub - graph feature matrix between formations; F f is the number of features of the formation; V Li Denote the longitudinal speed of the leading vehicle of the \(i\) - th formation; Y ai Denote the longitudinal position of the leading vehicle of the \(i\) - th formation; Y bi Denote the longitudinal position of the leading vehicle of the \(i\) - th formation; Denote the average speed within the \(i\) - th formation; σ vi Denote the average acceleration within the \(i\) - th formation; T i Denote the collision time between the leading vehicle of the \(i\) - th formation and the vehicle in front; T i =Δs i / Δv i ; Δs i Denote the relative displacement along the lane traveling direction between two vehicles, Δv i Denote the relative speed along the lane traveling direction between two vehicles; I i Denote the vehicle classification.
[0125] When the formations can share information by themselves and all formations can share information in the traffic scenario within the specified range, based on the maximum and minimum distances between formations and the lanes, construct the lateral - dimension graph and the sub - graph adjacency matrix between formations. Its weight function is a distance function, and the function value decreases as the distance increases, ranging from 0 to 1. The specific formula is as follows:
[0126]
[0127] In the formula, Δy ij Denote the distance along the lane between the leading vehicle of formation \(i\) and the leading vehicle of formation \(j\); Y is the distance threshold between formations. In this embodiment, Y = 500m.
[0128] The sub - graph feature matrix between vehicles includes:
[0129]
[0130] In the formula, Denote the sub - graph feature matrix between vehicles; F v Denote the number of features of each vehicle; V i Denote the longitudinal speed of the \(i\) - th vehicle; Y i Denote the longitudinal position of the \(i\) - th vehicle, ΔV i Denote the relative speed between the \(i\) - th vehicle and the vehicle in front; ΔY i Denote the relative longitudinal distance between the \(i\) - th vehicle and the vehicle in front; a i Denote the acceleration of the \(i\) - th vehicle; T vi Denote the collision time between the \(i\) - th vehicle and the vehicle in front; I vi Denote the vehicle classification.
[0131] The calculation of the sub - graph adjacency matrix between vehicles includes four prerequisites: 1) All driverless vehicles within the control range of the same base station can share information, denoted as a v-ij = 1; 2) Information cannot be shared between IDM vehicles; 3) All driverless vehicles can share the information of IDM vehicles within the constraint range; 4) A vehicle can share information with itself, denoted as a v-ii = 1. The sub - graph connection matrix between vehicles is constructed based on the distance between vehicles and the difference in vehicle speeds, with a range between 0 and 1. The weight function of the sub - graph connection matrix between vehicles is as follows:
[0132]
[0133] In the formula, Δv ij represents the speed difference; Δd ij represents the distance along the lane direction between vehicle i and vehicle j; D represents the vehicle - to - vehicle distance threshold, which is 100m in this embodiment; V represents the speed threshold, which is 3m / s in this embodiment.
[0134] As Figure 2 shown, feature extraction includes three parts: feature extraction of the nested graph between vehicles, feature extraction of the nested graph between formations, and feature extraction of the fused features. The feature extraction network includes: a graph attention layer, a fully - connected layer, and an activation layer, and a multi - head attention mechanism is introduced to improve the feature extraction ability of the model.
[0135] The working process of the hierarchical graph fusion decision network module includes:
[0136] First, a layer of graph attention network is used to extract the features between formations and the features between vehicles; then another layer of graph attention network is used to comprehensively extract the features of the nested graph between formations and the nested graph between vehicles to obtain the fused features, and attention weight training is performed on the above three types of features.
[0137] After that, the above three types of features are independently extracted through two layers of fully - connected networks, and finally their features are aggregated. The aggregated features are uniformly extracted through a layer of multi - head graph attention network, and the actions that each RL vehicle needs to execute are output through a layer of fully - connected layer.
[0138] The reinforcement learning training module is used to optimize the actions using a reward function to obtain an intelligent vehicle decision - making method.
[0139] In this embodiment, the reward function is further designed from four aspects: safe driving, efficiency improvement, energy conservation, and passenger comfort. Specifically, the reward function includes: a safe - driving reward function, an efficiency - improvement reward function, an energy - conservation reward function, and a passenger - comfort reward function.
[0140] Spatially and temporally couple the driving safety. The safe driving reward function includes: a safe distance sub-reward function and a time to collision sub-reward function;
[0141] The safe distance sub-reward function includes:
[0142]
[0143] In the formula, S n-1,n represents the position difference (headway) between the leading vehicle n-1 and the ego vehicle n at time t; D s represents the safe distance;
[0144] Among them, when both the front and rear vehicles decelerate at the maximum deceleration, the safe distance between the two vehicles is:
[0145]
[0146] In the formula, v0 represents the vehicle's own speed; v f represents the speed of the leading vehicle; d0 represents the minimum safe distance; a max represents the maximum deceleration; τ represents the reaction time, which is 1.5 s in this embodiment.
[0147] When the gap distance is less than the safe distance, the intelligent vehicle should be penalized.
[0148] The time to collision sub-reward function includes:
[0149]
[0150] In the formula, TTC represents the time to collision;
[0151] For the time to collision sub-reward function, the time to collision is a parameter widely used in the traffic field to measure the safety of vehicle following. It represents the time from the current moment to the collision between the leading vehicle and the following vehicle. Numerically, it is equal to the ratio of the position difference to the speed difference between two consecutive vehicles (when the leading vehicle is slower than the following vehicle). Specifically,
[0152]
[0153] In the formula, V n (t) represents the absolute speed of the ego vehicle n at time t; V n-1 (t) represents the absolute speed of the leading vehicle n-1 at time t; V n-1,n (t) represents the relative speed between the leading vehicle n-1 and the ego vehicle n at time t; S n-1,n represents the position difference (headway) between the leading vehicle n-1 and the ego vehicle n at time t.
[0154] When TTC is greater than 0 s and less than 4 s, penalize the intelligent vehicle, and give a great penalty to too small TTC values.
[0155] The driving efficiency is coupled from the vehicle distance and the vehicle speed. The efficiency improvement reward function includes: a following distance sub-reward function and a vehicle speed efficiency sub-reward function;
[0156] For the following distance sub-reward function, during the following process, the following of the host vehicle to the leading vehicle is mainly reflected in that the following distance approaches the expected following distance Δd des = 50 m, and the speed of the host vehicle approaches the speed of the leading vehicle; the following distance sub-reward function includes:
[0157]
[0158] In the formula, represents the penalty weight; S n-1,n represents the position difference (headway) between the leading vehicle n-1 and the host vehicle n at time t; Δd des represents the expected following distance;
[0159] The vehicle speed efficiency sub-reward function includes:
[0160]
[0161] In the formula, V n-1,n (t) represents the relative speed between the leading vehicle n-1 and the host vehicle n at time t; Δv limit represents the limit of the vehicle speed difference between the two vehicles. In this embodiment, it is taken as 20 km / h.
[0162] For the driving comfort, it is decoupled based on the order of the derivative. The passenger comfort reward function is divided into: a second-order acceleration comfort sub-reward function and a third-order jerk comfort sub-reward function. For the second-order acceleration comfort sub-reward function, in the longitudinal motion, the smaller the absolute value of the acceleration, the higher the longitudinal comfort. Its expression is:
[0163]
[0164] In the formula, a(t) represents the acceleration. In this embodiment, the maximum absolute value of the acceleration is 4.5 m / s 2 .
[0165] For the third-order jerk comfort sub-reward function, the jerk is the rate of change of the acceleration. The driving comfort is measured by the jerk. In the longitudinal motion, the smaller the absolute value of the jerk, the higher the comfort. Therefore, the jerk is less than 2.944.5 m / s 3 . For the case where the jerk exceeds 2.944.5 m / s 3 , a penalty weight is given; the third-order jerk comfort sub-reward function is:
[0166]
[0167] In the formula, represents the penalty weight; j(t) represents the third-order jerk of the vehicle at time t.
[0168] The energy-saving reward function is: r Energy = -Energy; in the formula, Energy is the total energy consumed by the RL vehicle at each time step.
[0169] The continuous DRL algorithm is adopted: including DDPG and its variants, SAC, PPO algorithms and various graph neural networks: including network models such as GCN, GAT, etc. for training. The actions at each time step are transmitted to the controller of the intelligent vehicle for interaction with the environment, and the rewards and the next state from the environment are received. According to the collected sample data, the parameters are continuously optimized so that the intelligent vehicle can better perform tasks and obtain more rewards. When the performance of the intelligent vehicle reaches a satisfactory level, the training process ends.
[0170] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A formation decision-making method for connected intelligent vehicle based on nested graph reinforcement learning, characterized in that Including the following steps: S1. Collect the vehicle status information within the formation, and process the status information to obtain the inter-formation nested graph and the inter-vehicle nested graph; S2. Use the feature extraction network to extract features from the inter-formation nested graph and the inter-vehicle nested graph, and obtain the actions of each intelligent vehicle based on the extracted features; The feature extraction network includes: a graph attention layer, a fully connected layer, and an activation layer, and a multi-head attention mechanism is introduced to improve the feature extraction ability; S3. Optimize the actions using a reward function to obtain an intelligent vehicle decision-making method; The inter-formation nested graph includes: an inter-formation sub-graph feature matrix and an inter-formation sub-graph adjacency matrix; The inter-formation sub-graph feature matrix includes: , In the formula, represents the number of formations; represents the sub-graph feature matrix between formations; F f is the number of features of the formation; V Li represents the longitudinal speed of the leading vehicle of the i-th formation; Y ai represents the longitudinal position of the leading vehicle of the i-th formation; Y bi represents the longitudinal position of the leading vehicle of the i-th formation; represents the average speed within the i-th formation; represents the average acceleration within the i-th formation; T i represents the collision time between the leading vehicle of the i-th formation and the vehicle in front; I i represents vehicle classification; The inter-formation sub-graph adjacency matrix is constructed based on the maximum and minimum distances between formations and the construction of a horizontal dimension graph with lanes; The weight function of the inter-formation sub-graph adjacency matrix is: , where ∆y ij represents the distance along the lane between the leading vehicle of formation i and the leading vehicle of formation j; Y is the inter-formation distance threshold; The inter-vehicle nested graph includes: an inter-vehicle sub-graph feature matrix and an inter-vehicle sub-graph adjacency matrix; The inter-vehicle sub-graph feature matrix includes: , Where N represents the number of vehicles; represents the sub-graph feature matrix between vehicles; F v represents the number of features of each vehicle; V i represents the longitudinal speed of the i-th vehicle; Y i represents the longitudinal position of the i-th vehicle, ΔV i represents the relative speed of the i-th vehicle and the vehicle in front; ΔY i represents the relative longitudinal distance between the i-th vehicle and the vehicle in front; represents the acceleration of the i-th vehicle; T vi represents the collision time of the i-th vehicle and the vehicle in front; I vi represents vehicle classification; The inter-vehicle sub-graph adjacency matrix is constructed based on the distance between vehicles and the difference in vehicle speeds; The weight function of the inter-vehicle sub-graph adjacency matrix is: , where ∆v ij represents the speed difference; ∆d ij represents the distance between vehicle i and vehicle j along the lane direction; D represents the vehicle distance threshold; V represents the speed threshold.
2. The method for making a formation decision of a connected intelligent vehicle based on nested graph reinforcement learning according to claim 1, wherein The reward function includes: a safe driving reward function, an efficiency improvement reward function, an energy saving reward function, and a passenger comfort reward function.
3. The method for making a platoon decision of a connected intelligent vehicle based on nested graph reinforcement learning according to claim 2, wherein The safe driving reward function includes: a safe distance sub-reward function and a time to collision sub-reward function; The safe distance sub-reward function includes: , Where S n-1,n represents the position difference between the leading vehicle n - 1 and the ego vehicle n at time t; D s represents the safety distance. Wherein, , Wherein, v0 represents the vehicle's own speed; v f represents the speed of the vehicle ahead; d0 represents the minimum safety distance; a max represents the maximum deceleration; τ represents the reaction time; The time to collision sub-reward function includes: , In the formula, TTC represents the time to collision; Wherein, , where, V n (t) represents the absolute speed of the host vehicle n at time t; V n-1 (t) represents the absolute speed of the leading vehicle n - 1 at time t; V n-1,n (t) represents the relative speed between the leading vehicle n - 1 and the host vehicle n at time t; S n-1,n represents the position difference between the leading vehicle n - 1 and the host vehicle n at time t.
4. The method for platoon decision-making of connected intelligent vehicles based on nested graph reinforcement learning according to claim 2, characterized in that The efficiency improvement reward function includes: a following distance sub-reward function and a vehicle speed efficiency sub-reward function; The following distance sub-reward function includes: , where φ represents the penalty weight; S n-1,n represents the position difference between the leading vehicle n - 1 and the ego vehicle n at time t; ∆d des represents the desired headway; The vehicle speed efficiency sub-reward function includes: , where, V n-1,n (t) represents the relative speed between the leading vehicle n - 1 and the host vehicle n at time t; ∆v limit represents the limit of the speed difference between the two vehicles.
Citation Information
Patent Citations
Car client requirement information cluster analysis system
CN101256646A
Vehicle speed limit control method and device, storage medium and program product
CN113635898A