Vehicle formation control method fusing graph nerve and reinforcement learning
By constructing the graph structure between vehicles and fusion graph attention, long and short-term memory, and convolutional neural network to process sensor data, the adaptability and incomplete information of existing vehicle fleet control in complex environments is solved, and efficient and safe vehicle fleet control is achieved.
Patent Information
- Application Number
- CN202510695121.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-07-22
AI Technical Summary
The existing vehicle formation control methods are difficult to adapt to individual vehicles' differentiated behaviors in complex dynamic traffic environments, lack perception of environmental obstacles and traffic spaces, and there are communication bottlenecks and incomplete information in multi-vehicle systems.
A graph structure centered on each vehicle is constructed, a graph attention network is used to weight neighbor node information, and a long and short-term memory network and a convolutional neural network are used to process sensor data, and aggregation characteristics, environmental characteristics and vehicle state are integrated, and the policy network is optimized through a multi-agent depth deterministic strategy gradient algorithm.
It improves the accuracy and flexibility of local interaction between vehicles, enhances the adaptability to dynamic traffic environments, improves the convergence efficiency and control accuracy of the policy network, and realizes safe and reliable decision-making in complex scenarios.
Smart Images

Figure CN120356322A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle platoon control, and particularly relates to a vehicle platoon control method integrating graph neural network and reinforcement learning. Background Art
[0002] With the rapid development of intelligent transportation systems and autonomous driving technologies, vehicle platoon driving, as a key technology to improve road traffic efficiency, reduce energy consumption, and enhance driving safety, has received extensive attention. Vehicle platoon control technology aims to enable multiple vehicles to maintain a specific formation relationship during driving, and through information collaboration and behavior coordination, achieve intelligent and efficient traffic organization.
[0003] Existing vehicle platoon control methods mainly include rule-based control methods, centralized control methods, and distributed cooperative control methods. Rule-based methods rely on preset control laws, such as constant spacing or time-interval control rules, but have poor adaptability to complex dynamic traffic environments. Centralized methods rely on a central controller to uniformly plan and allocate tasks. Although better performance can be obtained, there are problems such as communication bottlenecks and single-point failures in actual multi-vehicle systems. In distributed methods, each vehicle makes autonomous decisions based on local perception information, which has stronger system robustness, but also faces challenges such as incomplete information and local optimality.
[0004] In the prior art, Chinese Patent CN118082805A discloses an interconnected intelligent vehicle platoon decision-making method based on nested graph reinforcement learning, belonging to the technical fields of vehicle networking and autonomous driving. The method includes: S1, collecting vehicle state information within the platoon and processing the state information to obtain an inter-platoon nested graph and an intra-vehicle nested graph; S2, using a feature extraction network to extract features from the inter-platoon nested graph and the intra-vehicle nested graph, and obtaining the actions of each intelligent vehicle based on the extracted features; the feature extraction network includes: a graph attention layer, a fully connected layer, and an activation layer, and a multi-head attention mechanism is introduced to improve the feature extraction ability; S3, optimizing the actions using a reward function to obtain an intelligent vehicle decision-making method.
[0005] However, this method has the following limitations: the node features are mainly static quantities such as average speed and relative position, and cannot fully reflect the differential behaviors of individual vehicles in a dynamic environment; and there is a lack of modeling means for perceiving environmental obstacles, lanes, passing spaces, etc. Relying only on structured state variables, it is difficult to adapt to complex road scenarios. There is no modeling of historical states through a time series model, and there is a lack of understanding of dynamic behavior trends. Real perception means such as vision or radar are not introduced, and the decision-making depends on structured data simulation, making it difficult to be directly transferred to actual autonomous driving scenarios for use. Summary of the Invention
[0006] The object of the present invention is to provide a vehicle formation control method that integrates graph neural networks and reinforcement learning to overcome the defects of the above-mentioned existing technologies.
[0007] The object of the present invention can be achieved by the following technical solutions:
[0008] The present invention provides a vehicle formation control method that integrates graph neural networks and reinforcement learning, including the following steps:
[0009] For each vehicle in the formation, a graph structure is constructed based on the observation data collected by its sensors. Each node in the graph structure corresponds to a vehicle, and the node features include the relative speed, relative position, heading angle, acceleration of neighboring vehicles relative to the current vehicle, and identification information indicating whether they belong to the same fleet.
[0010] The constructed graph structure is processed using a graph attention network, and the information of neighboring nodes is weighted through an attention mechanism to extract the aggregated feature information of neighbors at the current moment.
[0011] The aggregated feature information of neighbors at multiple historical time steps is input into a long short-term memory network to output the predicted aggregated feature representation of neighbors at the next moment.
[0012] A two-dimensional reflection image of the surrounding environment is obtained through the radar sensor carried by the current vehicle and input into a convolutional neural network CNN to extract the environmental space feature map, including obstacles, free areas, and passing paths.
[0013] The state information of the current vehicle is collected, including absolute speed, absolute position coordinates, distances from the left and right lanes, traffic light status, front vehicle blocking identification, and path point information.
[0014] The aggregated feature representation of neighbors at the next moment, the environmental feature map, and the state information of the current vehicle are fused to construct a joint state representation.
[0015] The joint state representation is input into the policy network of the current vehicle to output the control actions of the current vehicle, including steering angle, acceleration, or braking amount, to achieve vehicle formation control.
[0016] Interact with the environment according to the control actions to obtain an immediate reward, which is calculated by a joint reward function composed of traffic efficiency, formation stability, and driving safety.
[0017] According to the immediate rewards of each vehicle in the vehicle formation, the multi-agent deep deterministic policy gradient algorithm MADDPG is used for policy training, and a mechanism combining centralized training and decentralized execution is used to optimize the policy network, graph attention network, long short-term memory network, and convolutional neural network of each vehicle.
[0018] Further, the observation data includes the relative speed, relative position, identification information on whether it belongs to the same formation, heading angle, acceleration, and current traffic state information of neighboring vehicles within the sensor observation range of the current vehicle.
[0019] Further, constructing a graph structure based on the observation data collected by its sensors specifically includes:
[0020] For each vehicle in the convoy, iterate in sequence. Using the current vehicle as the central node in the graph structure, construct adjacent nodes based on the information of neighboring vehicles detected within the sensor observation range of the current vehicle, forming a one-to-many graph structure, where each adjacent node corresponds to a neighboring vehicle; the features of the nodes in the graph structure include the relative position, relative speed, heading angle difference, acceleration of the neighboring vehicle relative to the current vehicle, and the identification of whether they belong to the same convoy; the feature of the central node in the graph structure is the self-vehicle state information of the current vehicle, including absolute position, absolute speed, heading angle, and acceleration.
[0021] Further, processing the constructed graph structure using a graph attention network specifically includes:
[0022] The constructed graph structure is used as the input of the graph attention network GAT, and the following attention weighting method is used to aggregate the features of adjacent nodes to obtain the neighbor aggregation feature representation
[0023]
[0024] Among them, is the central node, that is, the neighbor aggregation feature representation of the current vehicle; h j is the feature of neighbor node j; W is the linear mapping weight matrix; is the set of neighbor nodes of the current vehicle; α ij is the attention weight of neighbor vehicle j to current vehicle i, defined as:
[0025]
[0026] Among them, is the parameter vector of the attention mechanism, || represents vector concatenation, and LeakyReLU is the activation function.
[0027] Further, inputting the neighbor aggregation feature information of multiple historical time steps into a long short-term memory network and outputting the predicted neighbor aggregation feature representation of the next moment specifically includes:
[0028] Collect the neighbor aggregation feature representations of the current vehicle within continuous T historical time steps, denoted as the sequence and use it as the input to the long short-term memory network LSTM;
[0029] The LSTM network models the temporal dependencies of the sequence, updates the internal memory state sequentially, and outputs the predicted neighbor aggregation feature representation at the next moment.
[0030] Furthermore, obtaining a two-dimensional reflection image of the surrounding environment through the radar sensor carried by the current vehicle and inputting it into a convolutional neural network CNN to extract the environmental space feature map specifically includes:
[0031] The current vehicle obtains a two-dimensional reflection image of the surrounding environment at the current moment through the millimeter-wave radar or lidar carried, which is represented as the input image tensor I t , including obstacle reflection intensity, free area distribution, and lane geometry information;
[0032] The input image tensor I t After preprocessing, including normalization and resampling, it is input into a multi-layer convolutional neural network CNN. The CNN network structure includes several convolutional layers, pooling layers, and non-linear activation function layers to extract spatial features of different scales and semantic levels. Its extraction process includes the following steps:
[0033] The convolutional layer extracts spatial features:
[0034] F (l) = σ(W (l) *F (l-1) + b (l) )
[0035] Among them, F (l) is the feature map output of the l-th layer; F (0) = I t ; W (l) , b (l) are the convolutional kernel parameters and bias terms of the l-th layer respectively, σ is the activation function, and * represents the convolution operation;
[0036] The pooling layer performs dimensionality reduction processing:
[0037] F (l+1) = Pooling(F (l) )
[0038] Among them, Pooling is the pooling process;
[0039] After multiple convolutional and pooling processes, the final output environmental space feature map is represented as:
[0040]
[0041] Among them, Denote the environmental space feature vector extracted at the current moment, which describes the obstacle distribution, lane structure and traffic space information around the vehicle, f CNN Denote the mapping function of the entire CNN network structure.
[0042] Further, the fusion of the neighbor aggregation feature representation at the next moment, the environmental feature map and the state information of the current vehicle to construct a joint state representation specifically includes:
[0043] The predicted neighbor aggregation feature representation at the next moment The environmental feature map representation extracted by the convolutional neural network And the state information vector of the current vehicle Are concatenated and fused to construct a joint state representation For input into the policy network for control decision-making:
[0044]
[0045] Among them, Is the joint state representation vector of the current vehicle i at time t, and Concat is the vector concatenation operation.
[0046] Further, the immediate reward is:
[0047]
[0048] Among them, Represents the immediate reward obtained by the current vehicle i at time t, Represents the traffic efficiency reward term, Represents the formation stability reward term, Represents the driving safety reward term, ω1, ω2, ω3 are weighting coefficients;
[0049] The traffic efficiency reward term Is:
[0050]
[0051] Among them, Is the actual speed of the current vehicle i at time t, v target Is the target formation speed, that is
[0052]
[0053] Among them, Is the speed of the leader vehicle, v roadlimit Is the road speed limit;
[0054] The formation stability reward term Is:
[0055]
[0056] Among them, is the set of neighbor vehicles belonging to the same formation as vehicle i, are the absolute positions of vehicle i and vehicle j respectively, is the expected relative position of vehicle i and vehicle j set in the formation mission;
[0057] The driving safety reward item is:
[0058]
[0059] Among them, is a collision event indicator function, which is 1 when a collision occurs and 0 otherwise; is a hard braking event indicator function, which is 1 when a hard braking occurs and 0 otherwise; is a too-close penalty term:
[0060]
[0061] Among them, N is the total number of vehicles in the fleet, is the set of neighbor vehicles observed by the i-th vehicle at time t, is the distance between the i-th vehicle and neighbor vehicle j at time t, d safe is the set minimum safety distance threshold, is an indicator function, which is 1 when the condition is met and 0 otherwise.
[0062] Furthermore, the policy training is carried out using the multi-agent deep deterministic policy gradient algorithm MADDPG according to the immediate rewards of each vehicle in the vehicle formation, specifically including:
[0063] In the training stage, a centralized training and decentralized execution mechanism is adopted to construct a globally joint centralized Critic value network and multiple decentralized Actor networks; among them, each vehicle is equipped with an independent policy network Actor i Graph Attention Network GAT i Long Short-Term Memory Network LSTM i and Convolutional Neural Network CNN i to extract the local state and output the control action;
[0064] Concatenate the local observation information of each vehicle, including the neighbor aggregation feature representation, the environmental space feature map, and the vehicle's own state information, to construct the joint state representation of each vehicle, and input it into the corresponding policy network to generate control actions; input the joint state representations of all vehicles and their corresponding actions into the centralized Critic value network to calculate the state-action value function of each vehicle; the Critic value network is trained by minimizing the following mean squared error loss:
[0065]
[0066] y = r i + γQ′(s′, μ′1(s′1), …, μ′ N (s′ N ))
[0067] where are the parameters of the Critic value network, Q i (s, a1, …, a N ) is the state-action value function of vehicle i, r i is the immediate reward obtained by the i-th vehicle at the current time step, γ is the discount factor, s ′ is the joint state representation at the next moment, μ ′ 1(s1 ′ ) is the action made by the policy network under the joint state representation at the next moment;
[0068] The policy network is updated by maximizing the following expected function:
[0069]
[0070] where J(μ i ) is the objective function of the policy network of vehicle i, denotes the expectation, a j denotes the actions of other vehicles;
[0071] In each round of training iteration, the error signal is sequentially transmitted through the backpropagation mechanism, and the parameters of the feature extraction modules including the graph attention network GAT i , the long short-term memory network LSTM i and the convolutional neural network CNN i are jointly updated to achieve end-to-end policy optimization.
[0072] Compared with the prior art, the present invention has the following advantages:
[0073] 1. The present invention constructs a graph structure centered on each vehicle, where the node features include relative speed, position, acceleration, heading angle, and identification information indicating whether they belong to the same fleet, thereby solving the problems of coarse modeling granularity and difficulty in reflecting dynamic interactions between individuals in existing nested graphs or global graphs. This approach significantly improves the accuracy and flexibility of local interaction modeling between vehicles, enabling the system to better adapt to dynamic and complex traffic environments, and having stronger generalization ability and real-time performance.
[0074] 2. The present invention uses a graph attention network to perform weighted aggregation on the adjacent node features, dynamically allocating attention weights according to the importance of neighboring vehicles to the current vehicle, thereby solving the problem that traditional graph neural networks treat neighboring nodes equally and lack difference in information extraction. This mechanism can focus on key neighbor information, improve the effectiveness of state representation, and thus enhance the convergence efficiency and control accuracy of the policy network.
[0075] 3. The present invention uses a long short-term memory network (LSTM) to process the neighbor aggregation feature information of multiple historical time steps for predicting the state change trend at future moments, thereby solving the problem of insufficient modeling of behavior evolution trends in existing methods. The introduction of LSTM enhances the system's ability to understand dynamic traffic patterns and improves the temporal coherence and forward-looking of the policy.
[0076] 4. The present invention uses the two-dimensional reflection images obtained by millimeter-wave or lidar, and uses a CNN to extract the environmental space feature map, including information such as obstacles and free areas, effectively overcoming the problem of insufficient modeling ability of existing methods for unstructured environments. This feature enhancement mechanism significantly improves the model's ability to judge passage in complex scenarios, making vehicle decisions safer and more reliable.
[0077] 5. The present invention fuses the neighbor prediction features (LSTM output), the environmental space feature map (CNN output), and the vehicle's own state information to construct a unified joint state input for policy network decision-making, thereby solving the problems of scattered traditional decision input information and lack of relevance. This fusion modeling method improves the expression ability and decision-making accuracy of the policy, and helps to achieve integrated control in complex scenarios.
[0078] 6. The present invention adopts the multi-agent deep deterministic policy gradient algorithm (MADDPG), combines centralized training with decentralized execution, and jointly optimizes the policy network and feature extraction modules (GAT, LSTM, CNN) of each vehicle, thereby solving the problems of complex policy coupling and unstable training in multi-vehicle systems. This mechanism realizes efficient cooperative policy learning under global information sharing and has good scalability and engineering implementation value.
[0079] 7. The immediate reward function designed in the present invention consists of traffic efficiency, formation stability, and driving safety, systematically solving the problem of policy deviation caused by only optimizing a single objective in existing reinforcement learning methods. This multi-objective joint optimization mechanism makes the training more in line with actual driving needs, improving the rationality, stability, and safety of the overall vehicle behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 is a flowchart of the method of the present invention;
[0081] Figure 2 is a diagram of the evaluation index model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0082] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0083] This embodiment provides a vehicle formation control method integrating graph neural network and reinforcement learning, as Figure 1 shown, including the following steps:
[0084] Step S1: For each vehicle in the formation, a graph structure is constructed based on the observation data collected by its sensors. Each node in the graph structure corresponds to a vehicle, and the node features include the relative speed, relative position, heading angle, acceleration of neighboring vehicles relative to the current vehicle, and the identification information of whether they belong to the same fleet; the observation data includes the relative speed, relative position, identification information of whether they belong to the same formation, heading angle, acceleration of neighboring vehicles within the sensor observation range of the current vehicle, and the current traffic state information; the current traffic state information is parameters including traffic light signal status, traffic conditions of the lane where the vehicle is located and adjacent lanes, lane line type, presence of obstacles or traffic jams ahead, road speed limit information, intersection identification, lane change feasibility, road slope, and traffic density at the current time period. Among them, the vehicles in this embodiment are autonomous driverless vehicles.
[0085] Constructing a graph structure based on the observation data collected by its sensors specifically includes:
[0086] For each vehicle in the fleet, iterate sequentially. Use the current vehicle as the central node in the graph structure, and construct adjacent nodes based on the information of neighboring vehicles detected within the sensor observation range of the current vehicle, forming a one-to-many graph structure, where each adjacent node corresponds to a neighboring vehicle; the features of the nodes in the graph structure include the relative position, relative speed, heading angle difference, acceleration of the neighboring vehicle relative to the current vehicle, and the identifier of whether it belongs to the same fleet; the features of the central node in the graph structure are the ego-vehicle state information of the current vehicle, including the absolute position, absolute speed, heading angle, and acceleration; the absolute position and absolute speed refer to the physical space position and motion state of the vehicle in the global coordinate system, usually obtained by fusing high-precision GPS and an inertial navigation unit (IMU). Among them, the absolute position is the two-dimensional or three-dimensional coordinate value of the vehicle in the geographical coordinate system or the map coordinate system, and the absolute speed is the speed vector of the vehicle in the global coordinate system, reflecting the overall motion trend and operating speed of the vehicle.
[0087] Step S2: Process the constructed graph structure using a graph attention network, and extract the neighbor aggregation feature information at the current moment by weighting the neighbor node information through the attention mechanism. Specifically, it includes:
[0088] Take the constructed graph structure as the input of the graph attention network GAT, and aggregate the adjacent node features using the following attention weighting method to obtain the neighbor aggregation feature representation
[0089]
[0090] Among them, is the central node, that is, the neighbor aggregation feature representation of the current vehicle; h j is the feature of neighbor node j; W is the linear mapping weight matrix; is the set of neighbor nodes of the current vehicle; α ij is the attention weight of neighbor vehicle j to current vehicle i, defined as:
[0091]
[0092] Among them, is the parameter vector of the attention mechanism, || represents vector concatenation, and LeakyReLU is the activation function.
[0093] In step S2, a graph attention network (GAT) is used to process the constructed graph structure, aiming to address the deficiencies in neighbor information fusion in traditional vehicle platoon control. In the existing technology, vehicles usually perform information fusion through fixed weights or simple averaging of neighbor features, resulting in the inability to distinguish the importance of neighbor vehicles for the current vehicle's control decision-making and a lack of adaptability to dynamic environments and complex traffic conditions. This uniform weighting method is prone to introducing irrelevant or noisy information, affecting the stability and effectiveness of platoon control. The graph attention network enables the dynamic adjustment of the feature contribution of each neighbor node to the central vehicle by introducing an attention mechanism. Specifically, after concatenating the features of the central node and neighbor nodes, GAT calculates the attention weights through a linear mapping with LeakyReLU activation and normalizes all neighbor weights using softmax to obtain the relative importance weights of neighbor nodes for the current vehicle. Subsequently, these weights are used to weight the linear transformation results of neighbor node features to achieve differential fusion of neighbor information. The principle of this design is to automatically identify the neighbor vehicle information that is more critical for the current vehicle's decision-making through a learning mechanism, weaken the information with less influence or irrelevance, and make the aggregated neighbor features more accurate and representative. The technical effect is reflected in the ability to enhance the discriminability and dynamic adaptability of feature expression, and strengthen the perception and response capabilities of platooning vehicles to complex and changing traffic environments. In addition, the attention mechanism makes the information fusion process trainable and flexible, and can be combined with subsequent deep reinforcement learning strategies for end-to-end optimization, improving the performance and stability of the overall platoon control system.
[0094] Step S3: Input the neighbor aggregation feature information of multiple historical time steps into a long short-term memory network to output the predicted neighbor aggregation feature representation at the next moment, specifically including:
[0095] Collect the neighbor aggregation feature representations of the current vehicle within consecutive T historical time steps, denoted as the sequence and input it into the long short-term memory network LSTM as the input;
[0096] The LSTM network models the temporal dependence relationship of this sequence, sequentially updates the internal memory state, and outputs the predicted neighbor aggregation feature representation at the next moment
[0097] In step S3, the neighbor aggregation feature information of multiple historical time steps is input into a long short-term memory network (LSTM) for predicting the neighbor aggregation feature representation at the next moment. This is mainly to solve the problem of the lack of effective modeling of the dynamic changes in time series in traditional methods. In the prior art, vehicle formation control often only relies on the neighbor state information at the current moment, ignoring the temporal continuity and dynamic changes of the neighbor vehicle behavior, resulting in insufficient prediction ability for future states and affecting the foresight and stability of control strategies. As a recurrent neural network specifically designed for processing time series data, the LSTM network can capture the dependency relationships of neighbor aggregation features in the time dimension, effectively memorize and forget historical information, and avoid the vanishing gradient problem in traditional recurrent neural networks. By inputting the neighbor aggregation feature sequence of continuous T time steps, the LSTM can learn the laws of neighbor vehicle states changing over time, dynamically update the internal memory state, and thus accurately predict the neighbor feature representation at the next moment. The principle of this design is to utilize the powerful modeling ability of the LSTM for time series data to achieve the temporal dependency modeling of neighbor vehicle behavior and the prediction of future states, enhancing the vehicle's perception and understanding of the surrounding dynamic environment. The technical effect is manifested as improving the accuracy of neighbor state prediction, enhancing the predictability of control strategies, helping the vehicle better adapt to complex and changeable traffic conditions, and improving the collaborative efficiency and driving safety of the formation.
[0098] Step S4: Obtain a two-dimensional reflection image of the surrounding environment through the radar sensor carried by the current vehicle, and input it into a convolutional neural network CNN to extract the environmental space feature map, including obstacles, free areas, and passing paths. Specifically, it includes:
[0099] The current vehicle obtains a two-dimensional reflection image of the surrounding environment at the current moment through the millimeter-wave radar or lidar carried, denoted as the input image tensor I t , which includes obstacle reflection intensity, free area distribution, and lane geometric structure information;
[0100] The input image tensor I t After preprocessing, including normalization and resampling, it is input into a multi-layer convolutional neural network CNN. The CNN network structure includes several convolutional layers, pooling layers, and non-linear activation function layers to extract spatial features of different scales and semantic levels. The extraction process includes the following steps:
[0101] The convolutional layer extracts spatial features:
[0102] F (l) =σ(W (l) *F (l-1) +b (l) )
[0103] where F (l) is the feature map output of the l-th layer; F(0) = I t ; W (l) , b (l) are the convolution kernel parameters and bias terms of the l-th layer respectively, σ is the activation function, and * represents the convolution operation;
[0104] Dimensionality reduction processing of the pooling layer:
[0105] F (l+1) = Pooling(F (l) )
[0106] where Pooling is the pooling process;
[0107] After multiple layers of convolution and pooling processing, the final output environmental space feature map is expressed as:
[0108]
[0109] where represents the environmental space feature vector extracted at the current moment, describing the obstacle distribution, lane structure, and passing space information around the vehicle, and f CNN represents the mapping function of the entire CNN network structure.
[0110] In step S4, a two-dimensional reflection image of the surrounding environment is obtained by using the millimeter-wave radar or lidar mounted on the current vehicle, and used as an important data input for perceiving the environment. In the prior art, vehicle environment perception mostly relies on the simple processing of traditional sensor data, and it is difficult to fully extract complex spatial features in the environment, such as obstacle positions, free areas, and lane geometric structures, resulting in an incomplete and inaccurate understanding of the surrounding environment by the vehicle, thus limiting the effectiveness and safety of the formation control strategy. By using deep learning methods to accurately capture complex environmental features, the vehicle's perception ability of the surrounding environment is improved, which supports the path planning and safe driving of platoon vehicles in complex roads and changing environments. At the same time, the environmental features extracted based on CNN provide rich and reliable spatial information for subsequent control strategies, which helps to achieve more accurate and robust vehicle platoon control and enhance the safety and stability of driving.
[0111] Step S5: Collect the status information of the current vehicle, including the absolute speed, absolute position coordinates, distances from the left and right lanes, traffic light status, front vehicle blocking identification, and waypoint information; among them, the front vehicle blocking identification is a binary indication signal detected by the vehicle's front sensor (such as radar or camera) indicating whether the front vehicle blocks or affects the current vehicle's driving path. If there is a front vehicle blocking, the identification is 1, otherwise it is 0; the waypoint information is a series of key point coordinates on the planned path of the current vehicle, and these key points define the expected driving trajectory of the vehicle, including the curvature, turning radius, and target position of the path, helping the vehicle to accurately follow the predetermined route during platooning and achieve smooth and coordinated platooning control.
[0112] Step S6: Fuse the aggregated neighbor feature representation at the next moment, the environmental feature map, and the status information of the current vehicle to construct a joint status representation, specifically including:
[0113] The predicted aggregated neighbor feature representation at the next moment The environmental feature map representation extracted by the convolutional neural network And the status information vector of the current vehicle Perform concatenation fusion to construct a joint status representation For input into the policy network for control decision-making:
[0114]
[0115] Wherein, is the joint status representation vector of the current vehicle i at time t, and Concat is the vector concatenation operation.
[0116] Step S7: Input the joint status representation into the policy network of the current vehicle, and output the control actions of the current vehicle, including the steering angle, acceleration, or braking amount, to achieve platooning control of the vehicle;
[0117] Step S8: Interact with the environment according to the control actions to obtain an immediate reward, which is calculated by a joint reward function composed of traffic efficiency, formation stability, and driving safety;
[0118] The immediate reward is:
[0119]
[0120] Wherein, represents the immediate reward obtained by the current vehicle i at time t, represents the traffic efficiency reward term, represents the formation stability reward term, represents the driving safety reward term, and ω1, ω2, and ω3 are weighting coefficients.
[0121] Traffic efficiency reward item is:
[0122]
[0123] Among them, is the actual speed of the current vehicle i at time t, v target is the target formation speed, that is
[0124]
[0125] Among them, is the speed of the leading vehicle, v roadlimit is the road speed limit;
[0126] Formation stability reward item is:
[0127]
[0128] Among them, is the set of neighbor vehicles belonging to the same formation as vehicle i, are the absolute positions of vehicles i and j respectively, is the expected relative position of vehicles i and j set in the formation task;
[0129] Driving safety reward item is:
[0130]
[0131] Among them, is the collision event indicator function, which is 1 when a collision occurs and 0 otherwise; is the hard braking event indicator function, which is 1 when a hard braking occurs and 0 otherwise; is the too-close penalty term:
[0132]
[0133] Among them, N is the total number of vehicles in the fleet, is the set of neighbor vehicles observed by the i-th vehicle at time t, is the distance between the i-th vehicle and neighbor vehicle j at time t, d safe is the set minimum safety distance threshold, is the indicator function, which is 1 when the condition is satisfied and 0 otherwise.
[0134] Step S8 comprehensively evaluates and optimizes the vehicle platoon behavior by designing an immediate reward mechanism based on control actions and environmental interactions, taking into account traffic efficiency, formation stability, and driving safety. First, a traffic efficiency reward term is introduced to make the actual speed of the vehicle as close as possible to the target platoon speed. The target speed combines the speed of the lead vehicle and the road speed limit, ensuring that the platoon can drive efficiently and comply with road regulations, solving the problem in the prior art that the platoon speed is difficult to dynamically adapt to environmental changes, thereby improving the overall traffic flow. Second, the formation stability reward term promotes the vehicle to maintain the predetermined formation structure by punishing the deviation of the relative position between vehicles from the desired position, effectively avoiding the phenomenon of loose or misaligned formations, solving the technical problem in traditional platoon control that the formation is vulnerable to disturbances and difficult to maintain stability, and improving the coordination consistency and control accuracy of the platoon. Finally, the driving safety reward term comprehensively considers the punishment for collision events, hard braking behaviors, and overly close distances. By defining safety indicators, it strengthens the vehicle to maintain a safe distance and drive smoothly in the platoon, solving the problem in the prior art that there is a lack of effective constraints on safety hazards in complex traffic environments, thereby greatly improving the driving safety guarantee. This combined reward function realizes the multi-objective optimization of the behavior of platoon vehicles by weighted fusion of multi-dimensional performance indicators, taking into account efficiency, stability, and safety, with significant technical effects, and can effectively improve the practicability and robustness of the autonomous driving platoon system, meeting the dynamic adaptation requirements in complex traffic scenarios.
[0135] Step S9: Use the multi-agent deep deterministic policy gradient algorithm MADDPG to perform policy training according to the immediate rewards of each vehicle in the vehicle platoon, and use a mechanism that combines centralized training and decentralized execution to optimize the policy network, graph attention network, long short-term memory network, and convolutional neural network of each vehicle, specifically including:
[0136] In the training stage, a centralized training and decentralized execution mechanism is adopted to construct a globally joint centralized Critic value network and multiple decentralized Actor networks; among them, each vehicle is equipped with an independent policy network Actori, a graph attention network GAT i , a long short-term memory network LSTM i and a convolutional neural network CNN i for extracting the local state and outputting control actions;
[0137] Concatenate the local observation information of each vehicle, including the neighbor aggregation feature representation, the environmental space feature map, and the vehicle's own state information, to construct the joint state representation of each vehicle, and input it into the corresponding policy network to generate control actions; input the joint state representations of all vehicles and their corresponding actions into the centralized Critic value network to calculate the state-action value function of each vehicle; the Critic value network is trained by minimizing the following mean square error loss:
[0138]
[0139] y = r i + γQ′(s′, μ′1(s′1),..., μ′ N (s′ N ))
[0140] where are the parameters of the Critic value network, Q i (s, a1,…, a N ) is the state-action value function of vehicle i, r i is the immediate reward obtained by the i-th vehicle at the current time step, γ is the discount factor, s′ is the joint state representation at the next moment, and μ′1(s′1) is the action made by the policy network under the joint state representation at the next moment;
[0141] The policy network is updated by maximizing the following expected function:
[0142]
[0143] where J(μ i ) is the objective function of the policy network of vehicle i, denotes the expectation, and a j denotes the actions of other vehicles;
[0144] In each round of training iteration, the error signal is sequentially transmitted through the backpropagation mechanism, and the joint update includes the parameters of the feature extraction modules including the graph attention network GAT i , the long short-term memory network LSTM i and the convolutional neural network CNN i to achieve end-to-end policy optimization.
[0145] Step S9 uses the Multi-Agent Deep Deterministic Policy Gradient algorithm (MADDPG) for vehicle formation strategy training, which solves the limitations of traditional single-agent methods in multi-vehicle cooperative control and realizes the cooperative optimization of each vehicle in the formation. By adopting a mechanism that combines centralized training and decentralized execution, it fully utilizes global information to improve the training effect, and at the same time ensures that each vehicle can make independent decisions during actual application, overcoming the challenges brought by information asymmetry and instability in multi-agent environments. A centralized Critic value network that constructs a global joint is used, enabling the training process to evaluate the joint impact of the overall state and all vehicle actions, and solving the problem that a single vehicle relying only on local information leads to difficult training convergence and poor policy performance. The decentralized Actor network ensures that each vehicle can generate control actions independently relying only on local observations during execution, improving the robustness and real-time response ability of the system.
[0146] In addition, by integrating the Graph Attention Network (GAT), Long Short-Term Memory Network (LSTM), and Convolutional Neural Network (CNN) for feature extraction, it strengthens the perception ability of the dynamic relationships of neighboring vehicles, historical temporal dependencies, and environmental spatial features, and improves the understanding and adaptation ability of the policy network to complex traffic scenarios. Through the backpropagation of the error signal, end-to-end joint optimization of the feature extraction module and the policy network is achieved, ensuring the coordination and cooperation between different network modules and improving the overall learning efficiency and policy performance. This method effectively solves the problems of high-dimensional state space, complex dynamic environment, and cooperative decision-making in multi-vehicle formation control, realizes the efficient and stable training of vehicle formation strategies, significantly improves the traffic efficiency, safety, and formation stability of vehicle cooperative driving, and has good practical application prospects.
[0147] This embodiment also provides a vehicle formation control system based on deep reinforcement learning, as Figure 2 shown, including:
[0148] A sensor module for collecting information such as the relative position, speed, heading angle, acceleration, etc. of neighboring vehicles around the vehicle, as well as the two-dimensional radar reflection image of the surrounding environment and the vehicle's own state information;
[0149] A data processing module that constructs a graph structure between vehicles based on the collected sensor data and generates neighbor vehicle node features;
[0150] A feature extraction module, including a Graph Attention Network (GAT), Long Short-Term Memory Network (LSTM), and Convolutional Neural Network (CNN), which are respectively used to extract neighbor aggregation features, time series features, and environmental spatial features;
[0151] The control strategy module, based on the multi-agent deep reinforcement learning algorithm, adopts a centralized training and decentralized execution mechanism to jointly optimize the policy networks of each vehicle and achieve the cooperative control of vehicle formation;
[0152] The reward calculation module calculates the immediate reward based on the traffic efficiency, formation stability, and driving safety of the vehicle, serving as the feedback signal for policy training;
[0153] Through the collaborative work of each module, the system can achieve the efficient, safe, and stable driving of vehicle formation in a complex dynamic environment.
[0154] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.
[0155] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A vehicle formation control method integrating graph neural network and reinforcement learning, characterized in that, Including the following steps: For each vehicle in the formation, construct a graph structure based on the observation data collected by its sensors; Process the constructed graph structure using a graph attention network, weight the neighbor node information through the attention mechanism, and extract the neighbor aggregation feature information at the current moment; Input the neighbor aggregation feature information of multiple historical time steps into a long short-term memory network, and output the predicted neighbor aggregation feature representation at the next moment; Obtain a two-dimensional reflection image of the surrounding environment through the radar sensor carried by the current vehicle, and input it into a convolutional neural network to extract the environmental spatial feature map; Collect the state information of the current vehicle, including absolute speed, absolute position coordinates, distances from the left and right lanes, traffic light status, front vehicle blocking identification, and path point information; Fuse the neighbor aggregation feature representation at the next moment, the environmental feature map, and the state information of the current vehicle to construct a joint state representation; Input the joint state representation into the policy network of the current vehicle, and output the control actions of the current vehicle, including steering angle, acceleration, or braking amount, to achieve formation control of the vehicle; Interact with the environment according to the control actions to obtain an immediate reward, which is calculated by a joint reward function composed of traffic efficiency, formation stability, and driving safety; Perform policy training using the multi-agent deep deterministic policy gradient algorithm MADDPG according to the immediate rewards of each vehicle in the vehicle formation, and use a mechanism combining centralized training and decentralized execution to optimize the policy network, graph attention network, long short-term memory network, and convolutional neural network of each vehicle.
2. The vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that The observation data includes the relative speed, relative position, identification information of whether it belongs to the same formation, heading angle, acceleration, and current traffic state information of neighbor vehicles within the sensor observation range of the current vehicle.
3. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that Each node in the graph structure corresponds to a vehicle, and the node features include the relative speed, relative position, heading angle, acceleration of neighbor vehicles relative to the current vehicle, and identification information of whether they belong to the same fleet; The constructing of the graph structure based on the observation data collected by its sensors specifically includes: For each vehicle in the fleet, perform iterations in sequence. Take the current vehicle as the central node in the graph structure, and construct adjacent nodes based on the neighbor vehicle information detected within the sensor observation range of the current vehicle to form a one-to-many graph structure, where each adjacent node corresponds to a neighbor vehicle; the features of the nodes in the graph structure include the relative position, relative speed, heading angle difference, acceleration of neighbor vehicles relative to the current vehicle, and identification of whether they belong to the same fleet; the feature of the central node in the graph structure is the self-vehicle state information of the current vehicle, including absolute position, absolute speed, heading angle, and acceleration.
4. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that, The processing of the constructed graph structure using a graph attention network specifically includes: The constructed graph structure is used as the input of the Graph Attention Network (GAT). The following attention weighting method is adopted to aggregate the adjacent node features to obtain the neighbor aggregation feature representation Among them, is the central node, that is, the neighbor aggregation feature representation of the current vehicle; h j is the feature of neighbor node j; W is the linear mapping weight matrix; is the set of neighbor nodes of the current vehicle; α ij is the attention weight of neighbor vehicle j to the current vehicle i, defined as: Among them, is the parameter vector of the attention mechanism, || represents vector concatenation, and LeakyReLU is the activation function.
5. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that, The inputting of the neighbor aggregation feature information of multiple historical time steps into a long short-term memory network and outputting the predicted neighbor aggregation feature representation at the next moment specifically includes: Collect the neighbor aggregation feature representations of the current vehicle within consecutive T historical time steps, denoted as a sequence and feed it as input into the long short-term memory network LSTM; The LSTM network models the temporal dependencies of the sequence, updates the internal memory state sequentially, and outputs the predicted neighbor aggregation feature representation at the next moment.
6. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that, The obtaining of a two-dimensional reflection image of the surrounding environment through the radar sensor carried by the current vehicle and inputting it into a convolutional neural network CNN to extract the environmental spatial feature map specifically includes: The current vehicle obtains a two-dimensional reflection image of the surrounding environment at the current moment through the mounted millimeter-wave radar or lidar, which is represented as the input image tensor I t , including obstacle reflection intensity, free area distribution, and lane geometry information; The input image tensor I t After preprocessing, including normalization and resampling, it is input into a multi-layer convolutional neural network CNN. The CNN network structure includes several convolutional layers, pooling layers, and non-linear activation function layers to extract spatial features at different scales and semantic levels. The extraction process includes the following steps: The convolutional layer extracts spatial features: F (l) = σ(W (l) * F (l-1) + b (l) ) Among them, F (l) is the output of the feature map of the l-th layer; F (0) = I t ; W (l) , b (l) are the convolution kernel parameters and bias terms of the l-th layer respectively, σ is the activation function, and * represents the convolution operation; The pooling layer performs dimensionality reduction processing: F (l+1) = Pooling(F (l) ) Among them, Pooling is the pooling processing; After multiple layers of convolution and pooling processing, the final output environmental spatial feature map is expressed as: Among them, represents the environmental space feature vector extracted at the current moment, describing the distribution of obstacles around the vehicle, lane structure, and passing space information, f CNN represents the mapping function of the entire CNN network structure.
7. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that The fusion of the neighbor aggregation feature representation at the next moment, the environmental feature map, and the state information of the current vehicle to construct a joint state representation specifically includes: The predicted aggregated neighbor feature representation at the next moment The environmental feature map representation extracted by the convolutional neural network And the state information vector of the current vehicle Are concatenated and fused to construct a joint state representation For input into the policy network for control decision-making: Among them, is the joint state representation vector of the current vehicle i at time t, and Concat is the vector concatenation operation.
8. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that, The immediate reward is: Among them, represents the immediate reward obtained by the current vehicle i at time t, represents the reward item for traffic efficiency, represents the reward item for formation stability, represents the reward item for driving safety, and ω1, ω2, and ω3 are weighting coefficients.
9. The vehicle platoon control method integrating graph neural network and reinforcement learning according to claim 8, characterized in that The traffic efficiency reward item is as follows: Among them, is the actual speed of the current vehicle i at time t, v target is the target formation speed, that is Among them, is the speed of the leading vehicle, v roadlimit is the speed limit of the road; The formation stability reward item is as follows: Among them, is the set of neighbor vehicles belonging to the same formation as vehicle i, are the absolute positions of vehicle i and vehicle j respectively, is the expected relative position of vehicle i and vehicle j set in the formation mission; The driving safety reward items are as follows: Among them, is a collision event indication function, which is 1 when a collision occurs and 0 otherwise; is an emergency braking event indication function, which is 1 when an emergency braking occurs and 0 otherwise; is a penalty term for being too close: where N is the total number of vehicles in the fleet, is the set of neighboring vehicles observed by the i-th vehicle at time t, is the distance between the i-th vehicle and neighboring vehicle j at time t, d safe is the set minimum safety distance threshold, is the indicator function, which is 1 when the condition holds and 0 otherwise.
10. A vehicle formation control method integrating graph neural network and reinforcement learning according to claim 1, characterized in that The policy training is performed using the multi-agent deep deterministic policy gradient algorithm MADDPG according to the immediate rewards of each vehicle in the vehicle formation, specifically including: In the training phase, a centralized Critic value network that is globally united and multiple decentralized Actor networks are constructed by adopting a centralized training and decentralized execution mechanism; among them, each vehicle is equipped with an independent policy network Actor i , graph attention network GAT i , long short-term memory network LSTM i and convolutional neural network CNN i , which are used to extract local states and output control actions; The local observation information of each vehicle, including the neighbor aggregation feature representation, the environmental spatial feature map, and the vehicle's own state information, is concatenated to construct the joint state representation of each vehicle, and the corresponding policy network is input to generate control actions; the joint state representations of all vehicles and their corresponding actions are input into the centralized Critic value network together to calculate the state-action value function of each vehicle; the Critic value network is trained by minimizing the following mean square error loss: y = r i + γQ′(s′, μ′1(s′1),..., μ′ N (s′ N )) Among them, are the parameters of the Critic value network, Q i (s, a1, …, a N ) is the state-action value function of vehicle i, r i is the immediate reward obtained by the i-th vehicle at the current time step, γ is the discount factor, s′ is the joint state representation at the next moment, and μ′1(s′1) is the action made by the policy network under the joint state representation at the next moment; The policy network is updated by maximizing the following expected function: Among them, J(μ i ) is the objective function of the policy network of vehicle i, denotes expectation, a j denotes the actions of other vehicles; In each training iteration, the error signal is sequentially propagated through the backpropagation mechanism to jointly update the parameters of the feature extraction modules including the graph attention network GAT i , the long short-term memory network LSTM i and the convolutional neural network CNN i to achieve end-to-end policy optimization.
Citation Information
Patent Citations
Networked vehicle formation decision-making method based on nested graph reinforcement learning
CN118082805A
Cited By
Highway charging station intelligent guiding method based on multi-agent reinforcement learning
CN120952292A
Unmanned vehicle cooperative control method based on gating-hypergraph neural network
CN121069797A
An unmanned vehicle cooperative control method based on a gated-hypergraph neural network
CN121069797B