A Vehicle Trajectory Prediction Method and System Incorporating Large-Core Convolution and Attention Mechanism
Through the vehicle trajectory prediction method that integrates large-core convolution and attention mechanism, the problem of insufficient vehicle trajectory prediction accuracy and real-time performance in the prior art is solved, and more accurate multimodal trajectory prediction is achieved, especially in complex traffic environments.
Patent Information
- Application Number
- CN202510322044.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-19
AI Technical Summary
The prior art in vehicle trajectory prediction, especially in complex traffic environments, has low prediction accuracy, insufficient real-time performance, and fails to effectively capture the implicit interactions and time dependence between vehicles.
The vehicle trajectory prediction method that integrates large-core convolution and attention mechanisms is used to capture the time series features and spatial interactions of the vehicle's historical trajectory through the encoder module, multi-head spatial attention module, large-core convolution pooling module, convolution modulation module and decoder module, combined with the LSTM encoder, and realize multi-modal trajectory prediction.
It improves the accuracy and real-time performance of vehicle trajectory prediction, and can more accurately predict the future trajectory of the vehicle, especially in complex traffic scenarios, and enhances the modeling ability of spatial interaction and time dependence between vehicles.
Smart Images

Figure CN119848632B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent vehicle trajectory prediction, and specifically to a vehicle trajectory prediction method and system integrating large kernel convolution and attention mechanism. Background Art
[0002] In the field of trajectory prediction, researchers at home and abroad have proposed various trajectory prediction methods, which can be divided into three categories: physical model-based methods, traditional machine learning-based methods, and deep learning-based methods.
[0003] Physical model-based methods usually rely on the kinematic equation and dynamic equation of the vehicle, and assume that the vehicle moves at a constant speed or a constant yaw angular velocity. The advantage of this method is simple calculation, but it is obvious that regarding the vehicle movement as a constant speed or a constant acceleration or deceleration is divorced from reality, especially with poor adaptability in complex traffic environments. Traditional machine learning-based trajectory prediction methods mainly include Bayesian networks, hidden Markov models, support vector machines, and Gaussian processes. In vehicle trajectory prediction, Gaussian process methods are usually used to quantify the uncertainty of future trajectory prediction and obtain the probability distribution of multiple future trajectories to achieve multi-modal prediction. Bayesian networks, hidden Markov models, and support vector machines are often used to decompose the driving behavior of the vehicle horizontally and vertically, and then recombine the decomposed vehicle behaviors to finally achieve multi-modal prediction. Although these methods can quantify the uncertainty in trajectory prediction and achieve multi-modal trajectory prediction, they all ignore the impact of the interaction between vehicles on future trajectories.
[0004] Deep learning-based methods can not only consider physical and road-related factors, but also capture the implicit interaction between vehicles in space, thus being more suitable for more complex scenarios. Since the driving trajectory data of vehicles has time series characteristics, researchers initially adopted recurrent neural networks ( RNN ) and their variants, such as long short-term memory neural networks ( LSTM ) and gated recurrent units ( GRU ) to capture the time correlation of vehicle historical trajectory data. Although the methods based on RNN and their variants perform well in short-term domain (0 - 3 s ) trajectory prediction, their effects in long-term domain (3 - 5 s ) trajectory prediction are limited, and there is still room for improvement in capturing the implicit interaction between vehicles. With the continuous in-depth research, convolutional neural networks ( CNN ) and graph neural networks ( GNN ) have also been used to extract the hidden states in the trajectory dataset. Based on GNNThe method represents vehicles in the scene as nodes and the relationships between vehicles as connections between nodes by constructing a traffic map. It is applicable to multi-lane and traffic-congested trajectory prediction scenarios and can achieve high prediction accuracy. However, such methods rely strongly on map information.
[0005] In actual traffic scenarios, accurately predicting the future trajectories of surrounding vehicles still faces many challenges, mainly including: low prediction accuracy, weak real-time performance, implicit interaction between vehicles, and uncertainty in the future driving trajectories of vehicles, etc. Summary of the Invention
[0006] The present invention aims to provide a vehicle trajectory prediction method and system that integrates large kernel convolution and attention mechanism to solve the technical problems mentioned in the background art.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] Term Explanation:
[0009] Multi-modal Trajectory Prediction: Vehicle multi-modal trajectory prediction ( Trajectory Prediction ) refers to predicting multiple possible future positions of a vehicle based on its historical motion data (such as position, speed, acceleration, etc. information). Its main goal is to reasonably predict the future motion trajectory of the vehicle so that the vehicle can better make decision-making plans and controls;
[0010] Attention Mechanism: The attention mechanism ( Attention Mechanism ) is a technology that selectively focuses on key parts of input data by dynamically allocating different weights. It simulates the characteristics of human attention and is commonly used in fields such as natural language processing and computer vision. In deep learning, it allows the model to focus on more important parts when processing information, thereby improving prediction accuracy;
[0011] Large Kernel Convolution Pooling: Large kernel convolution pooling ( Large Kernel Convolution Pooling , LKCP ) is a method proposed in the present invention that combines multi-scale convolution kernels and pooling operations, which can effectively obtain information within the local neighborhood of vehicles and a larger receptive field, thereby modeling the spatial relationships between surrounding vehicles;
[0012] Convolution Modulation: Convolution modulation ( Convolution Modulation ) is a method that simulates and replaces the self-attention mechanism. It adjusts and transforms the input feature information through convolution operations, thereby weighting, enhancing, and changing the representation form of the features.
[0013] A vehicle trajectory prediction system integrating large-kernel convolution and attention mechanism, the system includes: an encoder module, a multi-head spatial attention module ( Spatial attention / SA ), a large-kernel convolution pooling module ( Large - kernel Convolution pooling / LKCP ), a convolutional modulation module ( Convolutional modulation / CM ), an intention recognition module and a decoder module;
[0014] A data acquisition module, configured to collect and track in real time the trajectory data and the lane information of the host vehicle and other surrounding vehicles within a preset range through an in-vehicle camera;
[0015] The encoder module is used to encode the vehicle historical trajectory-related data (historical trajectory x and y coordinates, lane information, speed, acceleration) into a shape format that can be embedded into the multi-head spatial attention module ( SA module) and LKCP module;
[0016] The multi-head spatial attention module is used to calculate the spatial correlation degree between the target vehicle and each vehicle within a preset range around it according to the calculation formula of the attention mechanism;
[0017] The large-kernel convolution pooling module is used to capture the interaction between other surrounding vehicles except the target vehicle within a preset range through convolutional layers and pooling layers with multiple different convolutional kernels;
[0018] The convolutional modulation module is used for feature extraction, and through convolution, LSTM after the encoder, it captures the temporal correlation in the vehicle historical trajectory information again;
[0019] The intention recognition module and the decoder module output the predicted trajectory of the target vehicle in the future.
[0020] As a further technical solution of the present invention, the encoder module includes a multi-layer perceptron MLP and a long short-term memory neural network LSTM ;
[0021] The multi-layer perceptron preprocesses the data. First, the vehicle historical state is linearly transformed as follows:
[0022] Formula (1): ;
[0023] Wherein: is the state information of the vehicle at time, and is the weight matrix of the linear layer, is the bias vector of the linear layer, and will be continuously updated and changed during the training process, and is the result after linear transformation;
[0024] The output of the linear layer in formula (1) is processed by the activation function to obtain the final output :
[0025] Formula (2): ;
[0026] Wherein: is the vector that can be embedded into the encoder, is an adjustable hyperparameter, usually defaulting to 1;
[0027] Since the vehicle historical trajectory data has time series characteristics, the present invention uses a long short-term memory neural network LSTM as the encoder to initially capture the hidden state in the historical trajectory and perform feature encoding. The target vehicle and surrounding vehicles are respectively encoded by LSTM the encoder, and each LSTM encoder shares weights, thereby reducing the number of network parameters, reducing the model complexity, and reducing the risk of overfitting;
[0028] The MLP output time embedding vector of the multi-layer perceptron ( LSTM is input into the encoder to obtain
[0029] the hidden state at time, and the process is as follows: ;
[0030] Formula (4): ;
[0031] Formula (5): ;
[0032] Formula (6): ;
[0033] Formula (7): ;
[0034] Formula (8): ;
[0035] Wherein: 、 、 are the outputs of the input gate, the forget gate, and the output gate respectively, 、 、 、 are all weight matrices, 、 、 、 are all weight matrices related to the hidden state at time t-1, 、 、 、 is the bias vector, is the activation function, is the vehicle at the state of the candidate cell at time, is the state of the candidate cell at time, is after the label of the vehicle output by the encoder is the hidden state of the vehicle at the historical time.
[0036] As a further technical solution of the present invention, the multi-head spatial attention module adopts a multi-head spatial attention mechanism to capture the interaction between the target vehicle and surrounding vehicles. At each moment , through the calculation formula of the attention mechanism, calculate the correlation between the target vehicle and each surrounding vehicle;
[0037] From the known the hidden state vector of the target vehicle output by the encoder at time is , the hidden state vector of the surrounding vehicles is { }, the labels of the surrounding vehicles are 1-N, respectively represent the hidden states of the surrounding vehicles numbered 1-N output by the LSTM encoder at time t. Respectively input them into the multi-head spatial attention module, and the calculation formula is as follows:
[0038] Formula (9): ;
[0039] Formula (10): ;
[0040] Formula (11): ;
[0041] Formula (12): ;
[0042] Wherein: 、 and are the query matrix, key matrix, and value matrix of the target vehicle obtained by linear layer transformation respectively, 、 and are the weight matrices of the linear transformation, 、 and are the corresponding bias terms, is the number of channels, and the output is the correlation between the target vehicle and other surrounding vehicles calculated at the moment by the first head in the multi-head attention;
[0043] In the present invention, the multi-head spatial attention module adopts four head attention layers, and finally the outputs of the four head attention layers are concatenated along the last dimension to obtain ;
[0044] Then, after passing through GLU the activation function and residual connection, layer normalization is performed to obtain the hidden state after considering the interaction between the target vehicle and the surrounding vehicles:
[0045] Formula (13): ;
[0046] Wherein, is the hidden state after considering the interaction between the target vehicle and the surrounding vehicles, LayerNorm is the layer normalization function, GLU is the activation function, is the attention obtained by concatenating the outputs of the four head attention layers along the last dimension, is the bias matrix, is the hidden state vector of the target vehicle output by the encoder at the moment.
[0047] As a further technical solution of the present invention, the large kernel convolutional pooling module is composed of three convolutional layers and an adaptive average pooling layer, and the input is ( LSTM the hidden state vector of the surrounding vehicles output by the encoder module), and the three convolutional layers are respectively 5 5 depth convolution 、7 7 dilated convolution DW-D-Conv, 1 1 convolution 1 1Conv;
[0048] First, through 5 5-depth convolutional ( ) to capture the detailed features of each surrounding vehicle in the local space;
[0049] Subsequently, 7 7 dilated convolutional layers ( DW - D - Conv ) are used to expand the receptive field, enabling the network to obtain vehicle interaction information in a larger range while maintaining the resolution;
[0050] Next, a 1 1 convolutional layer (1 1 Conv ) is used to fuse the features of different channels, integrate the information extracted by multiple convolutional layers, and enhance the feature expression ability;
[0051] Finally, through the adaptive average pooling layer ( Adaptive Average Pooling ), the extracted information is further compressed and mapped into a fixed-size feature representation for effective concatenation with the hidden features output by the multi-head spatial attention module, thereby improving the modeling ability of vehicle interaction relationships;
[0052] The above complete feature extraction process is shown as follows:
[0053] Equation (14): ;
[0054] Equation (15): ;
[0055] Equation (16): ;
[0056] Equation (17): ;
[0057] Equation (18): ;
[0058] Where: The input of the large-kernel convolutional pooling module is the hidden state vector of the surrounding vehicles output by the LSTM encoder module { }, the labels of the surrounding vehicles are 1-N, respectively representing the hidden states of the surrounding vehicles numbered 1-N output by the LSTM encoder at time t, , , and are the hidden states during the convolutional processing, , and is the convolution kernel, representing 5 respectively 5, 7 7 and 1 small local matrices of 1, , and is the network bias, is the hidden state output after feature extraction by the LKCP module.
[0059] As a further technical solution of the present invention, the convolution modulation module consists of layer normalization, 1 1 convolution, GELU activation function and 7 7 grouped convolutions. The input vector of the convolution modulation module is (concatenation of the hidden state vectors output by the multi-head spatial attention module ( SA ) and the large kernel convolution pooling module ( LKCP ));
[0060] First, is processed through the layer normalization function. Layer normalization helps to standardize the input features, ensure that the model stably learns the feature distribution at different time steps, and reduce the learning bias of the model;
[0061] Then, the results of the layer normalization processing are respectively subjected to feature extraction through two routes. The first route is processed through 1 1 convolution to simulate the linear transformation in the self-attention mechanism, reconstruct the feature representation, enhance the ability of the model to capture deep associations between time steps in the short term, and output the feature vector v , and the second route is processed through 1 1 convolution and 7 7 large kernel grouped convolutions to output the feature vector ;
[0062] Through operation, the extracted features are modulated, so as to dynamically identify and capture the key moments in the prediction process, enabling the model to more accurately predict the future trajectory;
[0063] Finally, the feature dimension of the modulated hidden features is corrected through 1 1 projection convolution to obtain the final output.
[0064] The output of the multi-head spatial attention mechanism and the output of the module are concatenated and then input into the convolution modulation module for time feature extraction. The above complete process is shown in the following formula:
[0065] Formula (19): ;
[0066] Formula (20): ;
[0067] Formula (21): ;
[0068] Formula (22): ;
[0069] Final output is the hidden state of the target vehicle's historical trajectory data generated on the basis of considering the vehicle space interaction and the time dependence of the trajectory data. The input vector of the convolutional modulation module is , is the hidden state after considering the interaction between the target vehicle and surrounding vehicles, is the hidden state output after feature extraction by the LKCP module, is the hidden state after the input vector of the convolutional modulation module is processed by layer normalization, is to through 1 1 convolution and 7 7 large kernel grouped convolutions for processing, and the output feature vector. v is to through 1 1 convolution for processing, and the output feature vector.
[0070] As a further technical solution of the present invention, the structure of the intention recognition module is mainly composed of a feed-forward layer, a normalization activation layer, and two linear classifiers;
[0071] Encode the information of the target vehicle output in the convolutional modulation module to extract all hidden state information of each time step of each sample's historical data;
[0072] Then, the extracted hidden state information is successively passed through the feed-forward layer and the normalization activation layer for feature extraction. The feed-forward layer consists of a fully connected layer ( FC ), and the fully connected layer uses ELU activation function. The normalization activation layer consists of a fully connected layer ( FC ), LayerNorm and ELU activation function;
[0073] Finally, the output of the normalization activation layer is respectively fed into the longitudinal classifier and the lateral classifier, and the weight probabilities of each longitudinal and lateral maneuver category at each future prediction time step are calculated through the fully connected layer ( FC ) and classification activation function. The longitudinal maneuver categories include vehicle acceleration, deceleration, and constant speed, and the lateral maneuver categories include vehicle left lane change, right lane change, and no lane change (lane keeping).
[0074] As a further technical solution of the present invention, the decoder module consists of two fully connected layers and a layer. The function of the last fully connected layer is to map the information encoded by the first fully connected layer and the LSTM layer from the feature space to the trajectory space. To consider the uncertainty of prediction, it is assumed that the result predicted by the decoder is considered to follow a binary Gaussian distribution. Therefore, the decoder can output the , , , and at each future prediction time step, where is the future prediction time step, . The multi-modal trajectory follows the total probability theorem, as shown in the following formula:
[0075] Formula (23): ;
[0076] Where: represents the coordinates at the future prediction time step, is all vehicle states at all times in the historical time period . The vehicle state at a certain moment in the historical time period is defined as , where represents the state quantity of the vehicle, represents the vehicle number, , where and represent the and coordinates of the target vehicle at the moment , and represent the speed and acceleration of the target vehicle at the moment , represents the lane where the target vehicle and all vehicles within a preset range around it are located at the moment , represents the vehicle type (encoded by numbers) of the target vehicle and all vehicles within a radius of 90 meters around it at the moment represents the lateral strategy, represents the longitudinal strategy, and represent the probabilities of the vehicle performing lateral strategy maneuvers and longitudinal strategy maneuvers based on the state X respectively, is the Gaussian distribution probability of the predicted future trajectory of the vehicle based on the vehicle's historical state and lateral and longitudinal maneuvers, , where , representing the parameters of the bivariate Gaussian distribution at each time step within the prediction horizon, and are respectively the expectations of the predicted trajectory points in the and directions, and are respectively the variances of the predicted trajectory points in the and directions, is the correlation coefficient, and respectively represent the x coordinate at the predicted time step and the y coordinate at the predicted time step , and respectively represent the set of Gaussian distribution parameters at the predicted time step and the set of Gaussian distribution parameters at the predicted time step .
[0077] Another object of the present invention is to provide a vehicle trajectory prediction method that combines large kernel convolution and attention mechanism, and the method includes:
[0078] S1. Real-time collect and track the trajectory data and lane information of the host vehicle and other surrounding vehicles within a preset range through an in-vehicle camera;
[0079] S2. Encode the collected vehicle historical trajectory-related data (historical trajectory x and y coordinates, lane information, speed, acceleration) into a shape format that can be embedded into the SA module and LKCP module;
[0080] S3. Calculate the spatial correlation degree between the target vehicle and each vehicle within a preset range according to the calculation formula of the attention mechanism;
[0081] S4. Capture the interaction between other surrounding vehicles except the target vehicle within a preset range through convolutional layers and pooling layers with multiple different convolutional kernels;
[0082] S5. Perform feature extraction, and capture the temporal correlation in the vehicle historical trajectory information again through convolution following the LSTM encoder;
[0083] S6. Output the predicted trajectory of the target vehicle in the future.
[0084] Compared with the prior art, the beneficial effects of the present invention are:
[0085] The present invention gives full play to LSTM , CNN and the advantages of the attention mechanism, and proposes a vehicle trajectory prediction model that combines a convolutional neural network and an attention mechanism. First, the model encodes the historical state information of the vehicle into an embeddable vector through a multi-layer perceptron and LSTM encoder, and initially captures its hidden state. Subsequently, the spatial interaction between the target vehicle and surrounding vehicles is modeled through a multi-head spatial attention mechanism. At the same time, the model creatively introduces LKCP module, an intention recognition module, and CM module. Among them, LKCP module can additionally consider the spatial interaction between surrounding vehicles, further enhancing the ability to model spatial interaction between vehicles; the intention recognition module infers the maneuver category of the vehicle based on the historical state information of the vehicle, providing a basis for multi-modal trajectory prediction; and CM module further captures the temporal dependence between each moment in the historical trajectory of the target vehicle on the basis of LSTM , enhancing the capture of hidden states in the time dimension. Finally, the decoder outputs and realizes the multi-modal trajectory prediction of the target vehicle. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 is the overall framework diagram of the vehicle trajectory prediction system in the embodiment of the present invention.
[0087] Figure 2 is the framework diagram of the large kernel convolutional pooling module in the embodiment of the present invention.
[0088] Figure 3 is the framework diagram of the convolutional modulation module in the embodiment of the present invention.
[0089] Figure 4 is the framework diagram of the intention recognition module in the embodiment of the present invention.
[0090] Figure 5 is the NGSIM dataset-based RMSE comparison result diagram of each model in the embodiment of the present invention.
[0091] Figure 6 is the NGSIM dataset-based NLL comparison result diagram of each model in the embodiment of the present invention.
[0092] Figure 7 is the visualization diagram of the first trajectory prediction result based on the NGSIM dataset in the embodiment of the present invention.
[0093] Figure 8 is the visualization diagram of the second trajectory prediction result based on the NGSIM dataset in the embodiment of the present invention. Detailed implementation mode
[0094] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0095] The purpose of the present invention is to predict the multi-modal trajectory of the target vehicle under the conditions of fully considering the interaction between vehicles in the traffic scenario and the time series characteristics of the vehicle historical trajectory.
[0096] The current vehicle trajectory prediction problems are basically to predict the future state information of the vehicle according to the historical state information of the vehicle. Assume that the length of the vehicle historical trajectory time period selected by the present invention is , and the length of the future trajectory time period to be predicted is . At a certain moment in the historical time period , the vehicle state is defined as , where represents the state quantity of the vehicle, and represents the vehicle number. The input of the prediction model in the present invention, that is, all vehicle states at all times in the historical time period , can be expressed as . Assume that the number of the target vehicle is , and the number range of other surrounding vehicles is . For a moment . For any vehicle, taking the target vehicle as an example, at the moment , its state quantity is specified as , where and respectively represent the and coordinates of the target vehicle at the moment , and respectively represent the speed and acceleration of the target vehicle at the moment , represents the lane where the target vehicle and all vehicles within a certain range around it are located at the moment , represents the vehicle type (encoded by numbers) of the target vehicle and all vehicles within a certain range around it at the moment . After prediction, the output is the trajectory of the vehicle in the time period . Assume that is any moment in the future prediction time period The state of the target vehicle at a moment is defined as The predicted state of the target vehicle within the future time period is expressed as .
[0097] The present invention provides a vehicle trajectory prediction method integrating large - kernel convolution and attention mechanism. The method includes:
[0098] S1. Real - time collect and track the trajectory data and lane information of the host vehicle and other surrounding vehicles within a preset range through an in - vehicle camera;
[0099] S2. Encode the vehicle historical trajectory - related data (historical trajectory MLP (multi - layer perceptron)+ LSTM (long short - term memory neural network)) into a shape format that can be embedded into the x and y modules through an encoder composed of; (coordinates, lane information, speed, acceleration); SA module and LKCP module;
[0100] S3. Then the SA module calculates the spatial correlation degree between the target vehicle and each vehicle within the preset surrounding range according to the calculation formula of the attention mechanism;
[0101] S4. The LKCP module and SA module are in a parallel relationship. Through convolutional layers and pooling layers with multiple different convolutional kernels, capture the interaction between other surrounding vehicles except the target vehicle within the preset range;
[0102] S5. Then splice the outputs of the SA module and LKCP module and input them into the CM module for further feature extraction. Through convolution, capture the temporal correlation in the vehicle historical trajectory information again after the LSTM encoder;
[0103] S6. Finally, output the predicted trajectory of the target vehicle in the future through an intention recognition module and a decoder module (composed of FC (fully - connected layer)+ LSTM (long short - term memory neural network)+ FC (fully - connected layer)).
[0104] As Figure 1 shown, the present invention provides a vehicle trajectory prediction system integrating large - kernel convolution and attention mechanism. The system includes: a data acquisition module, an encoder module, a multi - head spatial attention module( Spatialattention / SA ) Large kernel convolution pooling module ( Large - kernel Convolution pooling / LKCP ) Convolution modulation module ( Convolutional modulation / CM ) Intent recognition module and decoder module;
[0105] Data acquisition module, used to collect and track in real time the trajectory data and the lane information of the host vehicle and other surrounding vehicles within a preset range through an in-vehicle camera;
[0106] Specifically: collect and track in real time the driving trajectory data of the host vehicle and other vehicles within a certain range through an in-vehicle camera, and then store and clean these vehicle driving trajectory data to form a usable vehicle driving trajectory data set. Finally, divide the data set into a training set, a test set and a validation set, and plan to use the training set to train the prediction model system, and use the test set and the validation set to optimize the model system parameters and evaluate the prediction effect of the model system.
[0107] Encoder module, used to encode the vehicle historical trajectory related data (historical trajectory MLP (Multi-layer perceptron) + LSTM (Long short-term memory neural network)) into a shape format that can be embedded into the x module and the y module; SA module and LKCP module;
[0108] The encoder module includes a multi-layer perceptron MLP and a long short-term memory neural network LSTM ;
[0109] To extract the preliminary features of the historical trajectory data and encode it into a format suitable for LSTM encoder processing, the present invention first preprocesses the data through a multi-layer perceptron before the LSTM encoder, and first linearly transforms the vehicle historical state as follows:
[0110] Formula (1): ;
[0111] Where: is the state information of the vehicle at time , is the weight matrix of the linear layer, is the bias vector of the linear layer, and will be continuously updated and changed during the training process, is the result after linear transformation;
[0112] Considering that ELU the activation function can maintain a smoother curve compared to ReLU the activation function, in the entire domain, therefore, the present invention passes the output of the linear layer in formula (1) through the activation function for processing to obtain the final output :
[0113] Formula (2): ;
[0114] Wherein: is the vector that can be embedded into the encoder, is an adjustable hyperparameter, usually defaulting to 1;
[0115] Since the vehicle historical trajectory data has time series characteristics, therefore, the present invention adopts a long short-term memory neural network LSTM as the encoder to initially capture the hidden state in the historical trajectory and perform feature encoding. The target vehicle and surrounding vehicles are respectively encoded through LSTM the encoder, and each LSTM encoder shares weights, thereby reducing the number of network parameters, reducing the model complexity, and reducing the risk of overfitting;
[0116] Input the MLP time embedding vector output by the multi-layer perceptron ( into LSTM the encoder to obtain the hidden state at time, and the process is as follows:
[0117] Formula (3): ;
[0118] Formula (4): ;
[0119] Formula (5): ;
[0120] Formula (6): ;
[0121] Formula (7): ;
[0122] Formula (8): ;
[0123] Wherein: 、 , are the outputs of the input gate, the forget gate, and the output gate respectively, , , , are all weight matrices, , , , are all weight matrices related to the hidden state at time t-1, , , , are bias vectors, is an activation function, is the vehicle at the state of the candidate cell at time, is the candidate cell state at time, is after the encoder output labeled the hidden state of the vehicle at historical time.
[0124] The multi-head spatial attention module is used to calculate the spatial correlation degree between the target vehicle and each vehicle within a preset range around it according to the calculation formula of the attention mechanism (with the target vehicle as the center and a radius range of 90 meters around);
[0125] The core of the attention mechanism lies in dynamically adjusting the degree of attention to the input information through the relationship between the query ( Q ), the key ( K ), and the value ( V ). The present invention adopts a multi-head spatial attention mechanism to capture the interaction between the target vehicle and the surrounding vehicles. At each moment , the correlation between the target vehicle and each surrounding vehicle is calculated through the calculation formula of the attention mechanism;
[0126] From the known the hidden state vector of the target vehicle output by the encoder at time is , the hidden state vectors of the surrounding vehicles are { }, the labels of the surrounding vehicles are 1-N, respectively represent the hidden states of the surrounding vehicles numbered 1-N output by the LSTM encoder at time t. They are respectively input into the multi-head spatial attention module, and the calculation formula is as follows:
[0127] Formula (9): ;
[0128] Formula (10): ;
[0129] Formula (11): ;
[0130] Formula (12): ;
[0131] Wherein: 、 and are the query matrix, key matrix, and value matrix of the target vehicle obtained by linear layer transformation respectively, 、 and are the weight matrices of the linear transformation, 、 and are the corresponding bias terms, is the number of channels, and the output is the correlation between the target vehicle and other surrounding vehicles calculated at the moment by the first head in the multi-head attention;
[0132] In the present invention, the multi-head spatial attention module adopts 4 head attention layers, and finally the outputs of the four head attention layers are concatenated along the last dimension to obtain ;
[0133] Then is passed through GLU the activation function and residual connection and then layer normalization is performed to obtain the hidden state after considering the interaction between the target vehicle and the surrounding vehicles:
[0134] Formula (13): ;
[0135] Wherein, is the hidden state after considering the interaction between the target vehicle and the surrounding vehicles, LayerNorm is the layer normalization function, GLU is the activation function, is the attention obtained by concatenating the outputs of the four head attention layers along the last dimension, is the bias matrix, is the hidden state vector of the target vehicle output by the encoder at the moment.
[0136] The large kernel convolutional pooling module is used to capture the interaction between other surrounding vehicles except the target vehicle within a preset range through convolutional layers and pooling layers with multiple different convolutional kernels;
[0137] To more accurately predict the future position of a vehicle, the spatial interaction of other surrounding vehicles cannot be ignored. The present invention proposes a large kernel convolution pooling module, which consists of a multi-scale convolution and an adaptive average pooling layer, and can effectively obtain the information within the local neighborhood and the larger receptive field of the vehicle, thereby modeling the spatial relationship between vehicles. Figure 2 The structural schematic diagram of the module is shown;
[0138] The large kernel convolution pooling module consists of three convolutional layers and an adaptive average pooling layer, and the input is ( LSTM The hidden state vector of the surrounding vehicles output by the encoder module), and the three convolutional layers are respectively 、 DW - D - Conv 、1 1 Conv ;
[0139] First, 5 5 depth convolution ( ) is used to capture the detailed features of each surrounding vehicle in the local space;
[0140] Subsequently, a 7 7 dilation convolution layer ( DW - D - Conv ) is adopted to expand the receptive field, enabling the network to obtain vehicle interaction information within a larger range while maintaining the resolution;
[0141] Then, a 1 1 convolution (1 1 Conv ) is used to fuse the features of different channels, integrate the information extracted by multiple layers of convolution, and enhance the feature expression ability;
[0142] Finally, through the adaptive average pooling layer ( Adaptive Average Pooling ), the extracted information is further compressed and mapped into a fixed-size feature representation for effective concatenation with the hidden features output by the multi-head spatial attention module, thereby improving the modeling ability of vehicle interaction relationships;
[0143] The above complete feature extraction process is shown as follows:
[0144] Formula (14): ;
[0145] Formula (15): ;
[0146] Formula (16): ;
[0147] Formula (17): ;
[0148] Formula (18): ;
[0149] Wherein: The input of the large-kernel convolution pooling module is the hidden state vector of the surrounding vehicles output by the LSTM encoder module { }, and the labels of the surrounding vehicles are 1-N, respectively representing the hidden states of the surrounding vehicles numbered 1-N output by the LSTM encoder at time t, , , and are the hidden states during the convolution process, , and are the convolution kernels, respectively representing 5 5, 7 7 and 1 1 small local matrices, , and are the network biases, is the hidden state output after feature extraction by the LKCP module.
[0150] The convolution modulation module is used for feature extraction. By means of convolution, it captures the temporal correlation in the vehicle historical trajectory information again after the LSTM encoder;
[0151] Since there is a close association between the future driving path of a vehicle and its past historical trajectory, therefore, the present invention adopts a convolution modulation module to simulate and replace the self-attention mechanism to capture the hidden state of the target vehicle historical trajectory in the time dimension. Convolution modulation calculates the product between the large-kernel convolution output and the value ( V ), simulating the calculation process of the self-attention mechanism, that is, modulating Hadamard using the convolution features, which can reduce the computational complexity while maintaining the v framework advantages, thereby reducing the inference time. Transformer The architecture of the convolution modulation module is shown in Figure 3 ;
[0152] The convolution modulation module consists of layer normalization, 1 1 convolution, GELU activation function and 7 7 grouped convolutions. The input vector of the convolution modulation module is (multi-head spatial attention module (SA ), and the concatenation of the hidden state vectors output by the large kernel convolution pooling module ( LKCP );
[0153] First, is processed through a layer normalization function. Layer normalization helps to standardize the input features, ensuring that the model stably learns the feature distribution at different time steps and reducing the learning bias of the model;
[0154] Then, the results of the layer normalization processing are respectively subjected to feature extraction through two routes. The first route is processed through a 1 1 convolution to simulate the linear transformation in the self-attention mechanism, reconstruct the feature representation, enhance the model's ability to capture deep associations between time steps in the short term, and output a feature vector v , and the second route is processed through a 1 1 convolution and a 7 7 large kernel grouped convolution to output a feature vector ;
[0155] Through operation, the extracted features are modulated to dynamically identify and capture the key moments in the prediction process, enabling the model to more accurately predict future trajectories;
[0156] Finally, the feature dimension of the modulated hidden features is corrected through a 1 1 projection convolution to obtain the final output.
[0157] According to Figure 1 shown in the architecture, the output of the multi-head spatial attention mechanism and the output of the module are concatenated and then input into the convolutional modulation module for time feature extraction. The above complete process is shown in the following formula:
[0158] Formula (19): ;
[0159] Formula (20): ;
[0160] Formula (21): ;
[0161] Formula (22): ;
[0162] The final output is the hidden state of the historical trajectory data of the target vehicle generated on the basis of considering the spatial interaction of vehicles and the time dependence of trajectory data. The input vector of the convolutional modulation module is , is the hidden state after considering the interaction between the target vehicle and surrounding vehicles, is the hidden state output after feature extraction by the LKCP module, is the hidden state after the input vector of the convolutional modulation module is processed by layer normalization, is to through 1 1 convolution and 7 7 large-kernel grouped convolutions for processing, and the output feature vector. v is to through 1 1 convolution for processing, and the output feature vector.
[0163] The intention recognition module and the decoder module (consisting of FC (fully connected layer) + LSTM (long short-term memory neural network) + FC (fully connected layer)) output the predicted trajectory of the target vehicle in the future.
[0164] The intention recognition module in the present invention can capture the lateral (left lane change, right lane change, lane keeping) and longitudinal intentions (acceleration, deceleration, constant speed) of the vehicle, thereby helping the model better meet the prediction requirements of multi-modal trajectories and improving the accuracy and stability of the prediction. The structure of the intention recognition module mainly consists of a feed-forward layer, a normalization activation layer, and two linear classifiers. As Figure 4 shown;
[0165] Encode the information of the target vehicle output in the convolutional modulation module to extract all hidden state information of each time step of each sample's historical data;
[0166] Then, the extracted hidden state information is successively passed through the feed-forward layer and the normalization activation layer for feature extraction. The feed-forward layer consists of a fully connected layer ( FC ), and the fully connected layer uses ELU the activation function. The normalization activation layer consists of a fully connected layer ( FC ), LayerNorm and ELU the activation function;
[0167] Finally, the output of the normalization activation layer is respectively fed into the longitudinal classifier and the lateral classifier, and the weight probabilities of each longitudinal and lateral maneuver category at each future prediction time step are calculated through the fully connected layer ( FC ) and the classification activation function. The longitudinal maneuver categories include three types: vehicle acceleration, deceleration, and constant speed, and the lateral maneuver categories include three types: vehicle left lane change, right lane change, and no lane change (lane keeping).
[0168] Due to the uncertainty of the future driving trajectory of vehicles in reality, the present invention believes that the future predicted trajectory presents a multimodal distribution. Therefore, the present invention uses a bivariate Gaussian distribution to quantify the multimodal trajectory. Thus, it is defined that and are respectively the expectations of the predicted trajectory points in the and directions, and are respectively the variances of the predicted trajectory points in the and directions. is the correlation coefficient.
[0169] It is known from Figure 1 that the decoder module is composed of two fully connected layers and a layer. The function of the last fully connected layer is to map the information encoded by the first fully connected layer and the LSTM layer from the feature space to the trajectory space. To consider the uncertainty of the prediction, it is assumed that the result predicted by the decoder is considered to follow a bivariate Gaussian distribution. Therefore, the decoder can output the , , , and for each future prediction time step, where is the future prediction time step, , the multimodal trajectory follows the total probability theorem, as shown in the following formula:
[0170] Formula (23): ;
[0171] Wherein: represents the coordinates at the future prediction time step, is all vehicle states at all moments in the historical time period . The vehicle state at a certain moment in the historical time period is defined as , where represents the state quantity of the vehicle, represents the vehicle number, , where and respectively represent the and coordinates of the target vehicle at the moment, and respectively represent the speed and acceleration of the target vehicle at the moment, represents the lane where the target vehicle and all vehicles within the preset range around it are located at the , represents the vehicle types (coded numerically) of the target vehicle and all vehicles within a radius of 90 meters around it at a certain moment; represents the lateral strategy, represents the longitudinal strategy, and respectively represent the probabilities of the vehicle performing lateral strategy maneuvers and longitudinal strategy maneuvers in the case of state X, is the Gaussian distribution probability of the future trajectory of the vehicle predicted based on the vehicle's historical state and lateral and longitudinal maneuvers, , where , represents the parameters of the bivariate Gaussian distribution at each time step within the prediction time domain, and are respectively the expectations of the predicted trajectory points in the and directions, and are respectively the variances of the predicted trajectory points in the and directions, is the correlation coefficient, and respectively represent the x coordinate at the prediction time step and the y coordinate at the prediction time step , and respectively represent the set of Gaussian distribution parameters at the prediction time step and the set of Gaussian distribution parameters at the prediction time step .
[0172] Evaluation metrics:
[0173] In the evaluation stage, the accuracy of the predicted trajectory is evaluated by the root mean square error ( RMSE ) and the negative log-likelihood ( NLL ) within the 5-second future prediction time period. The following are the MSE , NLL and RMSE formulas used:
[0174] Formula (24): ;
[0175] Formula (25): ;
[0176] Formula (26): ;
[0177] Where: It represents the error between the predicted future coordinate points and the actual future coordinate points. For the specific explanations of other relevant parameters, please refer to Equation (23). RMSE is used to evaluate the error between the predicted trajectory and the actual trajectory. NLL is used to evaluate the driving maneuver error.
[0178] Comparison models:
[0179] (1) CV : Early model-based methods used a constant-speed Kalman filter, assuming that the vehicle moves at a constant speed.
[0180] (2) S - LSTM : This model separately inputs each traffic participant through LSTM layers, and then models the interaction between traffic participants through a social pooling layer, predicting the future coordinate positions only relying on historical coordinate position information.
[0181] (3) PIP : This method was proposed by H . Song et al. Different from the traditional prediction method that only relies on historical position data, it couples prediction and planning together, and assists prediction by introducing the planning information of the ego vehicle.
[0182] (4) Comparison model four: This method was proposed by Guo , Hongyan et al. This method predicts the future trajectory of the target vehicle by introducing an attention mechanism and a convolutional pooling layer, on the premise of considering the movement trend of the ego vehicle.
[0183] (5) STDAN : Predict the multi-modal trajectory of the target vehicle based on LSTM an encoder-decoder, a spatial attention mechanism, and a temporal attention mechanism.
[0184] (6) Att - GCN - LSTM : Predict the multi-modal trajectory of the target vehicle based on LSTM an encoder-decoder, a graph convolutional neural network, and an attention mechanism, and use the prediction information for the vehicle path planning task.
[0185] Simulation experiment results and visualization: As shown in Table 1-2, Figures 5 - 8 as follows.
[0186] Table 1 NGSIM Values of each model based on the RMSE dataset
[0187]
[0188] Table 2 NGSIMEach model of the dataset NLL Value
[0189]
[0190] Tables 1 and 2 show the prediction accuracy of the prediction system proposed by the present invention. Figure 5 And Figure 6 are the comparison charts of the corresponding RMSE and NLL prediction accuracies. It can be seen that the vehicle trajectory prediction system proposed by the present invention has a high prediction accuracy. Figure 7 And Figure 8 In, the historical trajectories of surrounding vehicles and the target vehicle are represented by cyan and red respectively, while the predicted trajectory of the target vehicle and its future real trajectory are represented by green and magenta with 50% transparency respectively. Figure 7 In (a), (b), and (c) represent lane keeping in the cases of sparse, medium, and dense vehicles respectively. Figure 8 In (a), (b), and (c) represent left lane changes in the cases of sparse, medium, and dense vehicles respectively. Figure 8 In (d), (e), and (f) represent right lane changes in the cases of sparse, medium, and dense vehicles respectively. The horizontal and vertical coordinates of the trajectory prediction in the picture are in meters (m) units.
[0191] It should be noted that in the present invention, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including that element.
[0192] The above is only the preferred embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. A vehicle trajectory prediction system integrating large-core convolution and attention mechanism, characterized in that The system includes: a data acquisition module, an encoder module, a multi-head spatial attention module, a large kernel convolution pooling module, a convolution modulation module, an intention recognition module, and a decoder module; The data acquisition module is used to collect and track in real time the trajectory data and the lane information of the target vehicle and other surrounding vehicles within a preset range through an in-vehicle camera; The encoder module is used to encode the data related to the historical trajectories of the vehicles collected into a shape format that can be embedded into the spatial attention module and the large kernel convolution pooling module. The encoder module uses an LSTM encoder; The multi-head spatial attention module is used to calculate the spatial correlation degree between the target vehicle and each vehicle within a preset range according to the calculation formula of the attention mechanism; The large kernel convolution pooling module is used to capture the interaction between other surrounding vehicles except the target vehicle within a preset range through convolutional layers and pooling layers with multiple different convolutional kernels; The convolution modulation module inputs the outputs of the multi-head spatial attention module and the large kernel convolution pooling module after splicing in the last dimension into the convolution modulation module for feature extraction. By means of convolution, the temporal correlation in the vehicle historical trajectory information is captured again after the LSTM encoder; The intention recognition module mainly consists of a feed-forward layer, a normalization activation layer, and two linear classifiers, which encode the information of the target vehicle output by the convolutional modulation module to extract all the hidden state information at each time step of each sample's historical data. Then, the extracted hidden state information is successively passed through the feed-forward layer and the normalization activation layer for feature extraction. Finally, the outputs of the normalization activation layer are respectively fed into the longitudinal classifier and the lateral classifier, and the weights of each maneuver category in the longitudinal and lateral directions at each future prediction time step are calculated through the fully connected layer and the softmax classification activation function. After combining the output of the convolutional modulation module and the output of the intention recognition module as the input of the decoder, the decoder module consists of two fully connected layers and an LSTM layer. The role of the last fully connected layer is to map the information encoded by the first fully connected layer and the LSTM layer from the feature space to the trajectory space to obtain the multi-modal prediction trajectory.
2. The vehicle trajectory prediction system integrating large-core convolution and attention mechanism according to claim 1, characterized in that The encoder module includes a multi-layer perceptron MLP and a long short-term memory neural network LSTM; The multi-layer perceptron preprocesses the data. First, the historical state of the vehicle is linearly transformed as follows: Formula (1): Wherein: is the state information of vehicle n at time t, W mlp is the weight matrix of the linear layer, b mlp is the bias vector of the linear layer, W mlp and b mlp will be continuously updated and changed during the training process, is the result after linear transformation; Process the output of the linear layer in formula (1) through the ELU activation function to obtain the final output Formula (2): Wherein: is a vector that can be embedded in the LSTM encoder, and α is an adjustable hyperparameter; The long short-term memory neural network LSTM is used as the encoder to initially capture the hidden state in the historical trajectory and perform feature encoding. The target vehicle and the surrounding vehicles are respectively encoded through the LSTM encoder, and each LSTM encoder shares weights, thereby reducing the number of network parameters; The embedding vector at time t-1 output by the multi-layer perceptron is input into the LSTM encoder to obtain the hidden state at time t, and the process is as follows: Formula (3): Formula (4): Formula (5): Formula (6): Formula (7): Formula (8): Wherein: are respectively the outputs of the input gate, the forget gate, and the output gate, and W i , W f , W c , W o are all weight matrices, and U i , U f , U c , U o are all weight matrices related to the hidden state at time t-1, b i , b f , b c , b o are bias vectors, σ is an activation function, is the state of the candidate cell of vehicle n at time t-1, is the state of the candidate cell at time t, is the hidden state of vehicle n labeled at time t in the history output by the LSTM encoder.
3. A vehicle trajectory prediction system integrating large-core convolution and attention mechanism according to claim 1, characterized in that, The multi-head spatial attention module adopts a multi-head spatial attention mechanism to capture the interaction between the target vehicle and the surrounding vehicles. At each moment t, the correlation between the target vehicle and each surrounding vehicle is calculated according to the calculation formula of the attention mechanism; The hidden state vector of the target vehicle at time t output by the known LSTM encoder is The hidden state vectors of the surrounding vehicles are The labels of the surrounding vehicles are 1-N, which respectively represent the hidden states of the surrounding vehicles numbered 1-N output by the LSTM encoder at time t. They are respectively input into the multi-head spatial attention module, and the calculation formula is as follows: Formula (9): Formula (10): Formula (11): Formula (12): Where: Q t , and are the query matrix, key matrix, and value matrix of the target vehicle obtained through linear layer transformation respectively, W Q , W K and W V are the weight matrices of the linear transformation, b Q , b K and b V are the corresponding bias terms, d is the number of channels, output is the correlation between the target vehicle and other surrounding vehicles calculated at the t-th moment for the first head in the multi-head attention; The multi-head spatial attention module uses 4 head attention layers, and finally the outputs of the four head attention layers are concatenated according to the last dimension to obtain Then Att t After passing through the GLU activation function and residual connection, layer normalization is performed to obtain the hidden state after considering the interaction between the target vehicle and surrounding vehicles: Formula (13): Among them, is the hidden state after considering the interaction between the target vehicle and surrounding vehicles, LayerNorm is the layer normalization function, GLU is the activation function, Att t is the attention obtained by concatenating the outputs of the four-head attention layers along the last dimension, W GLU is the bias matrix, is the hidden state vector of the target vehicle at time t output by the LSTM encoder.
4. A vehicle trajectory prediction system integrating large-core convolution and attention mechanism according to claim 3, characterized in that, The large-core convolution pooling module consists of three convolutional layers and an adaptive average pooling layer, with the input being The three convolutional layers are 5×5 depthwise convolution DW-Conv, 7×7 dilated depthwise convolution DW-D-Conv, and 1×1 convolution 1×1Conv respectively; First, 5×5 depthwise convolution DW-Conv is used to capture the detailed features of each surrounding vehicle in the local space; Subsequently, 7×7 dilated convolution DW-D-Conv is adopted to expand the receptive field, enabling the network to obtain vehicle interaction information within a larger range while maintaining the resolution; Then, a 1×1 convolution 1×1Conv is used to fuse the features of different channels, integrate the information extracted by multiple layers of convolution, and enhance the feature expression ability; Finally, the information extracted is further compressed through an adaptive average pooling layer and mapped into a feature representation of a fixed size for effective splicing with the hidden features output by the multi-head spatial attention module, thereby improving the modeling ability of the vehicle interaction relationship; The feature extraction process is shown as follows: Formula (14): Formula (15): Formula (16): H3 = Conv 1*1 (H2) = W3 * H2 + b3; Formula (17): H4 = AdaptiveAvgPool1d(H3); Formula (18): Among them: The input of the large kernel convolution pooling module is the hidden state vector of surrounding vehicles output by the LSTM encoder module The labels of the surrounding vehicles are 1-N, which respectively represent the hidden states of the surrounding vehicles numbered 1-N output by the LSTM encoder at time t. H1, H2, H3, and H4 are the hidden states during the convolution process. W1, W2, and W3 are the convolution kernels, which respectively represent small local matrices of 5×5, 7×7, and 1×1. b1, b2, and b3 are the network biases, and is the hidden state output after feature extraction by the large kernel convolution pooling module.
5. The vehicle trajectory prediction system integrating large-core convolution and attention mechanism according to claim 4, wherein The convolutional modulation module consists of layer normalization, 1×1 convolution, GELU activation function, and 7×7 grouped convolution. The input vector of the convolutional modulation module is First, process it through a layer normalization function, which helps standardize the input features; Then, the results of layer normalization are used for feature extraction through two routes respectively. The first route is processed by 1×1 convolution to simulate the linear transformation in the self-attention mechanism, reconstruct the feature representation, enhance the model's ability to capture deep associations between time steps in the short term, and output the feature vector v. The second route is processed by 1×1 convolution and 7×7 large kernel grouped convolution to output the feature vector a; Through the a·v operation, the extracted features are modulated to dynamically identify and capture the key moments in the prediction process, enabling the model to more accurately predict future trajectories; Finally, the feature dimension of the modulated hidden features is corrected through 1×1 projection convolution to obtain the final output; The output of the multi-head spatial attention mechanism and the output of the LKCP module are concatenated and then input into the convolutional modulation module for time feature extraction, as shown in the following formula: Formula (19): Formula (20): Formula (21): Formula (22): Final output is the hidden state of the historical trajectory data of the target vehicle generated on the basis of considering the vehicle space interaction and the time dependence of the trajectory data. The input vector of the convolutional modulation module is the hidden state after considering the interaction between the target vehicle and surrounding vehicles the hidden state output after feature extraction by the large-kernel convolutional pooling module the hidden state after layer normalization of the input vector of the convolutional modulation module. a is processed by 1×1 convolution and 7×7 large-kernel grouped convolution, and the output feature vector. v is processed by 1×1 convolution, and the output feature vector.
6. The vehicle trajectory prediction system integrating large-core convolution and attention mechanism according to claim 5, characterized in that The feed-forward layer in the intention recognition module consists of a fully connected layer, and the fully connected layer uses the ELU activation function. The normalization activation layer consists of a fully connected layer, LayerNorm, and the ELU activation function; The longitudinal maneuver categories include vehicle acceleration, deceleration, and constant speed, and the lateral maneuver categories include vehicle left lane change, right lane change, and no lane change.
7. The vehicle trajectory prediction system integrating large-core convolution and attention mechanism according to claim 3, characterized in that, The decoder module consists of two fully connected layers and an LSTM layer. The role of the last fully connected layer is to map the information encoded by the first fully connected layer and the LSTM layer from the feature space to the trajectory space. To consider the uncertainty of prediction, it is assumed that the prediction result of the decoder is considered to follow a binary Gaussian distribution. Therefore, the decoder can output μ for each future prediction time step t′,x , μ t′,y , σ 2 t′,x , σ 2 t′,y and where t′ is the future prediction time step, t′ ∈ [0, F′]. The multi-modal trajectory follows the total probability theorem, as shown in the following equation: Formula (23): where: Y = (x t′ , y t′ ) represents the coordinates at the future prediction time step, X = {X0, X1, X2, …, X T is all vehicle states at all times in the historical time period T. The vehicle state at a certain time t in the historical time period T is defined as where τ represents the state quantity of the vehicle, {0, 1, 2, …, n} represents the vehicle number, where and respectively represent the x and y coordinates of the target vehicle at time t, and respectively represent the speed and acceleration of the target vehicle at time t, lane t represents the lane id of the target vehicle and all vehicles within a preset range around it at time t, class t represents the vehicle types of the target vehicle and all vehicles within a radius of 90 meters around it at time t; lat represents the lateral strategy, lon represents the longitudinal strategy, P(lat|X) and P(lon|X) respectively represent the probabilities of the vehicle executing lateral strategy maneuvers and longitudinal strategy maneuvers based on the state X, P γ (Y|X, lat, lon) is the Gaussian distribution probability of the predicted future trajectory of the vehicle based on the vehicle's historical state and lateral and longitudinal maneuvers, γ = {γ 1′ , γ 2′ , γ t′ , …, γ F′} where represent the parameters of the bivariate Gaussian distribution at each time step within the prediction horizon, μ t′,x and μ t′,y are the expectations of the predicted trajectory points in the x and y directions respectively, σ 2 t′,x and σ 2 t′,y are the variances of the predicted trajectory points in the x and y directions respectively, is the correlation coefficient, X t′ and y t′ represent the x - coordinate at the predicted time step t′ and the y - coordinate at the predicted time step t′ respectively, γ t′ and γ F′ represent the set of Gaussian distribution parameters at the predicted time step t′ and the set of Gaussian distribution parameters at the predicted time step F′ respectively.
8. A vehicle trajectory prediction method integrating large-core convolution and attention mechanism, characterized in that, The method includes: S1. Real-time collect and track the trajectory data and lane information of the target vehicle and other surrounding vehicles within a preset range through an in-vehicle camera; S2. The encoder module encodes the vehicle historical trajectory-related data collected into a shape format that can be embedded into the spatial attention module and the large kernel convolutional pooling module. The encoder module uses an LSTM encoder; S3. The multi-head spatial attention module calculates the spatial correlation degree between the target vehicle and each vehicle within a preset range according to the calculation formula of the attention mechanism; S4. The large kernel convolutional pooling module captures the interaction between other surrounding vehicles except the target vehicle within the preset range through convolutional layers and pooling layers with multiple different convolutional kernels; S5. Then, the outputs of the multi-head spatial attention module and the large kernel convolutional pooling module are concatenated in the last dimension and input into the convolutional modulation module for further feature extraction. Through convolution, the time correlation in the vehicle historical trajectory information is captured again after the LSTM encoder; S6. The intention recognition module is mainly composed of a feed-forward layer, a normalization activation layer, and two linear classifiers, which encode the information of the target vehicle output by the convolutional modulation module, extract all the hidden state information of each sample historical data at each time step, and then sequentially pass the extracted hidden state information through the feed-forward layer and the normalization activation layer for feature extraction. Finally, the output of the normalization activation layer is fed into the longitudinal classifier and the lateral classifier respectively, and the weight probabilities of each longitudinal and lateral maneuver category at each future prediction time step are calculated through the fully connected layer and the softmax classification activation function; after combining the output of the convolutional modulation module and the output of the intention recognition module as the input of the decoder, the decoder module is composed of two fully connected layers and an LSTM layer. The role of the last fully connected layer is to map the information encoded by the first fully connected layer and the LSTM layer from the feature space to the trajectory space to obtain the multi-modal prediction trajectory. The output of the convolutional modulation module and the output of the intention recognition module are combined as the input of the decoder. The decoder module consists of two fully connected layers and an LSTM layer. The function of the last fully connected layer is to map the information encoded by the first fully connected layer and the LSTM layer from the feature space to the trajectory space, resulting in a multi-modal prediction trajectory.
Citation Information
Patent Citations
Vehicle interactive perception trajectory prediction system and method based on graph convolution
CN118570763A
Vehicle multi-modal trajectory prediction method
CN119167302A