Vehicle trajectory prediction method based on fusion of space-time attention and kinematic model
By constructing a vehicle trajectory prediction method that integrates spatiotemporal attention and kinematic models, the problem of insufficient long-term prediction accuracy in existing technologies is solved. The accuracy of vehicle trajectory prediction is improved through spatiotemporal feature extraction and probabilistic fusion network.
Patent Information
- Application Number
- CN202510519898.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-09-09
AI Technical Summary
Existing vehicle trajectory prediction methods lack accuracy in long-term predictions, mainly because they ignore vehicle interaction characteristics and kinematic constraints, resulting in prediction results that are inconsistent with actual motion laws.
A vehicle trajectory prediction method that integrates spatiotemporal attention and kinematic models is constructed. Through the spatiotemporal feature extraction module, the spatiotemporal Transformer trajectory prediction network, the vehicle kinematics prediction trajectory network and the probability fusion network, deep learning and physical models are combined to design a weighted loss function for training to generate the final predicted trajectory.
It improves the overall accuracy of vehicle trajectory prediction, effectively expresses the interactive relationship between traffic participants, and combines the advantages of deep learning and physical models to enhance the accuracy of prediction.
Smart Images

Figure CN120606865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of vehicle trajectory prediction, and in particular to a vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic models. Background Art
[0002] With the development of connected vehicles and autonomous driving technologies, accurate vehicle trajectory prediction has become a key technology for ensuring road traffic safety and improving the decision-making capabilities of intelligent connected vehicles. Vehicle trajectory prediction methods primarily fall into two categories: physics-based and data-driven. Traditional physics-based prediction methods typically assume that vehicle motion is based on fixed control variables, such as the constant acceleration (CA) model or the constant turning rate and acceleration (CTRA) model. While these simplified methods can effectively predict trajectories in the short term, they lack the ability to adapt to dynamic changes in the traffic environment, resulting in lower long-term prediction accuracy. Furthermore, currently mainstream data-driven trajectory prediction methods typically learn features directly from historical data and then predict future trajectories. Most of these methods ignore the vehicle's inherent kinematic characteristics, which can result in predicted trajectories that differ from actual motion patterns.
[0003] Patent No. CN202310272329.2 discloses a deep learning prediction method for vehicle trajectories that takes into account physical constraints. This method uses the trajectory predicted by a constant yaw rate and variable acceleration model as part of the loss function to constrain deep learning. However, the kinematic model adopted by this method is only applicable to high-speed lane-changing scenarios and does not consider the interaction characteristics between vehicles. Patent No. CN202411372241.9 discloses a vehicle trajectory prediction method based on spatiotemporal fusion attention. This method uses a graph neural network and a gated recurrent unit to extract the spatiotemporal interaction characteristics between vehicles, but this method ignores the motion constraints of the vehicle itself.
[0004] In summary, existing physical model-based methods, while offering high short-term prediction accuracy, struggle to cope with long-term predictions in dynamic traffic environments. While purely data-driven deep learning methods excel in capturing complex interactions, they often overlook vehicle kinematic constraints. Therefore, integrating the advantages of physical models with deep learning is a key issue that needs to be addressed in the current field of vehicle trajectory prediction. Summary of the Invention
[0005] In view of the shortcomings of the existing technology, the purpose of the present invention is to propose a vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model, so as to solve the problem of insufficient prediction accuracy caused by insufficient consideration of vehicle interaction and motion characteristics in the existing trajectory prediction methods.
[0006] The purpose of the present invention can be achieved through the following technical solutions:
[0007] A vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model includes the following steps:
[0008] Step 1: Obtain real traffic data;
[0009] Step 2: Construct a spatiotemporal feature extraction module for modeling and extracting the spatiotemporal interaction features between vehicles. The output of the spatiotemporal feature extraction module is a spatiotemporal feature vector, and its input is traffic data.
[0010] Step 3: Construct a prediction trajectory generation module, whose input is the spatiotemporal feature vector and whose output is the final prediction trajectory;
[0011] The predicted trajectory generation module includes:
[0012] The spatiotemporal Transformer trajectory prediction network, whose output is the spatiotemporal Transformer network output trajectory Y STT , whose input is the spatiotemporal feature vector;
[0013] Vehicle kinematics prediction trajectory network, whose output is the vehicle prediction trajectory Y STTKM , whose input is the spatiotemporal feature vector;
[0014] Probabilistic fusion network, whose output is fusion probability and whose input is spatiotemporal feature vector;
[0015] And the probability fusion model, whose input is Y STT 、Y STTKM and fusion probability, the output of which is the final predicted trajectory;
[0016] Step 4: Design a weighted loss function and train the network module consisting of the spatiotemporal feature extraction module and the predicted trajectory generation module based on real traffic data;
[0017] Step 5: Input real-time traffic data into the trained network module to obtain the predicted vehicle trajectory.
[0018] Preferably, step 1 specifically includes the following steps:
[0019] Step 1-1: Data segmentation: Divide the original data into time segments of fixed length, which are used for training and prediction;
[0020] Step 1-2: Create a feature matrix:
[0021] The vehicle motion state contained in the feature matrix is expressed as follows:
[0022]
[0023] Where: Represents the status information of vehicle i at the tth second;
[0024] in l represents the vehicle's longitudinal position, lateral position, longitudinal speed, lateral speed, heading angle and vehicle length information respectively;
[0025] Step 1-3: Create a mask matrix corresponding to the vehicle;
[0026] Steps 1-4: Feature normalization.
[0027] Preferably, step 2 specifically includes the following steps:
[0028] Step 2-1: Use the spatiotemporal graph G = (V, E) to represent the spatiotemporal interaction relationship of traffic participants;
[0029] Step 2-2: Use spatial self-attention to model the space of vehicles within the same frame time;
[0030] Step 2-3: Use the TCN network to perform convolution operations on the time edge of the spatiotemporal graph;
[0031] Step 2-4: Use the improved Transformer to encode and decode the obtained spatiotemporal feature vector to further extract the dynamic features of the time dimension.
[0032] Preferably, the specific steps of step 3 are as follows:
[0033] Step 3-1: Build a spatiotemporal Transformer trajectory prediction network;
[0034] Step 3-2: Construct vehicle kinematics prediction trajectory network;
[0035] Step 3-3: Design probability fusion network and probability fusion model;
[0036] The probability fusion network outputs a fusion probability α∈[0,1];
[0037] The probability fusion model is:
[0038] Y PFN-STTKM =α·Y STT +(1-α)·Y STTKM .
[0039] Preferably, step 4 is as follows:
[0040] Step 4-1: Design a weighted loss function that considers the prediction errors of the three trajectories;
[0041] Step 4-2: Use adaptive learning rate optimization method for training.
[0042] Compared with the prior art, the present invention has the following beneficial effects:
[0043] 1) Construct a spatiotemporal graph to effectively express the interactive relationships between traffic participants, and leverage the Transformer’s attention mechanism and the feature extraction capabilities of a temporal convolutional network to efficiently model and extract the spatiotemporal interaction features between vehicles.
[0044] 2) A vehicle kinematic model is introduced, and a probabilistic fusion network is designed to adaptively fuse the trajectory predicted by the spatiotemporal Transformer with the trajectory recursively derived from the kinematic model. By combining the advantages of deep learning and physical models, the overall prediction accuracy of the vehicle trajectory is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is the overall framework diagram of the spatiotemporal attention and vehicle kinematic model probability fusion network;
[0046] Figure 2 To improve the Transformer structure diagram;
[0047] Figure 3(a) is a schematic diagram of spatial self-attention in spatiotemporal attention;
[0048] Figure 3(b) is a schematic diagram of temporal self-attention in spatiotemporal attention;
[0049] Figure 4 It is the probability fusion curve;
[0050] Figure 5 It is the learning rate change curve;
[0051] Figure 6 is the feature matrix;
[0052] Figure 7 is the mask matrix;
[0053] Figure 8 Predict trajectory graph for the vehicle. DETAILED DESCRIPTION
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0055] A vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model, that is, a trajectory prediction method for intelligent connected vehicles (hereinafter referred to as vehicles) based on a probability fusion network of spatiotemporal attention and kinematic model, the overall structure of which is as follows Figure 1 As shown, the specific steps include:
[0056] Step 1: Obtain real traffic data and preprocess it to form a dataset.
[0057] The dataset includes normalized feature vectors and mask matrices.
[0058] Furthermore, the step 1 is specifically as follows:
[0059] Step 1-1: Data segmentation.
[0060] Select typical traffic scenarios (such as highways, ramp merging, roundabouts, intersections, etc.) and divide the original data into time windows.
[0061] Each data segment is 8 seconds long, of which the first 3 seconds of data is used as model input to predict the vehicle trajectory for the next 5 seconds.
[0062] The "data" in the data segment refers to
[0063] Step 1-2: Create a feature matrix, such as Figure 6 shown.
[0064] A fixed-dimensional matrix structure is used to store vehicle motion state information to construct a feature matrix as input for subsequent models.
[0065] The vehicle motion state contained in the feature matrix is expressed as follows:
[0066]
[0067] Where: Represents the status information of vehicle i at the tth second;
[0068] in l represents the vehicle's longitudinal position, lateral position, longitudinal velocity, lateral velocity, heading angle and vehicle length information respectively.
[0069] Figure 6 In , the horizontal axis represents the element number in the vehicle motion state, and the vertical axis represents the vehicle number.
[0070] Step 1-3: Create a mask matrix, such as Figure 7 shown.
[0071] The value of 1 in the matrix represents valid data, and the value of 0 represents invalid data. The mask matrix is mainly used to eliminate incomplete trajectories (insufficient number of trajectory points or interruptions) and invalid data. Such data will not participate in the subsequent model training and testing.
[0072] like Figure 6In the example, vehicle number 17 and the data corresponding to element numbers 0 to 5 are all 0, which is called invalid data.
[0073] like Figure 7 In the figure, the vertical axis represents the vehicle number.
[0074] Steps 1-4: Feature normalization.
[0075] In order to ensure the stability of model training, the extracted feature vectors are normalized to obtain the normalized feature vectors.
[0076] Among them, the eigenvector The set of , that is, the feature matrix.
[0077] Step 2: Construct a spatiotemporal feature extraction module to model and extract the spatiotemporal interaction features of vehicles.
[0078] The spatiotemporal feature extraction module includes: sequentially connected spatial attention layers, temporal convolutional networks (TCNs), improved Transformer encoders and decoders.
[0079] Further, the step 2 is specifically as follows:
[0080] Step 2-1: Use the spatiotemporal graph G = (V, E) to represent the N vehicle participants in the trajectory sequence at time T obs The space-time relationship within.
[0081] The spatiotemporal attention between vehicles is as follows Figure 2 shown.
[0082] Furthermore, the node set V contains the feature vectors of all time frames of each vehicle, which is expressed as follows:
[0083] V={x it |t∈(1,T obs ),i∈(1,N)}
[0084] Where: x it Represents the feature vector of vehicle i in the tth frame.
[0085] Among them, the tth frame, tth second, and time step t have the same meaning.
[0086] Furthermore, the edge set E contains two types of edges, including spatial edges E s Used to indicate intra-frame connection, time edge E t Used to indicate inter-frame connections, as follows:
[0087] E S ={(x it ,x jt)|i,j∈(1,N),t∈(1,T obs )}
[0088] E T ={(x it ,x it+1 )|i∈(1,N),t,t+1∈(1,T obs )}
[0089] Where: (x it ,x jt ) represents the relationship between node vehicle i and node vehicle j in the spatial dimension at time frame t; (x it ,x it+1 ) represents the relationship between node vehicle i in the time dimension between time t and t+1.
[0090] Step 2-2: Use the spatial self-attention layer to perform spatial modeling of vehicles within the same frame time.
[0091] Spatial attention can be viewed as a spatial edge in the spatiotemporal graph. A message passing mechanism is introduced on the spatial edge to calculate the input using linear transformation. Get the query vector Key Vector Sum value vector The expression is:
[0092]
[0093] Where: and are the corresponding linear projection matrices respectively.
[0094] The calculation of the query vector, key vector, and value vector of the node vehicle j is also applicable to the above formula.
[0095] Further, in and The dot product is used to calculate the attention weight score between node vehicle i and node vehicle j.
[0096] use The message sent from vehicle j to vehicle i represents the weighted score of the spatial edge, which can be expressed as:
[0097]
[0098] Where T represents transpose.
[0099] Furthermore, node i receives messages from all neighboring nodes. After that, the value vector Perform weighted average summation to obtain a single attention head, which is then concatenated and linearly transformed to obtain spatial multi-head attention, specifically:
[0100]
[0101] Where: H represents the number of attention heads; W o is the linear layer weight; concat represents the vector concatenation operation.
[0102] Step 2-3: Use the TCN network to perform convolution operations on the time edge of the spatiotemporal graph.
[0103] Given an input of shape dimension (T, N, C), where T represents the size of the history frame, N represents the number of nodes, and C represents the feature map dimension of each node, a convolution operation is performed using a (K×1) convolution kernel as follows:
[0104]
[0105] Furthermore, in order to ensure that the convolution operation does not use future information during the feature extraction process, causal convolution is used, that is, the output of the model at time step t only depends on the input at time step t and before, specifically:
[0106]
[0107] Where: H t is the value of the output sequence at time step t; W k is the convolution kernel parameter.
[0108] Step 2-4: Use the improved Transformer to encode and decode the obtained spatiotemporal feature vector to further extract the dynamic features of the time dimension.
[0109] like Figure 2 As shown in the figure, the present invention replaces the forward connection layer in the traditional Transformer encoder and decoder with a separable convolution layer. The improved separable convolution layer is more suitable for processing time series data, and its local convolution kernel can better capture the feature relationship between adjacent time steps or spatial neighborhoods, thereby enhancing the feature expression capability.
[0110] Different from the spatial self-attention mechanism in Figure 3(a), this step in Figure 3(b) calculates the attention weight independently along the time dimension for the trajectory features of each node. The temporal self-attention of node vehicle i is calculated as follows:
[0111]
[0112] Where: Q i 、 and V iThey are the query matrix, key matrix, and value matrix learned from the input of node i through the embedding layer, and their calculation method is the same as step 2-2.
[0113] W u is the linear layer weight;
[0114] d k is the dimension size of the key vector;
[0115] h is the number of attention heads.
[0116] Step 3: Construct a prediction trajectory generation module, which generates the final prediction trajectory based on the prediction results of the fusion probability network, the spatiotemporal Transformer trajectory prediction network, and the kinematic trajectory prediction network.
[0117] The predicted trajectory generation module includes: a spatiotemporal Transformer trajectory prediction network, a fusion probability network, a kinematic trajectory prediction network, and a fusion probability model.
[0118] Further, the step 3 is specifically as follows:
[0119] Step 3-1: Build a spatiotemporal Transformer trajectory prediction network, which uses a fully connected layer (trajectory prediction FC layer).
[0120] The spatiotemporal Transformer trajectory prediction network, whose output is the spatiotemporal feature vector, directly generates the spatiotemporal Transformer predicted trajectory Y at the tth second STT .
[0121] Spatiotemporal Transformer predicts trajectory Y STT , that is, the horizontal and vertical position coordinates (x, y) of the vehicle in the future prediction time domain).
[0122] Step 3-2: Construct a vehicle kinematics prediction trajectory network, using another fully connected layer (control quantity prediction FC layer).
[0123] The vehicle kinematics prediction trajectory network takes the spatiotemporal feature vector as input and outputs the vehicle control sequence u=[a0,δ f,0 ……a T ,δ f,T ], including acceleration and front wheel angle.
[0124] The control sequence is input into the kinematic model after being limited (control quantity constraint). The control quantity constraint is:
[0125] a min ≤a≤a max
[0126] δ min≤δ≤δ max
[0127] Where: a min and a max are the minimum and maximum accelerations of the vehicle, δ min and δ max are the minimum and maximum front wheel steering angles.
[0128] Then the predicted vehicle trajectory Y at the tth second is presented STTKM , the kinematic model is as follows:
[0129] x k+1 =x k +v k cos(ψ k +β k )Δt
[0130] y k+1 =y k +v k sin(ψ k +β k )Δt
[0131]
[0132] v k+1 =v k +a k Δt
[0133]
[0134] Where: x k+1 、y k+1 , ψ k+1 and v k+1 Respectively represent the longitudinal position, lateral position, heading angle and velocity of the vehicle at the next moment; where velocity is the sum of the longitudinal velocity and lateral velocity vectors;
[0135] β k Indicates the current vehicle center of mass sideslip angle;
[0136] Δt is the discrete step length;
[0137] a k and δ f,k They represent the control input acceleration and front wheel angle at the current moment respectively.
[0138] Step 3-3: Design the probability fusion network and probability fusion model.
[0139] The probability fusion network adopts a third fully connected network (fusion probability FC layer).
[0140] The probability fusion network takes the spatiotemporal feature vector as input and uses the sigmoid activation function. Its output is the fusion probability α∈[0,1], which provides a basis for weighted fusion.
[0141] Probability distribution such as Figure 4 As shown in the figure, the probabilistic fusion network reasonably combines the advantages of the vehicle kinematic model and the Transformer prediction framework. It tends to rely more on the kinematic model in short-term predictions and gradually increases its trust in the spatiotemporal Transformer in the long-term prediction stage.
[0142] Furthermore, the probability fusion model, whose output is the final predicted trajectory at the tth second, is expressed as follows:
[0143] Y PFN-STTKM =α·Y STT +(1-α)·Y STTKM
[0144] Step 4: Design a weighted loss function for network model training.
[0145] Further, the step 4 is specifically as follows:
[0146] Step 4-1: Design a weighted loss function for the network module and train the network module based on the dataset obtained in step 1.
[0147] The network module includes a spatiotemporal feature extraction module and a prediction trajectory generation module.
[0148] In step 3, a total of three trajectories can be obtained: the trajectory directly predicted by the spatiotemporal Transformer network, the trajectory derived and calculated by the kinematic prediction network, and the final trajectory output by the probabilistic fusion of the trajectories output by these two networks. Since the accuracy of the final output predicted trajectory also depends on the accuracy of the first two predicted trajectories, in order to optimize the network performance, the loss function needs to consider the prediction errors of these three trajectories at the same time. The expression is as follows:
[0149]
[0150] Where: β1, β2 and β3 represent the weights of the three prediction trajectory errors respectively;
[0151] Y t represents the true trajectory at the tth second, which can be obtained by step 1;
[0152] represents the spatiotemporal Transformer network output trajectory at the tth second;
[0153] represents the output trajectory of the kinematic prediction network at the tth second;
[0154] represents the final trajectory of the probability fusion network output at the tth second.
[0155] Step 4-2: Use the NoamOpt optimizer to train the model. The learning rate is dynamically adjusted as follows:
[0156]
[0157] Among them, such as Figure 5 As shown, step is the horizontal axis, which is the current training step number.
[0158] Furthermore, set the model dimension d model =32, warmup_steps=5000, learning rate change curve is as follows Figure 5 As shown in the figure, during the warmup phase, the learning rate gradually increases with the number of training steps to prevent instability caused by an excessively high initial learning rate. Once the number of training steps exceeds warmup_steps, the learning rate begins to decrease, helping the model converge more smoothly. Later in training, the learning rate stabilizes, reducing fluctuations and improving the model's generalization ability.
[0159] Step 5: After the real-time traffic data undergoes the same processing as in step 1, it is input into the trained network module to obtain the predicted vehicle trajectory.
[0160] like Figure 8 As shown, the numbers represent the vehicle numbers, the colored dotted lines represent the predicted vehicle trajectories, the colored solid lines represent the actual trajectories, and the gray solid lines represent the road edges.
[0161] The above description is only a description of the preferred embodiments of the present application and does not limit the scope of the present application. Any changes or modifications made by any person skilled in the art based on the above disclosed technical content should be regarded as equivalent valid embodiments and fall within the scope of protection of the technical solution of the present application.
Claims
1. A vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model, characterized in that: The steps include: Step 1: Obtain real traffic data; Step 2: Construct a spatiotemporal feature extraction module for modeling and extracting the spatiotemporal interaction features between vehicles. The output of the spatiotemporal feature extraction module is a spatiotemporal feature vector, and its input is traffic data. Step 3: Construct a prediction trajectory generation module, whose input is the spatiotemporal feature vector and whose output is the final prediction trajectory; The predicted trajectory generation module includes: The spatiotemporal Transformer trajectory prediction network, whose output is the spatiotemporal Transformer network output trajectory Y STT , whose input is the spatiotemporal feature vector; Vehicle kinematics prediction trajectory network, whose output is the vehicle prediction trajectory Y STTKM , whose input is the spatiotemporal feature vector; Probabilistic fusion network, whose output is fusion probability and whose input is spatiotemporal feature vector; And the probability fusion model, whose input is Y STT 、Y STTKM and fusion probability, the output of which is the final predicted trajectory; Step 4: Design a weighted loss function and train the network module consisting of the spatiotemporal feature extraction module and the predicted trajectory generation module based on real traffic data; Step 5: Input real-time traffic data into the trained network module to obtain the predicted vehicle trajectory.
2. The vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model according to claim 1 is characterized in that: Step 1 specifically includes the following steps: Step 1-1: Data segmentation: Divide the original data into time segments of fixed length, which are used for training and prediction; Step 1-2: Create a feature matrix: The vehicle motion state contained in the feature matrix is expressed as follows: Where: Represents the status information of vehicle i at the tth second; in l represents the vehicle's longitudinal position, lateral position, longitudinal speed, lateral speed, heading angle and vehicle length information respectively; Step 1-3: Create a mask matrix corresponding to the vehicle; Steps 1-4: Feature normalization.
3. The vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model according to claim 1 is characterized in that: Step 2 specifically includes the following steps: Step 2-1: Use the spatiotemporal graph G = (V, E) to represent the spatiotemporal interaction relationship of traffic participants; Step 2-2: Use spatial self-attention to model the space of vehicles within the same frame time; Step 2-3: Use the TCN network to perform convolution operations on the time edge of the spatiotemporal graph; Step 2-4: Use the improved Transformer to encode and decode the obtained spatiotemporal feature vector to further extract the dynamic features of the time dimension.
4. The vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model according to claim 1 is characterized in that: Step 3: Step 3-1: Build a spatiotemporal Transformer trajectory prediction network; Step 3-2: Construct vehicle kinematics prediction trajectory network; Step 3-3: Design probability fusion network and probability fusion model; The probability fusion network outputs a fusion probability α∈[0,1]; The probability fusion model is: AND PFN-STTKM =α·Y STT +(1-α)·Y STTKM 。 5. The vehicle trajectory prediction method based on the fusion of spatiotemporal attention and kinematic model according to claim 1 is characterized in that: Step 4 is as follows: Step 4-1: Design a weighted loss function that considers the prediction errors of the three trajectories; Step 4-2: Use adaptive learning rate optimization method for training.
Citation Information
Patent Citations
Vehicle track deep learning prediction method considering physical constraint
CN116495007A
Vehicle trajectory prediction method based on time-space fusion attention
CN119380532A
Cited By
Deep learning trajectory prediction method introducing vehicle kinematics constraint
CN121386791A
A deep learning trajectory prediction method introducing vehicle kinematic constraints
CN121386791B