Transform and end point induction-based multi-modal trajectory prediction method for automatic driving vehicle

By adopting a Transformer-based multimodal trajectory prediction method in the autonomous driving system, combining attention mechanism and dynamic weight MLP, the problem of difficulty in accurately predicting multimodal trajectory in complex traffic scenarios in the prior art is solved, and more efficient and accurate trajectory prediction is achieved.

CN119975390AActive Publication Date: 2025-05-13SOUTHEAST UNIV

Patent Information

Application Number
CN202510079290.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The existing methods of autonomous driving vehicle trajectory prediction are difficult to accurately predict multimodal trajectories in complex traffic scenarios, and ignore road structure, environmental information and vehicle interaction relationships.

Method used

The multimodal trajectory prediction method based on Transformer is adopted to extract road scene and vehicle information through an encoder, and the attention mechanism and Transformer are used to fusion feature, combining dynamic weight MLP and attention mechanism interaction endpoint information to achieve trajectory refinement and prediction.

Benefits of technology

It improves the accuracy and efficiency of multi-modal trajectory prediction of autonomous driving vehicles, can better adapt to complex and changeable traffic environments, and enhances driving safety and prediction rationality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119975390A_ABST
    Figure CN119975390A_ABST
Patent Text Reader

Abstract

The invention provides an automatic driving vehicle multi-mode track prediction method based on Transform and end point induction. The method comprises the following steps: firstly, extracting historical track characteristics of surrounding vehicles, historical track characteristics of a target vehicle, relative position information and scene map characteristics; carrying out local and global fusion on the extracted features by utilizing Transform, and efficiently and comprehensively fusing the features of each element in the scene; secondly, predicting a multi-modal trajectory end point through a dynamic weight multi-layer perceptron (MLP), simplifying a model architecture and reducing the quantity of model parameters; and finally, performing trajectory refinement by using attention mechanism interaction end point information, and outputting multi-modal trajectory prediction, so that the predicted future trajectory of the vehicle is reasonable and accurate. The method not only can improve the multi-modal trajectory prediction accuracy of the autonomous vehicle, but also can improve the model prediction efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction. Background Art

[0002] Autonomous driving has broad application prospects in solving problems such as traffic accidents, traffic jams, and emissions. In recent years, the perception, decision-making, and control modules of autonomous driving have achieved rapid development. However, the full deployment of autonomous driving requires verification of its safety. In complex traffic scenarios, even if the driver's final intention is determined, the instant changes in the dynamic driving paths of other vehicles during the decision-making process will cause the vehicle's path to present multimodal properties, that is, the vehicle's future trajectory has multiple possibilities. Trajectory prediction plays a connecting role in autonomous driving. It is not only an extension of environmental perception and the basis for decision-making planning, but also the basis for risk assessment and a bridge for human-computer interaction. Through accurate trajectory prediction, the autonomous driving system can better adapt to complex and changing traffic environments and improve driving safety and efficiency.

[0003] Existing vehicle trajectory prediction methods can be mainly divided into three categories: prediction methods based on physical models, prediction methods based on machine learning, and prediction methods based on deep learning. The core of the vehicle trajectory prediction method based on physical models is to simplify the prediction process of the target vehicle into a vehicle dynamics or kinematic model. This method accurately calculates the future state of the vehicle through iterative calculation based on the input parameters of the model, such as acceleration, steering angle, etc., as well as external conditions, such as the road friction coefficient. Although this method has low computational complexity and good real-time performance, it not only ignores the influence of prior knowledge such as road structure and environmental information, but also ignores the influence of the subjective intention and driving style of other vehicle drivers in the scene. It is only applicable to trajectory prediction within a short time domain of 1s. Most of the methods based on machine learning are based on driving behavior prediction. First, the driving behavior is judged, and then the future motion trajectory is predicted based on the behavior. However, in the prediction process, it is assumed that the vehicles are independent of each other, and the influence of the driving behavior of other vehicles in the driving scene is not fully considered. Ignoring this interactive relationship may generate unreasonable motion trajectories. Trajectory prediction methods based on deep learning can not only use time series networks to extract historical trajectory features, but also process trajectory features through different network structures to extract interaction information and road information of traffic participants. Compared with physics-based methods and classic machine learning-based methods, long-term predictions are more accurate. The application of deep learning in multimodal trajectory prediction tasks is mainly divided into three mainstream directions: methods based on sequence networks, methods based on graph neural networks, and methods based on generative models. However, compared with the classic sequence network convolutional neural network (CNN) and recurrent neural network (RNN), Transformer is a non-cyclic network structure based entirely on the self-attention mechanism. It breaks the limitation of sequence order input through encoder-decoder and attention mechanism, thereby solving the problem of long short-term memory (LSTM) and gated recurrent unit (GRU) that cannot be trained in parallel. At the same time, the Transformer model has better memory, can solve long-distance dependency problems, and has better interpretability when dealing with interaction problems. Therefore, there is an urgent need for a Transformer method that can quickly adapt to autonomous driving scenarios and ensure the completion of multimodal trajectory prediction tasks under the conditions of strong safety, comfort, timeliness and generalization. Therefore, it is of great significance to study the multimodal trajectory prediction method of autonomous driving vehicles based on Transformer and endpoint induction. Summary of the invention

[0004] In response to the current related issues of multimodal trajectory prediction in the field of autonomous driving, the present invention provides a multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction. The method uses Transformer to locally and globally fuse the extracted features, and then uses a multi-layer perceptron (MLP) based on dynamic weights to predict the multimodal trajectory endpoint. At the same time, the attention mechanism is used to interact with the endpoint information for trajectory refinement. This method improves the accuracy and efficiency of multimodal trajectory prediction for autonomous driving vehicles.

[0005] To achieve the above purpose, the technical solution of the present invention is as follows:

[0006] A multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction, the method comprising the following steps:

[0007] Step 1: Use the encoder to encode the road scene information, surrounding vehicle information including the current state and historical trajectory of the vehicle, the current state and historical trajectory and relative position information of the target vehicle, and generate a high-dimensional feature tensor.

[0008] Step 2: Perform hierarchical feature fusion on the encoded high-dimensional feature tensor. First, use the attention mechanism to perform local feature fusion on vehicle features and map features respectively, and then use Transformer to perform global feature fusion on vehicle features, map features and relative position information.

[0009] Step 3: Input the fused feature tensor into the dynamic weight MLP to output multimodal trajectory endpoint predictions in different driving scenarios.

[0010] Step 4: Use the attention mechanism to interact with the trajectory endpoint information and output the trajectory endpoint offset, so as to achieve feature interaction in a longer time and space range to obtain more accurate endpoint information, and finally complete the trajectory to output a complete multimodal prediction trajectory.

[0011] Step 5: Design a suitable loss function based on the multimodal trajectory prediction model. The loss function includes three items: regression loss of all trajectories, classification loss, and endpoint regression loss.

[0012] The feature extraction module in step 1 includes a vehicle trajectory encoder, a map information encoder and a relative position information encoder, and the specific steps are as follows:

[0013] S1: The trajectory information is a time series containing the vehicle position. The vehicle trajectory encoder input is represented by displacement. In map encoding, the position and direction of the midpoint of the lane centerline are used to represent the position and direction of each lane segment.

[0014] S2: The historical trajectory information of the i-th vehicle in H time steps is represented as a displacement sequence ΔX i ={ΔX i,-H+1 ,...,ΔX i,0}, where ΔX i,t Represents the displacement of the i-th vehicle from time t-1 to time t. In order to improve the effectiveness of multi-scale feature extraction and the efficiency of parallel computing, one-dimensional residual convolution and feature pyramid network (FPN) are used to process the historical trajectory input, and the vehicle feature tensor V output by the encoder is a The size is [N a ,128], where N a Indicates the number of vehicles in the history step.

[0015] S3: Considering that the lane segments are static during the observation time, a simple two-layer MLP is used to encode the map information to improve the operation efficiency. The map feature tensor V output by the encoder is m The size is [N l ,128], where N l Represents the number of lane centerlines in the map during the observation time.

[0016] S4: In relative position encoding, the heading angle difference α is used i→j , relative azimuth β i→j and distance ||d i→j ||These three quantities describe the relative position and posture between element i and element j, and the relative position information r between scene element i and element j i→j It can be represented by a five-dimensional vector:

[0017]

[0018] For the scenario with N elements in the observation time, MLP is used to encode the relative position information to obtain a relative position tensor rpe of size [N,N,128].

[0019] In step 2, the multi-feature fusion method based on hierarchical feature fusion includes: local feature fusion of vehicles and lanes based on a single-head self-attention mechanism, and global feature fusion of map features, vehicle trajectory features, and relative position information based on symmetric fusion transformer (SFT). Because there are too many features in complex traffic scenes, comprehensive and efficient fusion is required. Based on multiple considerations, the present invention designs a hierarchical feature fusion module.

[0020] 1) Local feature fusion

[0021] The behaviors of traffic participants in complex traffic scenes are often complex. In order to make the model focus on more important features and improve the efficiency and accuracy of trajectory prediction, local feature fusion is performed first. In local feature fusion, a single-head attention mechanism is used to realize trajectory feature interaction and map feature interaction respectively, which can realize the interaction function and improve the efficiency of interaction.

[0022] The trajectory feature tensor and map feature tensor output by the feature extraction module are input separately, and different weights are assigned to the scene element features through the single-head self-attention mechanism, where F represents the input feature tensor and F' represents the feature tensor after the attention mechanism. The specific formula is:

[0023] Q=W q F,K=W k F,V=W V F

[0024]

[0025] Where W q ,W k ,W V is the weight matrix. Finally, the interactive feature tensor is obtained through the layer normalization network and the feedforward network. Q, K, and V are the input vectors of the attention mechanism.

[0026] 2) Global feature fusion

[0027] Local feature fusion does not consider the relative position relationship of each element in the scene and the interaction between vehicles and lane segments. The Transformer model is widely used in the field of computer vision and feature fusion. It can easily process and fuse features of different scales and reduce the number of model parameters. To this end, the symmetric fusion Transformer (SFT) is used for global feature fusion. This feature fusion method performs directional information transfer in a symmetrical manner, enabling the network to predict the future movement of all road users in one feedforward pass, which not only enables effective interaction but also improves model efficiency. The specific steps are:

[0028] S1: Concatenate the local fused vehicle trajectory feature tensor and map feature tensor to obtain tensor V n , the tensor size is [N,128], where N is the sum of vehicles and lane segments in the scene.

[0029] S2: V n Repeat and stack along two different dimensions to obtain the target tensor and source tensor, both of size [N,N,128].

[0030] S3: The target tensor, source tensor and relative position feature tensor are concatenated and passed through an MLP to obtain the input K, V of the multi-head attention. The Q of the multi-head attention is the tensor Vn .

[0031] S4: Output the updated vehicle trajectory feature tensor V after the multi-head attention mechanism a 'and the relative position feature tensor rpe', where the number of attention heads is 8. Tensor V a 'Size is [N a ,128], the size of tensor rpe' is [N,N,128], where N a is the number of vehicles in the scene. The global feature fusion module effectively integrates the scene traffic participant information and road structure information.

[0032] The trajectory endpoint prediction method based on dynamic weight MLP in step 3 includes:

[0033] The traditional trajectory endpoint prediction method describes the uncertainty of vehicle trajectory based on probability distribution. However, when facing data from different regions and different collection methods, the generalization performance of the model is poor and overfitting is prone to occur. The present invention uses two MLPs with dynamic weights to generate different vehicle features to adapt to different traffic scenarios. The specific steps are as follows:

[0034] S1: Vehicle trajectory feature tensor V output by the multi-feature fusion module a 'Use an MLP layer to generate a tensor V w ' a .

[0035] S2: Introducing the trainable weight parameters W in the two MLP layers w1 and W w2 , so that V w ' a Two dynamic weight parameters W1 and W2 are generated by these two MLP layers, as shown in the formula:

[0036]

[0037] S3: The trajectory endpoint prediction process is transformed into a process that first passes through an MLP layer with a weight of W1, then passes through a normalization layer and an activation layer, and finally passes through an MLP layer with a weight of W2 to generate a multimodal endpoint y pred .

[0038] The endpoint refinement method based on trajectory endpoint information interaction in step 4 includes:

[0039] Traditional trajectory endpoint prediction methods do not perform effective endpoint refinement, and only focus on the interaction with historical observation information, ignoring the interaction of endpoint information. This paper uses the attention mechanism to efficiently interact with the feature tensor containing trajectory endpoint information, extracts effective features in a longer time and space range, and achieves accurate refinement of the trajectory endpoint, thereby enhancing the rationality and accuracy of the predicted trajectory. The specific steps are as follows:

[0040] S1: Encode the trajectory endpoint output by the endpoint prediction module to obtain the trajectory endpoint feature tensor E pred :

[0041] E pred =MLP(y pred )

[0042] S2: The concatenated vehicle trajectory feature tensor V a ' and the trajectory endpoint feature tensor E pred In the input attention mechanism, the output trajectory end offset b pred :

[0043] b pred =MHA(concat(V a ',E pred ))

[0044] S3: Use a simple MLP to complete the remaining trajectory points on the vehicle trajectory feature tensor and predict the probability of each mode.

[0045] The loss function in step 5 includes three items: regression loss of all trajectories, classification loss and endpoint regression loss. The regression loss of all trajectories can optimize the overall shape and continuity of the trajectory, making the predicted trajectory closer to the real trajectory; the classification loss can improve the model's recognition accuracy of the vehicle's motion intention, thereby providing a more reliable decision-making basis for trajectory prediction; the endpoint regression loss can optimize the endpoint position of the trajectory, making the predicted endpoint closer to the real endpoint and improving the accuracy of trajectory prediction.

[0046] Specifically:

[0047] 1) Regression loss of all trajectories

[0048] The regression loss of all trajectories includes the vehicle position coordinate regression loss and heading angle loss The introduction of heading angle loss makes the predicted trajectory smoother and more feasible, as shown in the formula:

[0049]

[0050] Where T is the prediction time step; is the position coordinate of the trajectory of the optimal mode at time t; is the position coordinate of the true trajectory at time t; CosSim is the cosine similarity metric, which takes the value of 1 when the two vectors are in the same direction and -1 when they are in opposite directions; The heading angle of the vehicle when it is in the best mode; is the actual heading angle of the vehicle.

[0051] 2) Classification loss

[0052] The classification loss function is used to measure the difference between the category predicted by the model and the true category. The cross entropy loss is used to calculate the classification loss, and the performance of the model is evaluated by calculating the difference between the predicted probability distribution and the true label.

[0053] 3) End point regression loss

[0054] The endpoint regression loss is a loss function used to measure the difference between the endpoint position predicted by the model and the actual endpoint position. Its core purpose is to optimize the prediction performance of the model by minimizing this difference, as shown in the formula:

[0055]

[0056] In the formula The position coordinates of the end point of the trajectory representing the best mode; is the position coordinate of the end point of the real trajectory.

[0057] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for multimodal trajectory prediction of an autonomous driving vehicle based on Transformer and endpoint induction is implemented.

[0058] A computer-readable storage medium stores computer instructions, which, when executed by a processor, implement the multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction.

[0059] Based on the above technical solution, the present invention has the following beneficial technical effects:

[0060] 1) Use the encoder to encode the road scene information, surrounding vehicle information including the current state and historical trajectory of the vehicle, the current state and historical trajectory and relative position information of the target vehicle, and generate a high-dimensional feature tensor.

[0061] 2) The encoded high-dimensional feature tensor is subjected to hierarchical feature fusion. First, the attention mechanism is used to perform local feature fusion on vehicle features and map features respectively. Then, Transformer is used to perform global feature fusion on vehicle features, map features and relative position information, thereby enhancing the interaction of scene element features.

[0062] 3) Using dynamic weight MLP to output multimodal trajectory endpoint predictions in different driving scenarios not only simplifies the model structure but also improves the accuracy of endpoint prediction.

[0063] 4) The attention mechanism is used to interactively output the trajectory endpoint information and the trajectory endpoint offset, so as to achieve feature interaction in a longer time and space range to obtain more accurate endpoint information. Finally, the trajectory is completed to output a complete multimodal prediction trajectory, which improves the rationality and comfort of the output trajectory.

[0064] 5) Design a suitable loss function based on the multimodal trajectory prediction model. The loss function includes three items: the weighted sum of the regression loss of all trajectories, the classification loss and the endpoint regression loss, so that the predicted multimodal trajectory is closer to the true trajectory. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 This is a structural diagram of a multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction according to the present invention;

[0066] Figure 2 This is a local feature fusion structure diagram based on the attention mechanism of the present invention;

[0067] Figure 3 This is a diagram of the global feature fusion structure based on Transformer of the present invention;

[0068] Figure 4 This is a structural diagram of the endpoint prediction based on the dynamic weight MLP of the present invention;

[0069] Figure 5 It is the result of multimodal trajectory prediction of the present invention in the embodiment. DETAILED DESCRIPTION

[0070] The following will take an intersection without traffic lights in an urban traffic scenario as an example, and combine the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0071] Example 1: Figure 1 As shown, a multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction includes the following steps:

[0072] Step 1: Use the encoder to encode the road scene information, surrounding vehicle information including the current state and historical trajectory of the vehicle, the current state and historical trajectory and relative position information of the target vehicle, and generate a high-dimensional feature tensor.

[0073] Step 2: Perform hierarchical feature fusion on the encoded high-dimensional feature tensor. First, use the attention mechanism to perform local feature fusion on vehicle features and map features respectively, and then use Transformer to perform global feature fusion on vehicle features, map features and relative position information.

[0074] Step 3: Input the fused feature tensor into the dynamic weight MLP to output multimodal trajectory endpoint predictions in different driving scenarios.

[0075] Step 4: Use the attention mechanism to interact with the trajectory endpoint information and output the trajectory endpoint offset, so as to achieve feature interaction in a longer time and space range to obtain more accurate endpoint information, and finally complete the trajectory to output a complete multimodal prediction trajectory.

[0076] Step 5: Design a suitable loss function based on the multimodal trajectory prediction model. The loss function includes three items: regression loss of all trajectories, classification loss, and endpoint regression loss.

[0077] The local feature fusion in step 2 is based on the attention mechanism structure as follows Figure 2 As shown, specifically:

[0078] In local feature fusion, a single-head attention mechanism is used to realize trajectory feature interaction and map feature interaction respectively, which can not only realize the interactive function but also improve the efficiency of the interaction.

[0079] The trajectory feature tensor and map feature tensor output by the feature extraction module are input separately, and different weights are assigned to the scene element features through the single-head self-attention mechanism, where F represents the input feature tensor and F' represents the feature tensor after the attention mechanism. The specific formula is:

[0080] Q=W q F,K=W k F,V=W V F

[0081]

[0082] Where W q ,W k ,W V is the weight matrix. Finally, the interactive feature tensor is obtained through the layer normalization network and the feedforward network. Q, K, and V are the input vectors of the attention mechanism.

[0083] The global feature fusion in step 2 is based on the structure of symmetric fusion Transformer (SFT) as follows Figure 3 As shown, the specific steps are:

[0084] S1: Concatenate the local fused vehicle trajectory feature tensor and map feature tensor to obtain tensor V n , the tensor size is [N,128], where N is the sum of vehicles and lane segments in the scene.

[0085] S2: V n Repeat and stack along two different dimensions to obtain the target tensor and source tensor, both of size [N,N,128].

[0086] S3: The target tensor, source tensor and relative position feature tensor are concatenated and passed through an MLP to obtain the input K, V of the multi-head attention. The Q of the multi-head attention is the tensor V n .

[0087] S4: Output the updated vehicle trajectory feature tensor V after the multi-head attention mechanism a 'and the relative position feature tensor rpe', where the number of attention heads is 8. Tensor V a 'Size is [N a ,128], the size of tensor rpe' is [N,N,128], where N a is the number of vehicles in the scene. The global feature fusion module effectively integrates the scene traffic participant information and road structure information.

[0088] The structure of the endpoint prediction based on dynamic weight MLP in step 3 is as follows Figure 4 As shown, the specific steps are:

[0089] S1: Vehicle trajectory feature tensor V output by the multi-feature fusion module a 'Use an MLP layer to generate a tensor V w ' a .

[0090] S2: Introducing the trainable weight parameters W in the two MLP layers w1 and W w2 , so that V w ' a Two dynamic weight parameters W1 and W2 are generated by these two MLP layers, as shown in the formula:

[0091]

[0092] S3: The trajectory endpoint prediction process is transformed into a process that first passes through an MLP layer with a weight of W1, then passes through a normalization layer and an activation layer, and finally passes through an MLP layer with a weight of W2 to generate a multimodal endpoint ypred .

[0093] In step 4, the endpoint information of the trajectory is interactively refined based on the attention mechanism, and the specific steps are as follows:

[0094] S1: Encode the trajectory endpoint output by the endpoint prediction module to obtain the trajectory endpoint feature tensor E pred :

[0095] E pred =MLP(y pred )

[0096] S2: The concatenated vehicle trajectory feature tensor V a ' and the trajectory endpoint feature tensor E pred In the input attention mechanism, the output trajectory end offset b pred :

[0097] b pred =MHA(concat(V a ',E pred ))

[0098] S3: Use a simple MLP to complete the remaining trajectory points on the vehicle trajectory feature tensor and predict the probability of each mode.

[0099] The loss function designed in step 5 includes three parts: regression loss of all trajectories, classification loss and endpoint regression loss, specifically:

[0100] 1) Regression loss of all trajectories

[0101] The regression loss of all trajectories includes the vehicle position coordinate regression loss and heading angle loss The introduction of heading angle loss makes the predicted trajectory smoother and more feasible, as shown in the formula:

[0102]

[0103] Where T is the prediction time step; is the position coordinate of the trajectory of the optimal mode at time t; is the position coordinate of the true trajectory at time t; CosSim is the cosine similarity metric, which takes the value of 1 when the two vectors are in the same direction and -1 when they are in opposite directions; The heading angle of the vehicle when it is in the best mode; is the actual heading angle of the vehicle.

[0104] 2) Classification loss

[0105] The classification loss function is used to measure the difference between the category predicted by the model and the true category. The cross entropy loss is used to calculate the classification loss, and the performance of the model is evaluated by calculating the difference between the predicted probability distribution and the true label.

[0106] 3) End point regression loss

[0107] The endpoint regression loss is a loss function used to measure the difference between the endpoint position predicted by the model and the actual endpoint position. Its core purpose is to optimize the prediction performance of the model by minimizing this difference, as shown in the formula:

[0108]

[0109] In the formula The position coordinates of the end point of the trajectory representing the best mode; is the position coordinate of the end point of the real trajectory.

[0110] like Figure 5 As shown, the completion process of multimodal trajectory prediction of autonomous driving vehicles based on Transformer and endpoint induction in the present invention in urban scenes is demonstrated. In the present invention, red represents the vehicle itself, blue represents the surrounding vehicles, the real trajectory endpoint is represented by a star, the historical trajectory is represented by a solid line, the predicted multimodal trajectory is represented by a dotted line, and the final posture of the vehicle is represented by an arrow. It can be seen from the figure that the algorithm in the present invention can predict the future trajectories of all vehicles at the same time, and in a complex traffic environment with dense vehicles, the model accurately predicts to avoid collisions.

[0111] In summary, the present invention proposes a multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction. During the process of autonomous driving trajectory prediction, it can focus on important vehicles in the scene and ignore vehicles that are not related to itself. At the same time, it improves the accuracy of trajectory prediction and reduces the complexity of the model.

[0112] It should be noted that the above embodiments are not intended to limit the protection scope of the present invention, and equivalent changes or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction, characterized in that: The method comprises the following steps: Step 1: Use the encoder to encode the road scene information, surrounding vehicle information including the current state and historical trajectory of the vehicle, the current state and historical trajectory of the target vehicle, and the relative position information to generate a high-dimensional feature tensor. Step 2: Perform hierarchical feature fusion on the encoded high-dimensional feature tensor. First, use the attention mechanism to perform local feature fusion on vehicle features and map features respectively, and then use Transformer to perform global feature fusion on vehicle features, map features and relative position information. Step 3: Input the fused feature tensor into the dynamic weight MLP to output the multimodal trajectory endpoint prediction in different driving scenarios. Step 4: Use the attention mechanism to interact with the trajectory endpoint information and output the trajectory endpoint offset, so as to achieve feature interaction in a longer time and space range to obtain more accurate endpoint information, and finally complete the trajectory to output a complete multimodal prediction trajectory. Step 5: Design a suitable loss function based on the multimodal trajectory prediction model. The loss function includes three items: the weighted sum of the regression loss of all trajectories, the classification loss, and the endpoint regression loss.

2. The multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction according to claim 1, characterized in that: In step 1, the feature extraction module includes a vehicle trajectory encoder, a map information encoder and a relative position information encoder. The specific steps are as follows: S1: The trajectory information is a time series containing the vehicle position. The input of the vehicle trajectory encoder is expressed in displacement. In map encoding, the position and direction of the midpoint of the lane centerline are used to represent the position and direction of each lane segment. S2: The historical trajectory information of the i-th vehicle in H time steps is represented as a displacement sequence ΔX i ={ΔX i,-H+1 ,...,ΔX i,0 }, where ΔX i,t Represents the displacement of the i-th vehicle from time t-1 to time t. In order to improve the effectiveness of multi-scale feature extraction and the efficiency of parallel computing, one-dimensional residual convolution and feature pyramid network (FPN) are used to process the historical trajectory input. The vehicle feature tensor V output by the encoder is a The size is [N a ,128], where N a represents the number of vehicles in the historical step, S3: Considering that the lane segment is static during the observation time, a simple two-layer MLP is used to encode the map information to improve the operation efficiency. The map feature tensor V output by the encoder is m The size is [N l ,128], where N l represents the number of lane centerlines in the map during the observation time, S4: In relative position encoding, the heading angle difference α is used i→j , relative azimuth β i→j and distance ||d i→j ||These three quantities describe the relative position and posture between element i and element j, and the relative position information r between scene element i and element j i→j It can be represented by a five-dimensional vector: For the scenario with N elements in the observation time, MLP is used to encode the relative position information to obtain a relative position tensor rpe of size [N,N,128].

3. The multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction according to claim 1, characterized in that: In step 2, the multi-feature fusion method based on hierarchical feature fusion includes: 1) Local feature fusion, The behaviors of traffic participants in complex traffic scenes are often complex. In local feature fusion, the single-head attention mechanism is used to realize trajectory feature interaction and map feature interaction respectively, which can not only realize the interaction function but also improve the efficiency of interaction. The trajectory feature tensor and map feature tensor output by the feature extraction module are input separately, and different weights are assigned to the scene element features through the single-head self-attention mechanism, where F represents the input feature tensor and F' represents the feature tensor after the attention mechanism. The specific formula is: Q=W q F,K=W k F,V=W V F Where W q ,W k ,W V is the weight matrix, and finally the interactive feature tensor is obtained through the layer normalization network and the feedforward network. Q, K, and V are the input vectors of the attention mechanism. 2) Global feature fusion, Local feature fusion does not consider the relative position relationship of each element in the scene and the interaction between vehicles and lane segments. The Transformer model is widely used in the field of computer vision and feature fusion. Symmetric Fusion Transformer (SFT) is used for global feature fusion. This feature fusion method performs directional information transfer in a symmetrical manner, enabling the network to predict the future movement of all road users in one feedforward pass, which not only enables effective interaction but also improves model efficiency. The specific steps are: S1: Concatenate the local fused vehicle trajectory feature tensor and map feature tensor to obtain tensor V n , the tensor size is [N,128], where N is the sum of vehicles and lane segments in the scene, S2: V n Repeat and stack along two different dimensions to get the target tensor and source tensor, both of size [N,N,128], S3: The target tensor, source tensor and relative position feature tensor are concatenated and passed through an MLP to obtain the input K, V of the multi-head attention. The Q of the multi-head attention is the tensor V n , S4: Output the updated vehicle trajectory feature tensor V after the multi-head attention mechanism a 'and relative position feature tensor rpe', where the number of attention heads is 8 and the tensor V a 'Size is [N a ,128], the size of tensor rpe' is [N,N,128], where N a The global feature fusion module effectively integrates the scene traffic participant information and road structure information.

4. The multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction according to claim 1, characterized in that: The trajectory endpoint prediction method based on dynamic weight MLP in step 3 includes: S1: Vehicle trajectory feature tensor V output by the multi-feature fusion module a 'Use an MLP layer to generate a tensor V w ' a , S2: Introducing the trainable weight parameters W in the two MLP layers w1 and W w2 , so that V w ' a Two dynamic weight parameters W1 and W2 are generated by these two MLP layers, as shown in the formula: S3: The trajectory endpoint prediction process is transformed into a process that first passes through an MLP layer with a weight of W1, then passes through a normalization layer and an activation layer, and finally passes through an MLP layer with a weight of W2 to generate a multimodal endpoint y pred .

5. The multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction according to claim 1, characterized in that: The endpoint refinement method based on trajectory endpoint information interaction in step 4 includes: S1: Encode the trajectory endpoint output by the endpoint prediction module to obtain the trajectory endpoint feature tensor E pred : E pred =MLP(y pred ) S2: The concatenated vehicle trajectory feature tensor V a ' and the trajectory endpoint feature tensor E pred In the input attention mechanism, the output trajectory end offset b pred : b pred =MHA(concat(V a ',E pred )) S3: Use a simple MLP to complete the remaining trajectory points on the vehicle trajectory feature tensor and predict the probability of each mode.

6. The multimodal trajectory prediction method for autonomous driving vehicles based on Transformer and endpoint induction according to claim 1, characterized in that: In step 5, the loss function includes three items: regression loss of all trajectories, classification loss, and endpoint regression loss, specifically: 1) Regression loss of all trajectories, The regression loss of all trajectories includes the vehicle position coordinate regression loss and heading angle loss The introduction of heading angle loss makes the predicted trajectory smoother and more feasible, as shown in the formula: Where T is the prediction time step; is the position coordinate of the trajectory of the optimal mode at time t; is the position coordinate of the true trajectory at time t; CosSim is the cosine similarity metric, which takes the value of 1 when the two vectors are in the same direction and -1 when they are in opposite directions; The heading angle of the vehicle when it is in the best mode; is the actual heading angle of the vehicle, 2) Classification loss, The classification loss function is used to measure the difference between the category predicted by the model and the actual category. The cross entropy loss is used to calculate the classification loss. The performance of the model is evaluated by calculating the difference between the predicted probability distribution and the actual label. 3) End point regression loss, The endpoint regression loss is a loss function used to measure the difference between the endpoint position predicted by the model and the actual endpoint position. Its core purpose is to optimize the prediction performance of the model by minimizing this difference, as shown in the formula: In the formula The position coordinates of the end point of the trajectory representing the best mode; is the position coordinate of the end point of the real trajectory.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction is implemented as described in any one of claims 1 to 6 above.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instruction is executed by the processor, the multimodal trajectory prediction method for an autonomous driving vehicle based on Transformer and endpoint induction as described in any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Multi-modal trajectory prediction method

    CN117302256A

  • Automatic driving vehicle motion planning method based on rule enhanced trajectory prediction

    CN117571011A

  • Vehicle trajectory prediction method and model based on improved Transform model and target point guidance, and electronic equipment

    CN118823731A

  • Trajectory prediction method based on hierarchical progressive interaction and target lane segment

    CN119099654A

  • Road traffic speed prediction method fusing multi-feature neural network

    WO2024244300A1

Cited By

  • Vehicle multi-modal trajectory prediction method taking graph as center

    CN120182937A

  • Internet of vehicles channel prediction method based on multi-modal fusion and related equipment

    CN120342527A

  • Transform-based automatic driving long time sequence track prediction method and system

    CN120408105A

  • Method and device for obtaining structured data of road scene, medium and product

    CN120744021A

  • Vehicle control method, device, equipment and medium

    CN120792858A