Automatic driving track prediction method based on random covering of vehicle and road information
By constructing graph structure and masking mechanism, the prediction capability of the trajectory prediction model in the data sparse area is optimized, and the problems of insufficient prediction capability of the existing model in the data sparse area and insufficient modeling of multi-vehicle interaction are solved, achieving higher trajectory prediction accuracy and safety.
Patent Information
- Application Number
- CN202510488979.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-01
AI Technical Summary
The existing trajectory prediction model performs well in training data-intensive areas, but the prediction capability in sparse data areas has decreased, and the interaction modeling of multiple vehicles is insufficient. Especially when some target information is missing, the interaction reasoning capability decreases, resulting in low trajectory prediction accuracy.
Using a method based on random masking vehicle road information, a graph structure is constructed by obtaining vehicle historical trajectory points and lane data, a local encoder and global encoder are used to capture local and global information, and a multimodal decoder is used to generate multiple prediction trajectories, and data loss is simulated through a masking mechanism, and prediction is optimized using multimodal decoder and loss function.
It improves the prediction ability in complex traffic scenarios, enhances the model's learning ability of unseen data, improves the safety and accuracy of trajectory prediction, and provides safety guarantees for autonomous vehicles.
Smart Images

Figure CN120236404A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving, and in particular to an autonomous driving trajectory prediction method based on randomly masking vehicle-road information. Background Art
[0002] With the rapid development of autonomous driving technology, the intelligence level of vehicles continues to improve. As the core component of the autonomous driving system, the trajectory prediction module is becoming increasingly important.
[0003] At present, mainstream trajectory prediction models perform well in areas with dense training data (such as intersections), but their prediction ability drops sharply in areas with sparse training data (such as roundabouts and construction sections). The essence is that the feature representation is not sensitive enough to the scene area, which leads to low accuracy in target vehicle trajectory prediction. In addition, the existing models do not adequately model the implicit interactions between multiple vehicles, especially when some target information is missing, the interactive reasoning ability is reduced.
[0004] Therefore, in trajectory prediction models, the interpretability of multimodal output (i.e., generating multiple possible future trajectories) is crucial. Strengthening regional semantic representation, avoiding confusion between similar trajectories, simulating the lack of real data, and forcing the model to explore deeper interaction logic have become important research directions for improving the performance of autonomous driving systems. Summary of the invention
[0005] The purpose of the present invention is to provide an automatic driving trajectory prediction method based on randomly masking vehicle-road information to solve the problem of low accuracy of existing trajectory prediction models.
[0006] In order to solve the above technical problems, the present invention provides an automatic driving trajectory prediction method based on randomly masking vehicle-road information, comprising:
[0007] Acquire historical trajectory points and lane data of the vehicle, convert the acquired coordinates of the historical trajectory points into a coordinate system with the vehicle as the origin, and generate historical trajectory data; then filter out the lane data and historical trajectory data within a local range with the vehicle as the center and a set distance as the radius;
[0008] The historical trajectory points in the filtered historical trajectory data are used as nodes of a graph structure, and the interactions between the historical trajectory points are used as edges of the graph structure to construct a graph structure; a local encoder is used to capture local information of the vehicle within the local range according to the graph structure; and then a global encoder is used to fuse the local information output by the local encoder according to the node and edge information of the graph structure;
[0009] A masking mechanism is used to mask the historical trajectory data and lane data in the fused local information;
[0010] The masked information is decoded using a multi-modal decoder to generate multiple predicted trajectories, and a confidence score is assigned to each of the predicted trajectories.
[0011] Further, the method for converting the historical trajectory point coordinates to a coordinate system with the vehicle as the origin specifically includes: obtaining the visible historical trajectory points in the historical time steps, then taking the last historical position of the vehicle as the origin, and rotating all the historical trajectory points around the vehicle's orientation to convert the coordinates of each historical trajectory point to a coordinate system with the vehicle as the origin.
[0012] Further, the lane data includes lane centerlines, intersection information, steering information, and traffic control information; the method for obtaining the lane data includes: extracting the centerline coordinates of the lane segments from a high-precision map, then in a coordinate system with the vehicle as the origin, representing the shape of the lane using a vector sequence, and attaching intersection information, steering information, and traffic control information to the lane represented by the vector sequence.
[0013] Further, the local encoder includes:
[0014] A vehicle-vehicle encoding module for capturing the interaction information between vehicles within the local range;
[0015] A spatio-temporal encoding module for capturing the temporal behavior information of vehicles within the local range;
[0016] A vehicle-lane encoding module for capturing the interaction information between vehicles and lanes within the local range.
[0017] Further, the global encoder uses a multi-head attention mechanism to extract the global relationships between nodes based on the node and edge information of the graph structure, uses a linear layer to embed the trajectory encoding as the query matrix Q, and then respectively uses two linear layers to embed the neighbor nodes and historical trajectory encodings as the key matrix K and value matrix V of the multi-head attention mechanism. The query matrix Q, key matrix K, and value matrix V are input into the multi-head attention mechanism, and the attention output is fused with the local information through a residual connection; the attention output is:
[0018]
[0019] where K node is the key generated by linearly transforming the feature x of the neighbor node j of the neighbor node, K edge is the key generated by linearly transforming the edge feature, V node is the value generated by linearly transforming the feature x of the neighbor node j of the neighbor node, and V edge is the value generated by linearly transforming the edge feature.
[0020] Further, the masking of the historical trajectory data in the fused local information by using the masking mechanism includes: for the historical trajectory data of each vehicle, randomly select a set proportion of time steps for masking, and the time steps of the masked historical trajectory data of two adjacent vehicles are complementary;
[0021] The masking of the lane data in the fused local information by using the masking mechanism includes: randomly masking the lane vector to simulate the missing lane data in the actual scenario.
[0022] Further, the total loss function of the masking mechanism is:
[0023] L mask = αL trajectory + βL lane
[0024] where L trajectory is the masking loss function of the historical trajectory data, and L lane is the masking loss function of the lane data; α and β are weight coefficients respectively;
[0025]
[0026] where i is the vehicle index, t is the time step index, x i,t is the original trajectory data, is the trajectory data reconstructed by the model; j is the lane index, l j is the original lane data, is the lane data reconstructed by the model.
[0027] Further, the generation formula for generating multiple predicted trajectories by using the multi-modal decoder is:
[0028] y loc = MLP loc (MLP aggr (h local + h global ))
[0029] where h local , h globa are the dynamic behavior information of the vehicle and the interaction information between the vehicle and the surrounding environment respectively; MLP aggr is a multi-layer perceptron that outputs the two-dimensional coordinates of the next T time steps.
[0030] Further, the total loss function of the multi-modal decoder is:
[0031] L total = λ1L ASC + λ2L iou + λ3Lcls
[0032] Among them, L ASC is the regional similarity comparison loss function, L iou is the classification function, and L cls is the regression function; λ1, λ2, and λ3 are weight coefficients respectively.
[0033] Furthermore, the regional similarity comparison loss function is as follows:
[0034]
[0035] Among them, s n,t is the cosine similarity between the two-modal prediction trajectory vectors of the nth sample at time step t, P is the set of positive sample pairs, and N is the set of negative sample pairs.
[0036] The beneficial effects of the present invention are as follows:
[0037] This method not only optimizes the feature extraction process in complex lane structures, but also masks part of the data through a random masking mechanism, forcing the model to learn to recover information from incomplete inputs, thereby improving the prediction ability for unseen data.
[0038] Using a multi-modal trajectory decoder and dynamically adjusting the weights through the ASC loss function, regression, and classification losses increases the judgment of the weights of the output trajectory features, meets the needs of multi-modal trajectory prediction, can balance safety, accuracy, and adaptability in complex traffic scenarios, provides a strong guarantee for the safety of autonomous driving vehicles, and thus ensures the safety of autonomous driving vehicles. Description of the Drawings
[0039] The drawings described herein are used to provide a further understanding of the present application, form a part of the present application, and the same reference numerals are used to represent the same or similar parts in these drawings. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0040] Figure 1 is the overall network structure diagram of an embodiment of the present invention;
[0041] Figure 2 is the trajectory data masking mechanism diagram of an embodiment of the present invention;
[0042] Figure 3 is the lane information masking mechanism diagram of an embodiment of the present invention;
[0043] Figure 4 is the complementary masking mechanism diagram of the trajectory data masking of an embodiment of the present invention. Detailed Embodiments
[0044] The present disclosure relates to an autonomous driving trajectory prediction method based on randomly masking vehicle-road information, including:
[0045] Obtain the historical trajectory points and lane data of the vehicle, convert the coordinates of the obtained historical trajectory points to a coordinate system with the vehicle as the origin, and generate historical trajectory data; then screen out the lane data and historical trajectory data within a local range centered on the vehicle with a set distance (such as 50 meters) as the radius; historical trajectory data is usually stored in the form of a CSV file, containing information such as the coordinates of trajectory points, timestamps, trajectory IDs, object types, etc. Read these data, organize all the extracted and processed data into a dictionary, and return it to the dataset class;
[0046] Use the historical trajectory points in the screened historical trajectory data as nodes of the graph structure, and use the interactions between the historical trajectory points as edges of the graph structure to construct the graph structure; use a local encoder to capture the local information of the vehicle within the local range according to the graph structure; then use a global encoder to fuse the local information output by the local encoder based on the node and edge information of the graph structure.
[0047] Adopt a masking mechanism to mask the historical trajectory data and lane data in the fused local information.
[0048] Use a multi-modal decoder to decode the information processed by the masking mechanism to generate multiple predicted trajectories, and assign a confidence score to each of the predicted trajectories.
[0049] This method not only optimizes the feature extraction process in complex lane structures, but also masks part of the data through a random masking mechanism, forcing the model to learn to recover information from incomplete inputs, thereby improving the prediction ability for unseen data.
[0050] According to an embodiment of the present application, the method for converting the coordinates of the historical trajectory points to a coordinate system with the vehicle as the origin specifically includes: obtaining the visible historical trajectory points in the historical time steps (the first 20 time steps), then using the last historical position of the vehicle (the 19th time step) as the origin, and rotating all the historical trajectory points around the orientation of the vehicle to convert the coordinates of each historical trajectory point to a coordinate system with the vehicle as the origin.
[0051] According to an embodiment of the present application, the lane data includes lane centerlines, intersection information, turning information, and traffic control information; the centerline coordinates of the lane segment are extracted from the high-precision map, and in the vehicle-centered coordinate system, the lane shape is represented by a vector sequence (lane_vectors), and the intersection information (is_intersections), turning information (turn_directions), and traffic control information (traffic_controls) are appended.
[0052] According to an embodiment of the present application, the local encoder includes:
[0053] A vehicle-vehicle encoding module for modeling local interactions between vehicles; calculating the attention weights between vehicles through the multi-head attention mechanism, embedding vehicle and edge features using single feature input SingleInputEmbedding and multi-feature input MultipleInputEmbedding, calculating the attention weights between vehicles through the multi-head attention mechanism (Multi-HeadAttention), and enhancing the model stability using residual connections and layer normalization (LayerNorm).
[0054] A spatio-temporal encoding module for modeling the temporal behavior of vehicles; using a Transformer encoder to capture temporal dependencies, combining historical trajectory data with positional encoding, adding positional encoding and CLS tokens to the input to retain temporal information and global context;
[0055] A vehicle-lane encoding module for modeling the interaction between vehicles and lanes; calculating the attention weights between vehicles and lanes through the multi-head attention mechanism, embedding lane and edge features using multi-feature input MultipleInputEmbedding, generating interaction features through weighted summation, and enhancing the stability of the features using residual connections and layer normalization.
[0056] The local encoder enables it to be flexibly applied to different scenarios and tasks through modular design, and at the same time supports expansion to handle more complex interaction relationships.
[0057] According to an embodiment of the present application, the global encoder uses the multi-head attention mechanism to extract the global relationships between nodes based on the node and edge information of the graph structure, uses a linear layer to embed the trajectory encoding as the query matrix Q (Query), and then uses two linear layers to embed the neighbor node and historical trajectory encoding as the key matrix K (Key) and value matrix V (Value) of the multi-head attention mechanism respectively. The query matrix Q, key matrix K, and value matrix V are input into the multi-head attention mechanism, and the attention output is fused with the input features through residual connections; the attention output is:
[0058]
[0059] Among them, K node is the key generated by the linear transformation of the feature x of the neighbor node j K, which is the key generated by the linear transformation edge V is the key generated by the linear transformation of the edge feature node which is the x of the neighbor node j V is the value generated by the linear transformation of the feature of the neighbor node edge V is the value generated by the linear transformation of the edge feature; the softmax function ensures the normalization of the attention weights in the neighbor node dimension, and the dropout function is used to prevent overfitting. Finally, the attention output is fused with the input feature through a residual connection to retain the original information, and then passed into the masking module
[0060] According to an embodiment of the present application, the masking module uses a masking mechanism to mask the historical trajectory data (create_trajectory_masks), including: for the historical trajectory data of each vehicle, randomly select a set proportion of time steps (controlled by masking_rate) for masking, and the time steps of the masked historical trajectory data of two adjacent vehicles are complementary, that is, the time steps masked by one vehicle are retained in the other vehicle, and vice versa
[0061] Using a masking mechanism to mask the lane data (create_lane_masks) includes: randomly masking the lane vector to simulate the missing lane data in the actual scenario
[0062] In addition, the decoder also supports uncertainty modeling and can predict the confidence of the trajectory
[0063] According to an embodiment of the present application, the total loss function of the masking mechanism is
[0064] L mask =αL trajectory +βL lane
[0065] Among them, L trajectory is the masking loss function of the historical trajectory data, and L lane is the masking loss function of the lane data; α and β are weight coefficients respectively, and the weight coefficients α and β are used to balance the two parts of the loss
[0066]
[0067] Among them, i is the vehicle index, t is the time step index, and x i,t is the original trajectory data is the trajectory data reconstructed by the model; j is the lane index, and lj is the original lane data, is the lane data reconstructed by the model.
[0068] By introducing a complementary masking mechanism, the model in this paper can still maintain a high prediction accuracy in the case of partial information loss, providing an effective training strategy for dynamic spatio-temporal data modeling. The loss calculation of the masking mechanism guides the model to learn to recover the masked data by calculating the mean square error between the predicted value and the true value of the masked part.
[0069] According to an embodiment of the present application, the generation formula for generating multiple prediction trajectories using a multi-modal decoder is:
[0070] y loc = MLP loc (MLP aggr (h local + h global ))
[0071] where h local , h globa are the dynamic behavior information of the vehicle and the interaction information between the vehicle and the surrounding environment respectively; MLP aggr is a multi-layer perceptron that outputs two-dimensional coordinates for the next T time steps. The final trajectory prediction result can output a total of F different trajectories, and the probabilities of different trajectories can be estimated. In addition, the decoder also supports uncertainty modeling and can predict the confidence of the trajectory.
[0072] According to an embodiment of the present application, the total loss function of the multi-modal decoder is:
[0073] L total = λ1L ASC + λ2L iou + λ3L cls
[0074] where, L ASC is the regional similarity contrast loss function, L iou is the classification function, L cls is the regression function; λ1, λ2, λ3 are the weight coefficients respectively.
[0075] The total loss function of the multi-modal decoder consists of the ASC loss function, the regression loss function, and the classification loss function; after the decoder outputs the predicted value (y_hat), the ASC loss is calculated in combination with the input data (data). The ASC loss is weighted with the regression loss and the classification loss according to dynamic weights to obtain the total loss and perform backpropagation. This design enables the model to simultaneously achieve accurate trajectory prediction, reliable behavior classification, and robust regional consistency perception in complex autonomous driving scenarios.
[0076] According to an embodiment of the present application, the regional similarity contrast loss function is as follows:
[0077]
[0078] where s n,t is the cosine similarity between the two-modal prediction trajectory vectors of the nth sample at time step t, P is the set of positive sample pairs, and N is the set of negative sample pairs.
[0079] The regional similarity contrast loss (ASC) is a novel loss function. It learns the consistency information between different modalities by maximizing the predicted similarity between different modalities, and learns the difference information between different modalities by minimizing the similarity between negative sample pairs. It can effectively utilize unlabeled multi-modal data to improve the model performance. The model can enhance the learning of information between different modalities and the improvement of prediction consistency by introducing the regional similarity contrast loss (ASC). The ASC loss uses cross-modal information and the prediction consistency between different modalities for contrastive mutual learning, which helps to improve the model's ability to capture complementary information in multi-modal data. The ASC loss can also be combined with local encoders and global interaction modules to optimize the feature fusion process, enabling the model to better understand and predict the behaviors of traffic participants.
[0080] The regional similarity contrast loss is combined with the regression loss and classification loss, and a training strategy of dynamic weight scheduling is added during the training process. In the early stage of training, the focus is on the ASC loss weight, forcing the model to learn regional consistency features; in the later stage of training, the focus is on the regression and classification losses, fine-tuning the trajectory details and pattern probabilities. Through the AdamW optimizer, with the learning rate gradually decreasing as the main learning strategy, a multi-modal trajectory prediction module is obtained, which provides policy information for the path planning module in the autonomous driving system.
[0081] The autonomous driving trajectory prediction method is evaluated on the large-scale Argoverse motion prediction benchmark. According to the results of the three core evaluation indicators of ADE, FDE, and MR in the trajectory prediction task, the average displacement error of the existing model for ADE is 0.69; the final displacement error of FDE is 1.03; the miss detection rate of MR is 0.103; the results obtained by processing the same data set with this method are: the average displacement error of ADE is 0.68; the final displacement error of FDE is 1.02; the miss detection rate of MR is 0.101; it can be seen that this method has achieved excellent performance in trajectory prediction.
[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. An automatic driving trajectory prediction method based on randomly masking vehicle-road information, characterized in that: include: Acquire historical trajectory points and lane data of the vehicle, convert the acquired coordinates of the historical trajectory points into a coordinate system with the vehicle as the origin, and generate historical trajectory data; then filter out the lane data and historical trajectory data within a local range with the vehicle as the center and a set distance as the radius; The historical trajectory points in the filtered historical trajectory data are used as nodes of a graph structure, and the interactions between the historical trajectory points are used as edges of the graph structure to construct a graph structure; Using a local encoder to capture local information of the vehicle within the local range according to the graph structure; then using a global encoder to fuse the local information output by the local encoder according to the node and edge information of the graph structure; A masking mechanism is used to mask the historical trajectory data and lane data in the fused local information; A multimodal decoder is used to decode the information processed by the masking mechanism to generate multiple prediction trajectories, and a confidence score is assigned to each of the prediction trajectories.
2. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 1, characterized in that: The method used to convert the coordinates of the historical trajectory points into a coordinate system with the vehicle as the origin specifically includes: obtaining the visible historical trajectory points in the historical time step, then taking the last historical position of the vehicle as the origin, and rotating all the historical trajectory points around the direction of the vehicle, and converting the coordinates of each historical trajectory point into a coordinate system with the vehicle as the origin.
3. The automatic driving trajectory prediction method based on randomly masking vehicle-road information according to claim 2, characterized in that: The lane data includes lane centerline, intersection information, turn information and traffic control information; The lane data acquisition method includes: extracting the centerline coordinates of the lane segment from the high-precision map, then using a vector sequence to represent the shape of the lane in a coordinate system with the vehicle as the origin, and adding intersection information, turning information and traffic control information to the lane represented by the vector sequence.
4. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 3, characterized in that: The local encoder comprises: A vehicle-vehicle encoding module is used to capture the interaction information between vehicles in the local range, calculate the attention weights between vehicles through a multi-head attention mechanism, and embed the historical trajectory points and edge features of the vehicles using single feature input and multi-feature input; A spatiotemporal coding module, used to capture the temporal behavior information of vehicles within the local range, combine historical trajectory points with position coding, and add position coding and CLS tags to the input; The vehicle-lane encoding module is used to capture the interaction information between the vehicle and the lane within the local range, calculate the attention weight between the vehicle and the lane using a multi-head attention mechanism, embed the lane and edge features using multi-feature inputs, and generate interaction features through weighted summation.
5. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 4, characterized in that: The global encoder uses a multi-head attention mechanism to extract the global relationship between nodes according to the nodes and edge information of the graph structure, uses a linear layer to embed the trajectory code as the query matrix Q, and then embeds the neighbor nodes and historical trajectory codes as the key matrix K and value matrix V of the multi-head attention mechanism through two linear layers. The query matrix Q, key matrix K and value matrix V are input into the multi-head attention mechanism, and the attention output is fused with the local information through residual connection; the attention output is: Among them, K node is the feature x of the neighbor node j The key generated by linear transformation, K edge is the key generated by linear transformation of edge features, V node is the feature x of the neighboring nodes j The value generated by the linear transformation, V edge is the value generated by linear transformation of edge features.
6. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 5, characterized in that: The masking mechanism is used to mask the historical trajectory data in the fused local information, including: randomly selecting a set proportion of time steps for masking the historical trajectory data of each vehicle, and the time steps of the masked historical trajectory data of two adjacent vehicles are complementary; The use of a masking mechanism to mask the lane data in the fused local information includes: randomly masking the lane vectors to simulate the lack of lane data in an actual scene.
7. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 6, characterized in that: The total loss function of the masking mechanism is: L mask =αL trajectory +βL lane Among them, L trajectory is the masking loss function of historical trajectory data, L lane is the masking loss function of lane data; α and β are weight coefficients respectively; Where i is the vehicle index, t is the time step index, and x i,t is the original trajectory data, is the trajectory data reconstructed by the model; j is the lane index, l j is the original lane data, Lane data reconstructed for the model.
8. The automatic driving trajectory prediction method based on randomly masking vehicle-road information according to claim 7, characterized in that: The generation formula for generating multiple prediction trajectories using the multimodal decoder is: y loc =MLP loc (MLP aggr (h local +h global )) where h local ,h globa They are the dynamic behavior information of the vehicle and the interaction information between the vehicle and the surrounding environment; MLP aggr is a multi-layer perceptron that outputs the two-dimensional coordinates of the next T time steps.
9. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 8, characterized in that: The total loss function of the multimodal decoder is: L total =λ1L ASC +λ2L iou +λ3L cls Among them, L ASC is the region similarity comparison loss function, L iou is the classification function, L cls is the regression function; λ1, λ2, λ3 are weight coefficients respectively.
10. The method for predicting automatic driving trajectory based on randomly masking vehicle-road information according to claim 9, characterized in that: The region similarity comparison loss function L ASC for: Among them, s n,t The cosine similarity between the two modal prediction trajectory vectors of the nth sample at time step t, P is the set of positive sample pairs, and N is the set of negative sample pairs.
Citation Information
Cited By
Multi-vehicle trajectory prediction method for complex dynamic traffic scene
CN121148174A
Online automatic thematic map making method and system based on artificial intelligence
CN121616699A