Vehicle multi-modal trajectory prediction method taking graph as center

By constructing local and global lane maps, combining lane convolution and multi-head attention layers, the problems of low spatial interaction capabilities and insufficient utilization of environmental information in the existing methods are solved, and accurate trajectory prediction in complex traffic environments is achieved.

CN120182937AActive Publication Date: 2025-06-20JILIN UNIVERSITY

Patent Information

Application Number
CN202510637447.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-06-20
Estimated Expiration
2045-05-19

AI Technical Summary

Technical Problem

The existing deep learning methods ignore the interaction relationships between traffic participants and the impact of environmental information in trajectory prediction, make it difficult to deal with complex traffic scenarios, and lack a flexible global interaction mechanism, affecting the accuracy and reliability of predictions.

Method used

Using a graph-centric approach, by constructing local lane maps and global lane maps, using a technology combining lane convolution and multi-head attention layer to capture the relationship between traffic participants and lanes and complex interactions between agents, and perform feature extraction and prediction.

Benefits of technology

In complex road environments, it is possible to accurately capture the multiple target positions of traffic-participated vehicles, improve the accuracy and robustness of trajectory prediction, and enhance the structured modeling capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182937A_ABST
    Figure CN120182937A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of automatic driving, and discloses a graph-centered vehicle multi-modal trajectory prediction method, which comprises the following steps of: constructing a local graph of each traffic participating vehicle; after the geometric features and the type features of the lane segments and the dynamic interaction relation between the lane segments and the traffic participating vehicles are spliced, feature extraction is carried out, and node features of a local graph are obtained; projecting all the coded local maps to a global lane map to obtain a global map; after node features of all the local graphs are aggregated, lane convolution operation is carried out, and local feature fusion vectors are obtained; performing feature extraction on the local feature fusion vector to obtain a global interaction feature vector; performing weighted combination on the local feature fusion vector and the global interaction feature vector to obtain node features of a global graph; projecting node features of the global graph back to the local graph to obtain an enhanced local graph; and according to the enhanced local map, obtaining a prediction track of each traffic participating vehicle and the confidence coefficient of each prediction track.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to a graph-centric multi-modal trajectory prediction method for vehicles. Background Art

[0002] With the development of intelligent transportation systems, the application of autonomous vehicles (AVs) in complex traffic environments is increasing. Precise trajectory prediction is one of the key technologies to ensure the safety and efficiency of autonomous driving. The goal of trajectory prediction is to predict the possible movement trajectories of traffic participants (such as vehicles, pedestrians, etc.) within a certain future time based on their historical movement data and environmental information. This process is crucial for realizing functions such as autonomous driving decision-making, path planning, and obstacle avoidance control.

[0003] Existing trajectory prediction methods can be roughly divided into three categories: models based on traditional methods, models based on deep learning, and models based on reinforcement learning. Traditional methods generally rely on physical models or rule-based methods to predict trajectories. These methods are often limited by simple assumptions and are difficult to handle dynamic and complex traffic scenarios. In recent years, deep learning methods, especially trajectory prediction models based on technologies such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and graph neural networks (GNNs), have made significant progress. By learning historical trajectories and environmental information, these models can better capture the movement patterns of agents and complex traffic interactions.

[0004] However, existing deep learning methods usually have some deficiencies. First, most models represent the historical trajectories of traffic participants and environmental information as single vectors or feature vectors respectively. This way ignores the interaction relationships between different traffic participants and the influence of environmental information (such as lanes, traffic lights, etc.). Second, when dealing with complex scenarios, traditional models fail to effectively utilize the spatial and topological structures in traffic scenarios, especially in scenarios with complex traffic rules such as intersections and roundabouts. Finally, existing models often lack a flexible global interaction mechanism and are difficult to accurately capture long-distance interactions between agents, thus affecting the accuracy and reliability of trajectory prediction. Summary of the Invention

[0005] The purpose of the present invention is to provide a graph-centric multi-modal trajectory prediction method for vehicles, which overcomes the problems of low spatial interaction ability and poor combination with lane environment in existing methods, can capture multiple target positions of traffic participating vehicles in complex road environments, and finally output multi-modal prediction trajectories of vehicles.

[0006] The technical solution provided by the present invention is as follows: A graph-centered multi-modal trajectory prediction method for vehicles, comprising: Obtain the historical trajectories of traffic participating vehicles and the lane lines corresponding to the historical trajectories, and segment the lane lines corresponding to the historical trajectories into lane segments; Using the center points of the lane segments as nodes and the positional relationships between the center points of adjacent lane segments as edges, construct a local graph for each traffic participating vehicle; After splicing the geometric features, type features of the lane segments and the dynamic interaction relationships between the lane segments and traffic participating vehicles, perform feature extraction to obtain the node features of the local graph; Construct a global lane graph containing all lanes in the scene, encode each local graph of the traffic participating vehicles respectively, and project all the encoded local graphs onto the global lane graph to obtain a global graph; After aggregating the node features of all local graphs, perform lane convolution operations on the aggregated feature vectors to obtain a locally feature-fused vector; Extract features from the locally feature-fused vector through a global interaction relationship encoding layer to obtain a globally interactive feature vector; Weightedly combine the locally feature-fused vector and the globally interactive feature vector to obtain the node features of the global graph; Project the node features of the global graph back to the local graphs of each traffic participating vehicle to obtain enhanced local graphs; Obtain the predicted trajectories of each traffic participating vehicle and the confidence of each predicted trajectory according to the enhanced local graphs.

[0007] Preferably, the method for encoding each local graph of the traffic participating vehicle is: Perform lane convolution operations on the node features of the local graph, perform average pooling operations after each layer of convolution, and add the results of the convolution operations and the average pooling operations to obtain the encoded local graph node feature vectors.

[0008] Preferably, the global interaction relationship encoding layer is composed of a multi-head attention layer.

[0009] Preferably, obtaining the predicted trajectories of each traffic participating vehicle and the confidence of each predicted trajectory according to the enhanced local graphs includes the following steps: Step 1: Use a first prediction unit to predict each node in the enhanced local graph to determine the probability of the node being a target point and the position residual from the node to the true target point; Step 2: Select the K nodes with the highest probability as candidate target points, perform information interaction on the candidate target points through a multi-head attention layer to obtain interaction features; input the interaction features into the second prediction unit to obtain the probability of the candidate target point being the target point and the position residual between the candidate target point and the true target point. Step 3: Use a Bessel curve to fit the trajectory from the current moment coordinate point to the candidate target point to obtain a reference trajectory. Step 4: Fine-tune the reference trajectory to obtain a predicted trajectory, and use the probability of the candidate target point being the target point as the confidence of the corresponding predicted trajectory.

[0010] Preferably, both the first prediction unit and the second prediction unit adopt a multi-layer perceptron.

[0011] Preferably, the method for fine-tuning the reference trajectory to obtain a predicted trajectory in Step 4 is as follows: Obtain the positions of the vehicle at multiple future moments through the reference trajectory, determine the offset at each moment through a multi-layer perceptron, and fine-tune the positions at the multiple future moments according to the offset to obtain a predicted trajectory.

[0012] The beneficial effects of the present invention are: The graph-centered vehicle multi-modal trajectory prediction method provided by the present invention makes full use of the lane environment around the intelligent agent, can well capture the relationship between the intelligent agent and the lane, and thus accurately predicts the future trajectory of the intelligent agent in a complex traffic environment.

[0013] Through the combination of lane convolution and a multi-head attention layer, the present invention can more comprehensively model the complex relationships between traffic participating vehicles, improve the accuracy of trajectory prediction; and perform weighted fusion on the features output by LaneGCN and the attention mechanism, determine the weights through a learning method, enable the model to adjust the attention degree to local and global features according to the dynamics of the actual scenario, and enhance the robustness of the prediction model.

[0014] In the prediction head part of the present invention, the local graph of the intelligent agent is combined for optimization, which can make the predicted trajectory more in line with the road characteristics and make the prediction more accurate and reasonable.

[0015] The present invention adopts a two-stage decoder, and predicts the target point and trajectory in stages through "coarse screening + fine tuning", effectively improving the accuracy, diversity, stability and expression ability of the prediction, and endowing the decoder with stronger structured modeling ability. Description of the Drawings

[0016] Figure 1 It is a flowchart of the graph-centered vehicle multi-modal trajectory prediction method described in the present invention.

[0017] Figure 2 This is the framework diagram of the multi-vehicle joint trajectory prediction model of the present invention. Specific embodiments

[0018] The following further elaborates on the present invention with reference to the accompanying drawings, enabling those skilled in the art to implement it based on the description in the specification.

[0019] As Figure 1-2 shown, the present invention provides a graph-centered multi-modal vehicle trajectory prediction method, and the specific implementation process is as follows.

[0020] I. Construct a multi-modal vehicle trajectory prediction model and train the multi-modal vehicle trajectory prediction model The multi-modal vehicle trajectory prediction model includes: An encoder, a global interaction module, and a decoder module.

[0021] The encoder includes a multi-layer perceptron and a lane convolution layer. The global interaction module includes a multi-layer perceptron 、a multi-layer perceptron a lane convolution layer, and a multi-head attention layer. The decoder module includes a first-stage prediction head and a second-stage prediction head; the first-stage prediction head uses a multi-layer perceptron ; the second-stage prediction head consists of a multi-head attention layer, a multi-layer perceptron and a multi-layer perceptron .

[0022] The specific training process is as follows: 1. Scene initialization Represent the past motion of the th traffic participating vehicle as a set of two-dimensional point sets , encoding the central positions of the traffic participating vehicles within the past time steps, that is: ; Among them, represents the absolute horizontal and vertical coordinates of the center point of the th traffic participating vehicle at the past moment in the map coordinate system, and the superscript represents that the coordinate is a historical trajectory.

[0023] At the same time, the representation of the th lane line in the scene is also represented as a set of two-dimensional point sets , and the center line of each lane line is fixedly divided into Segments are generally divided at intervals of 2m along the centerline of the lane lines. The center point coordinates of the divided lane segments are used as the position coordinates of the lane segments, encoding the lane features near the agent: ; Among them, represents the absolute horizontal and vertical coordinates of the center point of the th lane segment of the th lane line in the map coordinate system.

[0024] The target to be predicted in the present invention is the coordinate points of the possible driving trajectories of N participating vehicles in the scene within the next T time steps: ; Among them, represents the absolute horizontal and vertical coordinates of the center point of the th traffic participating vehicle at the future th moment in the map coordinate system. The superscript represents that the coordinate is a predicted trajectory.

[0025] 2. Construction of the local map of All lanes that the traffic participating vehicle may reach within the predicted time T are obtained through the existing data set , and at the same time, all lanes that the vehicle has passed through during the observation historical time L are also extracted. Then, these lanes are divided into multiple consecutive lane segments at intervals of 2m along the lane centerline. The center point of each lane segment is used as the node representation of the graph , and these nodes are pairwise connected to each other in structure, forming the following directed graph: ; Among them, each node in the local map represents the position of the th lane segment of the th traffic participating vehicle in the local map. represents the topological structure of the lane, encoding the structural relationships of predecessors, successors, left neighbors, and right neighbors, which are represented by four types of connecting edges respectively. represents the number of nodes in each local map.

[0026] For the embedding feature vector representation of each node : ; Among them, represents the geometric feature of the lane segment (node). It represents the type feature of the lane segment. It represents the dynamic interaction relationship between the lane segment and the agent, where C represents the dimension of the node embedding feature.

[0027] In this embodiment, it is set that ; In the above formula, represents the central point coordinates of the center line of the lane segment, represents the orientation of the lane segment, represents the curvature of the lane segment. In other embodiments, some road coefficients of the lane segment such as the friction coefficient can also be considered. If factors are added, it will cause the geometric feature dimension to change, represents the geometric feature dimension.

[0028] For the type feature of the lane segment it is as follows: ; In the above formula, represents the lane end category, including straight lane, turning lane, fork lane and ramp, these four types; after encoding them, they correspond to the vectors [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 1, 0], [0, 0, 0, 1] respectively. represents whether the lane segment is controlled by a red light, represents that the lane segment is controlled by a red light; represents whether the lane segment is the target lane, represents that the lane segment is the target lane of the agent. S represents the type feature dimension.

[0029] For since the past movement of traffic participating vehicles can be represented as a set of coordinate points , that is, it defines the vehicle center position at consecutive time steps. Therefore, these coordinate points are transformed into a displacement sequence , as shown in the following formula, and is projected onto the coordinate system of the local graph node to encode the movement information of the participants in the case related to the map.

[0030] ; ; In the above formula, represents the displacement of the vehicle at time, represents the coordinate information of the displacement sequence in the Frenet coordinate system of the lane segment, ​Represents the Frenet projection operation of the local graph node in the Cartesian coordinate system. At the same time, it stipulates that when the distance between the traffic participating vehicle and the lane segment exceeds 5 meters, the intelligent body motion embedding of the node is replaced by a full 0 vector .

[0031] 3. LaneG encoder (modeling of traffic participating vehicles in time sequence) The three features of the node are embedded and spliced ​​together, and projected into the final node features through the multi-layer perceptron (MLP) in the LaneG encoder: ; In summary, the local graph of the agent Depend on and Together they constitute: , indicating the The graph structure of the traffic participating vehicles encodes the trajectory information of the traffic participating vehicles and the surrounding environment information related to the prediction. Represents the node embedding feature vector A collection of yes The number of nodes in .

[0032] For each local graph constructed , the lane convolution layer in the LaneG encoder is used to transfer information with surrounding nodes to obtain the updated node embedding , as follows: ; In the above formula, and are all learnable parameters. Is and The nonlinear layer consists of a summation operation on all relations m and hop number n. The hop number n represents the distance between nodes. This multi-hop operation simulates the dilated convolution and effectively expands the receptive field. D represents the embedding feature dimension of each local graph node after convolution.

[0033] At the same time, in order to solve the long-term dependency, a graphics shortcut layer is added after each lane convolution. After each lane convolution, an average pooling operation is added to aggregate all node features into a global embedding and pass it directly to the graph. Node embedding In the following formula: ; ; In the above formula, Represents the aggregated global embedding, represents the number of nodes in the graph and is the final node representation after lane convolution and average pooling in the encoder and then summing them up.

[0034] The role of the encoder is to perform temporal modeling on the historical trajectory of the agent. It can be that the nodes in the encoded graph structure contain historical trajectory information, realizing the interaction of the agent in time series.

[0035] 4. Global Interaction (Spatial Interaction Modeling among Traffic Participating Vehicles) After encoding each agent's information, a graph that fuses the surrounding environment information is obtained , and then a global lane graph that contains all lanes in the scene is constructed . The nodes of the global lane graph are constructed by projecting all through lane pooling operation onto . The specific operation is as follows: 4.1 Construction of the Global Graph Node and Connection Edge Structure: ; Among them, N represents the number of local graphs , that is, the number of traffic participating vehicles in the scene, and M represents the number of nodes in the global graph.

[0036] Meanwhile, the edge relation matrix is determined by the topological structure of the lanes and is consistent with the connection relationship of each local graph .

[0037] 4.2 Construction of Node Feature Embedding : For each node in the global graph, retrieve all node sets related to from all graphs. For , there are: ; ; Among them, is the node of the local graph , represents the real distance between node and ,

[0038] ​​For all local neighbor nodes of the node features are aggregated to generate global nodes features : ; Among them, represents a multi-layer perceptron (MLP) used to process node features and relative position information, and is used to generate an enhanced node-position feature representation, represents the concatenation of local node features and relative positions; is also a multi-layer perceptron, which is used to further globally process the information after aggregating neighbor features to generate the final feature representation of the target node. E represents the embedding feature dimension of the nodes in the global graph; represents the set of neighbor nodes of

[0039] Finally, the aggregated features are used as the feature embedding of the global graph node : , .

[0040] 4.3 Information transfer of the global graph: Convolution operations can quickly extract effective local structure information, such as the movement of participants in the current lane, local obstacles, etc. However, when comparing long-term and long-distance context features, they cannot well capture global scene features. Different from convolution, the core of the attention mechanism lies in assigning different weights to different positions, enabling the model to selectively focus on important information as needed, especially the global dependencies in long-time spans and long-distance interactions. Therefore, the attention mechanism enables the model to learn and capture the interaction relationships between multiple targets and multiple time steps, and is especially suitable for modeling long-range dependencies and complex interactions. Therefore, in the global interaction process, on the basis of using lane convolution operations, a self-attention mechanism is introduced to further perform message passing on the information generated by convolution, thereby strengthening global information.

[0041] In the global graph, first use the lane convolution layer in the global interaction module to fuse the local features between nodes to capture the local context representation of the nodes: ; Then, the updated node features are input into the global interaction relationship encoding layer composed of multi-head attention layers in the global interaction module, and the spatial interaction relationships between vehicles and vehicles, and between vehicles and roads are captured through multi-head self-attention MHSA, denoted as the global interaction feature vector : Among them, PE represents the position feature vector of each node.

[0042] The final node embedding of the final global graph is the combination of the vectors obtained by the two operations. The combination method is to splice the two vectors and generate attention weights using a fully connected layer. and , as the combination weights, are as follows: ; ; where, and represent the combination weights of the two vectors, Q represents the node embedding feature dimension after the lane convolution layer and the multi-head attention layer, represents the fused feature vector after global interaction.

[0043] In summary, the feature vector obtained by getting the vectors through the lane convolution and the multi-head attention layer and then weighted combination of the two is used as the node feature vector of the final global graph. This feature vector highly fuses the local lane features and the global context features, and well realizes the interaction between agents and between agents and lanes. The finally obtained global graph is represented as .

[0044] 5. Lane Trajectory Decoder 5.1 The multi-modal trajectory prediction of the target agent is carried out on the basis of the local graph of each agent. Therefore, the node features obtained from the global graph need to be projected back to the local graph . To enhance the feature representation of the local graph so that it contains the context information after global interaction.

[0045] This part is the inverse process of 4.2 and will not be elaborated here.

[0046] Finally, the node features of each local graph obtain a new representation .

[0047] 5.2 Composition of the prediction head 5.2.1. The first-stage prediction head: Apply a simple multi-layer perceptron to predict each node in the local graph , predicting the probability of the node as the target point and the position residual of the node to the true target point: ; where, represents a two-layer multi-layer perceptron, and its activation function is ; represents The probability set where the middle node is the target point; Indicates the set of position residuals from the node to the true target point, where contains Relative position deviation and orientation residual.

[0048] 5.2.2. Second prediction head: According to the probabilities in, select the top K nodes with the highest scores as candidate target points TOP-K, and then use the feature set of the candidate target points , and perform information interaction between the candidate points through a streamlined multi-head attention layer to obtain the interaction feature : ; ; Through a multi-layer perceptron finally classify the interaction feature and estimate the position residual: ; Among them, indicates the probability set where the candidate target point in is the target point; Indicates the set of position residuals from the candidate target point to the true target point.

[0049] 5.2.3 Bezier curve fitting: Fit the trajectory curve between the obtained target point and the initial position of the vehicle. Use the Bezier curve to fit the trajectory from the current moment coordinate point to the target point. Through the final position, that is, the predicted target point coordinates and the position coordinates of the vehicle at the predicted moment , and assume that the vehicle moves with a kinematic model of constant acceleration in a short time. The final trajectory curve (in the Frenet coordinate system) is: ; Among them, ; Assume that the vehicle is moving with uniform acceleration, so the position of the vehicle at each moment can be calculated : ; ; ; 5.2.3.2 Trajectory optimization: Since the movement of the vehicle is not always with uniform acceleration. Therefore, we can use the Bezier curve as a reference trajectory to predict the position of the vehicle in the next sixty frames, that is, the future trajectory of the vehicle and through a multi-layer perceptron Regression fine-tuning offset to achieve accurate prediction.

[0050] ; Among them, , represents the driving direction offset and vertical direction offset of the vehicle trajectory in the next T frames, where T represents the number of frames for predicting the future trajectory. represents the node feature representation projected onto the graph after passing through the global interaction module. represents the application of a multi-layer perceptron.

[0051] The finally fine-tuned trajectory points are: ; Among them, represents the point at the th moment in the initial trajectory. represents the tangent vector direction at the th moment, represents

[0052] Finally, the model outputs K multi-modal trajectories and their confidence scores for each graph: ; ; Among them, a time series of each trajectory represents the future position. represents the rd predicted trajectory, which contains the position information of the vehicle in the next T frames. . represents the set of its corresponding confidence scores.

[0053] The entire model outputs the future trajectories of N traffic participating vehicles and their corresponding confidence functions as: ; ; 6. Loss function The training of the model is to reduce the error between the predicted value and the true value, and update the parameters in the model network by backpropagating the loss, guiding the model to perform end-to-end optimization, so that the trajectory predicted by the prediction model is closer to the true value.

[0054] The hyperparameters of each module in the multi-vehicle joint trajectory prediction model are shown in Table 1.

[0055] Table 1 Hyperparameters of the multi-vehicle joint trajectory prediction model structure

[0056] The loss function adopted for training the entire multi-vehicle joint trajectory prediction model is as follows: ; Preferably, set , , , ; Among them, represents the model classification loss. Each node determines whether it is the final target point, and the cross-entropy loss is adopted: ; Among them, N is the total number of nodes, represents whether the th node is the target point. is the probability that the node is predicted as the target.

[0057] represents the residual loss between each node and the end point (target point), and the loss is expressed as: ; ; ; Among them, represents the position and orientation difference between the candidate point and the true target point.

[0058] represents the trajectory regression loss. After the candidate target point is refined, K trajectories are output, and the trajectory regression loss is calculated by selecting the one closest to the GT. The loss is expressed as: ; Among them, T represents the number of time steps of trajectory prediction represents the predicted th trajectory at time step , where represents the one closest to the true trajectory, represents the position coordinate of the true trajectory at time step , represents the L1 norm, that is, the sum of the absolute values of the main elements of the position coordinates.

[0059] represents the loss of fine-tuning the trajectory points under Frenet. The loss value of correcting the vehicle trajectory at each time step is expressed as: ; Among them, represents the fine-tuned trajectory point, Denote the real trajectory points T Denote the number of frames (time steps) for predicting the future trajectory of the vehicle

[0060] Denote the multi-modal trajectory probability NLL loss (top-k trajectories + confidence), which models the probability that the target trajectory appears in multiple predicted trajectories. The loss value is as follows: ; After the training is completed, the multi-vehicle joint trajectory prediction model can be applied to the trajectory prediction of traffic participating vehicles in the actual driving scenario

[0061] II. Predict the multi-vehicle joint trajectory using the trained multi-vehicle joint trajectory prediction model Except for the input data processing process, the prediction process is basically the same as the training process

[0062] Data processing The input data is mainly obtained by in-vehicle sensors. Using a data acquisition system based on , combined with lidar, millimeter-wave radar, depth camera, GPS positioning, and high-precision maps, to obtain the traffic environment information of surrounding vehicles and obstacles. Then, through processes such as object detection and classification (YOLOv8), obstacle detection and tracking, lane detection (LaneATT), and traffic signal recognition (YOLOv5 + Traffic light head), the ICP algorithm is used to spatially align the objects in the Lidar point cloud data and camera images to obtain high-precision position information, so as to obtain the complete and accurate historical trajectory information of the ego vehicle and traffic participating vehicles, as well as the position information of the lane lines in this traffic scenario. Finally, the ego vehicle state and historical trajectory, the historical trajectory of other vehicles, and the lane line topology and lane line position information are output

[0063] Construct the graph structure The historical trajectory of each traffic participating vehicle and the position information of each surrounding lane line are represented as follows: ; ; According to the current position and scenario information of the traffic participating vehicle to be predicted, extract the lane segments passed by this traffic participating vehicle within the past time T range, generally select 2s (20 frames), and the lane segments that this traffic participating vehicle may pass through in the future 6s (60 frames). Divide the lane lines into continuous lane segments at intervals of 2m along the center line

[0064] Take the center points of each lane segment as the nodes of the local graph to form a node matrix , its topological structure in the real scenario, as the connection edge relationship between nodes, has four types: left neighbor, right neighbor, predecessor, and successor, forming an edge relationship matrix, and taking the characteristic information variables of each node as the embedding vector of the node, forming the embedding vector matrix of the node . Finally, the local graph is represented as , as the graph structure input information of the subsequent prediction model.

[0065] For each node , the embedding feature vector representation is: ; Among them, represents the geometric features of the lane segment (node), represents the lane segment type features, represents the dynamic interaction relationship between the lane segment and the agent, and C represents the node embedding feature dimension.

[0066] Set ; In the above formula, represents the center point coordinates of the lane segment center line, represents the orientation of the lane segment, represents the lane segment curvature. For the lane segment type features it is: ; In the above formula, represents the lane end category, which has four types: straight lane, turning lane, fork lane, and ramp; after encoding it, they correspond to the vectors [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 1, 0], [0, 0, 0, 1] respectively.

[0067] For , since the past movement of traffic participating vehicles can be represented as a set of coordinate points that is, it defines the vehicle center position at consecutive time steps. Therefore, these coordinate points are converted into a displacement sequence , as shown in the following formula, and is projected onto the coordinate system of the local graph node to encode the movement information of the participants in the case related to the map.

[0068] ; ; In the above formula, represents the displacement of the vehicle at moment, represents the coordinate information of the displacement sequence in the coordinate system of the lane segment, represents the projection operation from the Cartesian coordinate system to the local graph node under . It is also stipulated that when the distance between the traffic participant vehicle and the lane segment exceeds 5 meters, the intelligent agent motion embedding of this node is replaced with a vector of all zeros .

[0069] 3. Encoder concatenates the three feature embeddings of the node, and projects them into the final node feature through the multi-layer perceptron (MLP) in the encoder: ; To sum up, the intelligent agent local graph is composed of and together, that is: , representing the graph structure of the th traffic participant vehicle, encoding the traffic participant vehicle trajectory information and surrounding environment information related to the prediction. represents the set of node embedding feature vectors , is the number of nodes in

[0070] For each constructed intelligent agent's local graph , apply the lane convolution layer in the encoder to perform information transfer with surrounding nodes to obtain the updated node embedding , as shown in the following formula: ; In the above formula, and are both learnable parameters, is a non-linear layer composed of and . The summation operation is performed over all relationships m and hop numbers n. The hop number n represents the distance between nodes. By simulating dilated convolution through this multi-hop operation, the receptive field is effectively expanded. D represents the embedding feature dimension of each local graph node after convolution.

[0071] After each layer of lane convolution, a graphical shortcut layer is added. After each layer of lane convolution, an average pooling operation is added to aggregate all node features into a global embedding, which is directly passed to the node embedding of the graph of the graph as follows: ; ; In the above formula, represents the aggregated global embedding, represents the graph number of nodes, is the final node representation after lane convolution, average pooling, and summation in the encoder.

[0072] 4. Global Interaction (Spatial Interaction Modeling among Traffic Participating Vehicles) After encoding each agent's obtain the graph that fuses the surrounding environment information, and then construct a global lane graph G that contains all lanes in the scene. The nodes of the global lane graph G are constructed by projecting all onto through a lane pooling operation. The specific operation is as follows: 4.1 Construction of the Global Graph Node and Connection Edge Structure: ; where N represents the number of local graphs , that is, the number of traffic participating vehicles in the scene, and M represents the number of nodes in the global graph.

[0073] At the same time, the edge relation matrix is determined by the topological structure of the lanes and is consistent with the connection relationship of each local graph .

[0074] 4.2 Construction of the Node Feature Embedding F: For each node in the global graph, retrieve all node sets related to from all graphs containing . For there is: . .

[0075] Among them, is the node of the local graph , represents the node and true distance, Represents the distance threshold of neighbor nodes, generally taking 5m.

[0076] For all local neighbor nodes of the node features are aggregated to generate global node features : ; Among them, represents a multi-layer perceptron (MLP) used to process node features and relative position information, and is used to generate an enhanced node-position feature representation. represents the concatenation of local node features and relative positions; is also a multi-layer perceptron, used to further globally process the information after aggregating neighbor features to generate the final feature representation of the target node. E represents the embedding feature dimension of nodes in the global graph.

[0077] Finally, the aggregated features are used as the feature embedding of the global graph nodes : .

[0078] 4.3 Information transfer of the global graph: Convolution operations can quickly extract effective local structure information, such as the movement of participants in the current lane, local obstacles, etc. However, when comparing long-term and long-distance context features, they cannot well capture global scene features. Different from convolution, the core of the attention mechanism lies in assigning different weights to different positions, enabling the model to selectively focus on important information as needed, especially the global dependencies in long-time spans and long-distance interactions. Therefore, the attention mechanism enables the model to learn and capture the interaction relationships between multiple targets and multiple time steps, especially suitable for modeling long-range dependencies and complex interactions. Therefore, in the global interaction process, on the basis of using graph convolution operations, a self-attention mechanism is introduced to further perform message passing on the information generated by convolution, thereby strengthening global information.

[0079] In the global graph, first use the lane convolution layer in the global interaction module to fuse the local features between nodes, for capturing the local context representation of nodes: ; Then, the updated node features are input into the global interaction relationship encoding layer composed of multi-head attention layers in the global interaction module. The spatial interaction relationships existing between vehicles and between vehicles and roads are captured through multi-head self-attention MHSA, denoted as the global interaction feature vector: ; Among them, PE represents the position feature vector of each node.

[0080] Finally, the final node embedding of the global graph is the combination of the vectors obtained by the two operations. The combination method is to splice the two vectors and generate attention weights using a fully connected layer , , as the combination weight, as shown in the following formula: ; ; Among them, , represents the combination weight of the two vectors, Q represents the node embedding feature dimension after the lane convolution layer and the multi-head attention layer, represents the fused feature vector after global interaction.

[0081] In summary, the feature vector obtained by getting vectors through the lane convolution and the multi-head attention layer and then weighted combination of the two is used as the node feature vector of the final global graph. This feature vector highly integrates local lane features and global context features, and well realizes the interaction between agents and between agents and lanes. The finally obtained global graph is represented as .

[0082] 5. Lane Trajectory Decoder 5.1 The multi-modal trajectory prediction of the target agent is carried out based on the local graph of each agent . Therefore, the node features obtained from the global graph need to be projected back to the local graph to enhance the feature representation of the local graph so that it contains the context information after global interaction.

[0083] Finally, the node features of each local graph obtain a new representation .

[0084] 5.2 Composition of the prediction head 5.2.1. The first-stage prediction head: Apply a multi-layer perceptron to predict each node in the local graph , predicting the probability that the node is the target point and the position residual from the node to the true target point: ; Among them, represents a two-layer multi-layer perceptron, and its activation function is ; represents the set of probabilities that the nodes in are the target points; Denote the set of position residuals from nodes to the true target points, where contains relative position deviations and orientation residuals.

[0085] 5.2.2. Second-stage prediction head: According to the probabilities in, select the top K nodes with the highest scores as candidate target points TOP-K, and then use the feature set of the candidate target points , pass through a streamlined multi-head attention layer to perform information interaction between the candidate points to obtain interaction features , ; ; Through a multi-layer perceptron finally classify the interaction features and estimate the position residuals: ; Among them, denotes the set of probabilities that the candidate target points in are the target points; denotes the set of position residuals from the candidate target points to the true target points.

[0086] 5.2.3 Bezier curve fitting: Fit the trajectory curve between the obtained target points and the initial position of the vehicle. Use the Bezier curve to fit the trajectory from the current moment's coordinate points to the target points. Through the final position, that is, the predicted target point coordinates and the position coordinates of the vehicle at the predicted moment , and assuming that the vehicle moves with a kinematic model of constant acceleration in a short time, the obtained final trajectory curve (in the Frenet coordinate system) is:

[0087] Among them, .

[0088] Assume that the vehicle is moving with uniform acceleration, so the position of the vehicle at each moment can be calculated : ; ; ; 5.2.3.2 Trajectory optimization: Since the vehicle's motion is not always with uniform acceleration. So we can use the Bezier curve as a reference trajectory, predict the position of the vehicle in the next sixty frames, that is, the future trajectory of the vehicle and through a multi-layer perceptron Regression fine-tuning offset to achieve accurate prediction.

[0089] ; Among them, represents the driving direction offset and vertical direction offset of the vehicle trajectory in the next T frames. T represents the number of frames for predicting the future trajectory. represents the node feature representation projected onto the graph after passing through the global interaction module. represents the application of a multi-layer perceptron.

[0090] The finally fine-tuned trajectory points are:

[0091] Among them, represents the point at the th moment in the initial trajectory. represents the tangent vector direction at the represents the normal vector direction at the

[0092] Finally, the model outputs K multi-modal trajectories and their confidence scores for each graph: ; ; Among them, a time series of each trajectory represents the future position. represents the th predicted trajectory, which contains the future T-frame position information of the vehicle. . represents the set of its corresponding confidence scores.

[0093] The entire model outputs the future trajectories of N traffic participating vehicles and their corresponding confidence functions as: .

[0094] The multi-modal trajectory prediction method for multi-vehicle association centered on graphs provided by the present invention constructs a local lane graph to represent the context information (surrounding environment information) and historical trajectory data of traffic participants, and models agents in a distributed and map-aware manner. During the construction of the local graph, the topological structure of the map is retained and fine-grained environment information is extracted. During global interaction, a combination of lane convolution and multi-head attention layers is adopted. The lane convolution network is used to capture the local interaction between the agent and its neighborhood environment, while the global interaction between agents is modeled through the self-attention mechanism, further enhancing the prediction ability of the model. A two-stage decoding is designed in the decoder, dividing the decoding process into a stage of generating initial target candidates and a stage of refined classification and regression. By such means, the predicted trajectory has higher accuracy and stronger robustness. Through this innovative graph structure and information interaction mechanism and the unique two-stage decoder structure, the future trajectory of the agent can be predicted more accurately and with greater robustness, especially having significant advantages in scenarios such as complex intersections.

[0095] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the examples shown and described herein.

Claims

1. A graph-centric vehicle multimodal trajectory prediction method, characterized in that: include: Acquire historical trajectories of vehicles participating in the traffic and lane lines corresponding to the historical trajectories, and divide the lane lines corresponding to the historical trajectories into lane segments; Taking the center point of the lane segment as a node and the positional relationship between the center points of adjacent lane segments as an edge, a local graph of each traffic participant vehicle is constructed; After combining the geometric features and type features of the lane segment and the dynamic interaction relationship between the lane segment and the participating vehicles, feature extraction is performed to obtain the node features of the local graph. Construct a global lane map containing all lanes in the scene, encode the local map of each traffic participant vehicle respectively, and project all the encoded local maps onto the global lane map to obtain a global map; After aggregating the node features of all local graphs, perform lane convolution on the aggregated feature vector to obtain a local feature fusion vector; Performing feature extraction on the local feature fusion vector through a global interaction relationship encoding layer to obtain a global interaction feature vector; The local feature fusion vector and the global interaction feature vector are weightedly combined to obtain node features of the global graph; Projecting the node features of the global graph back to the local graph of each traffic participant vehicle to obtain an enhanced local graph; The predicted trajectory of each traffic participant vehicle and the confidence level of each predicted trajectory are obtained according to the enhanced local graph.

2. The graph-centric vehicle multimodal trajectory prediction method according to claim 1, characterized in that: The method for encoding the local graph of each traffic participant vehicle is: A lane convolution operation is performed on the node features of the local graph, and an average pooling operation is performed after each layer of convolution, and the results of the convolution operation and the average pooling operation are added to obtain an encoded local graph node feature vector.

3. The graph-centric vehicle multimodal trajectory prediction method according to claim 2, characterized in that: The global interaction relationship encoding layer is composed of a multi-head attention layer.

4. The graph-centric vehicle multimodal trajectory prediction method according to claim 3, characterized in that: The predicted trajectory of each traffic participant vehicle and the confidence of each predicted trajectory are obtained according to the enhanced local graph, including the following steps: Step 1: Use the first prediction unit to predict each node in the enhanced local graph, determine the probability of the node being a target point and the position residual from the node to the true target point; Step 2: Select K nodes with the highest probability of being the target point as candidate target points, and perform information interaction on the candidate target points through a multi-head attention layer to obtain interaction features; Inputting the interactive features into a second prediction unit to obtain a probability of the candidate target point being a target point and a position residual from the candidate target point to the true target point; Step 3: Use the Bezier curve to fit the trajectory from the current coordinate point to the candidate target point to obtain a reference trajectory; Step 4: fine-tune the reference trajectory to obtain a predicted trajectory, and use the probability of the candidate target point being the target point as the confidence of the corresponding predicted trajectory.

5. The graph-centric vehicle multimodal trajectory prediction method according to claim 4, characterized in that: The first prediction unit and the second prediction unit both use a multi-layer perceptron.

6. The graph-centric vehicle multimodal trajectory prediction method according to claim 4 or 5, characterized in that: In step 4, the reference trajectory is fine-tuned to obtain the predicted trajectory as follows: The position of the vehicle at multiple moments in the future is obtained through the reference trajectory, the offset at each moment is determined through a multi-layer perceptron, and the position of the multiple moments in the future is fine-tuned according to the offset to obtain a predicted trajectory.

Citation Information

Patent Citations

  • Target trajectory prediction system and prediction method

    CN113095504A

  • Multi-vehicle joint trajectory prediction method based on intent learning of surrounding vehicles

    CN118968441A

  • Trajectory prediction method based on hierarchical progressive interaction and target lane segment

    CN119099654A

  • Transform and end point induction-based multi-modal trajectory prediction method for automatic driving vehicle

    CN119975390A

Cited By

  • Vehicle track prediction method and system

    CN121191117A