A Multimodal Trajectory Prediction Method Based on Spatiotemporal Decoupled Encoding of Vehicle Interaction Graphs

The vehicle interaction graph spatiotemporal decoding method improves trajectory prediction by integrating vehicle and scene features, enhancing accuracy and generalization in complex traffic scenarios for autonomous driving.

CN119329558BActive Publication Date: 2025-07-15CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411370622.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-29
Publication Date
2025-07-15
Estimated Expiration
2044-09-29

AI Technical Summary

Technical Problem

The existing trajectory prediction methods are difficult to effectively consider the interaction between traffic participants and multi-source information in complex traffic scenarios, resulting in the trajectory prediction results being inaccurate enough to meet the safe and efficient decision-making needs of autonomous vehicles.

Method used

A multimodal trajectory prediction method based on space-time decoupling encoding of vehicle interaction graph is adopted. By constructing graph structure data, the time series changes of interaction graph node characteristics are captured, and the loss function is designed for model training in combination with the timing encoding module and the trajectory decoding module, and the vehicle historical characteristics, interaction characteristics and map characteristics are comprehensively considered.

Benefits of technology

It realizes more refined multi-agent interactive multi-modal trajectory prediction in dynamic traffic scenarios, improves the accuracy and generalization of predictions, and can better support the safety decisions of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119329558B_ABST
    Figure CN119329558B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-modal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs, belonging to the technical field of autonomous vehicles. The method includes: generating graph structure data according to a traffic scene trajectory data set; constructing an interaction graph considering multiple traffic factors based on the graph structure data; capturing the changes in the node features of the interaction graph in the time series through a temporal encoding module when the interaction graph updates features; completing multi-modal prediction by constructing a trajectory decoding module and a score decoding module; designing a loss objective function for the prediction model and training it. The present invention combines the influence of multi-source factors on motion prediction through a heterogeneous graph attention method, and at the same time decouples and encodes spatio-temporal features to better capture the features in the time series. Compared with traditional vehicle trajectory prediction methods, the present invention can complete more refined predictions and has better scene generalization performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving vehicles, and relates to a multi-modal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs. Background Technique

[0002] Autonomous driving vehicle technology focuses on solving traffic challenges such as safety, efficiency, and energy, and thus has received high attention from the academic and industrial communities. Its autonomous driving ability of the vehicle is achieved through modules such as environmental perception, trajectory prediction, decision-making, and planning control. An efficient and accurate trajectory prediction method can obtain the future states of itself and nearby traffic participants, and transmit the information to the downstream decision-making module to ensure the safe driving of the vehicle in a dynamic environment.

[0003] Currently, trajectory prediction can be mainly divided into physics-based methods, classical machine learning-based methods, deep learning-based methods, and reinforcement learning-based methods. Most traditional prediction methods are only applicable to simple traffic scenarios or short-term predictions, while the actual traffic scenario is a task with high complexity and strong interactivity. Deep learning-based methods, through data-driven and efficient extraction of target features, have powerful non-linear modeling capabilities and have received increasing attention. Deep learning-based methods can be further divided into specific methods such as those based on recurrent neural networks, convolutional neural networks, graph neural networks, and generative adversarial networks. In traffic tasks, the interactivity between traffic participants is a key link, and graph neural networks can efficiently express the interaction between node information and have strong transfer applicability, which is very suitable for solving trajectory prediction problems based on interaction-related factors.

[0004] However, in the trajectory prediction task, in addition to paying attention to interaction factors, the characteristics of traffic participants themselves are also important factors affecting the trajectory prediction results, and the constraints of the scene map and traffic rules on traffic are continuous. Therefore, there is an urgent need for a trajectory prediction method that comprehensively considers multi-source information to achieve multi-modal prediction of trajectories and provide information support for downstream safe and efficient decision-making and planning. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a multi-modal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A multi-modal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs, which includes the following steps:

[0008] S1. Generate graph structure data according to the traffic scene trajectory data set;

[0009] S2. Construct an interaction graph considering multiple traffic factors based on graph-structured data;

[0010] S3. When the interaction graph updates features, capture the changes in the features of the interaction graph nodes over time through a temporal encoding module;

[0011] S4. Complete multimodal prediction by constructing a trajectory decoding module and a score decoding module;

[0012] S5. Design the loss objective function of the prediction model and train it.

[0013] Furthermore, in step S1, the graph-structured data is represented as where the data sources used to construct the graph-structured data include the historical motion states of the agents and the high-definition map information

[0014]

[0015]

[0016]

[0017] where i represents the number of agents considered for interactive prediction within a certain range, represents the coordinate information of the agent at time t and the heading angle

[0018] Furthermore, in step S2, the following steps are included:

[0019] S21. Extract historical vehicle node features according to the historical motion states ; Output the vehicle node features by passing the historical motion states through a ModuleList model list composed of a fully connected layer, a regularization layer, an activation layer, a GRU encoding layer, and a residual connection module:

[0020]

[0021]

[0022] where, is the historical feature of a certain agent at time t; represents encoding through a sequential model list of a specific agent type, with the superscript at indicating a specific type of agent (agent-type) and the subscript hist indicating historical information; N t is the historical feature of all agents output by this module.

[0023] S22. Extract lane line node features according to the global high-definition map information For the high-definition map information Establish lane objects through lanelet2, extract lane centerlines, determine lane line nodes, and calculate the distances between the agent and the lane line nodes:

[0024]

[0025]

[0026] Where represents the start and end indices of each lane centerline within the threshold distance from agent i at time t.

[0027] S23. Establish interactions between vehicle agents and between agents and lanes; among them, for the interactions between vehicle agents, other agents within the specified range of the target agent node features are used as neighbor nodes to embed their own node features, construct directed edges pointing from neighbor nodes to the target node, and use an adjacency matrix to represent the connection with neighbor agents. At the same time, calculate the relative position information between nodes and embed it into the edge attribute variables constructed in the PyTorch_Geometric library to represent the interactions between agent nodes, that is:

[0028]

[0029]

[0030] Among them, is embedded as the set of edge attributes between nodes i and j to represent the relative position relationship between nodes, and E t is a set containing edge indices, edge attributes, and edge types.

[0031] For the implicit interaction between vehicle agents and lane lines, the interaction between agents and lanes is represented in the form of nodes and edges in the graph structure. Given the node features of the agent and edge features Under the condition, filter out the indices of the target lane line nodes according to the distance And embed the distances between the corresponding agent nodes and lane line nodes into the edge attribute variables constructed in the PyTorch_Geometric library to capture the implicit interaction between agents and lanes.

[0032] S24. Apply the multi-head self-attention mechanism to perform attention-weighted feature updates on node features, and realize the construction and update of the interaction graph considering multi-source information:

[0033] MH(Q, K, V) = [head1, head2, head3]W

[0034]

[0035] In the multi - head self - attention mechanism, three attention heads are set. Q, K, and V are the query vector, key vector, and value vector respectively, where is the weight matrix; when calculating the attention score α, scaled dot - product attention is adopted:

[0036]

[0037] where soft max(·) represents the normalization layer, (·) T represents matrix transpose, and d k represents the dimension of the key vector.

[0038] Furthermore, in step S3, the following steps are included:

[0039] S31. The target node completes the update of its own feature matrix N t by aggregating neighbor information. When aggregating and updating, the self - attention mechanism between nodes is used to focus on the dynamic update result of features over time is expressed as:

[0040]

[0041] where is the linear transformation of the feature matrix at the previous time step, serves as the attention - weighted feature matrix of the target agent among them,

[0042] S32. For the attention - weighted feature matrix, the time information is explicitly encoded through the temporal encoding module to obtain the final feature matrix of the target agent at the current time step:

[0043]

[0044] where g0, g1 are composed of a two - layer MLP network, is the temporal feature fusion module. The temporal feature fusion module encodes the timestamp through a specific linear layer, encodes the time information into a high - dimensional vector. Then the time encoding is repeated to be the same as the feature dimension, added to the hidden state and a feature mask is set, and finally the hidden state is updated through the residual module, and the final output is the node temporal feature at the current time step.

[0045] Moreover, the state features corresponding to the time steps are stored by the data replay module, which provides historical sequence information to the temporal encoding module. Among them, at least a data replay module is set up to separately track and store the target agent features and the scene features.

[0046] Furthermore, in step S32, for the data replay module used to track and store the target agent features, the LSTM is used to sequentially track the changes in the representation of the time features and the hidden state of the target agent at time t, and the formula is expressed as:

[0047]

[0048] where h t is the hidden state in the LSTM, h t-1 is the input target node feature, represents the feature matrix of the target agent at time t; in the interactivity prediction, the target agents are alternating. When the target agent of interest is confirmed, the remaining traffic participants act as neighbor agents to participate in the calculation of the interaction time series graph, and each agent for prediction stores the node features as the target agent in the data replay module.

[0049] Furthermore, in step S32, when building the data replay module for the scene features, at time step t, the linear layer initialization module is applied to the feature matrix N t to initialize it, and then a multi-layer self-attention operation network is constructed to store the scene features:

[0050]

[0051]

[0052] where g2 is the linear layer for extracting the feature matrix, l represents the number of layers, and the storage of the scene features can also ensure that the subsequent decoding module fully considers the map features when outputting the trajectory;

[0053] Layer normalization is performed on each self-attention layer built. After the last layer L, the node features are aggregated through the max-pooling operation to summarize the scene-related information and the information is also correlated through time by the LSTM module to complete the data storage of the scene features:

[0054]

[0055] where, h t ' represents the map hidden state tracked by the module at time t; h t ' -1 represents the map hidden state at time t - 1.

[0056] Further, in step S4, a trajectory decoding module and a score decoding module are constructed to model the probability distribution of multimodal trajectory prediction, which specifically includes the following steps:

[0057] S41. Considering the agent's own characteristics, interaction characteristics, and map characteristics comprehensively, complete the task of trajectory decoding through the GRU network structure, and output the coordinates of the trajectory at each time step:

[0058]

[0059] Among them, is the agent's multimodal future decoder, is the concatenation of features in the traffic task, is the predicted future trajectory of the k-th mode of the agent;

[0060] S42. The score decoder follows the same structure as the trajectory decoder, and finally outputs the confidence score of the proposed trajectory, and further generates the generation probability distribution of the multimodal trajectory through the softmax layer:

[0061]

[0062] Among them, c i represents the trajectory confidence score, represents the confidence range.

[0063] Further, in step S5, a loss function is designed according to the trajectory decoding module and the score decoding module for model training.

[0064] For the trajectory loss, use smooth loss at all predicted time steps, and the trajectory regression loss is defined as:

[0065]

[0066] Among them is the true trajectory position at time t, is the predicted trajectory closest to the true trajectory position; k * is the index in the multimodal predicted trajectories that is closest to the true trajectory position, which is obtained by calculating the distance between the end point of the multimodal predicted trajectory and the end point of the true trajectory:

[0067]

[0068] For the loss function of the predicted score, design a cross-entropy loss function Calculate the loss between the true score and the predicted probability distribution P, and define the predicted score loss as:

[0069]

[0070] Define the score distribution corresponding to the true trajectory:

[0071]

[0072] The total loss is the weighted sum of the trajectory regression loss and the score loss:

[0073]

[0074] where δ is a parameter set to balance the two learning objectives.

[0075] The beneficial effects of the present invention are as follows:

[0076] The present invention performs spatio-temporal decoupling encoding on driving tasks in vehicle trajectory prediction, better learns the time representation in dynamic traffic scenarios, and simultaneously realizes interactive trajectory prediction considering multi-source factors. Compared with traditional vehicle motion prediction methods, the present invention can more precisely complete multi-agent interactive multi-modal trajectory prediction and has stronger generalization ability for scenarios.

[0077] (1) The present invention designs a trajectory prediction model encoding module based on a graph neural network and a self-attention mechanism. Considering vehicle historical features, vehicle-interaction features, and traffic scene features, the constructed interaction graph can more accurately and effectively express the traffic tasks of the current time frame and effectively represent the interaction features between traffic factors.

[0078] (2) The present invention designs a temporal encoding module based on target information data replay. By receiving the updated interaction graph node features from the upstream, the temporal encoding module is constructed to capture the features in the time series and update them into the vehicle state features, thereby better decoupling the time features and space features of observing traffic tasks.

[0079] (3) The present invention constructs a multi-modal trajectory prediction trajectory decoding module and a score decoding module. The GRU module is applied to track agent features and map features, and based on this, the probability distribution of multi-modal trajectories is modeled, enabling the decoded output to comprehensively consider map factors and agent features. The designed trajectory smoothing loss and cross-entropy loss fusion objective function can comprehensively and efficiently train the model.

[0080] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, they will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the accompanying drawings, where:

[0082] Figure 1 is the overall flowchart of the method proposed by the present invention;

[0083] Figure 2 is the flowchart for constructing an interactive graph representation of traffic tasks based on a graph neural network;

[0084] Figure 3 is the working flowchart of the time series encoding module and the data playback module. Specific Embodiments

[0085] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0086] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0087] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0088] Please refer to Figures 1 to 3, the present invention relates to an interactive multi-modal trajectory prediction method based on a dynamic temporal graph neural network. First, an interactive graph considering the motion characteristics, interaction characteristics, and map attributes of traffic agents is constructed through a graph attention network. Secondly, a data replay module is constructed to store the key information of past frames, and a temporal encoding module is used to capture the changes in the features of the constructed interactive graph over time series; finally, a trajectory decoding module and a score decoding module are constructed to model the probability distribution of multi-modal trajectory prediction, and a loss function is designed accordingly for model training. The present invention combines the influence of map factors on motion prediction, and at the same time fully considers the interactive characteristics in traffic scenarios. Compared with traditional vehicle motion prediction methods, the present invention can complete multi-agent interactive multi-modal trajectory prediction with high confidence, has high accuracy, and has a wider range of application scenarios. The method specifically includes the following steps:

[0089] S1. Generate graph structure data according to the traffic scenario trajectory data set;

[0090] S2. Construct an interactive graph considering multiple traffic factors based on the graph structure data;

[0091] S3. When the interactive graph updates features, capture the changes in the features of the interactive graph nodes over time series through a temporal encoding module;

[0092] S4. Complete multi-modal prediction by constructing a trajectory and a score decoding module;

[0093] S5. Design a loss objective function for the prediction model and conduct training.

[0094] Embodiment

[0095] In step S1, in this embodiment, graph structure data is constructed through the historical trajectory features of vehicles. Specifically, the historical trajectory features of vehicles include the historical motion states of vehicle agents and high-definition map information, and the constructed graph structure data is used for the training, verification, and testing of subsequent prediction models.

[0096] Specifically, in step S1, the graph structure data is represented as where the data source for constructing the graph structure data includes the historical motion states of agents and high-definition map information

[0097]

[0098]

[0099]

[0100] where i represents the number of agents considering interactive prediction within a certain range, Represents the coordinate information of the agent at time t and the heading angle To ensure the translational and rotational invariance of the predicted agent, before feature extraction, the historical trajectory data needs to be subjected to a relative pose transformation with respect to the position and heading angle of the current frame target agent.

[0101] In step S2, according to the graph structure data of the historical motion state and the high-definition map information Extract the historical vehicle node features and lane line node features respectively, and then, based on the graph structure data composed of vehicle nodes, lane line nodes, and connecting edges, apply the multi-head self-attention mechanism to perform attention-weighted feature updates on the node features. Specifically, it includes the following steps:

[0102] S21. Extract historical vehicle node features according to the historical motion state The historical trajectory input received by the prediction model passes through a ModuleList composed of a fully connected layer, a regularization layer, an activation layer, a GRU encoding layer, and finally a residual connection module. The result obtained through the above steps is used as the proxy graph node feature of the network, as shown in the formula:

[0103]

[0104]

[0105] where is the historical feature of a certain agent at time t, and its encoding is through an ordered model list of a specific agent type: N t as the historical features of all agents output by this module.

[0106] S22. Extract lane line node features according to the global high-definition map information For the high-definition map data, establish lane objects through lanelet2 and extract the lane centerlines, determine the lane line nodes, and calculate the distance between the agent and the lane line nodes.

[0107]

[0108]

[0109] where represents the start and end indices of each lane centerline within the threshold distance from agent i at time t.

[0110] S23. The interaction graph mainly includes the interaction between agents (Agent-Agent) and the implicit interaction between agents and lanes (Agent-Lane). When establishing the A-A interaction, As the target agent node features, at this time, other agents within the specified range are used as neighbor nodes to embed their own node features, construct a directed edge pointing from the neighbor node to the target node, and use an adjacency matrix to represent the connection with the neighbor agent. At the same time, calculate the relative position information between nodes and embed it into the edge attribute variable constructed in the PyTorch_Geometric library to represent the interaction between agent nodes, that is:

[0111]

[0112]

[0113] Among them, Embed the relative position relationship between nodes as the edge attribute set of nodes i and j, and E t is a set containing edge indices, edge attributes, and edge types.

[0114] For the implicit interaction between agents and lanes, that is, the A-L interaction, the interaction between agents and lanes is also represented in the form of nodes and edges in the graph structure. After obtaining the node features of the agents and edge features In this case, filter out the indices of the target lane line nodes according to the distance and embed the distance between the corresponding agent nodes and lane line nodes into the edge attribute variable constructed in the PyTorch_Geometric library to capture the implicit interaction between agents and lanes.

[0115] S24. According to the above steps, obtain the graph structure data composed of vehicle nodes, lane line nodes, and connecting edges. Apply the multi-head self-attention mechanism to perform attention-weighted feature update on the node features to complete the construction and update of the interaction graph that comprehensively considers multi-source information:

[0116] MH(Q, K, V) = [head1, head2, head3]W

[0117]

[0118] In the multi-head self-attention mechanism, set three attention heads. Q, K, and V are the query vector, key vector, and value vector respectively, where is a trainable weight matrix. When calculating the attention score α, use scaled dot-product attention:

[0119]

[0120] In step S3, the data playback module of this embodiment stores the key information of past frames, and captures the changes in the characteristics of the interaction graph nodes built by the time series encoding module in the time series. The specific steps are as follows:

[0121] S31. Update the graph structure data received by the interaction graph. The nodes complete the update of their own feature matrix N by aggregating neighbor information (including the interaction between agents and the interaction between agents and the map lanes). t During the aggregation update, this embodiment uses the self-attention mechanism between connected nodes to focus on the dynamic update results of features over time. It is expressed as:

[0122]

[0123] Among them, is the linear transformation of the feature matrix of the previous time step. As the attention-weighted feature matrix of the target agent among them.

[0124] S32. Next, for the obtained feature matrix weighted by attention scores, explicitly encode the time information through the time series encoding module to obtain the final feature matrix of the target agent at the current time step:

[0125]

[0126] Among them, g0 and g1 are composed of a two-layer MLP network. is the time series feature fusion module. This module encodes the timestamp through a specific linear layer, encodes the time information into a high-dimensional vector. Then repeat the time encoding to the same dimension as the feature dimension, add it to the hidden state and set the feature mask, and then update the hidden state through the residual module. Finally, the output is the node time series feature at the current time step.

[0127] When performing time series graph encoding in step S32, a key module is the data playback module, which is used to store the state features corresponding to the time step and provide accurate historical sequence information to the time series encoding module. In the present invention, two specific modules are built to track and store the features of the target agent and the scene features respectively.

[0128] S321. In the data playback module for tracking and storing the features of the target agent, for the time features and hidden states of the target agent at time t, there are non-learnable solutions such as taking the latest information or selecting the average information. In this embodiment, a learnable solution is selected: use LSTM to sequentially track the changes in its representation, and the formula is expressed as:

[0129]

[0130] where h t is the hidden state in the LSTM, and the input is the feature of the target node. Here, in this embodiment, only the sequential data replay module is constructed for the target agent of interest. In the interactivity prediction, the target agents are alternating. When the target agent of interest is confirmed, the remaining traffic participants act as neighbor agents to participate in the interactive time series graph calculation. Each agent for prediction stores the node features in the data replay module as the target agent.

[0131] S322. In the construction of the data replay module for scene features, at time step t, initialize the module by applying a linear layer to the feature matrix N t and then construct a multi-layer self-attention operation network to store the scene features:

[0132]

[0133]

[0134] where g2 is the linear layer for extracting the feature matrix, l represents the number of layers, and the storage of scene features can also ensure that the subsequent decoding module fully considers the map features when outputting the trajectory.

[0135] For each self-attention layer constructed in S322, layer normalization is performed in this embodiment. After the last layer L, aggregate the node features through a max pooling operation to summarize the scene-related information and also use the LSTM module to associate the information over time to complete the data storage of scene features:

[0136]

[0137] where, h t ' represents the map hidden state tracked by the module at time t; h t ' -1 represents the map hidden state at time t - 1.

[0138] Step S323. Merge the node state and time encoding through the time series feature fusion module in the time series encoding module. This module encodes the timestamp through a specific linear layer, encodes the time information into a high-dimensional vector. Then repeat the time encoding to the same dimension as the feature dimension, add it to the hidden state and set the feature mask, and update the hidden state through the residual module. Finally, the output is the node time series feature at the current time step.

[0139] In step S4, construct a trajectory decoding module and a score decoding module to model the probability distribution of multi-modal trajectory prediction, specifically including the following steps:

[0140] S41. Considering the agent's own characteristics, interaction characteristics, and map characteristics comprehensively, the task of trajectory decoding is completed through the GRU network structure, and the coordinates of the trajectory at each time step are output at the end:

[0141]

[0142] Among them, is the multi-modal future decoder of the agent, is the concatenation of features in the traffic task, is the predicted future trajectory of the k-th mode of the agent.

[0143] S42. The score decoder follows the same structure as the trajectory decoder, and the confidence score of the proposed trajectory is output at the end. Further, a softmax layer is used to generate the generation probability distribution of the multi-modal trajectory:

[0144]

[0145] In step S5, a loss function is designed according to the trajectory decoding module and the score decoding module for model training. For the trajectory loss, in this embodiment, a smooth loss is used at all predicted time steps. The trajectory regression loss is defined as:

[0146]

[0147] Among them is the true trajectory position at time t. In this embodiment, only the loss between the true trajectory and the closest predicted trajectory is calculated. k * is the index in the multi-modal predicted trajectories that is closest to the true trajectory position. k * is obtained by calculating the distance between the end point of the multi-modal predicted trajectory and the end point of the true trajectory:

[0148]

[0149] For the loss function of the predicted score, in this embodiment, a cross-entropy loss function is designed to calculate the loss between the true score (probability distribution) P gt and the predicted probability distribution P, and the predicted score loss is defined as:

[0150]

[0151] Here, in this embodiment, the score distribution corresponding to the true trajectory also needs to be defined:

[0152]

[0153] The total loss is the weighted sum of the trajectory regression loss and the score loss:

[0154]

[0155] Where δ is a parameter set to balance two learning objectives, and the model is trained with this as the objective function.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A multimodal trajectory prediction method based on spatio-temporal decoupling encoding of vehicle interaction graphs, characterized in that: It includes the following steps: S1. Generate graph structure data according to the traffic scene trajectory dataset; S2. Construct an interaction graph considering multiple traffic factors based on the graph structure data; S3. When the interaction graph updates features, capture the changes in the node features of the interaction graph over time through the temporal encoding module; S4. Complete multimodal prediction by constructing a trajectory decoding module and a score decoding module; S5. Design the loss objective function of the prediction model and conduct training; In step S1, the graph structure data is represented as where the data source for constructing the graph structure data includes the historical motion states of the agent and the high-definition map information Among them, i represents the number of agents considering interactive prediction within a certain range. represents the coordinate information of the agent at time t and the heading angle In step S4, construct a trajectory decoding module and a score decoding module, and model the probability distribution of multimodal trajectory prediction, which specifically includes the following steps: S41. Comprehensively consider the agent's own features, interaction features, and map features, and complete the task of trajectory decoding through the GRU network structure, and output the coordinates of the trajectory at each time step at the end: Among them, is the surrogate multi-modal future decoder, is the concatenation of features in the traffic task, f t i |k is the predicted future trajectory of the k-th mode of the agent; S42. The score decoder follows the same structure as the trajectory decoder, and outputs the confidence score of the proposed trajectory at the end, and further generates the generation probability distribution of the multimodal trajectory through the softmax layer: Among them, c i represents the trajectory confidence score, indicating the confidence range.

2. The multi-modal trajectory prediction method based on spatio-temporal decoupling encoding of vehicle interaction graphs according to claim 1, characterized in that: In step S2, it includes the following steps: S21. Extract historical vehicle node features according to the historical motion state Extract historical vehicle node features; the historical motion state Output vehicle node features through a ModuleList model list composed of a fully connected layer, a regularization layer, an activation layer, a GRU encoding layer, and a residual connection module: Among them, is the historical feature of an agent at time t; denotes encoding through a list of sequential models of a specific agent type, where the superscript at represents a specific type of agent and the subscript hist represents historical information; N t is the historical feature of all agents output by this module; S22. According to the global high-definition map information Extract the lane line node features; for the high-definition map information Establish lane objects through lanelet2, extract the lane centerlines, determine the lane line nodes, and calculate the distance between the agent and the lane line nodes: Among them represent the start and end indices of each lane centerline within the threshold distance from the agent i at time t; S23. Establish the interactions between vehicle agents and the interactions between the agents and the lanes; among them, for the interactions between vehicle agents, the other agents within the specified range of the target agent node features are used as neighbor nodes and embedded into their own node features, a directed edge pointing from the neighbor node to the target node is constructed, and the adjacency matrix is used to represent the connection with the neighbor agent. At the same time, the relative position information between nodes is calculated and embedded into the edge attribute variable constructed in the PyTorch_Geometric library to represent the interaction between agent nodes, that is: Among them, embeds the relative position relationship between nodes as the set of edge attributes of nodes i and j, E t is a set containing edge indices, edge attributes, and edge types; For the implicit interaction between the vehicle agent and the lane lines, the interaction between the agent and the lane is represented in the form of nodes and edges in the graph structure. Given the node features of the agent and the edge features the indices of the target lane line nodes are filtered according to the distance and the distances between the corresponding agent nodes and lane line nodes are embedded into the edge attribute variables constructed in the PyTorch_Geometric library to capture the implicit interaction between the agent and the lane; S24. Apply the multi-head self-attention mechanism to update the attention-weighted features of the node features, and realize the construction and update of the interaction graph considering multi-source information: MH(Q, K, V) = [head1, head2, head3]W head i = α(QW i Q ,KW i K ,VW i V ) In the multi-head self-attention mechanism, three attention heads are set, and Q, K, and V are the query vector, key vector, and value vector respectively, where W, W i Q , W i K , W i V are weight matrices; when calculating the attention score α, scaled dot-product attention is used: Among them, softmax(·) represents the normalization layer, and (·) T represents matrix transpose, and d k represents the dimension of the key vector.

3. The multimodal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs according to claim 1, wherein: In step S3, it includes the following steps: S31. The target node completes the update of its own feature matrix N by aggregating neighbor information, and during the aggregation update, the self-attention mechanism between connected nodes is used to focus on the dynamic update results of features over time t which is expressed as follows: It is expressed as: Among them, is the linear transformation of the feature matrix of the previous time step, as the attention-weighted feature matrix of the target agent therein, S32. For the attention-weighted feature matrix, explicitly encode the time information through the temporal encoding module to obtain the final feature matrix of the target agent at the current time step: Among them, g0 and g1 are composed of a double-layer MLP network. It is a temporal feature fusion module. The temporal feature fusion module encodes timestamps through a specific linear layer, encodes time information into high-dimensional vectors; then repeats the time encoding to be the same as the feature dimension, adds it to the hidden state and sets a feature mask, and then updates the hidden state through a residual module, and finally outputs the node temporal features of the current time step. Moreover, store the state features corresponding to the time step through the data replay module, and provide historical sequence information to the temporal encoding module; among them, at least set data replay modules for tracking and storing the target agent features and scene features respectively.

4. The multimodal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs according to claim 3, characterized in that: In step S32, for the data replay module used to track and store the target agent features, for the time features and hidden states of the target agent at time t, use LSTM to sequentially track the changes in its representation, and the formula is expressed as: where h t is the hidden state of the agent in the LSTM, and h t-1 is the input target node feature, denotes the feature matrix of the target agent at time t; in the interactivity prediction, the target agents are alternating. When the target agent of interest is confirmed, the remaining traffic participants act as neighbor agents to participate in the calculation of the interactive time series graph, and each agent for prediction acts as a target agent to store the node features in the data replay module.

5. The multimodal trajectory prediction method based on spatio-temporal decoupled encoding of vehicle interaction graphs according to claim 4, characterized in that: In step S32, when building the data playback module of the scene features, at time step t, by applying the linear layer initialization module to the feature matrix N t and then constructing a multi-layer self-attention operation network to store the scene features: where g2 is the linear layer for extracting the feature matrix, l represents the number of layers, and the storage of scene features can also ensure that the subsequent decoding module fully considers the map features when outputting the trajectory; Layer normalization is performed on each self-attention layer built. After the last layer L, the node features are aggregated through a max-pooling operation to summarize the scene-related information And the information is also temporally correlated through an LSTM module to complete the data storage of the scene features: where h' t represents the hidden state of the map tracked by the module at time t; h' t-1 represents the hidden state of the map at time t - 1.

6. The multi-modal trajectory prediction method based on spatio-temporal decoupling coding of vehicle interaction graphs according to claim 1, characterized in that: In step S5, design a loss function according to the trajectory decoding module and the score decoding module to train the model; For the trajectory loss, a smoothed loss is used over all predicted time steps, and the trajectory regression loss is defined as: Among them is the true trajectory position at time t, is the predicted trajectory closest to the true trajectory position; k * is the index in the multi-modal predicted trajectory that is closest to the true trajectory position, which is obtained by calculating the distance between the end point of the multi-modal predicted trajectory and the end point of the true trajectory as follows: Design the cross-entropy loss function for the loss function of the predicted score Calculate the true score P gt Calculate the loss between the true score P and the predicted probability distribution P, and define the predicted score loss as: Define the score distribution corresponding to the true trajectory: The total loss is the weighted sum of the trajectory regression loss and the score loss: where δ is a parameter set to balance the two learning objectives.

Citation Information

Patent Citations

  • Vehicle track prediction method based on intention perception space-time attention network

    CN117141518A

  • Multi-modal space-time model for accurate motion prediction based on visual fusion

    CN117315603A