Automatic driving track prediction method based on graph neural network
Through a graph neural network-based method, combined with graph converter and attention mechanism, the problem of map feature loss in convolutional neural network in autonomous driving trajectory prediction is solved, the prediction accuracy and safety are improved, and the intelligence of autonomous driving vehicles is enhanced.
Patent Information
- Application Number
- CN202311438465.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-31
- Publication Date
- 2025-07-08
AI Technical Summary
The existing convolutional neural networks lose map features in the prediction of autonomous driving trajectory, resulting in a decrease in prediction accuracy, and the single-modal prediction results do not meet the uncertainty and driving safety requirements of human drivers.
Using a graph neural network-based method, combining long-term network structure, graph converter and attention mechanism, an autonomous driving trajectory prediction model is constructed by obtaining the historical trajectory of the target vehicle and surrounding agents, and using graph converter to extract road features and embed node encoding to enhance the feature weight judgment of the motion capture attention module.
It improves the accuracy of trajectory prediction, enhances the safety and intelligence of autonomous vehicles, and can better retain road structure data and information.
Smart Images

Figure CN120279512A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to, and in particular to, an automatic driving trajectory prediction method based on a graph neural network. Background Art
[0002] In the current field of autonomous driving, the intelligence of vehicles is gradually increasing, and the importance of trajectory prediction modules is also increasing. At present, the network structure that appears more frequently for vehicle trajectory prediction is the temporal neural network to encode the historical trajectory of the vehicle intelligent body, and the road topology data of the high-precision map is processed by the convolutional neural network, and the output is mostly single-modal trajectory information. However, the convolutional neural network has a lot of map feature loss in the processing, resulting in a decrease in the contribution of map information to prediction accuracy and poor prediction results. At the same time, the single-modal prediction results do not meet the uncertainty of human drivers driving vehicles and the requirements for driving safety. The graph neural network can use the road vector as an information source to extract features, which can retain the road detail features to a greater extent, further improve the role of the graph neural network in the application of road trajectories in trajectory prediction, and improve prediction accuracy has become an important development direction. Summary of the invention
[0003] In view of this, the present invention provides an autonomous driving trajectory prediction method based on graph neural network, which improves the accuracy of trajectory prediction by combining long- and short-time network structure, graph converter and attention mechanism.
[0004] The deep learning prediction algorithm of the present invention includes a complete workflow from upstream module data input to target vehicle trajectory output, which is composed of a decoding and fusion module based on LSTM and graph converter, and a motion capture attention module from top to bottom. It aims to improve the feature extraction capability of complex lane structures in complex urban roads, optimize the problem of low sensitivity of the decoder to timing information, and further improve the accuracy of target vehicle trajectory prediction, thereby ensuring driving safety.
[0005] The present invention provides an automatic driving trajectory prediction method based on a graph neural network, comprising the following steps:
[0006] S1. Obtain a sample data set, the sample data set includes: the historical trajectory of the target vehicle, the historical trajectory of the intelligent body around the target vehicle, and the road vector map;
[0007] S2. Build an autonomous driving trajectory prediction model;
[0008] S3. Input the sample data set into the autonomous driving trajectory prediction model for training;
[0009] S4. Determine whether the autonomous driving trajectory prediction model is trained. If so, proceed to step S5; if not, update the parameters in the autonomous driving trajectory prediction model and return to step S3;
[0010] S5. Collect the historical trajectory of the target vehicle, the historical trajectories of the agents around the target vehicle, and the vector map of the road on which the target vehicle travels in real time, and input the collected data into the trained autonomous driving trajectory prediction model to output the predicted trajectory.
[0011] Further, in step S1, the historical trajectory A of the target vehicle 0:t-1 =[A0, A1, …, A t-3 , A t-2 , A t-1 , where t represents the sequence number of the current time step, and t - 1 represents the sequence number of the previous time step of the current time step, and represent the absolute coordinates of the target vehicle from the bird's-eye view at the (t - 1)-th time step, represents the speed of the target vehicle at the (t - 1)-th time step, represents the acceleration of the target vehicle at the (t - 1)-th time step;
[0012] The historical trajectory of the agents around the target vehicle where i represents the agent sequence number, and represent the absolute coordinates of the i-th agent around the target vehicle from the bird's-eye view at the (t - 1)-th time step, represents the speed of the i-th agent around the target vehicle at the (t - 1)-th time step, represents the acceleration of the i-th agent around the target vehicle at the (t - 1)-th time step, D i represents the i-th agent;
[0013] The road vector map
[0014] where N represents the nodes in the lane map data, and H represents the historical trajectory sequence number, and represent the absolute coordinates of the lane nodes in the H-th historical trajectory from the bird's-eye view, represents the geometric angle of the lane line in the H-th historical trajectory, represents the binary vector of the road in the H-th historical trajectory.
[0015] Further, in step S2, the autonomous driving trajectory prediction model includes: a sequential encoder, an aggregation encoding module, and a motion capture attention module;
[0016] Among them, the output end of the sequential encoder is connected to the input end of the aggregation encoding module, and the output end of the aggregation encoding module is connected to the input end of the motion capture attention module;
[0017] The aggregation encoding module includes: a multi-head attention mechanism module, a splicing module, a graph transformer, and a fusion traversal module. The output end of the multi-head attention mechanism module is connected to the input end of the splicing module, the output end of the splicing module is connected to the input end of the graph transformer, the output end of the graph transformer is connected to the input end of the fusion traversal module, and the output end of the fusion traversal module is connected to the input end of the motion capture attention module.
[0018] Furthermore, the graph transformer includes two layers of graph transformation networks, and the output end of the first layer of graph transformation network is connected to the input end of the second layer of graph transformation network;
[0019] The graph transformation network includes a Product layer, a Scaling layer, a Softmax function layer, a dot product module, a SUM module, two Dropout layers, two Add modules, two Norm modules, and a feed-forward neural network;
[0020] Input Q, K, and V into the graph transformation network. Among them, Q and K are input into the Product layer, V is input into the dot product module, the output end of the Product layer is connected to the input end of the Scaling layer, the output end of the Scaling layer is connected to the input end of the Softmax function layer, the output end of the Softmax function layer is connected to the input end of the dot product module, the output end of the dot product module is connected to the input end of the SUM module, the output end of the SUM module is connected to the input end of the first Dropout layer, the output end of the first Dropout layer is connected to the first Add module, the output end of the first Add module is connected to the input end of the first Norm module, the output end of the first Norm module is connected to the input end of the second Add module and the input end of the feed-forward neural network, the output end of the feed-forward neural network is connected to the input end of the second Dropout layer, the output end of the second Dropout layer is connected to the input end of the second Add module, and the output end of the second Add module is connected to the input end of the second Norm module.
[0021] Furthermore, the lane graph data is fused with the output of the dot product module in the form of a mask.
[0022] Furthermore, in step S3, set the initial learning rate to 1×10 -4 , adopt the Nadma optimizer with a selected batch size of 32, and train using the following loss function L:
[0023] L = ω1L1 + ω2L2 + ω3L3
[0024]
[0025]
[0026]
[0027] Among them, ω1, ω2, and ω3 represent weights, M represents the number of predictions, s represents the time step number, and S represents the total number of time steps. represents the predicted position at the s-th time step. represents the actual position at the s-th time step. represents the sum of the Euclidean distances between the predicted position and the actual position, and L1 represents the minimum value of the average of the Euclidean distances between the predicted positions and the actual positions in the previous M predictions. represents that the error of the Euclidean distance between the predicted position and the actual position exceeds two meters, and L2 represents the ratio of the number of times the error between the predicted position and the actual position exceeds two meters in M predictions; ξ represents the nodes of the lane map data, γ represents the edges of the lane map data, q represents the initial node in the lane map data, r represents any node other than the initial node in the lane map data, and C(q,r) represents the probability score generated by the multi-layer perceptron from the initial node q to the node r.
[0028] Furthermore, in step S4, when the parameters in the autonomous driving trajectory prediction model converge to the optimal and no overfitting phenomenon occurs, the training of the autonomous driving trajectory prediction model is completed.
[0029] Among them, when the loss function obtains the minimum value, the parameters in the autonomous driving trajectory prediction model converge to the optimal.
[0030] Advantages of the present invention: By building a sequential encoder, a graph converter, and a motion capture attention module, different from the traditional method of extracting road features based on convolution, the graph converter is introduced to extract graph data features and embed them into node encodings, better retaining road structure data and information. And the motion capture attention module is used to increase the judgment of the output trajectory feature weights. It meets the needs of multi-modal trajectory prediction, further improves the prediction accuracy, and enhances the safety and intelligence of autonomous driving vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The present invention will be further described below in conjunction with the drawings and embodiments:
[0032] Figure 1 is the flowchart of the present invention;
[0033] Figure 2 is the schematic diagram of the autonomous driving trajectory prediction model of the present invention;
[0034] Figure 3 This is a schematic diagram of the graph transformation network of the present invention. Detailed implementation manners
[0035] The present invention will be further described below in conjunction with the accompanying drawings of the specification:
[0036] The present invention provides an autonomous driving trajectory prediction method based on a graph neural network, which includes the following steps:
[0037] S1. Obtain a sample data set, which includes: the historical trajectory of the target vehicle, the historical trajectories of intelligent agents around the target vehicle, and the road vector map;
[0038] S2. Build an autonomous driving trajectory prediction model; the autonomous driving trajectory prediction model includes: a sequential encoder, an aggregation encoding module, and a motion capture attention module;
[0039] Among them, the output end of the sequential encoder is connected to the input ends of the masked multi-head attention mechanism module, the splicing module, and the fusion traversal module. The output end of the masked multi-head attention mechanism module is connected to the input end of the splicing module. The output end of the splicing module is connected to the input end of the graph converter. The output end of the graph converter is connected to the input end of the fusion traversal module. The output end of the fusion traversal module is connected to the input end of the motion capture attention module;
[0040] S3. Input the sample data set into the autonomous driving trajectory prediction model for training;
[0041] S4. Determine whether the autonomous driving trajectory prediction model is trained. If so, enter step S5. If not, update the parameters in the autonomous driving trajectory prediction model and return to step S3;
[0042] S5. Real-time collect the historical trajectory of the target vehicle, the historical trajectories of intelligent agents around the target vehicle, and the vector map of the road on which the target vehicle travels, and input the collected data into the trained autonomous driving trajectory prediction model to output the predicted trajectory. Through the above method, road structure data and information can be better retained, the accuracy of trajectory prediction can be improved, and thus the safety and intelligence level of autonomous driving vehicles can be improved.
[0043] In this embodiment, in step S1, the sample data set is obtained through 6 cameras, 1 lidar, 5 radars, and GPS, IMU, etc. The sample data set includes: the historical trajectory of the target vehicle, the historical trajectories of intelligent agents around the target vehicle, and the road vector map;
[0044] The historical trajectory A of the target vehicle 0:t-1 =[A0, A1,..., A t-3 , A t-2 , A t-1, Among them, t represents the serial number of the current time step, and t - 1 represents the serial number of the previous time step of the current time step. and represent the absolute coordinates of the target vehicle from an aerial view at the (t - 1)-th time step. represents the speed of the target vehicle at the (t - 1)-th time step. represents the acceleration of the target vehicle at the (t - 1)-th time step;
[0045] The historical trajectories of the agents around the target vehicle Among them, i represents the agent serial number. and represent the absolute coordinates of the agents around the target vehicle from an aerial view at the (t - 1)-th time step for the i-th agent. represents the speed of the agents around the target vehicle at the (t - 1)-th time step for the i-th agent. represents the acceleration of the agents around the target vehicle at the (t - 1)-th time step for the i-th agent, D i represents the i-th agent. When D i = 0, the agent is a person. When D i = 1, the agent is a vehicle;
[0046] The historical trajectories of the agents around the target vehicle include the trajectories of the agents around the target vehicle sensed by the target vehicle using sensors and the historical trajectories of the agents around the target vehicle used during the training process;
[0047] Road vector map
[0048] Among them, N represents the lane nodes in the lane map data, and H represents the historical trajectory serial number. and represent the absolute coordinates of the lane nodes in the H-th historical trajectory from an aerial view. represents the geometric angle of the lane line in the H-th historical trajectory. represents the binary vector of the road in the H-th historical trajectory, and the information represented is whether the current lane node N is on a stop line or a zebra crossing; among them, the road vector map is a vector in which road features are embedded in node encoding, and the road feature vector is obtained from the lane map data. The lane map data contains information about lane lines and lane nodes. Therefore, the essence of encoding the road vector map is to encode the lane lines. Using the above data, the interaction patterns and relationships between the target vehicle and other vehicles and pedestrians can be captured as much as possible, thereby improving the prediction accuracy.
[0049] In this embodiment, in step S2, an autonomous driving trajectory prediction model is built. The autonomous driving trajectory prediction model includes: a sequential encoder, an aggregation encoding module, and a motion capture attention module; as Figure 2 shown;
[0050] Among them, the output end of the sequential encoder is connected to the input end of the aggregation encoding module, and the output end of the aggregation encoding module is connected to the input end of the motion capture attention module;
[0051] The sequential encoder includes three parallel encoders. The encoder includes a linear layer (Linear), an activation function layer, and a long short-term memory network. Among them, the output end of the linear layer is connected to the input end of the activation function layer, and the output end of the activation function layer is connected to the input end of the long short-term memory network (Long Short-term Memory Networks, LSTM). The activation function layer uses the LeakyRelu non-linear activation function. Compared with the Relu function, the LeakyRelu non-linear activation function can avoid being completely suppressed when the input is less than 0, thus preventing neurons from becoming completely inactive.
[0052] Specifically, the historical trajectory of the target vehicle, the historical trajectories of the agents around the target vehicle, and the road vector map are respectively input into the encoders of the three parallel encoders for encoding. The encoding formula is:
[0053] Encoding=LSTM(LeakyRelu(Linear(F)))
[0054] Among them, Encoding represents the encoder, and F represents the input data, including the historical trajectory of the target vehicle, the historical trajectories of the agents around the target vehicle, and the road vector map;
[0055] The first encoder obtains the encoding of the historical trajectory of the target vehicle The second encoder obtains the encoding of the historical trajectories of the agents around the target vehicle The third encoder obtains the encoding of the road vector map
[0056] The aggregation encoding module includes: a multi-head attention mechanism module, a splicing module, a graph converter, and a fusion traversal module. The output end of the multi-head attention mechanism module is connected to the input end of the splicing module, the output end of the splicing module is connected to the input end of the graph converter, the output end of the graph converter is connected to the input end of the fusion traversal module, and the output end of the fusion traversal module is connected to the input end of the motion capture attention module;
[0057] Input the encoding output by the sequential encoder into the aggregation encoding module, which specifically includes the following content:
[0058] The multi-head attention mechanism module adopts an existing module, and its specific structure will not be elaborated here;
[0059] Encode the historical trajectories of the agents around the target vehicle Respectively embed them through a linear layer as the Key and Value of the multi-head attention mechanism module, and encode the road vector map Embed it through a linear layer as the Query of the multi-head attention mechanism module, and then input the Query (Q), Key (K) and Value (V) into the multi-head attention mechanism module to obtain the multi-head attention features and the weights of the road vector map encoding;
[0060] The calculation formula of the multi-head attention mechanism module is:
[0061] MultiHead(Q, K, V) = Concat(head1, head2,..., head g )
[0062]
[0063] Among them, MultiHead represents the multi-head attention mechanism, Concat represents concatenation, head g represents the g-th head, softmax represents the normalized exponential function, Q represents the query, K represents the key, K T represents the transpose of the key, V represents the value, and d k represents the scaled factor;
[0064] Input the road vector map encoding and the weights of the road vector map encoding into the splicing module, splice them and then input them into the graph transformer. The splicing module is an existing technology and will not be elaborated here;
[0065] The graph transformer includes two layers of graph transformation networks. The output end of the first layer of graph transformation network is connected to the input end of the second layer of graph transformation network. The structure of the graph transformation network is as Figure 3 shown, including a Product layer, a Scaling layer, a Softmax function layer, a dot product module, a SUM module, two Dropout layers, two Add modules, two Norm modules, and a Feedforward Neural Nettwork (FFN);
[0066] Specifically, encode them respectively through an embedding layer to obtain the Query (Q), Key (K) and Value (V) input to the graph transformer; among them, this embedding layer is notFigure 2 is drawn in, and the embedding layer adopts the existing technology;
[0067] Among them, Q and K are input into the Product layer, V is input into the dot product module. The output end of the Product layer is connected to the input end of the Scaling layer. The output end of the Scaling layer is connected to the input end of the Softmax function layer. The output end of the Softmax function layer is connected to the input end of the dot product module. The output end of the dot product module is connected to the input end of the SUM module. The output end of the SUM module is connected to the input end of the first Dropout layer. The output end of the first Dropout layer is connected to the first Add module. The encoding obtained by adding the output of the position encoding module and the output of the splicing module is also input into the first Add module. The output end of the first Add module is connected to the input end of the first Norm module. The output end of the first Norm module is connected to the input end of the second Add module and the input end of the FFN. The output end of the FFN is connected to the input end of the second Dropout layer. The output end of the second Dropout layer is connected to the input end of the second Add module. The output end of the second Add module is connected to the input end of the second Norm module; the output of the second Norm module is transformed into the input of the second graph transformation network through the embedding layer;
[0068] Among them, the lane graph data is fused with the output of the dot product module in the form of a mask, and the lane graph data is represented by an adjacency matrix G(ξ, γ); the SUM module is used to sum the outputs of each head of the masked multi-head attention module;
[0069] The calculation formula of the graph transformation network is:
[0070]
[0071] Among them, Output represents the output of the graph transformation network, Add represents addition, FFN represents a feed-forward neural network, h represents the input tensor of the graph transformation network, represents the input tensor of the FFN feed-forward neural network;
[0072] Compared with the existing graph transformation networks (Graph Transformer Networks, GTNs), the graph transformation network in this application enhances the learning of lane graph data and also increases the dependence relationship on surrounding nodes. The above structure can aggregate nodes more effectively and improve the prediction accuracy. Since the performance of a single-layer graph transformation network is poor and the three-layer graph transformation network will have an overfitting phenomenon, therefore, the graph converter in this application adopts a two-layer graph transformation network; through the above structure, it can further capture the complex mutual relationships between nodes and enhance the processing ability for nodes and edges.
[0073] The fusion traversal module includes two splicing modules, an MLP probability traversal module, a position encoding module, and an attention mechanism module. Among them, the input features of the first splicing module are the features output by the graph transformer and the features output by the first encoder. The output end of the first splicing module is connected to the input end of the MLP probability traversal module. The features output by the MLP probability traversal module and the features output by the position encoding module are added together and then input into the multi-head attention mechanism module. The input features of the multi-head attention mechanism module also include the output features of the first encoder. The output end of the multi-head attention mechanism module is connected to the input end of the second splicing module. The input features of the second splicing module also include the output features of the first encoder;
[0074] Among them, the MLP probability traversal module includes an MLP (multi-layer perceptron) and a trajectory traversal module;
[0075] In the fusion traversal module, the historical trajectory of the target vehicle and the road vector map encoding generate interactive nodes, and the probability related to the nodes during the driving process of the target vehicle is obtained through a multi-layer perceptron, so as to obtain the multi-modal features that appear during the driving process of the target vehicle;
[0076] The motion capture attention module includes a position encoding module, a multi-head attention mechanism module, two Add modules (addition modules), a feed-forward neural network, two linear layers, an activation function layer, and a clustering module. Among them, the output features of the position encoding module and the input features of the motion capture attention module are added together and then input into the multi-head attention mechanism module and the first Add module. The output end of the multi-head attention mechanism module is connected to the input end of the first Add module. The output end of the first Add module is connected to the input end of the feed-forward neural network and the input end of the second Add module. The output end of the feed-forward neural network is connected to the input end of the second Add module. The output end of the second Add module is connected to the input end of the first linear layer. The output end of the first linear layer is connected to the input end of the activation function layer. The output end of the activation function layer is connected to the input end of the second linear layer. The output end of the second linear layer is connected to the input end of the clustering module;
[0077] Specifically, the motion capture attention module first performs positional encoding on the input encoding; then inputs the encoded encoding after positional encoding into the multi-head attention mechanism module for feature extraction; secondly, forms a residual connection between the extracted features and the encoding input to the multi-head attention mechanism module through the Add module to obtain a tensor combining attention weights and encoding; then, inputs the tensor into the feed-forward neural network and then into the Add module. The feed-forward neural network maps the tensor into a feature space of different dimensions, which can enhance the modeling ability of the autonomous driving trajectory prediction model for complex relationships, and the non-linear activation function in the feed-forward neural network helps the autonomous driving trajectory prediction model perform dimensional transformation; finally, sequentially inputs the output features of the Add module into the linear layer, the activation function layer, and the linear layer for decoding, and inputs the decoded features into the clustering module for clustering to obtain the final trajectory prediction result. In this embodiment, 10 prediction results are set to be output; among them, the activation function layer uses the LeakyRelu non-linear function.
[0078] In this embodiment, in step S3, the sample data set is input into the autonomous driving trajectory prediction model for training, and the initial learning rate is set to 1×10 -4 , and the Nadma optimizer with a batch size of 32 is used, and the following loss function L is used for training:
[0079] L = ω1L1 + ω2L2 + ω3L3
[0080]
[0081]
[0082]
[0083] Among them, ω1, ω2, and ω3 represent weights, ω1 = 1, ω2 = 0.25, ω3 = 0.5, M represents the number of predictions, M = 5, s represents the time step number, S represents the total number of time steps, represents the predicted position at the s-th time step, represents the true position at the s-th time step, represents the sum of the Euclidean distances between the predicted position and the true position, and L1 represents the minimum value of the average of the Euclidean distances between the predicted positions and the true positions of the first M predictions; It is indicated that the error of the Euclidean distance between the predicted position and the true position exceeds two meters. L2 represents the ratio of the number of times that the error between the predicted position and the true position exceeds two meters in M predictions; ξ represents the nodes of the lane map data, γ represents the edges of the lane map data, q represents the initial node in the lane map data, r represents any node other than the initial node in the lane map data, and C(q, r) represents the probability score generated by the multi-layer perceptron from the initial node q to the node r.
[0084] Through the above three loss functions, different weights are assigned to evaluate the average consistency, prediction stability of the model prediction, and the probability of the edges in the lane map, so that the parameters of the model converge to the optimal as quickly as possible.
[0085] In this embodiment, in step S4, it is judged whether the autonomous driving trajectory prediction model is trained. When the parameters in the autonomous driving trajectory prediction model converge to the optimal and no overfitting phenomenon occurs, the autonomous driving trajectory prediction model is trained.
[0086] Among them, when the loss function obtains the minimum value, the parameters in the autonomous driving trajectory prediction model converge to the optimal. Through the above method, it can be ensured that the trained autonomous driving trajectory prediction model is optimal.
[0087] In this embodiment, in step S5, the historical trajectory of the target vehicle, the historical trajectories of the intelligent agents around the target vehicle, and the vector map of the road on which the target vehicle travels are collected in real time, and the collected data is input into the trained autonomous driving trajectory prediction model to output the predicted trajectory of the target vehicle.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for predicting an autonomous driving trajectory based on a graph neural network, characterized in that: It includes the following steps: S1. Obtain a sample data set, which includes: the historical trajectory of the target vehicle, the historical trajectories of the agents around the target vehicle, and the road vector map; S2. Build an autonomous driving trajectory prediction model; S3. Input the sample data set into the autonomous driving trajectory prediction model for training; S4. Determine whether the autonomous driving trajectory prediction model is trained. If so, go to step S5; if not, update the parameters in the autonomous driving trajectory prediction model and return to step S3; S5. Real-time collect the historical trajectory of the target vehicle, the historical trajectories of the agents around the target vehicle, and the vector map of the road on which the target vehicle travels, and input the collected data into the trained autonomous driving trajectory prediction model to output the predicted trajectory.
2. The method for predicting an autonomous driving trajectory based on a graph neural network according to claim 1, wherein: In step S1, the historical trajectory A of the target vehicle 0:t-1 = [A0, A1, …, A t-3 , A t-2 , A t-1 , where t represents the sequence number of the current time step, and t - 1 represents the sequence number of the previous time step of the current time step. and represent the absolute coordinates of the target vehicle from an aerial view at the (t - 1)-th time step, represents the speed of the target vehicle at the (t - 1)-th time step, represents the acceleration of the target vehicle at the (t - 1)-th time step; Historical trajectories of agents around the target vehicle where i represents the agent number, and represents the absolute coordinates of the agents around the target vehicle from an aerial view at the (t - 1)-th time step for the i-th agent, represents the speed of the agents around the target vehicle at the (t - 1)-th time step for the i-th agent, represents the acceleration of the agents around the target vehicle at the (t - 1)-th time step for the i-th agent, D i represents the i-th agent; Road vector map Wherein, N represents the nodes in the lane map data, and H represents the historical trajectory number. and represents the absolute coordinates of the lane nodes in the H-th historical trajectory from the bird's-eye view. represents the geometric angle of the lane line in the H-th historical trajectory. represents the binary vector of the road in the H-th historical trajectory.
3. The method for predicting an autonomous driving trajectory based on a graph neural network according to claim 1, wherein: In step S2, the autonomous driving trajectory prediction model includes: a sequential encoder, an aggregation encoding module, and a motion capture attention module; Among them, the output end of the sequential encoder is connected to the input end of the aggregation encoding module, and the output end of the aggregation encoding module is connected to the input end of the motion capture attention module; The aggregation encoding module includes: a multi-head attention mechanism module, a splicing module, a graph transformer, and a fusion traversal module. The output end of the multi-head attention mechanism module is connected to the input end of the splicing module, the output end of the splicing module is connected to the input end of the graph transformer, the output end of the graph transformer is connected to the input end of the fusion traversal module, and the output end of the fusion traversal module is connected to the input end of the motion capture attention module.
4. The method for predicting an autonomous driving trajectory based on a graph neural network according to claim 3, wherein: The graph transformer includes two layers of graph transformation networks, and the output end of the first layer of graph transformation network is connected to the input end of the second layer of graph transformation network; The graph transformation network includes a Product layer, a Scaling layer, a Softmax function layer, a dot product module, a SUM module, two Dropout layers, two Add modules, two Norm modules, and a feed-forward neural network; Input Q, K, and V into the graph transformation network. Among them, Q and K are input into the Product layer, V is input into the dot product module. The output end of the Product layer is connected to the input end of the Scaling layer, the output end of the Scaling layer is connected to the input end of the Softmax function layer, the output end of the Softmax function layer is connected to the input end of the dot product module, the output end of the dot product module is connected to the input end of the SUM module, the output end of the SUM module is connected to the input end of the first Dropout layer, the output end of the first Dropout layer is connected to the first Add module, the output end of the first Add module is connected to the input end of the first Norm module, the output end of the first Norm module is connected to the input end of the second Add module and the input end of the feed-forward neural network, the output end of the feed-forward neural network is connected to the input end of the second Dropout layer, the output end of the second Dropout layer is connected to the input end of the second Add module, and the output end of the second Add module is connected to the input end of the second Norm module.
5. The method for predicting an autonomous driving trajectory based on a graph neural network according to claim 4, wherein: The lane graph data is fused with the output of the dot product module in the form of a mask.
6. The method for predicting an autonomous driving trajectory based on a graph neural network according to claim 1, wherein: In step S3, set the initial learning rate to 1×10 -4 , adopt the Nadma optimizer with a selected batch size of 32, and perform training using the following loss function L: L = ω1L1 + ω2L2 + ω3L3 Among them, ω1, ω2, and ω3 represent weights, M represents the number of predictions, s represents the time step number, and S represents the total number of time steps. represents the predicted position at the s-th time step. represents the true position at the s-th time step. represents the sum of the Euclidean distances between the predicted position and the true position. L1 represents the minimum value of the average of the Euclidean distances between the predicted positions and the true positions in the first M predictions. represents that the error of the Euclidean distance between the predicted position and the true position exceeds two meters. L2 represents the ratio of the number of times the error between the predicted position and the true position exceeds two meters in M predictions. ξ represents the node of the lane map data, γ represents the edge of the lane map data, q represents the initial node in the lane map data, r represents any node other than the initial node in the lane map data, and C(q,r) represents the probability score generated by the multi-layer perceptron from the initial node q to the node r.
7. The method for predicting an autonomous driving trajectory based on a graph neural network according to claim 6, wherein: In step S4, when the parameters in the autonomous driving trajectory prediction model converge to the optimal values and no overfitting phenomenon occurs, the training of the autonomous driving trajectory prediction model is completed; Among them, when the loss function reaches the minimum value, the parameters in the autonomous driving trajectory prediction model converge to the optimal values.