An automatic trajectory prediction method based on graph spatiotemporal pyramid
By constructing an automatic trajectory prediction model based on graph space-time pyramids, combining space-time pyramid network, graph convolution network and Transformer network, the existing methods have solved the shortcomings in trajectory prediction accuracy and category prediction, and achieved higher prediction accuracy and category prediction capabilities.
Patent Information
- Application Number
- CN202310287277.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-03-22
AI Technical Summary
Existing trajectory prediction methods based on deep learning are difficult to meet the accuracy requirements and category prediction of traffic participants' prediction trajectory, especially in the case of dynamic changes in the number of surrounding traffic participants.
The automatic trajectory prediction method based on the graph space-time pyramid is adopted to build models of the space-time pyramid network, graph convolution network, graph attention network and Transformer network. The bird's-eye view is processed through the space-time pyramid network, and trajectory prediction and category prediction are combined with the graph convolution network and Transformer network.
The accuracy of trajectory prediction of traffic participants around autonomous driving vehicles is improved, and category prediction of traffic participants is realized, reducing the impact of motion uncertainty on trajectory prediction.
Smart Images

Figure CN116128930B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of trajectory prediction, and in particular to an automatic trajectory prediction method based on a graph spatiotemporal pyramid. Background Art
[0002] In order to ensure the safe navigation of autonomous vehicles in traffic, the autonomous driving system needs to accurately predict the future position or movement of surrounding traffic participants to avoid traffic accidents. Previously, the trajectory prediction algorithm for traffic participants around the target vehicle was mainly based on vehicle dynamics, and the future position of traffic participants was predicted by identifying maneuvers (changing lanes, braking, or continuing to move forward, etc.). Now it is based on deep learning, using various neural networks to extract data features of historical trajectories, and then fusing different features into long-term series operations to obtain historical trajectories that are close to the real ones.
[0003] At present, deep learning-based methods are mainly divided into two categories: one is to assume that the target vehicle is only affected by a fixed number of nearest traffic participants; the other is to apply a circle to cover the geometric center of the target vehicle to select the surrounding traffic participants. However, both methods only consider the interaction between a fixed number of traffic participants, or force the input size to be the same. However, since the number of surrounding traffic participants is in dynamic change in actual scenarios, there are certain errors when these two methods are applied to actual data.
[0004] The above-mentioned method of trajectory prediction using deep learning cannot simultaneously meet the accuracy requirements of predicted trajectories of traffic participants and category prediction. Summary of the invention
[0005] To solve the above problems, the present invention provides an automatic trajectory prediction method based on a graph spatiotemporal pyramid, and constructs an automatic trajectory prediction model based on a graph spatiotemporal pyramid, which includes a spatiotemporal pyramid network, a graph convolutional network, a graph attention network and a transformer network; so as to improve the accuracy of predicting the trajectories of traffic participants around an autonomous driving vehicle and realize the category prediction of traffic participants.
[0006] The method for automatic trajectory prediction using the model includes the following steps:
[0007] S1. Preprocess the radar point cloud data collected in real time by the sensors on the autonomous driving vehicle to obtain a bird's-eye view;
[0008] S2. Process the bird's-eye view image through the spatiotemporal pyramid network and output scene features;
[0009] S3. Screening and processing the conventional historical trajectory data collected by the sensors on the autonomous driving vehicle to obtain the original data;
[0010] S4. Further process the original data to obtain input representation and fixed graph, and input the input representation and fixed graph into the graph convolution network to obtain graph features;
[0011] S5. Input the graph features and scene features into the graph attention network to integrate them to obtain the spatiotemporal graph;
[0012] S6. Input the spatiotemporal graph into the transformer network and output the predicted trajectories and categories of traffic participants around the autonomous driving vehicle.
[0013] Furthermore, step S2 processes the bird's-eye view output scene features through a spatiotemporal pyramid network, including:
[0014] S21. Input the bird's-eye view image into the first spatiotemporal convolution block, and the output of the first spatiotemporal convolution block is subjected to time pooling to obtain the first spatiotemporal feature; input the first spatiotemporal feature into the second spatiotemporal convolution block, and the output of the second spatiotemporal convolution block is subjected to time pooling to obtain the second spatiotemporal feature; input the second spatiotemporal feature into the third spatiotemporal convolution block, and the output of the third spatiotemporal convolution block is subjected to time pooling to obtain the third spatiotemporal feature; input the third spatiotemporal feature into the fourth spatiotemporal convolution block, and the output of the fourth spatiotemporal convolution block is subjected to time pooling to obtain the fourth spatiotemporal feature;
[0015] S22. The fourth spatiotemporal feature and the third spatiotemporal feature are subjected to time pooling respectively and then feature fused to obtain a first fused feature; the second spatiotemporal feature is subjected to time pooling and then feature fused with the first fused feature to obtain a second fused feature; the first spatiotemporal feature is subjected to time pooling and then feature fused with the second fused feature to obtain a third fused feature; the bird's-eye view is subjected to time pooling and then feature fused with the third fused feature to obtain the final scene feature.
[0016] Furthermore, the process of obtaining the original data in step S3 is as follows:
[0017] S31. Obtaining conventional historical trajectory data collected by sensors on the autonomous driving vehicle and removing abnormal data therein;
[0018] S32. Parse the data format of the conventional historical trajectory data after the elimination process and convert it into a data frame;
[0019] S33. Perform data cleaning on the data frame obtained in S32, including removing unnecessary columns and rows, filling missing values and outliers, and converting data types;
[0020] S34. Extracting relevant information of surrounding vehicles or pedestrians from the data frame after the data cleaning in S33;
[0021] S35. Convert the relevant information extracted by S34 into a processing format for the machine learning model to obtain the original data.
[0022] Furthermore, step S4 further processes the original data to obtain an input representation and a fixed graph, including:
[0023] S41 extracts the location information of each surrounding vehicle from the raw data;
[0024] S42. Convert the position information of each surrounding vehicle into the offset of the current target vehicle;
[0025] S43. Arrange the offsets of all surrounding vehicles obtained in S42 in chronological order to obtain a trajectory;
[0026] S44. Divide the trajectory into multiple time windows, each time window containing a number of trajectory points;
[0027] S45. Integrate the trajectory points of each time window, and merge all the time windows after the integrated trajectory points to form an input representation;
[0028] S46. Map all trajectory points in each time window to a fixed graph structure through a specific function. The nodes in the graph structure represent surrounding vehicles, and the edges represent the relationships between surrounding vehicles. Finally, all time windows are merged to obtain a fixed graph.
[0029] Furthermore, step S4 simultaneously inputs the input representation and the fixed graph into the graph convolutional network to obtain graph features, including:
[0030] S51. The input representation is normalized by a batch normalization layer, and the normalized input representation is passed through a two-dimensional convolution layer to obtain a two-dimensional feature;
[0031] S52. Input the two-dimensional feature and the fixed graph into the first graph convolution layer to obtain the first feature; add the first feature and the two-dimensional feature, and input the addition result and the fixed graph into the second graph convolution layer to obtain the second feature; add the second feature and the first feature, and input the addition result and the fixed graph into the third graph convolution layer to obtain the third feature;
[0032] S53. Add the third feature to the second feature to obtain a graph feature.
[0033] Furthermore, the first graph convolution layer includes a graph operation layer, a batch normalization layer, a temporal two-dimensional convolution layer and a batch normalization layer that are cascaded in sequence; the structures of the second graph convolution layer and the third graph convolution layer are the same as those of the first graph convolution layer.
[0034] Furthermore, the expression of each graph operation layer is:
[0035]
[0036] Among them, f graphrepresents the output of the graph operation layer, f conv represents the output of the previous layer, represents the training graph of the jth layer in the graph operation layer, Represents the diagonal transformation fixed graph of the jth layer in the graph operation layer, and its calculation formula is:
[0037]
[0038] Among them, A j represents the jth layer of the input fixed graph A, ∧ j Indicated by A j is a matrix of eigenvectors.
[0039] Furthermore, in step S1, the radar point cloud data is quantized into regular voxels and a three-dimensional voxel grid is formed. The occupancy of each voxel grid is represented by a binary state, and the height dimension of the three-dimensional voxel grid corresponds to the image channel of the two-dimensional pseudo image, thereby converting the three-dimensional radar point cloud data into a two-dimensional pseudo image, i.e., the desired bird's-eye view.
[0040] Beneficial effects of the present invention:
[0041] The present invention proposes an autonomous driving trajectory prediction method based on a graph spatiotemporal pyramid, which combines the spatiotemporal pyramid network and the graph convolutional network, and then uses the Transformer mechanism to predict the trajectory and category of traffic participants around the autonomous driving car. The rich neural network model increases the prediction accuracy. At the same time, the Transformer mechanism uses position encoding to connect the input embedding with the position encoding vector, so that computational parallelism can be achieved even in the case of long-term sequence input, reducing the time required for model training.
[0042] The present invention is not limited to common radar point cloud data in the training set, but also adds conventional data on the roads where autonomous driving vehicles travel, so that the content contained in the training set is richer, thereby effectively enhancing the training effect and improving the reliability of the trained model.
[0043] The present invention uses Transformer to model positional relationships, making up for the lack of positional features in the spatiotemporal pyramid network. It not only takes into account the influence of surrounding traffic participants on each other, but also can perform category prediction and trajectory prediction on traffic participants, reducing the impact of motion uncertainty on trajectory prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A flowchart of an autonomous driving trajectory prediction method based on a graph spatiotemporal pyramid provided by the present invention;
[0045] Figure 2A framework diagram of an autonomous driving trajectory prediction model based on a graph spatiotemporal pyramid provided by the present invention;
[0046] Figure 3 A spatiotemporal pyramid network architecture diagram in an autonomous driving trajectory prediction method based on a spatiotemporal pyramid provided by the present invention;
[0047] Figure 4 A graph convolutional network architecture diagram in an autonomous driving trajectory prediction method based on a graph spatiotemporal pyramid provided by the present invention;
[0048] Figure 5 A graph attention network architecture diagram in an autonomous driving trajectory prediction method based on a graph spatiotemporal pyramid provided by the present invention;
[0049] Figure 6 This is a schematic diagram of the Transformer mechanism in an autonomous driving trajectory prediction method based on a graph spatiotemporal pyramid provided by the present invention. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] The present invention provides an automatic trajectory prediction method based on graph spatiotemporal pyramid, such as Figure 1 As shown, the following steps are included:
[0052] S1. Preprocess the radar point cloud data collected in real time by the sensors on the autonomous driving vehicle to obtain a bird's-eye view;
[0053] Specifically, the radar point cloud data is quantized into regular voxels to form a three-dimensional voxel grid, a binary state is used to represent the occupancy of each voxel grid, and the height dimension of the three-dimensional voxel grid corresponds to the image channel of the two-dimensional pseudo image, thereby converting the three-dimensional radar point cloud data into a two-dimensional pseudo image, i.e., the required bird's-eye view.
[0054] S2. Process the bird's-eye view image through the spatiotemporal pyramid network and output scene features;
[0055] S3. Screening and processing the conventional historical trajectory data collected by the sensors on the autonomous driving vehicle to obtain the original data;
[0056] Specifically, the process of obtaining the original data in step S3 is as follows:
[0057] S31. Obtaining conventional historical trajectory data collected by sensors on the autonomous driving vehicle and removing abnormal data therein;
[0058] S32. Parse the data format of the conventional historical trajectory data after the elimination process and convert it into a data frame;
[0059] S33. Perform data cleaning on the data frame obtained in S32, including removing unnecessary columns and rows, filling missing values and outliers, and converting data types;
[0060] S34. Extracting relevant information of surrounding vehicles or pedestrians from the data frame after the data cleaning in S33, the relevant information includes location, speed and direction, etc.;
[0061] S35. Convert the relevant information extracted by S34 into a processing format for the machine learning model to obtain the original data.
[0062] S4. Further process the original data to obtain input representation and fixed graph, and input the input representation and fixed graph into the graph convolution network to obtain graph features;
[0063] Specifically, step S4 further processes the original data to obtain an input representation and a fixed graph, including:
[0064] S41 extracts the location information of each surrounding vehicle from the raw data;
[0065] S42. Converting the position information of each surrounding vehicle into the offset of the current target vehicle; ie, moving the coordinate system of the surrounding vehicles to the position of the current target vehicle;
[0066] S43. Arrange the offsets of all surrounding vehicles obtained in S42 in chronological order to obtain a trajectory;
[0067] S44. Divide the trajectory into multiple time windows, each time window containing a number of trajectory points;
[0068] S45. Integrate the trajectory points of each time window, and merge all the time windows after the integrated trajectory points to form an input representation, which is a multi-layer structure;
[0069] S46. Map all trajectory points in each time window to a fixed graph structure through a specific function. The nodes in the graph structure represent surrounding vehicles, and the edges represent the relationships between surrounding vehicles. Finally, all trajectory points in all time windows are mapped to the fixed graph structure to obtain a fixed graph.
[0070] S5. Input the graph features and scene features into the graph attention network to integrate them to obtain the spatiotemporal graph;
[0071] S6. Input the spatiotemporal graph into the transformer network and output the predicted trajectories and categories of traffic participants around the autonomous driving vehicle.
[0072] In one embodiment, the automatic trajectory prediction model based on the graph spatiotemporal pyramid constructed by the present invention is as follows: Figure 2 As shown, it includes a spatiotemporal pyramid network, a graph convolutional network, a graph attention network, and a transformer network.
[0073] Specifically, Figure 3 As shown, step S2 processes the bird's-eye view output scene features through the spatiotemporal pyramid network, including:
[0074] S21. Input the bird's-eye view of size T×C×H×W into the first spatiotemporal convolution block, and the output of the first spatiotemporal convolution block is obtained by temporal pooling The first spatiotemporal feature; the first spatiotemporal feature is input into the second spatiotemporal convolution block, and the output of the second spatiotemporal convolution block is obtained by time pooling The second spatiotemporal feature; the second spatiotemporal feature is input into the third spatiotemporal convolution block, and the output of the third spatiotemporal convolution block is obtained through time pooling The third spatiotemporal feature; the third spatiotemporal feature is input into the fourth spatiotemporal convolution block, and the output of the fourth spatiotemporal convolution block is obtained by time pooling The fourth space-time characteristic;
[0075] S22. The fourth spatiotemporal feature and the third spatiotemporal feature are respectively subjected to time pooling and feature fusion to obtain The first fusion feature; the second spatiotemporal feature is obtained by fusing the second spatiotemporal feature with the first fusion feature after time pooling The second fusion feature: The first spatiotemporal feature is obtained by fusion with the second fusion feature after time pooling. The third fusion feature: After temporal pooling, the bird's-eye view image is fused with the third fusion feature to obtain the final 1×C×H×W scene feature.
[0076] Specifically, the first spatiotemporal convolution block, the second spatiotemporal convolution block, the third spatiotemporal convolution block, and the fourth spatiotemporal convolution block have the same structure, and all extract features along the spatial dimension and the temporal dimension in a hierarchical manner, such as Figure 3 As shown in the figure, in the spatial dimension, the feature maps at different scales are calculated with a proportional step size of 2 to obtain spatial features of different scales; in the temporal dimension, the temporal resolution is gradually reduced by a ratio of 1 / 2 after each temporal convolution to obtain temporal features of different scales, namely scene features.
[0077] Specifically, in step S22, the fourth spatiotemporal feature is further deconvolved after time pooling before being fused with the third spatiotemporal feature after time pooling; the first fusion feature, the second fusion feature and the third fusion feature must all be deconvolved before fusion.
[0078] In one embodiment, if Figure 4 As shown, step S4 simultaneously inputs the input representation and the fixed graph into the graph convolutional network to obtain graph features, including:
[0079] S51. The input representation is normalized by a batch normalization layer, and the normalized input representation is passed through a two-dimensional convolution layer to obtain a two-dimensional feature;
[0080] S52. Input the two-dimensional feature and the fixed graph into the first graph convolution layer to obtain the first feature; add the first feature and the two-dimensional feature, and input the addition result and the fixed graph into the second graph convolution layer to obtain the second feature; add the second feature and the first feature, and input the addition result and the fixed graph into the third graph convolution layer to obtain the third feature;
[0081] S53. Add the third feature to the second feature to obtain a graph feature.
[0082] Specifically, the first graph convolution layer includes a graph operation layer, a batch normalization layer, a temporal two-dimensional convolution layer and a batch normalization layer that are cascaded in sequence; the structures of the second graph convolution layer and the third graph convolution layer are the same as those of the first graph convolution layer.
[0083] Specifically, the specific processing process of the first graph convolution layer includes:
[0084] S511. Input the two-dimensional features into the graph operation layer to obtain the graph features in its space;
[0085] Specifically, the graph operation layer consists of a fixed graph and a trainable graph, expressed as:
[0086]
[0087] Among them, f graph represents the output of the graph operation layer, f conv represents the output of the previous layer, represents the training graph of the jth layer in the graph operation layer, It represents the diagonal transformation fixed graph of the jth layer in the graph operation layer. The diagonal transformation fixed graph is the result of the diagonal transformation of the input fixed graph A. Its calculation formula is:
[0088]
[0089] Among them, A j represents the jth layer of the input fixed graph A, ∧ j Yes A jThe matrix of eigenvectors.
[0090] S512. performing batch normalization processing on the graph features of the two-dimensional features obtained in S511 to ensure that the range of the feature values does not change;
[0091] S513. Input the batch normalized features into the temporal two-dimensional convolutional layer to obtain useful temporal features;
[0092] S514. Perform batch normalization on the time features obtained in S513.
[0093] Specifically, the expression formula for the standardization processing of the batch normalization layer is:
[0094]
[0095] Among them, x represents the input data of the batch normalization layer, μ represents the mean of the input data x, and σ 2 represents the variance of the input data x, ∈ is the correction factor, γ represents the scale factor, β represents the translation factor, and γ and β are learned by themselves when training the network.
[0096] In one embodiment, the structure of the graph attention network is as follows Figure 5 As shown, step S5 inputs the graph features and scene features into the graph attention network to integrate and obtain a spatiotemporal graph, including:
[0097] S61. Input the graph features and scene features into a batch normalization layer for normalization to obtain normalized features;
[0098] S62. Input the normalized features into the graph attention network head for convolution operation, and regularize the convolution operation result through the Dropout layer to avoid overfitting;
[0099] S63. Input the output result of the Dropout layer into the graph attention network head for convolution operation, and finally obtain the spatiotemporal graph.
[0100] In one embodiment, step S6 inputs the spatiotemporal graph into the transformer network to perform the following operations:
[0101] S71. Input the spatiotemporal graph into the transformer encoder and output the high-dimensional feature representation of the historical trajectory; the structure of the transformer encoder is as follows Figure 6 As shown, it includes multi-head attention mechanism, feedforward neural network and layer normalization;
[0102] S72. Input the high-dimensional feature representation of the historical trajectory into the transformer decoder, and output the predicted trajectory and category of the traffic participants around the autonomous driving vehicle.
[0103] The present invention adds conventional historical trajectory data to enable the graph convolutional network to obtain the interactive features of participants around the autonomous driving vehicle. The combination of the graph convolutional network and the spatiotemporal pyramid network greatly enriches the feature information, greatly enhancing the prediction accuracy of the model and the interpretability of the input data. Although the lidar data collected by the sensor is used as the input of the autonomous driving system and outputs the control signal, this method does not require a detailed one-to-one mapping process, which is conducive to the rapid response of the autonomous driving system, but lacks explanatory power and verifiable robustness.
[0104] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An automatic trajectory prediction method based on graph spatiotemporal pyramid, characterized in that: Construct an automatic trajectory prediction model based on graph spatiotemporal pyramid, which includes spatiotemporal pyramid network, graph convolution network, graph attention network and transformer network; The method for automatic trajectory prediction using the model includes the following steps: S1. Preprocess the radar point cloud data collected in real time by the sensors on the autonomous driving vehicle to obtain a bird's-eye view; S2. Process the bird's-eye view image through the spatiotemporal pyramid network and output scene features; S3. Screening and processing the conventional historical trajectory data collected by the sensors on the autonomous driving vehicle to obtain the original data; S4. Further process the original data to obtain input representation and fixed graph, and input the input representation and fixed graph into the graph convolution network to obtain graph features; Step S4 simultaneously inputs the input representation and the fixed graph into the graph convolutional network to obtain graph features, including: S51. The input representation is normalized by a batch normalization layer, and the normalized input representation is passed through a two-dimensional convolution layer to obtain a two-dimensional feature; S52. Input the two-dimensional feature and the fixed graph into the first graph convolution layer to obtain the first feature; add the first feature and the two-dimensional feature, and input the addition result and the fixed graph into the second graph convolution layer to obtain the second feature; add the second feature and the first feature, and input the addition result and the fixed graph into the third graph convolution layer to obtain the third feature; S53. Adding the third feature to the second feature to obtain a graph feature; The first graph convolution layer includes a graph operation layer, a batch normalization layer, a temporal two-dimensional convolution layer, and a batch normalization layer that are cascaded in sequence; the structures of the second graph convolution layer and the third graph convolution layer are the same as those of the first graph convolution layer; The expression of each graph operation layer is: Among them, f graph represents the output of the graph operation layer, f conv represents the output of the previous layer, represents the training graph of the jth layer in the graph operation layer, Represents the diagonal transformation fixed graph of the jth layer in the graph operation layer, and its calculation formula is: Among them, A j represents the jth layer of the input fixed graph A, ∧ j Yes A j The matrix of eigenvectors; S5. Input the graph features and scene features into the graph attention network to integrate them to obtain the spatiotemporal graph; S6. Input the spatiotemporal graph into the transformer network and output the predicted trajectories and categories of traffic participants around the autonomous driving vehicle.
2. The automatic trajectory prediction method based on graph spatiotemporal pyramid according to claim 1, characterized in that: Step S2 processes the bird's-eye view output scene features through the spatiotemporal pyramid network, including: S21. Input the bird's-eye view image into the first spatiotemporal convolution block, and the output of the first spatiotemporal convolution block is subjected to time pooling to obtain the first spatiotemporal feature; input the first spatiotemporal feature into the second spatiotemporal convolution block, and the output of the second spatiotemporal convolution block is subjected to time pooling to obtain the second spatiotemporal feature; input the second spatiotemporal feature into the third spatiotemporal convolution block, and the output of the third spatiotemporal convolution block is subjected to time pooling to obtain the third spatiotemporal feature; input the third spatiotemporal feature into the fourth spatiotemporal convolution block, and the output of the fourth spatiotemporal convolution block is subjected to time pooling to obtain the fourth spatiotemporal feature; S22. The fourth spatiotemporal feature and the third spatiotemporal feature are subjected to time pooling respectively and then feature fused to obtain a first fused feature; the second spatiotemporal feature is subjected to time pooling and then feature fused with the first fused feature to obtain a second fused feature; the first spatiotemporal feature is subjected to time pooling and then feature fused with the second fused feature to obtain a third fused feature; the bird's-eye view is subjected to time pooling and then feature fused with the third fused feature to obtain the final scene feature.
3. The automatic trajectory prediction method based on graph spatiotemporal pyramid according to claim 1, characterized in that: The process of obtaining the original data in step S3 is as follows: S31. Obtaining conventional historical trajectory data collected by sensors on the autonomous driving vehicle and removing abnormal data therein; S32. Parse the data format of the conventional historical trajectory data after the elimination process and convert it into a data frame; S33. Perform data cleaning on the data frame obtained in S32, including removing unnecessary columns and rows, filling missing values and outliers, and converting data types; S34. Extracting relevant information of surrounding vehicles or pedestrians from the data frame after the data cleaning in S33; S35. Convert the relevant information extracted by S34 into a processing format for the machine learning model to obtain the original data.
4. The automatic trajectory prediction method based on graph spatiotemporal pyramid according to any one of claims 1 or 3, characterized in that: The raw data is further processed to obtain input representation and fixed graph, including: S41 extracts the location information of each surrounding vehicle from the raw data; S42. Convert the position information of each surrounding vehicle into the offset of the current target vehicle; S43. Arrange the offsets of all surrounding vehicles obtained in S42 in chronological order to obtain a trajectory; S44. Divide the trajectory into multiple time windows, each time window containing a number of trajectory points; S45. Integrate the trajectory points of each time window, and merge all the time windows after the integrated trajectory points to form an input representation; S46. Map all trajectory points in each time window to a fixed graph structure through a specific function. The nodes in the graph structure represent surrounding vehicles, and the edges represent the relationships between the surrounding vehicles, and finally a fixed graph is obtained.
5. The automatic trajectory prediction method based on graph spatiotemporal pyramid according to claim 1, characterized in that: In step S1, the radar point cloud data is quantized into regular voxels to form a three-dimensional voxel grid, a binary state is used to represent the occupancy of each voxel grid, and the height dimension of the three-dimensional voxel grid corresponds to the image channel of the two-dimensional pseudo image, thereby converting the three-dimensional radar point cloud data into a two-dimensional pseudo image, that is, the required bird's-eye view.
Citation Information
Patent Citations
Automatic driving vehicle track prediction method and device and electronic equipment
CN113705636A
Vehicle track prediction method based on graph attention interaction mechanism
CN114692762A
Automatic driving track prediction method based on space-time pyramid
CN115049130A