Visual end-to-end trajectory prediction system fusing historical prediction trajectory features

By integrating historical predicted trajectory features into the autonomous driving system, using technologies such as Transformer and convolutional neural networks, the problem of insufficient utilization of visual semantic information in existing systems is solved, and higher trajectory prediction and object detection accuracy is achieved, and the accuracy and robustness of the system are improved.

CN120047909APending Publication Date: 2025-05-27SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510013198.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Trajectory prediction tasks in existing autonomous driving systems are difficult to effectively utilize visual semantic information, resulting in insufficient prediction accuracy and difficult training, making it difficult to achieve high accuracy and robustness.

Method used

A visual end-to-end trajectory prediction system that integrates historical prediction trajectory features is proposed. Through the BEV feature encoding module, the timing fusion module, the trajectory prediction feature encoding module, the trajectory prediction decoding module and the image feature position encoding module, the Transformer and convolutional neural network are used to improve the spatial characterization of 2D image features and generate richer BEV features, ultimately improving the trajectory prediction and object detection accuracy.

Benefits of technology

By introducing historical prediction trajectory features, the system can more accurately predict the trajectory of the target in a certain time in the future, significantly improving the accuracy and robustness of the autonomous driving perception system, so that the accuracy of target detection and trajectory prediction is greatly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047909A_ABST
    Figure CN120047909A_ABST
Patent Text Reader

Abstract

A visual end-to-end trajectory prediction system fusing historical prediction trajectory features comprises a BEV feature coding module, a time sequence fusion module, a trajectory prediction feature coding module, a trajectory prediction decoding module and an image feature position coding module.The historical prediction trajectory is used for coding 2D image features, the spatial representation of the 2D image features is improved, and the prediction precision of the 2D image features is improved. Therefore, richer BEV features are obtained, the track prediction and target detection precision is finally improved, and a sensing system is more accurate and robust.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of autonomous driving perception, specifically a visual end-to-end trajectory prediction system that fuses historical prediction trajectory features. Background Art

[0002] The perception system in an autonomous driving system is a key module, mainly including two independently studied modules: object detection and trajectory prediction. Combining these two modules end-to-end can avoid information transmission loss. For example, dense bird's-eye view (BEV) features contain rich scene information during the detection process and can be accurately and efficiently transmitted to the trajectory prediction branch. The existing trajectory prediction tasks for autonomous driving mainly use the historical target tracking trajectory as input and estimate the future state of the target in combination with a high-precision map. However, this method will lose the visual semantic information in the object detection process and it is difficult to further improve the prediction accuracy. Summary of the Invention

[0003] Aiming at the deficiencies of the existing technology, such as poor interpretability, lack of corresponding closed-loop training data for incorrect inference results, and high training difficulty, the present invention proposes a visual end-to-end trajectory prediction system that fuses historical prediction trajectory features. It encodes 2D image features using historical prediction trajectories to enhance the spatial representation of 2D image features, and then obtains BEV features with richer features, ultimately improving the trajectory prediction and object detection accuracy, and making the perception system more accurate and robust.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a visual end-to-end trajectory prediction system that fuses historical prediction trajectory features, including: a BEV feature encoding module, a temporal fusion module, a trajectory prediction feature encoding module, a trajectory prediction decoding module, and an image feature position encoding module. Among them: the BEV feature encoding module uses the attention mechanism of the Transformer to convert the 2D features extracted from the visual image into BEV features; the temporal fusion module selects time intervals in a variable step size manner and uses a residual-like neural network to fuse BEV features at different time sequences; the trajectory prediction feature encoding module extracts trajectory prediction features with a larger receptive field from the fused features using a feature pyramid network; the trajectory prediction decoding module uses a convolutional neural network with an attention mechanism to convert the output result dimension to the same as the trajectory prediction result dimension according to the trajectory prediction features, and finally obtains the predicted trajectory result of the target; the image feature position encoding module performs coordinate transformation processing according to the historical prediction trajectory of the previous moment to achieve spatio-temporal consistency, combines the three-dimensional space grid coordinates and the trajectory prediction distribution, and uses a convolutional neural network to obtain the 2D position encoding result of the image features. Technical Effects

[0006] The present invention uses the prediction result of the historical trajectory containing 3D grids as the 4D expression to participate in the Transformer attention mechanism process of BEV feature encoding, and enhances the spatial expression of the image in the form of position encoding. It focuses on the consistency of time and space in detection and prediction. The starting point of trajectory prediction needs to be aligned with the target detection result in time and space. Compared with the prior art, the present invention avoids the ambiguity problem of inconsistent target spaces at the same time when the downstream of the perception module in autonomous driving, such as the planning and control module, uses perception information, and effectively obtains high-precision target detection results for the current frame data and accurately predicts the trajectory of the target within a certain period of time in the future. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1 It is a schematic diagram of the system of the present invention;

[0008] Figure 2 It is a flowchart of an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0009] As Figure 1 shown, this embodiment relates to a vision end-to-end trajectory prediction system integrating historical prediction trajectory features, including: a BEV feature encoding module, a temporal fusion module, a trajectory prediction feature encoding module, a trajectory prediction decoding module, and an image feature position encoding module.

[0010] The BEV feature encoding module adds the position encoding PE key of the image features from the image feature position encoding module and F img bitwise, and encodes them into the features at position p=(x, y) in the BEV space through a Transformer. It includes: a picture feature extraction unit and a feature conversion unit, where: the picture feature extraction unit processes the original picture information using a deep neural network such as ResNet-50 to obtain the 2D feature result of the picture, and the feature conversion unit obtains the BEV feature using the cross-attention mechanism according to the 2D feature result of the picture.

[0011] The temporal fusion module stacks the BEV features of multiple frames at different times along the channel dimension through a residual network, then extracts the fusion features using a multi-layer convolutional neural network and compresses the features using a 1×1 neural network respectively, and finally adds the two features bitwise to obtain the final fusion feature. This module includes: a variable-step temporal feature screening unit and a residual neural network fusion unit, where: the variable-step temporal feature screening unit selects the historical BEV features at a random step distance from the current frame for the historical features, and the residual neural network unit extracts the fusion features using the residual network according to the BEV feature at the current moment and the selected historical BEV features, and finally uses a 1×1 convolutional neural network to shrink the feature dimension to 256.

[0012] The so-called temporal fusion specifically refers to: using the attention mechanism to dynamically weight the detected features and predicted features at the target level obtained by extraction to obtain fused target features, which are used to predict the unique starting point (the detection result of the current frame) of the trajectory and the multi-modal predicted end point, and complete the accurate and robust multi-modal trajectory prediction task, specifically including: A s = sigmoid(ξ(A det )) * A pred , A fusion = A s + A pred , where: ξ is a multi-layer perceptron, and A det , A pred , A fusion are the detected, predicted, and fused features at the target level respectively. A det is used as the input, and after passing through a small multi-layer perceptron network ξ and sigmoid calculation, the attention weight is obtained. A pred After multiplying by the attention weight in the form of element-wise multiplication, the fused feature A fysion is obtained through a class residual network, and then the target detection result is optimized using the fused feature. For A fusion a multi-layer perceptron is used to generate dynamic weights to optimize the detection.

[0013] The so-called dynamic weighting specifically refers to: calculating the dynamically learnable weight matrix W 1 = linear(ξ(A fusion ))), where: A fusion is used as the input source after passing through a multi-layer perceptron, and the dynamically learnable weight matrix W 1 is generated using the fully connected layer in the neural network, and finally the final detection refine result is obtained using matrix multiplication.

[0014] During the training stage, the dynamic weighting only regresses the central position information x, y in the BEV plane of the target and the yaw angle of the target.

[0015] The described trajectory prediction feature encoding module, based on the feature pyramid module in the BEV space, extracts deep BEV features using 2D convolution, then restores the dimension of the feature tensor to the same tensor dimension as the BEV feature using transposed convolution, and then concatenates the pyramid BEV features along the channel dimension. This module includes: a feature pyramid unit, a multi-dimensional feature fusion unit, and a trajectory prediction feature extraction unit, where: the feature pyramid unit performs convolution processing on the BEV feature information after temporal fusion to obtain BEV feature results with different receptive fields; the multi-dimensional feature fusion unit performs stacking processing on the BEV feature information of different dimensions to obtain multi-dimensional fusion feature results; the trajectory prediction feature extraction unit uses the object detection result as an index based on the fused BEV feature to obtain the trajectory prediction feature. The described trajectory prediction decoding module transfers the historical trajectory to the current coordinate system, elevates it to the 3D real space using the height of the object detection result, takes the historical object state estimation as a prior, and selects the prediction trajectory with the highest score in the multi-modal trajectory prediction to obtain the trajectory prediction trajectory at T frames in the global coordinate system. Where: represents the pose transformation matrix of the vehicle itself and the world coordinate system at time T. represents the pose transformation matrix of the vehicle itself and the lidar coordinate system at the same time T. represents the i-th trajectory prediction result in the world coordinate system. represents the i-th trajectory prediction result in the lidar coordinate system at time T.

[0016] The described image feature position encoding module generates the position encoding of the image feature through a convolutional neural network. Where: HF and WF are the height and width of the image feature map respectively, D is the depth discretization amount in the camera coordinate system, N is the number of multi-views, dimension 1 is the probability of the target existing in the corresponding space. This module includes: a four-dimensional space feature generation unit and a feature encoding unit, where: the four-dimensional space feature generation unit performs aggregation processing based on the three-dimensional coordinate system grid information and the trajectory prediction result to obtain the four-dimensional space feature; the feature encoding unit performs convolution processing based on the four-dimensional space feature information to obtain the image feature position encoding result.

[0017] After specific actual experiments, in the Nuscenes dataset, as Figure 2 shown, with a batch size of 4 and after training for 24 epochs, the test results are NDS: 0.420, mAP: 0.309.

[0018] Compared with the prior art, the present system uses the trajectory prediction result to generate position encoding for the image, thereby introducing prior information on the possible positions of the target into the current frame BEV feature, achieving a maximum increase of 7.3% in mAP and a maximum increase of 8.1% in NDS.

[0019] The above specific implementation can be locally adjusted by those skilled in the art in different ways without departing from the principles and purposes of the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific implementation. All implementation solutions within its scope are subject to the present invention.

Claims

1. A visual end-to-end trajectory prediction system integrating historical prediction trajectory features, characterized in that: include: BEV feature encoding module, time series fusion module, trajectory prediction feature encoding module, trajectory prediction decoding module and image feature position encoding module, wherein: the BEV feature encoding module uses the attention mechanism of Transformer to convert the 2D features extracted from the visual image into BEV features, the time series fusion module selects the time interval by variable step size, and uses the residual neural network to fuse the BEV features under different time series, the trajectory prediction feature encoding module uses the feature pyramid network to extract trajectory prediction features with a larger receptive field from the fused features, the trajectory prediction decoding module uses a convolutional neural network with an attention mechanism to convert the output result dimension to the same dimension as the trajectory prediction result according to the trajectory prediction feature, and finally obtains the predicted trajectory result of the target, the image feature position encoding module performs coordinate transformation processing according to the historical predicted trajectory of the previous moment to achieve time-space consistency, and uses a convolutional neural network to obtain the 2D position encoding result of the image feature by combining the three-dimensional space grid coordinates and the trajectory prediction distribution; The temporal fusion specifically refers to: using the attention mechanism to dynamically weight the extracted target-level detection features and prediction features to obtain fused target features, which are used to predict the unique starting point of the trajectory (the current frame detection result) and the multi-modal prediction end point, and complete the accurate and robust multi-modal trajectory prediction task, specifically including: A s =sigmoid(ξ(A det ))*A pred , A fusion =A s +A pred , where: ξ is a multi-layer perceptron, A det ,A pred ,A fusion They are respectively the detection, prediction and fusion features of the target level, A det As input, it passes through a small multi-layer perceptron network ξ and sigmoid calculation to obtain the attention weight, A pred After multiplying the attention weight in the form of element multiplication, the fusion feature A is obtained through a residual network fusion , and then use the fusion features to optimize the target detection results. fusion Multilayer perceptron is used to generate dynamic weights to optimize detection.

2. The visual end-to-end trajectory prediction system integrating historical prediction trajectory features according to claim 1 is characterized in that: The dynamic weighting specifically refers to: calculating the dynamic learnable weight matrix W1=linear(ξ(A fusion )),in: A fusion After passing through a multi-layer perceptron as the input source, the fully connected layer in the neural network is used to generate a dynamic learnable weight matrix W1, and finally matrix multiplication is used to obtain the final detection refinement result.

3. The visual end-to-end trajectory prediction system integrating historical prediction trajectory features according to claim 1 is characterized in that: The BEV feature encoding module encodes the position of the image feature from the image feature position encoding module PE key With F img After bitwise addition, the feature at position p = (x, y) in BEV space is encoded by Transformer. The module includes: a picture feature extraction unit and a feature conversion unit, wherein: the picture feature extraction unit uses a deep neural network to process the original picture information to obtain a 2D feature result of the picture, and the feature conversion unit uses a cross-attention mechanism to obtain a BEV feature based on the 2D feature result of the picture.

4. The visual end-to-end trajectory prediction system integrating historical prediction trajectory features according to claim 1 is characterized in that: The time series fusion module includes: a variable step-size time series feature screening unit and a residual neural network fusion unit, wherein: the variable step-size time series feature screening unit selects historical BEV features with a random step length from the current frame for historical features, and the residual neural network unit uses the residual network to extract fusion features based on the BEV features at the current moment and the selected historical BEV features, and finally uses a 1×1 convolutional neural network to shrink the feature dimension to 256.

5. The visual end-to-end trajectory prediction system integrating historical prediction trajectory features according to claim 1 is characterized in that: The trajectory prediction feature encoding module uses 2D convolution to extract deep BEV features according to the feature pyramid module in the BEV space, and then uses transposed convolution to restore the feature tensor dimension to the same tensor dimension as the BEV feature, and then connects the pyramid BEV features along the channel dimension. The module includes: a feature pyramid unit, a multi-dimensional feature fusion unit and a trajectory prediction feature extraction unit, wherein: the feature pyramid unit performs convolution processing according to the BEV feature information after time series fusion to obtain BEV feature results with different receptive fields, the multi-dimensional feature fusion unit performs stacking processing on the feature dimension according to the BEV feature information of different dimensions to obtain a multi-dimensional fusion feature result, and the trajectory prediction feature extraction unit uses the target detection result as an index to obtain the trajectory prediction feature according to the fused BEV feature.

6. The visual end-to-end trajectory prediction system integrating historical prediction trajectory features according to claim 1 is characterized in that: The trajectory prediction decoding module transfers the historical trajectory to the current coordinate system, uses the height of the target detection result to lift it to the 3D actual space, takes the historical target state estimation as a priori, selects the predicted trajectory with the highest score in the multimodal trajectory prediction, and obtains the trajectory prediction trajectory of frame T in the global coordinate system. in: Represents the pose transformation matrix between the vehicle and the world coordinate system at time T, Represents the pose transformation matrix between the vehicle at time T and the laser radar coordinate system at the same time, Represents the i-th trajectory prediction result in the world coordinate system, Represents the i-th trajectory prediction result in the lidar coordinate system at time T.

7. The visual end-to-end trajectory prediction system integrating historical prediction trajectory features according to claim 1 is characterized in that: The image feature position encoding module generates the position encoding of the image feature through a convolutional neural network in: H F , W F are the height and width of the image feature map, D is the depth discrete quantity in the camera coordinate system, N is the number of multiple perspectives, and dimension 1 is the probability of the target existing in the corresponding space. The module includes: a four-dimensional space feature generation unit and a feature encoding unit, wherein: the four-dimensional space feature generation unit is based on the three-dimensional coordinate system grid information And the trajectory prediction results are aggregated to obtain four-dimensional space features. The feature encoding unit is based on the feature information of the thinking space. Perform convolution processing to obtain the image feature position encoding result.

Citation Information

Cited By

  • Target distance and speed measuring method based on monocular vision

    CN120926942A