Composite scene obstacle non-graphical trajectory prediction method

By integrating 4D mmWave radar and surround-view camera data for obstacle trajectory prediction, the problem of obstacle trajectory prediction in areas without high-precision maps in autonomous driving is solved, and efficient and low-latency prediction is achieved in composite scenarios, suitable for high-speed and urban scenarios.

CN120440068APending Publication Date: 2025-08-08ANHUI JIANGHUAI AUTOMOBILE GRP CORP LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510606856.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, it is difficult to predict obstacle trajectory in complex traffic scenarios in areas without high-precision map coverage in autonomous driving, and as the number of obstacles increases, it is impossible to cover both high-speed and urban scenarios at the same time.

Method used

By combining 4D mmWave radar and surround-view camera data, multimodal spatial features and obstacle trajectory characteristics are extracted, and obstacle trajectory prediction is used using a multi-layer perceptron network and self-attention mechanism to get rid of the dependence on high-precision maps and are suitable for composite scenarios.

Benefits of technology

It realizes obstacle trajectory prediction in complex traffic scenarios without high-precision maps, reduces prediction delay, is suitable for high-speed and urban scenarios, and expands the popularization of autonomous driving systems and scene adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120440068A_ABST
    Figure CN120440068A_ABST
Patent Text Reader

Abstract

The invention discloses a composite scene obstacle non-graphical trajectory prediction method, which is used for predicting future trajectories of obstacles around a vehicle by taking a composite scene as a center based on fusion of data of a 4D millimeter wave radar and an all-round camera. Acquiring 4D millimeter wave radar data of urban areas and highways, surround view camera data of surroundings and historical track data of surrounding obstacles during vehicle driving; and multi-modal spatial features and obstacle trajectory features are extracted, fusion interaction is carried out to obtain final interaction features, and then a trained obstacle trajectory prediction model is used to output a future trajectory of an obstacle. The method gets rid of dependence on a high-precision map, makes full use of the advantage of multi-mode perception, achieves obstacle trajectory prediction in a complex traffic scene under the condition of no graph, and does not increase the prediction time delay along with the increase of the number of surrounding obstacles. Therefore, a reliable implementation path is provided for popularization of an automatic driving system and expansion of scene adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method for predicting complex scene obstacle trajectories without graphics. Background Art

[0002] The primary function of the prediction module in an autonomous driving system is to predict the future state of surrounding obstacles (including vehicles, pedestrians, cyclists, etc.). Predicting the future trajectory of obstacles and pedestrians around highways and urban areas in real time is crucial for improving the safety and comfort of autonomous driving.

[0003] Traditional trajectory prediction methods primarily rely on physical and statistical models. Physical models predict a vehicle's future position by modeling its vehicle dynamics. However, this approach requires a large number of vehicle dynamics parameters and struggles with complex road environments and traffic conditions. Statistical models analyze historical data to build statistical models of vehicle and pedestrian behavior and predict future trajectories. However, this approach struggles with real-time data and unexpected situations.

[0004] In recent years, with the rapid development of artificial intelligence technology, trajectory prediction methods based on deep learning have been widely used. Deep learning has powerful feature extraction and pattern recognition capabilities and can handle complex data and real-time changes. Trajectory prediction methods based on deep learning mainly include the following:

[0005] Trajectory prediction based on convolutional neural networks (CNNs), recurrent neural networks (RNNs), attention mechanisms, and multimodal fusion can improve prediction accuracy by combining data from different modalities. This approach is suitable for processing data from multiple sensors, such as lidar and cameras. Trajectory prediction also requires high-precision map information or a map information acquisition module, which is very expensive.

[0006] The above can be further expanded upon: With the rapid development of artificial intelligence technology and its continued application in various fields, obstacle trajectory prediction methods based on deep learning are also gaining widespread application in autonomous driving systems. Deep learning has powerful feature extraction and pattern recognition capabilities, and is capable of extracting both deep and shallow information, as well as interactive information, between various features. Deep learning-based trajectory prediction methods primarily include the following:

[0007] ● Sequence modeling: Use models such as RNN, LSTM or Transformer to process time series data and capture the dynamic patterns of driving behavior.

[0008] ● Graph Neural Network (GNN): Models road scenes as graph structures, combines the interaction between vehicles and roads, and improves the modeling capabilities of complex traffic environments.

[0009] ● Multimodal data fusion: Combine environmental perception data (such as cameras, lidar) and map information to improve the perception and prediction accuracy of the surrounding dynamic environment.

[0010] Currently, the main algorithms used for autonomous driving trajectory prediction are obstacle-centric algorithms, such as MLP, VectorNet, and LSTM. These algorithms primarily transform surrounding environmental information into the corresponding obstacle coordinate system, then extract features from the surrounding environment and trajectory, and finally predict the future trajectory. However, existing technical solutions are obstacle-centric and can only predict the trajectory of a single obstacle. Furthermore, as the amount of obstacle data increases, the prediction time increases, significantly increasing latency. In autonomous driving scenarios, especially in urban areas, latency requirements are particularly high.

[0011] Methods based on deep learning for predicting the trajectory of surrounding obstacles generally rely heavily on high-precision map information or map information acquisition modules, and fail in areas without high-precision map coverage. Solutions that rely heavily on high-precision maps are costly and cannot predict obstacle trajectories in areas without them. Furthermore, as the amount of obstacle data increases, prediction time increases, significantly increasing latency. These models are generally only suitable for highway scenarios and cannot cover both highway and urban areas. They also require high-precision map coverage, without which trajectory prediction is impossible.

[0012] Furthermore, deep learning-based methods for predicting surrounding obstacle trajectories are typically categorized by scenario, such as highways, urban areas, or cut-in and cut-out scenarios. The more scenarios there are, the more models need to be trained. This means that as the number of scenarios increases, the number of model types also increases. Summary of the Invention

[0013] In view of the above, the present invention aims to provide a method for non-graphic trajectory prediction of obstacles in complex scenarios to solve the above-mentioned technical problems.

[0014] The technical solution adopted in the present invention is as follows:

[0015] The present invention provides a method for predicting the trajectory of obstacles in a complex scene without graphics, which includes:

[0016] While the vehicle is driving, it acquires 4D millimeter-wave radar data in urban areas and highways, surround-view camera data of the surrounding environment, and historical trajectory data of surrounding obstacles;

[0017] Extract multimodal spatial features and obstacle trajectory features from the acquired data, and perform fusion interaction to obtain the final interaction features;

[0018] Based on the final interaction features, the future trajectory of the obstacle is output using the trained obstacle trajectory prediction model.

[0019] In at least one possible implementation, obtaining the final interaction feature includes:

[0020] After voxelizing the 4D millimeter-wave radar data, extracting spatial structural information and projecting it into a bird's-eye view space; and extracting visual information from the surround-view camera data and projecting it into a bird's-eye view space.

[0021] After the multimodal bird's-eye view spatial features are fused, they are input into a multi-layer perceptron network, and by modeling the spatiotemporal dependency of obstacle trajectories, fused features are obtained and output. The fused features have a dynamic representation of interactive semantics.

[0022] In at least one possible implementation manner, obtaining the fusion feature includes extracting interaction features between the environment and the obstacle, between the environment and the environment, between the environment and the obstacle, and between obstacles and obstacles.

[0023] In at least one possible implementation, outputting the future trajectory of the obstacle includes: capturing the long-range dependency between environmental elements and obstacles from the final interaction features through a self-attention mechanism, then combining multi-head cross-attention to perform multimodal feature alignment and spatiotemporal information fusion operations, and ultimately outputting the multimodal trajectory parameters of the future obstacle with probability distribution characteristics.

[0024] In at least one possible implementation, outputting the future trajectory of the obstacle includes:

[0025] Based on the final interaction feature, obtaining the initial trajectory endpoint;

[0026] Correct the initial trajectory endpoint to obtain the corrected trajectory endpoint;

[0027] Combining the final interaction feature with the corrected trajectory endpoint to obtain an initial predicted trajectory;

[0028] The initial predicted trajectory is spliced with the corrected trajectory endpoints to obtain the complete predicted trajectory.

[0029] In at least one possible implementation, the training process of the obstacle trajectory prediction model includes:

[0030] Training data for highway scenarios and urban scenarios are configured according to a first predetermined ratio, and noisy driving data of a second predetermined ratio is introduced during the training process, and the labels of the noisy driving data are corrected to driving along the center line of the lane.

[0031] Compared with the existing technology, the main design concept of the present invention is to perform multimodal prediction of the future trajectory of obstacles around the vehicle based on the fusion of 4D millimeter-wave radar and surround-view camera data, with complex scenes such as highways and urban areas as the center. Specifically, 4D millimeter-wave radar data of urban areas and highways, surround-view camera data of the surrounding environment, and historical trajectory data of surrounding obstacles are obtained while the vehicle is driving; multimodal spatial features and obstacle trajectory features are extracted from them, and fused and interacted to obtain the final interactive features, and then the trained obstacle trajectory prediction model is used to output the future trajectory of the obstacle. The present invention gets rid of the dependence on high-precision maps, fully utilizes the advantages of multimodal perception, realizes obstacle trajectory prediction in complex traffic scenes under map-free conditions, and does not increase the prediction delay as the number of surrounding obstacles increases, thereby providing a reliable implementation path for the popularization of autonomous driving systems and the expansion of scene adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:

[0033] Figure 1 A schematic diagram of a method for predicting complex scene obstacle trajectories without graphics provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.

[0035] The present invention proposes an embodiment of a method for predicting the trajectory of obstacles in a complex scene without graphics. Specifically, Figure 1 shown, including:

[0036] Step S1: While the vehicle is driving, obtain 4D millimeter-wave radar data of urban areas and highways, surround-view camera data of the surrounding environment, and historical trajectory data of surrounding obstacles;

[0037] Acquire 4D millimeter-wave radar data in urban areas and highways, and surrounding surround-view camera data to replace high-precision maps while enabling cross-scene trajectory prediction and corresponding data preprocessing of the trajectory of surrounding obstacles.

[0038] Specifically, (1) 4D millimeter-wave radar data point cloud acquisition: 11 frames of millimeter-wave radar point cloud data with a time interval of 500ms can be selected for data cropping to extract important features. The preprocessing of each frame of point cloud data is as follows:

[0039]

[0040] in, The current frame The position coordinates, radial velocity, compensated point cloud velocity, characteristic Power value and signal-to-noise ratio of the point cloud are the original point cloud. , respectively, retain the point cloud Range maximum and minimum values.

[0041] (2) Surround view camera image data acquisition: The data acquired by the surrounding surround view cameras can be selected and the data acquired by multiple cameras can be spliced at the channel level. For example, the data dimension acquired by each camera is (C, W, L). There are n cameras in total. The total camera input data finally acquired by data channel-level splicing is (n×C, W, L).

[0042] (3) Acquisition of motion trajectory data of surrounding obstacles: Considering the complexity and uncertainty of the future trajectory of surrounding obstacles, the present invention proposes that the historical motion trajectory of surrounding obstacles should be used as the input of the model to further improve the prediction accuracy of the multimodality of future trajectories. After actual verification and statistics, it is found that the best trajectory prediction is to input 13 frames of the historical trajectory of the obstacle, where each frame of data includes (pos_x, pos_y, vel_x, vel_y).

[0043] Step S2: extracting multimodal spatial features and obstacle trajectory features from the acquired data, and fusing and interacting them to obtain the final interactive features;

[0044] First, the 4D millimeter-wave radar point cloud data is voxelized (i.e., the three-dimensional space is discretized into a uniform small cube grid), and the spatial structure information is extracted through voxel feature encoding and projected into the bird's-eye view (BEV) space. At the same time, the surround-view camera images are preprocessed by distortion correction and multi-view stitching, and a convolutional neural network is used to extract visual features and project them into the bird's-eye view (BEV) space. Then, the BEV features of the above two modalities are fused and input into a multi-layer perceptron (MLP) network. By modeling the spatiotemporal dependencies of obstacle trajectories, the interaction features between obstacles (such as collision risk and motion coupling relationship) are extracted. Finally, a dynamic representation of the surrounding environment with interactive semantics is output to provide support for behavior prediction and decision planning.

[0045] (1) 4D point cloud data feature extraction: The selected point cloud is voxelized and converted into a voxel grid according to the set voxel size and detection distance, where the height direction grid size is constant at 1. The voxelized grid features are sent to the feature extraction module for feature extraction. The feature extraction module includes ResNet18 and an upsampling module. First, the grid features are downsampled through the ResNet18 network, and the features are downsampled by 2 times, 4 times, and 8 times respectively to obtain the downsampled features. The formula is as follows:

[0046]

[0047] The characteristics after downsampling by 1, 2, 4, and 8 times, is the millimeter wave point cloud feature after voxelization, It is a convolutional neural network model.

[0048] Then After sampling, After channel splicing, convolution is used to keep the feature channels consistent, and the first upsampling result is obtained. :

[0049]

[0050] Among them, upasmple is the upsampling operation, cat is the data splicing operation, and conv is the convolution operation, including convolution, normalization and activation function.

[0051] Will Upsampling and After splicing along the channel, convolution is used to keep the feature channels consistent, and the second upsampling result is obtained. :

[0052]

[0053] Will Upsampling and After splicing along the channel, convolution is used to keep the feature channels consistent, and the second upsampling result is obtained. :

[0054]

[0055] Finally, After upsampling as well as After splicing, the BEV features are extracted through convolution output:

[0056] .

[0057] (2) Feature extraction of surround-view camera image data: Image feature extraction is performed through CNN, and the image data obtained by the surround-view camera is spliced at the channel level to obtain data with a shape of (n×C, L, W), which is then converted into data of (n×C, 224, 224); then a standard 3×3 convolution layer is used for preliminary feature extraction, and then a sequence of inverted residual blocks (composed of multiple inverted residual blocks, each containing an expanded 1×1 convolution to increase the number of channels; depthwise convolution 3×3, stride=1, for spatial feature extraction; linear bottleneck 1×1 convolution, dimensionality reduction; jump links) is used for further feature extraction; then, a convolution layer (1×1 kernel) is used for the extraction of the final feature map, and finally, average pooling and fully connected layers are used to generate (n, 1000) features.

[0058] (3) Obstacle historical trajectory feature extraction: RNN can be used to extract the temporal continuity features of the historical trajectory.

[0059] (4) Multi-dimensional feature interaction: It is preferred to use an interactive feature extractor to extract interactive features between the environment and obstacles, the environment and the environment, the environment and obstacles, and the obstacles and obstacles, and output the final interactive features as input for subsequent steps. The most critical module in the interaction process of each modality and dimension is Attention_Block, such as but not limited to: multi-head-attention to extract the above interactive features, and LA layer normalization, so as to enhance robustness:

[0060] .

[0061] Step S3: Based on the final interaction features, the trained obstacle trajectory prediction model is used to output the future trajectory of the obstacle.

[0062] Continuing from the previous article, in practice, a spatiotemporal interaction modeling module can be constructed based on the Transformer architecture to jointly model the BEV spatial features output by the multimodal feature fusion layer and the interaction features of obstacles. Subsequently, a self-attention mechanism is used to capture the long-range dependencies between environmental factors and obstacles from the final interaction features. Multi-head cross-attention is then combined to achieve multimodal feature alignment and spatiotemporal information fusion. The final output is a multimodal trajectory prediction of the future obstacle (including parameters such as position, velocity, and confidence) with probabilistic distribution characteristics, providing the core input for trajectory deduction and risk quantification of the autonomous driving system.

[0063] In addition, the trajectory prediction module can use Dynamic_Trajectory_Decoder to predict the endpoints of obstacles. If multi-modality is used, multiple trajectory endpoints will be input. The formula is as follows:

[0064]

[0065] Where final_features is the final interaction feature extracted by step S2, and endpoints is the endpoint of the predicted multimodal trajectory.

[0066] Then, endpoint_refiner is used to further modify the endpoint. The formula is as follows:

[0067]

[0068] Among them, meta_info is (v, x, y, pre_x, pre_y), endpoints is the output of Dynamic_Trajectory_Decoder, and refiner_endpoints is the endpoint after further correction. Then use trajectory_decoder to predict the complete trajectory. The formula is as follows:

[0069]

[0070] Among them, the input of trajectory_decoder is the concatenation of the upstream interactive features final_features and refiner_endpoints, and the output is the trajectory except for the last point. The Decoder method first uses the endpoint_predictor adaptive head to predict multiple endpoints.

[0071] Finally, the predicted trajectory predictions_traj and endpoint refiner_endpoints are concatenated to obtain the predicted complete future multimodal trajectory. The formula is as follows:

[0072] .

[0073] Finally, it should be noted that the training process of the aforementioned obstacle trajectory prediction model involves adjusting the urban and highway scene data according to a predetermined ratio. After training with this ratio, an obstacle trajectory prediction model that is suitable for both highway and urban scenes is obtained. Model performance can then be evaluated using metrics such as average displacement error, final displacement error, and future prediction duration.

[0074] The key step is to downsampling each case and adjust the training data ratio between each case. According to statistics, the data ratio between highway scenes and urban scenes is roughly 1:3, and the total data volume is over 500,000. At the same time, noisy data can be added during the training process and the label of this noisy data must be corrected to driving along the centerline of the lane. The approximate proportion of noisy data is 0.1%. Model training based on this configuration of training data can enable the model to make more comprehensive, non-image-based predictions of obstacle trajectories in highway and urban scenes.

[0075] In summary, the main design concept of the present invention is to perform multimodal prediction of the future trajectory of obstacles around the vehicle based on the fusion of 4D millimeter-wave radar and surround-view camera data, with complex scenes such as highways and urban areas as the center. Specifically, 4D millimeter-wave radar data of urban areas and highways, surround-view camera data of the surrounding environment, and historical trajectory data of surrounding obstacles are obtained while the vehicle is driving; multimodal spatial features and obstacle trajectory features are extracted from them, and fused and interacted to obtain the final interactive features, and then the trained obstacle trajectory prediction model is used to output the future trajectory of the obstacle. The present invention gets rid of the dependence on high-precision maps, fully utilizes the advantages of multimodal perception, realizes obstacle trajectory prediction in complex traffic scenes under map-free conditions, and does not increase the prediction delay as the number of surrounding obstacles increases, thereby providing a reliable implementation path for the popularization of autonomous driving systems and the expansion of scene adaptability.

[0076] If the expressions expressing directions are mentioned in the embodiments of the present invention, they are relative concepts based on the embodiments. In addition, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a, b and c, where a, b, c can be single or multiple.

[0077] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings, but the above is only a preferred embodiment of the present invention. It should be noted that the technical features involved in the above embodiments and their preferred modes can be reasonably combined and matched into a variety of equivalent schemes by those skilled in the art without departing from or changing the design ideas and technical effects of the present invention; therefore, the scope of implementation of the present invention is not limited to what is shown in the drawings. Any changes made in accordance with the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the scope of protection of the present invention.

Claims

1. A method for predicting obstacles' trajectories without graphics in complex scenes, characterized by: include: While the vehicle is driving, it acquires 4D millimeter-wave radar data in urban areas and highways, surround-view camera data of the surrounding environment, and historical trajectory data of surrounding obstacles; Extract multimodal spatial features and obstacle trajectory features from the acquired data, and perform fusion interaction to obtain the final interaction features; Based on the final interaction features, the future trajectory of the obstacle is output using the trained obstacle trajectory prediction model.

2. The method for predicting complex scene obstacle trajectories without graphics according to claim 1, characterized in that: Obtaining the final interaction features includes: After voxelizing the 4D millimeter-wave radar data, extracting spatial structural information and projecting it into a bird's-eye view space; and extracting visual information from the surround-view camera data and projecting it into a bird's-eye view space. After the multimodal bird's-eye view spatial features are fused, they are input into a multi-layer perceptron network, and by modeling the spatiotemporal dependency of obstacle trajectories, fused features are obtained and output. The fused features have a dynamic representation of interactive semantics.

3. The method for predicting complex scene obstacle trajectories without graphics according to claim 2, characterized in that: Acquiring fusion features includes extracting interaction features between the environment and obstacles, between the environment and the environment, between the environment and obstacles, and between obstacles and obstacles.

4. The method for predicting complex scene obstacle trajectories without graphics according to claim 1, characterized in that: Outputting the future trajectory of the obstacle includes: capturing the long-range dependency between environmental elements and obstacles from the final interaction features through a self-attention mechanism, and then combining multi-head cross attention to perform multimodal feature alignment and spatiotemporal information fusion operations, and finally outputting the multimodal trajectory parameters of the future obstacle with probability distribution characteristics.

5. The method for predicting complex scene obstacle trajectories without graphics according to claim 1, characterized in that: The output obstacle's future trajectory includes: Based on the final interaction feature, obtaining the initial trajectory endpoint; Correct the initial trajectory endpoint to obtain the corrected trajectory endpoint; Combining the final interaction feature with the corrected trajectory endpoint to obtain an initial predicted trajectory; The initial predicted trajectory is spliced with the corrected trajectory endpoint to obtain the complete predicted trajectory.

6. The method for predicting complex scene obstacle trajectories without graphics according to any one of claims 1 to 5, characterized in that: The training process of the obstacle trajectory prediction model includes: Training data for highway scenarios and urban scenarios are configured according to a first predetermined ratio, and noisy driving data of a second predetermined ratio is introduced during the training process, and the labels of the noisy driving data are corrected to driving along the center line of the lane.

Citation Information

Patent Citations

  • Method for generating future trajectory of obstacle

    CN114399743A

  • Obstacle trajectory prediction method and device, equipment and storage medium

    CN116152782A

  • Data fusion method and device, vehicle and storage medium

    CN116363615A

  • Vehicle trajectory prediction method based on attention mechanism

    CN118953406A

  • Automatic driving model training data set construction method and module and application thereof

    CN119046686A