Online map construction method based on deep learning

Through the encoder-decoder model and dispersed query technology based on Transformers, the problems of accuracy and efficiency in online map construction are solved, and high-precision and efficient online map construction effect are achieved.

CN120259991APending Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410007168.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing online map construction methods are difficult to achieve high precision and high efficiency in complex environments, especially the traditional SLAM method has complex calculations and insufficient segmentation results. Deep learning-based methods do not fully utilize the content and location information of query, resulting in insufficient accuracy.

Method used

A Transformers-based encoder-decoder model is designed. By extracting vehicle ring view feature images, using scattered query and position embedding decoder, combined with BEV features and random initialization query, the map structure is decoded to improve the map prediction accuracy.

Benefits of technology

The accuracy of 72% and above was achieved on the nuScenes and Argoverse2 datasets, and the efficiency was maintained at 17.9Hz, which significantly improved the accuracy and efficiency of map construction, and improved the accuracy of 1.5% compared to the GKT model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259991A_ABST
    Figure CN120259991A_ABST
Patent Text Reader

Abstract

The invention discloses an online map construction method based on deep learning, and the method comprises the steps: extracting a feature image of a vehicle surround view, and obtaining a feature map based on a 2D feature extraction network; the method comprises the following steps of: constructing an encoder-decoder model based on Transforms; after the feature pattern is coded through the coder-decoder model of the Transformers, BEV features are obtained; and taking the BEV features and the randomly initialized query as the input of an encoder-decoder model of the Transformers, taking the reference point of the query as prior information to carry out position coding, and carrying out decoding through a decoder by utilizing a bulk aggregation query to obtain a map structure diagram represented by each query. According to the method, the threshold value of a nuScenes data set {0.5 m, 1.0 m, 1.5 m} reaches the precision of 72% or above, and high efficiency (17.9 Hz) is maintained; a threshold value of a nuScenes data set {0.2 m, 0.5 m, 1.0 m} reaches the precision of more than 50%; the threshold value of an Argoverse2 data set {0.5 m, 1.0 m, 1.5 m} reaches the precision of 68% or above; under the condition that the same decoder is used, the precision of the BEV decoder provided by the method can be improved by 1.5% compared with that of GKT.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous driving online map construction, and is an online map construction method based on deep learning. Background Art

[0002] Although the traditional high-definition map construction method based on SLAM (Simultaneous Localization and Mapping) has certain effects, it involves a large amount of calculation and data processing, which increases the complexity of the technology and the implementation cost. Secondly, these methods often have difficulty obtaining accurate results, especially in complex roads and environments, and are prone to problems such as positioning errors or insufficient map accuracy.

[0003] Many deep learning-based works define map construction as a semantic segmentation task in the bird's-eye view (BEV) space to generate a rasterized map. Some methods further group the segmentation results into a vectorized representation or use autoregression for point sequence prediction. Although they have achieved great success, they face limitations due to the need for a large amount of post-processing to obtain vectorized information.

[0004] To overcome the limitations of the segmentation-based methods, new methods have emerged to predict point sets to construct maps, using a structure similar to DETR for end-to-end map construction. MapTR regards online map construction as a point set prediction problem and designs a framework similar to DETR, achieving state-of-the-art performance, but its accuracy is not sufficient to meet the actual needs. There are also some methods that use complex models to fit curves, resulting in inefficiency in practical scenarios.

[0005] Existing end-to-end map construction models do not fully utilize the role of queries, and do not add enough content and location information to the queries, resulting in the inability to achieve a very high level of accuracy under the premise of meeting certain efficiency requirements. Therefore, a model that can fully utilize the role of queries is needed to meet the requirements of map accuracy while ensuring high efficiency. Summary of the Invention

[0006] To overcome the above-mentioned deficiencies of the prior art, the present invention designs a scatter-gather query of a new decoder by using the content information and location information of the query. Therefore, in cross-attention, each point query of the same instance shares the same content information, and different location information can be embedded from different reference points. By applying the proposed decoder to other models, the effectiveness of the decoder is further proven. With the improvement of our BEV encoder, our method has achieved excellent results in the nuScenes data.

[0007] The technical solution of the present invention is as follows:

[0008] An online map construction method based on deep learning, characterized in that it includes the steps:

[0009] S1. Extract the feature image of the vehicle's surround view to obtain a feature map based on a 2D feature extraction network, wherein the surround view is acquired by a multi-view camera;

[0010] S2. Construct an encoder-decoder model based on Transformers;

[0011] S3. After encoding the feature map through the encoder-decoder model of Transformers, obtain the BEV feature;

[0012] S4. Use the BEV feature and randomly initialized query as the input of the encoder-decoder model of Transformers, perform position encoding on the reference point of the query as prior information, and use scatter-gather query to decode through the decoder to obtain the map structure diagram represented by each query.

[0013] Furthermore, in step S2 of constructing the encoder-decoder model based on Transformers, it specifically includes;

[0014] Release the height in the feature space for the encoder to obtain more accurate BEV features;

[0015] Design scatter-gather query for the decoder to obtain more accurate map predictions when only using instance queries.

[0016] Furthermore, in step S3 of obtaining the BEV feature after encoding the feature map through the BEV encoder, it specifically includes:

[0017] S3.1 Project the 2D position of the surround view and sample 3D reference points from a fixed height for each BEV query;

[0018] S3.2 Learn an adaptive offset from the projected 2D position

[0019] S3.3 Use a fixed kernel to extract features around the projected 2D image position;

[0020] S3.4 Predict an adaptive height offset to the original fixed height when generating 3D reference points for each bev query. Furthermore, step S4 specifically includes:

[0021] S4.1 Randomly initialize a set of queries and use them as input together with the BEV feature;

[0022] S4.2 Perform Self-attention between queries;

[0023] The output of S4.3 self-attention is used as the input of cross-attention and interacts with the BEV features.

[0024] The output of S4.4 cross-attention passes through an MLP to obtain the final query, which corresponds to each map instance.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] 1) Using the scatter-gather query and only instance queries in the decoder:

[0027] a) At the thresholds of the nuScenes dataset {0.5m, 1.0m, 1.5m}, it achieves an accuracy of over 72% and maintains a high efficiency (17.9Hz);

[0028] b) At the thresholds of the nuScenes dataset {0.2m, 0.5m, 1.0m}, it achieves an accuracy of over 50%;

[0029] c) At the thresholds of the Argoverse 2 dataset {0.5m, 1.0m, 1.5m}, it achieves an accuracy of over 68%;

[0030] 2) The BEV decoder using the release height can improve the accuracy by 1.5% compared to GKT. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flowchart of a method for online map construction based on deep learning according to the present invention;

[0032] Figure 2 is a model structure diagram of a method for online map construction based on deep learning according to the present invention

[0033] Figure 3 is a decoder structure diagram in the present invention

[0034] Figure 4 is an encoder structure diagram in the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The technical solution of the present invention will be further described below with reference to the drawings and embodiments, but the protection scope of the present invention should not be limited thereby.

[0036] Please refer to Figure 1 , Figure 1 which is a flowchart of a method for online map construction based on deep learning according to the present invention, as shown in Figure 1As shown, first, the image features of the vehicle surround view are extracted. The surround view is obtained by a multi-view camera, and the features are extracted by a general 2D backbone network such as ResNet50; then, through a BEV encoder, BEV features are obtained; the BEV features and randomly initialized queries are used as the input to the Transformer decoder, where one query represents one map instance; the reference point of the query is used as prior information for position encoding, and at the same time, the designed scatter-and-gather query is used to obtain the map instance represented by each query through the decoder. As Figure 2 shown, it is the specific model structure diagram.

[0037] Based on MapTr, in the decoder, the content information and position information of the queries are used, as Figure 3 shown. The decoder consists of stacked Transformer layers, and the interaction and cooperation between these layers enable the decoder to gradually extract the features of the input data and convert them into the internal representation of the model. The design of this structure aims to effectively capture the features of the input data and convert them into meaningful representations for subsequent model use. Its detailed structure is as follows Figure 4 shown. The improvement of the decoder mainly focuses on query design, including

[0038] Scatter-and-Gather and its compatible position embedding. Only one type of query, i.e., instance query, is used as the input to the decoder. The instance query is a special type of query that queries for each input instance. This method can more effectively capture the features of the input data and provide more accurate representations. In addition, by using instance queries, the decoder can better handle input sequences of different lengths. MapTR uses point queries as the input to the decoder layer and calculates the self-attention scores between them. This design makes the computational complexity of self-attention relatively high because a large number of queries are involved. In contrast, the query design of the present invention is different. This design ensures the consistency elements of the content in the same mapping and reduces the number of queries, thereby reducing the computational consumption of self-attention. Therefore, it can participate in more instance queries without significantly increasing the memory consumption.

[0039] The BEV encoder can learn to implicitly or explicitly adapt to the 3D space. BEVFormer samples 3D reference points from a fixed height for each BEV query, while learning adaptive offsets from the projected 2D positions. Similar to BEVFormer, GKT uses a fixed kernel to extract features around the projected 2D image positions. Since GKT relies on a fixed 3D transformation, it lacks a certain degree of flexibility. Nevertheless, its performance is still very excellent compared to BEVFormer. To further improve the flexibility of GKT, the present invention predicts an adaptive height offset to the original fixed height when generating 3D reference points for each bev query, as Figure 4 shown. The offset is predicted using a learnable linear projection, and 3D points are sampled from the BEV space as reference points. After projecting onto the two-dimensional image, the model still uses a fixed kernel to extract features.

[0040] Following most of the settings in the MapTR series, the modifications made by the proposed model will be emphasized. The BEV feature is set to 200×100, for perceiving [-30m, 30m] from back to front and [-15m, 15m] from left to right. 100 instance queries are used to detect map element instances, and each instance is modeled by 20 sequential points. (N = 100 and n = 20 are set). The model is trained on 8 NVIDIA 4090 GPUs with a batch size of 8×4. The learning rate is set to 6×10-4.

[0041] In the performance evaluation of the present invention, the key metric of Chamfer Distance is introduced to more accurately quantify the matching degree between the prediction result and the actual ground situation. Through this metric, the accuracy and integrity of the map construction algorithm for each map element in different scenarios can be comprehensively evaluated. To gain a deeper understanding of the algorithm performance, two different sets of Chamfer Distance thresholds are adopted. First, the standard thresholds used by MapTR in previous studies, namely {0.5m, 1.0m, 1.5m}, are borrowed, and FPS is used to reflect the efficiency

[0042]

[0043]

[0044] Using smaller thresholds, namely {0.2m, 0.5m, 1.0m}, the results are shown in the following table:

[0045]

[0046] To further demonstrate the effectiveness of the two components in the decoder, an ablation study of the disaggregation query and positional embedding is provided in the following table. The first row corresponds to the original MapTR, which uses point queries and learnable positional embeddings. mAP1 represents the thresholds of {0.2m, 0.5m, 1.0m}, and mAP1 represents the thresholds of {0.5m, 1.0m, 1.5m}.

[0047]

[0048]

Claims

1. An online map construction method based on deep learning, characterized in that Including the steps: S1. Extract the feature image of the vehicle surround view to obtain the feature map based on the 2D feature extraction network, where the surround view is acquired by a multi-view camera; S2. Construct an encoder-decoder model based on Transformers; S3. After encoding the feature map through the encoder-decoder model of Transformers, obtain the BEV feature; S4. Use the BEV feature and the randomly initialized query as the input of the encoder-decoder model of Transformers, perform position encoding on the reference point of the query as prior information, and use scatter-gather queries to decode through the decoder to obtain the map structure diagram represented by each query.

2. The online map construction method based on deep learning according to claim 1, wherein The step S2. Constructing an encoder-decoder model based on Transformers specifically includes: Release the height in the feature space for the encoder to obtain a more accurate BEV feature; Design scatter-gather queries for the decoder to obtain a more accurate map prediction in the case of only using instance queries.

3. The online map construction method based on deep learning according to claim 1, wherein, The step S3. After encoding the feature map through the BEV encoder, obtain the BEV feature, specifically including: S3.1 Project the 2D position of the surround view and sample 3D reference points from a fixed height for each BEV query; S3.2 Learn the adaptive offset from the projected 2D position S3.3 Use a fixed kernel to extract the features around the projected 2D image position; S3.4 Predict the adaptive height offset to the original fixed height when generating 3D reference points for each bev query.

4. The online map construction method based on deep learning according to claim 1, characterized in that The step S4. Specifically includes: S4.1 Randomly initialize a group of queries and use them as the input together with the BEV feature; S4.2 Perform Self-attention among the queries; S4.3 Use the output of Self-attention as the input of Cross-attention and interact with the BEV feature; S4.4 The output of Cross-attention passes through the MLP to obtain the final query, that is, corresponding to each map instance.