Online vectorization high-precision map generation method based on point features

By improving the PointMapNet network of the BEV feature encoding and decoding module, using the point features and distance attention masks that can be learned by location, it solves the problems of slow speed, high cost and poor range expansion in the existing technology of online vectored high-precision map generation, and realizes high-precision map generation with high precision and high real-time.

CN120339429APending Publication Date: 2025-07-18SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510400566.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing online vectorized high-precision map generation method has problems such as slow inference speed, high training cost and poor perceived range scalability, which is difficult to meet the requirements of autonomous driving systems for high precision and high real-time.

Method used

The BEV feature encoding module is improved by using PointMapNet network, using point features that can be learned instead of fixed position sampling, and optimizing point feature locations through binary map matching and distance loss; using distance-based attention masks in the map element decoding module to reduce redundant information and improve network convergence speed.

Benefits of technology

It realizes more efficient vectorized high-precision map generation, reduces training time and cost, and improves the scalability and generation effect of the perceived range, meeting the real-time needs of the autonomous driving system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339429A_ABST
    Figure CN120339429A_ABST
Patent Text Reader

Abstract

The invention discloses an online vectorization high-precision map generation method based on point features, which is characterized in that real-time high-quality generation of an online vectorization high-precision map is realized by utilizing a Point MapNet network, and the Point MapNet network is obtained by improving a BEV feature coding module and a map element decoding module of an original MapTR network; the BEV feature coding module is improved in the following steps: replacing BEV features sampled at fixed positions with position learnable point features, naming the improved BEV coding module as a point feature coding module, enabling the point features to learn a spatial position containing a detection object in a sensing area through bipartite graph matching and distance loss calculation, and updating the position of the point features; the map element decoding module is improved as follows: attention masks based on distance are used, so that map element features are only interacted with point features of a near space. Through the improvement, the reasoning speed is increased and the training cost is reduced while higher map generation quality is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online vectorized high-precision map generation, and in particular to an online vectorized high-precision map generation method based on point features. Background Art

[0002] High-precision maps around the vehicle, such as lane lines, crosswalks, and drivable area boundaries, are crucial for autonomous driving systems. For example, the system needs to plan the route according to the drivable area boundaries and adjust the driving direction according to the distance between the vehicle and the lane lines on both sides. In traditional driving behavior, navigation software displays high-precision maps around the vehicle through satellite positioning, relying on human drivers to deal with problems such as inaccurate vehicle positioning, road conditions not matching offline maps, and offline maps not being covered. In autonomous driving systems, it is necessary to solve the above problems by generating maps online. The online generated map is centered on the vehicle, and can assist in positioning in combination with offline maps. Online generation can ensure real-time and safety. Traditional online vectorized high-precision map generation methods usually use a detection model based on BEV features. The model is usually divided into three parts: image feature extraction module, BEV feature encoding module, and map element decoding module. In the BEV feature encoding module, by evenly setting sampling points in the detected BEV space, these BEV space points are projected to the image plane for sampling through camera internal and external parameters, or depth prediction is added to the image features to project the image features to the BEV space to obtain BEV features. All BEV-based methods need to be divided in the BEV space. The denser the division, the smaller the space represented by each BEV feature, and the more accurate the implicit spatial location information. The online vectorized high-precision map generation task requires very accurate spatial location information, so the number of BEV features will be very high. The original vectorized high-precision map generation method MapTR network has as many as 20,000 BEV features. This causes the current vectorized high-precision map generation methods to generally have problems such as slow reasoning speed and high training cost. Slow reasoning speed means that the autonomous driving system reacts slowly to sudden road conditions, posing serious safety hazards. The high training cost means that each update iteration requires high training costs. In addition, when the detection range needs to be expanded, in order to maintain the same spatial division accuracy, the detection space range grows linearly, and the number of BEV features needs to grow in the form of a square, which seriously restricts the expansion of the perception range. The expansion of the perception range means the improvement of the safety and comfort of the autonomous driving system, which is a goal that the autonomous driving system must pursue.

[0003] Based on the above discussion, an online vectorized high-precision map generation method that meets the requirements of high precision, good real-time performance and scalability is invented, which has high practical application value. Summary of the invention

[0004] The purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art, and propose an online vectorized high-precision map generation method based on point features, which can effectively solve the problems of slow inference speed, high training cost, and poor scalability of the perception range in traditional online vectorized high-precision map methods, and meet the requirements of high real-time while achieving high precision.

[0005] To achieve the above purpose, the technical solution provided by the present invention is: an online vectorized high-precision map generation method based on point features. This method uses the PointMapNet network to achieve high-quality generation of real-time online vectorized high-precision maps. The PointMapNet network is an improved MapTR network, which improves the BEV feature encoding module and the map element decoding module of the original MapTR network. The improvement of the BEV feature encoding module is: using position-learnable point features to replace the BEV features sampled at fixed positions, and calling the improved BEV encoding module the point feature encoding module. By bipartite graph matching and calculating the distance loss, the point features learn the spatial positions of the detection objects in the perception area and update their own positions, thereby reducing redundant information. The improvement of the map element decoding module is: using a distance-based attention mask to make the map element features only interact with the point features in the adjacent space, accelerating the convergence of the network;

[0006] The specific implementation of the online vectorized high-precision map generation method includes the following steps:

[0007] 1) Collect multi-view images of the vehicle-mounted surround camera at the same time and the internal and external camera parameters. Generate the road map around the vehicle at the current time through vehicle positioning and offline high-precision maps, and correct it to obtain an accurate road map. Convert the road map into a vectorized representation to obtain a vectorized high-precision map. Integrate the multi-view images, camera internal and external parameters, and vectorized high-precision map into a data set, and divide the data set into a training set and a test set;

[0008] 2) Feed the training set into the PointMapNet network for training. During training, first perform data augmentation on the multi-view images in the training set, and then input the augmented multi-view images into the PointMapNet network. Extract the multi-view multi-scale image features of the multi-view images through the feature extraction module of the PointMapNet network. Input the extracted multi-view multi-scale image features into the point feature encoding module. The point feature encoding module uses the internal and external camera parameters to perform multi-view multi-scale fusion and generates point features to be input into the map element decoding module to generate a vectorized high-precision map. The map elements in the generated vectorized high-precision map are subjected to bipartite graph matching with the map elements in the vectorized high-precision map in the training set to obtain map element matching. The point features are subjected to bipartite graph matching with the points sampled from the map elements on the vectorized high-precision map in the training set to obtain point feature matching. Calculate the classification loss and distance loss for the point feature matching and map element matching respectively, and use backpropagation to optimize the network parameters. After multiple iterations until the loss value is minimized, obtain the optimal network.

[0009] 3) Input the test set into the optimal network, and the vectorized high-precision map around the vehicle can be obtained.

[0010] Furthermore, step 1) includes the following steps:

[0011] 1.1) Data acquisition: Install the surround-view cameras at fixed positions on the vehicle, record the internal and external camera parameters, set the cameras to sample at the same frequency and start at the same time, collect multi-view images during the vehicle's driving process, and record the vehicle positioning and vehicle attitude at the same frequency.

[0012] 1.2) Image preprocessing: After data acquisition, take the multi-view images collected at the same time, their corresponding internal and external camera parameters, and the vehicle position as a set of data, and eliminate the data with a large time gap between the shooting times of the multi-view images within the set and obvious offsets in the camera positions.

[0013] 1.3) Data annotation: For each set of data, crop the road map around the vehicle at the current moment from the offline high-precision map through the vehicle position and vehicle attitude, and rotate it to match the vehicle orientation. Convert the cropped map into a vectorized high-precision map represented by an ordered point set, and set category labels according to the map elements.

[0014] 1.4) Dataset division: After data annotation is completed, integrate the multi-view images, internal and external camera parameters, and vectorized high-precision map into a dataset, and divide it into a training set and a test set according to a certain proportion.

[0015] Further, in step 2), the data augmentation methods include: image scaling, random brightness, random contrast, random saturation, random hue, and GridMask data augmentation. GridMask data augmentation means adding a square mask at a random position on the image to occlude part of the image information, thereby increasing the data diversity and improving the robustness of the model.

[0016] Further, in step 2), the feature extraction module includes a Resnet50 network and an FPN network. The Resnet50 network is a multi-scale image feature extraction network, and the FPN network is a multi-scale enhancement module. The multi-view images are input into the Resnet50 network to obtain multi-view multi-scale image features. After information exchange through the FPN network, the feature dimension is uniformly scaled to 256 to obtain multi-view multi-scale image features with a unified feature dimension, which are then input into the point feature encoding module.

[0017] The point feature encoding module performs view transformation and fusion on the multi-view multi-scale image features output by the feature extraction module of the PointMapNet network to generate point features that focus on the spatial position of the detection object: First, a set of learnable feature representation points are used as point features. A two-dimensional coordinate in the bird's-eye view space is predicted through a multi-layer linear network, and the coordinate is bound to the point feature. The two-dimensional coordinate is sampled at multiple heights to obtain a three-dimensional coordinate, which is projected onto the image planes of multiple views through the internal and external camera parameters. Local image features are collected through deformable attention and then projected through a linear layer to obtain point features. The newly obtained point features obtain a new two-dimensional coordinate through a multi-layer linear network, and the point features are updated through a new round of projection sampling. After multiple layers of projection sampling, the point features all focus on the spatial position of the detection object, thereby obtaining concise and efficient point features.

[0018] The map element decoding module efficiently detects map elements from the point features using global attention with spatial constraints. The map element decoding module is implemented through a Transformer decoder. The point features obtained by the point feature encoding module are used as the keys and values of the decoder. By combining m line features and n anchor features, m*n queries can be generated, and each query corresponds to a map element anchor.

[0019] For each map element, a fixed number of queries interact with a variable number of point features associated with multiple map elements, resulting in the re-arranged coordinates of n anchor points. Using a distance-based attention mask, which allows the queries to compute attention only with specific point features, each query is input into a multi-layer perceptron to predict a two-dimensional point coordinate as the anchor coordinate. Then, the Euclidean distances from all anchor coordinates to all point feature binding coordinates are calculated. The minimum distance among the distances from the point feature binding coordinates to each set of anchors is taken as the distance from the point feature to the map element. By pre-defining a threshold, a distance attention mask is generated, enabling the queries to interact only with the point features that are spatially close to the map element they belong to. The distance attention mask not only speeds up the network's convergence rate but also forces each set of queries to focus on the same set of point features, thereby enhancing the anchor queries' perception ability of the map elements.

[0020] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0021] 1. The data input of the present invention only depends on the relatively low-cost camera sensor and satellite positioning sensor, which is conducive to the application and popularization of the method.

[0022] 2. The present invention uses object-oriented point features instead of region-oriented BEV features, achieving a significant reduction in the number of encoded layer features while improving the generation effect. Specifically, BEV features use fixed-position sampling in space, encoding a large number of spatial positions that do not contain detection objects into BEV features at the same cost, resulting in a large amount of redundant calculation, and its position information is also limited by the division accuracy of the BEV space. Point features focus on the spatial positions containing detection objects, saving a large amount of calculation while also containing more accurate position information of the detection objects.

[0023] 3. Due to the significant reduction in the number of encoded layer features, the training time of the present invention is only half of that of the MapTR network. A shorter training time means a higher update frequency and lower training cost, which has significant economic benefits for companies researching autonomous driving technology that usually use extremely large datasets. Brief Description of the Drawings

[0024] Figure 1 It is the overall architecture diagram of the method of the present invention. In the figure, (x, y) represents the coordinates of the point feature in the xy plane of the lidar coordinate system, (x, y, z') represents one of the three-dimensional coordinates after multi-height sampling of the point feature, and (x', y') represents the updated coordinates of the point feature.

[0025] Figure 2This is a comparison of the embodiments of the present invention with existing methods in terms of efficiency. The vertical axis of the left and right figures is the mean average precision based on chamfer distance tested on the validation set of the nuScenes dataset, which reflects the accuracy of the method for generating vectorized high-precision maps, and the higher the value, the better. The horizontal axis of the left figure represents the inference speed, in frames per second. The higher the number of frames per second, the faster the inference speed. As shown in the figure, the PointMapNet network achieves a better balance between accuracy and inference speed. The horizontal axis of the right figure represents the video memory occupancy during model training, which reflects the training cost of the model. The closer to the upper left corner, the higher the accuracy while the lower the training cost. The PointMapNet network demonstrates an obvious benefit advantage.

[0026] Figure 3 This is a comparison chart of the main innovation points of the present invention - point features and BEV features used in the original method. BEV features are encoded in a fixed area and contain a large amount of information unrelated to the task, while point features focus on detecting objects through bipartite graph matching, achieving more efficient vectorized high-precision map generation.

[0027] Figure 4 This is a result chart of vectorized high-precision map generation in different weather conditions for the embodiments of the present invention. Each row shows the generation effect comparison under a specific weather condition. The first three columns are the multi-view images input to the network, the fourth column is the vectorized high-precision map generated by the PointMapNet network, the fifth column is the vectorized high-precision map generated by the MapTR network, and the sixth column is the real vectorized high-precision map. The areas with generation errors are highlighted with shadows in the figure, and shadows are also added to the corresponding areas of the real vectorized high-precision map. It can be seen that compared with the MapTR network, the PointMapNet network shows better map generation effects under various weather conditions. Even on a very dark night with poor lighting conditions, the PointMapNet network can generate a high-precision map around the vehicle as accurately as possible. Detailed implementation manners

[0028] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0029] As Figures 1 to 4 shown, this embodiment discloses an online vectorized high-precision map generation method based on point features. This method uses the PointMapNet network to achieve high-quality real-time online vectorized high-precision map generation. The PointMapNet network is an improved MapTR network, which improves the BEV feature encoding module and the map element decoding module of the original MapTR network; As Figure 3As shown in the figure, the improvement of the BEV feature encoding module is as follows: using position-learnable point features instead of BEV features sampled at fixed positions, and calling the improved BEV encoding module the point feature encoding module. By bipartite graph matching and calculating the distance loss, the point features learn to perceive the spatial positions of the detection objects in the region and update their own positions, thereby reducing redundant information. The improvement of the map element decoding module is as follows: using a distance-based attention mask to enable the map element features to interact only with the point features in the adjacent space, accelerating the network convergence.

[0030] The specific implementation of the above online vectorized high-precision map generation method based on point features includes the following steps:

[0031] 1) Collect multi-view images of the vehicle-mounted surround-view camera at the same time and the internal and external camera parameters. Generate the road map around the vehicle at the current time through vehicle positioning and offline high-precision map, and make corrections to obtain an accurate road map. Convert the road map into a vectorized representation to obtain a vectorized high-precision map. Integrate the multi-view images, internal and external camera parameters, and vectorized high-precision map into a data set, and divide the data set into a training set and a test set; including the following steps:

[0032] 1.1) Data collection: Install the surround-view camera at a fixed position on the vehicle, record the internal and external camera parameters, and set the camera to sample at the same frequency and start at the same moment. Collect the surround-view images (i.e., multi-view images) during the vehicle's driving process, and record the vehicle positioning and vehicle attitude at the same frequency;

[0033] 1.2) Image preprocessing: After data collection, take the multi-view images collected at the same time, their corresponding internal and external camera parameters, and the vehicle position as a set of data, and eliminate the data with a large time gap between the multi-view images in the set and obvious offsets in the camera positions;

[0034] 1.3) Data annotation: For each set of data, crop the road map around the vehicle at the current time from the offline high-precision map through the vehicle position and vehicle attitude, and rotate it to match the vehicle orientation. Convert the cropped map into a vectorized high-precision map represented by an ordered point set, and set category labels according to the map elements;

[0035] 1.4) Dataset division: After data annotation, integrate the multi-view images, internal and external camera parameters, and vectorized high-precision map into a data set, and divide it into a training set and a test set according to a certain proportion.

[0036] 2) Send the training set into the PointMapNet network for training. During training, first perform data augmentation on the multi-view images in the training set, and then input the augmented multi-view images into the PointMapNet network. As Figure 1As shown in the figure, the PointMapNet network mainly includes three modules: a feature extraction module, a point feature encoding module, and a map element decoding module. The multi-view multi-scale image features of the multi-view images are extracted through the feature extraction module of the PointMapNet network, and the extracted multi-view multi-scale image features are input into the point feature encoding module. The point feature encoding module uses the internal and external camera parameters to perform multi-view multi-scale fusion and generates point features to be input into the map element decoding module to generate a vectorized high-precision map; the map elements in the generated vectorized high-precision map are subjected to bipartite graph matching with the map elements in the vectorized high-precision map in the training set to obtain map element matching; the point features are subjected to bipartite graph matching with the points sampled from the map elements on the vectorized high-precision map in the training set to obtain point feature matching; the classification loss and the distance loss are calculated for the point feature matching and the map element matching respectively, and the network parameters are optimized using backpropagation. After multiple iterations until the loss value is minimized, the optimal network is obtained.

[0037] The data augmentation methods include: image scaling, random brightness, random contrast, random saturation, random hue, and GridMask data augmentation. GridMask data augmentation means adding a square mask at a random position on the image to block part of the image information, thereby increasing the diversity of the data and improving the robustness of the model.

[0038] The multi-view images are input into the feature extraction module of the PointMapNet network. The feature extraction module of the PointMapNet network includes a Resnet50 network and an FPN network. The Resnet50 network is a multi-scale image feature extraction network, and the FPN network is a multi-scale enhancement module. The multi-view images obtain multi-view multi-scale image features through the Resnet50 network. After information exchange through the FPN network, the feature dimension is uniformly scaled to 256 to obtain multi-view multi-scale image features with a unified feature dimension, which are input into the point feature encoding module.

[0039] The point feature encoding module performs view transformation and fusion on the multi-view multi-scale image features output by the feature extraction module of the PointMapNet network to generate point features focused on the spatial location of the detection object: First, a set of learnable feature representations are used to represent point features. A two-dimensional coordinate in the bird's-eye view space is predicted through a multi-layer linear network, and the coordinate is bound to the point feature. The two-dimensional coordinate is sampled at multiple heights to obtain a three-dimensional coordinate, and is projected onto the image planes of multiple views through the internal and external camera parameters. Local image features are collected through deformable attention, and then projected through a linear layer to obtain point features. The newly obtained point features pass through a multi-layer linear network to obtain a new two-dimensional coordinate, and the point features are updated through a new round of projection sampling. After multiple rounds of projection sampling, the point features all focus on the spatial location of the detection object, thus obtaining concise and efficient point features.

[0040] The map element decoding module uses global attention with spatial constraints to efficiently detect map elements from point features. The map element decoding module is implemented through a three-layer Transformer decoder, and the point features obtained by the point feature encoding module are used as the keys and values of the decoder. Specifically, by combining 50 line features and 20 anchor point features, 50 * 20, that is, 1000 queries can be generated, and each query corresponds to a map element anchor point.

[0041] For each map element, a fixed number of queries interact with a varying number of point features associated with multiple map elements. Thus, the rearranged 20 anchor point coordinates are obtained. The original Transformer decoder initially assigns equal attention weights to each point feature, resulting in slow model convergence. In this method, a distance-based attention mask is used, and the attention mask enables the queries to calculate attention only with specific point features. Specifically, each query is input into a multi-layer perceptron to predict a two-dimensional point coordinate as the anchor point coordinate, and then the Euclidean distance from all anchor point coordinates to the coordinates bound to all point features is calculated. The minimum value of the distance of the point feature bound coordinates relative to each group of anchors is used as the distance from the point feature to the map element. By pre-defining a threshold, a distance attention mask is generated, enabling the queries to interact only with the point features that are spatially close to the map element they belong to. The distance attention mask not only speeds up the convergence rate of the network but also forces each group of queries to focus on the same group of point features, thereby enhancing the perception ability of the anchor point queries for map elements.

[0042] After obtaining the features output by the map element decoding module, each feature is input into a multi-layer linear layer to predict the x and y coordinates. Every twenty points are used as the vector representation of a map element. There are a total of 50 map element prediction results, and each map element consists of twenty two-dimensional points. These 50 map element prediction results are subjected to bipartite graph matching with the annotated vector map elements. After obtaining the matching results, the classification loss and coordinate loss of all forward-matched map element predictions are calculated. Then, the network is trained through backpropagation. The learning rate used in this method is 6e-4, and it is trained using 8 A800 GPUs, with the batch size set to 4 on each GPU.

[0043] 3) Input the data in the test set into the optimal network obtained through training to obtain the vectorized high-precision map around the vehicle.

[0044] The experimental results of this experiment are described in detail below:

[0045] The PointMapNet network is evaluated on the large-scale nuScenes dataset. The annotation frequency of the nuScenes dataset is 2 Hz. Each frame contains RGB images from 6 surround cameras, covering a 360-degree field of view. There are 1000 scenes in the dataset, and each scene contains approximately 40 frames. The dataset is divided into 28000 frames for training and 6000 frames for validation.

[0046] Referring to existing methods, three types of map elements are considered, namely lane lines, crosswalk lines, and drivable area boundary lines, to evaluate the construction of high-definition maps. The range of the generated map is defined as 30 meters before and after the vehicle and 15 meters to the left and right. The average precision (AP) based on the chamfer distance is used as the evaluation metric. The average precision at the three thresholds of {0.5, 1.0, 1.5} meters is used. Only when the distance between the predicted value and the true value is less than the specified threshold, the prediction result is regarded as a true positive (TP).

[0047] The experimental results on the nuScenes dataset are shown in Table 1 below. AP divider represents the average precision of lane lines, AP ped represents the average precision of crosswalk lines, AP boundary represents the average precision of drivable area boundary lines. mAP represents the average precision of these three types of objects, reflecting the accuracy of the method for generating vectorized high-precision maps. FPS represents how many times the map can be generated per second by the method, reflecting the inference speed of the method. The values in the FPS column are all tested on the device of the same RTX 4090 GPU. As can be seen from Table 1, the PointMapNet network achieves the highest vectorized high-precision map generation accuracy with fewer parameters and can reach 30.3 FPS at the real-time level.

[0048] Table 1

[0049] Method <![CDATA[AP divider > <![CDATA[AP ped > <![CDATA[AP boundary > mAP FPS Number of parameters HDMapNet 21.7 14.4 33.0 23.0 1.9 69.8M VectorMapNet 47.3 36.1 39.3 40.9 6.2 19.4M MapTR 56.2 59.8 60.1 58.7 24.2 35.9M PointMapNet 61.6 61.6 63.3 62.2 30.3 33.0M

[0050] As Figure 2 shown, it can be seen from the two scatter plots that the PointMapNet network exceeds the original method in terms of both generation accuracy and inference speed while consuming less resources than the original method.

[0051] Comparison of the vectorized high-precision map generation effects of the PointMapNet network and the MapTR network under different weather conditions is as Figure 4 shown. The fourth column in the figure is the generation result of the PointMapNet network, the fifth column is the generation result of the MapTR network, and the sixth column is the ground truth given by the dataset. It can be seen that at many relatively complex details, the PointMapNet network can generate accurate road maps, and the errors are mostly in the problem of the extension distance, which are errors with relatively small impacts. It can be seen from the last row that even at night, the PointMapNet network can still generate relatively accurate road maps.

[0052] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. An online vectorization high-precision map generation method based on point features, characterized in that, This method uses the PointMapNet network to achieve the high-quality generation of real-time online vectorized high-precision maps. The PointMapNet network is an improved MapTR network, which improves the BEV feature encoding module and the map element decoding module of the original MapTR network. The improvement of the BEV feature encoding module is as follows: using position-learnable point features instead of BEV features sampled at fixed positions, and calling the improved BEV encoding module the point feature encoding module. By bipartite graph matching and calculating the distance loss, the point features learn to perceive the spatial positions of detection objects in the region and update their own positions, thereby reducing redundant information. The improvement of the map element decoding module is as follows: using a distance-based attention mask to enable the map element features to interact only with the point features in the adjacent space, accelerating network convergence. The specific implementation of the online vectorized high-precision map generation method includes the following steps: 1) Collect multi-view images of the vehicle-mounted surround camera at the same time and the internal and external camera parameters. Generate the road map around the vehicle at the current time through vehicle positioning and offline high-precision maps, and make corrections to obtain an accurate road map. Convert the road map into a vectorized representation to obtain a vectorized high-precision map. Integrate the multi-view images, the internal and external camera parameters, and the vectorized high-precision map into a data set, and divide the data set into a training set and a test set. 2) Feed the training set into the PointMapNet network for training. During training, first perform data augmentation on the multi-view images in the training set, and then input the augmented multi-view images into the PointMapNet network. Extract the multi-view multi-scale image features of the multi-view images through the feature extraction module of the PointMapNet network. Input the extracted multi-view multi-scale image features into the point feature encoding module. The point feature encoding module performs multi-view multi-scale fusion using the internal and external camera parameters and generates point features to be input into the map element decoding module to generate a vectorized high-precision map. The map elements in the generated vectorized high-precision map are subjected to bipartite graph matching with the map elements in the vectorized high-precision map in the training set to obtain map element matching. The point features are subjected to bipartite graph matching with the points sampled from the map elements on the vectorized high-precision map in the training set to obtain point feature matching. Calculate the classification loss and the distance loss for the point feature matching and the map element matching respectively, and use backpropagation to optimize the network parameters. After multiple iterations until the loss value is minimized, obtain the optimal network. 3) Input the test set into the optimal network to obtain the vectorized high-precision map around the vehicle.

2. The method for generating a high-precision map by online vectorization based on point features according to claim 1, wherein Step 1) includes the following steps: 1.1) Data collection: Install the surround camera at a fixed position on the vehicle, record the internal and external camera parameters, set the camera to sample at the same frequency and start at the same moment, collect multi-view images during the vehicle's driving process, and record the vehicle positioning and vehicle attitude at the same frequency. 1.2) Image preprocessing: After collecting data, multi-view images collected at the same time, their corresponding internal and external camera parameters, and the vehicle position are taken as a set of data. Data with a large time gap between multi-view images within the set and obvious offsets in the camera position are removed; 1.3) Data annotation: For each set of data, the road map around the vehicle at the current moment is cropped from the offline high-precision map based on the vehicle position and vehicle attitude, and rotated to match the vehicle orientation. The cropped map is converted into a vectorized high-precision map represented by an ordered point set, and category labels are set according to map elements; 1.4) Dataset division: After data annotation is completed, the multi-view images, internal and external camera parameters, and vectorized high-precision map are integrated into a dataset and divided into a training set and a test set according to a certain proportion.

3. The method for generating an online vectorized high-precision map based on point features according to claim 1, wherein In step 2), the data augmentation methods include: image scaling, random brightness, random contrast, random saturation, random hue, and GridMask data augmentation. GridMask data augmentation means adding a square mask at a random position on the image to occlude part of the image information, thereby increasing data diversity and improving the robustness of the model.

4. The online vectorization high-precision map generation method based on point features according to claim 1, characterized in that In step 2), the feature extraction module includes a Resnet50 network and an FPN network. The Resnet50 network is a multi-scale image feature extraction network, and the FPN network is a multi-scale enhancement module. The multi-view images obtain multi-view multi-scale image features through the Resnet50 network. After information exchange through the FPN network, the feature dimension is uniformly scaled to 256 to obtain multi-view multi-scale image features with a unified feature dimension, which are input into the point feature encoding module; The point feature encoding module performs view transformation and fusion on the multi-view multi-scale image features output by the feature extraction module of the PointMapNet network to generate point features focused on the spatial position of the detection object: First, a set of learnable feature representations are used to represent point features, and a two-dimensional coordinate in the bird's-eye view space is predicted through a multi-layer linear network. The coordinate is bound to the point feature, and the two-dimensional coordinate is sampled at multiple heights to obtain a three-dimensional coordinate, which is projected onto the multi-view image plane through the internal and external camera parameters. Local image features are collected through deformable attention, and then projected through a linear layer to obtain point features. The newly obtained point features obtain a new two-dimensional coordinate through a multi-layer linear network, and the point features are updated through a new round of projection sampling. Through multiple layers of projection sampling, the point features all focus on the spatial position of the detection object, thereby obtaining concise and efficient point features; The map element decoding module uses global attention with spatial constraints to efficiently detect map elements from point features. The map element decoding module is implemented through a Transformer decoder. The point features obtained by the point feature encoding module are used as the keys and values of the decoder. By combining m line features and n anchor features, m*n queries can be generated, and each query corresponds to a map element anchor; For each map element, a fixed number of queries interact with a variable number of point features associated with multiple map elements, resulting in the re-arranged coordinates of n anchor points. Using a distance-based attention mask, which enables the queries to compute attention only with specific point features, each query is input into a multi-layer perceptron to predict a two-dimensional point coordinate as the anchor coordinate. Then, the Euclidean distances from all the anchor coordinates to all the point feature binding coordinates are calculated. The minimum distance within each group of anchors for the point feature binding coordinates is taken as the distance from the point feature to the map element. By predefined a threshold, a distance attention mask is generated, enabling the queries to interact only with the point features whose spatial positions are close to the corresponding map element. The distance attention mask not only accelerates the network's convergence speed but also forces each group of queries to focus on the same group of point features, thereby enhancing the ability of the anchor queries to perceive the map elements.

Citation Information

Cited By

  • Vectorized high-precision map construction method and device, electronic equipment and medium

    CN122066820A