A high-precision map construction method and system based on the fusion of lidar and camera

By using multimodal fusion technology of lidar and camera in the online construction method of high-precision maps, and using the Transformer perspective conversion module to process feature maps, the problem of insufficient sensor complementarity of the single-modal method and data alignment of the multimodal method is solved, and the accurate construction and efficient calculation of high-precision maps are realized.

CN117008150BActive Publication Date: 2025-06-27FUDAN UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310518436.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-10
Publication Date
2025-06-27
Estimated Expiration
2043-05-10

AI Technical Summary

Technical Problem

The existing online construction method of high-precision maps has a single-modal method that can only utilize one sensor, but cannot utilize the complementary advantages between sensors; the multi-modal method cannot effectively construct the conversion relationship between perspective feature maps and bird's-eye feature maps, and the multi-modal feature map fusion method is simple, resulting in data aberration and hindering the correct prediction of the model.

Method used

Using a high-precision map construction method based on lidar and camera fusion, data is processed in parallel through image branches and point cloud branches, the multi-modal fusion module is used to fuse images and point cloud features, and the Transformer viewing angle conversion module is used to reduce the difficulty of model learning and calculation.

Benefits of technology

It realizes the construction of high-precision maps using the complementary advantages of cameras and lidar sensors, improves the prediction accuracy and computing efficiency of the model, and achieves the optimal performance of the current online construction method of high-precision maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117008150B_ABST
    Figure CN117008150B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-precision map construction method based on the fusion of lidar and camera, which includes: obtaining multi-view images, internal and external parameters of the camera, and point cloud data at the same timestamp from multiple perspective cameras, and respectively preprocessing the obtained multi-view and point cloud data to obtain preprocessed multi-view images and point cloud data. Inputting all the preprocessed point cloud data, multi-view images, and the projection matrix corresponding to the multi-view images into a pre-trained multi-modal high-precision map construction model to obtain prediction results corresponding to different road categories, and the prediction results include coordinate point sets of lane lines, curbs, and sidewalks. The present invention can solve the technical problem that the existing single-modal high-precision map online construction method can only use one sensor to construct the high-precision map online and cannot utilize the complementary advantages between sensors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of deep learning and visual perception, and more specifically, relates to a high-precision map construction method based on the fusion of lidar and camera. Background Art

[0002] The online construction of high-precision maps plays a crucial role in the field of autonomous driving research. The accurate perception of roads is often the first step in downstream tasks such as route planning, vehicle steering, and lane keeping. Therefore, the research on the online construction of high-precision maps is an important part of current deep learning, demonstrating great research potential and application value. High-precision maps usually contain information on various road categories such as lane lines, sidewalks, and curbs.

[0003] Existing online construction methods for high-precision maps can be divided into single-modal and multi-modal online construction methods for high-precision maps. Single-modal online construction methods for high-precision maps generally rely on a single sensor, usually a camera sensor. Multi-modal online construction methods for high-precision maps generally rely on two or more sensors. For single-modal online construction methods for high-precision maps, the general process is to first generate a bird's-eye view feature map (BEV feature map) from the data collected by the sensor, and then perform online construction of the high-precision map based on the generated BEV feature map. For multi-modal online construction methods for high-precision maps, BEV feature maps of different modalities are generated, and the BEV feature maps of different modalities are fused to obtain a fused BEV feature map, and then the online construction of the high-precision map is further performed based on the fused BEV feature map. In the process of online construction of high-precision maps, the process of generating a BEV feature map from a perspective (PV) feature map is very crucial, and the quality of the generated BEV feature map directly determines the accuracy of the constructed high-precision map. In existing methods for online construction of high-precision maps, a PV image is generally converted into a BEV feature map through a linear layer and a Transformer. The linear layer converts the PV feature map into a BEV feature map by constructing the relationship between each pixel in the PV feature map and each pixel in the BEV feature map. The Transformer takes each pixel in the BEV feature map as a query vector Query, and uses the PV feature map as the Key and Value to query the required semantic information.

[0004] However, the above - mentioned methods all have some non - negligible defects: First, the single - modality method can only use one type of sensor to online construct a high - precision map and cannot utilize the complementary advantages between sensors. Second, for the method of generating BEV feature maps in the above - mentioned multi - modality high - precision map online construction method, the linear - layer mapping is relatively simple and usually cannot effectively construct the conversion relationship between PV feature map pixels and BEV feature map pixels. And the method based on Transformer regards the conversion between PV feature maps and BEV feature maps as a graph - to - graph conversion, which is difficult for the model to learn and has a large computational cost. Third, the existing multi - modality high - precision map online construction methods simply splice along the channel dimension when fusing multiple - modality BEV feature maps. This way may have the phenomenon of misalignment between different - modality data, which will hinder the model from making correct predictions. Summary of the Invention

[0005] In view of the above - mentioned defects or improvement requirements of the prior art, the present invention provides a high - precision map construction method based on the fusion of lidar and camera. The purpose is to solve the technical problems that the existing single - modality high - precision map online construction method can only use one type of sensor to online construct a high - precision map and cannot utilize the complementary advantages between sensors, and the existing multi - modality high - precision map online construction method cannot effectively construct the conversion relationship between PV feature map pixels and BEV feature map pixels, and the method based on Transformer regards the conversion between PV feature maps and BEV feature maps as a graph - to - graph conversion, which is difficult for the model to learn and has a large computational cost, and the existing multi - modality high - precision map online construction method simply splices along the channel dimension when fusing multiple - modality BEV feature maps, and this way may have the phenomenon of misalignment between different - modality data, thus hindering the model from making correct predictions.

[0006] To achieve the above object, according to one aspect of the present invention, there is provided a high - precision map construction method based on the fusion of lidar and camera, including the following steps:

[0007] (1) Obtain multi - perspective images, internal and external parameters of the camera, and point - cloud data at the same timestamp from multiple - perspective cameras, and pre - process the obtained multi - perspective and point - cloud data respectively to obtain pre - processed multi - perspective images and point - cloud data.

[0008] (2) Input all the pre - processed point - cloud data, multi - perspective images, and the projection matrix corresponding to the multi - perspective images in step (1) into a pre - trained multi - modality high - precision map construction model to obtain prediction results corresponding to different road categories, and the prediction results include coordinate point sets of lane lines, road edges, and sidewalks.

[0009] Preferably, the internal and external parameters of the camera include the internal parameters required for the projection conversion from the camera coordinate system to the image coordinate system, and the external parameters required for the conversion from the camera coordinate system to the ego-vehicle coordinate system.

[0010] The preprocessing of the multi-view images in step (1) includes a scaling operation, which scales the multi-view images from the original size to 448×800×3 using the bilinear interpolation method.

[0011] The preprocessing operation of the point cloud data in step (1) is to filter out the point cloud data within the detection range, where the detection range is a rectangular area centered on the vehicle body, 30 meters in the front and back directions and 15 meters in the left and right directions.

[0012] Preferably, the high-precision map construction model includes an image branch, a point cloud branch, a multi-modal fusion module, and a detection head network implemented based on Transformer, which are sequentially connected. The image branch and the point cloud branch process the image data and the point cloud data in parallel, and the obtained image features and point cloud features are fused in the multi-modal fusion module, and finally sent to the detection head network to predict the high-precision map.

[0013] Preferably, the image branch includes an image encoder for extracting image features and a perspective view (PV) to bird's-eye view (BEV) perspective conversion module PV2BEV for converting the PV feature map into a BEV feature map. The input of the image branch is the multi-view image, and the output is the set of PV feature maps corresponding to the multi-view image and the complete image BEV feature map.

[0014] The image encoder includes a backbone network Efficientnet and a neck network FPN. Its input is the multi-view image I = {I1, I2,..., I N}, I i ∈R 3×H×W captured by N cameras at the same time, and the output is the set of PV feature maps corresponding to the multi-view image F i,u represents the PV feature map of the u-th scale corresponding to the image captured by the i-th camera, where i ∈ [1, N], u ∈ [1, 3], H pv,u 、W pv,u 、K respectively represent the height, width and number of channels of the u-th scale image feature map in the PV feature map F i,u corresponding to the image captured by the i-th camera, W represents the width of the multi-view image, H represents the height of the multi-view image, and the downsampling multiples of the finally obtained multi-scale PV image feature map F i,u are 1 / 16, 1 / 32, 1 / 32 respectively.

[0015] The PV2BEV module consists of two layers of Transformer decoder layers and a sampling layer. Its input is the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images output by the image encoder, which are transformed column-wise to obtain the set of image column features {F 1,u,w , F 2,u,w ,..., F N,u,w}, the set of image column position encodings and the set of position encodings of BEV polar rays The output is the complete image BEV feature map Where C bev represents the number of channels of the complete image BEV feature map, H bev represents the height of the complete image BEV feature map, and W bev represents the width of the complete image BEV feature map.

[0016] For the PV feature map F i,u corresponding to the image taken by the i-th camera at the u-th scale, the PV2BEV module takes its corresponding as Q, K, and V respectively for cross-attention calculation to obtain the polar ray features corresponding to this PV feature map Then, is concatenated along the azimuth axis to obtain the BEV feature map corresponding to this PV feature map in the polar coordinate system Finally, the BEV feature maps corresponding to the PV feature maps of all the images taken by the cameras are concatenated and sampled to obtain the complete image BEV feature map in the Cartesian coordinate system

[0017] Preferably, the point cloud branch assigns the point cloud data to the corresponding point cloud columns in the voxelization layer according to the given detection range and voxel resolution, and then encodes the point cloud data in each point cloud column using the HardVFE method to obtain the point cloud column features Then, all the point cloud columns are evenly dispersed in the BEV grid of the point cloud branch through the PointPillarsScatter module in the point cloud branch to obtain the preliminary point cloud BEV feature map. Finally, the preliminary point cloud BEV feature map obtained by the PointPillarsScatter module is encoded using the SECOND and SECOND FPN modules in the point cloud branch to obtain the final point cloud BEV feature map Where N point represents the number of point clouds, and P numIndicates the number of point cloud columns;

[0018] The input of the multi-modal fusion module is the set of PV feature maps {F corresponding to the multi-view images output by the image branch 1,u , F 2,u ,..., F N,u} and the complete image BEV feature map and the point cloud BEV feature map output by the point cloud branch The output is a fused BEV feature map that combines image information and point cloud information

[0019] The multi-modal fusion module first performs cross-attention calculation on the set of PV feature maps {F corresponding to the multi-view images 1,u , F 2,u ,..., F N,u} and the point cloud BEV feature map output by the point cloud branch to obtain their interaction features Then, the interaction features and the point cloud BEV feature map are concatenated along the channel dimension to obtain an enhanced point cloud BEV feature map Then, the enhanced BEV features are concatenated with the complete image BEV feature map along the channel dimension to obtain the concatenated feature map of the two Subsequently, the concatenated feature map is sequentially fed into a convolutional unit and a 1×1 convolutional block to obtain the offset of the image BEV feature map of Then, according to the offset Δ, resampling is performed again from the image BEV feature map to obtain an image BEV feature map aligned with the point cloud BEV feature map Finally, after concatenating the image BEV feature map and the point cloud BEV feature map along the channel dimension, it is fed into a 1×1 convolutional block for dimensionality reduction processing, thereby obtaining the final fused BEV feature map This module helps the model make full use of the information of both modalities, thereby constructing a more accurate high-precision map.

[0020] Preferably, the input of the detection head network is the fused BEV feature map output by the multi-modal fusion module and the query vector responsible for querying high-precision map instances where M represents the number of query vectors, d model ​Denote the dimension of the query vector, and the output is the prediction results of M high-precision map instances;

[0021] The detection head network takes the query vector and the fused BEV feature map to perform deformable attention calculation to obtain the prediction results of M high-precision map instances where respectively represent the class confidence of the i-th high-precision map instance and the coordinate point set of the high-precision map instance.

[0022] Preferably, the high-precision map construction model is trained through the following steps:

[0023] (2-1) Obtain a high-precision map autonomous driving dataset, which includes multi-view images, point cloud data, and camera internal and external parameter matrices at the same timestamp. Preprocess the obtained high-precision map autonomous driving dataset to obtain a preprocessed high-precision map autonomous driving dataset, and divide it into a training set and a validation set according to a ratio.

[0024] (2-2) Input the multi-view images and point cloud data in the training set obtained in step (2-1) into the image branch and point cloud branch of the high-precision map construction model respectively to obtain a set of PV feature maps {F 1,u ,F 2,u ,...,F N,u} corresponding to the multi-view images, the image BEV feature map and the point cloud BEV feature map

[0025] (2-3) Use the set of PV feature maps {F 1,u ,F 2,u ,...,F N,u} corresponding to the multi-view images obtained in step (2-) to enhance the point cloud BEV feature map to obtain an enhanced point cloud BEV feature map Concatenate the enhanced point cloud BEV feature map and the image BEV feature map along the channel dimension to obtain the final fused feature map

[0026] (2-4) Input the fused BEV feature map obtained in step (2-3) into the detection head network to obtain the prediction results of multiple high-precision map instances where respectively represent the class confidence of the i-th high-precision map instance and the coordinate point set of the high-precision map instance.

[0027] Use the prediction results of the high-precision map construction model obtained in (2-4) in (2-5). and the high-precision map labels Obtain the loss function of the high-precision map construction model;

[0028] (2-6) Use the loss function obtained in step (2-5) to iteratively train the high-precision map construction model until the high-precision map construction model converges, thereby obtaining a trained high-precision map construction model;

[0029] Preferably, step (2-2) includes the following sub-steps.

[0030] (2-2-1) Input the multi-view images in the training set obtained in step (2-1) into the image branch, and output the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images and the image BEV feature map

[0031] This step (2-2-1) includes the following sub-steps:

[0032] (2-2-1-1) Input the multi-view images into the shared multi-view image encoder to obtain the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images.

[0033] (2-2-1-2) Input the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images obtained in step (2-2-1-1) into the view transformation module of the image branch to obtain the image BEV feature map

[0034] This step (2-2-1-2) includes the following sub-steps:

[0035] (2-2-1-2-1) First, use the position information of the feature in the image and the BEV polar ray, the position information of a certain column in the image in the image, and the position information of the BEV polar ray in the BEV to encode the position encoding required by the view transformation module.

[0036] (2-2-1-2-2) After obtaining the position encoding in step (2-2-1-2-1) and the PV feature map set {F 1,u , F 2,u ,..., F N,u} The PV2BEV module inputs are combined to obtain a BEV feature map for obtaining epipolar ray features

[0037] (2-2-1-2-3) The epipolar ray features obtained in (2-2-1-2-2) are concatenated along the azimuth axis to obtain the BEV feature map in the polar coordinate system converted from the u-scale feature map of the i-th image:

[0038]

[0039] The BEV feature maps corresponding to the PV feature maps of the images captured by all cameras are concatenated and sampled to obtain the complete image BEV feature map in the Cartesian coordinate system

[0040] (2-2-2) Through the point cloud branch, according to the given detection range and voxel resolution, the point cloud data is assigned to the corresponding point cloud columns in the voxelization layer, and the HardVFE method is used to encode the point cloud data in each point cloud column to obtain the point cloud column features All point cloud columns are evenly scattered in the BEV grid of the point cloud branch through the PointPillarsScatter module in the point cloud branch to obtain a uniform feature map. The SECOND and SECOND FPN modules in the point cloud branch are used to encode the preliminary point cloud BEV feature map obtained by the PointPillarsScatter module to obtain the final point cloud BEV feature map where N point represents the number of point clouds, and P num represents the number of point cloud columns.

[0041] Preferably, step (2-5) is specifically as follows. First, the places where the labels are less than M are filled with empty sets, where y i = (c i , V i , Γ i ), c i , V i , Γ i respectively represent the i-th instance category label, the point set label, and the point set equivalent permutation. Then, the predicted values are matched with the label sheets by the Hungarian method to obtain a permutation method that minimizes the loss Finally, the predicted values and the label sheets are used to calculate the loss according to the permutation calculated above to obtain the final total loss of the model. The loss of the model includes three parts: classification loss, point loss, and orientation loss. The three losses have different weights, and the total loss is expressed as follows:

[0042] L = λlcls +αL p2p +βL dir

[0043]

[0044]

[0045]

[0046] where L cls ,L p2p ,L dir represent classification loss, point loss, and direction loss respectively. In the direction loss

[0047] The parameter settings of the high-precision map construction model during training are as follows:

[0048] The range for the high-precision map construction model to construct the high-precision map is set to [-15.0, -30.0, -2.0, 15.0, 30.0, 2.0], that is, a range of 30 meters in the front and back and 15 meters on the left and right centered on the vehicle body;

[0049] The generated BEV resolution in the high-precision map construction model is set to 200×100; the weights of the model's classification loss, point loss, and direction loss are 2, 5, and 0.005 respectively; where α = 0.25 and γ = 2.0 in the classification loss Focal Loss;

[0050] The Batch Size used for training is 16; the initial learning rate is 2e-4 and a cosine annealing learning strategy is adopted;

[0051] A total of 50 Epochs are trained;

[0052] The model optimizer is AdamW.

[0053] According to another aspect of the present invention, a high-precision map construction system based on the fusion of lidar and camera is provided, including:

[0054] A first module for obtaining multi-view images, internal and external parameters of the camera, and point cloud data at the same timestamp from multiple perspective cameras, and respectively preprocessing the obtained multi-view and point cloud data to obtain preprocessed multi-view images and point cloud data.

[0055] A second module for inputting all the point cloud data, multi-view images, and the projection matrix corresponding to the multi-view images preprocessed by the first module into a pre-trained multi-modal high-precision map construction model to obtain prediction results corresponding to different road categories, and the prediction results include coordinate point sets of lane lines, road edges, and sidewalks.

[0056] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention can achieve the following beneficial effects:

[0057] (1) In the present invention, by using the data of two sensors, namely a camera and a lidar, in step (1) and the multi-modal fusion module in step (2), the complementary advantages of the camera and the lidar are utilized, and the image data generated by the camera and the point cloud data generated by the lidar are appropriately fused together to generate a feature map with richer information. Therefore, the technical problem that the existing single-modal method can only use one sensor to online construct a high-precision map and cannot utilize the complementary advantages between sensors can be solved. Compared with the single-modal high-precision map online construction method, the performance is higher, and the optimal performance of the current high-precision map online construction method is achieved on the Nuscenes dataset;

[0058] (2) In the present invention, due to the adoption of the perspective conversion module in the image branch of step (2), the conversion from the PV feature map to the BEV feature map is regarded as the conversion between columns, which reduces the learning difficulty of the model and the computational amount of the model while;

[0059] (3) In the present invention, due to the adoption of the multi-modal fusion module in step (2), in which a method for aligning the image BEV feature map and the point cloud BEV feature map is designed, and the image BEV feature map and the point cloud BEV feature map are aligned and then fused. Therefore, the technical problem that the existing multi-modal high-precision map online construction method simply stitches along the channel dimension when fusing multiple modal BEV feature maps, which will cause misalignment between different modal data and hinder the model from making correct predictions, can be solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a flowchart of the high-precision map construction method based on the fusion of lidar and camera of the present invention;

[0061] Figure 2 is a schematic structural diagram of the multi-modal high-precision map construction model of the present invention.

[0062] Figure 3 is a schematic structural diagram of the image branch in the multi-modal high-precision map construction model of the present invention.

[0063] Figure 4 is a schematic structural diagram of the point cloud branch in the multi-modal high-precision map construction model of the present invention.

[0064] Figure 5 and Figure 6 is a schematic diagram of the multi-modal fusion module in the multi-modal high-precision map construction model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0065] In order to make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0066] In order to utilize the advantages between different sensors, the present invention designs a high-precision map construction model based on the fusion of lidar and camera, and respectively designs a point cloud branch for lidar point cloud data and an image branch for camera image data to process data of two modalities. Finally, a multi-modal fusion module is designed to fuse information of the two modalities to achieve a better detection effect.

[0067] The present invention gives full play to the advantages of multi-sensor complementarity and uses Transformer to regard the conversion of PV feature maps to BEV feature maps as a conversion between columns, which has a stronger construction ability compared with the above-mentioned linear layer method, and has the advantages of being easier to learn and having a smaller computational amount compared with the above-mentioned Transformer graph-to-graph conversion.

[0068] As Figure 1 shown, the present invention provides a high-precision map construction method based on the fusion of lidar and camera, including the following steps:

[0069] (1) Obtain multi-view images, internal and external parameters of the camera, and point cloud data at the same timestamp from multiple perspective cameras, and respectively preprocess the obtained multi-view images and point cloud data to obtain preprocessed multi-view image data and point cloud data. Finally, obtain the internal and external parameters of the camera, the ego vehicle calibration parameters, and the preprocessed point cloud data and multi-view image data.

[0070] In this step, the internal and external parameters of the camera include the camera internal parameters required for the projection conversion from the camera coordinate system to the image coordinate system, and the external parameters required for the conversion from the camera coordinate system to the ego vehicle coordinate system.

[0071] In this step, the preprocessing of the multi-view images includes a scaling operation and a normalization operation. The scaling operation uses bilinear interpolation to scale the multi-view images from the original size to 448×800×3, which can reduce the computational amount of the network model. The purpose of the normalization operation is to remove the average brightness value in the multi-view images after the scaling operation, which can more prominently highlight the individual differences between samples. The preprocessing operation of the point cloud data is specifically to screen out the point cloud data within the detection range, where the detection range is a rectangular area centered on the vehicle body, 30 meters in the front and back, and 15 meters on the left and right.

[0072] (2) Input all the preprocessed point cloud data, multi-view images, and the projection matrices corresponding to the multi-view images in step (1) into a pre-trained multi-modal high-precision map construction model to obtain prediction results corresponding to different road categories. The prediction results include the coordinate point sets of lane lines, curbs, and sidewalks.

[0073] As Figure 2 shown, the high-precision map construction model of the present invention includes an image branch, a point cloud branch, a multi-modal fusion module, and a detection head network implemented based on Transformer, which are sequentially connected. Among them, the image branch and the point cloud branch process image data and point cloud data in parallel. The obtained image features and point cloud features are fused in the multi-modal fusion module and finally sent to the detection head network to predict the high-precision map.

[0074] The image branch includes an image encoder for extracting image features and a perspective conversion module for converting a perspective (PV) feature map into a bird's-eye view (BEV) feature map. Its input is multi-view images, and the output is a set of PV feature maps corresponding to the multi-view images and a complete image BEV feature map.

[0075] The image encoder is composed of a backbone network Efficientnet and a neck network FPN. Its input is a multi-view image I = {I1, I2,..., I N} taken by N cameras at the same time, where I i ∈R 3×H×W , and the output is a set of PV feature maps corresponding to the multi-view images F i,u represents the PV feature map of the u-th scale corresponding to the image taken by the i-th camera, where i ∈ [1, N], u ∈ [1, 3], H pv,u , W pv,u , and K respectively represent the height, width, and number of channels of the u-th scale image feature map in the PV feature map F i,u corresponding to the image taken by the i-th camera. W represents the width of the multi-view image, H represents the height of the multi-view image, and the downsampling multiples of the finally obtained multi-scale PV image feature map F i,u are 1 / 16, 1 / 32, and 1 / 32 respectively.

[0076] As Figure 3 shown, the perspective conversion module (hereinafter referred to as the PV2BEV module) includes two layers of Transformer decoder layers and a sampling layer. Its input is a set of PV feature maps {F 1,u , F 2,u ,..., F N,uThe set of image column feature {F} obtained by deforming column by column 1,u,w , F 2,u,w ,..., F N,u,w}, the set of image column position encodings and the set of position encodings of BEV polar rays are output as a complete image BEV feature map Among them C bev represents the number of channels of the complete image BEV feature map, H bev represents the height of the complete image BEV feature map, and W bev represents the width of the complete image BEV feature map.

[0077] Specifically, for the PV feature map F of the u-th scale corresponding to the image captured by the i-th camera i,u , the view transformation module takes its corresponding as Q, K, and V respectively for cross-attention calculation to obtain the polar ray feature corresponding to this PV feature map Then, is concatenated along the azimuth axis to obtain the BEV feature map corresponding to this PV feature map in the polar coordinate system Finally, the BEV feature maps corresponding to the PV feature maps of all the images captured by the cameras are concatenated and sampled to obtain the complete image BEV feature map in the Cartesian coordinate system In this module, the conversion from the PV feature map to the BEV feature map is regarded as the conversion between columns, which not only reduces the learning difficulty of the model but also reduces the computational amount of the model.

[0078] The structure of the point cloud branch is as Figure 4 shown. It allocates the point cloud data to the corresponding point cloud columns in the voxelization layer according to the given detection range and voxel resolution, and then encodes the point cloud data in each point cloud column using the HardVFE method to obtain the point cloud column feature Then, all the point cloud columns are evenly scattered in the BEV grid of the point cloud branch through the PointPillarsScatter module in the point cloud branch to obtain a preliminary point cloud BEV feature map. Finally, the preliminary point cloud BEV feature map obtained by the PointPillarsScatter module is encoded using the SECOND and SECOND FPN modules in the point cloud branch to obtain the final point cloud BEV feature map Among them, N point represents the number of point clouds, and P num represents the number of point cloud columns.

[0079] Multimodal fusion modules such as Figure 4 As shown, its input is the PV feature map set corresponding to the multi-view image output by the image branch {F 1,u ,F 2,u ,...,F N,u} and the complete image BEV feature map And the point cloud BEV feature map output by the point cloud branch The output is a fused BEV feature map that combines image information and point cloud information.

[0080] Specifically, the multimodal fusion module first converts the PV feature map set {F 1,u ,F 2,u ,...,F N,u} and the point cloud BEV feature map output by the point cloud branch Perform cross attention calculation to obtain the interactive features of the two Then the interaction feature And point cloud BEV feature map Splice along the channel dimension to obtain the enhanced point cloud BEV feature map (like Figure 5 As shown in Figure 2, Channel Concatenate refers to concatenation along the channel dimension), and then the enhanced BEV features are BEV feature map with complete image Splice along the channel dimension to obtain the feature map of the two spliced ​​together Then, the concatenated feature map The image is fed into a convolution unit (the convolution unit consists of a 3×3 convolution, a BatchNorm layer, and a ReLu layer) and a 1×1 convolution block to obtain the image BEV feature map. The offset Then, according to the offset Δ, the BEV feature map of the image is reconstructed. Sampling in order to obtain the point cloud BEV feature map Aligned image BEV feature map Finally, the image BEV feature map BEV feature map with point cloud After concatenation along the channel dimension, it is sent to a 1×1 convolution block for dimensionality reduction to obtain the final fused BEV feature map. This module helps the model make full use of information from both modalities to build more accurate high-precision maps.

[0081] The input of the detection head network is the fused BEV feature map output by the multimodal fusion module. and a query vector responsible for querying high-precision map instances where M represents the number of query vectors, and d model represents the dimension of the query vector, and the output is the prediction results of M high-precision map instances.

[0082] Specifically, the detection head network takes the query vector and the fused BEV feature map to perform deformable attention calculation to obtain the prediction results of M high-precision map instances where respectively represent the class confidence of the i-th high-precision map instance and the coordinate point set of the high-precision map instance. Among the prediction results of all high-precision map instances, the instances with a confidence greater than the set threshold retain their instance coordinate point sets, while those less than the threshold are discarded. The instance point sets that are retained are concatenated to obtain the complete vectorized representation of the high-precision map.

[0083] Specifically, the high-precision map construction model of the present invention is trained through the following steps:

[0084] (2-1) Obtain a high-precision map autonomous driving dataset, which includes multi-view images, point cloud data, and camera internal and external parameter matrices at the same timestamp. Preprocess the obtained high-precision map autonomous driving dataset to obtain the preprocessed high-precision map autonomous driving dataset, and divide it into a training set and a validation set according to a ratio.

[0085] Specifically, the high-precision map autonomous driving dataset used in this step is the nuScenes dataset. The nuScenes dataset is collected in a total of four regions: the Boston Seaport area, Queenstown in Singapore, Yibei, and Holland Village. It contains a total of 1000 autonomous driving scenes, including rainy days, nights, and foggy days, etc. During training and testing, the official dataset division method is used, and it is divided into a training set and a test set according to a ratio of 4.7:1, that is, the training set has a total of 28,130 timestamp samples, and the validation set has a total of 6,019 timestamp samples. Each timestamp includes image data of 6 cameras. Calculate the camera internal and external parameters and the predefined bird's-eye view space coordinate system to obtain the projection matrix corresponding to each image data. In addition, this step only focuses on three road information: lane lines, sidewalks, and curbs.

[0086] It should be noted that the multi-view image preprocessing method and point cloud data preprocessing method used in this step are exactly the same as those in the above step (1), so they will not be elaborated here.

[0087] (2-2) Input the multi-view images and point cloud data in the training set obtained in step (2-1) into the image branch and the point cloud branch of the high-precision map construction model respectively, so as to obtain the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images, the image BEV feature map , and the point cloud BEV feature map

[0088] This step (2-2) includes the following sub-steps.

[0089] (2-2-1) Input the multi-view images in the training set obtained in step (2-1) into the image branch, and output the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images and the image BEV feature map

[0090] This step (2-2-1) includes the following sub-steps:

[0091] (2-2-1-1) Input the multi-view images into the shared multi-view image encoder to obtain the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images.

[0092] (2-2-1-2) Input the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images obtained in step (2-2-1-1) into the view transformation module of the image branch to obtain the image BEV feature map

[0093] This step (2-2-1-2) includes the following sub-steps:

[0094] (2-2-1-2-1) First, use the position information of the feature in the image and the BEV polar ray, the position information of a certain column in the image in the image, and the position information of the BEV polar ray in the BEV to encode the position encoding required by the view transformation module.

[0095] The present invention encodes the two kinds of position information using sine and cosine functions as follows:

[0096]

[0097]

[0098] where pos represents the mathematical representation of two kinds of position information, d model represents the feature dimension, and i represents the i-th dimension of the feature. If the height of a column in the image feature is H, then the mathematical representation pos of the first position encoding position information is pos = [1, 2, 3..., H]. The mathematical representation of the second position encoding position information is based on the imaging principle of the camera. Let (x C , y C , z C ) be a point in the camera coordinate system, and (x I , y I ) be the corresponding coordinates in the pixel system. According to the camera imaging principle, the conversion relationship between the two can be described as:

[0099]

[0100] where f x , f y , u0 and v0 are parameters in the camera internal parameter matrix. Then the mathematical representation pos of the second position encoding position information is calculated as:

[0101]

[0102] (2-2-1-2-2) After obtaining the position encoding in step (2-2-1-2-1), the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images is input into the PV2BEV module to obtain the BEV feature map in order to obtain the epipolar ray feature

[0103] Before input, the operations on the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images and the position encoding are the same as those described in the view transformation module in step (2).

[0104] As Figure 3 shown, the PV2BEV module is constructed by two Transformer decoder layers to establish the conversion relationship between the columns of the image and the BEV epipolar rays. Let represent the w-th column of the u-th scale of the i-th image. Using the position encoding method introduced above, the position encodings of the columns of the image and the epipolar rays can be obtained where H pv,u and R u represent the height of the image column and the length of the epipolar ray respectively. Then the conversion relationship between the columns of the image and the BEV epipolar rays can be expressed as:

[0105]

[0106] wherein

[0107]

[0108] wherein are mapping parameters, d q = d k = d v = d model / h, d represents the feature dimension, and h represents the number of heads in multi-head attention.

[0109] (2-2-1-2-3) splices the polar ray features obtained in (2-2-1-2-2) along the azimuth axis to obtain the BEV feature map in the polar coordinate system converted from the feature map of the u-th scale of the i-th image:

[0110]

[0111] The BEV feature maps corresponding to the PV feature maps of all the images captured by the cameras are spliced and sampled to obtain the complete image BEV feature map in the Cartesian coordinate system

[0112] (2-2-2) According to the given detection range and voxel resolution, the point cloud data is assigned to the corresponding point cloud pillars in the voxelization layer through the point cloud branch, and the point cloud data in each point cloud pillar is encoded using the HardVFE method to obtain the point cloud pillar features All the point cloud pillars are evenly dispersed in the BEV grid of the point cloud branch through the PointPillarsScatter module in the point cloud branch to obtain a uniform feature map, and the preliminary point cloud BEV feature map obtained by the PointPillarsScatter module is encoded using the SECOND and SECOND FPN modules in the point cloud branch to obtain the final point cloud BEV feature map where N point represents the number of point clouds, and P num represents the number of point cloud pillars.

[0113] (2-3) Using the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images obtained in step (2-) to enhance the point cloud BEV feature map to obtain the enhanced point cloud BEV feature map The enhanced point cloud BEV feature map and the image BEV feature map Concatenate along the channel dimension to obtain the final fused feature map

[0114] Specifically, in the method of enhancing the point cloud BEV, it is first necessary to find the correspondence between each feature in the point cloud BEV feature and the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} of the PV features corresponding to the multi-view images. The present invention constructs a pixel-to-pixel mapping relationship between the PV features in the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the point cloud BEV feature map and the perspective images. Specifically, let (i p , j p ) be a pixel point in the point cloud BEV feature map. The point cloud {(x, y, z)} within the point cloud column located at (i p , j p ) is projected onto the pixel system using the camera intrinsic matrix and the extrinsic matrix to obtain the pixel points {(i, j)} on the PV features in the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images. This mapping from the point cloud BEV feature to the PV features in the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} can be expressed as:

[0115] M p→c (i p , j p ) = {(i, j)}

[0116] Secondly, after finding the correspondence, it is necessary to perform attention calculation on the corresponding features to obtain the interaction features. Specifically, let be the feature of the point cloud BEV at (i p , j p ). According to the mapping relationship between the point cloud BEV feature and the PV features in the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images, the set of PV feature maps

[0117] {F 1,u , F 2,u ,..., F N,u} of the PV features such as Figure 6As shown, taking as q, and taking as k, v, the interaction between the point cloud BEV feature and the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images can be expressed as follows:

[0118]

[0119] Through the interaction between the point cloud BEV feature and the set of PV feature maps

[0120] {F 1,u , F 2,u ,..., F N,u}, the interaction feature Finally, the interaction feature and the point cloud BEV feature are concatenated along the channel dimension and then dimension-reduced to obtain the enhanced point cloud BEV feature

[0121] After obtaining the enhanced point cloud BEV feature , it is concatenated with the image BEV feature along the channel dimension to obtain the final fused BEV feature map The specific method is as Figure 5 shown. First, is concatenated with along the channel dimension to obtain Then, after passing through a convolutional unit and a 1×1 convolution, the offset of the image BEV feature map is obtained Next, according to the offset, the image BEV feature map is resampled from the image BEV feature map again to obtain the image BEV feature map aligned with the point cloud BEV feature map The sampling process can be expressed as follows:

[0122]

[0123] where F′ wh represents the feature of the offset image BEV feature map at (w, h), and F w′h′ represents the feature of the image BEV feature map before offset at (w′, h′). Among them, Δ 1wh , Δ 2wh represents the offset predicted by the network for the image BEV feature map at the (w, h) position. After obtaining the aligned image BEV feature map, it is concatenated with the point cloud BEV feature map along the channel dimension and dimension-reduced using a 1×1 convolution to obtain the fused BEV feature map

[0124] (2 - 4) Input the fused BEV feature map obtained in step (2 - 3) into the detection head network to obtain the prediction results of multiple high - definition map instances where

[0125] respectively represent the class confidence of the i - th high - definition map instance and the coordinate point set of the high - definition map instance.

[0126] For the specific structure of the detection head network in the present invention, please refer to the article "MAPTR: STRUCTURED MODELING AND LEARNING FOR ONLINE VECTORIZED HD MAP CONSTRUCTION".

[0127] (2 - 5) Utilize the prediction results of the high - definition map construction model obtained in (2 - 4) and the high - definition map label to obtain the loss function of the high - definition map construction model;

[0128] Specifically, first fill the places where the label is less than M with empty sets, where y i =(c i , V i , Γ i ), c i , V i , Γ i respectively represent the i - th instance class label, point set label, and point set equivalent permutation. Then, perform Hungarian matching between the predicted value and the label paper to obtain an arrangement method that minimizes the loss Finally, calculate the loss between the predicted value and the label paper according to the arrangement calculated above to obtain the final total loss of the model. The loss of the model includes three parts: classification loss, point loss, and direction loss. The three losses have different weights, and the total loss is expressed as follows:

[0129] L = λL cls +αL p2p +βL dir

[0130]

[0131]

[0132]

[0133] where L cls , L p2p , L dir ​respectively represent the classification loss, the point loss, and the direction loss. In the direction loss

[0134] (2-6) Use the loss function obtained in step (2-5) to iteratively train the high-precision map construction model until the high-precision map construction model converges, thereby obtaining a trained high-precision map construction model;

[0135] Among them, the parameter settings of the high-precision map construction model during training are as follows: the range for the model to construct the high-precision map is set to [-15.0, -30.0, -2.0, 15.0, 30.0, 2.0], that is, a range of 30 meters before and after and 15 meters left and right centered on the vehicle body; the generated BEV resolution in the model is set to 200×100; the weights of the model classification loss, point loss, and direction loss are 2, 5, and 0.005 respectively; among them, α = 0.25 and γ = 2.0 in the classification loss Focal Loss; the applicable Batch Size for training is 16; the initial learning rate is 2e-4 and a cosine annealing learning strategy is adopted; a total of 50 Epochs are trained; the model optimizer is AdamW.

[0136] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A high-precision map construction method based on the fusion of lidar and camera, characterized in that Including the following steps: (1) Obtain multi-view images, the internal and external parameters of the cameras, and point cloud data at the same timestamp from multiple perspective cameras, and preprocess the obtained multi-view and point cloud data respectively to obtain preprocessed multi-view images and point cloud data; (2) Input all the preprocessed point cloud data, multi-view images, and the projection matrices corresponding to the multi-view images in step (1) into a pre-trained multi-modal high-precision map construction model to obtain prediction results corresponding to different road categories, and the prediction results include coordinate point sets of lane lines, curbs, and sidewalks; the high-precision map construction model includes an image branch, a point cloud branch, a multi-modal fusion module, and a detection head network implemented based on Transformer connected in sequence, where the image branch and the point cloud branch process image data and point cloud data in parallel, and the obtained image features and point cloud features are fused in the multi-modal fusion module and finally sent to the detection head network to predict the high-precision map; The image branch includes an image encoder for extracting image features and a perspective transformation module PV2BEV for converting a perspective PV feature map into a bird's-eye view BEV feature map. The input of the image branch is the multi-view image, and the output is a set of PV feature maps corresponding to the multi-view image and a complete image BEV feature map; The image encoder includes a backbone network Efficientnet and a neck network FPN. Its input is the multi-view image I = {I1, I2,..., I N}, where I i ∈ R 3×H×W . Its output is the set of PV feature maps {F 1,u , F 2, u,..., F N,u} corresponding to the multi-view image. F i,u represents the PV feature map of the u-th scale corresponding to the image captured by the i-th camera, where i ∈ [1, N], u ∈ [1, 3], H pv,u , W pv,u , and K respectively represent the height, width, and number of channels of the u-th scale image feature map in the PV feature map F i,u corresponding to the image captured by the i-th camera. W represents the width of the multi-view image, and H represents the height of the multi-view image. The downsampling multiples of the finally obtained multi-scale PV image feature map F i,u are 1 / 16, 1 / 32, and 1 / 32 respectively; The PV2BEV module includes two layers of Transformer decoder layers and a sampling layer. Its input is the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images output by the image encoder, which are deformed column by column to obtain the set of image column features {F 1,u,w , F 2,u,w ,..., F N,u,w}, the set of image column position encodings and the set of position encodings of BEV polar rays The output is the complete image BEV feature map Among them C bev represents the number of channels of the complete image BEV feature map, H bev represents the height of the complete image BEV feature map, and W bev represents the width of the complete image BEV feature map; For the PV feature map F at the u-th scale corresponding to the image captured by the i-th camera i,u the PV2BEV module performs cross-attention calculations using the corresponding as Q, K, and V respectively to obtain the epipolar ray features corresponding to this PV feature map Then, they are concatenated along the azimuth axis to obtain the BEV feature map corresponding to this PV feature map in the polar coordinate system Finally, the BEV feature maps corresponding to the PV feature maps of all the images captured by the cameras are concatenated and sampled to obtain the complete image BEV feature map in the Cartesian coordinate system 2. The method for constructing a high-precision map based on the fusion of lidar and camera according to claim 1, wherein The internal and external parameters of the camera include the camera internal parameters required for the projection conversion from the camera coordinate system to the image coordinate system and the external parameters required for the conversion from the camera coordinate system to the ego-vehicle coordinate system; The preprocessing of the multi-view image in step (1) includes a scaling operation, which uses bilinear interpolation to scale the multi-view image from the original size to 448×800×3; The preprocessing operation of the point cloud data in step (1) is to filter out the point cloud data within the detection range, where the detection range is a rectangular area centered on the vehicle body, 30 meters in the front and back, and 15 meters on the left and right.

3. The method for constructing a high-precision map based on the fusion of lidar and camera according to claim 2, wherein The point cloud branch assigns the point cloud data to the corresponding point cloud pillars in the voxelization layer according to the given detection range and voxel resolution, and then encodes the point cloud data in each point cloud pillar using the HardVFE method to obtain the point cloud pillar features Then, all the point cloud pillars are evenly scattered in the BEV grid of the point cloud branch through the PointPillarsScatter module in the point cloud branch to obtain a preliminary point cloud BEV feature map. Finally, the preliminary point cloud BEV feature map obtained by the PointPillarsScatter module is encoded using the SECOND and SECOND FPN modules in the point cloud branch to obtain the final point cloud BEV feature map where N represents the number of point clouds, and P point represents the number of point cloud pillars; num ​ The input of the multi-modal fusion module is the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images output by the image branch and the complete image BEV feature map and the point cloud BEV feature map output by the point cloud branch The output is the fused BEV feature map that fuses image information and point cloud information The multi-modal fusion module first performs cross-attention calculation on the set of PV feature maps {F 1,u , F 2, u,..., F N,u} corresponding to the multi-view images and the point cloud BEV feature map output by the point cloud branch to obtain their interaction features Then, the interaction features and the point cloud BEV feature map are concatenated along the channel dimension to obtain the enhanced point cloud BEV feature map Then, the enhanced BEV features are concatenated with the complete image BEV feature map along the channel dimension to obtain the concatenated feature map of the two Subsequently, the concatenated feature map is sequentially fed into a convolutional unit and a 1×1 convolutional block to obtain the offset of the image BEV feature map Then, according to the offset Δ, resampling is performed again from the image BEV feature map to obtain an image BEV feature map aligned with the point cloud BEV feature map Finally, after concatenating the image BEV feature map and the point cloud BEV feature map along the channel dimension, it is fed into a 1×1 convolutional block for dimensionality reduction processing to obtain the final fused BEV feature map This module helps the model make full use of the information of the two modalities, thereby constructing a more accurate high-precision map.

4. The method for constructing a high-precision map based on the fusion of lidar and camera according to claim 3, wherein The input of the detection head network is the fused BEV feature map output by the multi-modal fusion module and the query vector responsible for querying high-precision map instances where M represents the number of query vectors, and d model represents the dimension of the query vector, and the output is the prediction results of M high-precision map instances; The detection head network will query the vector and the fused BEV feature map to perform deformable attention calculation to obtain the prediction results of M high-precision map instances where respectively represent the class confidence of the i-th high-precision map instance and the coordinate point set of the high-precision map instance.

5. The high-precision map construction method based on the fusion of lidar and camera according to claim 4, characterized in that, The high-precision map construction model is trained through the following steps: (2-1) Obtain a high-precision map autonomous driving dataset, which includes multi-view images, point cloud data, and camera internal and external parameter matrices at the same timestamp; preprocess the obtained high-precision map autonomous driving dataset to obtain a preprocessed high-precision map autonomous driving dataset, and divide it into a training set and a validation set according to a ratio; (2-2) Input the multi-view images and point cloud data in the training set obtained in step (2-1) into the image branch and the point cloud branch of the high-precision map construction model respectively, so as to obtain the PV feature map set {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images, the image BEV feature map and the point cloud BEV feature map (2-3) Use the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images obtained in step (2-2) to enhance the point cloud BEV feature map , so as to obtain the enhanced point cloud BEV feature map Concatenate the enhanced point cloud BEV feature map and the image BEV feature map along the channel dimension to obtain the final fused feature map (2-4) Input the fused BEV feature map obtained in step (2-3) into the detection head network to obtain the prediction results of multiple high-precision map instances where respectively represent the class confidence of the i-th high-precision map instance and the coordinate point set of the high-precision map instance; (2-5) Use the prediction result of the high-precision map construction model obtained in (2-4) and the high-precision map label Obtain the loss function of the high-precision map construction model; (2-6) Use the loss function obtained in step (2-5) to iteratively train the high-precision map construction model until the high-precision map construction model converges, thereby obtaining a trained high-precision map construction model.

6. The high-precision map construction method based on the fusion of lidar and camera according to claim 5, wherein Step (2-2) includes the following sub-steps; (2-2-1) Input the multi-view images in the training set obtained in step (2-1) into the image branch, and output the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images and the image BEV feature map This step (2-2-1) includes the following sub-steps: (2-2-1-1) Input the multi-view images into a shared multi-view image encoder to obtain a set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images; (2-2-1-2) Input the set of PV feature maps {F 1,u , F 2, u,..., F N,u} corresponding to the multi-view images obtained in step (2-2-1-1) into the view transformation module of the image branch to obtain the image BEV feature map This step (2-2-1-2) includes the following sub-steps: (2-2-1-2-1) First, use the position information of the features in the image and the BEV polar rays, the position information of a certain column in the image in the image, and the position information of the BEV polar rays in the BEV to encode the position encoding required by the view transformation module; (2-2-1-2-2) After obtaining the position encoding in step (2-2-1-2-1), the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images is input into the PV2BEV module together to obtain the BEV feature map, so as to obtain the epipolar ray feature (2-2-1-2-3) The polar ray features obtained from (2-2-1-2-2) are stitched along the azimuth axis to obtain the BEV feature map in the polar coordinate system converted from the feature map of the u-th scale of the i-th image: The BEV feature maps corresponding to the PV feature maps of the images captured by all cameras are stitched and sampled to obtain a complete image BEV feature map in the Cartesian coordinate system (2-2-2) The point cloud data is assigned to the corresponding point cloud pillars in the voxelization layer according to the given detection range and voxel resolution by the point cloud branch, and the HardVFE method is used to encode the point cloud data in each point cloud pillar to obtain the point cloud pillar features All the point cloud pillars are evenly scattered in the BEV grid of the point cloud branch through the PointPillarsScatter module in the point cloud branch to obtain a uniform feature map, and the SECOND and SECOND FPN modules in the point cloud branch are used to encode the preliminary point cloud BEV feature map obtained by the PointPillarsScatter module to obtain the final point cloud BEV feature map Among them, N represents the number of point clouds, and P point represents the number of point cloud pillars. num ​ 7. The high-precision map construction method based on lidar and camera fusion according to claim 6, characterized in that Step (2-5) specifically is to first fill the places where the labels are less than M with empty sets, where y i =(c i , V i , Γ i ), c i , V i , Γ i respectively represent the i-th instance category label, point set label, and point set equivalent permutation. Then, the predicted value is matched with the label paper using the Hungarian algorithm to obtain an arrangement that minimizes the loss Finally, the predicted value and the label paper are used to calculate the loss according to the arrangement obtained above to obtain the final total loss of the model. The loss of the model consists of three parts: classification loss, point loss, and orientation loss. The three losses have different weights, and the total loss is expressed as follows: L = λL cls + αL p2p + βL dir where L cls , L p2p , L dir represent the classification loss, point loss, and direction loss respectively; in the direction loss The parameter settings of the high-precision map construction model during training are as follows: The range for the high-precision map construction model to construct the high-precision map is set to [-15.0, -30.0, -2.0, 15.0, 30.0, 2.0], that is, a range of 30 meters before and after and 15 meters left and right centered on the vehicle body; The BEV resolution generated in the high-precision map construction model is set to 200×100; the weights of the model classification loss, point loss, and direction loss are 2, 5, and 0.005 respectively; among them, α = 0.25 and γ = 2.0 in the classification loss Focal Loss; The Batch Size used for training is 16; the initial learning rate is 2e-4 and a cosine annealing learning strategy is adopted; A total of 50 Epochs are trained; The model optimizer is AdamW.

8. A high-precision map construction system based on the fusion of lidar and camera, characterized in that, Including: The first module is used to obtain multi-view images, the internal and external parameters of the camera, and the point cloud data at the same timestamp from multiple perspective cameras, and preprocess the obtained multi-view and point cloud data respectively to obtain the preprocessed multi-view images and point cloud data; The second module is used to input all the preprocessed point cloud data, multi-view images, and the projection matrices corresponding to the multi-view images in the first module into a pre-trained multi-modal high-precision map construction model to obtain prediction results corresponding to different road categories, and the prediction results include the coordinate point sets of lane lines, curbs, and sidewalks; the high-precision map construction model includes an image branch, a point cloud branch, a multi-modal fusion module, and a detection head network implemented based on Transformer connected in sequence, where the image branch and the point cloud branch process image data and point cloud data in parallel, and the obtained image features and point cloud features are fused in the multi-modal fusion module, and finally sent to the detection head network to predict the high-precision map; The image branch includes an image encoder for extracting image features and a view transformation module PV2BEV for converting the perspective PV feature map into a bird's-eye view BEV feature map. The input of the image branch is the multi-view image, and the output is the set of PV feature maps corresponding to the multi-view image and the complete image BEV feature map; The image encoder includes a backbone network Efficientnet and a neck network FPN. Its input is the multi-view image I = {I1, I2,..., I N}, where I i ∈ R 3×H×W . The output is the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view image. F i,u represents the PV feature map of the u-th scale corresponding to the image captured by the i-th camera, where i ∈ [1, N], u ∈ [1, 3], and H pv,u , W pv,u , and K respectively represent the height, width, and number of channels of the u-th scale image feature map in the PV feature map F i,u corresponding to the image captured by the i-th camera. W represents the width of the multi-view image, and H represents the height of the multi-view image. The downsampling multiples of the finally obtained multi-scale PV image feature map F i,u are 1 / 16, 1 / 32, and 1 / 32 respectively. The PV2BEV module includes two layers of Transformer decoder layers and a sampling layer. Its input is the set of PV feature maps {F 1,u , F 2,u ,..., F N,u} corresponding to the multi-view images output by the image encoder, which are deformed column by column to obtain the set of image column feature maps {F 1,u,w , F 2,u,w ,..., F N,u,w}, the set of image column position encodings and the set of position encodings of BEV polar rays The output is the complete image BEV feature map where C bev represents the number of channels of the complete image BEV feature map, H bev represents the height of the complete image BEV feature map, and W bev represents the width of the complete image BEV feature map; For the PV feature map F at the u-th scale corresponding to the image captured by the i-th camera i,u the PV2BEV module takes its corresponding as Q, K, and V respectively for cross-attention calculation to obtain the epipolar ray feature corresponding to this PV feature map Then, they are concatenated along the azimuth axis to obtain the BEV feature map corresponding to this PV feature map in the polar coordinate system Finally, the BEV feature maps corresponding to the PV feature maps of all the images captured by the cameras are concatenated and sampled to obtain the complete image BEV feature map in the Cartesian coordinate system

Citation Information

Patent Citations

  • Transform-based high-precision map real-time prediction method and system

    CN116071721A