High-precision semantic map construction method and system based on strong coupling network

The strong coupling network addresses limitations in existing methods by enhancing modality fusion and element interactions, achieving high-precision semantic maps with expanded range and efficient storage through a semantic-geometric and point-element coupling module.

CN120318442APending Publication Date: 2025-07-15Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510381351.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the existing high-precision semantic map construction methods, multimodal fusion fails to capture the synergies and differences between different modes, resulting in limited scope; classification and positioning based on point information ignore the interaction between elements and points, resulting in low accuracy; modeling strategies based on uniform points are difficult to balance computational complexity and construction accuracy, resulting in computing/storage redundancy.

Method used

Using a method based on strong coupling network, the semantic-geometric coupling module and point-element coupling module are used to strongly couple semantic and geometric information from the camera and LiDAR, respectively, and long-distance BEV features are generated, and high-precision classification and positioning are achieved through three-level interactive units. Map elements are modeled as ordered key points.

Benefits of technology

It realizes high-precision semantic map construction, expands the perception range, improves the classification and positioning accuracy of map elements, and realizes efficient storage and calculation through the modeling of ordered key points.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318442A_ABST
    Figure CN120318442A_ABST
Patent Text Reader

Abstract

The invention discloses a high-precision semantic map construction method and system based on a strong coupling network, and the method comprises the steps: 1, inputting a camera image and a LiDAR point cloud into a preset semantic-geometric coupling module, and obtaining a fused aerial view feature; 2, inputting the fused aerial view features and preset learnable point query and element query into a point-element coupling module to obtain combined features; and 3, constructing a high-precision semantic map according to the combined features. According to the method, the synergistic effect and difference between different modalities are captured, the interaction between elements and points is concerned, and efficient storage and calculation are realized by modeling map elements into ordered key points.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of high-precision semantic map construction, and particularly to a high-precision semantic map construction method and system based on a strongly coupled network. Background Art

[0002] High-precision semantic maps (HD Maps) play a crucial role in the safety, reliability, and intelligence of autonomous driving systems, especially in perception, prediction, and planning. High-precision semantic maps are composed of various map elements, such as road boundaries, lane dividers, and crosswalks, providing comprehensive geometric and semantic information about the surrounding environment. High-precision semantic map construction is formulated as the classification and localization of these map elements in a bird's-eye view (BEV).

[0003] Traditional high-precision map construction relies on time-consuming and labor-intensive manual annotation. With the rapid development of deep learning and the great interest in bird's-eye view (BEV), constructing local high-precision maps using data from on-vehicle sensors is becoming a promising solution. Most existing works regard high-precision map construction as a semantic segmentation task, which rasterizes the high-precision map and assigns a label to each pixel. However, the raster is not an ideal representation. It contains redundant information for each pixel and requires a large amount of storage, especially when the map has a wide range. It assumes independence between pixels and elements, resulting in broken or missing shapes. In addition, vectorization requires complex post-processing, increasing additional computational complexity and cumulative errors.

[0004] Recently, end-to-end vectorized high-precision map construction methods, which model map elements in the form of a unified ordered point set, have received increasing attention and achieved remarkable success. However, this strategy is difficult to balance computational complexity and construction accuracy. Current methods only use a single modality for bird's-eye view (BEV) feature generation, and the limited perception ability results in a limited map range. In addition, using only point information for classification and localization makes it difficult to handle element failures, such as incorrect element shapes or entanglements between elements, resulting in low construction accuracy.

[0005] The main workflow of the deep learning-based high-definition map construction method is divided into two steps: the generation of bird's-eye view (BEV) features and the classification and localization of map elements.

[0006] The generation of BEV features aims to transform data from different perspectives into BEV and extract features in the BEV space. According to the modalities used, existing methods are divided into single-modal and multi-modal fusion methods. For cameras, the main challenge lies in the transformation from the perspective view (PV) to BEV. Currently, the main methods are based on geometric projection, explicitly estimating the depth distribution to obtain frustum features, and assigning them to the BEV space grid through the internal and external camera parameters. With the widespread popularity of Transformers, researchers have proposed Transformer-based methods that implicitly utilize the prior geometry of the camera to map the transformation from PV to BEV. Camera-based BEV features have rich semantic information, but due to the lack of precise 3D geometric information, the larger the map range, the lower the construction accuracy. For LiDAR, transforming PV to BEV is much simpler than for cameras because LiDAR always has precise 3D geometric information. However, there is less BEV generation based solely on LiDAR because its effective perception range is only 30m. Due to the limitations of single-modal perception, multi-modal fusion has become a trend in recent years, and the most common method is to directly connect the BEVs of the two modalities. However, these simple fusion strategies fail to consider the synergy and differences between different modalities.

[0007] The classification and localization of map elements aim to determine the shape, location, and category of each map element. Based on the representation of map elements, the construction of high-precision semantic maps can be divided into rasterized high-precision semantic map estimation and vectorized high-precision semantic map construction. Rasterized high-precision semantic map estimation is defined as a segmentation task of predicting the category of each pixel. The core innovation is the generation of BEV features. The core of vectorized high-precision semantic map construction is the localization of map elements in the BEV space, represented by a series of ordered point sets.

[0008] In summary, the existing methods for constructing high-precision semantic maps face the following challenges: (1) Multi-modal fusion based on direct stitching fails to capture the synergy and differences between different modalities, resulting in limited range; (2) Classification and localization based on point information ignore the interaction between elements and points, resulting in low accuracy; (3) The modeling strategy based on uniform points is difficult to balance computational complexity and construction accuracy, resulting in computational / storage redundancy. Summary of the Invention

[0009] In order to at least partially solve the problems in the construction of existing high-precision semantic maps. The direct splicing of multimodal fusion cannot capture the synergistic effects and differences between different modalities, resulting in limited scope. The classification and positioning based on point information ignore the interaction between elements and points, resulting in low accuracy. And the modeling strategy based on uniform points is difficult to balance computational complexity and construction accuracy, resulting in computational / storage redundancy. The present invention provides a method and system for constructing a high-precision semantic map based on a strongly coupled network. The present invention proposes a semantic-geometry coupling module and a point-element coupling module, and uses the semantic-geometry coupling module and the point-element coupling module as a strongly coupled network (SuperMapNet) for constructing a high-precision semantic map. The core of the present invention is to strongly couple semantic and geometric information from cameras and LiDAR through the collaborative enhancement unit and the difference alignment unit in the semantic-geometry coupling module to generate long-range BEV features, which can capture the synergistic effects and differences between different modalities and expand the perception ability. And through the three-level interaction of the point-to-point interaction unit, element-to-element interaction unit, and point-to-element interaction unit in the point-element coupling module, local and global features are coupled from points and elements respectively to achieve high-precision classification and positioning. In addition, by modeling map elements as ordered key points, efficient storage and calculation are achieved.

[0010] To achieve the above object, the technical solution of the present invention is:

[0011] The first aspect of the present invention proposes a method for constructing a high-precision semantic map based on a strongly coupled network, including:

[0012] Step 1: Input the camera image and LiDAR point cloud into a preset semantic-geometry coupling module to obtain fused bird's-eye view features, which is convenient for enhancing the perception ability of rich information and long distances;

[0013] Step 2: Input the fused bird's-eye view features, a preset learnable point query, and an element query into a point-element coupling module to obtain combined features, which is convenient for enhancing the mutual constraints between points, between elements, and between points and elements;

[0014] Step 3: Construct a high-precision semantic map according to the combined features.

[0015] Further, the semantic-geometry coupling module includes a camera encoder, a 2D to bird's-eye view transformation unit, a lidar encoder, a collaborative enhancement unit, a difference alignment unit, and a connection unit;

[0016] The camera encoder is used to encode the camera image;

[0017] The 2D to bird's-eye view transformation unit is used to transform the output of the camera encoder into camera bird's-eye view features, which is convenient for subsequent collaborative enhancement and difference alignment;

[0018] The LiDAR encoder is used to encode the LiDAR point cloud to obtain LiDAR features;

[0019] The collaborative enhancement unit is used to enhance the collaborative effect of the camera bird's-eye view features and the LiDAR features, obtaining refined camera bird's-eye view features and LiDAR features, facilitating the enhancement of the collaborative effect of the camera bird's-eye view features and the LiDAR features, enriching semantic information and filling in the gaps of the LiDAR features, and at the same time adding accurate 3D geometric information to the camera bird's-eye view features;

[0020] The difference alignment unit is used to perform difference alignment on the refined camera bird's-eye view features and the LiDAR features, facilitating the elimination of the differences between the camera bird's-eye view features and the LiDAR features;

[0021] The connection unit is used to connect the aligned refined camera bird's-eye view features and LiDAR features and perform basic convolution processing to obtain fused bird's-eye view features.

[0022] Furthermore, the collaborative enhancement unit is represented by the following formula:

[0023]

[0024]

[0025] Among them, mod1 is the camera bird's-eye view feature, mod2 is the LiDAR feature, is the complementary information of the LiDAR feature, is the query obtained after the camera bird's-eye view feature undergoes three-layer fully connected processing, is the key obtained after the camera bird's-eye view feature undergoes three-layer fully connected processing, is the value obtained after the camera bird's-eye view feature undergoes three-layer fully connected processing, d k is the dimension of the query and key vectors, is the refined camera bird's-eye view feature, conv is the convolution operation, cat is the connection operation, is the query obtained after the LiDAR feature undergoes three-layer fully connected processing, is the key obtained after the LiDAR feature undergoes three-layer fully connected processing, is the value obtained after the LiDAR feature undergoes three-layer fully connected processing, is the complementary information of the camera bird's-eye view feature.

[0026] Furthermore, the difference alignment unit is used to perform difference alignment on the refined camera bird's-eye view features and the LiDAR features, specifically including:

[0027] After connecting the refined camera bird's-eye view features and lidar features, perform multiple convolutional processes to obtain a difference flow;

[0028] Input the difference flow and the refined camera bird's-eye view features into the transform to obtain the aligned camera bird's-eye view features.

[0029] Further, the point-element coupling module includes multiple point-to-point interaction units, multiple element-to-element interaction units, and a point-to-element interaction unit; among them, multiple said point-to-point interaction units are connected in sequence, the structures of multiple point-to-point interaction units are the same, multiple said element-to-element interaction units are connected in sequence, and the structures of each element-to-element interaction unit are the same;

[0030] The point-to-point interaction unit is used to process the features after connecting the fused bird's-eye view features and a preset learnable point query to obtain a point descriptor, which is convenient for comprehensively learning the external geometric relationship between element points and the internal local coordinate information of each point;

[0031] The element-to-element interaction unit is used to process the features after connecting the fused bird's-eye view features and a preset learnable element query to obtain an element descriptor, which is convenient for comprehensively learning the overall shape and semantic relationship between different elements;

[0032] The point-to-element interaction unit is used to fuse the point descriptor and the element descriptor to obtain a combined feature, which is convenient for using the overall shape and semantic knowledge at the element level to supplement the point-level information only through the local position knowledge, so as to integrate the overall information while considering the details.

[0033] Further, the point-to-point interaction unit is used to process the features after connecting the fused bird's-eye view features and a preset learnable point query to obtain a point descriptor, specifically including:

[0034] Connect a group of point queries in a preset element and input them into a multi-layer perceptron to obtain point query features;

[0035] Multiply the point query features and the fused bird's-eye view features to obtain a point mask;

[0036] Input the point mask, the fused bird's-eye view features, and the point query into a cross-attention sub-module to obtain point features

[0037] Input the point features into a self-attention sub-module and a feed-forward network connected in sequence for processing to obtain a point descriptor.

[0038] Further, the element-to-element interaction unit is used to process the features after connecting the fused bird's-eye view features and a preset learnable element query to obtain an element descriptor, specifically including:

[0039] Query and concatenate a group of elements in the preset elements, and input them into a multi-layer perceptron to obtain element query features;

[0040] Multiply the element query features by the fused bird's-eye view features to obtain an element mask;

[0041] Input the element mask, the fused bird's-eye view features, and the element query into the cross-attention sub-module to obtain element features;

[0042] Input the element features into the sequentially connected self-attention sub-module and feed-forward network for processing to obtain an element descriptor.

[0043] Furthermore, the point-to-element interaction unit is used to fuse the point descriptor and the element descriptor to obtain a combined feature, which specifically includes:

[0044] Update the point descriptor and the element descriptor to obtain the updated point descriptor and element descriptor;

[0045] Connect the updated point descriptor and the element descriptor to obtain a combined feature;

[0046] Update the point descriptor and the element descriptor according to the following formula:

[0047]

[0048] Among them, is the updated point descriptor, is the point descriptor, is the query obtained after the point descriptor passes through two fully connected layers, is the key obtained after the element descriptor passes through two fully connected layers, is the element descriptor, d k is the dimension of the query and key vectors, is the updated element descriptor, is the query obtained after the element descriptor passes through two fully connected layers, is the key obtained after the point descriptor passes through two fully connected layers.

[0049] Furthermore, the specific steps of step three include:

[0050] Input the combined feature into multiple decoders respectively to obtain map elements; among them, the multiple decoders include a decoder with a category head, a key-point head decoder with a dynamic matching module, and a decoder with a mask head;

[0051] Model according to the map elements in the form of ordered key points to complete the construction of the high-precision semantic map.

[0052] The second aspect of the present invention proposes a high-precision semantic map construction system based on a strongly coupled network, including:

[0053] A collaborative enhancement alignment module, which is used to input camera images and LiDAR point clouds into a preset semantic-geometry coupling module to obtain fused bird's-eye view features, facilitating the enhancement of the perception ability for rich information and long distances;

[0054] A point-element module, which is used to input the fused bird's-eye view features, preset learnable point queries, and element queries into a point-element coupling module to obtain combined features, facilitating the enhancement of the mutual constraints between points, between elements, and between points and elements;

[0055] A construction module, which is used to construct a high-precision semantic map according to the combined features.

[0056] Advantages of the present invention:

[0057] The present invention proposes a high-precision semantic map construction method based on a strongly coupled network. By using a semantic-geometry coupling module and a point-element coupling module to construct a strongly coupled network (SuperMapNet), high-precision semantic map construction is realized through the strongly coupled network. Among them, the core of the semantic-geometry coupling module is to strongly couple semantic and geometric information from a camera (camera) and a lidar (LiDAR) respectively through a collaborative enhancement unit and a difference alignment unit to generate long-distance BEV features. And through the three-level interactions of the point-to-point interaction unit, element-to-element interaction unit, and point-to-element interaction unit in the difference alignment unit, local and global features are coupled from points and elements respectively to achieve high-precision classification and positioning. In addition, by modeling map elements as ordered key points, efficient storage and calculation are realized. Description of the Drawings

[0058] Figure 1 It is a flowchart of a high-precision semantic map construction method based on a strongly coupled network provided by an embodiment of the present invention.

[0059] Figure 2 It is a schematic diagram of the overall architecture of a high-precision semantic map construction method based on a strongly coupled network provided by an embodiment of the present invention.

[0060] Figure 3 It is a partial schematic diagram of the semantic-geometry coupling module provided by an embodiment of the present invention.

[0061] Figure 4 It is a schematic diagram of the point-element coupling module provided by an embodiment of the present invention.

[0062] Figure 5 It is an architecture diagram of a high-precision semantic map construction system based on a strongly coupled network provided by an embodiment of the present invention. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0064] Embodiment 1

[0065] As Figure 1 and Figure 2 shown, a high-precision semantic map construction method based on a strongly coupled network includes:

[0066] S101: Input camera images and LiDAR point clouds into a preset semantic-geometry coupling module to obtain fused bird's-eye view features.

[0067] Specifically, LiDAR point clouds and camera images each have their own advantages and disadvantages. LiDAR point clouds provide accurate 3D geometric information, but have problems of disorder and sparsity, with an effective range of about 30m. Camera images capture rich environmental semantic information at a long distance, but lack accurate 3D geometric information. Multimodal fusion can effectively complement each other to generate long-distance features with rich semantic and geometric information. However, due to errors between different sensors, directly connecting features from different modalities will result in lower construction accuracy. Therefore, as Figure 3 shown ( Figure 3 including a collaborative enhancement unit, a differential alignment unit, and a connection unit), a semantic-geometry coupling module is proposed to generate long-distance BEV features with rich information while considering the collaboration and differences between different modalities.

[0068] The semantic-geometry coupling module simultaneously receives camera images and LiDAR point clouds as inputs to enhance the perception ability of rich information and long distances. For multi-view images with rich semantic information, Swin Transformer is used as a shared backbone network to encode image features, and then these features are connected and transformed into a unified bird's-eye view (BEV) representation through a deformable Transformer with geometric priors between the camera and BEV. For LiDAR point clouds with precise geometric information, a lidar encoder is used to process them to improve computational efficiency. Specifically, the random sample consensus algorithm (RANSAC) is first used to filter out non-ground points because BEV only focuses on ground information. Then, PointPillars with dynamic voxelization are used to generate LiDAR BEV features (lidar features). Finally, the coupling of semantic information from camera images and geometric information from LiDAR point clouds is applied, including a collaborative enhancement unit, a difference alignment unit, and a connection module for generating fused bird's-eye view features.

[0069] Specifically, the semantic-geometry coupling module includes a camera encoder, a 2D-to-bird's-eye view transformation unit, a lidar encoder, a collaborative enhancement unit, a difference alignment unit, and a connection unit.

[0070] The camera encoder is used to encode the camera image. The 2D-to-bird's-eye view transformation unit is used to transform the output of the camera encoder into camera bird's-eye view features. The lidar encoder is used to encode the LiDAR point cloud to obtain lidar features (LiDAR BEV features). The collaborative enhancement unit is used to enhance the collaborative effect of the camera bird's-eye view features and the lidar features to obtain refined camera bird's-eye view features and lidar features. The difference alignment unit is used to perform difference alignment on the refined camera bird's-eye view features and lidar features. The connection unit is used to connect the aligned refined camera bird's-eye view features and lidar features and perform basic convolution processing to obtain fused bird's-eye view features.

[0071] S102: Input the fused bird's-eye view features, a preset learnable point query, and an element query into the point-element coupling module to obtain combined features.

[0072] Specifically, for map elements modeled as ordered point sets, there are three levels of information: (1) point information, which represents the local coordinates of each point and the geometric relationship between adjacent points; (2) element information, which represents the overall shape and semantic category of each element and the relationship between adjacent elements; (3) point-element information, where the information of the element provides global constraints for the affiliated points, and the affiliated points provide specific refinements for their elements. These three levels of information are coordinated with each other. Therefore, as Figure 4As shown, the present invention introduces three levels of interaction to fully couple point and element information.

[0073] The point-element coupling module includes L point-to-point interaction units, L element-to-element interaction units, and a point-to-element interaction unit. Among them, the L point-to-point interaction units are connected in sequence, and the structures of the L point-to-point interaction units are the same. The L element-to-element interaction units are connected in sequence, and the structures of each element-to-element interaction unit are the same.

[0074] The point-to-point interaction unit is used to process the fused bird's-eye view features and the features after connecting with the preset learnable point queries to obtain point descriptors. The element-to-element interaction unit is used to process the fused bird's-eye view features and the features after connecting with the preset learnable element queries to obtain element descriptors. The point-to-element interaction unit is used to fuse the point descriptors and the element descriptors to obtain combined features.

[0075] Specifically, in order to enhance the mutual constraints between points, between elements, and between points and elements, the point-element coupling module is applied to generate combined features with local and global knowledge. It receives a set of predefined point queries element queries and the fused bird's-eye view feature B∈R C×H×W as inputs. Through three levels of interaction (point-to-point, element-to-element, point-to-element), combined features with global and local information are generated.

[0076] S103: Construct a high-precision semantic map according to the combined features.

[0077] Specifically, the combined features are sent to three decoders: a class head for labeling the classes of elements, a key point head with a dynamic matching module for regressing the coordinates and order of key points, and a mask head for predicting the masks of elements. Map elements are modeled in the form of ordered key points instead of a unified set of ordered points to achieve efficient storage and calculation.

[0078] The present invention processes the camera image and the LiDAR point cloud through constructing a semantic-geometry coupling module to obtain the fused bird's-eye view features. Then, the fused bird's-eye view features, the preset learnable point queries, and element queries are input into the point-element coupling module to obtain combined features, facilitating the mutual coordination of information at three levels. Finally, the combined features are respectively input into the three decoders to obtain map elements, and the map elements are modeled in the form of ordered key points to complete the construction of the high-precision semantic map. The present invention captures the synergies and differences between different modalities, pays attention to the interaction between elements and points, and realizes efficient storage and calculation by modeling map elements as ordered key points.

[0079] Embodiment 2

[0080] Based on the above embodiments, the present invention proposes a collaborative enhancement unit, specifically including:

[0081] To enhance the synergistic effect of two modalities (camera bird's-eye view features and lidar features), enrich semantic information and fill in the gaps in LiDAR BEV features, and at the same time add accurate 3D geometric information to camera BEV features, a collaborative enhancement unit based on cross-attention is proposed for information exchange and feature enhancement. For the BEV features of each modality. First, three fully connected layers are used to obtain the query Q, key K, and value V of the camera bird's-eye view features and lidar features respectively. Then, the attention matrix A of each modality is obtained by softmax normalization of the inner product between the Q and K of different modalities, and then it is multiplied by the corresponding V of the other modality to obtain complementary information C. Finally, the original V of each modality is concatenated with the complementary information C of the other modality and fed into a convolution operation to learn the refined BEV features of each modality with semantic and geometric information, as shown in the following formula:

[0082]

[0083]

[0084] where mod1 is the camera bird's-eye view feature and mod2 is the lidar feature, is the complementary information of the lidar feature, is the query obtained after three fully connected processes of the camera bird's-eye view feature, is the key obtained after three fully connected processes of the camera bird's-eye view feature, is the value obtained after three fully connected processes of the camera bird's-eye view feature, d k is the dimension of the query and key vectors, is the refined camera bird's-eye view feature, conv is the convolution operation, and cat is the concatenation operation, is the query obtained after three fully connected processes of the lidar feature, is the key obtained after three fully connected processes of the lidar feature, is the value obtained after three fully connected processes of the lidar feature, is the complementary information of the camera bird's-eye view feature.

[0085] Embodiment 3

[0086] Based on the above embodiments, the present invention proposes a difference alignment unit, specifically including:

[0087] For the differences between the two modalities caused by sensor errors, since the attitude accuracy of lidar is usually higher than that of cameras, the present invention uses an alignment module based on differential flow to register the improved camera BEV features to the improved lidar BEV features. The refined camera bird's-eye view feature B obtained in the collaborative enhancement unit cam ∈R C×H×W and the lidar feature b lidar ∈R C×H×W are first concatenated and fed into multiple basic convolutional blocks to obtain a differential flow (Δx, Δy) ∈ R 2×H×W . Then, the coordinates of the refined camera bird's-eye view feature B cam are corrected by adding the differential flow (Δx, Δy) to the original coordinates and resampling to generate the aligned camera bird's-eye view feature B' cam ∈R C×H×W .

[0088] Embodiment 4

[0089] Based on the above embodiments, the present invention proposes a point-element coupling module, which specifically includes:

[0090] The point-element coupling module includes L point-to-point interaction units, L element-to-element interaction units, and a point-to-element interaction unit. Among them, the L point-to-point interaction units are connected in sequence, the structures of the L point-to-point interaction units are the same, the L element-to-element interaction units are connected in sequence, and the structures of each element-to-element interaction unit are the same.

[0091] Specifically, the point-to-point interaction unit aims to comprehensively learn the external geometric relationships between element points and the internal local coordinate information of each point. For each element m, this unit takes the fused bird's-eye view feature B ∈ R C×H×W and a set of preset learnable point queries as inputs (the point queries are set according to the actual situation and experience), and outputs a set of point descriptors This module contains L layers, and each layer includes a cross-attention sub-module for external point interaction, a self-attention sub-module for internal point interaction, and a feed-forward network (FFN) for final point feature learning. In each layer l, a set of learnable point queries of the preset element m are concatenated and input into a multi-layer perceptron (the preset element m is set according to the actual situation and experience, such as: the elements include lane dividers, crosswalks, and road boundaries) to learn the point query feature F m ∈R C . Then, the point query feature F m and the fused bird's-eye view feature B ∈ R C×H×W are multiplied to obtain a point mask M point ∈R H×W。Mask M point Subsequently, it is input into the cross-attention layer together with the fused bird's-eye view feature B and the point query to learn the internal point features and further enhance the subordinate prior information. Then, a self-attention sub-module (including a self-attention layer) and an FFN are respectively used to learn the inter-point features and obtain the final point descriptor

[0092] The element-to-element interaction unit aims to comprehensively learn the overall shape and semantic relationships between different elements. For all map elements in the map, this unit takes the fused bird's-eye view feature B ∈ R C×H×W and a set of preset learnable element queries as inputs (the preset element m is set according to the actual situation and experience, and is the same as the element described in the point-to-point interaction unit), and outputs a set of element descriptors This unit contains L layers and has the same structure as the point-to-point interaction unit (Point2Point Interactor). For all elements, a set of preset learnable element queries are concatenated and input into a multi-layer perceptron to learn the element query feature F ∈ R C . Then, the element query feature F and the fused bird's-eye view feature B ∈ R C×H×W are multiplied to obtain the element mask M element ∈ R H×W 。Mask M element Subsequently, it is used in the cross-attention layer, together with the fused bird's-eye view feature B and the element query to generate the overall intra-element features Then, a self-attention layer and a feed-forward network are used to learn the inter-element features and generate the final element descriptor

[0093] The point-to-element interaction unit between an element and its constituent points aims to utilize the overall shape and semantic knowledge at the element level to supplement the point-level information only through local position knowledge, so as to integrate the overall information while considering the details. A cross-attention layer with position embeddings is applied between the point feature (simplified to D points ) and the overall element feature (simplified to D elements ) for information exchange. The update steps of the two features can be expressed as follows:

[0094]

[0095] where, is the updated point descriptor, is a point descriptor, is the query obtained after the point descriptor passes through two fully connected layers, is the key obtained after the element descriptor passes through two fully connected layers, is the element descriptor, d k is the dimension of the query and key vectors, is the updated element descriptor, is the query obtained after the element descriptor passes through two fully connected layers, is the key obtained after the point descriptor passes through two fully connected layers.

[0096] Then, the updated point descriptor and element descriptor are concatenated to obtain a combined feature.

[0097] Embodiment 5

[0098] Based on the above embodiments, as Figure 5 shown, the present invention proposes a high-precision semantic map construction system based on a strong coupling network, including:

[0099] A collaborative enhancement alignment module, configured to input camera images and LiDAR point clouds into a preset semantic-geometry coupling module to obtain fused bird's-eye view features.

[0100] A point-element module, configured to input the fused bird's-eye view features, a preset learnable point query, and an element query into a point-element coupling module to obtain a combined feature.

[0101] A construction module, configured to construct a high-precision semantic map according to the combined feature.

[0102] It should be noted that the high-precision semantic map construction system based on a strong coupling network provided by the embodiments of the present invention is to implement the above-mentioned high-precision semantic map construction method based on a strong coupling network. Its functions can be specifically referred to the above method embodiments and will not be elaborated here.

[0103] In summary, the present invention proposes a high-precision semantic map construction method based on a strong coupling network. By using a semantic-geometry coupling module and a point-element coupling module to construct a strong coupling network (SuperMapNet), high-precision semantic map construction is achieved through the strong coupling network. Among them, the core of the semantic-geometry coupling module is to strongly couple semantic and geometric information from a camera (camera) and a lidar (LiDAR) respectively through a collaborative enhancement unit and a difference alignment unit to generate long-range BEV features. And through the three-level interaction of the point-to-point interaction unit, element-to-element interaction unit, and point-to-element interaction unit in the difference alignment unit, local and global features are coupled from points and elements respectively to achieve high-precision classification and positioning. In addition, by modeling map elements as ordered key points, efficient storage and calculation are realized.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a high-precision semantic map based on a strongly coupled network, characterized in that, Including: Step 1: Input the camera image and LiDAR point cloud into a preset semantic-geometry coupling module to obtain fused bird's-eye view features; Step 2: Input the fused bird's-eye view features, a preset learnable point query, and an element query into a point-element coupling module to obtain combined features; Step 3: Construct a high-precision semantic map based on the combined features.

2. The method for constructing a high-precision semantic map based on a strongly coupled network according to claim 1, characterized in that, The semantic-geometry coupling module includes a camera encoder, a 2D to bird's-eye view transformation unit, a LiDAR encoder, a collaborative enhancement unit, a difference alignment unit, and a connection unit; The camera encoder is used to encode the camera image; The 2D to bird's-eye view transformation unit is used to transform the output of the camera encoder into camera bird's-eye view features; The LiDAR encoder is used to encode the LiDAR point cloud to obtain LiDAR features; The collaborative enhancement unit is used to enhance the synergistic effect of the camera bird's-eye view features and the LiDAR features to obtain refined camera bird's-eye view features and LiDAR features; The difference alignment unit is used to perform difference alignment on the refined camera bird's-eye view features and the LiDAR features; The connection unit is used to connect the aligned refined camera bird's-eye view features and the LiDAR features and perform basic convolution processing to obtain fused bird's-eye view features.

3. A high-precision semantic map construction method based on a strongly coupled network according to claim 2, characterized in that The collaborative enhancement unit is represented by the following formula: Among them, mod1 is the camera bird's-eye view feature, and mod2 is the lidar feature. is the complementary information of the lidar feature. is the query obtained after the camera bird's-eye view feature undergoes three fully connected processes. is the key obtained after the camera bird's-eye view feature undergoes three fully connected processes. is the value obtained after the camera bird's-eye view feature undergoes three fully connected processes, d k is the dimension of the query and the key vector. is the refined camera bird's-eye view feature, conv is the convolution operation, and cat is the concatenation operation. is the query obtained after the lidar feature undergoes three fully connected processes. is the key obtained after the lidar feature undergoes three fully connected processes. is the value obtained after the lidar feature undergoes three fully connected processes. is the complementary information of the camera bird's-eye view feature.

4. A high-precision semantic map construction method based on a strongly coupled network according to claim 2, characterized in that, The difference alignment unit is used to perform difference alignment on the refined camera bird's-eye view features and the LiDAR features, specifically including: Connect the refined camera bird's-eye view features and the LiDAR features and perform multiple convolution processes to obtain a difference flow; Input the difference flow and the refined camera bird's-eye view features into a transform to obtain the aligned camera bird's-eye view features.

5. A high-precision semantic map construction method based on a strongly coupled network according to claim 1, characterized in that The point-element coupling module includes multiple point-to-point interaction units, multiple element-to-element interaction units, and a point-to-element interaction unit; among them, multiple point-to-point interaction units are connected in sequence, the structures of multiple point-to-point interaction units are the same, multiple element-to-element interaction units are connected in sequence, and the structures of each element-to-element interaction unit are the same; The point-to-point interaction unit is used to process the feature obtained by connecting the fused bird's-eye view features and a preset learnable point query to obtain a point descriptor; The element-to-element interaction unit is used to process the feature obtained by connecting the fused bird's-eye view features and a preset learnable element query to obtain an element descriptor; The point-to-element interaction unit is used to fuse the point descriptor and the element descriptor to obtain combined features.

6. The high-precision semantic map construction method based on a strongly coupled network according to claim 5, characterized in that The point-to-point interaction unit is used to process the feature obtained by connecting the fused bird's-eye view features and a preset learnable point query to obtain a point descriptor, specifically including: Connect a set of point queries in a preset element and input them into a multi-layer perceptron to obtain point query features; Multiply the point query features and the fused bird's-eye view features to obtain a point mask; Input the point mask, the fused bird's-eye view features, and the point query into a cross-attention sub-module to obtain point features Input the point features into a self-attention sub-module and a feed-forward network connected in sequence for processing to obtain a point descriptor.

7. A method for constructing a high-precision semantic map based on a strongly coupled network according to claim 5, characterized in that The element-to-element interaction unit is used to process the feature obtained by connecting the fused bird's-eye view features and the preset learnable element queries to obtain an element descriptor, specifically including: Connect a set of element queries in the preset elements and input them into a multi-layer perceptron to obtain element query features; Multiply the element query features and the fused bird's-eye view features to obtain an element mask; Input the element mask, the fused bird's-eye view features, and the element queries into a cross-attention sub-module to obtain element features; Input the element features into a self-attention sub-module and a feed-forward network connected in sequence for processing to obtain an element descriptor.

8. A high-precision semantic map construction method based on a strongly coupled network according to claim 5, characterized in that The point-to-element interaction unit is used to fuse the point descriptor and the element descriptor to obtain a combined feature, specifically including: Update the point descriptor and the element descriptor to obtain the updated point descriptor and element descriptor; Connect the updated point descriptor and the element descriptor to obtain a combined feature; Update the point descriptor and the element descriptor according to the following formula: Among them, is the updated point descriptor, is the point descriptor, is the query obtained after the point descriptor passes through two fully connected layers, is the key obtained after the element descriptor passes through two fully connected layers, is the element descriptor, d k is the dimension of the query and key vectors, is the updated element descriptor, is the query obtained after the element descriptor passes through two fully connected layers, is the key obtained after the point descriptor passes through two fully connected layers.

9. A high-precision semantic map construction method based on a strongly coupled network according to claim 1, characterized in that The specific content of step three includes: Input the combined feature into multiple decoders respectively to obtain map elements; wherein, the multiple decoders include a decoder with a class head, a key point head decoder with a dynamic matching module, and a decoder with a mask head; Model according to the map elements in the form of ordered key points to complete the construction of a high-precision semantic map.

10. A high-precision semantic map construction system based on a strongly coupled network, characterized in that, Including: A collaborative enhancement alignment module, which is used to input the camera image and the LiDAR point cloud into a preset semantic-geometry coupling module to obtain fused bird's-eye view features; A point-element module, which is used to input the fused bird's-eye view features, the preset learnable point queries, and the element queries into a point-element coupling module to obtain a combined feature; A construction module, which is used to construct a high-precision semantic map according to the combined feature.

Citation Information

Cited By

  • Semantic map construction method and robot control system

    CN120635902A