Laser radar and vision fusion location identification method based on cross-modal attention mechanism

By adopting cross-modal attention mechanism and depth separation convolution technologies in the fusion of lidar and visual information, the problem of poor fusion of lidar and visual information in the existing technology is solved, and rapid and accurate location recognition is achieved when environmental changes are achieved.

CN120219780APending Publication Date: 2025-06-27上海智能新能源汽车科创功能平台有限公司 +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311789152.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate lidar and visual information, making it difficult to quickly and accurately identify the location when the environment changes.

Method used

Using a method based on a cross-modal attention mechanism, the lidar point cloud data and visual image data are feature extraction and fusion, and through technologies such as multi-head attention mechanism and depth separation convolution, robust feature representations are learned and fusion descriptors are generated.

Benefits of technology

It realizes the effective fusion of lidar and visual information, is not affected by the environment, can quickly and accurately identify the location, and improves the robustness of single-modal descriptors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219780A_ABST
    Figure CN120219780A_ABST
Patent Text Reader

Abstract

The invention relates to a laser radar and vision fusion location identification method based on a cross-modal attention mechanism, and the method comprises the following steps: carrying out the feature extraction of laser radar point cloud data through employing a voxel format, and obtaining the multi-scale BEV features of the laser radar point cloud data; performing feature extraction on the image data to obtain multi-scale 2D features of the image; fusing the multi-scale BEV features of the laser radar point cloud data and the multi-scale 2D features of the image by introducing a cross-modal fusion method and a local and global fusion method to obtain fusion features; mapping the fusion feature into a fusion descriptor; and querying the fusion descriptor in a map database to complete location identification. Compared with the prior art, the method has the advantages that the laser radar and visual information are effectively fused, the method is not affected by the environment, site identification can be rapidly and accurately carried out, and the problem that site identification is carried out by effectively fusing the laser radar and camera data in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-sensor fusion, and in particular to a method for lidar and vision fusion location recognition based on a cross-modal attention mechanism. Background Art

[0002] Location recognition is an important task with practical applications in various fields such as robotics, simultaneous localization and mapping (SLAM), and autonomous driving. This task is typically formulated as an instance retrieval problem. Given a query instance, the goal is to find the closest match in a large database. Over the years, researchers have utilized different sensors to improve discrimination ability and recognition accuracy. LiDAR sensors provide accurate point cloud data and reliable depth information, making them ideal for capturing geometric features such as building facades and tree canopies. In contrast, cameras can capture rich visual information, including color, texture, and semantic content, which can help identify landmarks and recognize objects.

[0003] However, point cloud-based methods may have difficulty identifying places with small geometric changes and occlusions, while camera-based methods are easily affected by lighting changes and appearance variations. Combining these two sensors has the potential to mitigate some of the challenges faced by each sensor, thus providing more reliable and powerful performance for location recognition. Existing location recognition fusion methods involve simply adding the two modalities together without fully considering the potential correlations between them. However, the simple stacking of multi-modal features may lead to information redundancy. To address the above problems, a method that can effectively fuse lidar and visual information, is not affected by the environment, and can perform location recognition quickly and accurately is needed. Summary of the Invention

[0004] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a method for lidar and vision fusion location recognition based on a cross-modal attention mechanism.

[0005] The purpose of the present invention can be achieved through the following technical solutions:

[0006] On the one hand, the present invention provides a method for lidar and vision fusion location recognition based on a cross-modal attention mechanism, including the following steps:

[0007] Step S1: Extract features from the lidar point cloud data in the format of voxels to obtain BEV features;

[0008] Step S2: Use the ResNet algorithm to extract features from the image data to obtain 2D image features;

[0009] Step S3: Introduce a cross-modal fusion method and a local-global fusion method to fuse the BEV features and the image 2D features to obtain cross-modal fusion features;

[0010] Step S4: Map the cross-modal fusion features to fusion descriptors;

[0011] Step S5: Query the fusion descriptors in the map database to complete location recognition.

[0012] Further, the specific steps of Step S1 include:

[0013] Step S11: Use the VoxelNet algorithm to divide the lidar point cloud data into 3D regular grids to obtain lidar point cloud voxels, perform feature encoding and feature aggregation on the points in each voxel to obtain voxel features;

[0014] Step S12: Use the point cloud feature extraction network of the feature pyramid network to abstract the voxel features to obtain the BEV features.

[0015] Further, Step S2 includes: Extracting 2D image feature maps through the ResNet algorithm, and the ResNet algorithm includes four stages, and in the four stages, the resolution of the 2D image feature maps is reduced and the number of channels of the 2D image feature maps is increased.

[0016] Further, the local-global fusion method uses the multi-head attention mechanism to fuse local features and global features.

[0017] Further, Step S3 includes the following steps:

[0018] Step S31: Perform local-global feature fusion on the BEV features and the image 2D features respectively to obtain the local-global fusion features of the lidar point cloud data and the local-global fusion features of the image;

[0019] Step S32: Perform cross-modal fusion on the local-global fusion features of the lidar point cloud data and the local-global fusion features of the image to obtain cross-modal fusion features;

[0020] Step S33: Perform depthwise separable convolution on the local-global fusion features of the lidar point cloud data to obtain context representations in the same hidden space;

[0021] Step S34: Connect the cross-modal fusion features and the lidar features after depthwise separable convolution to generate the cross-modal fusion features.

[0022] Further, obtaining cross-modal fusion features specifically includes the following steps:

[0023] Step S321: Linearly transform the local-global fusion feature of the lidar point cloud data into a query feature, and linearly transform the local-global fusion feature of the image into a key feature and a value feature;

[0024] Step S322: Perform a scaled dot product on the query feature and the key feature to obtain an attention affinity matrix;

[0025] Step S323: Normalize the attention affinity matrix through a softmax operation;

[0026] Step S324: Divide the query feature, the key feature, and the value feature into multiple heads through a multi-head attention mechanism, and perform the attention process in parallel;

[0027] Step S325: Connect the output values of each head and perform a linear projection to obtain the cross-modal fusion feature of the two modalities.

[0028] Furthermore, adopt a descriptor generation algorithm to map the fusion feature into a fusion descriptor.

[0029] Furthermore, query the fusion descriptor with the minimum Euler distance in the map database to complete location recognition.

[0030] Second invention, the present invention provides an electronic device, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the method described above.

[0031] Third aspect, the present invention provides a storage medium, on which a program is stored, and characterized in that the program, when executed, implements the method described above.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] 1. The present invention can effectively fuse lidar and visual information, is not affected by the environment, can quickly and accurately perform location recognition, and solves the problem of effectively fusing lidar and camera data for location recognition in the prior art.

[0034] 2. The present invention proposes a local-global feature fusion method, which uses local fine-grained features to supplement global coarse-grained features, and improves the robustness of the single-modal descriptor.

[0035] 3. The present invention proposes a cross-modal fusion method, which uses a cross-attention mechanism to find the correlation between modalities and can learn a robust feature representation. Description of the Drawings

[0036] Figure 1 Schematic diagram of the process of the present invention;

[0037] Figure 2 Schematic diagram of point cloud feature extraction in the specific embodiment of the present invention;

[0038] Figure 3 Schematic diagram of transmembrane state fusion in the specific embodiment of the present invention. Specific embodiment

[0039] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and the detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0040] This embodiment provides a lidar and vision fusion location recognition method based on a cross-modal attention mechanism, as Figure 1 shown, including the following steps:

[0041] Step 1, lidar point cloud feature extraction: Use the voxel format to extract features from the lidar point cloud data to obtain the BEV features of the lidar point cloud data.

[0042] Specifically, as Figure 2 shown, select VoxelNet as the backbone network for point cloud feature extraction, and add a feature pyramid network. Among them, the backbone network divides the lidar point cloud data into 3D regular grids to obtain a certain number of voxels, encodes and aggregates the features of the points in each voxel to obtain voxel features. Then, use the point cloud feature extraction network based on the feature pyramid network to further abstract the point cloud features. Specifically, three voxel downsampling blocks composed of 3D convolution - Batch Norm - ReLU activation functions are used to process the sparse voxel features in the bottom - up direction to generate 3D feature maps with continuously increasing receptive fields and decreasing resolutions. Then, the features generated by the last convolution block are transformed using a 1×1×1 3D convolution and then upsampled using a transposed convolution. The upsampled feature map is connected with the features passed from bottom to top at the corresponding layer after being transformed by a 1×1×1 3D convolution to obtain the final BEV features of the point cloud data; this design produces a feature map with relatively high spatial resolution and large receptive field.

[0043] Step 2, image feature extraction: Extract features from the image data to obtain the 2D features of the image.

[0044] Specifically, ResNet18 is used as the backbone network to extract the feature maps of 2D images. The backbone network is divided into four stages, and these four stages are used to extract 2D feature maps at four scales with 32, 64, 128, and 256 channels. Each stage reduces the resolution of the input feature maps and increases the number of channels of the feature maps.

[0045] Step 3, Feature Fusion: By introducing cross-modal fusion methods and local-global fusion methods, fuse the BEV features of the lidar point cloud data and the 2D features of the images obtained in Step 1 and Step 2 to obtain fused features.

[0046] Specifically, it includes the following steps:

[0047] S31. For the multi-scale BEV features of the lidar point cloud data and the multi-scale 2D features of the images, use the multi-head attention mechanism to fuse the local features extracted in the first stage and the global features extracted in the last stage of the lidar modality and the image modality together. It includes the following steps: First, linearly transform the local features into local query features local key features and local value features Linearly transform the global features into global query features global key features and global value features

[0048] For the lidar point cloud, N l and N g represent the number of voxels of the local features and the global features respectively. In this example, the number of voxels of the local features of the lidar point cloud is 576000, and the number of voxels of the global features of the lidar point cloud is 72000. For the image, N l and N g represent the product of the height and width of the local features and the global feature maps respectively. In this example, the number of voxels of the local features of the image is 226440, and the number of voxels of the global features of the lidar point cloud is 28305. d q , d k , d v are the dimensions of the query features, the key features, and the value features respectively. In this example, d q = d k = d v = 128; Then, fuse the local features and the global features through multi-head attention. Input the local query features, the global key features, and the global value features into a multi-head attention module respectively to obtain the fusion feature F guided by the local features l , and input the global query features, the local key features, and the local value features into another multi-head attention module to obtain the fusion feature F guided by the global features g, the locally feature-guided fused feature and the globally feature-guided fused feature are concatenated together to obtain the final local-global fused feature F, and the above process is expressed as:

[0049] F l = MultiHead(Q l , K g , V g )

[0050] F g = MultiHead(Q g , K l , V l )

[0051] F = F l + F g

[0052] where, is the local query feature, is the local key feature, is the local value feature, is the global query feature, is the global key feature, is the global value feature.

[0053] The local-global fused feature of the lidar point cloud data and the local-global fused feature of the image are respectively generated through the above formulas and the local-global fused feature of the image

[0054] S32. Cross-modal fusion is performed on the local-global fused feature of the lidar point cloud data and the local-global fused feature of the image to obtain the cross-modal fused feature of the two modalities. The specific main steps include:

[0055] First, linearly transform the local-global fused feature of the lidar point cloud data into the query feature Linearly transform the local-global fused feature of the image into the key feature and the value feature In this example, N1 = 648000, N2 = 254745, d1 = d2 = d qk = d v = 128;

[0056] After that, use the query feature and the key feature to perform scaled dot product to obtain the attention affinity matrix, which represents the N1×N2 correlation between the point cloud feature and all corresponding N2 RGB camera features. The scaled dot product is expressed as:

[0057]

[0058] After normalizing the attention affinity matrix through the softmax operation, the attention affinity matrix is used to weight and aggregate V2 to dynamically capture the correlation between the two modalities:

[0059] CrossAttn(Q1, K2, V2) = softmax(AffinityMatrix)V2

[0060] Q1, K2, and V2 are divided into multiple heads through the multi-head attention mechanism, and the attention process is executed in parallel;

[0061] Finally, the output values of each head are concatenated and linearly projected to form the final output, that is, the cross-modal fusion features of the two modalities;

[0062] S32. The fusion features of the lidar point cloud data are subjected to depthwise separable convolution to obtain context representations in the same hidden space. In the depth convolution stage, each input channel of the fusion features of the lidar point cloud data is convolved with an independent set of filters. In the point convolution stage, the output feature maps of the depth convolution stage are combined by applying a 1×1 convolution kernel to merge the channels of the depth convolution stage into the final output channels. Then, in order to further capture features in different ranges and improve the generalization ability, units with three convolution kernel sizes of 7, 5, and 3 are used for parallel processing. For element i and output dimension c, the weight is The output is The above process is represented as follows:

[0063]

[0064] where d is the hidden layer dimension, which is 128 in this example, and softmax() is the softmax operation.

[0065] S34. The cross-modal fusion features and the lidar features after depthwise separable convolution are concatenated to generate the final fusion features.

[0066] Step 4. Apply a descriptor generation algorithm to the final fusion features to generate the final fusion descriptor.

[0067] Specifically, the descriptor generation algorithm can be global average pooling, VLAD, or max pooling. Preferably, in this embodiment, the method of global average pooling is used to generate a 256-dimensional fusion descriptor.

[0068] Step 5. Query the generated fusion descriptor in the map database to find the descriptor with the minimum Euler distance in the map database to complete place recognition.

[0069] Specifically, taking the location recognition of the park scene as an example for illustration:

[0070] S51. Establish a park fusion descriptor database: Obtain a set of lidar point cloud and camera image data every 1 meter in the park, and each set of data is attached with the true GPS coordinates of the acquisition location. Input the collected park data into the above method to obtain the fusion descriptor database.

[0071] S52. Location recognition: Obtain lidar point cloud and camera image data in the park and input them into the network, extract the fusion descriptor, calculate the Euclidean distance through all the fusion descriptors in the fusion descriptor database, and sort the distances. The position represented by the lidar and camera combination data frame with the smallest distance is the position of the query data. Return the GPS coordinates of the combination data frame as the geographical location of the query data to complete the location recognition.

[0072] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A lidar and vision fusion location recognition method based on a cross-modal attention mechanism, characterized in that, It includes the following steps: Step S1: Extract features from the lidar point cloud data in the format of voxels to obtain BEV features; Step S2: Use the ResNet algorithm to extract features from the image data to obtain 2D image features; Step S3: Introduce a cross-modal fusion method and a local-global fusion method to fuse the BEV features and the 2D image features to obtain cross-modal fusion features; Step S4: Map the cross-modal fusion features to fusion descriptors; Step S5: Query the fusion descriptors in the map database to complete location recognition.

2. The lidar and vision fusion location recognition method based on cross-modal attention mechanism according to claim 1, characterized in that The specific content of step S1 includes: Step S11: Use the VoxelNet algorithm to divide the lidar point cloud data into 3D regular grids to obtain lidar point cloud voxels, perform feature encoding and feature aggregation on the points in each voxel to obtain voxel features; Step S12: Use the point cloud feature extraction network of the feature pyramid network to abstract the voxel features to obtain the BEV features.

3. The lidar and vision fusion location recognition method based on a cross-modal attention mechanism according to claim 1, characterized in that, Step S2 includes: Extracting a 2D image feature map through the ResNet algorithm. The ResNet algorithm includes four stages, and in each of the four stages, the resolution of the 2D image feature map is reduced and the number of channels of the 2D image feature map is increased.

4. A lidar and vision fusion location recognition method based on a cross-modal attention mechanism according to claim 1, characterized in that, The local-global fusion method uses the multi-head attention mechanism to fuse local features and global features.

5. A lidar and vision fusion location recognition method based on a cross-modal attention mechanism according to claim 4, characterized in that, Step S3 includes the following steps: Step S31: Perform local-global feature fusion on the BEV features and the 2D image features respectively to obtain the local-global fusion features of the lidar point cloud data and the local-global fusion features of the image; Step S32: Perform cross-modal fusion on the local-global fusion features of the lidar point cloud data and the local-global fusion features of the image to obtain cross-modal fusion features; Step S33: Perform depthwise separable convolution on the local-global fusion features of the lidar point cloud data to obtain a context representation in the same hidden space; Step S34: Connect the cross-modal fusion features and the lidar features after depthwise separable convolution to generate the cross-modal fusion features.

6. A lidar and vision fusion location recognition method based on a cross-modal attention mechanism according to claim 1, characterized in that, The specific steps for obtaining cross-modal fusion features include the following: Step S321: Linearly transform the local-global fusion features of the lidar point cloud data into query features, and linearly transform the local-global fusion features of the image into key features and value features; Step S322: Perform a scaled dot product on the query features and the key features to obtain an attention affinity matrix; Step S323: Normalize the attention affinity matrix through softmax operation; Step S324: Divide the query features, the key features, and the value features into multiple heads through the multi-head attention mechanism and perform the attention process in parallel; Step S325: Connect the output values of each head and perform a linear projection to obtain cross-modal fusion features of the two modalities.

7. A lidar and vision fusion location recognition method based on a cross-modal attention mechanism according to claim 1, characterized in that, Use a descriptor generation algorithm to map the fusion features to fusion descriptors.

8. A lidar and vision fusion location recognition method based on a cross-modal attention mechanism according to claim 1, characterized in that, Query the fusion descriptor with the minimum Euler distance in the map database to complete location recognition.

9. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that the electronic device executes the method according to any one of claims 1-8.

10. A storage medium, on which a program is stored, characterized in that, When the program is executed, it implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Attention-based 4D millimeter wave radar and vision fusion method

    CN116129234A

  • Road scene adaptive three-dimensional target detection method based on Leiyu fusion registration

    CN116402677A

  • Multi-modal fusion sensing method and device of vehicle, vehicle and storage medium

    CN116543361A

  • Target detection method for autonomous vehicle based on radar and vision fusion

    CN116958934A