A point cloud image fusion method based on sparse fusion and dense fusion

By designing a sparse-intensive fusion model, combining a sparse fusion module and a dense fusion module, the problem of information loss of point cloud and RGB images in autonomous driving scenarios is solved, and more efficient feature fusion and detection accuracy are achieved.

CN116468981BActive Publication Date: 2025-07-25SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Patent Information

Application Number
CN202310368978.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-06
Publication Date
2025-07-25
Estimated Expiration
2043-04-06

AI Technical Summary

Technical Problem

The existing point cloud and RGB image fusion methods have the problem of information loss in autonomous driving scenarios, sparse fusion of lost semantic information and dense fusion of lost geometric information, making it difficult to effectively combine the advantages of both.

Method used

A sparse-density fusion model (SDF) is designed, which includes a sparse fusion module and a dense fusion module. The sparse fusion module is used to maintain the geometric prior of the point cloud, and the dense fusion module is used to utilize the semantic information of the image to perform feature interaction through a deformable Transformer.

Benefits of technology

It significantly improves the accuracy and efficiency of 3D object detection, improves the perceptual accuracy in autonomous driving scenarios, especially the performance on the nuScenes dataset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468981B_ABST
    Figure CN116468981B_ABST
Patent Text Reader

Abstract

The present invention discloses a point cloud image fusion method based on sparse fusion and dense fusion. The method includes: for a target scene, acquiring data of two modalities, namely an image and a point cloud; inputting the image and the point cloud into a trained fusion model to obtain fusion features. The fusion model includes a first branch, a second branch, a sparse fusion module, and a dense fusion module. The first branch is used to extract image features and respectively transmit them to the sparse fusion module and the dense fusion module; the second branch is used to extract voxel features of the point cloud, and further generate a bird's-eye view feature based on the output of the sparse fusion module; the sparse fusion module is used to fuse the image features into valid voxels from the second branch; the dense fusion module is used to generate a fused bird's-eye view feature as the fusion feature based on the bird's-eye view from the second branch and the image features from the first branch. The present invention improves the accuracy of target detection by designing a simple and effective complementary structure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and more specifically, to a point cloud image fusion method based on sparse fusion and dense fusion. Background Art

[0002] In the field of 3D object detection, the point cloud from lidar and the RGB image from the camera are two complementary general perception sources. Taking the autonomous driving scenario as an example, the point cloud can provide an accurate 3D spatial structure around the vehicle, and the RGB image can provide rich semantic information of the perception scene. However, due to the huge differences in the structures of the point cloud and the image, it is difficult to fully fuse the data of these two modalities.

[0003] Multi-modal learning is crucial in the autonomous driving scenario because the sensor configuration on the vehicle is very complex, and information interaction between sensors is an effective solution to improve the perception accuracy. However, due to the significant differences between the sparsity of radar data and the density of camera data, the current fusion solutions still have deficiencies. According to the different feature representations in the fusion module, the existing data fusion methods can be divided into two ways: only sparse fusion and only dense fusion.

[0004] In sparse-only fusion, the point cloud is projected onto the image plane, and semantic labels or image features are collected around the projection positions. In terms of collecting semantic labels, two schemes, point-painting and point-augmenting, project the point cloud onto the image plane and attach semantic labels to 3D points around the projection positions. Recently, the research trend has shifted towards fusing features, such as AutoAlign (Zehui Chen, Zhenyu Li, Shiquan Zhang, Liangji Fang, Qinghong Jiang, Feng Zhao, Bolei Zhou, and Hang Zhao. Autoalign: Pixel-instance feature aggregation for multimodal 3d object detection. arXiv preprint arXiv:2201.06493, 2022.) which focuses on pixel-instance features. In the sparse-only fusion approach, the point cloud is projected onto the image plane, and semantic labels or image features are collected around the projection positions, thereby fusing image information into the point cloud representation. Although this approach can retain 3D geometric information in the point cloud features, the rich semantic information provided by the image features is damaged. This is because a 3D object described by several points in the point cloud data corresponds to hundreds of pixels in the image data. The problem caused by this sparse density inconsistency is that when attaching image features to the voxel representation, approximately one to two orders of magnitude of image features are deprecated through LiDAR camera projection.

[0005] In dense-only fusion, the image features are either sampled into the voxel space or projected into the BEV (Bird's Eye View) space and fused with the point cloud features in this representation. For example, the image features and the point cloud features are respectively transformed into the BEV space and then fused grid by grid in the BEV. Although transforming the image features into the BEV space still retains semantic information, compressing the point cloud features into the BEV space loses the geometric prior in the 3D space. In addition, projecting the image features into the BEV space is ill-posed because the camera does not capture any 3D geometry, resulting in grid inconsistencies of the same object between the image BEV and the point cloud BEV features, increasing the difficulty of fusion.

[0006] In summary, whether it is sparse-only fusion or dense-only fusion, information loss is inevitable. Therefore, a new data fusion mechanism needs to be designed so that sparse-only fusion and dense-only fusion can jointly contribute to the network. Summary of the Invention

[0007] The objective of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a point cloud image fusion method based on sparse fusion and dense fusion. The method comprises the following steps:

[0008] For a target scene, obtain data of two modalities, namely images and point clouds;

[0009] Input the images and point clouds into a trained fusion model to obtain fusion features;

[0010] Wherein, the fusion model includes a first branch, a second branch, a sparse fusion module, and a dense fusion module. The first branch is used to extract image features and respectively transmit them to the sparse fusion module and the dense fusion module; the second branch is used to extract voxel features of the point cloud and then generate bird's-eye view features based on the output of the sparse fusion module; the sparse fusion module is used to fuse the image features into valid voxels from the second branch; the dense fusion module is used to generate fused bird's-eye view features as the fusion features based on the bird's-eye view from the second branch and the image features from the first branch.

[0011] Compared with the existing technologies, the advantages of the present invention are as follows: a sparse-dense fusion model (SDF) is proposed. This model has a complementary structure connecting the sparse fusion module and the dense fusion module. The sparse fusion module is responsible for collecting image features around non-empty voxels to aggregate geometric metric priors, aiming to maintain the geometric priors provided by the point cloud features. The dense fusion module projects the point cloud features into a bird's-eye view and queries the image features in the perspective view to utilize the rich semantic information in the images. Through this simple and effective design, the defects of only sparse or only dense fusion can be compensated for, and the advantages of both can be organically combined.

[0012] Through the following detailed description of the exemplary embodiments of the present invention with reference to the accompanying drawings, other features and advantages of the present invention will become clear. Description of the Drawings

[0013] The accompanying drawings incorporated in and constituting a part of this specification illustrate embodiments of the present invention and, together with the description thereof, are used to explain the principles of the present invention.

[0014] Figure 1 is a flowchart of a point cloud image fusion method based on sparse fusion and dense fusion according to an embodiment of the present invention;

[0015] Figure 2 is a schematic diagram of the process of a point cloud image fusion method based on sparse fusion and dense fusion according to an embodiment of the present invention;

[0016] In the accompanying drawings, Image - Image; Visual Encoder - Visual Encoder; Camera feature - Camera feature; Voxelization - Voxelization; Voxel feature - Voxel feature; Sparse Encoder - Sparse Encoder; SparseFusion - Sparse Fusion; Dense Fusion - Dense Fusion; Fusion Output - Fusion Output; BEV Output - Bird's Eye View Output; Feed Forward - Feed Forward layer; Add&Norm - Add and Normalization layer; query - Query; DetectionHead - Detection Head. Detailed implementation manners

[0017] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and numerical values set forth in these embodiments do not limit the scope of the present invention.

[0018] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way serves as a limitation to the present invention, its application, or its use.

[0019] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such techniques, methods, and devices should be regarded as part of the specification.

[0020] In all the examples shown and discussed herein, any specific values should be construed as merely exemplary and not as a limitation. Accordingly, other examples of the exemplary embodiments may have different values.

[0021] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it need not be further discussed in subsequent drawings.

[0022] In this article, existing data fusion schemes are divided into sparse fusion and dense fusion. Among them, only the sparse fusion scheme retains the 3D geometric prior but loses rich semantic information, while only the dense fusion scheme retains semantic continuity but lacks precise geometric information. The present invention aims to fully combine the advantages of the sparse fusion and dense fusion schemes, propose a sparse - dense fusion strategy, and achieve the purpose of fully utilizing point cloud and image information for 3D object detection. The present invention can be used for object detection in various scenarios. For the sake of clarity, the 3D detection feature fusion strategy will be mainly described below by taking the autonomous driving scenario as an example.

[0023] See Figure 1As shown, the provided point cloud image fusion method based on sparse fusion and dense fusion includes the following steps:

[0024] Step S110, obtain data of two types of perception sources, namely an image and a point cloud of the target scene.

[0025] LiDAR and cameras are two important sensors for 3D object detection in the autonomous driving scenario. An RGB image of the target scene can be captured by a camera, and point cloud data can be obtained by a LiDAR system.

[0026] It should be understood that other methods can also be used to obtain the image and the point cloud. For example, a depth camera can be used to obtain the point cloud.

[0027] Step S120, input the image and the point cloud into a trained fusion model, fuse the image and the point cloud from both sparse and dense perspectives, and obtain fused features.

[0028] See Figure 2 As shown, the constructed fusion model is also called a sparse-dense fusion model, which generally includes an image branch (or called a camera branch), a point cloud branch (or called a LiDAR branch), a sparse fusion module, and a dense fusion module. The image branch is used to extract image features and implement information interaction with the sparse fusion module and the dense fusion module. The point cloud branch is used to voxelize the input point cloud, extract voxel features, and implement information interaction with the sparse fusion module and the dense fusion module.

[0029] The core idea of the present invention is to fuse image data and point cloud data from both sparse and dense perspectives, and then design a sparse fusion module and a dense fusion module. Considering that the deformable Transformer has linear complexity and a fast convergence speed, a deformable Transformer is used to collect and fuse image features and point cloud features. Generally speaking, the sparse fusion module fuses image features into valid voxels. Here, valid voxels refer to voxels with LiDAR points inside. The fused voxel features maintain the original 3D spatial structure and are further sent to the subsequent network. For the dense fusion module, BEV queries are respectively used to integrate image features, point cloud features, and past BEV features. The bird's-eye view query is a parameter in the shape of a grid and can be learned during training.

[0030] 1) Sparse fusion module

[0031] The purpose of the sparse fusion module is to maintain the 3D geometric prior during the fusion process. In the pipeline of the entire LiDAR branch, the voxel space can maintain geometric information. Since the voxel space is large and sparse, it is useless and inefficient to use all voxels as fusion queries. Therefore, in an embodiment of the present invention, an improved deformable attention module is proposed to establish an effective connection between voxels and image features.

[0032] Combination Figure 2 As described above, the sparse fusion module includes a voxel-to-image cross-attention module (V2C Cross-attention), a feed-forward layer (Feed Forward), and an add-and-normalization layer (Add&Norm). The voxel-to-image cross-attention module is used to calculate the weights of voxel features and image features during the sparse fusion process. In the sparse fusion module, non-empty voxel queries interact with image features.

[0033] 2) Dense fusion module

[0034] In one embodiment, the dense fusion module includes 6 Transformer layers. Still combined with Figure 2 As shown, each Transformer layer includes a temporal cross-attention module (Temporal Cross-attention), a LiDAR cross-attention module (LiDAR Cross-attention), a bird's-eye view to image cross-attention module (B2C Cross-attention), a feed-forward layer, and an add-and-normalization layer. The purpose of these cross-attention modules is to collect temporal information, LiDAR information, and image information respectively. In the dense fusion module, each BEV query alternately interacts with historical BEV features, LiDAR BEV features, and image features.

[0035] 3) Training of the fusion model

[0036] During the training of the fusion model, the camera branch uses, for example, a backbone and a feature pyramid network to extract multi-camera image features. The LiDAR branch converts the original LiDAR points into sparse voxel features through voxelization. Sparse fusion is performed in the voxel space, and dense fusion is performed in the bird's-eye view space to generate fused bird's-eye view features. The training process of the fusion model can use the nuScenes dataset or other datasets.

[0037] Step S130, perform object detection using the obtained fusion features.

[0038] After obtaining the fusion features, that is, the fused bird's-eye view features, they can be used for object detection in the target scene. For example, using the final fused bird's-eye view feature as the input, a 3D detection head is used to predict the 3D bounding box of the object.

[0039] To further verify the effectiveness of the present invention, experiments were conducted on the nuScenes dataset, which contains 1000 scenes. Each scene consists of RGB images from 6 cameras, with a 360° horizontal FOV (field of view) and a LiDAR point cloud, containing 1.4 million 3D bounding boxes and annotated with ten categories. In the experiment, the average precision (mAP) and nuScenes detection score (NDS) evaluation metrics provided by the dataset were used to evaluate the effectiveness of the present invention.

[0040] To prove the effectiveness of the present invention, the BEVFormer model was used for the camera branch and the TransFusion model was used for the LiDAR branch to construct a baseline network. It was verified that the SDF framework significantly improved their joint performance, exceeding the baseline by 4.3% and 2.5% in mAP and NDS respectively, and reaching the highest level of the nuScenes benchmark. In addition, ablation experiments showed that only the dense fashion has advantages in texture-intensive categories such as trucks and cars, while only the sparse fashion can play an advantage in position-sensitive categories such as trailers and pedestrians.

[0041] In summary, compared with the prior art, the present invention has the following advantages:

[0042] 1) The present invention designs a simple and effective fusion model framework that fuses image data and point cloud data from both sparse and dense perspectives. This complementary structure can share the advantages of sparse and dense fusion, thus taking into account the accuracy and efficiency of feature fusion.

[0043] 2) The sparse fusion module and dense fusion module designed by the present invention are easy to apply to existing baseline models. The sparse fusion module fuses image features into valid voxels, and the fused voxel features maintain the original 3D spatial structure. The dense fusion module can use BEV queries to integrate image, point cloud, and past bird's-eye view features respectively.

[0044] 3) It has been experimentally verified that the model framework provided by the present invention significantly improves the existing baseline models and ranks first in the 3D tasks of the nuScenes data, and can be used in various object detection fields, such as 3D detection feature fusion strategies in autonomous driving.

[0045] The present invention can be a system, method, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present invention.

[0046] A computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0047] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0048] The computer program instructions for carrying out the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, Python, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present invention.

[0049] Aspects of the present invention are described herein with reference to the flowchart and / or block diagram of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by computer - readable program instructions.

[0050] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0051] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other devices to implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0052] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As will be apparent to those of ordinary skill in the art, implementations in hardware, in software, and in a combination of software and hardware are all equivalent.

[0053] The embodiments of the present invention have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein. The scope of the present invention is defined by the appended claims.

Claims

1. A point cloud image fusion method based on sparse fusion and dense fusion, comprising the following steps: For a target scene, obtain data of two modalities, namely images and point clouds; Input the images and point clouds into a trained fusion model to obtain fusion features; Among them, the fusion model includes a first branch, a second branch, a sparse fusion module, and a dense fusion module. The first branch is used to extract image features and respectively transmit them to the sparse fusion module and the dense fusion module; the second branch is used to extract voxel features of the point cloud, and then generate bird's-eye view features based on the output of the sparse fusion module; the sparse fusion module is used to fuse the image features into the valid voxels from the second branch; the dense fusion module is used to generate fused bird's-eye view features as the fusion features based on the bird's-eye view features from the second branch and the image features from the first branch; Among them, the sparse fusion module includes a voxel-to-image cross-attention module, a feed-forward layer, and an addition and normalization layer. The voxel-to-image cross-attention module is used to calculate the weights of voxel features and image features during the sparse fusion process, and in the sparse fusion module, non-empty voxel queries interact with the image features; Among them, the voxel-to-image cross-attention module included in the sparse fusion module is used to calculate the weights of voxel features from the second branch and image features from the first branch; Among them, the dense fusion module includes multiple Transformer layers, and each Transformer layer includes a temporal cross-attention module, a radar cross-attention module, and a bird's-eye view-to-image cross-attention module; Among them, the second branch converts the point cloud into sparse voxel features through voxelization, performs sparse fusion in the voxel space by interacting with the sparse fusion module, and performs dense fusion in the bird's-eye view space by interacting with the dense fusion module, and then generates the fused bird's-eye view features; Among them, the dense fusion module uses bird's-eye view queries to integrate image features, point cloud features, and bird's-eye view features of past time series.

2. The method according to claim 1, wherein The first branch uses the BEVFormer model to extract image features and convert the image features into the bird's-eye view space.

3. The method according to claim 1, wherein The second branch uses the TransFusion model to extract voxel features of the point cloud.

4. The method according to claim 1, characterized in that, It further includes: Using the fused bird's-eye view features as input, and using a 3D detection head to predict the 3D bounding boxes of objects in the target scene.

5. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 4.

6. A computer device, comprising a memory and a processor, and a computer program capable of running on the processor is stored on the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Bird-eye view feature generation method based on multi-modal fusion

    CN115578705A

  • Multi-sensor fusion sensing method and device

    CN115861601A

Cited By

  • Multi-agent collaborative cognitive calculation method based on point cloud feature marking

    CN121053500A