3D target detection method based on point cloud and image multi-stage fusion

Through the methods of data preprocessing, feature extraction and multi-stage fusion, the feature alignment and information registration problems in 3D object detection are solved, and high-precision 3D object detection is realized, which is suitable for autonomous driving and mobile robot tasks in complex scenarios.

CN120260024APending Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510313354.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing 3D object detection methods have difficulties in feature alignment and information registration, which are difficult to meet the application needs in a complex open world, and are insufficient in computing efficiency and real-time.

Method used

The data preprocessing module is used to improve the quality of multimodal data, extract multi-scale features of images and point clouds through the feature extraction module, deeply fusion is used for multi-stage fusion module, and 3D object detection is combined with the SSD detection head to output high-precision bounding box, category and posture information.

Benefits of technology

It significantly improves the feature expression richness and alignment accuracy of multimodal data fusion, improves the accuracy and robustness of 3D object detection, and can effectively deal with detection tasks in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260024A_ABST
    Figure CN120260024A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D target detection method based on point cloud and image multi-stage fusion, and the method comprises the following modules: a data preprocessing module which is used for improving the quality of original laser point cloud data and image data; the feature extraction module is used for extracting features of the input point cloud and the image; the multi-stage fusion module is used for carrying out multi-stage feature fusion on the point cloud features and the image features and enhancing the richness and alignment precision of feature expression; and the 3D target detection module is used for outputting the 3D bounding box, category and attitude information of the target to obtain a final 3D target detection result. According to the invention, image features of image data and local features and global features of point cloud data are respectively extracted through the feature extraction module, and multi-level and multi-dimensional feature information in input data is fully mined; and the multi-stage fusion module fuses the features of the multi-modal data at different levels, so that the richness of feature expression is enhanced, and the precision and robustness of 3D target detection are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly relates to a 3D object detection method based on multi-stage fusion of point cloud and image. Background Art

[0002] With the rapid development of technologies such as autonomous driving, mobile robots, and virtual reality, 3D object detection, as one of the core tasks in these applications, has received extensive attention. 3D object detection aims to locate and classify objects of interest in three-dimensional space and generate corresponding 3D bounding boxes, providing accurate information about the position, size, and orientation of the objects. This information is crucial for subsequent scene understanding, path planning, and decision-making.

[0003] In recent years, significant progress has been made in 3D object detection methods based on deep learning. According to the processing methods of point clouds, existing methods can be divided into methods based on raw point clouds, based on voxels, based on projection views, based on graph structures, and multi-modal fusion methods. Methods based on raw point clouds (such as PointNet and PointNet++) directly process point cloud data, retaining the geometric information of the point cloud, but with a large computational overhead. Methods based on voxels (such as VoxelNet and SECOND) divide the point cloud into regular voxel grids and use 3D convolution to extract features, improving the computational efficiency, but fine-grained information may be lost during the voxelization process. Methods based on projection views (such as VeloFCN and BirdNet) project the point cloud onto a 2D plane and use a 2D convolutional network for detection, but information loss may be introduced during the projection process. Methods based on graph structures (such as Point-GNN and GraphR-CNN) use graph neural networks to process point cloud data, capable of capturing complex relationships between points, but with low computational efficiency. Multi-modal fusion methods (such as F-PointNet and MVX-Net) combine point cloud and image data, making full use of the advantages of different modalities and improving the detection performance.

[0004] At present, significant progress has been made in 3D object detection technology, but there are still many challenges. First, existing datasets have limitations in terms of scene diversity and the number of object categories, making it difficult to meet the application requirements in complex open worlds. Second, the sparsity and occlusion problems of point clouds make it difficult to detect distant or objects with fewer points. In addition, multi-modal fusion methods still face challenges in feature alignment and information registration. Finally, the real-time requirement of 3D object detection is crucial in practical applications, and how to improve the computational efficiency while ensuring accuracy still needs further research. Summary of the Invention

[0005] Aiming at the problem that the 3D object detection method based on multi-modal fusion has difficulties in feature alignment and information registration, resulting in unsatisfactory 3D object detection results, the present invention provides a 3D object detection method based on multi-stage fusion of point cloud and image, specifically including the following modules:

[0006] A data preprocessing module for improving the quality of the original lidar point cloud data and image data;

[0007] A feature extraction module for extracting the features of the input point cloud and image;

[0008] A multi-stage fusion module for performing multi-stage feature fusion on the point cloud features and image features to enhance the richness and alignment accuracy of feature expression;

[0009] A 3D object detection module for outputting the 3D bounding box, category, and pose information of the object to obtain the final 3D object detection result.

[0010] The data preprocessing module improves the quality of the original data by enhancing the sparsity of the lidar point cloud, completing the virtual point cloud, filtering the noise, and performing data augmentation on the image.

[0011] The feature extraction module uses an image feature extraction network to extract the image features of the input image, and uses a point cloud feature extraction network to extract the local point cloud features and global point cloud features of the input point cloud.

[0012] The feature fusion module adopts a three-stage fusion strategy to gradually perform deep fusion on the image features, local point cloud features, and global point cloud features extracted by the feature extraction module, fully mining the complementary information of the point cloud and image, and generating the target detection features with rich semantics and spatial consistency.

[0013] The 3D object detection module uses an SSD detection head to process the target detection features generated by the feature fusion module, predicts the 3D bounding box, category label, and pose information of the object through regression and classification tasks, and finally outputs a high-precision 3D object detection result.

[0014] This paper discloses a 3D object detection method based on multi-stage fusion of point cloud and image. This method improves the quality of the original multi-modal data through a data preprocessing module; extracts multi-scale features of the image, as well as local and global features of the point cloud through a feature extraction module, fully capturing the details and context information of the multi-modal data; and fuses the image features, point cloud local features, and global features through a multi-stage fusion module for multi-stage and deep fusion to generate high-quality target features to be detected, effectively solving the problems of feature alignment and information registration in multi-modal data fusion, and significantly improving the model's understanding ability of images and point clouds. The precise 3D bounding boxes, categories, and pose information are output through a 3D object detection module, achieving high-precision 3D object detection and being able to effectively handle detection tasks in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 . Flowchart of the 3D object detection method based on multi-stage fusion of point cloud and image of the present invention;

[0016] Figure 2 . Flowchart of the feature extraction module of the present invention;

[0017] Figure 3 . Flowchart of the multi-stage fusion module of the present invention. DETAILED IMPLEMENTATION MANNER

[0018] As Figure 1 shown is the 3D object detection flowchart of the method of the present invention: The present invention provides a 3D object detection method based on multi-stage fusion of point cloud and image, specifically including the following modules:

[0019] A data preprocessing module for improving the quality of the original lidar point cloud data and image data;

[0020] Specifically, the data preprocessing module enhances the sparsity, complements the virtual point cloud, and filters the noise of the original lidar point cloud to improve the quality of the original lidar point cloud data, and enhances the data of the original image to improve the quality of the original image data.

[0021] A feature extraction module for extracting features of the input point cloud and image;

[0022] Specifically, the feature extraction module consists of two parts: an image feature extraction network and a point cloud feature extraction network. The image feature extraction network uses a multi-scale feature extraction network composed of a ResNet-50 network and a Feature Pyramid Network (FPN) to extract image features; the point cloud feature extraction network uses PointNet++ as the basic network to extract local set features of the point cloud. Through the farthest point sampling (FPS) and K-nearest neighbor (KNN) algorithms, the point cloud is sampled and grouped. Each local region extracts features through a multi-layer perceptron (MLP) to obtain local features of the point cloud. By converting the point cloud into a sparse voxel grid and using 3D sparse convolution to extract global context information, global features of the point cloud are obtained.

[0023] The multi-stage fusion module is used to perform multi-stage feature fusion on point cloud features and image features to enhance the richness and alignment accuracy of feature expressions;

[0024] Specifically, the multi-stage fusion module consists of three stages. Each stage gradually integrates the feature information of the image and the point cloud through a specific fusion strategy to generate high-quality object detection features. In the first stage, the image features and the local features of the point cloud are initially fused through a cross-modal attention mechanism. In this process, the similarity between the image features and the local features of the point cloud is calculated first to generate an attention weight matrix, which can reflect the spatial and semantic correlation between the image features and the local features of the point cloud. Subsequently, the image features are weighted using this weight matrix to align them with the local features of the point cloud in terms of space and semantics, thereby achieving the initial fusion of the two-modal features and generating low-level fusion features. In the second stage, the low-level fusion features and the global features of the point cloud are secondarily fused through feature concatenation and MLP encoding. In this process, the low-level fusion features and the global features of the point cloud are concatenated in the channel dimension to form a feature tensor containing multi-scale information. Subsequently, the concatenated features are non-linearly transformed through a multi-layer perceptron (MLP) to further explore the deep association between the two modalities and extract more discriminative feature representations, generating multi-level fusion features. In the third stage, the multi-level fusion features and the image features are finally fused through a dual-branch attention mechanism. The dual-branch attention mechanism independently calculates the attention for the image features and the multi-level fusion features respectively to generate two sets of attention weights. Subsequently, the image features and the multi-level fusion features are fused through weighted summation to generate the features of the object to be detected.

[0025] The 3D object detection module is used to output the 3D bounding box, category, and pose information of the object to obtain the final 3D object detection result.

[0026] Specifically, the 3D object detection module uses an SSD detection head to process the features of the object to be detected, predicts the 3D bounding box, class label, and pose information of the object through regression and classification tasks, and finally outputs high-precision 3D object detection results.

Claims

1. A 3D object detection method based on multi-stage fusion of point cloud and image, specifically including the following modules: A data preprocessing module, which is used to improve the quality of the original lidar point cloud data and image data; A feature extraction module, which is used to extract the features of the input point cloud and image; A multi-stage fusion module, which is used to perform multi-stage feature fusion on the point cloud features and image features to enhance the richness and alignment accuracy of feature expression; A 3D object detection module, which is used to output the 3D bounding box, category and pose information of the object to obtain the final 3D object detection result.

2. The 3D object detection method based on multi-stage fusion of point cloud and image according to claim 1, wherein The feature extraction module consists of two parts: an image feature extraction network and a point cloud feature extraction network.

3. The 3D object detection method based on multi-stage fusion of point cloud and image according to claim 2, characterized in that, The image feature extraction network uses a multi-scale feature extraction network composed of a ResNet-50 network and a Feature Pyramid Network (FPN) to extract image features.

4. The 3D object detection method based on multi-stage fusion of point cloud and image according to claim 2, wherein The point cloud feature extraction network uses PointNet++ as the basic network to extract the local set features of the point cloud. Through the farthest point sampling (FPS) and K-nearest neighbor (KNN) algorithms, the point cloud is sampled and grouped. Each local region extracts features through a multi-layer perceptron (MLP) to output the local features of the point cloud. By converting the point cloud into a sparse voxel grid and using 3D sparse convolution to extract global context information, the global features of the point cloud are output.

5. The 3D object detection method based on multi-stage fusion of point cloud and image according to claim 1, characterized in that, The multi-stage fusion module consists of three stages: the first stage preliminarily fuses the image features with the local features of the point cloud to generate low-level fusion features; the second stage performs secondary fusion on the low-level fusion features and the global features of the point cloud to generate multi-level fusion features; the third stage performs final fusion on the multi-level fusion features and the image features to generate the features of the object to be detected.

Citation Information

Cited By

  • Dam global hidden danger detection method based on radar image interpretation and visual large model

    CN122023341A