3D small target detection method based on deep learning

Through multi-scale feature fusion and attention mechanism, combined with point cloud data augmentation and RPN generation candidate regions, the problem of low detection accuracy of small and low targets in the existing 3D object detection methods is solved, and efficient end-to-end object detection is achieved, which is suitable for a variety of scenarios.

CN120388161APending Publication Date: 2025-07-29EAST CHINA UNIV OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510451671.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing 3D object detection methods are difficult to effectively detect small objects in three-dimensional space, are susceptible to background noise interference and have low detection accuracy.

Method used

Using multi-scale feature fusion, attention mechanism and point cloud data enhancement, we use multi-level feature extraction network, calculate the importance of feature maps using channel attention mechanism, and generate candidate areas for classification and regression, and combine RPN to generate candidate areas to achieve end-to-end object detection.

Benefits of technology

It improves the detection accuracy and robustness of small targets, simplifies the detection process, improves the detection efficiency, and is suitable for the detection of small targets and large targets, as well as three-dimensional target detection in different scenarios.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The invention provides a 3D small target detection method based on deep learning, and the method comprises the following steps: 1, data preprocessing, point cloud data collection, data cleaning and enhancement, and scale difference elimination; the feature extraction capability of the model on the small target is enhanced; and step 3, calculating the importance of each channel by using a channel attention mechanism, weighting a feature map, and highlighting the features of important areas and inhibiting the interference of irrelevant areas through feature weighting by the model. Through multi-scale feature fusion and the attention mechanism, the feature extraction capability of the model for small targets is enhanced, and the accuracy of feature extraction is improved. Through the technical means of data enhancement, joint optimization and the like, the robustness of the model to noise and interference is improved, the end-to-end target detection process is realized, the detection process is simplified, and the detection efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a 3D small target detection method based on deep learning. Background Art

[0002] With the rapid development of deep learning technology, object detection has made remarkable progress in the field of two-dimensional images. However, in three-dimensional scenes, especially for the detection of small targets, there are still many challenges. Small targets have less information in three-dimensional space, are easily interfered by background noise, and most existing 3D object detection methods are designed for large targets and are difficult to be directly applied to small target detection. Summary of the Invention

[0003] In view of this, the present invention proposes a 3D small target detection method based on deep learning. Through multi-scale feature fusion, attention mechanism, and point cloud data augmentation, the detection accuracy and robustness of small targets in three-dimensional space are significantly improved.

[0004] The specific technical solutions are as follows:

[0005] A 3D small target detection method based on deep learning, comprising the following steps:

[0006] Step 1: Data preprocessing, collecting point cloud data, performing data cleaning and augmentation, and eliminating scale differences

[0007] Step 2: By constructing a multi-level feature extraction network, fusing features of different scales, thereby enhancing the model's feature extraction ability for small targets;

[0008] Step 3: Using the channel attention mechanism to calculate the importance of each channel and weighting the feature map. Through feature weighting, the model can highlight the features of important regions and suppress the interference of irrelevant regions;

[0009] Step 4: Using the RPN to generate candidate regions, and classifying and regressing each candidate region. Through joint optimization of the classification and regression tasks, the model can achieve end-to-end object detection, perform non-maximum suppression on the detection results, remove overlapping detection frames, and output the final 3D small target detection results.

[0010] Further, in the first step, sensors such as lidar and depth cameras are used to collect point cloud data in three-dimensional space, removing useless data such as noise points and outliers to ensure the accuracy and integrity of the point cloud data. Through operations such as rotation, scaling, and translation, the diversity of the point cloud data is enhanced, and the generalization ability of the model is improved.

[0011] Further, in the first step, according to the characteristics of small targets, a data augmentation strategy is adopted to increase the number of small targets, adjust the position and size of small targets, etc., so as to further highlight the characteristics of small targets.

[0012] Further, in the second step, a multi-level feature extraction network is constructed using a convolutional neural network (CNN) or a point cloud neural network in deep learning to fuse features at different levels, so as to make full use of information at different scales. By fusing low-level spatial detail information and high-level semantic information, the feature extraction ability of the model for small targets is enhanced.

[0013] Further, in the fourth step, the RPN is used to generate candidate regions: a Region Proposal Network (RPN) is used to generate a series of candidate regions. The RPN can quickly screen out regions that may contain targets, providing a basis for subsequent classification and regression tasks. Each candidate region is subjected to classification and regression processing. The classification task is used to determine whether the candidate region contains a target, and the regression task is used to adjust the position and size of the candidate region to make it fit the target more accurately. By jointly optimizing the classification and regression tasks, an end-to-end object detection is achieved. During the training process, an appropriate loss function is used to optimize the model to improve the detection accuracy and robustness. The non-maximum suppression process is performed on the detection results to remove overlapping detection frames, ensuring that each target is detected only once, and the final 3D small target detection results are output, including the position, size, and category of the target.

[0014] Adopting the above technical solutions, the following beneficial effects are obtained:

[0015] Through multi-scale feature fusion and attention mechanism, the present invention enhances the feature extraction ability of the model for small targets, improves the detection accuracy. Through technical means such as data augmentation and joint optimization, the robustness of the model to noise and interference is improved, an end-to-end object detection process is realized, the detection process is simplified, and the detection efficiency is improved. This method is not only applicable to the detection of small targets, but also can be extended to the detection of large targets and 3D object detection tasks in different scenarios. Specific embodiments

[0016] The technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0017] A 3D small target detection method based on deep learning includes the following steps:

[0018] Step 1: Data preprocessing. Collect point cloud data, perform data cleaning and augmentation, and eliminate scale differences.

[0019] Step 2: By constructing a multi-level feature extraction network, fuse features of different scales to enhance the model's feature extraction ability for small targets.

[0020] Step 3: Use the channel attention mechanism to calculate the importance of each channel and weight the feature map. Through feature weighting, the model can highlight the features of important regions and suppress the interference of irrelevant regions.

[0021] Step 4: Use the RPN to generate candidate regions and perform classification and regression on each candidate region. By jointly optimizing the classification and regression tasks, the model can achieve end-to-end object detection. Perform non-maximum suppression on the detection results to remove overlapping detection boxes and output the final 3D small target detection results. In Step 1, use sensors such as lidar and depth cameras to collect point cloud data in three-dimensional space, remove useless data such as noise points and outliers to ensure the accuracy and integrity of the point cloud data, and enhance the diversity of the point cloud data through operations such as rotation, scaling, and translation to improve the generalization ability of the model. In Step 1, according to the characteristics of small targets, adopt a data augmentation strategy to increase the number of small targets, adjust the position and size of small targets, etc., to further highlight the features of small targets. In Step 2, use convolutional neural networks (CNNs) or point cloud neural networks in deep learning to construct a multi-level feature extraction network and fuse features at different levels to make full use of information at different scales. By fusing low-level spatial detail information and high-level semantic information, enhance the model's feature extraction ability for small targets. In Step 4, use the RPN to generate candidate regions: Use the Region Proposal Network (RPN) to generate a series of candidate regions. The RPN can quickly screen out regions that may contain targets and provide a basis for subsequent classification and regression tasks. Perform classification and regression processing on each candidate region. The classification task is used to determine whether the candidate region contains a target, and the regression task is used to adjust the position and size of the candidate region to make it fit the target more accurately. By jointly optimizing the classification and regression tasks, achieve end-to-end object detection. During the training process, use an appropriate loss function to optimize the model to improve the detection accuracy and robustness. Perform non-maximum suppression processing on the detection results to remove overlapping detection boxes and ensure that each target is detected only once. Output the final 3D small target detection results, including the position, size, category, etc. of the target.

[0022] In this embodiment, a channel attention mechanism (such as SE Block) is used to calculate the importance weights of each channel. The SE Block calculates the importance weights of each channel through steps such as global average pooling, fully connected layers, ReLU activation functions, and Sigmoid activation functions. These weights reflect the degree of importance of different channels in feature representation.

[0023] In this embodiment, significant performance improvements have been achieved on multiple datasets, especially in the detection task of small targets.

[0024] The present invention enhances the feature extraction ability of the model for small targets and improves the detection accuracy through multi-scale feature fusion and attention mechanisms. Through technical means such as data augmentation and joint optimization, the robustness of the model to noise and interference is improved, an end-to-end object detection process is realized, the detection process is simplified, and the detection efficiency is improved. This method is not only applicable to the detection of small targets, but can also be extended to the detection of large targets and 3D object detection tasks in different scenarios.

[0025] The basic principles and main features of the present invention have been described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the invention claimed is defined by the appended claims and their equivalents.

Claims

1. A 3D small target detection method based on deep learning, characterized in that, It includes the following steps: Step 1: Data preprocessing. Collect point cloud data, perform data cleaning and enhancement, and eliminate scale differences. Step 2: By constructing a multi-level feature extraction network, fuse features at different scales to enhance the model's feature extraction ability for small targets. Step 3: Use the channel attention mechanism to calculate the importance of each channel and weight the feature map. Through feature weighting, the model can highlight the features of important regions and suppress the interference of irrelevant regions. Step 4: Use the RPN to generate candidate regions and perform classification and regression on each candidate region. By jointly optimizing the classification and regression tasks, the model can achieve end-to-end object detection. Perform non-maximum suppression on the detection results, remove overlapping detection boxes, and output the final 3D small target detection results.

2. The 3D small target detection method based on deep learning according to claim 1, wherein In the above Step 1, use sensors such as lidar and depth cameras to collect point cloud data in three-dimensional space, remove useless data such as noise points and outliers, ensure the accuracy and integrity of the point cloud data, and enhance the diversity of the point cloud data through operations such as rotation, scaling, and translation to improve the generalization ability of the model.

3. A 3D small target detection method based on deep learning according to claim 2, characterized in that, In the above Step 1, according to the characteristics of small targets, adopt a data augmentation strategy to increase the number of small targets, adjust the position and size of small targets, etc., to further highlight the features of small targets.

4. A 3D small target detection method based on deep learning according to claim 1, characterized in that In the above Step 2, use a convolutional neural network (CNN) or a point cloud neural network in deep learning to construct a multi-level feature extraction network, fuse features at different levels to make full use of information at different scales, and enhance the model's feature extraction ability for small targets by fusing low-level spatial detail information and high-level semantic information.

5. A 3D small target detection method based on deep learning according to claim 1, characterized in that, In the above Step 4, use the RPN to generate candidate regions: adopt a region proposal network to generate a series of candidate regions. The RPN can quickly screen out regions that may contain targets, providing a basis for subsequent classification and regression tasks. Perform classification and regression processing on each candidate region. The classification task is used to determine whether the candidate region contains a target, and the regression task is used to adjust the position and size of the candidate region to make it fit the target more accurately. By jointly optimizing the classification and regression tasks, achieve end-to-end object detection. During the training process, use an appropriate loss function to optimize the model to improve the detection accuracy and robustness. Perform non-maximum suppression processing on the detection results, remove overlapping detection boxes, ensure that each target is detected only once, and output the final 3D small target detection results, including the position, size, and category of the target.