3D target detection method using scale-dependent receptive field

By designing a scale-related receptive field and a backbone network combining voxelization, sparse convolution and self-correcting convolution modules, the problem of difficulty in detecting small targets in 3D object detection is solved, and higher detection accuracy and real-time inference speed are achieved.

CN120071087APending Publication Date: 2025-05-30TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510149189.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In 3D object detection tasks, it is difficult for the prior art to effectively detect small targets while maintaining real-time inference speed. Especially in point cloud scenarios, the contradiction between downsampling multiples and calculation costs makes it difficult to detect small targets.

Method used

By designing scale-related receptive fields, combining voxelization, sparse convolution and self-correcting convolution modules, a backbone network is built to improve the effect of 3D object detection, especially for small objects detection, while maintaining real-time inference.

Benefits of technology

It achieves higher detection accuracy and real-time inference speed, significantly improving the detection effect of 3D object detection methods, especially when detecting small targets, meeting the real-time needs of autonomous driving scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071087A_ABST
    Figure CN120071087A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of unmanned driving environment perception, and discloses a 3D target detection method using a scale-dependent receptive field, and the method comprises the following steps: S1, carrying out the coding of a point cloud through employing voxels, and calculating a mean value as the initial feature of the voxels; s2, performing feature extraction and down-sampling on voxels by using sparse convolution; s3, based on sparse convolution, designing an additional sparse convolution layer to supplement features of the small target and relieve calculation pressure; s4, designing a self-correction module with a scale-related receptive field; and step S5, designing a residual error self-correction convolution module, and constructing a backbone network according to the residual error self-correction convolution module. According to the 3D target detection method using the scale-dependent receptive field, the detection effect of the 3D target detection method is remarkably improved, the high detection speed can be kept, the real-time detection requirement is met, and the method has definite theoretical significance and important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned driving environment perception, and particularly to a 3D object detection method using scale-related receptive fields. Background Art

[0002] Object detection of 3D point clouds provides reliable object information for many scenarios such as autonomous driving, robots, and virtual reality. In recent years, object detection technology based on point clouds has developed rapidly, but still faces many challenges. Compared with 2D images, point cloud distributions are sparse and uneven. Therefore, 2D detection methods cannot be directly applied to 3D tasks. An effective solution is to voxelize and quantize the point cloud into a bird's-eye view perspective. It is usually called a voxel-based method, and many related studies have shown the feasibility of this method. Generally, voxel-based methods can be divided into single-stage and two-stage methods. Although two-stage methods can obtain better detection accuracy, their complex structures often lead to low inference efficiency and are difficult to meet the real-time requirements of autonomous driving scenarios.

[0003] In object detection tasks, the receptive field is closely related to detection performance. For each input pixel, any position outside the network's receptive field will not affect the feature of that pixel. This indicates that expanding the network's receptive field is crucial. Self-calibrating convolution is an effective way to obtain a larger receptive field. Its main idea is to use a larger receptive field to calibrate the convolution operation and achieve information integration for each point in a larger range. However, in 3D object detection tasks, the sizes of various objects are similar and very small compared to the entire scene.

[0004] Therefore, when using a self-calibrating convolution module in 3D object detection tasks, the size of the receptive field should be carefully designed. Another key issue in 3D object detection is the contradiction between the downsampling factor and the computational cost. Point cloud scenes are usually very large, and it is crucial to balance detection effects and inference speed. On the one hand, using a feature map with a higher downsampling factor can effectively reduce the computational cost, but the detection performance for small objects is not ideal. On the other hand, using a feature map with a lower downsampling factor can improve detection accuracy, but it will bring excessive computational resources and memory occupancy. However, it is still very difficult to detect small objects in the current mainstream feature maps with downsampling factors. Summary of the Invention

[0005] The purpose of the present invention is to provide a 3D object detection method using scale-related receptive fields, which can improve the 3D object detection effect, especially for small objects, by designing scale-related receptive fields, while the network maintains real-time performance during inference, having important theoretical significance and application value.

[0006] To achieve the above object, the present invention provides a 3D object detection method using scale-related receptive fields, including the following steps:

[0007] Step S1: Encode the point cloud using voxels and calculate the mean value as the initial feature of the voxels;

[0008] Step S2: Use sparse convolution to extract features and downsample the voxels;

[0009] Step S3: Based on sparse convolution, design additional sparse convolution layers to supplement the features of small objects and relieve the computational pressure;

[0010] Step S4: Design a self-correction module with scale-related receptive fields;

[0011] Step S5: Design a residual self-correction convolution module and construct a backbone network therefrom.

[0012] Preferably, in step S1, the point cloud is encoded using voxels and the mean value is calculated as the initial feature of the voxels. The specific process is as follows:

[0013] Encode the entire point cloud scene using voxels with a length × width × height of 0.05m × 0.05m × 0.05m; wherein, the features of each point cloud are as follows:

[0014] P ∈ {X, Y, Z, I};

[0015] wherein, X represents the x-coordinate value, Y represents the y-coordinate value, Z represents the z-coordinate value, and I represents the point cloud reflection intensity;

[0016] The features of each voxel are as follows:

[0017] Voxel ∈ {X mean , Y mean , Z mean , I mean};

[0018]

[0019] wherein, X mean , Y mean , Z mean , I mean respectively represent the mean values of the point cloud features in the voxel; N represents the number of point clouds in the voxel; X i , Y i , Z i , I i respectively represent the features of the i-th point cloud in the voxel;

[0020] Convert the unstructured point cloud into a regular form through voxelization to obtain a fixed-length input, reduce the amount of input data, and retain the spatial structure of the point cloud.

[0021] Preferably, in step S2, sparse convolution is used to encode and downsample the voxels. The sparse convolution includes a sparse convolution layer and a submanifold sparse convolution layer, and a normalization layer and an activation layer are linked after each convolution layer.

[0022] Preferably, when using sparse convolution to extract features and downsample the voxels, first use a layer of sparse convolution layer to downsample the voxels, and then use two layers of submanifold sparse convolution layers for feature extraction; use three groups of the above structures to downsample the voxels by 8 times.

[0023] Preferably, in step S3, based on sparse convolution, an additional sparse convolution layer is designed to supplement the features of small targets and relieve the computational pressure. The specific process is as follows:

[0024] Add a group of additional sparse convolution layers to the voxels downsampled by 4 times as a branch of the network. This branch first uses a layer of sparse convolution layer to downsample along the height;

[0025] Then, use two layers of submanifold sparse convolution layers with a normalization layer and an activation layer for feature extraction;

[0026] Finally, use another layer of sparse convolution layer with a normalization layer and an activation layer to downsample along the height.

[0027] Preferably, in step S4, the self-correction module with a scale-related receptive field consists of two convolution layers;

[0028] The first convolution layer is a convolution with a kernel size of 3×3 and a dilation of 1; the second convolution layer is a dilated convolution with a kernel size of 3×3 and a dilation of 2; a normalization layer is connected after each convolution layer; the receptive field size jointly formed by them is 7×7.

[0029] Preferably, in step S5, a residual self-correction convolution module is designed and a backbone network is constructed therefrom;

[0030] The residual self-correction convolution module uses a residual connection to merge the features output by the convolution and the output of the self-correction structure, reducing the model size.

[0031] Preferably, the backbone network first uses two residual self-correction convolution modules for feature extraction; then uses a convolution layer for downsampling; finally uses two residual self-correction convolution modules to extract features.

[0032] Therefore, the present invention adopts the above 3D object detection method using scale-related receptive fields, proposes scale-related receptive fields. Compared with the traditional method of expanding receptive fields in 3D object detection, the scale-related receptive fields can contain the entire object while reducing the introduction of background information. On this basis, the designed method for 3D object detection using scale-related receptive fields has higher detection accuracy and real-time inference speed, and has great practical application potential.

[0033] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a flowchart of a 3D object detection method using scale-related receptive fields according to the present invention;

[0035] Figure 2 It is a schematic diagram of the method for encoding point cloud using voxels and calculating the mean as the initial feature of the voxels according to the present invention;

[0036] Figure 3 It is a schematic diagram of using sparse convolution to extract features and downsample voxels according to the present invention;

[0037] Figure 4 It is a schematic diagram of designing an additional sparse convolution layer to supplement the features of small objects and relieve the computational pressure based on the advantages of sparse convolution according to the present invention;

[0038] Figure 5 It is a schematic diagram of a self-correction module with scale-related receptive fields according to the present invention;

[0039] Figure 6 It is a schematic diagram of designing a residual self-correction convolution module and constructing a backbone network therefrom according to the present invention;

[0040] Figure 7 It is the comprehensive comparison result of the present invention with the existing 3D object detection methods. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0042] As Figure 1 shown, a 3D object detection method using scale-related receptive fields includes the following steps:

[0043] Step S1: Encode the point cloud using voxels and calculate the mean as the initial feature of the voxels;

[0044] Step S2: Use sparse convolution to extract features and downsample the voxels;

[0045] Step S3: Design additional sparse convolutional layers based on sparse convolution to supplement the features of small targets and relieve the computational pressure;

[0046] Step S4: Design a self-calibration module with a scale-related receptive field;

[0047] Step S5: Design a residual self-calibration convolutional module and construct a backbone network therefrom.

[0048] Embodiment

[0049] Based on a 3D object detection method using a scale-related receptive field provided by the present invention, this embodiment is tested and verified on the KITTI dataset. The dataset contains 7,481 training samples and 7,518 test samples. The training samples are divided into a training set consisting of 3,712 samples and a validation set consisting of 3,769 samples. This embodiment detects three categories in the dataset: cars, pedestrians, and cyclists.

[0050] Step S1: Encode the point cloud using voxels and calculate the mean as the initial feature of the voxel. The implementation process is as Figure 2 shown.

[0051] Encode the entire point cloud scene using voxels with a size of 0.05 m × 0.05 m × 0.05 m (length × width × height). Among them, the features of each point cloud are as follows:

[0052] P ∈ {X, Y, Z, I};

[0053] Among them, X represents the x-coordinate value, Y represents the y-coordinate value, Z represents the z-coordinate value, and I represents the point cloud reflection intensity.

[0054] The features of each voxel are as follows:

[0055] Voxel ∈ {X mean , Y mean , Z mean , I mean};

[0056]

[0057] Among them, X mean , Y mean , Z mean , I mean respectively represent the mean of the point cloud features in the voxel; N represents the number of point clouds in the voxel; X i , Y i , Z i , I i respectively represent the features of the i-th point cloud in the voxel.

[0058] Convert the unstructured point cloud into a regular form through voxelization to obtain a fixed-length input. At the same time, voxelization can reduce the amount of input data and provide the spatial structure of the point cloud.

[0059] Step S2: Use sparse convolution to perform feature extraction and downsampling on the voxels.

[0060] Sparse convolution is used to encode and downsample the voxels. Sparse convolution includes a sparse convolutional layer and a submanifold sparse convolutional layer. A normalization layer and an activation layer are linked after each convolutional layer.

[0061] As Figure 3 shown, first use a layer of sparse convolutional layer to downsample the voxels, and then use two layers of submanifold sparse convolutional layers for feature extraction. Use three groups of such structures to downsample the voxels by 8 times.

[0062] Step S3: Based on sparse convolution, design an additional sparse convolutional layer to supplement the features of small targets and relieve the computational pressure.

[0063] The additional sparse convolution module is as Figure 4 shown. In order to better extract the features of small targets and relieve the computational pressure, add a group of additional sparse convolutional layers to the voxels downsampled 4 times as a branch of the network. This branch first uses a layer of sparse convolutional layer to downsample along the height, then uses two layers of submanifold sparse convolutional layers with a normalization layer and an activation layer for feature extraction, and finally uses a layer of sparse convolutional layer with a normalization layer and an activation layer to downsample along the height.

[0064] Step S4: Design a self-correction module with a scale-related receptive field.

[0065] The self-correction module with a scale-related receptive field is as Figure 5 shown. It consists of two convolutional layers. The first convolutional layer is a convolution with a kernel size of 3×3 and a dilation of 1. The second convolutional layer is an atrous convolution with a kernel size of 3×3 and a dilation of 2. A normalization layer is connected after each convolutional layer. The receptive field size jointly formed by them is 7×7.

[0066] The structure of the self-correction module with a scale-related receptive field is simpler and the receptive field is smaller. It can completely perceive the entire target while reducing the influence of environmental information.

[0067] Step S5: Design a residual self-correction convolution module and construct a backbone network therefrom.

[0068] The residual self-correction convolution module is as Figure 6 shown. This module uses a residual connection to merge the features output by the convolution with the output of the self-correction structure, reducing the model size.

[0069] The backbone network first uses two residual self-correction convolution modules for feature extraction, then uses a convolutional layer for downsampling, and finally uses two residual self-correction convolution modules to extract features.

[0070] As shown in Table 1 and Figure 7 as shown, a comprehensive comparison of the method proposed in the present invention with the current leading 3D object detection methods is carried out quantitatively and qualitatively. The performance of three main categories under different difficulty levels is mainly compared.

[0071] The method proposed in the present invention shows highly competitive metrics in all three categories. Especially when detecting small objects such as pedestrians and cyclists, the method proposed in the present invention has obvious improvements. Compared with the baseline network SECOND, the method proposed in the present invention has a 5% improvement in the detection results of pedestrians and a 4% improvement in the detection results of cyclists. Compared with the latest network VoxSeT, the method proposed in the present invention still has advantages in most metrics. Therefore, the method proposed in the present invention has better detection effects and real-time inference speeds.

[0072] Table 1 Comprehensive comparison results of the present invention and existing 3D object detection methods

[0073]

[0074] Therefore, the present invention adopts the above-mentioned 3D object detection method using scale-related receptive fields, and proposes scale-related receptive fields. Compared with the traditional method of expanding receptive fields in 3D object detection, the scale-related receptive fields can contain the entire object while reducing the introduction of background information; on this basis, the designed method for 3D object detection using scale-related receptive fields has higher detection accuracy and real-time inference speed; the method proposed in the present invention not only significantly improves the detection effect of 3D object detection methods, but also can maintain a relatively high detection speed, meet the real-time detection requirements, and has clear theoretical significance and important application value.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A 3D object detection method using scale-dependent receptive field, characterized in that: The following steps are involved: Step S1, use voxels to encode the point cloud and calculate the mean as the initial feature of the voxel; Step S2, using sparse convolution to perform feature extraction and downsampling of voxels; Step S3: Based on sparse convolution, an additional sparse convolution layer is designed to supplement the features of small objects and relieve the calculation pressure; Step S4, designing a self-correction module with a scale-dependent receptive field; Step S5: design a residual self-correction convolution module and construct a backbone network based on it.

2. The 3D object detection method using scale-dependent receptive field according to claim 1, characterized in that: In step S1, the point cloud is encoded using voxels, and the mean is calculated as the initial feature of the voxels. The specific process is as follows: The entire point cloud scene is encoded using voxels with a length × width × height of 0.05m × 0.05m × 0.05m; the features of each point cloud are as follows: P∈{X,Y,Z,I}; Where X represents the x-coordinate value, Y represents the y-coordinate value, Z represents the z-coordinate value, and I represents the reflection intensity of the point cloud; The characteristics of each voxel are as follows: Voxel∈{X mean ,Y mean ,Z mean ,I mean }; Among them, X mean , Y mean , Z mean , I mean They represent the mean of the point cloud features in the voxel; N represents the number of point clouds in the voxel; X i , Y i , Z i , I i Respectively represent the characteristics of the i-th point cloud in the voxel; The unstructured point cloud is converted into a regular form through voxelization to obtain a fixed-length input, reduce the amount of input data, and retain the spatial structure of the point cloud.

3. The 3D object detection method using scale-dependent receptive field according to claim 1, characterized in that: In step S2, sparse convolution is used to encode and downsample voxels. The sparse convolution includes a sparse convolution layer and a submanifold sparse convolution layer. Each convolution layer is followed by a normalization layer and an activation layer.

4. The 3D object detection method using scale-dependent receptive field according to claim 3, characterized in that: When using sparse convolution to extract features and downsample voxels, one sparse convolution layer is first used to downsample the voxels, and then two layers of submanifold sparse convolution layers are used for feature extraction; using three sets of the above structures, the voxels are downsampled by 8 times.

5. The 3D object detection method using scale-dependent receptive field according to claim 1, characterized in that: In step S3, based on sparse convolution, additional sparse convolution layers are designed to supplement the features of small objects and relieve the computational pressure. The specific process is as follows: Add an additional set of sparse convolutional layers in the voxels that are downsampled by a factor of 4, as a branch of the network that first downsamples along the height using a single sparse convolutional layer; Then, two submanifold sparse convolutional layers with normalization and activation layers are used for feature extraction; Finally, a sparse convolutional layer with normalization and activation layers is used to downsample along the height.

6. The 3D object detection method using scale-dependent receptive field according to claim 1, characterized in that: In step S4, the self-correction module with scale-dependent receptive field consists of two convolutional layers; The first convolution layer is a convolution with a kernel of 3×3 and a dilation of 1; the second convolution layer is a dilated convolution with a kernel of 3×3 and a dilation of 2; each convolution layer is connected to a normalization layer; the size of the receptive field they together form is 7×7.

7. The 3D object detection method using scale-dependent receptive field according to claim 1, characterized in that: In step S5, a residual self-correction convolution module is designed, and a backbone network is constructed based on it; The residual self-correction convolution module uses residual connections to merge the features of the convolution output with the output of the self-correction structure to reduce the model size.

8. The 3D object detection method using scale-dependent receptive field according to claim 7, characterized in that: The backbone network first uses two residual self-correction convolution modules for feature extraction; then uses a convolution layer for downsampling; and finally uses two residual self-correction convolution modules to extract features.