A semi-supervised 3D object detection method based on noisy data

Through the anti-noise instance supervision module and the dense feature consistency constraint module, the interference problem of noise pseudo-labels on the model is solved, the noise tolerance and generalization ability of the model are improved, and high-precision three-dimensional object detection is achieved.

CN117115555BActive Publication Date: 2025-08-26UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311188737.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2025-08-26
Estimated Expiration
2043-09-14

AI Technical Summary

Technical Problem

The existing semi-supervised three-dimensional object detection method is susceptible to interference in the generation process of noise pseudo-labels, affecting the convergence and final performance of the model, and has poor generalization capabilities.

Method used

The anti-noise instance supervision module and the dense feature consistency constraint module are adopted to improve the tolerance and generalization of the model's tolerance and generalization ability of noise pseudo-labels by regularizing the unlabeled data sets.

Benefits of technology

The noise tolerance and generalization ability of the model are improved, and high-precision detection on the autonomous driving data set is achieved, exceeding the performance of the existing methods and achieving an improvement of 58.01 average accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117115555B_ABST
    Figure CN117115555B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised three-dimensional target detection method based on noisy data, comprising obtaining a target detection dataset, the dataset comprising a labeled dataset and an unlabeled dataset; training a teacher model in an average teacher framework using the labeled dataset; using the trained teacher model to infer the unlabeled dataset, generating pseudo labels on the unlabeled dataset, and obtaining a pseudo-labeled dataset; sampling from the labeled dataset and the pseudo-labeled dataset, supervising the noise using a noise-resistant instance supervision module and a dense feature consistency constraint module, obtaining useful information, and thereby training a student model; and using the trained student model to perform detection tasks. By using soft task supervision of unlabeled data and unsupervised feature consistency regularization, the model's tolerance for noisy pseudo labels is improved, and the model's generalization ability is improved. The method of the present invention can effectively detect three-dimensional targets and achieve high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a semi-supervised three-dimensional target detection method based on noise data. Background Art

[0002] Object detection is a traditional task in computer vision. It aims to identify objects in images or videos, assign their corresponding categories, and present the object's location using a minimum bounding box. Applications include autonomous driving, surveillance systems, robotic perception, medical image analysis, and aerospace. Object detection can be categorized into two-dimensional and three-dimensional object detection based on its dimensionality. Three-dimensional object detection targets objects in three-dimensional space and is crucial for a wide range of applications.

[0003] Compared to traditional 3D object detection methods, semi-supervised object detection has shown great promise in recent years due to its simplicity and lack of reliance on expensive annotations. Currently, mainstream semi-supervised object detection is based on two frameworks: Mean-Teacher (MT) and Pseudo-Labeling (PL).

[0004] Both frameworks have significant flaws: the Average Teacher (MT) model employs a teacher-student paradigm, generating supervisory signals on unlabeled data in an end-to-end training approach. However, this model is not model-agnostic, resulting in poor generalization. The Pseudo-Label (PL) model first trains the model on labeled data and then generates pseudo-labels on unlabeled data for subsequent training. It can be easily applied to any detector, but the final performance is often limited by the quality of the pseudo-labels. Although methods have emerged to improve the quality of pseudo-labels, the pseudo-label generation process inevitably generates noise, which interferes with model convergence and even affects the final performance. Summary of the Invention

[0005] To solve the above problems, the present invention provides a semi-supervised three-dimensional object detection method based on noisy data, in order to design a three-dimensional object detection model with good generalization ability and high tolerance to noisy pseudo labels.

[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0007] A semi-supervised 3D object detection method based on noisy data comprises the following steps:

[0008] Step 1: Obtain a target detection dataset, which includes a labeled dataset and an unlabeled dataset;

[0009] Step 2: Use the labeled dataset obtained in step 1 to train the teacher model in the average teacher framework;

[0010] Step 3: Use the teacher model trained in step 2 to infer the unlabeled dataset obtained in step 1, generate pseudo labels on the unlabeled dataset, and obtain a pseudo-labeled dataset;

[0011] Step 4: Sample from the labeled dataset obtained in step 1 and the pseudo-labeled dataset obtained in step 3, use the noise-resistant instance supervision module and the dense feature consistency constraint module to supervise the noise, obtain useful information, and use the classification loss function to Regression loss function L reg And the consistency loss function L consist Train the student model;

[0012] Step 5: Use the student model trained in step 4 to perform the detection task and obtain the detection results.

[0013] Furthermore, in step four, the noise-resistant instance supervision module is divided into a classification module and a regression module. Classification by the classification module and regression by the regression module are two processes in target detection, and there is no order of precedence. Classification determines the category of the detection target, and regression determines the specific detection frame of the detection target.

[0014] Furthermore, the classification module of the noise-resistant instance supervision module in step 4 uses the confidence c as an indicator to measure the quality of the pseudo-label. Based on the confidence c and the intersection-over-union ratio τ between the student model prediction result and the pseudo-label matched by the student model, the classification label is softened to a value in the range of 0 to 1, and is regarded as a combination of the quality of the ground truth box itself and the learning ability of the student model.

[0015] A variant of the cross entropy loss function is used to supervise non-discrete classification labels, which are represented by quality scores. The specific form is as follows:

[0016]

[0017]

[0018] in, represents the quality score predicted by the teacher model, y represents the quality score predicted by the student model, α is a configurable hyperparameter, β is a modulation parameter, This is the classification loss.

[0019] Furthermore, α is set to 0.75.

[0020] Furthermore, the regression module of the noise-resistant instance supervision module in step 2 performs network prediction in the student model for each bounding box, modeling it as a Gaussian distribution h given a feature vector x, which is as follows:

[0021]

[0022] where μ(x) and σ(x) represent the mean and variance of each regression term predicted by the network in the student model, The symbol for Gaussian distribution;

[0023] The regression loss L reg Converted into negative log-likelihood loss, the specific form is as follows:

[0024]

[0025] Furthermore, in step 4, the dense feature consistency constraint module uses the lidar point cloud data as input and enhances the input data by rotation and flipping operations. For a given point cloud frame P and a set of data enhancement strategies A, two transformations A1 and A2 are randomly extracted from A, and A1 and A2 are applied to P to generate two different point cloud views P1 and P2. The enhanced point cloud is then input into the point feature extractor to generate the features of the bird's-eye view; the obtained bird's-eye view features are returned to the original space, and the transformation process is recorded to obtain the returned features. and From this, the loss function is derived, which is the pixel-level feature consistency constraint L with standard Euclidean distance loss consist :

[0026]

[0027] Furthermore, a foreground focus mask is introduced to selectively regularize the enhanced bird’s-eye view features, where the center of each ground-truth (x i ,y i ) to draw a Gaussian distribution:

[0028]

[0029] where σ i is a constant representing the standard deviation of the object size, is the reference center point, φ i,x,y Represents the Gaussian distribution of the coordinate (x, y) position at the i-th latitude.

[0030] Furthermore, σ i =2.

[0031] Furthermore, by taking the maximum value in dimension i, all φ i,x,yMerged into a mask Φ, we get the final dense feature consistency constraint L consist :

[0032]

[0033] Where H and W represent the height and width of the feature image respectively, φ xy Represents the mask centered at (x, y) on the feature image.

[0034] Compared with the prior art, the present invention has the following beneficial effects:

[0035] 1. Based on the semi-supervised 3D object detection framework, this paper designs a noisy pseudo-labeling semi-supervised 3D object detection method. By treating the semi-supervised learning task as a noisy learning task, this paper proposes two core modules to overcome the fuzzy detection problem: a noise-resistant instance supervision module and a dense feature consistency constraint module. Through soft task supervision of unlabeled data and unsupervised feature consistency regularization, the model's generalization ability is improved; the model's tolerance to noise is also increased, reducing the impact of noise on model performance.

[0036] 2. The method described in this paper can effectively detect three-dimensional objects with high accuracy. By implementing our method on the sparsely embedded convolutional detection (SECOND) 3D object detector, we achieved an ultra-high accuracy of 58.01 mean average precision (mAP) on the currently mainstream autonomous driving dataset ONCE, outperforming previous semi-supervised detection methods and improving the mainstream self-training method NoisyStudent by 2.5 mAP. On the more powerful detector CenterPoint, our method also achieved a 1.8 mAP improvement over NoisyStudent. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is the main processing process of the method described in the present invention;

[0038] Figure 2 This is a framework diagram of the anti-noise instance supervision module of the present invention;

[0039] Figure 3 This is a framework diagram of the dense feature consistency constraint module described in the present invention. DETAILED DESCRIPTION

[0040] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0041] Explanation of terms:

[0042] (1) LiDAR point cloud data is a data set of spatial points scanned by a 3D LiDAR device. Each point contains 3D coordinate information, also known as the three elements X, Y, and Z. Some also contain color information, reflection intensity information, echo count information, etc.

[0043] (2) CenterPoint is a laser point cloud 3D target detection and tracking algorithm framework;

[0044] (3) ONCE (One millioN sCenEs) dataset is a large-scale autonomous driving dataset with 2D+3D object annotations open sourced by Huawei;

[0045] (4) ProficientTeacher is a semi-supervised 3D detection model;

[0046] (5) Quality Focal Loss is a variant of the cross entropy loss function that optimizes the continuous value labels of the joint classification-quality scores;

[0047] (6) Gaussian Focal Loss is a loss function used for target detection tasks, which is based on

[0048] An improved version of Focal Loss. Focal Loss is a loss function used to solve the problem of class imbalance, which focuses on samples that are difficult to classify by adjusting the weights of positive and negative samples.

[0049] (7)NLL Loss is called Negative Log-Likelihood Loss, which means negative log-likelihood loss.

[0050] This embodiment provides a semi-supervised three-dimensional object detection method based on noisy data. This method improves the model's tolerance to noisy labels by transforming instance supervision of an unlabeled dataset into noise-resistant supervision. In data augmentation, the bird's-eye-view (BEV) features are inverted according to data transformation, and then dense pixel-by-pixel regularization is performed. Consistency constraints are applied to the feature layer to avoid performance damage to the strategy caused by inaccurate labels.

[0051] 1. Semi-supervised 3D object detection method based on noisy data

[0052] In order to achieve the above technical objectives, the technical solution of the present invention is:

[0053] like Figure 1As shown in the figure, a semi-supervised 3D object detection method based on noisy data is proposed. After obtaining a dataset containing labeled and unlabeled data, the labeled dataset is first used to train a teacher model. The trained teacher model is then used to perform inference on the unlabeled dataset. Inference on the unlabeled dataset generates pseudo-labels, resulting in a pseudo-labeled dataset. Subsequently, uniform sampling is performed on the labeled and pseudo-labeled datasets, which are used as input to train the student model, ultimately resulting in a 3D detection model with good generalization capabilities. During the student model training process, the quality of the pseudo-labels is not directly improved. Instead, useful information is directly learned from the noise. Specifically, this is achieved through two core modules: the noise-resistant instance supervision module and the dense feature consistency constraint module. The two modules supervise the noise simultaneously during the training process:

[0054] 1.1 Noise-Resistant Instance Supervision Module

[0055] The noise-resistant instance supervision module improves the model's tolerance to noisy labels by transforming instance supervision on unlabeled datasets into noise-resistant supervision.

[0056] like Figure 2 As shown in Figure 2, the noise-resistant instance supervision module is mainly divided into a classification module and a regression module. Specifically:

[0057] a. In the classification module, the confidence level c is used as a metric to measure the quality of pseudo-labels. Based on the confidence level c and the intersection over union (IoU) τ between the student model's prediction and its matching pseudo-label, the classification label is softened to a value between 0 and 1. This is considered a combination of the quality of the ground truth (GT) box itself and the learning ability of the student model.

[0058] b. Use Quality Focal Loss to supervise non-discrete classification labels. Its specific form is as follows:

[0059]

[0060]

[0061] in, represents the quality score predicted by the teacher model, y represents the quality score predicted by the student model, α is a configurable hyperparameter, generally set to 0.75, and β is a modulation parameter. This loss function construction method can be easily extended to other continuous versions of cross entropy loss, such as Gaussian Focal Loss.

[0062] c. In addition to the classification loss, since the bounding box boundary target contains seven degrees of freedom and has fewer training samples, it may present higher ambiguity and produce misleading regression targets. To address this issue, the deterministic regression task is transformed into a probabilistic optimization task, which can effectively handle misleading regression targets. Specifically, the network prediction of each bounding box is modeled as a Gaussian distribution h for a given feature vector x, which is as follows:

[0063]

[0064] where μ(x) and σ(x) represent the mean and variance of each regression term predicted by the network.

[0065] d. Regression loss L reg Transformed into negative log-likelihood loss (NLL loss), its objective function is to maximize the likelihood value of each GT h in the predicted distribution. The specific form is as follows:

[0066]

[0067] By converting the deterministic regression task into a probability estimation problem, the model has a stronger tolerance to noise information in the training data, thus achieving better performance.

[0068] 1.2 Dense Feature Consistency Constraint Module

[0069] like Figure 3 As shown in the figure, based on the strategy of using unsupervised learning to obtain useful information about label-independent features, a dense feature consistency constraint module is designed. By inverting the BEV features according to data transformation in data augmentation and then performing dense pixel-by-pixel regularization, the consistency constraint is applied to the feature layer to avoid the performance damage of the strategy when the labels are not accurate enough.

[0070] a. Using the LiDAR point cloud as input, we can usually enhance the input data by using operations such as rotation and flipping. For a given point cloud frame P and a set of data enhancement strategies A, we randomly extract two transformations A1 and A2 from A and apply them to P to generate two different point cloud views P1 and P2. The enhanced point cloud is then input into the point feature extractor to generate the BEV feature F. Once the BEV feature is obtained, it is simply returned to the original space and the transformation process is recorded to obtain the returned feature. and From this, we derive the pixel-level feature consistency constraint L with standard Euclidean distance (L2) loss. consist :

[0071]

[0072] b. Considering that point-based 3D features can only preserve meaningful information when points exist, we further introduce a foreground focus mask to selectively regularize the enhanced BEV features. Specifically, for each GT center (x i ,y i ) plots a Gaussian distribution:

[0073]

[0074] where σ i is a constant (set to 2) representing the standard deviation of the object size, is the reference center point, φ i,x,y Represents the Gaussian distribution of the coordinate (x, y) position at the i-th latitude.

[0075] c. Since the feature map is class-independent, by taking the maximum value in the i dimension, all φ i,x,y Merged into a mask Φ, we get the final dense feature consistency constraint (loss function) L consist for:

[0076]

[0077] Where H and W represent the height and width of the feature image respectively, φ xy represents the mask centered at (x, y) on the feature image, L consist It is the consistency loss function used to supervise noisy data.

[0078] By aligning dense pixel-wise features on the BEV space, the model can gradually learn the ability to extract transformation-invariant features and fully utilize unlabeled data in a self-supervised manner.

[0079] 2. Test Method

[0080] In the implementation, the ONCE dataset was used for testing. This dataset contains 1 million LiDAR point clouds and 7 million paired images, of which only 15,000 samples have 3D bounding box annotations. During training, a teacher network was first trained on the ONCE dataset for 80 epochs (all data is fed into the network, completing a forward computation and backpropagation process). Pseudo-labels were then obtained on the unlabeled dataset using the Spatial-Temporal Ensemble (STE) module proposed in ProficientTeacher. Following the official ONCE benchmark, the student model was initialized from a pre-trained checkpoint on the full labeled dataset. The student model was trained on the small, medium, and large datasets of the ONCE dataset for 25, 50, and 75 epochs, respectively. The initial learning rate was 1e-4, and the pseudo-labels were updated every 25 epochs. The entire experiment was conducted on a machine with 8 NVIDIA V100 GPUs.

[0081] In summary, the present invention proposes a semi-supervised three-dimensional object detection method based on noisy pseudo-labeling. By treating semi-supervised learning as a noisy learning task, two core modules are proposed to overcome the problem of fuzzy detection: a noise-resistant instance supervision module and a dense feature consistency constraint module. Through soft task supervision of unlabeled data and unsupervised feature consistency regularization, the model's tolerance to noisy pseudo-labels is improved, and the model's generalization ability is improved. Finally, a large number of experiments on the ONCE dataset demonstrate the effectiveness and generalization of our method. This method can provide a new perspective for dealing with pseudo-labels with insufficient accuracy in semi-supervised three-dimensional object detection.

[0082] The above-described embodiments are merely descriptions of preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.

Claims

1. A semi-supervised 3D object detection method based on noisy data, comprising the following steps: Step 1: Obtain a target detection dataset, which includes a labeled dataset and an unlabeled dataset; Step 2: Use the labeled dataset obtained in step 1 to train the teacher model in the average teacher framework; Step 3: Use the teacher model trained in step 2 to infer the unlabeled dataset obtained in step 1, generate pseudo labels on the unlabeled dataset, and obtain a pseudo-labeled dataset; Step 4: Sample from the labeled dataset obtained in step 1 and the pseudo-labeled dataset obtained in step 3, use the noise-resistant instance supervision module and the dense feature consistency constraint module to supervise the noise, obtain useful information, and use the classification loss function to Regression loss function L reg And the consistency loss function L consist Train the student model; Step 5: Use the student model trained in step 4 to perform the detection task and obtain the detection results; The noise-resistant instance supervision module is divided into a classification module and a regression module. Classification by the classification module and regression by the regression module are two processes in target detection, with no order of precedence. Classification determines the category of the detected target, and regression determines the specific detection box of the detected target. The classification module uses confidence c as an indicator to measure the quality of pseudo labels. Based on confidence c and the intersection-over-union ratio τ between the student model prediction result and the pseudo labels matched by the student model, the classification label is softened to a value in the range of 0 to 1, and is regarded as a combination of the quality of the ground truth box itself and the learning ability of the student model. The regression module models each bounding box as a Gaussian distribution h given a feature vector x: σ(x)); where μ(x) and σ(x) represent the mean and variance of each regression term predicted by the network in the student model, The symbol for Gaussian distribution; The dense feature consistency constraint module uses lidar point cloud data as input and enhances the input data using rotation and flipping operations. For a given point cloud frame P and a set of data enhancement strategies A, two transformations A1 and A2 are randomly extracted from A and A1 and A2 are applied to P to generate two different point cloud views P1 and P2. The enhanced point cloud is then input into the point feature extractor to generate bird's-eye view features; the obtained bird's-eye view features are returned to the original space, and the transformation process is recorded to obtain the returned features. and 2. The semi-supervised three-dimensional object detection method based on noise data according to claim 1, characterized in that A variant of the cross entropy loss function is used to supervise non-discrete classification labels, which are represented by quality scores. The specific form is as follows: in, represents the quality score predicted by the teacher model, y represents the quality score predicted by the student model, α is a configurable hyperparameter, β is a modulation parameter, This is the classification loss.

3. The semi-supervised three-dimensional object detection method based on noise data according to claim 2, characterized in that: Alpha is set to 0.

75.

4. The semi-supervised 3D object detection method based on noise data according to claim 1, characterized in that: The regression loss L reg Converted into negative log-likelihood loss, the specific form is as follows:

5. The semi-supervised three-dimensional object detection method based on noise data according to claim 1, characterized in that By features and The loss function is derived, namely the pixel-level feature consistency constraint L with standard Euclidean distance loss consist :

6. The semi-supervised three-dimensional object detection method based on noise data according to claim 5, characterized in that: A foreground focus mask is introduced to selectively regularize the enhanced bird’s-eye view features, where the center of each ground-truth (x i ,y i ) to draw a Gaussian distribution: where σ i is a constant representing the standard deviation of the object size, is the reference center point, φ i,x,y Represents the Gaussian distribution of the coordinate (x, y) position at the i-th latitude.

7. The semi-supervised three-dimensional object detection method based on noise data according to claim 6, characterized in that: s i =2.

8. The semi-supervised three-dimensional object detection method based on noise data according to claim 7, characterized in that: By taking the maximum value in dimension i, all φ i,x,y Merged into a mask Φ, we get the final dense feature consistency constraint L consist : Where H and W represent the height and width of the feature image respectively, φ xy Represents the mask centered at (x, y) on the feature image.

Citation Information

Patent Citations

  • Training method of general semi-supervised target detection framework

    CN116563634A

  • Teacher-student network method for semi-supervised directive target detection

    CN116563687A