Semi-supervised 3D object detection method and system in indoor scene and storage medium

CN117671239BActive Publication Date: 2026-09-29UNIV OF CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311619574.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2026-09-29
Estimated Expiration
2043-11-30

AI Technical Summary

Technical Problem

尽管以上方法都在室内场景的半监督检测方面取得了较大的进步,但目前的方法仍存在以下问题:(1)没有对室内场景中的遮挡物体和小物体的检测提出针对性的解决方法;(2)没有充分利用已标注数据的有效信息;(3)在对教师网络和学生网络的预测结果进行匹配时,粗暴简单的匹配方法容易造成匹配歧义问题,匹配歧义是指在学生网络中的同一预测结果与教师网络中不同的预测结果相匹配并计算损失函数,会削弱网络的收敛速度,并降低检测精度

Benefits of technology

[0035]1、本发明采用随机删除网格内所有点的数据增强方法,模拟了室内场景中遮挡物体和小物体的稀疏性,以提升对于这些物体的检测性能。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117671239B_ABST
    Figure CN117671239B_ABST
Patent Text Reader

Abstract

The application relates to a semi-supervised 3D target detection method and system in an indoor scene and a storage medium, which comprises the following steps: acquiring an input image and a label, pretraining through a full-supervised network, and dividing all point cloud data into labeled data and unlabeled data; inputting the point cloud data into a student network after sequentially performing random down-sampling, random flipping and random rotation to calculate a prediction loss, inputting the point cloud data into a teacher network after only performing random down-sampling, matching the prediction points of the student network and the teacher network, and calculating a consistency loss; randomly deleting the point cloud in a grid of the labeled data, inputting the point cloud into an increased auxiliary network after random down-sampling to generate a prediction result, and calculating a voting consistency loss between the prediction result generated by the auxiliary network and the student network; adding all the calculated losses to obtain a total loss of the semi-supervised network, inputting a target image to be detected, marking a positioning frame and classification of the target through the weight obtained through training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a semi-supervised 3D target detection method, system and storage medium for indoor scenes. Background Technology

[0002] In recent years, with the development of deep learning technology and the massive growth of data, 3D object detection methods have made great progress. As the foundation of many downstream tasks, 3D object detection has attracted widespread attention from researchers. Since point cloud data can preserve the original features of 3D data to the greatest extent, current 3D object detection methods mainly use point cloud data as input. However, the sparsity, irregularity, and disorder inherent in point clouds bring a series of challenges to object detection. Existing 3D object detection methods all rely on large amounts of carefully labeled point cloud data, but labeling 3D scenes is time-consuming and laborious. Therefore, semi-supervised 3D object detection has gradually become a research hotspot and has achieved significant results. Current semi-supervised 3D object detection methods mainly adopt a general teacher-student network architecture, achieving detection of unlabeled scenes by constraining the consistency of prediction results between the two networks. However, compared to outdoor scenes, indoor scenes contain many small objects and occluded objects, making the detection of these objects particularly difficult in a semi-supervised architecture. Currently, there is a lack of research on semi-supervised detection methods for occluded and small objects in indoor scenes. The limited number of points on the surface of these objects exacerbates the sparsity of the point cloud, which greatly affects the performance of semi-supervised object detection.

[0003] The existing literature SESS (Zhao, N.; Chua, T.-S.; and Lee, GH2020. Sess: Self-ensembling semi-supervised 3D object detection. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 11079–11087) proposes a semi-supervised 3D object detection method for indoor scenes. This method uses a teacher-student network, with a small amount of labeled data and a large amount of unlabeled data as input. The network is trained to generate prediction results by learning the consistency of the prediction results of the student network and the teacher network. To improve detection accuracy, the existing paper 3DIoUMatch (Wang, H.; Cong, Y.; Litany, O.; Gao, Y.; and Guibas, LJ2021. 3dioumatch: Leveraging iou prediction for semisupervised 3dobjectdetection. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 14615–14624) proposes a two-stage semi-supervised object detection network based on SESS. By selecting and fine-tuning the predicted bounding boxes, it achieves better performance in indoor scene detection. Although the above methods have made significant progress in semi-supervised detection of indoor scenes, the current methods still have the following problems: (1) No specific solutions have been proposed for the detection of occluded objects and small objects in indoor scenes; (2) The effective information of the labeled data has not been fully utilized; (3) When matching the prediction results of the teacher network and the student network, the crude and simple matching method is prone to matching ambiguity. Matching ambiguity means that the same prediction result in the student network is matched with different prediction results in the teacher network and the loss function is calculated, which will weaken the convergence speed of the network and reduce the detection accuracy.

[0004] Based on the above analysis, researching semi-supervised detection methods to improve the accuracy of indoor scene detection is particularly important. Therefore, a new semi-supervised 3D object detection method is urgently needed to further target the detection of occluded and small objects, thereby improving detection performance. Summary of the Invention

[0005] To address the aforementioned problems, the purpose of this invention is to provide a semi-supervised 3D target detection method, system, and storage medium for indoor scenes, which can improve the detection performance of occluded objects and small objects, achieving results close to those of fully supervised target detection networks.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a semi-supervised 3D target detection method in indoor scenes, comprising: acquiring the input image and labels of the network; completing pre-training through a fully supervised network; dividing all point cloud data of the input image into labeled data and unlabeled data; sequentially performing random downsampling, random flipping, and random rotation on all point cloud data, and inputting it into the student network to calculate the prediction loss; simultaneously, performing only random downsampling on all point cloud data and inputting it into the teacher network; matching the prediction points of the student network and the teacher network and calculating the consistency loss; randomly deleting point clouds within the grid of the labeled data, performing random downsampling, and inputting it into an added auxiliary network to generate prediction results; calculating the voting consistency loss between the prediction results generated by the auxiliary network and the student network; summing all the calculated losses as the overall loss of the semi-supervised network; inputting the target image to be detected; and using the weights obtained through training, marking the target's bounding box and classification.

[0007] Furthermore, the student network is input to calculate the predicted loss, including:

[0008] The student network is a fully supervised 3D object detection network based on Hough voting;

[0009] The predicted loss includes voting loss L v Target confidence loss L o Classification loss L c and positioning loss L s The prediction loss for the labeled data is: L sup =L v +λ1L o +λ2L c +λ3L s , where λ1, λ2, and λ3 represent the weights of the loss function.

[0010] Furthermore, the teacher network maintains a completely identical structure to the student network, and its parameters are calculated from the student network parameters using the exponential moving average method. In the t-th iteration step, the teacher network parameters are updated as follows:

[0011]

[0012] Where β is the smoothing hyperparameter, Φ t-1 Let be the parameters of the teacher grid at iteration step t-1. Let be the parameters of the student network at the t-th iteration step.

[0013] Furthermore, the point cloud within the grid of the labeled data is randomly deleted, including:

[0014] A data augmentation method that randomly deletes points within a grid is used to divide the ground truth box and randomly delete point clouds within a certain grid.

[0015] Furthermore, the truth boxes for the labeled data are divided using one of the following methods:

[0016] The first method of partitioning is to divide the truth box into K grids along the y-axis and z-axis, and then randomly select a point in one of the grids to delete.

[0017] The second method of partitioning is to divide the truth box into a center box and an outer box. The center box is the bounding box whose center point coincides with the truth box and whose size is 1 / N of the truth box. The outer box is the part excluding the center box. Delete the points inside the center box.

[0018] Furthermore, the consistency loss between the predictions generated by the auxiliary network and the voting results of the student network is calculated, including:

[0019] For the center point c of the truth box g g The predicted voting point in the auxiliary network and its Euclidean distance are: The predicted polling points in the student network and their Euclidean distance are The predicted voting results are determined by the Smooth L1 loss function L. vc Supervision:

[0020]

[0021] Where G is the number of truth boxes.

[0022] Furthermore, the predicted points of the student network and the teacher network are matched to calculate the consistency loss, including finding matching loops by using the nearest point matching method;

[0023] Calculate the distance between the predicted center points of the student network and the teacher network. Take the point with index i in the predicted point set of the student network as the current point and find the point j in the teacher network that is closest to point i.

[0024] With j as the current point, find the point closest to j in the student network prediction point set, and alternately find the point closest to the current point in the other party's point set;

[0025] The points found in the teacher network and the student network are respectively formed into point set T and point set S, until the next nearest point found already exists in T or S, forming a matching loop;

[0026] Consistency losses include matching loss, centroid consistency loss, category consistency loss, and size consistency loss.

[0027] Calculate the Euclidean distance between points in point set T and point set S respectively, and sum the two distances as the matching loss L. match ;

[0028] The central consistency loss is calculated from the L2 norm of the center points of the aligned teacher network and student network. Alignment refers to finding the point in the other network that is closest to the predicted point of this network.

[0029] The class consistency loss is calculated from the KL divergence of the aligned point pairs;

[0030] The size consistency loss is calculated from the MSE loss between alignment point pairs.

[0031] A semi-supervised 3D object detection system for indoor scenes includes: a data acquisition module, which acquires the input image and labels of the network, performs pre-training through a fully supervised network, and divides all point cloud data of the input image into labeled and unlabeled data; a first loss calculation module, which performs random downsampling, random flipping, and random rotation on all point cloud data sequentially, inputs it into a student network to calculate the prediction loss, and simultaneously performs only random downsampling on all point cloud data before inputting it into a teacher network, matches the prediction points of the student network and the teacher network, and calculates the consistency loss; a second loss calculation module, which randomly deletes points from the grid in the labeled data, performs random downsampling, inputs it into an added auxiliary network to generate prediction results, and calculates the voting consistency loss between the prediction results generated by the auxiliary network and the student network; and a detection module, which sums all the calculated losses as the overall loss of the semi-supervised network, inputs the target image to be detected, and uses the weights obtained during training to mark the target's bounding box and classify it.

[0032] A computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform any of the methods described above.

[0033] A computing device includes: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing any of the methods described above.

[0034] The present invention has the following advantages due to the adoption of the above technical solutions:

[0035] 1. This invention employs a data augmentation method that randomly deletes all points within a grid to simulate the sparsity of occluded objects and small objects in an indoor scene, thereby improving the detection performance for these objects.

[0036] 2. This invention employs an auxiliary network approach, which constrains the consistency of voting between the auxiliary network and the student network to extract more effective information from the labeled data.

[0037] 3. This invention employs a method of matching predicted points between student networks and teacher networks. By finding matching loops, it more rigorously constrains the consistency of the prediction results of the two networks.

[0038] 4. Applying this invention to a deep learning-based 3D target detection model for indoor scenes can greatly improve the model's detection performance.

[0039] In summary, this invention can be applied to fully supervised 3D object detection models based on deep learning, using only partially labeled data to improve the model's detection performance for occluded and small objects, thereby achieving results close to those of fully supervised object detection networks. Attached Figure Description

[0040] Figure 1 This is a schematic diagram of the network structure of the semi-supervised 3D target detection method in an embodiment of the present invention;

[0041] Figure 2 This is a schematic diagram of the SESS test results;

[0042] Figure 3 This is a schematic diagram of the detection results according to an embodiment of the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0044] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0045] To improve the accuracy of indoor scene detection, this invention provides a semi-supervised 3D object detection method, system, and storage medium for indoor scenes, primarily targeting a semi-supervised detection method for deep learning-based 3D object detection in indoor scenes. This invention includes a data augmentation method for randomly deleting points within a grid, a method for adding an auxiliary network, and a method for finding matching loop closures. The specific implementation steps are as follows: Inputting preprocessed images and labels into a network for pre-training; dividing the data into labeled and unlabeled data; randomly flipping, rotating, and downsampling all data, and inputting it into a student network to calculate the prediction loss; randomly downsampling all data and inputting it into a teacher network; randomly deleting points within the grid from the labeled data and inputting it into the added auxiliary network; calculating the voting consistency between the auxiliary network and the student network; matching the predicted points of the student network and the teacher network using a matching loop closure method, and calculating the consistency loss; summing all calculated losses as the final loss; inputting the target image to be detected, and using the trained weights, marking the target's bounding box and classification.

[0046] In one embodiment of the present invention, a semi-supervised 3D target detection method for indoor scenes is provided. In this embodiment, as shown... Figure 1 As shown, the method includes the following steps:

[0047] 1) Obtain the input image and labels of the network, complete the pre-training through a fully supervised network, and divide all point cloud data of the input image into labeled data and unlabeled data;

[0048] 2) After randomly downsampling, randomly flipping and randomly rotating all point cloud data in sequence, input them into the student network to calculate the prediction loss. At the same time, after randomly downsampling all point cloud data, input them into the teacher network, match the prediction points of the student network and the teacher network, and calculate the consistency loss.

[0049] 3) Randomly delete the point cloud within the grid of the labeled data, perform random downsampling, input it into the added auxiliary network to generate prediction results, and calculate the voting consistency loss between the prediction results generated by the auxiliary network and the student network.

[0050] 4) Sum all the calculated losses as the overall loss of the semi-supervised network. Input the target image to be detected, and use the weights obtained during training to mark the target's bounding box and classify it.

[0051] In step 1) above, a fully supervised network based on Hough voting is used to complete the pre-training, and the labeled data includes all categories.

[0052] In step 2) above, for each image, it is randomly flipped around the x-axis, and its binarized representation is as follows:

[0053]

[0054] α is a random variable uniformly drawn from [0,1].

[0055] For each image, it is randomly flipped around the y-axis, consistent with the flipping process around the x-axis described above;

[0056] For each image, rotate it around the z-axis by an angle ω, which can be expressed as follows:

[0057]

[0058] For each image, it is randomly downsampled to a fixed number of point clouds.

[0059] In this embodiment, the original data input to the teacher network is only randomly downsampled to a fixed number of point clouds.

[0060] In step 2) above, the student network is used to calculate the predicted loss. Specifically, the student network is a fully supervised 3D object detection network based on Hough voting.

[0061] The predicted loss includes voting loss L v Target confidence loss L o Classification loss L c and positioning loss L s The prediction loss for the labeled data is: L sup =L v +λ1L o +λ2L c +λ3L s , where λ1, λ2, and λ3 represent the weights of the loss function.

[0062] Among them, voting loss L v The formula is:

[0063]

[0064] Among them, M pos Δx represents the number of positive samples. i This represents the distance from the point with index i to its predicted center point. This represents the distance from the point with index i to the center point of its corresponding truth box;

[0065] The target confidence loss is calculated using Cross Entropy Loss, and the classification loss L... c Focal Loss, Localization Loss L s IoU series loss functions or Smooth L1 loss function can be used.

[0066] In step 2) above, the teacher network and student network have completely identical structures. The parameters of the teacher network are calculated from the student network parameters using the exponential moving average method. In the t-th iteration step, the parameters of the teacher network are updated as follows:

[0067]

[0068] Where β is a smoothing hyperparameter that controls how much information the teacher obtains from the student network; it is empirically set to 0.99; Φ t-1 Here are the parameters of the teacher grid at iteration step t-1; Let be the parameters of the student network at the t-th iteration step.

[0069] In step 2) above, the predicted points of the student network and the teacher network are matched, the consistency loss is calculated, and the nearest point matching method is used to find the matching loop; this includes the following steps:

[0070] 2.1) Calculate the distance between the predicted center points of the student network and the teacher network. Take the point with index i in the predicted point set of the student network as the current point, and find the point j in the teacher network that is closest to point i.

[0071] 2.2) Taking j as the current point, find the point closest to j in the student network prediction point set, and alternately find the point closest to the current point in the other party's point set;

[0072] 2.3) The points found in the teacher network and the student network are respectively formed into point set T and point set S, until the next nearest point found already exists in T or S, forming a matching closed loop.

[0073] In this embodiment, the consistency loss includes matching loss, center point consistency loss, category consistency loss, and size consistency loss, wherein:

[0074] Calculate the Euclidean distance between points in point set T and point set S respectively, and sum the two distances as the matching loss L. match :

[0075]

[0076] Among them, |C s | represents the number of points in the point set S. and Let C be the three-dimensional coordinates of the points with indices i and j in the point set S; t |, and The same principle applies as described above.

[0077] The center consistency loss is calculated from the L2 norm of the center points of the aligned teacher and student networks. Alignment refers to finding the point in the other network that is closest to the predicted point of the current network. The center consistency loss L... center for:

[0078]

[0079] Where, r s and r t These represent the center points predicted for the student network and the teacher network, respectively. and The points representing the teacher network and student network aligned using the above alignment method, |r s |and|r t | Represents the number of prediction points in the student network and the teacher network.

[0080] The class consistency loss is calculated from the KL (Kullback-Leibler) divergence of the aligned point pairs. class for:

[0081]

[0082] Where, p t The confidence level representing the teacher's network prediction category. |p t | Represents the number of teacher network prediction points.

[0083] The dimensional consistency loss is calculated from the MSE loss between alignment point pairs, where L is the dimensional consistency loss. size for:

[0084]

[0085] Where, d t The teacher network predicts the size of the bounding box. The size of the bounding box representing the alignment of the student network and the teacher network, |d t | represents the number of bounding boxes predicted by the teacher network.

[0086] Therefore, the consistency loss L con for:

[0087] L con =λ4L match +λ5L center +λ6L class +λ7L size

[0088] Where λ4, λ5, λ6, and λ7 are the weights of the loss function.

[0089] In step 3) above, the point cloud within the grid of the labeled data is randomly deleted. Specifically, a data augmentation method that randomly deletes points within the grid is used to divide the ground truth box and randomly delete the point cloud within a certain grid.

[0090] In this embodiment, the truth boxes of the labeled data are divided in one of the following ways:

[0091] The first method of partitioning is to divide the truth box into K grids along the y-axis and z-axis, and then randomly select a point in one of the grids to delete.

[0092] The second method of partitioning is to divide the truth box into a center box and an outer box. The center box is the bounding box whose center point coincides with the truth box and whose size is 1 / N of the truth box. The outer box is the part excluding the center box. Delete the points inside the center box.

[0093] For the two deletion methods mentioned above, one is randomly selected and applied to the labeled data, and the formula is expressed as follows:

[0094]

[0095] Let α be a random variable uniformly drawn from [0, 1], and D g D represents the random deletion of points within a grid after average division. c This indicates that the point inside the center box will be deleted.

[0096] In step 3) above, the structure of the auxiliary network is consistent with that of the student network, and its input is the labeled data after randomly deleting the grid. Based on the Hough voting prediction results of the auxiliary network and the student network for the center point, the consistency loss between the prediction results generated by the auxiliary network and the voting results of the student network is calculated, specifically as follows:

[0097] For the center point c of the truth box g g The predicted voting point in the auxiliary network and its Euclidean distance are: The predicted polling points in the student network and their Euclidean distance are The predicted voting results are determined by the Smooth L1 loss function L. vc Supervision:

[0098]

[0099] Where G is the number of truth boxes.

[0100] In step 4) above, all calculated losses are summed to obtain the overall loss of the semi-supervised network. The overall loss function of the semi-supervised network consists of prediction loss and consistency loss.

[0101] Ltotal =L sup +μ1L con +μ2L vc

[0102] Here, μ1 is a weight parameter, whose value increases from 0 to 10 with the number of iterations, specifically determined by the sigmoid-shaped method. The network weights are calculated and trained using a loss function. During the testing phase, the weights obtained from training are used to identify and locate targets in the test images. μ2 is the weight parameter.

[0103] This embodiment is based on the semi-supervised 3D object detection model SESS using Hough voting, combined with the data augmentation method of randomly deleting points within the grid, the method of adding auxiliary networks, and the method of finding matching loop closures proposed in this invention. This embodiment uses the large-scale 3D indoor dataset ScanNet V2 (Dai, A.; Chang, AX; Savva, M.; Halber, M.; Funkhouser, T.; and Nieβner, M. 2017. Scannet: Richly-annotated 3d reconstructions of indoor scenes. In Proceedings of the IEEE conference on computer vision and pattern recognition), which contains 18 types of targets and includes various occluded objects and small objects. The specific implementation steps are as follows:

[0104] 1) In this embodiment, the number of point clouds in the input image is inconsistent. First, it is randomly sampled to 50,000 points and then input into the fully supervised network VoteNet (Qi, CR; Litany, O.; He, K.; and Guibas, LJ2019. Deepphough voting for 3d object detection in point clouds. In proceedings of the IEEE / CVF International Conference on Computer Vision) for pre-training.

[0105] 2) Divide all point cloud data into labeled data and unlabeled data at a ratio of 1:9.

[0106] 3) Randomly flip and rotate the data in 2), and randomly downsample it to 20,000 points.

[0107] 4) Randomly downsample the data in 2) to 20,000 points.

[0108] 5) Perform data augmentation on the labeled data from 2) by randomly deleting points within the grid. Specifically, for the truth boxes in the labeled data, there are two ways to divide them: one is to divide them into four grids along the y and z axes and randomly select one grid to delete; the other is to divide them into a central grid and an outer grid. The central grid is the grid that coincides with the center point of the truth box and has a size of 1 / 2 of the truth box. The outer grid is the rest of the grid excluding the central grid, and points within the central grid are deleted. The network randomly selects one of these two methods for data augmentation.

[0109] 6) Input the data from 3) into the student network VoteNet and calculate the prediction loss, where λ1, λ2, and λ_3 are set to 0.5, 1.0, and 0.1, respectively.

[0110] 7) Input the data from 4) into the teacher network VoteNet, whose training parameters are calculated by the student network using the exponential moving average method.

[0111] 8) Input the data from 5) into the auxiliary network VoteNet to generate voting prediction results.

[0112] 9) Perform voting consistency supervision on the voting prediction results generated in 6) and 8) to improve the prediction performance of the voting center point by constraining the distance between the two voting results.

[0113] 10) Match the prediction results generated in 6) and 7), i.e., alternately search for the point closest to the current point of this network from the prediction point set of another network until a matching loop is formed. Further constrain the Euclidean distance between points within the matching loop, thereby constraining the consistency of the prediction results of the teacher network and the student network in a more rigorous way. Further, align the two prediction results and calculate the center point consistency loss, category consistency loss, and size consistency loss, with λ4, λ5, λ6, and λ7 set to 0.1, 1, 2, and 1, respectively.

[0114] 11) Based on the losses in 6) and 10), calculate the overall loss of the network. μ1 is increased from 0 to 10 in the first 30 epochs of the training phase using a sigmoid-shaped method, and μ1 is set to 0.1.

[0115] In this embodiment, the hardware configuration for executing the algorithm is: CPU is an Intel i9, GPU is a GeForce 3090 with 12GB of memory; the software configuration is: computer operating system is Ubuntu 16.04, CUDA version is 11.0, and the neural network framework used is PyTorch version 0.8. The parameters for the pre-training phase are set as follows: initial learning rate, learning rate decay steps are set to 0.001, {80, 120}, {0.1, 0.1} respectively, with 180 training iterations. The parameters for the training phase are set as follows: initial learning rate, learning rate decay steps are set to 0.001, {100, 140, 180}, {0.1, 0.1, 0.1} respectively, with 220 training iterations. Other embodiments can adjust the parameters appropriately according to the selected object detection method and dataset. After training is complete, the network weights are obtained. In the testing phase, the input image is used to classify and locate the target using the weights.

[0116] In summary, this invention enables the detection of 3D targets in indoor scenes.

[0117] To verify the effectiveness and practicality of the method proposed in this invention, an example on the SESS dataset is given below. Table 1 shows the detection results of the example on the test set. The various metrics are AP (Average Precision) and mAP (mean Average Precision), which is the average AP value of all categories.

[0118] Table 1 shows the validation results of the examples on the dataset.

[0119]

[0120] As shown in Table 1, the detection performance of the SESS model is greatly improved after adding the data augmentation method of randomly deleting points within the grid, the auxiliary network method, and the matching loop finding method proposed in this invention. In particular, the detection accuracy for occluded objects (bookshelf) and small objects (picture) is greatly increased, proving the effectiveness of this invention.

[0121] For example Figure 2 and Figure 3 The visualization results shown demonstrate that, in a qualitative comparison, small objects missed by SESS were accurately detected after incorporating the method proposed in this invention. The method proposed in this invention can also be flexibly applied to other semi-supervised 3D object detection network frameworks.

[0122] In one embodiment of the present invention, a semi-supervised 3D target detection system for indoor scenes is provided, comprising:

[0123] The data acquisition module acquires the input images and labels of the network, completes pre-training through a fully supervised network, and divides all point cloud data of the input images into labeled data and unlabeled data.

[0124] The first loss calculation module performs random downsampling, random flipping, and random rotation on all point cloud data in sequence, and then inputs it into the student network to calculate the prediction loss. At the same time, it performs only random downsampling on all point cloud data and inputs it into the teacher network, matching the prediction points of the student network and the teacher network to calculate the consistency loss.

[0125] The second loss calculation module randomly deletes point clouds within the grid from the labeled data, performs random downsampling, inputs them into the added auxiliary network to generate prediction results, and calculates the consistency loss between the prediction results generated by the auxiliary network and the voting results of the student network.

[0126] The detection module sums all the calculated losses to form the overall loss of the semi-supervised network. It takes the target image to be detected as input, and uses the weights obtained during training to mark the target's bounding box and classify it.

[0127] In the above embodiments, the input student network calculates the predicted loss, including:

[0128] The student network is a fully supervised 3D object detection network based on Hough voting;

[0129] The predicted loss includes voting loss L v Target confidence loss L o Classification loss L c and positioning loss L s The prediction loss for the labeled data is: L sup =L v +λ1L o +λ2L c +λ3L s , where λ1, λ2, and λ3 represent the weights of the loss function.

[0130] In the above embodiments, the teacher network and student network have completely identical structures. The parameters of the teacher network are calculated from the student network parameters using the exponential moving average method. In the t-th iteration step, the parameters of the teacher network are updated as follows:

[0131]

[0132] Where β is the smoothing hyperparameter, Φ t-1 Let be the parameters of the teacher grid at iteration step t-1. Let be the parameters of the student network at the t-th iteration step.

[0133] In the above embodiments, randomly deleting point clouds within the grid from the labeled data includes:

[0134] A data augmentation method that randomly deletes points within a grid is used to divide the ground truth box and randomly delete point clouds within a certain grid.

[0135] In the above embodiments, the truth boxes of the labeled data are divided in one of the following ways:

[0136] The first method of partitioning is to divide the truth box into K grids along the y-axis and z-axis, and then randomly select a point in one of the grids to delete.

[0137] The second method of partitioning is to divide the truth box into a center box and an outer box. The center box is the bounding box whose center point coincides with the truth box and whose size is 1 / N of the truth box. The outer box is the part excluding the center box. Delete the points inside the center box.

[0138] In the above embodiments, calculating the consistency loss between the prediction results generated by the auxiliary network and the voting results of the student network includes:

[0139] For the center point c of the truth box g g The predicted voting point in the auxiliary network and its Euclidean distance are: The predicted polling points in the student network and their Euclidean distance are The predicted voting results are determined by the Smooth L1 loss function L. vc Supervision:

[0140]

[0141] Where G is the number of truth boxes.

[0142] In the above embodiments, the predicted points of the student network and the teacher network are matched and the consistency loss is calculated, including: finding matching loops by using the nearest point matching method;

[0143] Calculate the distance between the predicted center points of the student network and the teacher network. Take the point with index i in the predicted point set of the student network as the current point and find the point j in the teacher network that is closest to point i.

[0144] With j as the current point, find the point closest to j in the student network prediction point set, and alternately find the point closest to the current point in the other party's point set;

[0145] The points found in the teacher network and the student network are respectively formed into point set T and point set S, until the next nearest point found already exists in T or S, forming a matching loop;

[0146] Consistency losses include matching loss, centroid consistency loss, category consistency loss, and size consistency loss.

[0147] Calculate the Euclidean distance between points in point set T and point set S respectively, and sum the two distances as the matching loss L. match ;

[0148] The central consistency loss is calculated from the L2 norm of the center points of the aligned teacher network and student network. Alignment refers to finding the point in the other network that is closest to the predicted point of this network.

[0149] The class consistency loss is calculated from the KL divergence of the aligned point pairs;

[0150] The size consistency loss is calculated from the MSE loss between alignment point pairs.

[0151] The system provided in this embodiment is used to execute the above-described method embodiments. For specific processes and details, please refer to the above embodiments, which will not be repeated here.

[0152] In one embodiment of the present invention, a computing device is provided, which can be a terminal and may include: a processor, a communication interface, memory, a display screen, and an input device. The processor, communication interface, and memory communicate with each other via a communication bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs, which, when executed by the processor, implement the methods described in the above embodiments. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, a management network, NFC (Near Field Communication), or other technologies. The display screen can be a liquid crystal display or an e-ink display. The input device can be a touch layer covering the display screen, or buttons, a trackball, or a touchpad mounted on the casing of the computing device, or an external keyboard, touchpad, or mouse. The processor can call logical instructions stored in the memory.

[0153] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0154] In one embodiment of the present invention, a computer program product is provided, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to perform the methods provided in the above-described method embodiments.

[0155] In one embodiment of the present invention, a non-transitory computer-readable storage medium is provided, which stores server instructions that cause a computer to perform the methods provided in the above embodiments.

[0156] The computer-readable storage medium provided in the above embodiments has a similar implementation principle and technical effect to the above method embodiments, and will not be described again here.

[0157] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0158] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0159] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A semi-supervised 3D target detection method for indoor scenes, characterized in that, include: The input image and labels of the network are obtained, and pre-training is completed through a fully supervised network. All point cloud data of the input image are divided into labeled data and unlabeled data. After randomly downsampling, randomly flipping, and randomly rotating all point cloud data in sequence, the data is input into the student network to calculate the prediction loss. At the same time, after randomly downsampling all point cloud data, the data is input into the teacher network. The prediction points of the student network and the teacher network are matched to calculate the consistency loss. The point cloud within the grid of the labeled data is randomly deleted, and after random downsampling, it is input into the added auxiliary network to generate prediction results. The consistency loss between the prediction results generated by the auxiliary network and the voting results of the student network is calculated. The sum of all calculated losses is used as the overall loss of the semi-supervised network. The target image to be detected is input, and the target's bounding box and classification are marked using the weights obtained during training.

2. The semi-supervised 3D target detection method in indoor scenes as described in claim 1, characterized in that, Input the student network to calculate the predicted loss, including: The student network is a fully supervised 3D object detection network based on Hough voting; The predicted loss includes voting loss L v Target confidence loss L o Classification loss L c and positioning loss L s The prediction loss for the labeled data is: L sup =L v +λ1L o +λ2L c +λ3L s , where λ1, λ2, λ3 represent the weights of the loss function.

3. The semi-supervised 3D target detection method in indoor scenes as described in claim 1, characterized in that, The teacher network and student network have completely identical structures. The teacher network's parameters are calculated from the student network's parameters using the exponential moving average method. In the t-th iteration step, the teacher network's parameters are updated as follows: Where β is the smoothing hyperparameter, Φ t-1 Let be the parameters of the teacher grid at iteration step t-1. Let be the parameters of the student network at the t-th iteration step.

4. The semi-supervised 3D target detection method in indoor scenes as described in claim 1, characterized in that, Randomly delete point clouds within the grid from the labeled data, including: A data augmentation method that randomly deletes points within a grid is used to divide the ground truth box and randomly delete point clouds within a certain grid.

5. The semi-supervised 3D target detection method in indoor scenes as described in claim 4, characterized in that, The truth boxes for labeled data are divided using one of the following methods: The first method of partitioning is to divide the truth box into K grids along the y-axis and z-axis, and then randomly select a point in one of the grids to delete. The second method of partitioning is to divide the truth box into a center box and an outer box. The center box is the bounding box whose center point coincides with the truth box and whose size is 1 / N of the truth box. The outer box is the part excluding the center box. Delete the points inside the center box.

6. The semi-supervised 3D target detection method in indoor scenes as described in claim 1, characterized in that, The calculation of the consistency loss between the predictions generated by the auxiliary network and the votes of the student network includes: For the center point c of the truth box g g The predicted voting point in the auxiliary network and its Euclidean distance are: The predicted polling points in the student network and their Euclidean distance are The predicted voting results are determined by the Smooth L1 loss function L. vc Supervision: Where G is the number of truth boxes.

7. The semi-supervised 3D target detection method in indoor scenes as described in claim 1, characterized in that, Match the predicted points of the student network and the teacher network, and calculate the consistency loss, including finding matching loops by using the nearest point matching method; Calculate the distance between the predicted center points of the student network and the teacher network. Take the point with index i in the predicted point set of the student network as the current point and find the point j in the teacher network that is closest to point i. With j as the current point, find the point closest to j in the student network prediction point set, and alternately find the point closest to the current point in the other party's point set; The points found in the teacher network and the student network are respectively formed into point set T and point set S, until the next nearest point found already exists in T or S, forming a matching loop; Consistency losses include matching loss, centroid consistency loss, category consistency loss, and size consistency loss. Calculate the Euclidean distance between points in point set T and point set S respectively, and sum the two distances as the matching loss L. match ; The central consistency loss is calculated from the L2 norm of the center points of the aligned teacher network and student network. Alignment refers to finding the point in the other network that is closest to the predicted point of this network. The class consistency loss is calculated from the KL divergence of the aligned point pairs; The size consistency loss is calculated from the MSE loss between alignment point pairs.

8. A semi-supervised 3D target detection system for indoor scenes, characterized in that, include: The data acquisition module acquires the input images and labels of the network, completes pre-training through a fully supervised network, and divides all point cloud data of the input images into labeled data and unlabeled data. The first loss calculation module performs random downsampling, random flipping, and random rotation on all point cloud data in sequence, and then inputs it into the student network to calculate the prediction loss. At the same time, it performs only random downsampling on all point cloud data and inputs it into the teacher network, matching the prediction points of the student network and the teacher network to calculate the consistency loss. The second loss calculation module randomly deletes point clouds within the grid from the labeled data, performs random downsampling, inputs them into the added auxiliary network to generate prediction results, and calculates the consistency loss between the prediction results generated by the auxiliary network and the voting results of the student network. The detection module sums all the calculated losses to form the overall loss of the semi-supervised network. It takes the target image to be detected as input, and uses the weights obtained during training to mark the target's bounding box and classify it.

9. A computer-readable storage medium for storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to perform any of the methods described in claims 1 to 7.

10. A computing device, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described in claims 1 to 7.