Semi-supervised three-dimensional object detection method and device based on cross-scale distillation and medium

CN118135555BActive Publication Date: 2026-09-22FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410350563.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-26
Publication Date
2026-09-22
Estimated Expiration
2044-03-26

AI Technical Summary

Benefits of technology

[0042]一、本发明提出了面向半监督三维目标检测的跨尺度蒸馏模块,利于学生模型学习单一尺度下无法感知到的物体信息并将不同尺度的知识进行互补增强,增强学生模型对于物体尺度方差的鲁棒性与提高其检测的尺度一致性能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118135555B_ABST
    Figure CN118135555B_ABST
Patent Text Reader

Abstract

The application relates to a semi-supervised three-dimensional target detection method, device and medium based on cross-scale distillation, wherein the method comprises the following steps: constructing an average teacher semi-supervised learning framework of a multi-scale point cloud input; performing scale alignment on multi-scale outputs of a teacher model and a student model and inputting the multi-scale outputs into a cross-scale distillation module; performing voxel anchor point matching in the cross-scale distillation module, and performing double consistency learning between matched voxel anchor point pairs; constructing a labeled data loss function and an unlabeled data loss function, and training the student model; inputting a point cloud scene to be detected into the trained three-dimensional target detection student model, and obtaining a point cloud scene with predicted object bounding boxes. Compared with the prior art, the application has higher utilization of unlabeled training point cloud data, higher overall detection precision and better detection capability for small-scale objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of 3D point cloud understanding technology, and in particular to a semi-supervised 3D target detection method, apparatus and medium based on cross-scale distillation. Background Technology

[0002] 3D object detection is the task of automatically identifying the position, shape, and size of objects in a 3D scene, and it has been widely used in fields such as autonomous driving and robotics. Due to the high cost of 3D scene annotation, the scarcity of large-scale, high-quality 3D point cloud annotation data has become a key challenge for current deep learning-based 3D object detection methods. Achieving the same performance as fully supervised 3D object detection with less annotation data has become a major breakthrough challenge for existing technologies. Therefore, semi-supervised 3D object detection has become a highly anticipated issue in 3D object detection.

[0003] The paper "SESS: Self-Ensembling Semi-Supervised 3D Object Detection" presents pioneering work in the field of semi-supervised 3D object detection, designing asymmetric data augmentation suitable for 3D scenes and three consistency losses between teacher and student model predictions. The paper "3DIoUMatch: Leveraging IoUPrediction for Semi-Supervised 3D Object Detection" designs an IoU-aware VoteNet as the underlying architecture for both student and teacher models, filtering low-quality predictions from pseudo-labels by setting different types of confidence thresholds. The paper "DQS3D: Densely-matched Quantization-aware Semi-supervised 3D Detection" proposes a quantization-aware semi-supervised 3D object detection method based on dense matching, reducing quantization errors in the process of converting point clouds into voxels.

[0004] The aforementioned works share a common drawback: they fail to consider the learning imbalance caused by the variance of the three-dimensional object scale, and the use of teacher models to supervise student models at only the prediction level makes them susceptible to noise interference. Summary of the Invention

[0005] The purpose of this invention is to provide a semi-supervised three-dimensional target detection method, device, and medium based on cross-scale distillation that has high efficiency and performance in utilizing unlabeled training data.

[0006] The objective of this invention can be achieved through the following technical solutions:

[0007] A semi-supervised 3D target detection method based on cross-scale distillation includes the following steps:

[0008] Step 1: Construct an average teacher semi-supervised learning framework for multi-scale point cloud input. The average teacher semi-supervised learning framework includes a labeled branch and an unlabeled branch, which correspond to the learning of labeled point cloud data and unlabeled point cloud data, respectively. The labeled branch uses the ground truth labels of the labeled point cloud data to supervise the student model, and the unlabeled branch uses the teacher model to supervise the student model.

[0009] Step 2: Scale-align the multi-scale outputs of the teacher model and student model and input them into the cross-scale distillation module;

[0010] Step 3: Perform voxel anchor matching in the cross-scale distillation module and perform double consistency learning between the matched voxel anchor pairs;

[0011] Step 4: Construct labeled data loss functions and unlabeled data loss functions, and train the student model;

[0012] Step 5: Input the point cloud scene to be detected into the trained 3D object detection student model to obtain the point cloud scene with predicted object bounding boxes.

[0013] In step 1, the labelless 3D point cloud scene x u After asymmetric data augmentation, the point cloud scene is transformed into a strongly augmented data scene. And weak data augmentation point cloud scenarios The point cloud scene is used as the original input for the student model and the teacher model, respectively. The result after scaling As an additional input to the student model, rescale(·) represents the scaling transformation function.

[0014] In step 2, the cross-scale distillation module is divided into a cross-scale internal distillation module and a cross-scale external distillation module, depending on the distillation target.

[0015] For different scale inputs to the same unlabeled point cloud scene, the outputs of the teacher model and the student model contain voxel anchors of the corresponding scale, as well as the features and predictions corresponding to each voxel anchor.

[0016] Based on the scale transformation parameters, the voxel anchors of the two sets of student models at different scales are transformed to the same scale and then input into the cross-scale internal distillation module.

[0017] Based on the initial strong and weak data augmentation parameters, the output of the teacher model is transformed to the same view as the output of the student model, and then transformed to the same scale according to the scale transformation parameters. The teacher model output and the scale-transformed output of the student model are then input into the cross-scale external distillation module.

[0018] In step 3, voxel anchor matching specifically involves: in the source voxel anchor set V s and the target voxel anchor set V t One-to-one dense matching is performed between them, and the K-nearest neighbor algorithm is used to anchor each source voxel. s Match the nearest target voxel anchor point v t This yields one-to-one matching voxel anchor pairs, where the Euclidean distance metric is used to calculate the distance, and the specific formula is as follows:

[0019]

[0020] Where φ(·) represents the coordinates of the voxel anchor point.

[0021] In step 3, after obtaining n one-to-one matched voxel anchor pairs Then, using the target voxel anchor point Corresponding predicted center score and category scores To eliminate low-confidence matching voxel anchor pairs When the center score Greater than the center score threshold τ cen And category scores Greater than the classification score threshold τ cls If the matching voxel anchor pair is valid, it is retained; otherwise, it is discarded.

[0022] In step 3, the matching voxel anchor points are... Double consistency learning is performed between them, which includes prediction consistency learning and feature consistency learning.

[0023] In step 3, for the matched voxel anchor pairs (v s ,v t ), and the source voxel anchor point v s The corresponding predictions and features are represented as follows: and f s , and These represent the parameters of the 3D bounding box, center score, and classification score corresponding to the source voxel anchor point in the prediction, respectively, and are related to the target voxel anchor point v. t The corresponding predictions and features are represented as follows: and f t , and These represent the predicted 3D bounding box parameters, center score, and classification score corresponding to the matched target voxel anchor points, respectively.

[0024] A prediction consistency learning loss function and a feature consistency learning loss function are constructed separately, and then weighted and summed to obtain a double consistency learning loss function, which is used to train the model.

[0025] Predictive Consistency Learning Loss Function The formula is as follows:

[0026]

[0027] Where L box L cen and L cls Let these represent Huber loss, L2 loss, and KL divergence, respectively.

[0028] Feature consistency learning loss function The formula is as follows:

[0029]

[0030] Double consistency learning loss function L dual The formula is as follows:

[0031]

[0032] Where μ represents the prediction consistency learning loss. Feature consistency learning loss The balance coefficient between them.

[0033] In step 4, there is a label data loss function L. l The loss function L for unlabeled data is calculated based on the ground truth labels of the labeled training data. u Losses from transscale internal distillation and cross-scale external distillation losses Adding them together, we get:

[0034]

[0035] Among them, the cross-scale internal distillation loss and the cross-scale external distillation loss are the dual consistency learning losses of the cross-scale internal distillation module and the cross-scale external distillation module, respectively.

[0036] Based on the loss functions for labeled and unlabeled data, the overall loss function of the average teacher semi-supervised learning framework is determined as follows:

[0037] L = L l +λL u

[0038] Where λ is the labeled loss L l And unlabeled loss L uThe weighting coefficients between them.

[0039] A semi-supervised three-dimensional target detection device based on cross-scale distillation includes a memory, a processor, and a program stored in the memory, wherein the processor executes the program to implement the method described above.

[0040] A storage medium having a program stored thereon, which, when executed, implements the method described above.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] I. This invention proposes a cross-scale distillation module for semi-supervised 3D object detection, which helps the student model learn object information that cannot be perceived at a single scale and complements and enhances knowledge at different scales, thereby improving the robustness of the student model to object scale variance and improving its ability to detect scale consistency.

[0043] Second, this invention proposes a dual consistency learning paradigm, which integrates the consistency learning process of the feature level and the prediction level, providing more complex and robust supervision signals for the training of student models, and alleviating the problem that student models are easily misled by noise when using a single prediction level supervision signal.

[0044] Third, performing feature consistency learning under cross-scale conditions helps student models learn multi-scale target features that cannot be perceived under a single scale. Performing prediction consistency learning under cross-scale conditions helps student models ensure consistent predictions for inputs transformed to different scales, enhancing the detection consistency of student models for objects at different scales, especially improving the detection ability of student models for small-scale objects. Attached Figure Description

[0045] Figure 1 This is a flowchart of the method of the present invention;

[0046] Figure 2 This is a schematic diagram of the average teacher-supervised learning framework of the present invention;

[0047] Figure 3 This is a schematic diagram of the voxel anchor point matching process of the present invention;

[0048] Figure 4 This is a schematic diagram of the dual consistency learning method of the present invention. Detailed Implementation

[0049] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0050] This embodiment provides a semi-supervised 3D target detection method based on cross-scale distillation, such as... Figure 1 As shown, it includes the following steps:

[0051] Step 1: Construct an average teacher semi-supervised learning framework for multi-scale point cloud inputs.

[0052] like Figure 2 As shown, the average teacher-supervised learning framework includes labeled and unlabeled branches, corresponding to learning from labeled and unlabeled point cloud data, respectively. The labeled branch uses the ground truth labels of the labeled point cloud data to supervise the student model, while the unlabeled branch uses the teacher model to supervise the student model. The unlabeled point cloud data is then subjected to weak and strong data augmentation, respectively, and used as the basic input to both the teacher and student models. The strongly augmented point cloud data is then copied and scaled, and used as additional input to the student model.

[0053] Specifically, label-free 3D point cloud scene x u After asymmetric data augmentation, the point cloud scene is transformed into a strongly augmented data scene. And weak data augmentation point cloud scenarios The point cloud scene is used as the original input for the student model and the teacher model, respectively. The result after scaling As an additional input to the student model, `rescale(·)` represents the scaling function. Weak data augmentation methods include translation and rotation operations, while strong data augmentation methods include flipping, rotation, translation, and scaling operations. Both the student and teacher models use the same fully convolutional 3D object detection network, FCAF3D, as their basic structure.

[0054] Step 2: Scale-align the multi-scale outputs of the teacher model and student model and input them into the cross-scale distillation module.

[0055] For inputs of the same unlabeled point cloud scene at different scales, the outputs of the teacher and student models include voxel anchors at the corresponding scales, as well as the features and predictions corresponding to each voxel anchor. For example... Figure 2 As shown, depending on the distillation objective, the multiscale distillation module is divided into a multiscale internal distillation module and a multiscale external distillation module.

[0056] For inputs of the same unlabeled point cloud scene at different scales, the outputs of the teacher model and the student model contain voxel anchors at the corresponding scales, as well as the features and predictions corresponding to each voxel anchor.

[0057] The distillation objective of the cross-scale internal distillation module is the student's multi-scale output. Based on the scale transformation parameters, the voxel anchor points of the student model at two different scales are transformed to the same scale and then input into the cross-scale internal distillation module.

[0058] The distillation objective of the cross-scale external distillation module is the teacher's single-scale output and the student's multi-scale output. It corresponds to two distillation processes. Based on the initial strong and weak data augmentation parameters, the output of the teacher model is transformed to the same view as the output of the student model, and then transformed to the same scale according to the scale transformation parameters. The scale-transformed output of the teacher model and the student model are then input into the cross-scale external distillation module together.

[0059] Step 3: Perform voxel anchor matching in the cross-scale distillation module and perform double consistency learning between the matched voxel anchor pairs.

[0060] Both the cross-scale internal distillation module and the cross-scale external distillation module contain voxel anchor matching and dual consistency learning processes.

[0061] like Figure 3 As shown, in the source voxel anchor set V s and the target voxel anchor set V t One-to-one dense matching is performed between them, and the K-nearest neighbor algorithm is used to anchor each source voxel. s Match the nearest target voxel anchor point v t This yields one-to-one matching voxel anchor pairs, where the Euclidean distance metric is used to calculate the distance, and the specific formula is as follows:

[0062]

[0063] Where φ(·) represents the coordinates of the voxel anchor points. In the cross-scale intradistillation module, the source voxel anchor point set V s The output of the student model for a scale-transformed point cloud scene input is the target voxel anchor point set V. t This is the output of the student model for a strongly data-augmented point cloud scene input. In the cross-scale external distillation module, the source voxel anchor set V... s The output of the student model is the set of target voxel anchor points V, which is the input of two strongly augmented point cloud scenes at two different scales. t It is the output of the teacher model for weak data augmentation point cloud scenarios.

[0064] Because the predictions corresponding to the target voxel anchors contain low-quality samples, therefore, after obtaining n one-to-one matched voxel anchor pairs... Subsequently, it is necessary to apply a confidence-based filtering mechanism to eliminate low-quality matching voxel anchor pairs. Specifically, this involves utilizing the target voxel anchor... Corresponding predicted center score and category scores To eliminate low-confidence matching voxel anchor pairs When the center score Greater than the center score threshold τcen And category scores Greater than the classification score threshold τ cls If the matching voxel anchor pair is valid, it is retained; otherwise, it is discarded.

[0065] like Figure 4 As shown, in the matching voxel anchor point pair It performs dual consistency learning, namely prediction consistency learning and feature consistency learning.

[0066] For the matched voxel anchor pairs (v s ,v t ), and the source voxel anchor point v s The corresponding predictions and features are represented as follows: and f s , and These represent the parameters of the 3D bounding box, center score, and classification score corresponding to the source voxel anchor point in the prediction, respectively, and are related to the target voxel anchor point v. t The corresponding predictions and features are represented as follows: and f t , and These represent the predicted 3D bounding box parameters, center score, and classification score corresponding to the matched target voxel anchor points, respectively.

[0067] We construct prediction consistency learning loss function and feature consistency learning loss function respectively, and then perform a weighted summation of the two to obtain a double consistency learning loss function, which is used to train the model.

[0068] Predictive Consistency Learning Loss Function The formula is as follows:

[0069]

[0070] Where L box L cen and L cls Let Huber loss, L2 loss, and KL divergence be represented by the following formulas:

[0071]

[0072]

[0073]

[0074] Where δ is taken as 0.3.

[0075] Feature consistency learning loss function The formula is as follows:

[0076]

[0077] Then, the double consistency learning loss function L dual The formula is as follows:

[0078]

[0079] Where μ represents the prediction consistency learning loss. Feature consistency learning loss The balance coefficient between them is set to 1 in this embodiment.

[0080] Step 4: Construct labeled data loss functions and unlabeled data loss functions, and train the student model.

[0081] The overall loss of a semi-supervised learning framework consists of two distinct parts: labeled loss L. l And unlabeled loss L u These correspond to labeled and unlabeled branches, respectively. For the labeled loss L... l Because the truth labels are directly used to supervise the student model in the labeled branch, there is a labeled loss L. l The calculation method is consistent with that of the student model under fully supervised conditions. The unlabeled data loss function L... u Losses from transscale internal distillation and cross-scale external distillation losses Adding them together, we get:

[0082]

[0083] Among them, the cross-scale internal distillation loss and the cross-scale external distillation loss are the dual consistency learning losses of the cross-scale internal distillation module and the cross-scale external distillation module, respectively.

[0084] Therefore, the overall loss function of the average teacher-supervised learning framework is as follows:

[0085] L = L l +λL u

[0086] Where λ is the labeled loss L l And unlabeled loss L u The weighting coefficient between them is set to 1 in this embodiment.

[0087] The parameters of the student 3D object detection model are optimized using the backpropagation algorithm to complete the training of the student 3D object detection model. During the semi-supervised learning training process, the weights of the teacher model are updated using an exponential moving average of the student model weights.

[0088] In this embodiment, the number of training iterations is set to 6000 and 12000 for the ScanNet V2 and SUN RGB-D benchmark datasets, respectively. During the training phase, the student model is trained using the Adam optimizer and initialized with a learning rate of 0.001 and a weight decay of 0.0001.

[0089] Under semi-supervised conditions, each batch of training data consists of labeled and unlabeled training data. In this section, the batch size for both labeled and unlabeled data is set to 4, meaning each training batch contains 4 labeled 3D scene data points and 4 unlabeled 3D scene data points. Similar to previous semi-supervised 3D object detection work, this section sets the proportion of labeled training data to 5%, 10%, 20%, and 100%, respectively. When the proportion of labeled training data is set to 100%, all labeled training data will be copied and the labels discarded to serve as unlabeled training data.

[0090] Step 5: Input the point cloud scene to be detected into the trained 3D object detection student model to obtain the point cloud scene with predicted object bounding boxes.

[0091] This embodiment uses training data from the publicly available ScanNet V2 and SUN RGB-D benchmark datasets, and test data for testing. The test results are as follows: On the ScanNet V2 dataset, when the proportion of labeled training data is set to 10%, this invention outperforms the state-of-the-art DQS3D algorithm by 3.8% and 3.5% in mAP@0.25 and mAP@0.5, respectively. On the SUN RGB-D dataset, when the proportion of labeled training data is set to 5%, this invention outperforms the state-of-the-art DQS3D algorithm by 3.4% and 4.6% in mAP@0.25 and mAP@0.5, respectively; when the proportion of labeled training data is set to 10%, this invention outperforms the state-of-the-art DQS3D algorithm by 2.8% and 4.1% in mAP@0.25 and mAP@0.5, respectively. Therefore, this invention improves the accuracy of object detection.

[0092] The above is an introduction to the method embodiments. The following describes the solution of the present invention further through device embodiments.

[0093] This embodiment provides a semi-supervised three-dimensional target detection device based on cross-scale distillation, including a memory, a processor, and a program stored in the memory. When the processor executes the program, it implements the method described above.

[0094] Specifically, the device includes:

[0095] A semi-supervised learning framework module is used to construct an average teacher semi-supervised learning framework for multi-scale point cloud input. The average teacher semi-supervised learning framework includes a labeled branch and an unlabeled branch, which correspond to the learning of labeled point cloud data and unlabeled point cloud data, respectively. The labeled branch uses the ground truth labels of the labeled point cloud data to supervise the student model, and the unlabeled branch uses the teacher model to supervise the student model.

[0096] The scale alignment module is used to scale the multi-scale outputs of the teacher model and the student model and input them into the cross-scale distillation module.

[0097] The cross-scale distillation module is used for voxel anchor matching, performing double consistency learning between matched voxel anchor pairs;

[0098] The training module is used to construct labeled and unlabeled data loss functions to train the student model;

[0099] The detection module is used to input the point cloud scene to be detected into the trained 3D object detection student model to obtain the point cloud scene with the predicted object bounding boxes.

[0100] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the described module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0101] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0102] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.

Claims

1. A semi-supervised three-dimensional target detection method based on cross-scale distillation, characterized in that, Includes the following steps: Step 1: Construct an average teacher semi-supervised learning framework for multi-scale point cloud input. The average teacher semi-supervised learning framework includes a labeled branch and an unlabeled branch, which correspond to the learning of labeled point cloud data and unlabeled point cloud data, respectively. The labeled branch uses the ground truth labels of the labeled point cloud data to supervise the student model, and the unlabeled branch uses the teacher model to supervise the student model. Step 2: Scale-align the multi-scale outputs of the teacher model and student model and input them into the cross-scale distillation module; Step 3: Perform voxel anchor matching in the cross-scale distillation module and perform double consistency learning between the matched voxel anchor pairs; Step 4: Construct labeled data loss functions and unlabeled data loss functions, and train the student model; Step 5: Input the point cloud scene to be detected into the trained 3D object detection student model to obtain the point cloud scene with predicted object bounding boxes; Depending on the distillation objective, the multiscale distillation module is divided into the multiscale internal distillation module and the multiscale external distillation module. For different scale inputs to the same unlabeled point cloud scene, the outputs of the teacher model and the student model contain voxel anchors of the corresponding scale, as well as the features and predictions corresponding to each voxel anchor. The distillation objective of the cross-scale internal distillation module is the student's multi-scale output. Based on the scale transformation parameters, the voxel anchor points of the student model at two different scales are transformed to the same scale and then input into the cross-scale internal distillation module. The distillation objective of the cross-scale external distillation module is the teacher's single-scale output and the student's multi-scale output. It corresponds to two distillation processes. Based on the initial strong and weak data augmentation parameters, the output of the teacher model is transformed to the same view as the output of the student model, and then transformed to the same scale according to the scale transformation parameters. The scale-transformed output of the teacher model and the student model are then input into the cross-scale external distillation module together.

2. The semi-supervised three-dimensional target detection method based on cross-scale distillation according to claim 1, characterized in that, In step 1, the labelless 3D point cloud scene After asymmetric data augmentation, the point cloud scene is transformed into a strongly augmented data scene. And weak data augmentation point cloud scenarios These are used as the original inputs for the student model and the teacher model, respectively, in a strongly data-augmented point cloud scenario. The result after scaling As an additional input to the student model, This represents the scaling transformation function.

3. The semi-supervised three-dimensional target detection method based on cross-scale distillation according to claim 1, characterized in that, In step 3, voxel anchor matching specifically involves: in the source voxel anchor set... V s and target voxel anchor set V t One-to-one dense matching is performed between them, and the K-nearest neighbor algorithm is used to anchor each source voxel. v s Match the nearest target voxel anchor point v t This yields one-to-one matching voxel anchor pairs, where the Euclidean distance metric is used to calculate the distance, and the specific formula is as follows: in It is the coordinate representation of the voxel anchor point.

4. The semi-supervised three-dimensional target detection method based on cross-scale distillation according to claim 1, characterized in that, In step 3, after obtaining A one-to-one matched voxel anchor pair Then, using the target voxel anchor point Corresponding predicted center score and category scores To eliminate low-confidence matching voxel anchor pairs When the center scores Greater than the center score threshold And category scores Greater than the classification score threshold If the matching voxel anchor pair is valid, it is retained; otherwise, it is discarded.

5. The semi-supervised three-dimensional target detection method based on cross-scale distillation according to claim 1, characterized in that, In step 3, the matching voxel anchor points are... Double consistency learning is performed between them, which includes prediction consistency learning and feature consistency learning.

6. The semi-supervised three-dimensional target detection method based on cross-scale distillation according to claim 5, characterized in that, In step 3, for the matched voxel anchor pairs anchor point with source voxel The corresponding predictions and features are represented as follows: and , , and These represent the parameters of the 3D bounding box, center score, and classification score corresponding to the source voxel anchor point in the prediction, respectively, and the target voxel anchor point. The corresponding predictions and features are represented as follows: and , , and These represent the predicted 3D bounding box parameters, center score, and classification score corresponding to the matched target voxel anchor points, respectively. A prediction consistency learning loss function and a feature consistency learning loss function are constructed separately, and then weighted and summed to obtain a double consistency learning loss function, which is used to train the model. Predictive consistency learning loss function The formula is as follows: in , and Let these represent Huber loss, L2 loss, and KL divergence, respectively. Feature consistency learning loss function The formula is as follows: Double consistency learning loss function The formula is as follows: in Represents the learning loss for predictive consistency. and feature consistency learning loss The balance coefficient between them.

7. The semi-supervised three-dimensional target detection method based on cross-scale distillation according to claim 1, characterized in that, In step 4, there is a label data loss function. The loss function for unlabeled data is calculated based on the ground truth labels of the labeled training data. Losses from transscale internal distillation and cross-scale external distillation losses Adding them together, we get: Among them, the cross-scale internal distillation loss and the cross-scale external distillation loss are the dual consistency learning losses of the cross-scale internal distillation module and the cross-scale external distillation module, respectively. Based on the loss functions for labeled and unlabeled data, the overall loss function of the average teacher semi-supervised learning framework is determined as follows: in There is label loss and unlabeled loss The weighting coefficients between them.

8. A semi-supervised three-dimensional target detection device based on cross-scale distillation, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1-7.

9. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the method as described in any one of claims 1-7.