Millimeter wave point cloud semi-supervised open set target detection method based on unknown class generation

By introducing semi-supervised learning and speed information processing methods in millimeter-wave radar point cloud target detection, unknown target detection problems are solved, and efficient millimeter-wave point cloud target detection is achieved, suitable for real open scenarios.

CN120198646APending Publication Date: 2025-06-24XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510328274.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing millimeter-wave radar point cloud target detection methods have degraded performance in real open scenarios, especially when unknown targets exist, and lack effective utilization of millimeter-wave point cloud characteristics.

Method used

A semi-supervised open-set object detection method based on unknown classes is proposed. By obtaining the millimeter wave point cloud data to be detected, the basic model of trained open-set detection is extracted, and combining the semi-supervised learning framework and speed information can alleviate the problem of pseudo-label inaccuracy.

Benefits of technology

It realizes that while ensuring the detection and recognition accuracy of known object, it accurately detects and recognizes unknown object, and reduces labeling pressure. It has high mobility and is suitable for different point cloud object detection frameworks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198646A_ABST
    Figure CN120198646A_ABST
Patent Text Reader

Abstract

The invention discloses a millimeter wave point cloud semi-supervised open set target detection method based on unknown class generation, and relates to the technical field of target detection, and the method comprises the steps: obtaining to-be-detected millimeter wave point cloud data; the trained open set detection basic model is adopted to process millimeter wave point cloud data to be detected, part of features are extracted through a first branch to obtain first point-by-point features, part of features are extracted through a second branch to obtain second point-by-point features, the first point-by-point features and the second point-by-point features are spliced to obtain spliced features, and the spliced features are used for detecting the millimeter wave point cloud data to be detected. And according to the spliced features, extracting a target candidate region, and performing fine tuning on the target candidate region to obtain a detection result of the millimeter wave point cloud data to be detected. According to the method, the millimeter wave point cloud characteristics and the speed measurement capability of the millimeter wave radar are effectively utilized, and accurate detection and recognition of unknown targets can be realized while the detection and recognition precision of known targets is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and particularly relates to a millimeter-wave point cloud semi-supervised open-set target detection method based on unknown class generation. Background Technique

[0002] As one of the key sensors in the field of autonomous driving, millimeter-wave radar has received extensive attention due to its advantages such as low cost, high range / velocity resolution, and all-weather operation. The work of target detection using the point cloud data obtained by millimeter-wave radar has also become a research hotspot.

[0003] Most of the existing point cloud target detection methods are based on the closed-set assumption, that is, the environmental scene is fixed or similar, and the target categories are known in advance. However, in the real open scene, there may be unknown class targets without annotations in the training data, resulting in a decline in the performance of traditional closed-set detection models when applied to real open scenes. In addition, compared with optical images, point cloud data also has the problem of difficult annotation. Most of the existing open-set target detection and semi-supervised target detection work is for optical image data or lidar point cloud data, which may cause new problems when applied to millimeter-wave radar point cloud data, and lacks the utilization of the characteristics of millimeter-wave point clouds, resulting in poor method performance. In addition, there is relatively little existing research on open-set semi-supervised target detection. Millimeter-wave radar point clouds contain information such as the position and velocity of targets. How to make full use of the target motion characteristics contained in millimeter-wave radar point clouds and design a reasonable open-set semi-supervised target detection framework is an urgent problem to be considered. Summary of the Invention

[0004] In order to solve the above problems existing in the prior art, the present invention provides a millimeter-wave point cloud semi-supervised open-set target detection method based on unknown class generation. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0005] In the first aspect, the present invention provides a millimeter-wave point cloud semi-supervised open-set target detection method based on unknown class generation, including:

[0006] Obtain the millimeter-wave point cloud data to be detected;

[0007] Process the millimeter-wave point cloud data to be detected by using a trained open-set detection basic model, extract partial features through the first branch to obtain the first per-point feature, extract partial features through the second branch to obtain the second per-point feature, splice the first per-point feature and the second per-point feature to obtain the spliced feature, extract the target candidate region according to the spliced feature, and fine-tune the target candidate region to obtain the detection result of the millimeter-wave point cloud data to be detected;

[0008] Among them, the trained open-set detection basic model is obtained by training the initial open-set detection basic model with data of preset categories as the training data set.

[0009] Advantages of the present invention:

[0010] A millimeter-wave point cloud semi-supervised open-set object detection method based on unknown class generation provided by the present invention effectively utilizes the characteristics of millimeter-wave point clouds and the speed measurement ability of millimeter-wave radars, and can accurately detect and identify unknown class objects while ensuring the detection and recognition accuracy of known class objects. In addition, the present invention combines a semi-supervised training scheme, alleviates the impact of inaccurate pseudo-labels on semi-supervised learning based on speed information, and reduces the annotation pressure. In addition, the method proposed by the present invention has high transferability and can be conveniently applied to different point cloud object detection frameworks.

[0011] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0012] Figure 1 is a flowchart of a millimeter-wave point cloud semi-supervised open-set object detection method based on unknown class generation provided by an embodiment of the present invention;

[0013] Figure 2 is a schematic diagram of an open-set detection basic model provided by an embodiment of the present invention;

[0014] Figure 3 is a schematic diagram of a splicing method provided by an embodiment of the present invention;

[0015] Figure 4 is a schematic diagram of a cropping method provided by an embodiment of the present invention. Detailed Embodiments

[0016] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0017] In the prior art, when an unknown class object not labeled in the training data appears, a closed-set object detection model will misjudge the unknown class object as the background or a known class. At the same time, since millimeter-wave radar point clouds contain speed information, the network may learn an incorrect speed pattern during the training stage, that is, if the speed value of a certain object is particularly small, it will be determined as a certain known class, thus exacerbating the confusion between unknown classes and known classes. In addition, existing semi-supervised methods also lack the utilization of the characteristics of millimeter-wave radar point clouds themselves.

[0018] In view of this, a millimeter-wave point cloud semi-supervised open-set object detection method based on unknown class generation proposed by the present invention effectively utilizes the speed measurement ability of millimeter-wave radar and combines the characteristics of millimeter-wave radar point cloud for method design, improves the network's ability to detect / discriminate unknown class objects, and designs a semi-supervised learning framework to reduce the annotation pressure; in addition, the solution proposed by the present invention has high transferability and can be conveniently applied to different point cloud detection frameworks.

[0019] Please refer to Figure 1 , Figure 1 FIG. is a flowchart of a millimeter-wave point cloud semi-supervised open-set object detection method based on unknown class generation provided by an embodiment of the present invention. A millimeter-wave point cloud semi-supervised open-set object detection method based on unknown class generation provided by the present invention includes:

[0020] S101. Obtain millimeter-wave point cloud data to be detected.

[0021] S102. Process the millimeter-wave point cloud data to be detected by using a trained open-set detection basic model. Extract partial features through the first branch to obtain the first per-point feature, extract partial features through the second branch to obtain the second per-point feature, splice the first per-point feature and the second per-point feature to obtain the spliced feature, extract target candidate regions according to the spliced feature, and fine-tune the target candidate regions to obtain the detection result of the millimeter-wave point cloud data to be detected;

[0022] Among them, the trained open-set detection basic model is obtained by training an initial open-set detection basic model with data of a preset class as the training data set.

[0023] Specifically, please refer to Figure 2 , Figure 2 FIG. is a schematic diagram of an open-set detection basic model provided by an embodiment of the present invention. In this embodiment, before training the initial open-set detection basic model, it further includes:

[0024] Construct an open-set detection basic model; among them,

[0025] The open-set detection basic model includes a first branch, a second branch, a first multi-layer perceptron, a second multi-layer perceptron, a first classification head, a regression head, and a two-stage network;

[0026] The first branch includes a first feature extraction module, a first multi-layer perceptron combination, a localization head, a second multi-layer perceptron combination, a second classification head, a third multi-layer perceptron combination, and a foreground / background classification head; among them, the first multi-layer perceptron combination, the second multi-layer perceptron combination, and the third multi-layer perceptron combination each include a third multi-layer perceptron and a fourth multi-layer perceptron; among them, the first multi-layer perceptron, the second multi-layer perceptron, the third multi-layer perceptron, and the fourth multi-layer perceptron each include a fully connected layer, a batch normalization layer, and a non-linear function activation layer, and the ReLU function can be used; optionally, the input dimension of the third multi-layer perceptron and the fourth multi-layer perceptron is 128, and the output dimension is 64; the inputs of the localization head, the second classification head, and the foreground / background classification head are all 64, the output dimension of the localization head is 8, the output dimension of the second classification head is the number of classification categories, and the output dimension of the foreground / background classification head is 1;

[0027] The second branch includes a second feature extraction module; among them, the parameters of the first feature extraction module and the second feature extraction module are not shared.

[0028] Optionally, the two-stage network is built using the method proposed in the prior art literature "Shaoshuai Shi, Xiaogang Wang, Hongsheng Li. PointRCNN: 3D Object Proposal Generation and Detection from Point Cloud[C]. IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019" to build the two-stage network in the open-set object detection network for point clouds.

[0029] In this embodiment, the trained open-set detection base model is used to process the millimeter-wave point cloud data to be detected. Part of the features are extracted through the first branch to obtain the first point-by-point feature, and part of the features are extracted through the second branch to obtain the second point-by-point feature. The first point-by-point feature and the second point-by-point feature are concatenated to obtain the concatenated feature. Based on the concatenated feature, target candidate regions are extracted and fine-tuned to obtain the detection result of the millimeter-wave point cloud data to be detected, including:

[0030] Based on the three-dimensional coordinates and velocity features of the millimeter-wave point cloud data to be detected, the trained first feature extraction module is used for feature extraction to obtain the first point-by-point feature; the trained first multi-layer perceptron combination is used to process the first point-by-point feature to obtain the first feature after dimensionality reduction; the trained second multi-layer perceptron combination is used to process the first point-by-point feature to obtain the second feature after dimensionality reduction; the trained third multi-layer perceptron combination is used to process the first point-by-point feature to obtain the third feature after dimensionality reduction;

[0031] Weight the second feature after dimensionality reduction to obtain the weighted second feature; concatenate the first feature after dimensionality reduction, the weighted second feature, and the third feature after dimensionality reduction to obtain the first concatenated feature fmerge, and its expression is:

[0032]

[0033] Among them, f_loc represents the first feature after dimensionality reduction, f_K represents the second feature after dimensionality reduction, w represents the weighting value for weighting the second feature after dimensionality reduction, and f_foreground represents the third feature after dimensionality reduction. represents concatenation, and v represents the velocity of the millimeter-wave point cloud data to be detected;

[0034] Process the first concatenated feature using the trained first multi-layer perceptron to obtain the first concatenated feature after dimensionality reduction; optionally, the input dimension of the first multi-layer perceptron is 192, and the output dimension is 128;

[0035] Based on the three-dimensional spatial coordinates, scattering cross-section area, distance, azimuth angle, and elevation angle features of the millimeter-wave point cloud data to be detected, use the trained second feature extraction module to extract features to obtain the second per-point feature;

[0036] Concatenate the first concatenated feature after dimensionality reduction and the second per-point feature to obtain the second concatenated feature;

[0037] Process the second concatenated feature using the trained second multi-layer perceptron to obtain the second concatenated feature after dimensionality reduction; optionally, the input dimension of the second multi-layer perceptron is 256, and the output dimension is 128;

[0038] Process the second concatenated feature after dimensionality reduction using the trained first classification head to obtain the first feature; process the second concatenated feature after dimensionality reduction using the trained regression head to obtain the second feature; extract the target candidate region according to the first feature and the second feature;

[0039] Process the target candidate region using the trained two-stage network to obtain the detection result of the millimeter-wave point cloud data to be detected.

[0040] In this embodiment, training the initial open-set detection base model includes:

[0041] Obtain data of multiple preset categories; among them, the data of the preset categories is millimeter-wave point cloud data;

[0042] In the first-stage training, select data of some preset categories for annotation to obtain the annotation information of the data of these preset categories as the true labels, and use the data of these preset categories as the original samples in the first training dataset; generate pseudo-unknown class samples according to the original samples in the first training dataset; screen the pseudo-unknown class samples according to the preset first condition to update the samples in the first training dataset;

[0043] Input some samples in the updated first training dataset into the j-th open-set detection base model to be trained to obtain the prediction results output during the j-th training;

[0044] Calculate the first loss according to the prediction results output during the j-th training and the true labels of the samples used to train the j-th open-set detection base model, and use it as the first loss during the j-th training process;

[0045] Perform backpropagation according to the first loss during the j-th training process to update the network parameters of the j-th open-set detection base model to be trained, and obtain the (j + 1)-th open-set detection base model to be trained; iterate in this way until the number of training times or the convergence degree meets the preset conditions, and obtain the open-set detection base model trained in the first stage;

[0046] In the second-stage training, select data of some preset categories as the samples in the second training dataset; optionally, select the remaining preset category data in this stage;

[0047] Input the samples in the second training dataset into the open-set detection base model trained in the first stage to obtain the prediction results; if the confidence value in the prediction results is greater than the confidence threshold, use the prediction results as the pseudo-labels of the samples in the second training dataset;

[0048] Calculate the velocity-related vector and label-related vector of the samples in the second training dataset, obtain the mask vector according to the velocity-related vector and label-related vector; calculate the second loss according to the mask vector; perform backpropagation according to the second loss to update the network parameters of the open-set detection base model trained in the first stage, and iterate in this way until the number of training times or the convergence degree meets the preset conditions, and obtain the open-set detection base model trained in the second stage;

[0049] Use the updated open-set detection base model to obtain the pseudo-labels of the samples in the second training dataset; calculate the mask vector according to the sample velocity-related vector and label-related vector in the second training dataset; calculate the second loss according to the mask vector and perform backpropagation to update the parameters of the updated open-set detection base model until the number of repeated rounds of updating the parameters of the updated open-set detection base model meets the preset conditions, and obtain the trained open-set detection base model.

[0050] Optionally, for the samples in the second training dataset without annotation information, first calculate the speed-related vector, arrange it based on the speeds of all millimeter-wave point clouds, set the vector values of the k millimeter-wave point clouds with the minimum speed to 0, and the others to 1; secondly, calculate the annotation-related vector, set the vector values of the millimeter-wave point clouds located within the annotation box to 1, and the others to 0; perform an exclusive OR operation on the speed-related vector and the annotation-related vector to obtain a mask vector.

[0051] In this embodiment, in the first training stage, the expression of the first loss L1 is:

[0052]

[0053] where λ1, λ2, λ3, λ4, λ5, λ6, and λ7 respectively represent different weight factors, Lloc represents the loss output by the regression head, Lloc_v represents the loss output by the localization head, and Lstage_2 represents the loss output by the two-stage network;

[0054] Lcls represents the loss output by the first classification head, and its expression is:

[0055]

[0056] where N represents the total number of samples in the updated first training dataset, i represents the sample index, c represents the category, yic represents the prediction result output by the first classification head, represents the true label of the sample, and K represents the number of classification categories;

[0057] Lcls_v represents the loss output by the second classification head, and its expression is:

[0058]

[0059] where yic_v represents the prediction result output by the second classification head;

[0060] Lfbc_v represents the loss output by the foreground-background classification head, and its expression is:

[0061]

[0062] where, represents the prediction result output by the foreground-background classification head, and yi represents the true label of the foreground-background of the sample;

[0063] Lunk_cal represents the loss of unknown class score correction, and its expression is:

[0064]

[0065] Among them, yi,unk represents the prediction result output by the first classification head, M represents the total number of non-unknown class samples in the updated first training dataset, and represents the set of unknown class point clouds.

[0066] Lloc represents the loss output by the regression head, and Lloc_v represents the loss output by the localization head. The forms of the two loss functions are the same, and only the point clouds located within the target bounding box will participate in the calculation of the loss. The form of the loss function adopts the Smooth-L1 loss, and the specific formula is as follows:

[0067]

[0068] Among them, Ntarget represents the number of point clouds located within the bounding box, u represents the target parameter that needs to participate in the loss calculation, represents the true value of the target parameter, upred represents the predicted value of the target parameter output by the network, and μ is a hyperparameter. The specific target parameters include:

[0069]

[0070] Among them, xgt, ygt, ygt represent the three-dimensional space coordinates corresponding to the center point of the target bounding box where the point cloud is located, xp, yp, zp represent the three-dimensional space coordinates corresponding to the point cloud, l gt 、w gt 、h gt represent the lengths of the bounding box of the target where the point cloud is located in the x, y, and z directions, and θ gt represents the orientation angle of the target bounding box where the point cloud is located.

[0071] In this embodiment, in the second training stage, for the samples with a value of 0 in the mask vector, the expression of the second loss L2 is:

[0072]

[0073] Among them, λ8, λ9, λ10, λ11, λ12, and λ13 respectively represent different weight factors, Lloc represents the loss output by the regression head, Lloc_v represents the loss output by the localization head, and Lstage_2 represents the loss output by the two-stage network;

[0074] Lcls represents the loss output by the first classification head, and its expression is:

[0075]

[0076] Among them, N represents the total number of samples in the second-stage training, i represents the sample index, c represents the category, yic represents the prediction result output by the first classification head, represents the pseudo-label of the sample, and K represents the number of classification categories;

[0077] Lcls_v represents the loss output by the second classification head, and its expression is:

[0078]

[0079] Among them, yic_v represents the prediction result output by the second classification head;

[0080] Lfbc_v represents the loss output by the foreground-background classification head, and its expression is:

[0081]

[0082] Among them, represents the prediction result output by the foreground-background classification head, and yi represents the pseudo-label of the foreground-background of the sample.

[0083] Lloc represents the loss output by the regression head, and Lloc_v represents the loss output by the localization head. The two loss functions have the same form, and only the point cloud located within the target bounding box will participate in the calculation of the loss. The form of the loss function adopts the Smooth-L1 loss, and the specific formula is as follows:

[0084]

[0085]

[0086] Among them, Ntarget represents the number of point clouds located within the bounding box, u represents the target parameter that needs to participate in the loss calculation, represents the true value of the target parameter, upred represents the predicted value of the target parameter output by the network, and μ is a hyperparameter. The specific target parameters include:

[0087]

[0088] Among them, xgt, ygt, ygt represent the three-dimensional space coordinates corresponding to the center point of the target bounding box where the point cloud is located, xp, yp, zp represent the three-dimensional space coordinates corresponding to the point cloud, l gt 、w gt 、h gt represent the lengths of the bounding box of the target where the point cloud is located in the x, y, and z directions, and θ gt represents the orientation angle of the target bounding box where the point cloud is located.

[0089] In the second training stage, the samples with a value of 1 in the mask vector do not participate in the calculation of the loss output by the first classification head, and the target candidate regions generated by these samples are classified as the background class by the first classification head. The expression of the second loss L2 is:

[0090] L2 = λ13Lstage_2;

[0091] Among them, λ13 respectively represent different weight factors, and Lstage_2 represents the loss output by the two-stage network;

[0092] In the second training stage, for the samples with a value of 1 in the mask vector, they do not participate in the loss calculation of the output of the first classification head, and the target candidate regions generated by these samples, if not classified as the background class by the first classification head, do not participate in the loss calculation of the output of the two-stage network.

[0093] In this embodiment, screening pseudo-unknown class samples according to a preset first condition to update the samples in the first training dataset includes:

[0094] Obtain the bounding box of the pseudo-unknown class sample; optionally, use the minimum bounding rectangle of the pseudo-unknown class sample as the bounding box annotation, and the length and width values of the bounding box can be appropriately increased;

[0095] Calculate the intersection over union (IOU) between the bounding box of the pseudo-unknown class sample and the bounding boxes of the original samples in the first training dataset, and calculate the IOU between the bounding box of the pseudo-unknown class sample and the bounding boxes of other pseudo-unknown class samples;

[0096] If the IOU values are all 0, then use the annotation information of this pseudo-unknown class sample as the true label of this pseudo-unknown class sample, as the sample in the updated first training dataset, and delete the annotation information of the original sample corresponding to the generation of this pseudo-unknown class sample; otherwise, retain the annotation information of the original sample corresponding to the generation of this pseudo-unknown class sample.

[0097] This embodiment also includes: enhancing the samples in the updated first training dataset, and the enhancement methods include global rotation, local rotation, and horizontal rotation.

[0098] In this embodiment, generating pseudo-unknown class samples according to the original samples in the first training dataset includes:

[0099] Traverse the original samples in the first training dataset. If the current original sample is a target sample of a preset category and the millimeter wave point cloud data of the current original sample is greater than 3, then there is a preset probability of being selected as the original sample;

[0100] According to the preset probability, generate pseudo-unknown class samples by magnifying the selected original samples; among them, the magnifying methods include:

[0101] Use the size expansion method to generate pseudo-unknown class samples; or,

[0102] Use the splicing method to generate pseudo-unknown class samples.

[0103] Optionally, there is a 30% probability of generating pseudo-unknown class samples for the selected original samples, and a 35% probability of generating pseudo-unknown class samples by magnification.

[0104] Specifically, there is a 50% probability that the process of generating pseudo-unknown class samples by using the size expansion method is as follows:

[0105] Translate the center of the millimeter-wave point cloud cluster corresponding to the original sample to the coordinate origin, multiply the coordinates of the millimeter-wave point cloud corresponding to the original sample by a random number between 1.5 and 3, then translate the millimeter-wave point cloud corresponding to the original sample back to the original position, and update the distance and azimuth angle features of the millimeter-wave point cloud corresponding to the original sample to generate pseudo-unknown class samples.

[0106] There is a 50% probability that the process of generating pseudo-unknown class samples by using the splicing method is as follows:

[0107] Translate the millimeter-wave point cloud corresponding to the original sample along the target orientation direction and in the direction where the y-axis coordinate increases by a distance d to obtain a replicated millimeter-wave point cloud, where d represents the side length of the original sample's bounding box in the translation direction. Please refer to Figure 3 , Figure 3 is a schematic diagram of a splicing method provided by an embodiment of the present invention; update the distance and azimuth angle features of the replicated millimeter-wave point cloud, randomly discard 20% of the replicated millimeter-wave point cloud, and merge the remaining replicated millimeter-wave point cloud with the millimeter-wave point cloud corresponding to the original sample to generate pseudo-unknown class samples.

[0108] In this embodiment, generating pseudo-unknown class samples according to the original samples in the first training dataset includes:

[0109] Traverse the original samples in the first training dataset. If the current original sample is a target sample of a preset category and the millimeter-wave point cloud data of the current original sample is greater than 3, then there is a preset probability of being a selected original sample;

[0110] According to a preset probability, generate pseudo-unknown class samples by using a shrinking method according to the selected original samples; among them, the shrinking method includes:

[0111] Generate pseudo-unknown class samples by using the size reduction method; or,

[0112] Generate pseudo-unknown class samples by using the cropping method.

[0113] Optionally, there is a 30% probability of generating pseudo-unknown class samples for the selected original samples, and a 35% probability of generating pseudo-unknown class samples by using a shrinking method.

[0114] Specifically, there is a 50% probability that the process of generating pseudo-unknown class samples by using the size reduction method is as follows:

[0115] Translate the center of the millimeter-wave point cloud cluster corresponding to the original sample to the origin of coordinates, multiply the coordinates of the millimeter-wave point cloud corresponding to the original sample by a random number between 0.25 and 0.5, then translate the millimeter-wave point cloud corresponding to the original sample back to the original position, randomly discard 20% of the millimeter-wave point cloud, update the distance and azimuth angle features of the millimeter-wave point cloud corresponding to the original sample, and generate a pseudo-unknown class sample.

[0116] There is a 50% probability. When using the cropping method, the process of generating a pseudo-unknown class sample is as follows:

[0117] Delete the millimeter-wave point cloud above the straight line perpendicular to the target orientation where the center of the millimeter-wave point cloud cluster corresponding to the original sample is located. If the number of remaining millimeter-wave point clouds is less than 3, cancel this operation and return the millimeter-wave point cloud corresponding to the original sample. Otherwise, use the remaining millimeter-wave point cloud after deletion as a pseudo-unknown class sample. Please refer to Figure 4 , Figure 4 is a schematic diagram of a cropping method provided by an embodiment of the present invention.

[0118] In this embodiment, according to the original samples in the first training dataset, generating pseudo-unknown class samples includes:

[0119] Traverse the original samples in the first training dataset. If the current original sample is a target sample of a preset category and the millimeter-wave point cloud data of the current original sample is greater than 3, then there is a preset probability of being selected as the original sample;

[0120] According to the selected original sample, generate a pseudo-unknown class sample in a rotation and splicing manner according to a preset probability.

[0121] Optionally, there is a 30% probability of generating a pseudo-unknown class sample for the selected original sample, and a 30% probability of generating a pseudo-unknown class sample in a rotation and splicing manner.

[0122] Specifically, when using the rotation and splicing method, the process of generating a pseudo-unknown class sample is as follows:

[0123] Copy the millimeter-wave point cloud corresponding to the original sample, and perform the above-mentioned cropping operation. Rotate the remaining millimeter-wave point cloud by a random angle between -90° and 90°, rotate based on the center of the remaining millimeter-wave point cloud after cropping, update the distance and azimuth angle features of the rotated millimeter-wave point cloud, add Gaussian random noise with a mean of 0 and a variance of 0.04 to the velocity and cross-sectional area features, and merge the rotated millimeter-wave point cloud data with the millimeter-wave point cloud corresponding to the original sample to generate a pseudo-unknown class sample.

[0124] In summary, the millimeter-wave point cloud semi-supervised open-set target detection method based on unknown class generation provided by the present invention effectively utilizes the characteristics of millimeter-wave point clouds and the speed measurement ability of millimeter-wave radars, and can accurately detect and identify unknown class targets while ensuring the detection and recognition accuracy of known class targets. In addition, the present invention combines a semi-supervised training scheme, alleviates the impact of inaccurate pseudo-labels on semi-supervised learning based on speed information, and reduces the annotation pressure. In addition, the method proposed by the present invention has high transferability and can be conveniently applied to different point cloud target detection frameworks.

[0125] In an optional embodiment of the present invention, the effect of the millimeter-wave point cloud semi-supervised open-set target detection method based on unknown class generation provided in the above embodiment is verified through a simulation experiment, specifically as follows:

[0126] I. Simulation process;

[0127] The simulation experiment platform of this embodiment is configured as follows: CPU: Inter Xeon(R) CPU E5-2620 v3 @ 2.40GHz, GPU: GeForce GTX TITAN X, ubuntu 16.04, and the deep learning environment is CUDA 10.0, pytorch 1.2.0.

[0128] To verify the effectiveness of the method proposed by the present invention, experiments are carried out on the publicly available millimeter-wave radar point cloud dataset Radarscenes. The targets in the dataset are divided into 5 categories: cars, large vehicles, two-wheeled vehicles, pedestrians, and crowds, obtaining training data containing 130 sequences, test data containing 28 sequences, and the bounding box annotation information for each frame. The fixed input network point number in the training stage is set to 1024 points, and the input network point number in the test stage is 800 points. If the number of point clouds in the current frame is greater than the threshold, the point clouds are sorted according to the point cloud speed and the part with the minimum speed is discarded; if the number of point clouds in the current frame is greater than the threshold, some point clouds are randomly duplicated. In addition, the k value in the above is fixed to 600, the μ value is fixed to 1, and the weight term values in the first loss and the second loss are both fixed to 1. It should be noted that only the targets of the car category are used in this simulation experiment to generate pseudo-unknown classes.

[0129] According to the semi-supervised open-set experiment setting, 10% of the training data is randomly selected as labeled data, and the remaining data is used as unlabeled data. The two-wheeler is set as the unknown class, and the point cloud frames containing two-wheelers only appear in the unlabeled data and the test data. The labeled data training phase has a total of 20 epochs, the batch size is set to 32, and the learning rate is set to 0.01. The model update phase using unlabeled data has a total of 5 epochs, the batch size is set to 32, and the learning rate is set to 0.01. The self-learning process is repeated 4 times. The confidence threshold for pseudo-labels is set as follows: car 0.9, large vehicle 0.9, pedestrian 0.7, crowd 0.7, unknown class (two-wheeler) 0.7.

[0130] The experimental comparison methods include: 1. Training an open-set model using the method proposed in the present invention, but in the update phase using unlabeled data, only based on basic self-learning, without using the speed-based label mask proposed in the present invention; 2. Training an open-set detection basic model based on the existing point cloud open-set method MLUC and using unlabeled data for network update based on the method proposed in the present invention. The evaluation metrics include overall Precision, overall Recall, overall F1-score, and F1-score for the unknown class. Here, "overall" means calculating the scores by comprehensively considering the detection and recognition results of known and unknown classes, and the confidence threshold is set to 0.5.

[0131] II. Analysis of simulation experiment results;

[0132] As can be seen from Table 1, when not using the speed-based label mask, although the false alarm rate of the network is very low and the accuracy is very high, the recall rate is very low because the target will be learned as the background. The improved method effectively improves the recall rate and obtains a higher F1-score. Compared with other methods, both the overall F1-score and the F1-score for the unknown class corresponding to the method proposed in the present invention are higher, verifying that the method proposed in the present invention has better unknown class discrimination ability and can effectively utilize unlabeled data to improve the network detection performance.

[0133] Table 1 Results of the open-set semi-supervised experiment of the method proposed in the present invention

[0134]

[0135] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements includes not only those elements but also other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising said element. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "above", "below", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and thus should not be construed as a limitation on the present invention.

[0136] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0137] The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A semi-supervised open-set target detection method for millimeter-wave point clouds generated from unknown classes, characterized in that: include: Obtain millimeter wave point cloud data to be detected; The millimeter-wave point cloud data to be detected is processed using a trained open set detection basic model, a first branch is used to extract part of the features to obtain a first point-by-point feature, a second branch is used to extract part of the features to obtain a second point-by-point feature, the first point-by-point feature and the second point-by-point feature are spliced ​​to obtain a spliced ​​feature, a target candidate area is extracted according to the spliced ​​feature, the target candidate area is fine-tuned, and a detection result of the millimeter-wave point cloud data to be detected is obtained; The trained open set detection basic model is obtained by training an initial open set detection basic model using data of a preset category as a training data set.

2. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 1 is characterized in that: Before training the initial open set detection base model, it also includes: Construct an open set detection basic model; among them, The open set detection basic model includes a first branch, a second branch, a first multi-layer perceptron, a second multi-layer perceptron, a first classification head, a regression head and a two-stage network; The first branch includes a first feature extraction module, a first multi-layer perceptron combination, a positioning head, a second multi-layer perceptron combination, a second classification head, a third multi-layer perceptron combination and a foreground-background classification head; wherein the first multi-layer perceptron combination, the second multi-layer perceptron combination and the third multi-layer perceptron combination all include a third multi-layer perceptron and a fourth multi-layer perceptron; The second branch includes a second feature extraction module.

3. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 1 is characterized in that: The method uses the trained open set detection basic model to process the millimeter wave point cloud data to be detected, extracts part of the features through the first branch to obtain the first point-by-point features, extracts part of the features through the second branch to obtain the second point-by-point features, splices the first point-by-point features and the second point-by-point features to obtain spliced ​​features, extracts the target candidate area according to the spliced ​​features, and fine-tunes the target candidate area to obtain the detection result of the millimeter wave point cloud data to be detected, including: Based on the three-dimensional coordinates and speed features of the millimeter wave point cloud data to be detected, a trained first feature extraction module is used to perform feature extraction to obtain a first point-by-point feature; a trained first multi-layer perceptron combination is used to process the first point-by-point feature to obtain a first feature after dimensionality reduction; a trained second multi-layer perceptron combination is used to process the first point-by-point feature to obtain a second feature after dimensionality reduction; a trained third multi-layer perceptron combination is used to process the first point-by-point feature to obtain a third feature after dimensionality reduction; The second feature after dimensionality reduction is weighted to obtain a weighted second feature; the first feature after dimensionality reduction, the weighted second feature and the third feature after dimensionality reduction are concatenated to obtain a first concatenated feature fmerge, whose expression is: Among them, fposition represents the first feature after dimensionality reduction, fK represents the second feature after dimensionality reduction, w represents the weighted value of the second feature after dimensionality reduction, and fbackground represents the third feature after dimensionality reduction. Indicates splicing; Using the trained first multi-layer perceptron to process the first splicing feature to obtain a first splicing feature after dimensionality reduction; Based on the three-dimensional spatial coordinates, scattering cross-sectional area, distance, azimuth and elevation angle features of the millimeter wave point cloud data to be detected, a trained second feature extraction module is used to perform feature extraction to obtain a second point-by-point feature; Splicing the first splicing feature after dimensionality reduction and the second point-by-point feature to obtain a second splicing feature; Using the trained second multi-layer perceptron to process the second splicing feature to obtain a second splicing feature after dimensionality reduction; Using the trained first classification head to process the second splicing feature after dimensionality reduction to obtain a first feature; using the trained regression head to process the second splicing feature after dimensionality reduction to obtain a second feature; extracting the target candidate area according to the first feature and the second feature; The trained two-stage network is used to process the target candidate area to obtain the detection result of the millimeter wave point cloud data to be detected.

4. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 1 is characterized in that: The training of the initial open set detection basic model includes: Acquire a plurality of data of the preset categories; wherein the data of the preset categories are millimeter wave point cloud data; In the first stage of training, a portion of the data of the preset category is selected for labeling, and the labeling information of the data of the preset category is obtained as the true label, and the data of the preset category is used as the original sample in the first training data set; according to the original samples in the first training data set, a pseudo unknown class sample is generated; according to the preset first condition, the pseudo unknown class sample is screened to update the samples in the first training data set; Inputting part of the samples in the updated first training data set into the j-th open set detection basic model to be trained for training, and obtaining the prediction result output during the j-th training process; According to the prediction results output in the j-th training process and the true labels of the samples of the j-th open set detection basic model to be trained, the first loss is calculated and used as the first loss of the j-th training process; Back propagation is performed according to the first loss of the j-th training process to update the network parameters of the open set detection basic model to be trained for the j-th time, and the open set detection basic model to be trained for the j+1-th time is obtained; this is repeated until the number of training times or the degree of convergence meets the preset conditions, and the open set detection basic model trained in the first stage is obtained; In the second stage of training, a portion of the preset category of data is selected as samples in the second training data set; Input the samples in the second training data set into the open set detection basic model trained in the first stage to obtain a prediction result; if the confidence value in the prediction result is greater than the confidence threshold, use the prediction result as the pseudo label of the sample in the second training data set; Calculate the speed-related vector and the label-related vector of the samples in the second training data set, and obtain a mask vector according to the speed-related vector and the label-related vector; calculate the second loss according to the mask vector; perform back propagation according to the second loss to update the network parameters of the open set detection basic model trained in the first stage, and iterate until the number of training times or the degree of convergence meets the preset conditions, and obtain the updated open set detection basic model; The updated open set detection basic model is used to obtain pseudo labels for samples in the second training data set; a mask vector is calculated based on the sample speed-related vectors and label-related vectors in the second training data set; a second loss is calculated based on the mask vector, and backpropagation is performed to update the updated open set detection basic model parameters until repeated rounds of updating the updated open set detection basic model parameters meet preset conditions, thereby obtaining a trained open set detection basic model.

5. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 4 is characterized in that: In the first training stage, the first loss L1 is expressed as: Among them, λ1, λ2, λ3, λ4, λ5, λ6 and λ7 represent different weight factors, Lloc represents the loss of the regression head output, Lloc_v represents the loss of the positioning head output, and Lstage_2 represents the loss of the two-stage network output; Lcls represents the loss of the first classification head output, and its expression is: Where N represents the total number of samples in the updated first training data set, i represents the sample index, c represents the category, and yic represents the prediction result output by the first classification head. Represents the true label of the sample, and K represents the number of classification categories; Lcls_v represents the loss of the second classification head output, and its expression is: Among them, yic_v represents the prediction result output by the second classification head; Lfbc_v represents the loss of the foreground and background classification head output, and its expression is: in, represents the prediction result output by the foreground and background classification head, and yi represents the true label of the foreground and background of the sample; Lunk_cal represents the loss of unknown class score correction, and its expression is: Among them, yi,unk represents the prediction result output by the first classification head, M represents the total number of non-unknown class samples in the updated first training data set, and unk represents the unknown class point cloud set.

6. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 4, characterized in that: In the second training stage, for samples with a value of 0 in the mask vector, the expression of the second loss L2 is: Among them, λ8, λ9, λ10, λ11, λ12 and λ13 represent different weight factors, Lloc represents the loss of the regression head output, Lloc_v represents the loss of the positioning head output, and Lstage_2 represents the loss of the two-stage network output; Lcls represents the loss of the first classification head output, and its expression is: Where N represents the total number of samples in the second stage of training, i represents the sample index, c represents the category, and yic represents the prediction result output by the first classification head. Represents the pseudo label of the sample, and K represents the number of classification categories; Lcls_v represents the loss of the second classification head output, and its expression is: Among them, yic_v represents the prediction result output by the second classification head; Lfbc_v represents the loss of the foreground and background classification head output, and its expression is: in, represents the prediction result output by the foreground and background classification head, and yi represents the pseudo label of the foreground and background of the sample; In the second training stage, the sample with a value of 1 in the mask vector and the target candidate region generated by the sample is classified as the background class by the first classification head, and the expression of the second loss L2 is: L2 = λ13Lstage_2; Among them, λ13 represents different weight factors, Lstage_2 represents the loss of the second-stage network output; In the second training stage, the samples whose values ​​in the mask vector are 1 and the target candidate regions generated by the samples are not classified as background classes by the first classification head and do not participate in the calculation of the loss.

7. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 4 is characterized in that: The step of screening the pseudo unknown class samples according to the preset first condition to update the samples in the first training data set includes: Obtaining a bounding box of the pseudo unknown class sample; Calculating the intersection-and-union ratio of the bounding box of the pseudo unknown class sample and the bounding box of the original sample in the first training data set, and calculating the intersection-and-union ratio of the bounding box of the pseudo unknown class sample and the bounding boxes of other pseudo unknown class samples; If the intersection-union ratios are all 0, the labeling information of the pseudo-unknown class sample is used as the true label of the pseudo-unknown class sample, as the sample in the updated first training data set, and the labeling information of the original sample corresponding to the pseudo-unknown class sample is deleted; otherwise, the labeling information of the original sample corresponding to the pseudo-unknown class sample is retained.

8. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 4, characterized in that: The step of generating a pseudo unknown class sample according to the original sample in the first training data set includes: Traversing the original samples in the first training data set, if the current original sample is a target sample of a preset category, and the number of millimeter wave point cloud data of the current original sample is greater than 3, then there is a preset probability that the original sample is selected; According to the preset probability, based on the selected original sample, a pseudo unknown class sample is generated by an amplification method; wherein the amplification method includes: Use the size expansion method to generate pseudo unknown class samples; or, The splicing method is used to generate pseudo unknown class samples.

9. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 4, characterized in that: The step of generating a pseudo unknown class sample according to the original sample in the first training data set includes: Traversing the original samples in the first training data set, if the current original sample is a target sample of a preset category, and the number of millimeter wave point cloud data of the current original sample is greater than 3, then there is a preset probability that the original sample is selected; According to the preset probability, based on the selected original sample, a pseudo unknown class sample is generated by a reduction method; wherein the reduction method includes: Use size reduction methods to generate pseudo unknown class samples; or, The clipping method is used to generate pseudo unknown class samples.

10. The semi-supervised open set target detection method of millimeter wave point cloud based on unknown class generation according to claim 4, characterized in that: The step of generating a pseudo unknown class sample according to the original sample in the first training data set includes: Traversing the original samples in the first training data set, if the current original sample is a target sample of a preset category, and the number of millimeter wave point cloud data of the current original sample is greater than 3, then there is a preset probability that the original sample is selected; According to the preset probability, a pseudo unknown class sample is generated based on the selected original sample by adopting a rotation splicing method.