A Cross-Domain Object Detection Method Based on Dual-Domain Active Learning
By adopting a cross-domain object detection method based on dual-domain active learning in the field of autonomous driving, combined with domain perception and diversity sampling strategies, the problem of degradation of three-dimensional object detection accuracy across data sets is solved, and higher detection accuracy is achieved.
Patent Information
- Application Number
- CN202310115927.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2043-02-15
AI Technical Summary
The prior art is difficult to realize three-dimensional object detection across data sets in the field of autonomous driving, especially in the lidar parameters of different manufacturers and different urban environments, where the detection accuracy is reduced.
The cross-domain object detection method based on dual-domain active learning is adopted, and the three-dimensional object detection model is pre-trained dual-domain active learning, combined with the source domain sampling strategy based on domain perception and the target domain sampling strategy based on diversity, to improve the adaptability of the detector in the target domain.
The detection accuracy of cross-domain data detection tasks is greatly improved, and is better than the traditional unsupervised domain adaptation method and the active domain adaptation method based on 2D images.
Smart Images

Figure CN116311221B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of driverless, and particularly relates to a cross-domain object detection method based on dual-domain active learning. Background Art
[0002] The 3D object detection technology plays a very crucial role in the field of autonomous driving and can help vehicles perceive the surrounding environment. So far, the most advanced lidar-based 3D object detection methods are usually trained and evaluated in a single dataset, and rarely involve the research on cross-domain datasets. However, in many real scenarios of autonomous driving, due to the fact that different manufacturers often use lidars with different parameters and the environmental differences between different cities are huge, the three-dimensional object detection across datasets has become an urgent problem to be solved in autonomous driving.
[0003] Some researchers have attempted to address this cross-dataset performance degradation issue through Unsupervised Domain Adaptation (UDA) techniques. SPG (refer to Qiangeng Xu, Yin Zhou, Weiyue Wang, Charles R Qi, and Dragomir Anguelov. Spg: Unsupervised domain adaptation for 3d object detection via semantic point generation. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 15446–15456, 2021) designed a semantic point generation method and tried to recover the missing regions of given foreground instances. ST3D (refer to Jihan Yang, Shaoshuai Shi, Zhe Wang, Hongsheng Li, and Xiaojuan Qi. St3d: Self-training for unsupervised domain adaptation on 3d object detection. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 10368–10378, 2021) designed a self-supervised training-based framework to adapt a pre-trained detector from the source domain dataset to a new target domain dataset. LiDAR Distillation (refer to Yi Wei, Zibu Wei, Yongming Rao, Jiaxin Li, Jie Zhou, and Jiwen Lu. Lidar distillation: Bridging the beam-induced domain gap for 3d object detection. arXiv preprint arXiv:2203.14956, 2022) utilizes transferable knowledge obtained from high-beam LiDAR data to distill low-beam LiDAR data. Although these UDA detection methods have achieved success in cross-dataset tasks, there is still a significant detection accuracy gap between them and supervised learning using full annotations.To verify the scalability of the 2D image-based ADA method for 3D point clouds, we directly integrated existing 2D image-based ADA methods (such as TQS (see Bo Fu, Zhangjie Cao, Jianmin Wang, and Mingsheng Long. Transferable query selection for active domain adaptation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 7272–7281, 2021) and CLUE (see Viraj Prabhu, Arjun Chandrasekaran, Kate Saenko, and Judy Hoffman. Active domain adaptation via clustering uncertainty-weighted embeddings. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 8505–8514, 2021)) into many typical 3D baseline detectors for research, but it cannot achieve satisfactory results in solving the differences in cross-domain datasets. Summary of the Invention
[0004] The purpose of the embodiments of this specification is to provide a cross-domain object detection method based on dual-domain active learning.
[0005] To solve the above technical problems, the embodiments of this application are implemented as follows:
[0006] This application provides a cross-domain object detection method based on dual-domain active learning, which is characterized in that the method includes:
[0007] Obtain a frame of point cloud of the target domain to be detected;
[0008] Input the point cloud into the trained dual-domain active learning 3D object detection model to obtain the category and bounding box of each object in the point cloud scene;
[0009] Among them, the dual-domain active learning 3D object detection model is trained using the source domain dataset, and the source domain and the target domain to be detected are different domains.
[0010] In one of the embodiments, the dual-domain active learning 3D object detection model includes:
[0011] Foreground region discriminator, used to determine the domain label and instance-level description of the input point cloud data;
[0012] Domain-aware source domain sampling strategy, used to select class target domain samples from the source domain dataset according to the domain label;
[0013] Diversity-based target domain sampling strategy, used to select diverse and representative target domain data from the target domain dataset according to the instance-level description.
[0014] In one embodiment, the foreground region discriminator includes a detector and a discriminator;
[0015] The detector is used to determine the target score and instance-level description of the input point cloud data;
[0016] The discriminator determines the domain label according to the target score.
[0017] In one embodiment, the detector is used to determine the target score and instance-level description of the input point cloud data, including:
[0018] Extract the 3D features of the input point cloud data by the 3D backbone network, map the extracted features to the bird's-eye view, and obtain the 2D bird's-eye view features;
[0019] Use the 2D backbone network to extract features from the 2D bird's-eye view features, and the extracted features are operated by the region proposal network to obtain the target score and proposal boxes;
[0020] The proposal boxes pass through the detection head to obtain the instance-level description.
[0021] In one embodiment, the proposal boxes pass through the detection head to obtain the instance-level description, including:
[0022] The detection head outputs each interesting feature and the corresponding confidence score;
[0023] Determine the instance-level description according to the interesting features and confidence scores.
[0024] In one embodiment, the discriminator determines the domain label according to the target score, including:
[0025] Determine the entropy value according to the target score;
[0026] Determine the perceptual feature map of the foreground region according to the target score, entropy value, and 2D bird's-eye view features;
[0027] Determine the domain label according to the perceptual feature map of the foreground region.
[0028] In one embodiment, selecting class target domain samples from the source domain dataset according to the domain label includes:
[0029] Use a foreground region discriminator to calculate the scene-level domain feature scores for all source domain data in the source domain dataset;
[0030] Sort the scene-level domain feature scores of all source domain data, and sample the sorted domain feature scores according to preset conditions to obtain class target domain samples.
[0031] In one embodiment, select diverse and representative target domain data from the target domain dataset according to instance-level descriptions, including:
[0032] Calculate the similarity between all instance-level descriptions, and cluster the target data in the target domain dataset according to the similarity into multiple sub-clusters;
[0033] Dynamically update the prototypes of candidate region of interest features in each sub-cluster;
[0034] Use a foreground region discriminator to calculate the domain feature scores for all target domain data in the target domain dataset;
[0035] Select the unlabeled frames with the highest domain feature scores of the target domain data from each updated prototype of candidate region of interest features;
[0036] The unlabeled frames selected from all sub-clusters are combined into a set of data to be labeled;
[0037] Label the data to be labeled in the set of data to be labeled to obtain diverse and representative target domain data.
[0038] In one embodiment, train a dual-domain active learning 3D object detection model, including:
[0039] Obtain the source domain dataset;
[0040] Pre-train the detector on the source domain dataset;
[0041] Determine the domain labels of the discriminator according to the target scores determined by the pre-trained detector;
[0042] Select class target domain samples from the source domain dataset according to the domain labels based on a domain-aware source domain sampling strategy;
[0043] Fine-tune the detector according to the class target domain samples;
[0044] Based on a diversity-based target domain sampling strategy, select diverse and representative target domain data according to the class target domain samples and the fine-tuned detector;
[0045] Re-train the fine-tuned detector according to the class target domain samples and the diverse and representative target domain data until the number of training times is reached.
[0046] In one embodiment, the target loss function of the detector includes a region proposal network loss function, an optimization loss function, and a key point segmentation loss function.
[0047] As can be seen from the technical solutions provided in the embodiments of this specification above, this solution significantly improves the detection accuracy of cross-domain data detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0049] Figure 1 It is a schematic flowchart of a cross-domain object detection method based on dual-domain active learning provided by this application;
[0050] Figure 2 It is a schematic structural diagram of a dual-domain active learning three-dimensional object detection model provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts should belong to the scope protected by this specification.
[0052] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.
[0053] Without departing from the scope or spirit of this application, various improvements and changes can be made to the specific embodiments of this application specification, which are obvious to those skilled in the art. Other embodiments obtained from the specification of this application are obvious to those skilled in the art. The specification and embodiments of this application are only exemplary.
[0054] The terms "comprising", "including", "having", "containing", etc. used in this article are all open-ended terms, meaning including but not limited to.
[0055] Unless otherwise specified, the "parts" in this application are based on mass parts.
[0056] 3D object detection is a crucial part in the field of autonomous driving. So far, the most advanced lidar-based 3D object detection methods are usually trained and evaluated in a single dataset. However, in many real-world scenarios of autonomous driving, due to the differences in lidar parameters and the frequent changes in target scenes, cross-dataset 3D object detection has become an urgent problem to be solved in autonomous driving. Recent studies such as ST3D have extensively explored self-supervised training 3D frameworks under the Unsupervised Domain Adaptation (UDA) task, which helps to improve the cross-dataset adaptability of 3D detectors. However, there is still a large performance gap between the unsupervised learning achieved by these UDA works and the supervised learning using a fully supervised framework. In practice, some unlabeled frames can be sampled from the new target domain for manual annotation, that is, the Active Domain Adaptation (ADA) task. However, the active domain adaptation task is mainly for 2D scenarios and has poor effects when directly integrated into 3D point clouds. Compared with 2D images, 3D point clouds are highly sparse, which makes it difficult for traditional 2D models designed for dense pixels to extract scene-level features; in addition, the data for autonomous driving mainly consists of road scenes, which are relatively simple compared to 2D natural images, and 2D active domain adaptation methods usually select target data based on scene-level prototypes, making it easy to select similar data and resulting in labeling redundancy.
[0057] Based on the above deficiencies, this application first studies the active domain adaptation problem in 3D autonomous driving scenarios and proposes a cross-domain object detection method based on bi-domain active learning, which significantly improves the detection accuracy of cross-domain data detection tasks. This method uses a pre-trained bi-domain active learning 3D object detection (Bi-domain Active Learning for Cross-dataset 3D Object Detection, Bi3D) framework (or model) to predict the category and bounding box (or 3D box) of each object in the point cloud of the input target domain. The Bi3D framework includes a domain-aware source domain sampling strategy and a diversity-based target domain sampling strategy. Among them, the domain-aware source domain sampling strategy is a source domain sampling strategy that narrows the domain gap according to the domain characteristics corresponding to the class target source data. The diversity-based target domain sampling strategy selects diverse and representative target data based on instance-level features on the basis that the scene-level features are relatively similar to the 3D scene. Thereby, the accuracy of cross-domain data detection tasks can be improved.
[0058] Basic concept of active domain adaptation object detection task
[0059] Given a frame of point cloud \(X\in\mathbb{R}\) N×3 , the object detection task is to predict the information of a category and a 3D box for each object in the scene, where \(N\) is the number of points contained in a frame of point cloud, and each point in the point cloud contains the \((x, y, z)\) coordinates of the point in the ego-vehicle coordinate system. Existing methods usually use a Convolution Neural Network (CNN) to make end-to-end predictions. The cross-domain object detection task refers to the model being trained on the source dataset and its performance being transferred to the target dataset.
[0060] This application uses the method of Active Domain Adaptation (ADA) to solve this problem. Given a labeled source domain set n s representing the total amount of source domain data; an unlabeled target domain set with an annotation budget \(B\), where \(B\ll n\) t ,n t representing the total amount of target domain data. According to the standard active domain adaptation task setting, a diversity-based target domain sampling strategy is used to construct a labeled target dataset This dataset is initially empty and is updated during the \(R\) rounds of sampling. In the \(k\)-th sampling round, when \(k < R\), from (denoting the dataset \(D\) t excluding ) a subset And manually label. Then will be updated to After R rounds of sampling process, the number of data in reaches the upper limit of the annotation budget B, that is Note that different from the previous ADA methods, in this application, we construct a source subset using a domain-aware source domain sampling strategy by sampling from the original source dataset D s The goal of the proposed Bi3D in the present invention is: 1) select class target domain data from D s and 2) select the data with the largest amount of information from D t to form and and make the 3D detector better adapt to the target domain by jointly training on the sets of and
[0061] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0062] Referring to Figure 1 , which shows a schematic flow diagram of a cross-domain object detection method based on dual-domain active learning provided in the embodiments of the present application.
[0063] As Figure 1 shown, the cross-domain object detection method based on dual-domain active learning may include:
[0064] S110. Obtain a frame of point cloud of the target domain;
[0065] S120. Input the point cloud into the trained dual-domain active learning three-dimensional object detection model to obtain the category and bounding box of each object in the point cloud scene;
[0066] Among them, the dual-domain active learning three-dimensional object detection model is trained using the source domain dataset, and the source domain and the target domain are different domains.
[0067] Specifically, according to the previous research on cross-domain datasets for 3D object detection, we use PV-RCNN as the baseline model of the dual-domain active learning three-dimensional object detection model. PV-RCNN is a typical two-stage three-dimensional detection framework that combines the advantages of point-based and 3D voxel-based CNNs.
[0068] In one embodiment, the dual-domain active learning three-dimensional object detection model includes:
[0069] A foreground region discriminator for determining the domain label and instance-level description of the input point cloud data;
[0070] Domain-aware source domain sampling strategy for selecting class target domain samples from the source domain dataset according to domain labels;
[0071] Diversity-based target domain sampling strategy for selecting diverse and representative target domain data from the target domain dataset according to instance-level descriptions.
[0072] Among them, the foreground region discriminator includes a detector and a discriminator;
[0073] The detector is used to determine the target score and instance-level description of the input point cloud data;
[0074] The discriminator determines the domain label according to the target score.
[0075] Among them, the detector is used to determine the target score and instance-level description of the input point cloud data, including:
[0076] The 3D features of the input point cloud data are extracted by a 3D backbone network (or called a three-dimensional backbone network or simply 3D backbone), and the extracted features are mapped to a bird's-eye view to obtain 2D bird's-eye view features;
[0077] A 2D backbone network (or called a two-dimensional backbone network or simply 2D backbone) is used to extract features from the 2D bird's-eye view features, and the extracted features are operated by a region proposal network to obtain the target score and proposal boxes;
[0078] The proposal boxes pass through the detection head to obtain the instance-level description.
[0079] Among them, the proposal boxes pass through the detection head to obtain the instance-level description, including:
[0080] The detection head outputs each feature of interest and the corresponding confidence score;
[0081] According to the features of interest and the confidence score, the instance-level description is determined.
[0082] Given a frame of data from the target domain The k-th region of interest (ROI) feature is Then we can easily obtain the corresponding confidence score through a standard post-processing process For example, the non-maximum suppression (NMS) algorithm. We first recalculate the weights of all ROI instance-level features of the current frame using the confidence score to obtain a more accurate instance-level description Among them
[0083] Among them, the discriminator determines the domain label according to the target score, including:
[0084] Determine the entropy value according to the target score;
[0085] Determine the perceptual feature map of the foreground area according to the target score, entropy value, and 2D bird's-eye view features;
[0086] Determine the domain label according to the perceptual feature map of the foreground area.
[0087] Specifically, to effectively measure the properties of the source domain and the target domain, we first designed a foreground area discriminator (or called foreground area perception discriminator). Then, based on the discriminator, a dual-domain active sampling strategy was proposed to transfer the pre-trained 3D detector from the source domain to the new target domain.
[0088] Among them, the input point cloud data is the 3D point cloud data in the source domain dataset.
[0089] Considering that the instance-level features lose the spatial relationship with the original scene, and a large number of negative anchor boxes have a greater interference on the discriminator learning, we obtain the scene-level spatial representation by extracting the Bird Eye View (BEV) features. However, as mentioned before, due to the sparse distribution of the point cloud data, the BEV features extracted by the 3D sparse convolution are also highly sparse. Therefore, the traditional discriminator cannot focus on the information-rich foreground area, resulting in a deviation in the prediction of the domain-related representation under the cross-dataset feature differences.
[0090] To address this issue, this application designs a foreground area discriminator, aiming to measure the domain characteristics of the source domain data and the target data by judging the foreground feature area at the scene level. Specifically, let represent the input point cloud data, where d ∈ [s, t], indicating that the sample x comes from the source domain s or the target domain t. Next, the 3D features are first extracted by the 3D backbone F 3d and then converted into 2D BEV features f bev = R C×H×W , where C represents the number of channels, and H and W are the height and width of the features respectively. Then, the 2D backbone is used to extract features from the 2D BEV features.
[0091] To make the discriminator pay more attention to the foreground area, the features extracted by the 2D backbone are obtained through the Region Proposal Network (RPN) operation to get the target score (or called target score) S obj ∈ R C′×H×W, where C′ represents the number of anchor boxes at each position, and the object score represents the probability that the default anchor box belongs to the foreground object. To better extract the spatial features between instances and scenes, inspired by the method of using entropy to measure uncertainty by predecessors, we use the following formula to calculate the entropy value (or entropy score) S ent ∈R C′×H×W :
[0092] S ent =-S obj log S obj -(1 - S obj ) log(1 - S obj )
[0093] where S ent represents the uncertainty that the spatial position is divided into instance objects, which means the spatial relationship of instance objects distributed in the whole scene. According to the above formula, combining S obj and S ent we can obtain the scene-level attention map, making the model pay more attention to the foreground features. The calculation method of the perceptual feature map of the foreground area is as follows:
[0094]
[0095] where, represents the BEV feature of the foreground area perception (i.e., the perceptual feature map of the foreground area), and are the maximum values of S obj and S ent along the channel dimension respectively.
[0096] Based on the perceptual feature map of the foreground area , we use a domain discriminator with a typical convolutional structure to distinguish whether the data comes from the source domain or the target domain. Specifically, 0 and 1 represent the domain labels of the discriminator. When the value output by the discriminator is close to 0, it means the data comes from the source domain, and if it is close to 1, it means the data comes from the target domain. The loss function of the discriminator can be written as:
[0097]
[0098] where L dom is the domain loss, H represents the source-target domain discriminator, is the perceptual feature map of the foreground area of the source domain; is the perceptual feature map of the foreground area of the target domain.
[0099] Previous data analysis work mainly focused on how to make full use of representative data in the target domain, while ignoring the domain characteristics and effectiveness of source domain data. However, there is actually an overlap in the data distribution of a certain amount of data from the source domain and the target domain, which means that the feature distributions represented by these source domain data may be similar to the feature distributions represented by the target domain data.
[0100] Based on this, the present application proposes a simple and effective domain-aware source domain sampling strategy, aiming to select target domain-like samples from the source domain.
[0101] In one embodiment, selecting target domain-like samples from the source domain dataset according to domain labels includes:
[0102] Using a foreground region discriminator to calculate the scene-level domain characteristic scores of all source domain data in the source domain dataset;
[0103] Sorting the scene-level domain characteristic scores of all source domain data, and sampling the sorted domain characteristic scores according to preset conditions to obtain target domain-like samples.
[0104] Specifically, using the discriminator in the foreground region discriminator to calculate the scene-level domain characteristic scores of all source domain data where where, is the source domain foreground region perception feature map, which can be considered as a similarity metric between the source domain data and the target data. A higher value indicates that the i-th frame of data in the source domain dataset conforms to the data distribution of the target domain dataset. To select source domain data with high domain characteristic scores, we simply perform a descending sort on, and for the sorted data, we can sample according to preset conditions (the preset conditions can be set according to actual needs, such as by proportion or threshold, etc.) to construct Exemplarily, extract 30% of the larger data from the sorted .
[0105] Note and D t have a small domain difference, so by fine-tuning the detector on the performance of the model on the target domain can be improved. Therefore, this detector can extract more accurate instance-level features, which is further beneficial to selecting more informative target data.
[0106] To make the detector better adapt to the target domain, we first fine-tune the detector on Fine-tuning is performed on it, and representative data is selected from the target domain. However, due to the relatively similar semantics of adjacent frames in the autonomous driving scenario, traditional active learning methods (such as committee voting and uncertainty query) have encountered challenges. They often select samples with very small differences between classes, resulting in the problem of redundant annotations. For this reason, we designed a diversity-based target sampling strategy to select diverse and representative target domain data (or called labeled target domain data).
[0107] In one embodiment, selecting diverse and representative target domain data from the target domain dataset according to the instance-level description includes:
[0108] Calculate the similarity between all instance-level descriptions, cluster the target data in the target domain dataset according to the similarity, and cluster them into multiple sub-clusters;
[0109] Dynamically update the prototype of the candidate region of interest features in each sub-cluster;
[0110] Use the foreground region discriminator to calculate the domain characteristic scores of all target domain data in the target domain dataset;
[0111] Select the unlabeled frame with the highest domain characteristic score of the target domain data from each updated prototype of the candidate region of interest features;
[0112] The unlabeled frames selected from all sub-clusters are combined into a set of data to be labeled;
[0113] Label the data to be labeled in the set of data to be labeled to obtain diverse and representative target domain data.
[0114] Specifically, the basic idea of the diversity-based target sampling strategy is to maintain a similarity library, and cluster all unlabeled target data in the target domain dataset by judging the semantic similarity of the re-weighted ROI features, clustering them into multiple sub-clusters to ensure the diversity of the selected target data.
[0115] This application uses the following formula to dynamically update the prototype of the candidate ROI features in each sub-cluster:
[0116]
[0117] Where c i , c j are the i-th and j-th budget prototypes allocated according to the budget, which means that each budget is represented by a prototype, where c i , c j The initial values of are both instance-level descriptions. P i and P jDenote the similarity library of the above-mentioned $i$-th and $j$-th budget prototypes, which is used to cache unlabeled frames, and $num(\cdot)$ represents the number of unlabeled frames in the buffer.
[0118] To sample more diverse and representative target domain data from the target domain, we select an unlabeled frame with the highest domain feature score from each updated budget prototype $c$. Form a complete set of all data to be labeled (i.e., unlabeled target domain data). Then annotation can be performed, and manual annotation can be carried out.
[0119] Based on the diversity-based target sampling strategy in the embodiments of this application, diverse and representative target data are selected based on instance-level features, effectively improving the task accuracy.
[0120] In one embodiment, training a dual-domain active learning 3D object detection model includes:
[0121] Obtain the source domain dataset;
[0122] Pre-train the detector on the source domain dataset;
[0123] Determine the domain labels of the discriminator according to the object scores determined by the pre-trained detector;
[0124] Select class target domain samples from the source domain dataset according to the domain labels based on the domain-aware source domain sampling strategy;
[0125] Fine-tune the detector according to the class target domain samples;
[0126] Based on the diversity-based target domain sampling strategy, select diverse and representative target domain data according to the class target domain samples and the fine-tuned detector;
[0127] Retrain the fine-tuned detector according to the class target domain samples and the diverse and representative target domain data until the number of training times is reached.
[0128] Specifically, to enable the detector to adapt from the source domain to the target domain, the training of the dual-domain active learning 3D object detection model in the method of this application includes three steps: 1) Source-domain Pre-training: First, pre-train the detector on the source domain dataset $D$ s to ensure that the detector can learn sufficient knowledge for subsequent model transfer; 2) Active Sampling Source Domain: In this step, we select the source domain data of the class target domain (i.e., class target domain samples), and The detector on it is fine-tuned to reduce the domain difference; 3) Target domain active sampling: Based on the selected source domain data above and the fine-tuned detector, we further sample the most informative (i.e., diverse and representative) target domain data and and retrain the detector on it.
[0129] The structure of the dual-domain active learning 3D object detection model is as Figure 2 shown. When training this model: First, obtain the source domain dataset (this part is not shown in Figure 2 ), and first train the detector through the 3D point cloud data in the entire source domain dataset. The RPN in the detector outputs the target score S obj , and then train the discriminator through S obj . The discriminator is used to judge whether the data comes from the source domain or the target domain, and then calculate the scene-level domain feature scores of all source domain data in the source domain dataset according to the discriminator. Select the class target domain samples according to the scene-level domain feature scores of the source domain data (i.e., the source domain data similar to the target domain in Figure 2 ); The various interesting features and corresponding confidence scores output by the detection head in the detector are used to obtain the instance-level description. Cluster the target domain dataset according to the instance-level description, and then select the unlabeled frames from it according to the clustering result (i.e., the unannotated target domain data in Figure 2 ), and then through annotation, obtain the diverse and representative target domain data (i.e., the annotated target domain data in Figure 2 ); The selected class target domain samples and the diverse and representative target domain data are used as the dataset to train the detector again. Repeat this process until the preset conditions are met (such as reaching the number of training times or the training result converges, etc.) to complete the training.
[0130] In one embodiment, the target loss function of the detector includes a region proposal network loss function, an optimization loss function, and a key point segmentation loss function.
[0131] Specifically, the target loss function L det of the detector is:
[0132]
[0133] Among them, L rpn represents the RPN loss function composed of the anchor box classification loss function and the anchor box regression loss function . L rcnn represents the optimization loss function, which includes the prediction loss function guided by the intersection over union (IOU) And the bounding box optimization loss function L seg is the key point segmentation loss function.
[0134] The cross-domain object detection method based on dual-domain active learning provided by the embodiments of the present application has been experimented in many typical cross-dataset scenarios, including cross-beam, cross-country, and cross-sensor domain adaptation tasks, and achieved excellent object domain detection accuracy. Among them, Bi3D (89.63%) trained only on 1% labeled KITTI is better than the corresponding baseline model (88.98%) trained using 100% labeled KITTI data.
[0135] The cross-domain object detection strategy based on ADA provided by the present application greatly improves the task accuracy.
[0136] It should be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0137] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment.
Claims
1. A cross - domain object detection method based on dual - domain active learning, characterized in that, The method includes: Obtain a frame of point cloud of the target domain to be detected; Input the point cloud into the trained dual-domain active learning 3D object detection model to obtain the category and bounding box of each object in the point cloud scene; The dual-domain active learning 3D object detection model includes a foreground region discriminator for determining the domain label and instance-level description of the input point cloud data; The foreground region discriminator includes a detector and a discriminator; The detector is used to determine the object score and the instance-level description of the input point cloud data; The discriminator determines the domain label according to the object score; Among them, the dual-domain active learning 3D object detection model is trained using a source domain dataset and a target domain dataset, and the source domain and the target domain to be detected are different domains; Training the dual-domain active learning 3D object detection model includes: Obtain a source domain dataset; Pre-train the detector on the source domain dataset; Determine the domain label of the discriminator according to the object score determined by the pre-trained detector; Select class target domain samples from the source domain dataset according to the domain label based on the domain-aware source domain sampling strategy; Fine-tune the detector according to the class target domain samples; Based on the diversity-based target domain sampling strategy, select diverse and representative target domain data according to the class target domain samples and the fine-tuned detector; Retrain the fine-tuned detector according to the class target domain samples and the diverse and representative target domain data until the number of training times is reached.
2. The method according to claim 1, characterized in that, The dual-domain active learning 3D object detection model further includes: A domain-aware source domain sampling strategy for selecting class target domain samples from the source domain dataset according to the domain label; A diversity-based target domain sampling strategy for selecting diverse and representative target domain data from the target domain dataset according to the instance-level description.
3. The method according to claim 1, characterized in that, The detector is used to determine the object score and the instance-level description of the input point cloud data, including: Extract the 3D features of the input point cloud data by a 3D backbone network, and map the extracted features to a bird's-eye view to obtain 2D bird's-eye view features; Use a 2D backbone network to extract features from the 2D bird's-eye view features, and the extracted features are operated by a region proposal network to obtain an object score and proposal boxes; The proposal boxes pass through the detection head to obtain the instance-level description.
4. The method according to claim 3, characterized in that, The proposal boxes pass through the detection head to obtain the instance-level description, including: The detection head outputs each region of interest feature and the corresponding confidence score; Determine the instance-level description according to the region of interest feature and the confidence score.
5. The method according to claim 3, characterized in that, The discriminator determines the domain label according to the object score, including: Determine the entropy value according to the object score; Determine the perceptual feature map of the foreground region according to the object score, the entropy value, and the 2D bird's-eye view features; Determine the domain label according to the perceptual feature map of the foreground region.
6. The method according to any one of claims 2 - 5, characterized in that, The selecting class target domain samples from the source domain dataset according to the domain label includes: Calculate the scene-level domain characteristic scores of all source domain data in the source domain dataset using the foreground region discriminator; Sort the scene-level domain feature scores of all the source domain data, and sample the sorted domain feature scores according to preset conditions to obtain the class target domain samples.
7. The method according to any one of claims 2-5, characterized in that, Selecting the target domain data with diversity and representativeness from the target domain dataset according to the instance-level description includes: Calculate the similarity between all instance-level descriptions, and cluster the target data in the target domain dataset according to the similarity into multiple sub-clusters; Dynamically update the prototype of the candidate region of interest features in each of the sub-clusters; Use the foreground region discriminator to calculate the domain feature scores of all the target domain data in the target domain dataset; Select the unlabeled frame with the highest domain feature score of the target domain data from each updated prototype of the candidate region of interest features; The unlabeled frames selected from all the sub-clusters are combined into a set of data to be labeled; Label the data to be labeled in the set of data to be labeled to obtain the target domain data with diversity and representativeness.
8. The method according to any one of claims 4-5, characterized in that, The target loss function of the detector includes a region proposal network loss function, an optimization loss function, and a key point segmentation loss function.
Citation Information
Patent Citations
Domain adaptive target detection method and system considering category semantic matching
CN113807420A
Training method based on image-instance alignment network and cross-domain target detection method
CN114693983A