A Cross-Domain Object Detection Method and Device Based on Temporal Frame Active Learning
Through the method based on time-series frame active learning, the problem of degradation of 3D object detection performance across data sets is solved, and efficient detection of cross-domain data sets is achieved in autonomous driving scenarios, reducing dependence on massive annotated data, and significantly improving detection accuracy and robustness.
Patent Information
- Application Number
- CN202310304971.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-03-27
AI Technical Summary
The prior art is difficult to effectively solve the problem of degradation in 3D object detection performance across data sets, especially in autonomous driving scenarios. Due to the differences in LiDAR parameters of different manufacturers and environmental differences in different cities, the detection accuracy of cross-domain data sets is difficult to reach a satisfactory level.
A cross-domain object detection method based on timing frame active learning is adopted. By obtaining the timing point cloud information of the target domain to be detected, it is divided into single frame point cloud data, and inputting the trained timing frame active learning three-dimensional object detection model, the category and marking box of each object in the point cloud scene are obtained. This method actively learns the sampling strategy based on space-time continuous frames, and selects the most valuable timing frames from all unlabeled timing frames for annotation, reducing the dependence on massive annotated data.
This method greatly reduces the demand for the automatic driving perception model for the annotation of massive timing frames, realizes the performance that can be achieved by fully labeling all timing frames, and significantly improves the accuracy and robustness of 3D object detection in cross-data set scenarios.
Smart Images

Figure CN116740554B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of driverless, and particularly relates to a cross-domain object detection method and device based on active learning of time-series frames. Background Art
[0002] The 3D object detection technology plays a very crucial role in the field of autonomous driving and can help autonomous vehicles perceive the surrounding environment. So far, the most advanced LiDAR-based 3D object detection methods are usually trained and evaluated in a single dataset, and rarely involve the research of cross-domain datasets. However, in many real scenarios of autonomous driving, due to different manufacturers often using lidars with different parameters and the huge environmental differences in different cities, the 3D object detection scheme across datasets has become an urgent problem to be solved in autonomous driving.
[0003] The Active Domain Adaptation (ADA) task is a method of selecting a representative subset of data from the target domain and performing manual annotation. In 2D natural image scenarios such as AADA (see Jong-Chyi Su, Yi-Hsuan Tsai, Kihyuk Sohn, Buyu Liu, Subhransu Maji, and Manmohan Chandraker. Active adversarial domain adaptation. In Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, pages 739–748, 2020), TQS (see Bo Fu, Zhangjie Cao, Jianmin Wang, and Mingsheng Long. Transferable query selection for active domain adaptation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 7272–7281, 2021), CLUE (see Viraj Prabhu, Arjun Chandrasekaran, Kate Saenko, and Judy Hoffman. Active domain adaptation via clustering uncertainty-weighted embeddings. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 8505–8514, 2021), etc., the ADA method has been fully explored. However, it is still blank in the research of 3D point cloud data.
[0004] Some researchers have attempted to address this cross-dataset performance degradation issue through Unsupervised Domain Adaptation (UDA) techniques. SPG (refer to Qiangeng Xu, Yin Zhou, Weiyue Wang, Charles R Qi, and Dragomir Anguelov. Spg: Unsupervised domain adaptation for 3d object detection via semantic point generation. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pages 15446–15456, 2021) designed a semantic point generation method and tried to recover the missing regions of given foreground instances. ST3D (refer to Jihan Yang, Shaoshuai Shi, Zhe Wang, Hongsheng Li, and Xiaojuan Qi. St3d: Self-training for unsupervised domain adaptation on 3d object detection. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 10368–10378, 2021) designed a self-supervised training-based framework to adapt a pre-trained detector from the source domain dataset to a new target domain dataset. LiDAR Distillation (refer to Yi Wei, Zibu Wei, Yongming Rao, Jiaxin Li, Jie Zhou, and Jiwen Lu. Lidar distillation: Bridging the beam-induced domain gap for 3d object detection. arXiv preprint arXiv:2203.14956, 2022) utilizes transferable knowledge obtained from high-beam LiDAR data to distill low-beam LiDAR data. Although these UDA detection methods have achieved success in cross-dataset tasks, there is still a significant detection accuracy gap between them and supervised learning using full annotations.To verify the scalability of the ADA method based on 2D images to 3D point clouds, we directly integrated existing 2D-image-based ADA methods (such as TQS and CLUE) into many typical 3D baseline detectors for research, but it could not achieve satisfactory results in solving the differences in cross-domain datasets. Summary of the Invention
[0005] The purpose of the embodiments of this specification is to provide a cross-domain object detection method and device based on temporal frame active learning.
[0006] To solve the above technical problems, the embodiments of this application are implemented in the following ways:
[0007] In a first aspect, this application provides a cross-domain object detection method based on temporal frame active learning. The method includes:
[0008] Obtain the temporal point cloud information of the target domain to be detected, and divide the temporal point cloud information into single-frame point cloud data;
[0009] Input the single-frame point cloud data into a trained 3D object detection model for temporal frame active learning to obtain the category and bounding box of each object in the point cloud scene;
[0010] Among them, the 3D object detection model for temporal frame active learning is trained using a labeled temporal frame dataset. The labeled temporal frame dataset is obtained by returning the most valuable temporal frames from the full set of unlabeled temporal frames based on a spatio-temporally continuous frame active learning sampling strategy and annotating the most valuable temporal frames.
[0011] In one embodiment, returning the most valuable temporal frames from the full set of unlabeled temporal frames based on a spatio-temporally continuous frame active learning sampling strategy includes:
[0012] Obtain a number of unlabeled temporal frames;
[0013] Input the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames;
[0014] Sort the domain scores of the temporal frames in descending order, and select the temporal frames corresponding to the top preset proportion in the sorting as the most valuable temporal frames.
[0015] In one embodiment, inputting the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames includes:
[0016] Input the unlabeled temporal frames into a detector to obtain a temporal-scene-level saliency description feature map;
[0017] Obtain the domain scores of the temporal frames according to the temporal-scene-level saliency description feature map.
[0018] In one embodiment, an unannotated temporal frame input detector obtains a temporal-scene level saliency description feature map, including:
[0019] An unannotated temporal frame input three-dimensional backbone network obtains a three-dimensional feature description;
[0020] Map the three-dimensional feature description to extract bird's-eye view features to obtain a bird's-eye view feature map;
[0021] Input the bird's-eye view feature map into a two-dimensional backbone network to obtain a two-dimensional feature description;
[0022] Pass the two-dimensional feature description through a region proposal network to obtain an object score;
[0023] Determine the temporal-scene level saliency description feature map according to the object score.
[0024] In one embodiment, determining the temporal-scene level saliency description feature map according to the object score includes:
[0025] Determine an entropy score according to the object score;
[0026] Adopt a consistency evaluation function to determine the temporal information score of the two-dimensional feature descriptions of two adjacent temporal frames;
[0027] Determine the temporal-scene level saliency description feature map according to the object score, entropy score, temporal information score and two-dimensional feature description.
[0028] In one embodiment, the most valuable temporal frames are used to fine-tune the detector.
[0029] In one embodiment, the object loss function of the detector includes a region proposal network loss function, an optimization loss function and a key point segmentation loss function.
[0030] In a second aspect, the present application provides a cross-domain object detection device based on temporal frame active learning. The device includes:
[0031] An acquisition module, configured to acquire the temporal point cloud information of the target domain to be detected, and divide the temporal point cloud information into single-frame point cloud data;
[0032] An object detection module, configured to input the single-frame point cloud data into a trained three-dimensional object detection model for temporal frame active learning, and obtain the category and bounding box of each object in the point cloud scene;
[0033] Wherein, the three-dimensional object detection model for temporal frame active learning is trained using an annotated temporal frame data set. The annotated temporal frame data set is obtained by returning the most valuable temporal frames from the full amount of unannotated temporal frames based on a spatio-temporal continuous frame active learning sampling strategy and annotating the most valuable temporal frames.
[0034] In a third aspect, the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the cross-domain object detection method based on active learning of time-series frames as described in the first aspect.
[0035] In a fourth aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the cross-domain object detection method based on active learning of time-series frames as described in the first aspect.
[0036] As can be seen from the technical solutions provided in the embodiments of this specification above, this solution significantly reduces the need for massive time-series frame annotation in the autonomous driving perception model. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a schematic flowchart of the cross-domain object detection method based on active learning of time-series frames provided by the present application;
[0039] Figure 2 It is a schematic framework diagram of the active learning sampling strategy based on spatio-temporally continuous frames provided by the present application;
[0040] Figure 3 It is a schematic structural diagram of the cross-domain object detection device based on active learning of time-series frames provided by the present application;
[0041] Figure 4 It is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0043] In the following description, specific details such as specific system architectures, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0044] Without departing from the scope or spirit of the present application, various improvements and changes can be made to the specific embodiments of the specification of the present application, which are obvious to those skilled in the art. Other embodiments obtained from the specification of the present application are obvious to those skilled in the art. The specification and embodiments of the present application are merely exemplary.
[0045] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, meaning including but not limited to.
[0046] 3D object detection based on LiDAR sensors is a crucial part of achieving L4-level autonomous driving. However, due to the significant differences in LiDAR parameters of different versions produced by different manufacturers currently, the perception and description of the target scene often change dynamically, further resulting in significant 3D object detection perception results when manufacturers update the LiDAR sensor scheme. For this reason, recent research has focused on exploring some cross-dataset and cross-domain 3D object detection schemes. For example, ST3D has extensively studied cross-sensor 3D object detection schemes in unsupervised domain adaptation scenarios in recent years, improving the robustness of the autonomous driving perception model when dealing with different scenarios, different domains, and different sensor schemes. Although these 3D object detection schemes in unsupervised domain adaptation scenarios can alleviate the requirements for labeling target scenes with domain differences by the 3D object detector, there is still a large difference from the perception performance required for actual applications.
[0047] Existing LiDAR-based autonomous driving perception models need to label a large amount of data to achieve reliable perception performance, which will significantly increase the R & D costs of autonomous driving enterprises. To address the above defects, the present application first proposes an active learning sampling strategy based on spatio-temporal continuous frames, which aims to select a part of the most valuable temporal frames from the full amount of unlabeled temporal frames using an algorithm. Subsequently, the annotation team manually annotates this part of the most valuable temporal frames and feeds the labeled temporal frames into the 3D perception model for training. We find that it can achieve the performance that can be achieved by fully labeling all temporal frames.
[0048] Existing data annotation mitigation strategies mainly target the 2D natural image scenario and fail to consider the per-sequence annotation scenario in the 3D autonomous driving scenario. Therefore, the current data annotation mitigation strategies are difficult to align with the 3D autonomous driving annotation team, resulting in the mismatch between the sampling results and the annotation scheme of the annotation team. To address the above deficiencies, this application comprehensively evaluates the importance of an autonomous driving data sequence from three perspectives: target score, entropy score, and temporal information score, thereby further alleviating the data differences of different sensor schemes and improving the perception performance of the 3D autonomous driving perception model in the face of different scenarios and different domains.
[0049] The basic concept of the active domain adaptation object detection task
[0050] Given a frame of point cloud X ∈ R N×3 , the object detection task is to predict the information of a category and a 3D bounding box for each object in the scene, where N is the number of points contained in a frame of point cloud, and each point in the point cloud contains the (x, y, z) coordinates of the point in the ego-vehicle coordinate system. Existing methods usually use Convolution Neural Network (CNN) to make end-to-end predictions. The domain adaptation object detection task refers to the model being trained on the source dataset and transferring the performance to the target dataset.
[0051] This application adopts the task definition and method of Active Domain Adaptation (ADA) to alleviate the common sensor and data differences in the autonomous driving scenario. Given the labeled source domain set n s representing the total amount of source domain data; the unlabeled target domain set the annotation budget B of the annotation team, where B << n t , n t representing the total cost for the annotation team to fully annotate the target domain data. According to the standard active domain adaptation task setting, a labeled target dataset which is initially empty and updated during the R-round sampling process. In the k-th sampling round, when k < R, select a subset (denoting the dataset of D t excluding ) and manually label it. Then will be updated to After the R-round sampling process, the number of data in reaches the upper limit of the annotation budget B, that is Note that, different from the previous ADA methods, in this application, for the autonomous driving scenario, we adopt a sampling strategy at the temporal frame level, that is saved with temporal information and updated to The same annotation process is also carried out in time series.
[0052] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0053] Refer to Figure 1 , which shows a schematic flowchart of a cross-domain object detection method based on temporal frame active learning provided in an embodiment of the present application.
[0054] As Figure 1 shown, the cross-domain object detection method based on temporal frame active learning may include:
[0055] S110. Obtain the temporal point cloud information of the target domain to be detected, and divide the temporal point cloud information into single-frame point cloud data;
[0056] S120. Input the single-frame point cloud data into a trained three-dimensional object detection model based on temporal frame active learning to obtain the category and bounding box of each object in the point cloud scene;
[0057] Among them, the three-dimensional object detection model based on temporal frame active learning is trained using a labeled temporal frame data set, and the labeled temporal frame data set is obtained by returning the most valuable temporal frames from the full amount of unlabeled temporal frames based on a spatio-temporal continuous frame active learning sampling strategy and annotating the most valuable temporal frames.
[0058] Specifically, in order to verify the effectiveness of the method of the present application, the present application uses a common 3D object detection model, that is, PV-RCNN, as the baseline model of the three-dimensional object detection model based on temporal frame active learning. PV-RCNN is a typical two-stage three-dimensional object detection model that combines the advantages of CNN based on 3D Point and 3D voxel.
[0059] The spatio-temporal continuous frame active learning sampling strategy proposed in the embodiment of the present application aims to select a part of the most valuable temporal frames from the full amount of unlabeled temporal frames by an algorithm, and the annotation team only needs to annotate this part of the most valuable temporal frames output by our algorithm, greatly reducing the dependence of autonomous driving enterprises on data annotation.
[0060] After annotating the most valuable temporal frames, train the three-dimensional object detection model based on temporal frame active learning. Through experimental comparison, it can be found that it can achieve the performance that can be achieved by fully annotating all temporal frames.
[0061] In one embodiment, the most valuable temporal frames are returned from all unlabeled temporal frames based on a spatio-temporal continuous frame active learning sampling strategy, including:
[0062] Obtain a number of unlabeled temporal frames;
[0063] The unlabeled temporal frames are input into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames;
[0064] The domain scores of the temporal frames are sorted in descending order, and the temporal frames corresponding to the top preset proportion in the sorting are selected as the most valuable temporal frames.
[0065] Among them, the unlabeled temporal frames are input into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames, including:
[0066] The unlabeled temporal frames are input into a detector to obtain a spatio-temporal scene-level saliency description feature map;
[0067] According to the spatio-temporal scene-level saliency description feature map, the domain scores of the temporal frames are obtained.
[0068] Among them, the unlabeled temporal frames are input into a detector to obtain a spatio-temporal scene-level saliency description feature map, including:
[0069] The unlabeled temporal frames are input into a 3D backbone network to obtain a 3D feature description;
[0070] The 3D feature description is mapped to extract bird's-eye view features to obtain a bird's-eye view feature map;
[0071] The bird's-eye view feature map is input into a 2D backbone network to obtain a 2D feature description;
[0072] The 2D feature description is passed through a region proposal network to obtain object scores;
[0073] According to the object scores, a spatio-temporal scene-level saliency description feature map is determined.
[0074] Among them, according to the object scores, determining a spatio-temporal scene-level saliency description feature map includes:
[0075] According to the object scores, entropy scores are determined;
[0076] An agreement evaluation function is used to determine the temporal information scores of the 2D feature descriptions of two adjacent temporal frames;
[0077] According to the object scores, entropy scores, temporal information scores, and 2D feature descriptions, a spatio-temporal scene-level saliency description feature map is determined.
[0078] Among them, the detector is fine-tuned using the most valuable temporal frames.
[0079] Among them, the target loss function of the detector includes the region proposal network loss function, the optimization loss function, and the key point segmentation loss function.
[0080] Specifically, referring to Figure 2 , which shows a schematic diagram of the framework of the active learning sampling strategy based on spatio-temporal continuous frames, as Figure 2 In it, the whole method is divided into two stages: 1) the temporal frame sampling / annotation stage; 2) the temporal frame fine-tuning stage. In the temporal frame sampling / annotation stage, the method provided by this application can be used to summarize and generalize a large number of redundant and repetitive temporal frames, so as to select the most representative temporal frames from the large amount of data and provide them to the downstream annotation team for effective temporal frame annotation. For the temporal frame fine-tuning stage, the existing autonomous driving perception models, such as PV-RCNN, are used for model deployment.
[0081] Temporal frame sampling / annotation stage:
[0082] Suppose the autonomous driving system has collected a large amount of unlabeled temporal data within a certain period of time where n t represents the cost required for the autonomous driving manufacturer to perform full-scale data annotation. The temporal frame sampling / annotation stage mainly scores each unlabeled temporal frame one by one through a multi-granularity (object score, entropy score, temporal information score) evaluation criterion, and each temporal frame will be assigned an importance score s, which represents the amount of information contained in the current temporal information in the large amount of unlabeled data. Next, how to obtain the importance score s for each temporal frame will be introduced in detail.
[0083] As Figure 2 shown, the input temporal frame is sent into the 3D backbone network to obtain a 3D feature description, and then the 3D feature description is mapped to extract the Bird Eye View (BEV) feature, so as to obtain a scene-level spatial representation of the current temporal frame. However, due to the sparse distribution of point cloud data, the BEV feature extracted by 3D sparse convolution is also highly sparse. In addition, considering that the input is a sequence of temporal frames, it is necessary to model the temporal information to obtain a scene-level representation along the time dimension.
[0084] For the above two considerations, this application designs a multi-granularity temporal domain discriminator, aiming to merge the domain characteristics of the entire temporal frame by mining the foreground feature regions at the scene level. Specifically, let represent the input temporal frame, where d ∈ [s, t], indicating that the sample x comes from the source domain s or the target domain t. Next, first obtain the BEV feature map f bev = R C×H×W , where C represents the number of channels, and H and W are the height and width of the feature respectively.
[0085] In order to make the discriminator pay more attention to the foreground area, such as Figure 2 As shown, the target score S is first obtained through the Region Proposal Network (RPN) operation. obj ∈R C′×H×W , where C′ represents the number of anchor boxes at each position. The target score represents the probability that the default anchor box belongs to the foreground object. Since the input temporal frame sequence carries stronger uncertainty and domain differences, and in order to better model the spatial features between the instance level and the scene, inspired by previous studies using entropy to measure uncertainty, this application uses the following formula to calculate the entropy score S ent ∈R C′×H×W :
[0086] S ent =-S obj logS obj -(1-S obj )log(1-S obj )
[0087] Where S ent Represents the uncertainty evaluation of the BEV feature describing the current time frame, which describes the uncertainty relationship of all instance objects in the entire time frame to the entire scene. In addition, since the input data is time-series, it is also necessary to consider the time-series information to obtain a reasonable sampling importance score. Assume and It is the two frames before and after in a time series (BEV feature expression is used here). To calculate the time series information score, the consistency of the current two frames needs to be considered. If the consistency is strong, it is considered that the current frames are easy to be recognized by the perception model, so the score is reduced; on the contrary, if the consistency is poor, it is considered that the current frames are abnormal and need to be considered, so the score is increased, as follows:
[0088]
[0089] Among them, Consistent(,) represents the consistency evaluation function for two features. It should be noted that the calculation of the two features and There are many ways to evaluate the consistency, and cosine similarity is used here.
[0090] Finally, combined with S obj ,S ent and S con, a significance description map at the time series - scene level can be obtained, enabling the model to pay more attention to foreground features, uncertainty descriptions, and time series features. The calculation method of the significance description map at the time series - scene level is as follows:
[0091]
[0092] Among them, represents the significance description at the time series - scene level, and are respectively the maximum values along the channel dimension of the S obj , S ent and S con feature vectors.
[0093] Based on the significance description feature map at the time series - scene level , a domain discriminator with a typical convolutional structure is used to distinguish whether the data comes from the source domain or the target domain. In this way, it is modeled whether the current time series frame is close to the previous source domain data distribution, so as to judge whether the current time series frame is "safely migratable". The loss function of the domain discriminator can be written as:
[0094]
[0095] Among them, L dom is the domain loss, H represents the source - target domain discriminator, is the significance description feature map at the source domain time series - scene level, is the significance description feature map at the target domain time series - scene level, and 0 and 1 represent domain labels. If the domain score of the time series frame is closer to 0, it means that the current time series frame is close to the source domain and can be migrated, and no manual annotation is required; if the domain score of the time series frame is close to 1, it means that the current time series frame is close to the target domain, and manual annotation is required. Therefore, finally, the domain scores predicted by the domain discriminator are sorted from high to low, and a batch of higher (i.e., the top preset proportion in the sorting, and this preset proportion can be set according to actual needs) data (i.e., the most valuable time series frames) are selected for manual annotation.
[0096] Time series frame fine - tuning stage
[0097] When the most valuable time series frames are selected in the above - mentioned way, the most valuable time series frames are sent to the annotation team for annotation, and then the perception model is fine - tuned on this batch of annotated data. The overall goal of the fine - tuning process can be expressed as follows:
[0098]
[0099] Among them, L rpn represents the anchor box classification loss function and the anchor box regression loss function to form the RPN loss function. L rcnn represents the optimized loss function, including the prediction loss function guided by the Intersection over Union (IOU) and the bounding box optimization loss function L seg is the key point segmentation loss function.
[0100] To adapt the detector from the source domain to the target domain, the method provided in this application includes three steps. 1) Source-domain Pre-training: First, pre-train the detector on D s to ensure that the detector can learn sufficient knowledge for subsequent model transfer and effective temporal frame sampling; 2) Sequence-level Active Learning Sampling Strategy: In this step, judge whether a temporal frame is worthy of annotation from multiple granularities including the target score, entropy score, and temporal information score, and give the ranking result of the temporal-level importance of all input data, and manually annotate the most important batch of data; 3) Sequence-level Fine-tuning: Based on the above-mentioned manually annotated important temporal frames We use the baseline detector to fine-tune on to reduce the data domain differences between different sensors and manufacturers.
[0101] In the autonomous driving scenario, in order to align the annotation scheme of the annotation team (which often performs a temporal frame-level annotation), this application designs a sequence-level active learning sampling strategy, which can help the annotation team to preferentially annotate the most valuable temporal frames, thereby reducing the annotation cost of the autonomous driving annotation team.
[0102] This application proposes a multi-granularity evaluation criterion for the autonomous driving temporal frame scenario, comprehensively evaluating the importance of a temporal sequence from three perspectives: the target score, entropy score, and temporal information score. Such a multi-granularity evaluation method comprehensively alleviates the data differences between different sensor schemes, thereby improving the perception performance of the 3D autonomous driving perception model in the face of different scenarios and different domains.
[0103] There have been relevant numerical experiments verifying that, compared with existing 3D object detection methods, the cross-domain object detection method based on temporal frame active learning provided in this application has been experimented in many typical cross-dataset scenarios, including cross-beamline, cross-country, and cross-sensor domain adaptation tasks, achieving excellent object domain detection accuracy and verifying the robustness of the model under different scenario domain change conditions. By using the method of this application, a model trained only on 1% labeled KITTI can achieve a 3D perception performance of 89.63%, which is better than the result (88.98%) of training on 100% labeled KITTI.
[0104] Referring to Figure 3 , which shows a schematic structural diagram of a cross-domain object detection device 300 based on temporal frame active learning described according to an embodiment of this application.
[0105] As Figure 3 shown, the cross-domain object detection device based on temporal frame active learning may include:
[0106] An acquisition module 310, configured to acquire the temporal point cloud information of the target domain to be detected and divide the temporal point cloud information into single-frame point cloud data;
[0107] An object detection module 320, configured to input the single-frame point cloud data into a trained three-dimensional object detection model based on temporal frame active learning to obtain the category and bounding box of each object in the point cloud scene;
[0108] Among them, the three-dimensional object detection model based on temporal frame active learning is trained using a labeled temporal frame dataset, and the labeled temporal frame dataset is obtained by returning the most valuable temporal frames from the full amount of unlabeled temporal frames according to a spatio-temporally continuous frame active learning sampling strategy and annotating the most valuable temporal frames.
[0109] Optionally, the cross-domain object detection device based on temporal frame active learning is further configured to:
[0110] Acquire a number of unlabeled temporal frames;
[0111] Input the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames;
[0112] Sort the domain scores of the temporal frames in descending order, and select the temporal frames corresponding to the preset proportion in the sorting as the most valuable temporal frames.
[0113] Optionally, the cross-domain object detection device based on temporal frame active learning is further configured to:
[0114] Input the unlabeled temporal frames into a detector to obtain a temporal-scene level saliency description feature map;
[0115] According to the temporal-scene level saliency description feature map, obtain the domain score of the temporal frame.
[0116] Optionally, the cross-domain object detection device based on active learning of temporal frames is further configured to:
[0117] Input the unannotated temporal frame into the three-dimensional backbone network to obtain a three-dimensional feature description;
[0118] Map the three-dimensional feature description to extract the bird's-eye view feature to obtain a bird's-eye view feature map;
[0119] Input the bird's-eye view feature map into the two-dimensional backbone network to obtain a two-dimensional feature description;
[0120] Pass the two-dimensional feature description through the region proposal network to obtain an object score;
[0121] According to the object score, determine the temporal-scene level saliency description feature map.
[0122] Optionally, the cross-domain object detection device based on active learning of temporal frames is further configured to:
[0123] According to the object score, determine the entropy score;
[0124] Adopt a consistency evaluation function to determine the temporal information score of the two-dimensional feature descriptions of two adjacent temporal frames;
[0125] According to the object score, entropy score, temporal information score and two-dimensional feature description, determine the temporal-scene level saliency description feature map.
[0126] Optionally, the cross-domain object detection device based on active learning of temporal frames is further configured to: fine-tune the detector using the most valuable temporal frames.
[0127] Optionally, the object loss function of the detector includes a region proposal network loss function, an optimization loss function and a key point segmentation loss function.
[0128] The cross-domain object detection device based on active learning of temporal frames provided in this embodiment can execute the embodiments of the above method, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0129] Figure 4 This is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 4 shown, it shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present application.
[0130] As Figure 4As shown, the electronic device 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage section 408 into a random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the device 400 are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0131] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as needed. A removable medium 411, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 410 as needed so that a computer program read from it can be installed into the storage section 408 as needed.
[0132] Specifically, according to an embodiment of the present disclosure, the process described above with reference to Figure 1 can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program tangibly embodied on a machine-readable medium, and the computer program includes program code for performing the above-described cross-domain object detection method based on temporal frames. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from the removable medium 411.
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0134] The units or modules involved in the embodiments described in this application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. The names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0135] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a mobile phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0136] On the other hand, this application also provides a storage medium. The storage medium can be the storage medium included in the aforementioned device in the above embodiments; it can also exist separately and be not assembled into the device. The storage medium stores one or more programs, and the aforementioned programs are used by one or more processors to execute the cross-domain object detection method based on temporal frame active learning described in this application.
[0137] A storage medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0138] It should be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0139] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the description of the method embodiment.
Claims
1. A cross-domain object detection method based on active learning of temporal frames, characterized in that, the method includes: Obtain the temporal point cloud information of the target domain to be detected, and divide the temporal point cloud information into single-frame point cloud data; Input the single-frame point cloud data into the trained three-dimensional object detection model of active learning of temporal frames to obtain the category and bounding box of each object in the point cloud scene; Among them, the three-dimensional object detection model of active learning of temporal frames is trained using a labeled temporal frame dataset, and the labeled temporal frame dataset is obtained by returning the most valuable temporal frames from the full amount of unlabeled temporal frames based on the spatio-temporal continuous frame active learning sampling strategy and labeling the most valuable temporal frames; The spatio-temporal continuous frame active learning sampling strategy returns the most valuable temporal frames from the full amount of unlabeled temporal frames, including: Obtain a number of unlabeled temporal frames; Input the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames; Sort the domain scores of the temporal frames in descending order, and select the temporal frames corresponding to the previous preset proportion in the sorting as the most valuable temporal frames; Input the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames, including: Input the unlabeled temporal frames into a detector to obtain a spatio-temporal scene-level saliency description feature map; Obtain the domain scores of the temporal frames according to the spatio-temporal scene-level saliency description feature map; Input the unlabeled temporal frames into a detector to obtain a spatio-temporal scene-level saliency description feature map, including: Input the unlabeled temporal frames into a three-dimensional backbone network to obtain a three-dimensional feature description; Map the three-dimensional feature description to extract bird's-eye view features to obtain a bird's-eye view feature map; Input the bird's-eye view feature map into a two-dimensional backbone network to obtain a two-dimensional feature description; Pass the two-dimensional feature description through a region proposal network to obtain object scores; Determine the spatio-temporal scene-level saliency description feature map according to the object scores; The determining the spatio-temporal scene-level saliency description feature map according to the object scores includes: Determine entropy scores according to the object scores; Adopt a consistency evaluation function to determine the temporal information scores of the two-dimensional feature descriptions of two adjacent temporal frames; Determine the spatio-temporal scene-level saliency description feature map according to the object scores, the entropy scores, the temporal information scores and the two-dimensional feature descriptions.
2. The method according to claim 1, characterized in that, Fine-tune the detector using the most valuable temporal frames.
3. The method according to claim 2, characterized in that, The object loss function of the detector includes a region proposal network loss function, an optimization loss function and a key point segmentation loss function.
4. A cross-domain object detection device based on active learning of temporal frames, characterized in that, the device includes: An acquisition module for acquiring the temporal point cloud information of the target domain to be detected and dividing the temporal point cloud information into single-frame point cloud data; A target detection module, configured to input the single-frame point cloud data into a trained temporal frame active learning 3D target detection model to obtain the category and bounding box of each object in the point cloud scene; Wherein, the temporal frame active learning 3D target detection model is trained using a labeled temporal frame dataset, and the labeled temporal frame dataset is obtained by returning the most valuable temporal frames from the full amount of unlabeled temporal frames based on a spatio-temporal continuous frame active learning sampling strategy and labeling the most valuable temporal frames; The returning the most valuable temporal frames from the full amount of unlabeled temporal frames based on the spatio-temporal continuous frame active learning sampling strategy includes: Obtaining a plurality of unlabeled temporal frames; Inputting the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames; Sorting the domain scores of the temporal frames in descending order, and selecting the temporal frames corresponding to the top preset proportion in the sorting as the most valuable temporal frames; Inputting the unlabeled temporal frames into a multi-granularity temporal domain discriminator to obtain the domain scores of the temporal frames, including: Inputting the unlabeled temporal frames into a detector to obtain a temporal-scene level saliency description feature map; Obtaining the domain scores of the temporal frames according to the temporal-scene level saliency description feature map; Inputting the unlabeled temporal frames into a detector to obtain a temporal-scene level saliency description feature map, including: Inputting the unlabeled temporal frames into a 3D backbone network to obtain a 3D feature description; Mapping the 3D feature description to extract bird's-eye view features to obtain a bird's-eye view feature map; Inputting the bird's-eye view feature map into a 2D backbone network to obtain a 2D feature description; Passing the 2D feature description through a region proposal network to obtain object scores; Determining the temporal-scene level saliency description feature map according to the object scores; The determining the temporal-scene level saliency description feature map according to the object scores includes: Determining entropy scores according to the object scores; Adopting a consistency evaluation function to determine the temporal information scores of the 2D feature descriptions of two adjacent temporal frames; Determining the temporal-scene level saliency description feature map according to the object scores, the entropy scores, the temporal information scores and the 2D feature descriptions.
5. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the cross-domain target detection method based on temporal frame active learning as described in any one of claims 1-3.
6. A readable storage medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the cross-domain target detection method based on temporal frame active learning as described in any one of claims 1-3.
Citation Information
Patent Citations
Laser radar 3D real-time target detection method fusing multi-frame time sequence point cloud
CN111429514A
Target detection method for reducing labeling demand based on active domain adaptive learning
CN114724015A