Few-Shot Object Detection Method and Device Based on Feature Decoupling and Knowledge Transfer
By setting up dual-channel detection heads and knowledge migration technology in the basic detection model, the efficiency and effect problems of small sample object detection in complex backgrounds are solved, and more efficient detection results are achieved.
Patent Information
- Application Number
- CN202111405138.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-11-24
AI Technical Summary
Existing object detection methods are difficult to detect effectively under small sample data, especially in complex backgrounds, resulting in poor detection efficiency and effectiveness.
A small sample object detection method based on feature decoupling and knowledge transfer is proposed. By setting up a dual detection head in the basic detection model, the transfer learning model is adjusted using the labeled small sample data, and the adjusted model is jointly adjusted by inputting the adjusted small sample data without labeling to complete the detection.
The detection efficiency and effectiveness of small sample object detection are improved, and the generalization ability of new categories is improved by retaining the performance of basic categories and realizing learning between different tasks.
Smart Images

Figure CN114332449B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a few-shot object detection method and device based on feature decoupling and knowledge transfer. Background Art
[0002] In recent years, Computer Vision has become one of the research hotspots in the field of artificial intelligence. Object Detection is one of the basic tasks of computer vision, and its main task is to classify and locate objects in images. However, most of the current mature methods rely on large-scale labeled data, which inevitably leads to the simplification of application scenarios and covered tasks. In most scenarios, collecting labeled data that meets the requirements is a labor-intensive, material-intensive, and time-consuming task, which limits the implementation and promotion of existing object detection methods.
[0003] Based on this situation, few-shot learning has gradually attracted attention in the academic community. It studies how to use a small amount of labeled data to improve the generalization performance of the model by learning the commonalities in different sub-tasks, so as to meet actual needs. Few-shot object detection is a further expansion of few-shot learning, and the core problem is how to locate objects belonging to this category in complex backgrounds through a small amount of labels of unseen classes. Summary of the Invention
[0004] This application aims to at least solve one of the technical problems in the related technologies to some extent.
[0005] To this end, the first object of this application is to propose a few-shot object detection method based on feature decoupling and knowledge transfer to improve the detection efficiency and detection effect of few-shot object detection.
[0006] The second object of this application is to propose a few-shot object detection device based on feature decoupling and knowledge transfer.
[0007] To achieve the above object, a few-shot object detection method based on feature decoupling and knowledge transfer proposed in the first aspect embodiment of this application includes:
[0008] Set a dual-path detection head in the basic detection model to obtain a transfer learning model;
[0009] Use the small amount of labeled data to adjust the transfer learning model to obtain an adjusted transfer learning model;
[0010] Input the unlabeled small-sample data into the adjusted transfer learning model, and jointly adjust the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unlabeled small-sample data.
[0011] Optionally, in an embodiment of the present application, before setting a dual-path detection head in the basic detection model to obtain a transfer learning model, it further includes:
[0012] Determine the learning rate and the training length, and use the base class samples to train the basic detection model until the learning rate and the training length converge.
[0013] Optionally, in an embodiment of the present application, the adjusting the transfer learning model using the labeled small-sample data includes:
[0014] Use the base class samples to train the basic branch and the dynamic branch of the dual-path detection head, so as to initialize the network parameters of the basic branch and the dynamic branch;
[0015] Use the labeled small-sample data to update the initialized network parameters of the dynamic branch to obtain the first network parameters.
[0016] Optionally, in an embodiment of the present application, the inputting the unlabeled small-sample data into the adjusted transfer learning model and jointly adjusting the region proposal network and the dual-path detection head of the adjusted transfer learning model includes:
[0017] Input the unlabeled small-sample data into the adjusted transfer learning model, and determine the detection score of the unlabeled small-sample data;
[0018] Based on the detection score of the unlabeled small-sample data and the labeled small-sample data, determine a mixed data set;
[0019] Use the mixed data set to jointly adjust the region proposal network and the first network parameters of the dynamic branch of the adjusted transfer learning model, and fix the initialized network parameters of the basic branch to obtain the final transfer learning model.
[0020] Optionally, in an embodiment of the present application, the inputting the unlabeled small-sample data into the adjusted transfer learning model and determining the detection score of the unlabeled small-sample data includes:
[0021] The adjusted transfer learning model converts the unlabeled small-sample data into region of interest (RoI) features, and inputs the flattened RoI features into the dual-path detection head;
[0022] The basic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain base class scores;
[0023] The dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain new class scores;
[0024] Determine the detection scores of the unlabeled small-sample data according to the base class scores and the new class scores.
[0025] Optionally, in an embodiment of the present application, the base branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain base class scores; the dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain new class scores, including:
[0026] The base branch and the dynamic branch perform feature transformation on the flattened RoI features according to the regression sub-task and the classification sub-task;
[0027] The regression sub-task is used to perform regression on the coordinate offset of the expanded RoI features;
[0028] The classification sub-task is used to balance the mean-variance difference between the base class features and the new class features in the expanded RoI features.
[0029] Optionally, in an embodiment of the present application, the determining the detection scores of the unlabeled small-sample data according to the base class scores and the new class scores includes:
[0030] Use the base class scores to regularize and constrain the base class sub-scores in the new class scores to obtain the constrained new class scores;
[0031] Determine the detection scores of the unlabeled small-sample data according to the base class scores and the constrained new class scores.
[0032] Optionally, in an embodiment of the present application, the determining the mixed dataset based on the detection scores of the unlabeled small-sample data and the labeled small-sample data includes:
[0033] Sort the detection scores of the unlabeled small-sample data by confidence, and determine the pseudo-labeled small-sample data according to the detection scores of the first preset number;
[0034] Mix the pseudo-labeled small-sample data and the labeled small-sample data in a preset ratio to obtain a mixed dataset.
[0035] Optionally, in an embodiment of the present application, it further includes: setting an RoI feature discriminator parallel to the dual-path detection head in the basic detection model;
[0036] The transfer learning model converts the input small-sample data into RoI features, and inputs the flattened RoI features into the RoI feature discriminator;
[0037] The RoI feature discriminator classifies the base class features and new class features in the flattened RoI features by using at least two fully connected layers.
[0038] In summary, for the method proposed in the first aspect embodiment of this application, a transfer learning model is obtained by setting a dual-path detection head in the basic detection model; the transfer learning model is adjusted by using a small sample of labeled data to obtain an adjusted transfer learning model; the unlabeled small sample data is input into the adjusted transfer learning model, and the region proposal network and the dual-path detection head of the adjusted transfer learning model are jointly adjusted, so as to complete the detection of the unlabeled small sample data. This application retains the performance of the basic category through the dual-path detection head, realizes learning between different tasks, and improves the detection efficiency and detection effect of small sample object detection by jointly adjusting the region proposal network and the dual-path detection head during the process of small sample object detection.
[0039] To achieve the above object, a small sample object detection device based on feature decoupling and knowledge transfer proposed in the second aspect embodiment of this application includes:
[0040] A model determination module, configured to obtain a transfer learning model by setting a dual-path detection head in the basic detection model;
[0041] A model adjustment module, configured to adjust the transfer learning model by using a small sample of labeled data to obtain an adjusted transfer learning model;
[0042] A sample detection module, configured to input the unlabeled small sample data into the adjusted transfer learning model, and jointly adjust the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unlabeled small sample data.
[0043] In summary, for the device proposed in the second aspect embodiment of this application, the model determination module obtains a transfer learning model by setting a dual-path detection head in the basic detection model; the model adjustment module adjusts the transfer learning model by using a small sample of labeled data to obtain an adjusted transfer learning model; the sample detection module inputs the unlabeled small sample data into the adjusted transfer learning model, and jointly adjusts the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unlabeled small sample data. This application retains the performance of the basic category through the dual-path detection head, realizes learning between different tasks, and improves the detection efficiency and detection effect of small sample object detection by jointly adjusting the region proposal network and the dual-path detection head during the process of small sample object detection.
[0044] The additional aspects and advantages of this application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of this application. Brief Description of the Drawings
[0045] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, where:
[0046] Figure 1 is a flowchart of a few-shot object detection method based on feature decoupling and knowledge transfer provided by an embodiment of the present application;
[0047] Figure 2 is a schematic structural diagram of a dual-path detection head provided by an embodiment of the present application;
[0048] Figure 3 is a schematic flowchart of semi-supervised proposal region mining provided by an embodiment of the present application;
[0049] Figure 4 is a schematic structural diagram of a transfer learning model provided by an embodiment of the present application;
[0050] Figure 5 is a schematic structural diagram of a few-shot object detection device based on feature decoupling and knowledge transfer provided by an embodiment of the present application. Detailed Description of the Embodiments
[0051] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary only for explaining the present application and should not be construed as a limitation of the present application. On the contrary, the embodiments of the present application include all variations, modifications, and equivalents falling within the spirit and scope of the appended claims.
[0052] Currently, there are mainly the following two few-shot object detection methods:
[0053] The first method: Meta Learning method. To solve the few-shot problem, the goal of meta learning is to acquire task-level meta-knowledge to help the model quickly adapt to new tasks and environments using a small number of labeled samples. Such methods usually introduce a meta-learner and perform episodic training using a dataset containing a support set and a query set. Some works focus on obtaining a good parameter initialization to adapt to new tasks with a small number of random gradient descents; another direction is to use parameter generation to generate the weights of the classifier for new categories when adapting to new tasks.
[0054] The second method: the transfer learning method. This type of method first enables the model to acquire rich knowledge of basic categories, and then uses a small number of samples of unseen tasks to fine-tune the detection model. By minimizing the classification posterior probability gap between the source domain and the target domain, knowledge is transferred from a large dataset to a small dataset. There is also work that uses a balanced subset to fine-tune the last layer of the target detector, indicating that the feature representation learned from the base classes can be transferred to new classes, and simple adjustments to the detection network can provide a strong performance gain.
[0055] However, the first method introduces additional meta-learning model parameters during training, increasing the space complexity. Moreover, when the number of categories in the support set increases, the scenario learning used in the meta-learning method is inefficient, resulting in a high time complexity. The second method performs the transfer from the source domain to the target domain through small-sample fine-tuning. However, it only focuses on the performance of the model on new categories while ignoring the retention of knowledge of the model on the basic categories, leading to catastrophic forgetting during the transfer learning process. In addition, the second method freezes the backbone network part of the model and only fine-tunes the last classification layer, resulting in a large number of false negative samples in the region proposal network, and thus the optimal performance cannot be achieved.
[0056] The following will explain the present application in detail with specific embodiments.
[0057] Figure 1 It is a flowchart of a small-sample object detection method based on feature decoupling and knowledge transfer provided by an embodiment of the present application.
[0058] As Figure 1 shown, a small-sample object detection method based on feature decoupling and knowledge transfer provided by an embodiment of the present application includes the following steps:
[0059] Step 110, set a dual-path detection head in the basic detection model to obtain a transfer learning model;
[0060] Step 120, use the small-sample data with annotations to adjust the transfer learning model to obtain an adjusted transfer learning model;
[0061] Step 130, input the unannotated small-sample data into the adjusted transfer learning model, and jointly adjust the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unannotated small-sample data.
[0062] It should be noted that in small sample object detection, the categories of a dataset are divided into basic category samples (base class samples) and new category samples (new class samples). The base class samples contain rich annotated sample data, and the new class samples contain unannotated small sample data and a small amount of annotated small sample data. Small sample object detection refers to learning prior knowledge from sufficient base class samples, and then using a small amount of annotated small sample data in the new class samples to achieve good generalization on the new class.
[0063] In some embodiments, the basic detection model is based on Faster R-CNN and includes a backbone network, a feature pyramid network, a region proposal network, and a RoI pooling layer, where:
[0064] The input sample is successively subjected to feature extraction by the backbone network and the feature pyramid network to generate multi-scale image features with different downsampling rates. The downsampling rate can be, for example, from 2x to 64x, and the backbone network can be, for example, ResNet-101.
[0065] The region proposal network takes the output of the feature pyramid network as input and generates category-agnostic region proposal boxes through further feature calculation, including foreground-background classification prediction and regression coordinate prediction.
[0066] The RoI pooling layer is used to receive the region proposal boxes generated by the region proposal network and perform RoI pooling operations on the features obtained by the backbone network to obtain RoI features.
[0067] In some embodiments, during the process of adjusting the transfer learning model using the annotated small sample data and jointly adjusting the region proposal network and the dual-path detection head of the adjusted transfer learning model, the weights of the backbone network are fixed.
[0068] In some embodiments, when verifying the detection results of the unannotated small sample data, the detection accuracies of both the base class samples and the new class samples are detected, and mAP is used as the detection index.
[0069] In the embodiments of the present application, before setting up a dual-path detection head in the basic detection model to obtain a transfer learning model, it further includes:
[0070] Determine the learning rate and the training length, and use the base class samples to train the basic detection model until the learning rate and the training length converge.
[0071] In the embodiments of the present application, adjusting the transfer learning model using the annotated small sample data includes:
[0072] Use the base class samples to train the basic branch and the dynamic branch of the dual-path detection head, thereby initializing the network parameters of the basic branch and the dynamic branch;
[0073] Update the initialized network parameters of the dynamic branch using the labeled small sample data to obtain the first network parameters.
[0074] In the embodiments of the present application, input the unlabeled small sample data into the adjusted transfer learning model, and jointly adjust the region proposal network and the dual-path detection head of the adjusted transfer learning model, including:
[0075] Input the unlabeled small sample data into the adjusted transfer learning model to determine the detection score of the unlabeled small sample data;
[0076] Determine a mixed data set based on the detection score of the unlabeled small sample data and the labeled small sample data;
[0077] Use the mixed data set to jointly adjust the region proposal network of the adjusted transfer learning model and the first network parameters of the dynamic branch, and fix the initialized network parameters of the base branch to obtain the final transfer learning model.
[0078] In the embodiments of the present application, input the unlabeled small sample data into the adjusted transfer learning model to determine the detection score of the unlabeled small sample data, including:
[0079] The adjusted transfer learning model converts the unlabeled small sample data into region of interest (RoI) features, and inputs the flattened RoI features into the dual-path detection head;
[0080] The base branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain the base class score;
[0081] The dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain the new class score;
[0082] Determine the detection score of the unlabeled small sample data according to the base class score and the new class score.
[0083] In the embodiments of the present application, the base branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain the base class score; the dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain the new class score, including:
[0084] The base branch and the dynamic branch perform feature transformation on the flattened RoI features according to the regression sub-task and the classification sub-task;
[0085] The regression sub-task is used to perform regression on the coordinate offset of the expanded RoI features;
[0086] The classification sub-task is used to balance the mean-variance difference between the base class features and the new class features in the expanded RoI features.
[0087] In some embodiments, the structure of the dual-path detection head is as Figure 2 shown. Among them, the dual-path detection head includes two parallel detection branches, namely the base branch and the dynamic branch. The dual-path detection head decouples the RoI features of the base class samples and the new class samples through learning different tasks on the two detection branches. The base branch and the dynamic branch perform feature transformation through the fully connected layer (FC) of the regression sub-task and the fully connected layer of the classification sub-task to obtain 1024-dimensional ROI features. The base branch and the dynamic branch respectively obtain the base class score and the new class score based on the 1024-dimensional ROI features.
[0088] According to some embodiments, the regression sub-task can be, for example, a standard fully connected network. The classification sub-task can be, for example, a cosine classifier, so as to improve the accuracy of small sample object detection.
[0089] In some embodiments, during the small sample object detection process of the base branch, the initialized network parameters of the base branch are fixed, and the prior knowledge learned in the base class samples is retained to prevent the performance degradation of the base class caused by the learning of new class samples during the knowledge transfer process.
[0090] In some embodiments, during the small sample object detection process of the dynamic branch, the initialized network parameters of the dynamic branch are updated using the base class small sample data and the new class small sample data, so as to realize the learning of new class samples.
[0091] In the embodiments of the present application, determining the detection score of unlabeled small sample data according to the base class score and the new class score includes:
[0092] Regularizing and constraining the base class sub-score in the new class score using the base class score to obtain the constrained new class score;
[0093] Determining the detection score of unlabeled small sample data according to the base class score and the constrained new class score.
[0094] It should be noted that, in order to make full use of the prior knowledge learned in the base class samples to guide the learning of new class samples, the dynamic branch implements base class consistency regularization during the process of updating network parameters using small sample data. The base class consistency loss is determined according to the following formula:
[0095]
[0096] Among them, is the base class consistency loss, uses the KL divergence, represents the base class sub-score of the c-th category in the new class score obtained by the dynamic branch, Represents the base class score of the c-th category in the base branch.
[0097] In the embodiments of the present application, a mixed dataset is determined based on the detection scores of unlabeled small sample data and labeled small sample data, including:
[0098] Sort the detection scores of the unlabeled small sample data according to the confidence level, and determine the pseudo-labeled small sample data according to the detection scores of the top preset number;
[0099] Mix the pseudo-labeled small sample data and the labeled small sample data in a preset ratio to obtain a mixed dataset.
[0100] It is easy to understand that when jointly adjusting the region proposal network and the dual-path detection head of the adjusted transfer learning model, semi-supervised proposal region mining is adopted to further improve the detection accuracy of the transfer learning model on new class samples. The process of semi-supervised proposal region mining is as Figure 3 shown.
[0101] In some embodiments, if the confidence level of the detection score of the unlabeled small sample data is lower than the preset threshold, the weight of the pseudo-labeled small sample data is set to a low weight. Thus, the influence of the unlabeled small sample data during the parameter update of the transfer learning model in the small sample object detection process is weakened.
[0102] In the embodiments of the present application, it further includes: setting a RoI feature discriminator parallel to the dual-path detection head in the base detection model;
[0103] The transfer learning model converts the input small sample data into RoI features, and inputs the flattened RoI features into the RoI feature discriminator;
[0104] The RoI feature discriminator classifies the base class features and new class features in the flattened RoI features by using at least two fully connected layers.
[0105] According to some embodiments, a RoI feature discriminator is set parallel to the dual-path detection head in the base detection model, and the RoI features are input into the dual-path detection head and the RoI feature discriminator at the same time to output the detection result, as Figure 4 shown.
[0106] In some embodiments, when the RoI feature discriminator classifies the base class features and new class features in the flattened RoI features, the RoI feature discriminator loss is determined according to the following formula:
[0107]
[0108] Where N i,j represents the set of categories to which the j-th region proposal in the i-th training sample belongs, Ni,j = 0 is the base class sample, N i,j = 1 is the new class sample, p i,j represents the output of the j-th region proposal in the i-th training sample passing through the RoI feature discriminator.
[0109] Finally, the overall loss of the transfer learning model is determined according to the following formula:
[0110]
[0111] where is the classification loss of Faster R-CNN, is the regression loss of Faster R-CNN, is the base class consistency loss, is the RoI feature discrimination loss. λ 1 and λ 2 are the weighting coefficients of the base class consistency loss and the RoI feature discrimination loss, respectively.
[0112] In summary, the method proposed in the embodiments of the present application obtains a transfer learning model by setting a dual-path detection head in the basic detection model; uses small sample data with annotations to adjust the transfer learning model to obtain an adjusted transfer learning model; inputs the small sample data without annotations into the adjusted transfer learning model, and jointly adjusts the region proposal network and the dual-path detection head of the adjusted transfer learning model, thereby completing the detection of the small sample data without annotations. The present application retains the performance of the basic category through the dual-path detection head and realizes learning between different tasks; obtains a more generalized region proposal result for the new category by fine-tuning the region proposal network; improves the detection efficiency and detection effect of small sample object detection by jointly adjusting the region proposal network and the dual-path detection head during the process of small sample object detection. By using the RoI feature discriminator to add a new information flow, the discrimination degree between the base class features and the new class features in the transfer learning model is improved, and the task learning of the base class samples and the new class samples is further decoupled.
[0113] To implement the above embodiments, the present application also proposes a small sample object detection device based on feature decoupling and knowledge transfer.
[0114] Figure 5 is a schematic structural diagram of a small sample object detection device based on feature decoupling and knowledge transfer provided by an embodiment of the present application.
[0115] As Figure 5 shown, a small sample object detection device based on feature decoupling and knowledge transfer includes:
[0116] A model determination module 510, configured to obtain a transfer learning model by setting a dual-path detection head in the basic detection model;
[0117] A model adjustment module 520, configured to adjust a transfer learning model by using a small sample of labeled data to obtain an adjusted transfer learning model;
[0118] A sample detection module 530, configured to input an unlabeled small sample of data into the adjusted transfer learning model, and jointly adjust the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unlabeled small sample of data.
[0119] In summary, for the device proposed in the embodiment of the present application, the model determination module sets a dual-path detection head in the basic detection model to obtain a transfer learning model; the model adjustment module adjusts the transfer learning model by using a small sample of labeled data to obtain an adjusted transfer learning model; the sample detection module inputs an unlabeled small sample of data into the adjusted transfer learning model, and jointly adjusts the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unlabeled small sample of data. The dual-path detection head of the present application retains the performance of the basic category and realizes learning between different tasks. By jointly adjusting the region proposal network and the dual-path detection head during the small sample object detection process, the detection efficiency and detection effect of the small sample object detection are improved.
[0120] It should be noted that in the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0121] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the technical field to which the embodiments of the present application belong.
[0122] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0123] Those of ordinary skill in the art can understand that all or part of the steps carried by the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0124] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0125] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0126] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0127] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A small-sample object detection method based on feature decoupling and knowledge transfer, characterized in that, the method includes: Setting a dual-path detection head in the basic detection model to obtain a transfer learning model; Using the small-sample data with annotations to adjust the transfer learning model to obtain an adjusted transfer learning model. Among them, the base class samples are used to train the basic branch and the dynamic branch of the dual-path detection head, so as to initialize the network parameters of the basic branch and the dynamic branch. Using the small-sample data with annotations, the initialized network parameters of the dynamic branch are updated to obtain the first network parameters; Inputting the unannotated small-sample data into the adjusted transfer learning model, and jointly adjusting the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unannotated small-sample data. Among them, the adjusted transfer learning model converts the unannotated small-sample data into Region of Interest (RoI) features, and inputs the flattened RoI features into the dual-path detection head. The basic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain base class scores, and the dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain new class scores. Using the base class scores to regularize and constrain the base class sub-scores in the new class scores to obtain the constrained new class scores. Determining the detection scores of the unannotated small-sample data according to the base class scores and the constrained new class scores. Determining a mixed data set based on the detection scores of the unannotated small-sample data and the small-sample data with annotations. Using the mixed data set to jointly adjust the region proposal network and the first network parameters of the dynamic branch of the adjusted transfer learning model, and fixing the initialized network parameters of the basic branch to obtain the final transfer learning model.
2. The method according to claim 1, characterized in that, before setting a dual-path detection head in the basic detection model to obtain a transfer learning model, it further includes: Determining the learning rate and the training length, and using the base class samples to train the basic detection model until the learning rate and the training length converge.
3. The method according to claim 1, characterized in that, the basic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain base class scores; the dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain new class scores, including: The basic branch and the dynamic branch perform feature transformation on the flattened RoI features according to the regression sub-task and the classification sub-task; The regression sub-task is used to perform regression on the coordinate offset of the flattened RoI features; The classification sub-task is used to balance the mean-variance difference between the base class features and the new class features in the flattened RoI features.
4. The method according to claim 2, characterized in that, determining the mixed data set based on the detection scores of the unannotated small-sample data and the small-sample data with annotations, includes: Sorting the detection scores of the unannotated small-sample data according to the confidence level, and determining the pseudo-labeled small-sample data according to the detection scores of the first preset number. Mix the pseudo-labeled small sample data and the labeled small sample data in a preset ratio to obtain a mixed data set.
5. The method according to claim 1, characterized in that, further comprising: setting a RoI feature discriminator parallel to the dual-path detection head in the basic detection model; the transfer learning model converts the input small sample data into RoI features and inputs the flattened RoI features into the RoI feature discriminator; the RoI feature discriminator classifies the base class features and the new class features in the flattened RoI features by using at least two fully connected layers.
6. A small sample object detection device based on feature decoupling and knowledge transfer, characterized in that, the device includes: a model determination module, configured to set a dual-path detection head in the basic detection model to obtain a transfer learning model; a model adjustment module, configured to adjust the transfer learning model by using the labeled small sample data to obtain an adjusted transfer learning model, wherein the base class samples are used to train the basic branch and the dynamic branch of the dual-path detection head, so as to initialize the network parameters of the basic branch and the dynamic branch, and the labeled small sample data is used to update the initialized network parameters of the dynamic branch to obtain the first network parameters; a sample detection module, configured to input the unlabeled small sample data into the adjusted transfer learning model, and jointly adjust the region proposal network and the dual-path detection head of the adjusted transfer learning model, so as to complete the detection of the unlabeled small sample data, wherein the adjusted transfer learning model converts the unlabeled small sample data into region of interest (RoI) features, and inputs the flattened RoI features into the dual-path detection head, the basic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain base class scores, the dynamic branch of the dual-path detection head performs feature transformation on the flattened RoI features to obtain new class scores, the base class scores are used to regularize and constrain the base class sub-scores in the new class scores to obtain the constrained new class scores, the detection scores of the unlabeled small sample data are determined according to the base class scores and the constrained new class scores, a mixed data set is determined based on the detection scores of the unlabeled small sample data and the labeled small sample data, and the region proposal network and the first network parameters of the dynamic branch of the adjusted transfer learning model are jointly adjusted by using the mixed data set, and the initialized network parameters of the basic branch are fixed to obtain the final transfer learning model.
Citation Information
Patent Citations
Small sample target detection method based on attention and contrast learning
CN113392855A