Small sample cross-domain target detection method based on adversarial learning
By introducing adversarial learning branches and feature measurement losses in small sample object detection tasks, the problems of data scarcity and feature distribution differences are solved, and the efficient cross-domain adaptation and generalization capabilities of the model are improved.
Patent Information
- Application Number
- CN202510059822.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-30
AI Technical Summary
In small sample object detection tasks, due to data scarcity, the model's adaptability and generalization ability in a specific domain are limited, and the difference in feature distribution between the source domain and the target domain makes it difficult for the model to achieve good performance on the target domain.
A small sample cross-domain object detection method based on adversarial learning is adopted, and the feature distribution alignment of the source domain and the target domain is enhanced by introducing adversarial learning branches and feature measurement losses on the benchmark model.
It significantly improves the model's adaptability among different fields, improves detection accuracy and efficiency, and enhances the model's adaptability and robustness to new environments.
Smart Images

Figure CN120070956A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of few-shot cross-domain object detection, and particularly relates to a few-shot cross-domain object detection method and device based on adversarial learning. Background Art
[0002] Object detection is a research hotspot that has received much attention in the fields of machine learning and computer vision in recent years. It usually relies on a large amount of data to train deep learning models. However, in some specific application scenarios, due to factors such as privacy protection, security requirements, and high annotation costs, it has become extremely difficult to obtain and annotate certain types of data, resulting in a lack of data, which poses challenges to the further development of object detection technology. Especially under the condition of limited training samples, traditional deep learning algorithms cannot be fully trained, making the deep neural network model prone to overfitting, seriously affecting the generalization ability of the model, and it is difficult to meet the actual needs only relying on existing deep learning technologies. To address this challenge, researchers have begun to explore techniques and methods specifically for few-shot object detection (FSOD).
[0003] Regarding the problem of few-shot object detection, scholars at home and abroad have proposed a series of methods, mainly relying on existing mature detection frameworks and advanced few-shot learning technologies to build detection models suitable for the case of scarce samples. In the early research stage, commonly used methods included semi-supervised learning and weakly-supervised learning, which mainly relied on a small amount of labeled data. Their core strategy was to collect additional, easily annotated data to alleviate the problem of difficult annotation in object detection. However, due to the lack of sufficient supervision of training images and complex model designs in these methods, the models were difficult to generalize to new categories with less labeled data, resulting in poor detection performance for new categories.
[0004] In recent years, significant breakthroughs have been made in the research of few-shot object detection. From the perspective of working principles, these methods can mainly be divided into four categories: meta-learning-based methods, transfer-learning-based methods, data-augmentation-based methods, and metric-learning-based methods. Nevertheless, the technology for solving the problem of few-shot object detection is still in the stage of academic exploration, and there is still a large gap compared with the object detection technology under large-scale datasets, facing some problems that need to be solved urgently, such as overfitting, large computational amount, and difficult model design. The challenges of few-shot object detection mainly stem from data scarcity, which limits the adaptability and generalization ability of the model in a specific domain. This problem has promoted the development of cross-domain few-shot object detection. However, in the task of cross-domain few-shot object detection, there are significant differences in the feature distributions between the source domain and the target domain, which makes it difficult for the model trained on the source domain to achieve good performance on the target domain. Summary of the Invention
[0005] In view of this, the present invention provides a few-shot cross-domain object detection method based on adversarial learning. Based on a benchmark model, this method introduces an adversarial learning strategy and a feature metric loss to effectively perform a few-shot object detection task, improving the generalization ability of the model, with high detection accuracy and efficiency.
[0006] To solve the above technical problems, the present invention is implemented as follows.
[0007] A few-shot cross-domain object detection method based on adversarial learning, comprising:
[0008] Step 1: Construct a training set using source domain data and some target domain data;
[0009] Step 2: Input the training set into an object detection network model; the object detection network model realizes the alignment of the feature distributions of the source domain and the target domain by introducing an adversarial learning branch on the benchmark network; the improved object detection network model uses a feature extractor to extract features, and inputs them into the detection branch and the adversarial learning branch at the same time; the detection branch outputs the object detection result, and the adversarial learning branch outputs the domain class retrieval result of whether the sample belongs to the source domain or the target domain;
[0010] Step 3: Calculate the total loss of the object detection network model using the total loss function, and perform end-to-end training on the object detection network model using the total loss; wherein, the total loss function adds a cross-domain loss, including the domain classification loss of the adversarial learning branch;
[0011] Step 4: Perform object detection using the trained object detection network model.
[0012] Preferably, the cross-domain loss further includes: the feature metric loss of the feature extractor; the design goal of this feature metric loss is to minimize the difference in the feature distributions between the source domain and the target domain.
[0013] Preferably, the cross-domain loss loss dom includes the feature metric loss of the feature extractor and the domain classification loss of the adversarial learning branch
[0014]
[0015] where λ 3 、λ 4 are weight coefficients, N is the number of samples, d i is the domain class predicted by the adversarial learning branch, d i * is the domain class label; sample g i and g jThe output feature map of the last layer after being processed by the feature extractor, where i and j represent the i-th and j-th among the N samples input to the feature extractor respectively;
[0016] Domain classification loss Use binary cross-entropy loss;
[0017] Feature metric loss The calculation method is:
[0018]
[0019] Among them, y ij is the distribution gap label of the sample pair g i and g j If the domain labels of the two samples are the same, the distribution gap label y ij is 1, otherwise the distribution gap label y ij is 0; ||f i -f j || 2 is the Euclidean distance of the sample pair in the feature space, and m is a hyperparameter that controls the interval of the same-domain sample pairs.
[0020] Preferably, the total loss function further includes a detector loss loss dec ; The detector loss loss dec includes an object classification loss and a regression loss
[0021]
[0022] Among them, λ 1 、λ 2 are weight coefficients, N is the number of samples, c i is the class score calculated for the class predicted by the detection branch, t i is the predicted box parameter of the prediction branch, is the true box parameter; is the object score;
[0023] Object classification loss Use binary cross-entropy loss;
[0024] Regression loss includes Distribution Focal Loss DFL and Complete Intersection over Union Loss CIoU Loss.
[0025] Preferably, the reference network is based on YOLOv8 and uses MobileNetv3 as the feature extractor.
[0026] Preferably, the adversarial learning branch consists of a Gradient Reversal Layer (GRL) and a discriminator; the output of the discriminator is the domain category.
[0027] Preferably, the part of the target domain data that is not used as training samples is used as the test set; the trained object detection network model is used to detect the targets in the test set, and the performance of the model in actual applications is evaluated.
[0028] The present invention also provides a small-sample cross-domain object detection device based on adversarial learning, which includes a training set module, an object detection network model, and a training module;
[0029] The training set module is used to store the training set, which is constructed from source domain data and part of the target domain data;
[0030] The object detection network model aligns the feature distributions of the source domain and the target domain by introducing an adversarial learning strategy on the benchmark network; the improved object detection network model includes a feature extractor, a detection branch, and an adversarial learning branch; the features extracted by the feature extractor are input into the detection branch and the adversarial learning branch; the detection branch outputs the object detection results; the adversarial learning branch outputs the domain category detection results indicating whether the sample belongs to the source domain or the target domain; the trained object detection network model performs object detection;
[0031] The training module is used to input the training set into the object detection network model during training, calculate the total loss using the total loss function, and perform end-to-end training on the object detection network model using the total loss;
[0032] Among them, the total loss function incorporates a cross-domain loss; the cross-domain loss includes the domain classification loss of the adversarial learning branch and the feature metric loss of the feature extractor; the design goal of the feature metric loss is to minimize the difference in the feature distributions between the source domain and the target domain.
[0033] Preferably, the training module includes a loss calculation module and a network optimization module;
[0034] The loss calculation module is used to calculate the total loss of the network model; the network optimization module optimizes the network parameters according to the total loss;
[0035] The loss calculation module includes a detector loss calculation unit, a cross-domain loss calculation unit, and a fusion unit;
[0036] The detector loss calculation unit is used to calculate the detector loss loss dec , including the object classification loss and the regression loss and then perform weighting;
[0037] The cross-domain loss calculation unit is used to calculate the cross-domain loss, including the feature metric loss Domain classification loss Then perform weighting; the calculation formula is:
[0038]
[0039] Among them, λ 3 , λ 4 are weight coefficients, N is the number of samples, d i is the domain category predicted by the adversarial learning branch, d i * is the domain category label, f i and f j are the output feature maps of the last layer after the samples g i and g j are processed by the feature extractor. i and j respectively represent the i-th and j-th of the N samples input to the feature extractor;
[0040] Domain classification loss Adopts binary cross-entropy loss;
[0041] Feature metric loss The calculation method is:
[0042]
[0043] Among them, y ij is the distribution gap label of the sample pair g i and g j If the domain labels of the two samples are the same, the distribution gap label y ij is 1, otherwise the distribution gap label y ij is 0; ||f i -f j || 2 is the Euclidean distance of the sample pair in the feature space, and m is a hyperparameter that controls the interval of the same-domain sample pair;
[0044] Fusion unit, used to fuse the detector loss and the cross-domain loss into the total loss.
[0045] Preferably, the benchmark network is based on YOLOv8 and uses MobileNetv3 as the feature extractor; the adversarial learning branch consists of a gradient reversal layer GRL and a discriminator.
[0046] Beneficial effects:
[0047] (1) This paper introduces an adversarial learning branch specifically for predicting the domain category of samples. This design significantly improves the adaptability of the model across different domains. By predicting the domain category during training, the model is forced to learn feature representations that are critical for cross-domain tasks. This learning mechanism effectively narrows the distribution difference between the source domain and the target domain, enhances the model's adaptability to new environments, and thus achieves higher performance and stronger robustness in the target domain.
[0048] (2) The feature metric loss function proposed in the present invention deepens the model's understanding of sample similarities and differences by calculating the distance between sample pairs in the feature space. It brings together the feature representations of samples from different domains but belonging to the same category, which helps the model capture a wider range of generalized features. At the same time, it also ensures the compactness of sample features within the same domain by setting a preset boundary to control the maximum allowable difference between features of samples of the same category. This fine balancing mechanism not only improves the accuracy of the model in identifying similar samples, but also enhances the model's sensitivity to subtle differences between different samples. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flow chart of a small sample cross-domain target detection method based on adversarial learning of the present invention;
[0050] Figure 2 This is a schematic diagram of the network structure of small sample cross-domain target detection based on adversarial learning in the present invention;
[0051] Figure 3 Schematic diagram of source domain and target domain data feature distribution model before and after processing;
[0052] Figure 4 Comparison of AP50 (%) between the present invention and other detection models on the Cityscapes dataset;
[0053] Figure 5 The detection results of the model of the present invention on the target domain Cityscapes dataset;
[0054] Figure 6 It is a block diagram of the composition of the small sample cross-domain target detection device based on adversarial learning of the present invention; DETAILED DESCRIPTION
[0055] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0056] The present invention provides a few-shot cross-domain object detection method based on adversarial learning. The main idea is as follows: Using the adversarial learning strategy to predict the image domain categories, align the feature distributions of the source domain and the target domain, and design the feature metric loss for the feature representations of the same category among samples. Utilize the source domain data to provide the detection rate of the target domain and accelerate the model convergence.
[0057] The few-shot cross-domain object detection method of the present invention includes dataset construction, cross-domain object detection network model, loss function, and model training and inference. As Figure 1 shown, the specific steps are as follows:
[0058] Step 1, data division of the source domain and the target domain.
[0059] To construct a network model that can effectively transfer and adapt to the target domain, first, the datasets in different domains are respectively designated as the source domain dataset D s and the target domain dataset D T . Both contain the same target categories, but the source domain dataset has richer annotated samples. Then, randomly divide the target domain data according to a predetermined ratio, merge a part of it with all the data in the source domain to form a comprehensive training set for training the network model to learn general and robust feature representations. The remaining target domain data is used as the test set to evaluate the generalization ability of the model in the target domain.
[0060] Step 2, construct the object detection network model.
[0061] In this embodiment, YOLOv8 is used as the basic framework, and MobileNetv3 is used as the feature extractor to form a benchmark network. An adversarial learning strategy is introduced into the benchmark network, that is, an adversarial learning branch is added after the feature extractor to align the feature distributions of the source domain and the target domain. The schematic diagram of the network structure is as Figure 2 shown. The feature map output by the feature extractor will enter the detection branch and the adversarial learning branch simultaneously.
[0062] Among them, the detection branch consists of a Path Aggregation Network (PAN) and a detector.
[0063] The adversarial learning branch in the model consists of a Gradient Reversal Layer (GRL) and a discriminator. Among them, GRL is similar to a linear layer with a weight of 1 during the forward inference process and does not change the data distribution, but it will reverse the sign of the gradient during the backpropagation process. This operation helps to achieve the goal of adversarial training. The discriminator is composed of multiple convolutional layers and an output layer sigmoid. Its output is the domain category, that is, to predict whether the input feature belongs to the source domain or the target domain, which is a binary classification situation.
[0064] Step 3: Construct a loss function.
[0065] The loss function of the present invention includes a detector loss and a cross - domain loss. Among them, the detector branch includes a regression loss and an object classification loss, and the cross - domain loss includes a domain classification loss of the adversarial learning branch and a feature metric loss of the feature extractor.
[0066] Furthermore, the loss function is calculated by summing the detector loss and the cross - domain loss, and the loss Loss is obtained using the following formula:
[0067] Loss = loss dec + loss dom
[0068] (1) loss dec is the detector loss, including the object classification loss and the regression loss, as follows:
[0069]
[0070] where λ 1 , λ 2 are weight coefficients, N is the number of samples input into the network model at one time, c i is the class score obtained by calculating the class predicted by the detection branch, t i is the predicted box parameter of the prediction branch (cx i , cy i , w i , h i ), which are the coordinates, length, and width of the predicted box respectively, is the ground - truth box parameter is the object score, which is calculated from the intersection - over - union of the predicted box and the ground - truth box for positive samples, and is 0 for negative samples.
[0071] (1.1) is the object classification loss, using Binary CrossEntropy Loss (BCE Loss), as follows:
[0072]
[0073] (1.2) is the regression loss, including Distribution Focal Loss (DFL) and Complete Intersection over Union Loss (CIoU Loss), as follows:
[0074]
[0075] is the DFL loss, that is, the DFL of the rectangular box regression. The predicted value of the model obtains the prediction parameter l after the operation S(·) of the softmax layer of the detection branch i , r i , t i , b i 's discrete probability distribution, where is the probability relative to the left and right nearest integer boundaries of the true parameter ; l i , r i , t i , b i and are the distances from the center points of the predicted box and the true box to the left, right, top, and bottom boundaries of the box respectively. The DFL is obtained using the following formula:
[0076]
[0077] where y i is the true value of the parameter of the true box, and y i takes values of l i , r i , t i , b i ; and are obtained by rounding down and up from y i , representing the left and right integer boundaries closest to y i .
[0078] loss ciou (t i , t i * ) is the complete intersection over union loss, that is, the rectangular box regression loss. While considering the overlap rate of the predicted box and the true box, the Euclidean distance of the center point and the aspect ratio are introduced as follows:
[0079]
[0080] where b i and b i * represent the center points of the predicted box and the true box respectively, ρ 2 (b i , b i * ) is the Euclidean distance between the center points of the two predicted boxes and the true box, s i is the diagonal length of the smallest rectangle containing the predicted box and the true box, and IoU i is the overlap rate between the predicted box and the true box.
[0081] (2) loss dom is a cross - domain loss, including domain classification loss and feature metric loss, and the calculation formula is as follows:
[0082]
[0083] where λ 3 , λ 4 are weight coefficients, N is the number of samples, d i is the domain category predicted by the adversarial learning branch, d i * is the domain category label; the output feature maps of the last layer after the samples g i and g j are processed by the feature extractor, and i and j respectively represent the i - th and j - th of the N samples input to the feature extractor.
[0084] (2.1) is the domain classification loss, using binary cross - entropy loss, as shown below:
[0085]
[0086] (2.2) is the feature metric loss, and the calculation formula is as follows:
[0087]
[0088] where, y ij is the distribution gap label of the distribution gap between the sample pair g i and g j . If the domain labels of the two samples are the same, the distribution gap label y ij is 1, otherwise the distribution gap label y ij is 0. ||f i -f j || 2 is the Euclidean distance of the sample pair in the feature space, and m is a hyperparameter that controls the interval of the same - domain sample pair. The design goal of this feature metric loss is to minimize the distribution gap between the source domain and target domain features, reduce the feature distribution difference; and make the training based on source - domain data and the training based on target - domain data fused, enhancing the generalization ability of the network.
[0089] Step four, model training and validation.
[0090] In this embodiment, the publicly available datasets KITTI and Cityscapes are used as the source domain dataset and the target domain dataset respectively, and only the car category is considered. The target domain dataset Cityscapes has a total of 2,860 samples. Among them, 2,382 data samples are combined with 7,481 samples of the source domain dataset KITTI to form the training set, and the remaining 478 data samples in the target domain are used as the test set. The source domain data in the training set can provide sufficient information to help the model learn, while the target domain data ensures that the model can adapt to a specific application environment.
[0091] The training set data is input into the network model, the results output by the model are decoded, and then backpropagation is performed using the loss function to update the model parameters. After the model converges, a model weight file is obtained. Through the adversarial learning strategy, a model is trained that can not only learn general features from the source domain but also be fine-tuned through the target domain data to improve its adaptability and accuracy in the target domain. This helps to reduce the distribution difference between the source domain and the target domain, make the features extracted by the feature extractor conform to the same distribution, and enhance the generalization ability of the model, as Figure 3 shown. Finally, the trained model is used to detect the targets in the test set and evaluate the performance of the model in practical applications.
[0092] Step Five: Use the trained object detection network model for object detection. The detection branch outputs the object detection results.
[0093] In this embodiment, the intersection over union (IoU) of 50% is used as the threshold, and the average precision (AP), that is, AP50, is used as the evaluation metric. Compared with some other competitive few-shot cross-domain object detection models, the model proposed in the present invention performs excellently in terms of performance, as Figure 4 shown. The model weight file generated by model training is used for forward inference on the target domain images, and the detection results are as Figure 5 shown, where Ours is the model proposed in the present invention. It can be seen that the method proposed in the present invention can almost detect all targets, can efficiently identify and locate the vast majority of targets. Even if the targets are obscured by the background, the method of the present invention can still accurately identify and locate the targets. The method proposed in the present invention still maintains a high detection rate for complex scenes, demonstrating excellent robustness.
[0094] The present invention further provides a few-shot cross-domain object detection device based on adversarial learning, as Figure 6 shown. The device includes a training set module, an object detection network model, and a training module;
[0095] The training set module is used to store the training set, which is constructed from source domain data and partial target domain data;
[0096] The object detection network model aligns the feature distributions of the source domain and the target domain by introducing an adversarial learning strategy on the benchmark network; the improved object detection network model includes a feature extractor, a detection branch, and an adversarial learning branch; the features extracted by the feature extractor are input into the detection branch and the adversarial learning branch; the detection branch outputs the object detection results; the adversarial learning branch outputs the domain category detection results indicating whether the sample belongs to the source domain or the target domain; the trained object detection network model performs object detection;
[0097] The training module is used to input the training set into the object detection network model during training, calculate the total loss function, and perform end-to-end training on the object detection network model using the total loss function;
[0098] Among them, the total loss function includes a detector loss and a cross-domain loss; the detector loss includes a regression loss and an object classification loss of the detection branch; the cross-domain loss includes a feature metric loss of the feature extractor and a domain classification loss of the adversarial learning branch.
[0099] In a preferred embodiment, the benchmark network is based on YOLOv8 and uses MobileNetv3 as the feature extractor; the adversarial learning branch consists of a gradient reversal layer GRL and a discriminator.
[0100] In a preferred embodiment, the training module includes a loss calculation module and a network optimization module;
[0101] The loss calculation module is used to calculate the total loss of the network model; the network optimization module optimizes the network parameters according to the total loss;
[0102] The loss calculation module includes a detector loss calculation unit, a cross-domain loss calculation unit, and a fusion unit;
[0103] The detector loss calculation unit is used to calculate the detector loss loss dec , including the object classification loss and the regression loss and then performs weighting;
[0104] The cross-domain loss calculation unit is used to calculate the cross-domain loss, including the feature metric loss and the domain classification loss and then performs weighting; the calculation formula is:
[0105]
[0106] Among them, λ 3 、λ4 is the weight coefficient, N is the number of samples, and d i is the domain category predicted by the adversarial learning branch, is the domain category label, and f i and f j are the output feature maps of the last layer after the samples g i and g j are processed by the feature extractor. i and j respectively represent the i-th and j-th of the N samples input to the feature extractor;
[0107] is the domain classification loss, and the binary cross-entropy loss is adopted;
[0108] is the feature metric loss:
[0109]
[0110] where y ij is the distribution gap label of the sample pair g i and g j . If the domain labels of the two samples are the same, the distribution gap label y ij is 1, otherwise the distribution gap label y ij is 0; ||f i -f j || 2 is the Euclidean distance of the sample pair in the feature space, and m is the hyperparameter that controls the interval of the same-domain sample pairs;
[0111] The fusion unit is used to fuse the detector loss and the cross-domain loss into the total loss.
[0112] In summary, the above is only the preferred embodiment of the present invention and is not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A small sample cross-domain target detection method based on adversarial learning, characterized in that: include: Step 1: Use source domain data and part of target domain data to build a training set; Step 2: Input the training set into the target detection network model; the target detection network model realizes the feature distribution alignment of the source domain and the target domain by introducing an adversarial learning branch on the baseline network; The improved target detection network model uses a feature extractor to extract features and inputs them into the detection branch and the adversarial learning branch at the same time; The detection branch outputs the target detection result, and the adversarial learning branch outputs the domain category retrieval result of whether the sample belongs to the source domain or the target domain; Step 3: Calculate the total loss of the target detection network model using the total loss function, and perform end-to-end training on the target detection network model using the total loss; wherein the total loss function incorporates cross-domain loss, including domain classification loss of the adversarial learning branch; Step 4: Use the trained target detection network model to perform target detection.
2. The method according to claim 1, characterized in that The cross-domain loss further includes: a feature metric loss of a feature extractor; the design goal of the feature metric loss is to minimize the difference in feature distribution between the source domain and the target domain.
3. The method according to claim 2, characterized in that The cross-domain loss dom Include feature metric loss for feature extractor and domain classification loss for the adversarial learning branch Among them, λ3 and λ4 are weight coefficients, N is the number of samples, and d i To learn branch prediction domains adversarially, is the domain category label; sample g i and g j The output feature map of the last layer after being processed by the feature extractor, i and j represent the i-th and j-th samples of the N samples input to the feature extractor respectively; Domain Classification Loss Adopt binary cross entropy loss; Feature metric loss The calculation method is: Among them, y ij is the sample pair g i and g j If the domain labels of two samples are the same, then the distribution gap label y ij is 1, otherwise the distribution gap label y ij is 0; ||f i -f j ||2 is the Euclidean distance of sample pairs in the feature space, and m is a hyperparameter that controls the interval between pairs of samples in the same domain.
4. The method according to any one of claims 1 to 3, characterized in that The total loss function further includes the detector loss dec ; Detector loss loss dec Including target classification loss And the regression loss Among them, λ1 and λ2 are weight coefficients, N is the number of samples, c i is the category score calculated for the category predicted by the detection branch, t i is the prediction box parameter of the prediction branch, is the real frame parameter; is the target score; Object classification loss Adopt binary cross entropy loss; Regression Loss Including distribution focus loss DFL and complete intersection-over-union loss CIoU Loss.
5. The method according to claim 1, characterized in that The benchmark network uses YOLOv8 as the basic framework and MobileNetv3 as the feature extractor.
6. The method according to claim 1 or 5, characterized in that The adversarial learning branch consists of a gradient reversal layer GRL and a discriminator; the output of the discriminator is the domain category.
7. The method according to claim 1, characterized in that The target domain data that is not used as training samples is used as the test set. The trained target detection network model is used to detect the test set targets and evaluate the performance of the model in practical applications.