An Unsupervised Evaluation Method for Image Recognition Domain Adaptation

The ACM method addresses the challenge of evaluating UDA models without labeled validation sets by integrating source domain accuracy, data augmentation consistency, and classifier diversity, resulting in improved image classification performance and hyperparameter selection.

CN116883744BActive Publication Date: 2025-07-15ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310857344.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-13
Publication Date
2025-07-15
Estimated Expiration
2043-07-13

AI Technical Summary

Technical Problem

Existing unsupervised domain adaptation techniques often fail to select appropriate models and hyperparameters when evaluating the accuracy of image recognition in target domains, resulting in poor image recognition results, especially in the absence of labeled verification sets.

Method used

An unsupervised evaluation method is proposed to form an evaluation index ACM by combining the accuracy of the source domain, enhanced consistency after data augmentation and diversity terms of the classifier to evaluate the migration effect of the model, thereby selecting the optimal hyperparameters and model.

Benefits of technology

It can effectively select the optimal hyperparameters and unsupervised domain adaptation algorithm without the need for target verification set annotation, improve the accuracy of image recognition, surpass the results of manual debugging, and fight against deliberately designed training method attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883744B_ABST
    Figure CN116883744B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised evaluation method for image recognition domain adaptation, including: (1) learning a number of classification models using the training sets of the source domain and the target domain; (2) obtaining model predictions using the source domain through the classification model M and comparing them with the labels of the source domain validation set to obtain the accuracy A of the source domain S ; (3) retaining intermediate features when the source domain data passes through the classification model M, and training a classifier h using the intermediate features and the corresponding labels; (4) performing data augmentation on the validation set of the target domain, passing the data before and after augmentation through the classification model M to obtain corresponding intermediate features, and then passing them through the classifier h to obtain corresponding predictions, obtaining the augmentation consistency AC; (5) combining A S and AC, and adding a diversity term of the classifier h to obtain an evaluation metric ACM to evaluate the transfer effect of the classification model M. The present invention can evaluate the transfer effect of the model unsupervised, so as to select the hyperparameters and models with the best image recognition effect, thereby improving the effect of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image recognition, and particularly relates to an unsupervised evaluation method for image recognition domain adaptation. Background Art

[0002] Nowadays, image recognition algorithms are widely used, and their success mainly stems from the use of deep neural networks. In many practical scenarios, manually annotating sufficient images for training the deep neural network of image recognition is both expensive and time-consuming. To solve this problem and improve the accuracy of image recognition simultaneously, previous scholars introduced unsupervised domain adaptation (UDA) technology. This technology transfers the image recognition model trained on the source domain with class labels to the target domain without class labels. In recent years, many UDA methods have been proposed to address the situation where the source domain images and the target domain images are not similar during the transfer process, that is, the domain shift problem. Although these methods have improved the image recognition accuracy of the target domain, this improvement usually requires tuning the parameters using the class-labeled validation set of the target domain. However, in practice, obtaining labeled validation set images may be expensive, and different image datasets usually require different sets of hyperparameters to achieve ideal image recognition performance.

[0003] The prior art has studied evaluating the accuracy of image recognition in the target domain without a labeled validation set. For example, in the article "Domain-adversarial training of neural networks" published in the JMLR journal in 2016, the domain distribution distance was used as an evaluation metric to select the hyperparameters of their training method. However, this evaluation metric is tightly coupled with the training method. The article "Towards accurate model selection in deep unsupervised domain adaptation" published in the ICML conference in 2019 first proposed a general UDA evaluation metric, which uses the importance-weighted validation method and adds a variance control term. Subsequently, the article "Tune it the right way: Unsupervised validation of domain adaptation via soft neighborhood density" published in the ICCV conference in 2021 proposed that a good transfer model should have a compact neighborhood for each target feature and introduced the soft neighborhood density metric. However, after more comprehensive and detailed experiments on UDA evaluation metrics, it was found that previous evaluation metrics often failed to select the appropriate model in most cases, and thus could not improve the image recognition effect in the target domain. This is because the assumptions on which their metrics are based do not always hold in a wide range of scenarios. Summary of the Invention

[0004] The present invention provides an unsupervised evaluation method for image recognition domain adaptation, enabling the unsupervised evaluation of the transfer effect of a model during unsupervised domain adaptation, thereby selecting the hyperparameters and model with the best image recognition effect, and thus improving the effect of image classification.

[0005] An unsupervised evaluation method for image recognition domain adaptation, characterized by comprising the following steps:

[0006] (1) According to the existing unsupervised domain adaptation algorithm, use the image training sets of the source domain and the target domain for learning to obtain several classification models, and each classification model M will be evaluated in the subsequent steps;

[0007] (2) Use the image validation set of the source domain to obtain the predictions of the model through the classification model M, and compare them with the labels of the source domain validation set to obtain the accuracy rate A of the source domain S ;

[0008] (3) When the source domain data passes through the classification model M, retain the intermediate features, and use the intermediate features and the corresponding labels to train a multi-layer fully connected classifier h;

[0009] (4) The image validation set of the target domain is subjected to data augmentation to obtain the image data after data augmentation. Both the data before and after data augmentation pass through the classification model M to obtain the corresponding intermediate features, and then pass through the classifier h in step (3) to obtain the corresponding predictions. Compare these two predictions to obtain the augmentation consistency AC;

[0010] (5) Combine the accuracy rate A of the source domain S and the augmentation consistency AC, and add the diversity term of the classifier h to obtain the final evaluation index ACM, which evaluates the transfer effect of the classification model M;

[0011] (6) According to the evaluation index ACM corresponding to each model, select the model with the best transfer effect among several classification models, and use this classification model for image classification.

[0012] In step (1), the labeled source domain is: The unlabeled target domain is: represents the image of the i-th sample in the source domain, represents the label of the i-th sample, n s represents the total number of source domain samples, represents the image of the j-th sample in the target domain, n t represents the total number of target domain samples.

[0013] Existing domain adaptation algorithms include DANN, CDAN, MDD, and MCC. Several classification models are trained using an image training set, and each classification model consists of a feature extractor g and a linear classification layer f.

[0014] In step (2), the accuracy A of the source domain S The calculation formula is as follows:

[0015]

[0016] In the formula, denotes taking the mathematical expectation of the samples in the source domain validation set, denotes an image of a sample in the source domain validation set, denotes the label corresponding to this sample, p s denotes the data after passing through the model M, the prediction vector has K components, and K is the total number of categories; denotes p s the position of the largest component among the K components of ; I[·] represents the indicator function, which has a value of 1 if the content in the parentheses is true, otherwise 0.

[0017] In step (3), the intermediate features of the source domain data passing through the model M in step (2) are retained where g is the feature extractor of the model M; the intermediate features and labels are combined into an intermediate feature dataset: A classifier h is supervised and trained on this intermediate feature dataset, which consists of two fully connected layers.

[0018] In step (4), the image validation set data of the target domain is subjected to data augmentation to obtain the data after data augmentation Among them, the data augmentation used includes random cropping, random flipping, random color adjustment, and random blurring.

[0019] In step (4), the formula for obtaining the augmented consistency AC is:

[0020]

[0021] where q t , q t′ respectively denote the prediction vectors of the data and after passing through the feature extractor g and classifier h of the model M; denotes the position of the largest component among the components of q t ; denotes q t′The position of the largest component among the components of; I[·] represents the indicator function, which has a value of 1 if the content in the brackets is true, and 0 otherwise.

[0022] In step (5), the formula for calculating the evaluation metric ACM is:

[0023]

[0024] where the entropy function H(q) = ∑ k q k logq k , q t represents the predicted vector after the data passes through the feature extractor g and the classifier h of the model M, and K is the total number of categories.

[0025] Compared with the prior art, the present invention has the following beneficial effects:

[0026] 1. The unsupervised evaluation method proposed by the present invention can effectively take into account the source domain label information, and at the same time consider the diversity of categories, and can detect the situation of model collapse.

[0027] 2. The unsupervised evaluation method proposed by the present invention can resist the attacks of deliberately designed training methods and accurately reflect the transfer effect of the evaluated model in various situations.

[0028] 3. The present invention can search for the optimal hyperparameters and the optimal unsupervised domain adaptation algorithm (UDA) without the need for annotation of the target validation set, and the accuracy of its image recognition can exceed the results of previous manual debugging. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a schematic diagram for the unsupervised evaluation method of the present invention to select the optimal model hyperparameters;

[0030] Figure 2 is a schematic flow diagram of the unsupervised evaluation method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0031] The present invention will be further described in detail below with reference to the drawings and embodiments. It should be noted that the following embodiments are intended to facilitate the understanding of the present invention and do not limit it in any way.

[0032] As Figure 1 and Figure 2 shown, an unsupervised evaluation method for image recognition domain adaptation includes:

[0033] (a) The training dataset includes a labeled source domain: and an unlabeled target domain: Any off-the-shelf unsupervised domain adaptation algorithm (UDA), along with its hyperparameter search space, randomly or regularly selects a set of hyperparameters from the space during each training.

[0034] (b) Using the unsupervised domain adaptation algorithm (UDA) and a set of hyperparameters, train to obtain model M1. Then select another set of hyperparameters and train to obtain model M2. Repeat this process to obtain l trained models.

[0035] (c) Use the unsupervised evaluation method ACM proposed in the present invention to evaluate the models in turn and obtain the scores corresponding to these models.

[0036] (d) Select the model with the highest score, which is the model selected by the method of the present invention, and the corresponding hyperparameters are the optimal hyperparameters searched out.

[0037] The specific steps are elaborated as follows:

[0038] (a) The off-the-shelf unsupervised domain adaptation algorithms (UDA) considered in the present invention include source only, DANN, CDAN, MDD, and MCC, but are not limited to these 5 UDA training algorithms. There are different selection strategies according to the size of the hyperparameter space. For a sparse hyperparameter space, grid search is used, and each hyperparameter combination will be selected. For a dense hyperparameter space, the ACM evaluation score is set as the goal of hyperparameter search. The present invention uses the TPE parameter search algorithm to perform 50 searches and uses the median pruner of the Optuna library to accelerate the search. The specific details of the two hyperparameter spaces are shown in Table 1, and their meanings will be explained in detail in the subsequent experiments.

[0039] (b) Training algorithms: The present invention uses five popular UDA methods to obtain trained models. 1) Sourceonly: This model only receives supervised training on the source domain. 2) DANN: Train a domain discriminator, and the feature generator tries to deceive it. 3) CDAN: The domain discriminator takes the features and the predictions of the classifier as inputs. 4) MDD: Train an additional classifier to optimize the maximum mean discrepancy. 5) MCC: Aims to reduce the class confusion in the classifier predictions. The implementation of these methods and the selection of optimizers all follow the Transfer-Learning-Library code repository.

[0040] For the architecture of model M, the feature generator g contains a ResNet backbone pre-trained on the ImageNet dataset and a linear bottleneck layer, and the classifier f is a linear layer.

[0041] (c) As Figure 2As shown, for a model M to be evaluated, the ACM evaluation method of the present invention evaluates it from three aspects, and finally combines the three evaluation results to obtain an evaluation score. The validation set dataset consists of the labeled source domain validation set and the unlabeled target domain validation set .

[0042] The present invention studies the performance of unsupervised evaluation metrics on four popular image recognition UDA datasets, VisDA2017, DomainNet, OfficeHome, and Office31. The model to be evaluated is trained using five UDA methods and different hyperparameters. When evaluating these models, this paper examines whether the metrics are consistent with the image classification accuracy of the target domain.

[0043] VisDA-2017 is a large-scale dataset that poses challenges for the adaptation of unsupervised domains from simulation to reality. The dataset contains 152,397 synthetic images as the source domain and 55,388 real images as the target domain. These two domains share 12 object categories. The present invention evaluates all methods on the VisDA validation set.

[0044] DomainNet is a large-scale domain adaptation image dataset that contains 345 categories from six domains. Four of these domains were selected for the experiment: clipart (c), painting (p), real (r), and sketch (s). The present invention only studies the single-source domain configuration of DomainNet. There are 12 transfer tasks between these domains.

[0045] Office-31 is a commonly used dataset for unsupervised domain adaptation, which contains 4,652 images and 13 categories collected from the following three domains: Amazon (A), webcam (W), and DSLR (D). The present invention evaluates all methods in six domain adaptation tasks: A→W, D→W, W→D, A→D, D→A, and W→A.

[0046] Office-Home is a more difficult domain adaptation dataset than Office-31, which includes 15,500 images from four different domains: art images (Ar), clipart (Cl), product images (Pr), and real world (Rw). Each domain contains images of 65 object categories common in office and home scenes. The present invention evaluates the performance of all methods in 12 domain adaptation scenarios.

[0047] The present invention uses ResNet50 as the backbone for Office31 and OfficeHome, and ResNet101 as the backbone for VisDA and DomainNet. The present invention trains each model for 3000 steps on Office31, OfficeHome and VisDA, and a total of 6000 steps on DomainNet.

[0048] Hyperparameter set: The present invention finds that several hyperparameters are usually adjusted manually, and the present invention selects them to check the robustness of the metrics. In total, the present invention changes up to six hyperparameters of the training method: 1) Early stopping step: For the UDA problem, the model at the last step is usually not the best model during the training process. The present invention needs to evaluate the model regularly and select the best model during training. 2) Learning rate: The initial learning rate of the optimizer. 3) Weight decay: The weight decay of the optimizer. 4) Trade-off value: The trade-off between the supervised cross-entropy loss on the source domain and the objective loss of the UDA method. 5) Bottleneck dimension: The feature dimension output by the feature generator. 6) Hyperparameters related to the training method: The present invention selects the margin value γ for MDD and the temperature value T for MCC. For DANN and CDAN, the present invention adjusts the learning rate of the domain discriminator as a hyperparameter to balance the convergence of the discriminator and the generator. The present invention defines the learning rate ratio of D as the ratio of the discriminator learning rate to the generator learning rate.

[0049] Each time the model is trained, the present invention samples hyperparameters from its hyperparameter space. In the study of comparing various evaluation metrics, the present invention sets a rough hyperparameter space according to the default hyperparameters of the method. As shown in Table 1, the present invention uses the hyperparameters in the sparse hyperparameter space to train various algorithms, and then analyzes the consistency between each evaluation metric and the target accuracy rate of the trained result model. The present invention performs a grid search on the hyperparameter space of each algorithm and collects models during the search process. It should be noted that in order to obtain models with different early stopping steps, the present invention divides the total training steps into 10 rounds and evaluates the model at the end of each round.

[0050] Table 1

[0051]

[0052] For an evaluation method, given the trained model obtained during hyperparameter search Evaluation metric score should be consistent with the target classification accuracy To compare the effects of different evaluation methods. In the experiment, the present invention uses two measurement methods to measure the degree of consistency between the evaluation score and the target accuracy rate:

[0053] 1. Pearson correlation coefficient:

[0054]

[0055] where σ is the standard deviation.

[0056] 2. Bias of the best model:

[0057]

[0058] where l * = argmax l S l represents the best model according to the evaluation method. Metrics with higher correlation and lower bias are more consistent with the target error.

[0059] The consistency of each evaluation metric with the target accuracy rate on the four datasets is shown in Tables 2, 3, 4, and 5 respectively:

[0060] Table 2

[0061]

[0062] Table 3

[0063]

[0064] Table 4

[0065]

[0066] Tables 2, 3, and 4 respectively show the UDA metric results of five training methods on VisDA2017, DomainNet, and Office-home. The results show that it is difficult for previous metrics to represent the target accuracy in all training methods. Some metrics can perform well on the transfer tasks of one of the datasets, but do not perform well on all three datasets, which also indicates that testing on some datasets may lead to biased conclusions. It is worth noting that the ISM proposed in the present invention is consistent with the target accuracy of most training methods. The ACM of the present invention performs better on training methods that align the features of two domains (e.g., DANN and CDAN) because it can detect over-alignment problems.

[0067] Comparison between training methods: The present invention also studies the consistency between evaluation metrics and target accuracy when comparing different methods. Because in practice, the present invention needs to determine the best UDA method for the transfer task. The present invention collects all the models trained by five methods, their metric scores, and target accuracy. For each metric, we calculate the Pearson correlation and the deviation of the best model, and the results are shown in the "ALL" column. As shown in Tables 2, 3, and 4, when comparing all training methods, it becomes more difficult for most metrics to maintain consistency. Notably, the ISM and ACM of the present invention perform well on all three datasets, and the deviation ("dev") of the best model is less than 2%. Therefore, the present invention can use the proposed unsupervised metrics to determine the best training method for the dataset and its hyperparameters.

[0068] Hyperparameter search: Most UDA methods require manual adjustment of hyperparameters for different datasets. It would be ideal to automatically find suitable hyperparameters without supervision. In this section, the present invention shows that our ACM can be used for unsupervised hyperparameter search. We will perform unsupervised hyperparameter search for four algorithms: DANN, CDAN, MCC, and MDD. For each UDA training method, we first define its hyperparameter search space, as shown in the dense hyperparameter space in Table 1. Set ACM as the target of hyperparameter search. Simply use the TPE search algorithm for 50 trials and use Optuna's median pruner to accelerate the search. For each transfer task in the dataset, we report the target accuracy of the best model found by ACM. We compare this with the performance of the default hyperparameters of each method in TL-Lib. Tables 5, 6, 7, and 8 respectively show the target accuracies of the models found by the evaluation metrics of the present invention and the default models on the VisDA, DomainNet, Office-home, and Office-31 datasets. For all four training methods, the hyperparameters found by the present invention are better than those manually adjusted by TL-Lib. Different from previous supervised tuning, our search process does not require label information of the target domain.

[0069] Table 5

[0070]

[0071] Table 6

[0072]

[0073] Table 7

[0074]

[0075] Table 8

[0076]

[0077] The embodiments described above have elaborated in detail the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the present invention. Any modifications, supplements, and equivalent replacements made within the principle scope of the present invention should be included within the protection scope of the present invention.

Claims

1. An unsupervised evaluation method for image recognition domain adaptation, characterized in that, Including the following steps: (1) Use the image training sets of the source domain and the target domain to learn several classification models according to the existing unsupervised domain adaptation algorithms. Each classification model M will be evaluated in the subsequent steps; (2) Use the image validation set of the source domain to obtain the predictions of the model through the classification model M, and compare them with the labels of the source domain validation set to obtain the accuracy A of the source domain S ; (3) When the source domain data passes through the classification model M, the intermediate features are retained, and a multi-layer fully connected classifier h is trained using the intermediate features and the corresponding labels; (4) The image validation set of the target domain is augmented to obtain the augmented image data. Both the data before and after augmentation are passed through the classification model M to obtain the corresponding intermediate features, and then through the classifier h in step (3) to obtain the corresponding predictions. Comparing these two predictions to obtain the augmentation consistency AC; The formula for obtaining the augmentation consistency AC is: where q t and q t′ represent the prediction vectors after the feature extractor g and the classifier h of the model M for the data and respectively; denotes the position of the largest component among the components of q t ; denotes the position of the largest component among the components of q t′ ; I[·] represents the indicator function, which has a value of 1 if the content in the parentheses is true and 0 otherwise; (5) Combine the accuracy rate A of the source domain s and the enhanced consistency AC, and add the diversity term of the classifier h to obtain the final evaluation metric ACM, which evaluates the transfer effect of the classification model M; the formula for calculating the evaluation metric ACM is: Among them, the entropy function H(q) = ∑ k q k logq k , q t represents the prediction vector after the data has passed through the feature extractor g and the classifier h of the model M, and K is the total number of categories; (6) According to the evaluation index ACM corresponding to each model, select the model with the best transfer effect among several classification models, and use this classification model for image classification.

2. The unsupervised evaluation method for image recognition domain adaptation according to claim 1, characterized in that, In step (1), the source domain with annotations is: The target domain without annotations is: represents the image of the i-th sample in the source domain, represents the label of the i-th sample, n s represents the total number of source domain samples, represents the image of the j-th sample in the target domain, n t represents the total number of target domain samples; The existing domain adaptation algorithms include DANN, CDAN, MDD, MCC. Several classification models are trained from the image training sets. Each classification model consists of a feature extractor g and a linear classification layer f.

3. The unsupervised evaluation method for image recognition domain adaptation according to claim 2, characterized in that In step (2), the accuracy rate A of the source domain S The calculation formula is as follows: In the formula, denotes taking the mathematical expectation of the samples in the source domain validation set, denotes the image of a sample in the source domain validation set, denotes the label corresponding to this sample, p s denotes the data denotes the predicted vector after the data passes through the model M, which has K components, and K is the total number of categories; denotes p s the position of the largest one among the K components of ; I[·] denotes the indicator function, which has a value of 1 if the content in the parentheses is true, and a value of 0 otherwise.

4. The unsupervised evaluation method for image recognition domain adaptation according to claim 1, characterized in that In step (3), the intermediate features of the source domain data passing through model M at the time of step (2) are retained. Among them, g is the feature extractor of model M; the intermediate features and the labels are combined into an intermediate feature dataset: On this intermediate feature dataset, a classifier h is supervised and trained, which consists of two fully connected layers.

5. The unsupervised evaluation method for image recognition domain adaptation according to claim 1, characterized in that In step (4), the image validation set data of the target domain is augmented through data augmentation to obtain the data after data augmentation Among them, the data augmentation used includes random cropping, random flipping, random color adjustment, and random blurring.