A Cross-Domain Few-Shot Image Classification Method Based on Global-Local Knowledge Distillation

Through the global-local knowledge distillation of cross-domain small sample image classification model, the problem of insufficient semantic information capture in cross-domain small sample image classification is solved, and the efficient generalization and classification performance of the model in cross-domain is achieved.

CN115953630BActive Publication Date: 2025-07-18NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310038225.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-07-18
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively generalize in cross-domain small sample image classification because the domain difference between the source domain and the target domain makes the model unable to effectively capture the semantic information of cross-domain generalization.

Method used

A cross-domain small sample image classification model for global-local knowledge distillation is constructed. Through the combination of global branches and local branches, the global branches extract the global features, local branches extract local features, and through global-local knowledge distillation loss, promoting global features to focus on local areas, improving semantic information capture capabilities.

Benefits of technology

It improves the generalization performance of the model on cross-domain small sample tasks, and can be effectively tested on the target domain without fine-tuning after training in the source domain, achieving better classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_8
    Figure SMS_8
  • Figure SMS_13
    Figure SMS_13
  • Figure SMS_18
    Figure SMS_18
Patent Text Reader

Abstract

The present invention provides a cross-domain few-shot image classification method based on global-local knowledge distillation. A classification model composed of a global branch and a local branch is constructed. Among them, the global branch takes the original image as input and is used to extract the global features of the image, and the local branch takes the local patches of the original image as input and is used to extract the local features of the image. Between the two branches, by constructing a global-local knowledge distillation loss, the global features are promoted to pay attention to the local regions of the image, so that the global features can capture rich semantic information, thereby improving the generalization performance of the global features in cross-domain few-shot tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to a cross-domain few-shot image classification method based on global-local knowledge distillation. Background Art

[0002] Image processing is a key technology for machine vision to move towards industrial applications, and image classification is the basis of image processing technology. In various scenarios such as medicine and remote sensing, image data is often difficult to obtain, presenting typical few-shot characteristics. To alleviate the few-shot problem, an effective way is to use source domain data to learn transferable knowledge and generalize the learned knowledge to few-shot tasks in the target domain. However, due to the domain difference between the source domain and the target domain, the model trained on the source domain is difficult to effectively generalize to the target domain. Therefore, researching few-shot image classification techniques applicable to cross-domain scenarios has important application value. The literature "Snell J, Swersky K, Zemel R. Prototypical networks for few learning[C] / / Advances in Neural Information Processing Systems.2017:4077- ." proposed a prototypical few-shot image classification method. It first uses a deep neural network to extract the features of images, then constructs prototype representations of categories in the feature space using a small number of labeled samples in each few-shot task, and finally assigns category membership according to the distances between test samples and these category prototypes. However, due to the simplicity preference of the deep neural network, the prototypes constructed by this method often can only capture the most discriminative patterns, such as colors, shapes, etc., ignoring semantic information with cross-domain generalization ability. Therefore, this method performs poorly in cross-domain few-shot image classification tasks. Summary of the Invention

[0003] To overcome the deficiencies of the prior art, the present invention provides a cross-domain few-shot image classification method based on global-local knowledge distillation. A classification model composed of a global branch and a local branch is constructed. Among them, the global branch takes the original image as input and is used to extract the global features of the image, and the local branch takes the local blocks of the original image as input and is used to extract the local features of the image; between the two branches, by constructing a global-local knowledge distillation loss, the global features are promoted to pay attention to the local regions of the image, so that the global features can capture rich semantic information, thereby improving the generalization performance of the global features in cross-domain few-shot tasks.

[0004] A cross-domain few-shot image classification method based on global-local knowledge distillation, characterized by the following steps:

[0005] Step 1: Construct a small-sample task training dataset based on the existing image dataset, including the support set and the query set Among them, the support set includes N categories, each category with K supervised samples, and the query set also includes these N categories, each category with M unlabeled samples;

[0006] Step 2: Construct the global branch of the model, and its processing process is as follows:

[0007] First, obtain the prototype representation of the support set according to the following formula:

[0008]

[0009] Among them, represents the k-th sample of the n-th category in the support set , represents the feature extraction network in the global branch. In the present invention, the ResNet-10 network is adopted, and C n represents the prototype representation of the n-th category, n = 1, 2,..., N;

[0010] Then, based on the prototype representation, predict the class membership of each sample in the query set :

[0011]

[0012] Among them, represents the i-th query sample in the query set , i = 1, 2,..., N*M, represents the predicted score of this sample, and matching(·) is the similarity measurement function between two vectors. In the present invention, the Euclidean distance is used for similarity measurement;

[0013] Next, use the category corresponding to the maximum similarity in the predicted scores as the predicted label of this query sample and calculate the cross-entropy loss according to the predicted label and the true label of the query sample as follows:

[0014]

[0015] Among them, H(·) represents the cross-entropy loss function, represents the true label corresponding to the query sample , represents the predicted label of the query sample and the true label and the true label The cross-entropy loss between;

[0016] Step 3: Construct the local branch of the model, and its processing process is as follows:

[0017] For the query sample First, use random cropping to obtain its corresponding local image patches where r ∈ [1, R], represents the number of local image patches corresponding to each query image, represents the query sample the r-th local image patch of;

[0018] Then, use the feature extraction network in the local branch to extract the local features corresponding to each local image patch Among them, the feature extraction network in the local branch adopts the ResNet-10 network;

[0019] Next, use the prototypes calculated in Step 2 for the local features to perform class membership prediction, and obtain the prediction scores corresponding to each local image patch

[0020]

[0021] Among them, represents the query sample the r-th local image patch of the similarity score of, represents the query sample the r-th local image patch of the local feature of;

[0022] Step 4: Calculate the total loss of the model according to the following formula

[0023]

[0024] Among them, I represents the total number of query samples in the few-shot task, represents the global-local knowledge distillation loss of the query sample , represents the cross-image local-global distillation loss, λ1 represents the coefficient of the global-local knowledge distillation loss term, set λ1 to 1, λ2 represents the coefficient of the cross-image local-global distillation loss term, set λ2 to 0.15;

[0025] The said query sample the global-local knowledge distillation loss of is calculated according to the following formula:

[0026]

[0027] The local-global distillation loss across images is calculated as follows:

[0028]

[0029] where represents the j-th query sample in the query set and the predicted score of the r-th local image patch of, where j≠i means that j is a different sample of the same category as the i-th query sample, and j = 1, 2, …, N*M;

[0030] Step 5: According to the total loss of the model calculated in Step 4, use the stochastic gradient descent method to train the network parameters of the global branch end-to-end, and update the network parameters of the local branch according to the following formula:

[0031] θ T ←mθ T +(1 - m)θ S (8)

[0032] where θ T represents the network parameters in the local branch, m represents the momentum coefficient in the exponential moving average update, m is set to 0.998, and θ S represents the network parameters in the global branch, and ← represents the update operation;

[0033] Step 6: Input the image dataset to be processed into the global branch obtained after training in Step 5, predict the membership category of each image therein, and complete image classification.

[0034] The beneficial effects of the present invention are as follows: The global-local knowledge distillation framework constructed in the training stage promotes the global features to pay attention to the local information of the image, so that the model can learn semantic representations with strong generalization ability and improve the generalization performance on cross-domain few-shot tasks; adopting an end-to-end framework design method, once the model is trained on the source domain (training dataset), it can be tested on the few-shot tasks of any target domain (image dataset to be processed) without fine-tuning the feature extraction model; the present invention can obtain good classification results in cross-domain few-shot image classification. Specific embodiments

[0035] The present invention will be further described below in conjunction with embodiments, and the present invention includes but is not limited to the following embodiments.

[0036] The present invention provides a cross-domain few-shot image classification method based on global-local knowledge distillation, and its specific implementation process is as follows:

[0037] 1. Construct a training dataset for few-shot tasks

[0038] The cross-domain few-shot image classification task requires the model to be trained in the source domain and then process the few-shot tasks in the target domain Therefore, first construct a training dataset for few-shot tasks based on the existing image dataset. Specifically: randomly sample N classes from the dataset, and randomly sample K supervised samples for each class. These N*K samples form the support set At the same time, randomly sample M unlabeled samples from N classes. These N*M samples form the query set

[0039] 2. Global branch calculation

[0040] Construct the global branch of the model, and its processing process is as follows:

[0041] First, obtain the prototype representation of the support set according to the following formula:

[0042]

[0043] where represents the k-th sample of the n-th class in the support set , represents the feature extraction network in the global branch. In the present invention, the ResNet-10 network is adopted, and C n represents the prototype representation of the n-th class, n = 1, 2,..., N;

[0044] Then, based on the prototype representation, predict the class membership of each sample in the query set :

[0045]

[0046] where represents the i-th query sample in the query set , i = 1, 2,..., N*M, represents the predicted score of this sample, and matching(·) is the similarity measurement function between two vectors. In the present invention, the Euclidean distance is used for similarity measurement;

[0047] Next, according to the class corresponding to the maximum similarity in the predicted scores as the predicted label of this query sample and calculate the cross-entropy loss according to the predicted label and the true label of the query sample as follows:

[0048]

[0049] Among them, H(·) represents the cross-entropy loss function, represents the query sample corresponding true label, represents the query sample predicted label and the true label cross-entropy loss between.

[0050] 3. Local branch calculation

[0051] Construct the local branch of the model, and its processing process is as follows:

[0052] For the query sample First, use random cropping to obtain its corresponding local image patch where r ∈ [1, R], represents the number of local image patches corresponding to each query image, represents the query sample the r-th local image patch;

[0053] Then, similar to step 2, first use the feature extraction network in the local branch to extract the local features corresponding to each local image patch Among them, the feature extraction network in the local branch adopts the ResNet-10 network. Then use the prototype calculated in step 2 to perform class membership prediction on the local features to obtain the prediction scores corresponding to each local image patch

[0054]

[0055] Among them, represents the query sample the r-th local image patch similarity score, represents the query sample the r-th local image patch local feature.

[0056] 4. Calculate the total loss

[0057] Calculate the total loss of the model according to the following formula

[0058]

[0059] Among them, I represents the total number of query samples in the few-shot task, represents the global-local knowledge distillation loss of the query sample , Denote the local-global distillation loss across images. Let λ1 denote the coefficient of the global-local knowledge distillation loss term, and set λ1 to 1. Let λ2 denote the coefficient of the local-global distillation loss term across images, and set λ2 to 0.15;

[0060] The query sample mentioned above of the global-local knowledge distillation loss The predicted score (similarity score) corresponding to the global feature calculated according to Step 2 and the predicted score (similarity score) corresponding to the local feature calculated according to Step 3 are calculated as follows:

[0061]

[0062] The local-global distillation loss across images mentioned above is designed to constrain the semantic consistency across images, and its calculation formula is as follows:

[0063]

[0064] where denotes the query set the predicted score of the r-th local image patch of the j-th query sample in , j≠i means that j is a different sample of the same category as the i-th query sample, and j = 1, 2, …, N*M.

[0065] 5. Train the model

[0066] According to the total loss of the model calculated in Step 4, use the stochastic gradient descent method to train the network parameters of the global branch end-to-end. For the network parameters of the local branch, use the exponential moving average of the global branch to update the parameters, that is:

[0067] θ T ←mθ T +(1 - m)θ S (16)

[0068] where θ T denotes the network parameters in the local branch, m denotes the momentum coefficient in the exponential moving average update, set m to 0.998, and θ S denotes the network parameters in the global branch, and ← denotes the update operation.

[0069] 6. Image classification

[0070] After the model training is completed, discard the local branch and only retain the global branch to classify the few-shot images in the target domain, that is, input the image dataset to be processed into the global branch obtained after training in step 5, and according to the calculation process in step 2, predict the membership category of each image therein to complete the image classification task.

[0071] The present invention can achieve good classification performance in the cross-domain few-shot image classification task. For example, in this embodiment, the mini-ImageNet dataset is used as the training dataset in the source domain for model training, and then the remote sensing scene classification dataset EuroSAT and the medical image dataset ISIC are used as the target domain for classification processing. The method of the present invention achieves classification accuracies of 63.70% and 33.51% respectively in the 5-way 1-shot task (the support set contains 5 categories, and there is 1 sample in each category), which are 4.59% and 1.78% higher than the existing prototype-based few-shot image classification methods respectively.

Claims

1. A cross-domain few-shot image classification method based on global-local knowledge distillation, characterized in that The steps are as follows: Step 1: Construct a small-sample task training dataset based on the existing image dataset, including a support set and a query set Among them, the support set includes N categories, each category with K supervised samples, and the query set also includes these N categories, each category with M unlabeled samples; Step 2: Construct the global branch of the model, and its processing procedure is as follows: First, obtain the support set according to the following formula The prototype representation of: Among them, represents the support set the k-th sample of the n-th category in represents the feature extraction network in the global branch. In the present invention, the ResNet-10 network is adopted, C n represents the prototype representation of the n-th category, where n = 1, 2,..., N; Then, based on the prototype representation, the class membership of each sample in the query set is predicted: Among them, represents the i-th query sample in the query set where i = 1, 2, …, N*M, represents the predicted score of the sample, and matching(·) is a similarity measurement function between two vectors. In the present invention, the Euclidean distance is used for similarity measurement; Next, use the category corresponding to the maximum similarity in the predicted scores as the predicted label for this query sample And calculate the cross-entropy loss based on the predicted label and the true label of the query sample as follows: Among them, H(·) represents the cross-entropy loss function, represents the query sample corresponding true label, represents the query sample predicted label and the true label cross-entropy loss between; Step 3: Construct the local branch of the model, and its processing procedure is as follows: For query samples First, randomly crop to obtain its corresponding local image patches where r ∈ [1, R], and R represents the number of local image patches corresponding to each query image denotes the query sample and the r-th local image patch of Then, use the feature extraction network in the local branch Extract the local features corresponding to each local image patch Among them, the feature extraction network in the local branch Adopts the ResNet-10 network; Next, use the prototype calculated in step 2 to predict the class membership of the local features and obtain the prediction scores corresponding to each local image patch Among them, represents the similarity score of the r-th local image patch of the query sample , and represents the local feature of the r-th local image patch of the query sample . Step 4: Calculate the total loss of the model according to the following formula where I represents the total number of query samples in the small sample task, represents the query sample of the global-local knowledge distillation loss, represents the cross-image local-global distillation loss, λ1 represents the coefficient of the global-local knowledge distillation loss term, λ1 is set to 1, λ2 represents the coefficient of the cross-image local-global distillation loss term, λ2 is set to 0.15; the said query sample of the global-local knowledge distillation loss It is calculated according to the following formula: The local-global distillation loss across images It is calculated according to the following formula: Among them, Denote the query set The j-th query sample in The prediction score of the r-th local image patch of, where j≠i indicates that j is a different sample belonging to the same category as the i-th query sample, and j = 1, 2, …, N*M; Step 5: According to the total loss of the model calculated in Step 4, use the stochastic gradient descent method to train the network parameters of the global branch end-to-end, and update the network parameters of the local branch according to the following formula: θ T ←mθ T +(1 - m)θ S (8) Among them, θ T represents the network parameters in the local branch, m represents the momentum coefficient in the exponential moving average update, m is set to 0.998, θ S represents the network parameters in the global branch, and ← represents the update operation; Step 6: Input the image dataset to be processed into the global branch obtained after training in Step 5, predict the membership category of each image therein, and complete image classification.

Citation Information

Patent Citations

  • Small sample classification method based on twinborn knowledge distillation and self-supervised learning

    CN114298160A

  • Method of domain adaptation applied to image segmentation, apparatus, and storage medium

    JP2022164597A