Small-sample Image Classification Method Based on Prototype Complementation

By combining intra-class information and inter-class information in small sample image classification, we adaptively supplement task information and generate more distinctive prototypes, solving the problem of interference in the existing technology of prototype generation, and improving classification accuracy and model adaptability.

CN117132830BActive Publication Date: 2025-06-27CHONGQING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311121099.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-30
Publication Date
2025-06-27
Estimated Expiration
2043-08-30

AI Technical Summary

Technical Problem

When the existing small sample learning methods face too few samples and large data deviations, prototype generation is susceptible to background interference, resulting in poor classification effect.

Method used

A small sample image classification method based on prototype complementation is adopted, through in-class information extraction and inter-class information fusion, adaptively supplement information to the task to generate a more distinctive prototype.

Benefits of technology

It improves the accuracy of image classification, enhances the model's adaptability to tasks with different deviations, and has good migration and portability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117132830B_ABST
    Figure CN117132830B_ABST
Patent Text Reader

Abstract

The present invention discloses a small-sample image classification method based on prototype complementation, including: S1. Extracting the image features Z of the support set s and the image features Z of the query set q ; S2. Performing intra-class information extraction processing on the image features Z s to obtain the feature prototype c final ; S3. Performing inter-class information extraction processing on the image features Z q to obtain new samples; performing average pooling on the image features Z q to obtain the pooled feature image; S4. Inputting the feature prototype c final and the new samples into the first classifier to calculate the loss function L reg ; inputting the feature prototype c final and the pooled feature image into the second classifier to calculate the loss function L meta ; S5. Using the loss function L reg and the loss function L meta to construct the loss function L of the image classification model, inputting the image to be tested into the image classification model, and outputting the classification result of the image to be tested. The present invention can adaptively supplement information for each small-sample task by combining the sample characteristics in each small-sample task, improving the accuracy of image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image classification, and particularly to a few-shot image classification method based on prototype complementation. Background Art

[0002] Deep learning methods using a large amount of labeled data have achieved great results in the field of image classification. However, in many real-world scenarios, the labeled data is extremely limited, and the classification effect of the model will drop significantly when facing such task scenarios. In order to enable the neural network model to have strong classification ability, few-shot learning (FSL) has become an important and widely studied object. Different from traditional machine learning, few-shot learning, on the one hand, has to face the problem of too few samples, and on the other hand, it also needs to be able to quickly adapt to new class tasks with large deviations from the training set data. Such deviations mean that the categories, image domains, and fine-grainedness in the new class data are all different from the training data.

[0003] Currently, few-shot learning uses metric learning methods such as prototype networks. By learning the representation of images, the distance between the features of each image is calculated through a distance function in the feature space to complete classification. However, the prototype network simply obtains the prototype by averaging the samples in the support set, without considering the problem of sample correlation in the task. When facing a task with a large deviation from the training task, the generation of the prototype is easily affected by interference information such as the background, resulting in a poor classification effect. Therefore, a few-shot image classification method based on prototype complementation is needed to solve the above problems. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to overcome the defects in the prior art and provide a few-shot image classification method based on prototype complementation, which can combine the sample characteristics in each few-shot task, adaptively supplement information for the task, and improve the accuracy of image classification.

[0005] The few-shot image classification method based on prototype complementation of the present invention includes the following steps:

[0006] S1. Extract the image feature Z of the support set s and the image feature Z of the query set q ;

[0007] S2. Perform intra-class information extraction processing on the image feature Z s to obtain the feature prototype c final ;

[0008] S3. Perform inter-class information extraction processing on the image feature Z q to obtain new samples; perform average pooling on the image feature Z q to obtain the pooled feature image;

[0009] S4. Input the feature prototype c final and the new sample into the first classifier to calculate the loss function L reg ; Input the feature prototype c final and the pooled feature image into the second classifier to calculate the loss function L meta ;

[0010] S5. Use the loss function L reg and the loss function L meta to construct the loss function L of the image classification model, input the image to be tested into the image classification model, and output the classification result of the image to be tested.

[0011] Furthermore, perform intra-class information extraction processing on the image feature Z s to obtain the feature prototype c final , specifically including:

[0012] S21. Reconstruct the image feature Z s to obtain the reconstructed image feature Z s ';

[0013] S22. Construct an intra-class information extractor; the intra-class information extractor includes a 3×3 convolutional layer, a 1×1 convolutional layer, a Transformer layer, a LeakyRelu activation function, a max pooling function, and an average pooling function;

[0014] Input the reconstructed image feature Z s ' into the intra-class information extractor and perform the following operations:

[0015] Perform 3×3 convolution to obtain a fused feature map;

[0016] After passing the fused feature map through the LeakyReLu activation function and the max pooling operation, the size of the feature map is reduced to the size of the normal sample feature map;

[0017] Perform 1×1 convolution to obtain a second fused feature map; use the Transformer layer to perform self-attention calculation to further enhance the second fused feature map to obtain an enhanced feature;

[0018] Process the enhanced feature through the LeakyReLu activation function and global average pooling, and output the class description information;

[0019] S23. Take the sum of the class description information and the initial feature prototype as the feature prototype c final .

[0020] Furthermore, according to the following method, for the image feature Z sPerform reconstruction to obtain the reconstructed image feature Z s ′:

[0021] Transform the image feature Z s into the image feature Z s-1 :

[0022]

[0023] where N S is the number of support set samples, with a size of N×K, N represents the number of support set categories, and K represents the number of samples in each category of the support set; C, H, and W are the number of channels, height, and width of the sample feature vector respectively; → is the transformation symbol; is the set of real numbers;

[0024] Transform the image feature Z s-1 into the image feature Z s ′:

[0025]

[0026] where is the floor symbol.

[0027] Furthermore, determine the initial feature prototype according to the following method:

[0028] Perform global average pooling on the image feature Z s of the support set, then calculate the mean of the pooled features for each category, and use the obtained mean as the initial feature prototype.

[0029] Furthermore, perform inter-class information extraction processing on the image feature Z q to obtain new samples;

[0030] S31. Perform reconstruction on the image feature Z q to obtain the reconstructed image feature Z q ′;

[0031] S32. Construct an inter-class information extractor; the inter-class information extractor includes a 5×1 convolutional layer, a 1×1 convolutional layer, a Transformer layer, a LeakyRelu activation function, and an average pooling function;

[0032] Input the reconstructed image feature Z q ′ into the inter-class information extractor and perform the following operations:

[0033] Perform 5×1 convolution to obtain the fused feature map Z′ q-f ;

[0034] Input the fused feature map Z′ q-fProcessed by the LeakyReLu activation function and adjusted in image size to obtain the feature map Z' after size adjustment q-s ;

[0035] Perform a 1×1 convolution on the feature map Z' q-s to obtain the fused feature map

[0036] Perform self-attention calculation using the Transformer layer to further enhance the feature map and obtain the enhanced feature Z' q-e ;

[0037] Process the enhanced feature Z' through the LeakyReLu activation function and global average pooling q-e and output the new sample

[0038] Furthermore, reconstruct the image feature Z q according to the following method to obtain the reconstructed image feature Z q ':

[0039] Transform the image feature Z q into the image feature Z q ' according to the following formula

[0040]

[0041] where N q is the number of samples in the query set, with a size of N×m, N representing the number of query set categories, and m representing the number of samples in each category of the query set; C, H, and W are the number of channels, height, and width of the sample feature vector respectively; → is the transformation symbol is the set of real numbers

[0042] Furthermore, determine the loss function L reg according to the following formula

[0043]

[0044] where p(y q′ =k'|x q′ ) represents the probability that the q'-th new sample x q′ belongs to the k'-th class of samples, y q′ represents the predicted label of the q'-th new sample; Y q′ represents the true label of the q'-th new sample; N' is the number of new sample categories; m is the number of new samples

[0045] Furthermore, determine the loss function L meta according to the following formula

[0046]

[0047] Among them, P(y q =k|x q ) represents the probability that the q-th pooled feature image x q belongs to the k-th class of samples, and y q represents the predicted label of the q-th pooled feature image; Y q represents the true label of the q-th pooled feature image; N is the number of classes of the pooled feature images; M is the number of pooled feature images.

[0048] Furthermore, the loss function L is determined according to the following formula:

[0049] L = L meta + λL reg ;

[0050] where λ is the loss balance factor.

[0051] The beneficial effects of the present invention are as follows: A small-sample image classification method based on prototype complementation disclosed by the present invention uses intra-class information to complement the prototype, making the generated prototype more discriminative; uses inter-class information fusion to obtain generated samples, participates in classification and optimizes the model, and improves the quality of feature representation. The trained model can better quickly adapt to the task when facing small-sample tasks with different biases, can generate a small-sample classification model that does not rely on the attention mechanism, and has good transferability and portability. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The present invention will be further described below with reference to the drawings and embodiments:

[0053] Figure 1 is a schematic flow chart of the image classification method of the present invention;

[0054] Figure 2 is a schematic diagram of the principle of the image classification processing framework of the present invention;

[0055] Figure 3 (a) is a schematic diagram of the extraction principle of the intra-class information extractor facing large data distribution deviation;

[0056] Figure 3 (b) is a schematic diagram of the extraction principle of the intra-class information extractor facing small data distribution deviation;

[0057] Figure 4 (a) is a schematic diagram of the extraction principle of the inter-class information extractor facing large data distribution deviation;

[0058] Figure 4 (b) is a schematic diagram of the extraction principle of the inter-class information extractor facing small data distribution deviation. Detailed implementation manners

[0059] The following further describes the present invention in conjunction with the accompanying drawings of the specification, as shown in the figures:

[0060] The small-sample image classification method based on prototype complementation of the present invention includes the following steps:

[0061] S1. Extract the image feature Z of the support set s and the image feature Z of the query set q ; wherein, a backbone network can be used to extract the image features, such as using the existing ResNet12;

[0062] S2. Perform intra-class information extraction processing on the image feature Z s to obtain the feature prototype c final ;

[0063] S3. Perform inter-class information extraction processing on the image feature Z q to obtain new samples; perform average pooling on the image feature Z q to obtain the pooled feature image;

[0064] S4. Input the feature prototype c final and the new samples into the first classifier to calculate the loss function L reg ; input the feature prototype c final and the pooled feature image into the second classifier to calculate the loss function L meta ;

[0065] S5. Use the loss function L reg and the loss function L meta to construct the loss function L of the image classification model, input the image to be tested into the image classification model, and output the classification result of the image to be tested. Among them, the image classification model is constructed based on the existing deep learning network model. For example, an initial image classification model is constructed based on the existing prototype network, and the loss function L is generated. This loss function L is used as the loss function of the image classification model, so as to obtain the optimized image classification model, and the optimized image classification model is used to accurately and effectively classify the image to be tested.

[0066] In this embodiment, in step S1, the image features are extracted by constructing a small-sample classification task; specifically, a small-sample classification task is described in the form of N-way K-shot. After observing a small amount of data of N classes (observing K samples for each class), the model attempts to classify the remaining samples of the N classes. In this task, the observable samples are the support set The remaining test samples are the query set where, x iand x j For the images in the support set and the query set, y i and y j are the corresponding images of x i and x j The labels of, K is the number of samples in each support set category, and M is the number of samples in the query set.

[0067] Taking the 5-way 5-shot classification task as an example, 5 categories are randomly selected from the dataset, and 20 samples are randomly selected from each category of samples. Among them, 5 samples of the same category are used as the support set with known labels, and the remaining 15 samples of the same category are used as the unlabeled query set to be classified. During the few-shot learning process, the model constructs a series of classification tasks T on the training set train for training in order to obtain the generalization ability across tasks. During evaluation, a large number of classification tasks T are constructed on the test set test and the mean value of the classification accuracy is calculated as the accuracy of the final classification.

[0068] In this embodiment, in step S2, as Figure 2 shown, the intra-class information extraction process is performed on the image feature Z s to obtain the feature prototype c final , which specifically includes:

[0069] S21. Reconstruct the image feature Z s to obtain the reconstructed image feature Z s ';

[0070] The image feature Z s is reconstructed according to the following method to obtain the reconstructed image feature Z s ':

[0071] The image feature Z s is transformed into the image feature Z s-1 according to the following formula:

[0072]

[0073] where N S is the number of support set samples, with a size of N×K, N represents the number of support set categories, and K represents the number of samples in each category of the support set; C, H, and W are the number of channels, height, and width of the sample feature vector respectively; → is the transformation symbol; is the set of real numbers;

[0074] The image feature Z s-1 is transformed into the image feature Z s ' according to the following formula:

[0075]

[0076] Among them, is the floor symbol.

[0077] S22. Construct an Intra-Class Information Extractor (ICIE); the Intra-Class Information Extractor includes a 3×3 convolutional layer, a 1×1 convolutional layer, a Transformer layer, a LeakyRelu activation function, a max pooling function, and an average pooling function;

[0078] Input the reconstructed image feature Z s ' into the Intra-Class Information Extractor and perform the following operations, as Figure 3 shown:

[0079] Perform 3×3 convolution to obtain a fused feature map;

[0080] After passing the fused feature map through the LeakyReLu activation function and max pooling operation, the size of the feature map is reduced to the size of the normal sample feature map;

[0081] Perform 1×1 convolution to obtain a second fused feature map; use the Transformer layer for self-attention calculation to further enhance the second fused feature map and obtain an enhanced feature;

[0082] Process the enhanced feature through the LeakyReLu activation function and global average pooling, and output the class description information;

[0083] Of course, in combination with the actual working conditions, different intra-class information extraction methods can be adopted according to the degree of data distribution deviation. In the case of large data distribution deviation, the extraction principle shown in Figure 3 (a) can be used for intra-class information extraction; in the case of small data distribution deviation, the extraction principle shown in Figure 3 (b) can be used for intra-class information extraction;

[0084] S23. Use the sum of the class description information and the initial feature prototype as the feature prototype c final . Among them, the feature prototypes of each class can be combined into a set N represents the number of classes, and k represents the class number.

[0085] Among them, the initial feature prototype is determined according to the following method:

[0086] Perform global average pooling on the image feature Z s of the support set, then calculate the mean of the pooled features for each class, and use the obtained mean as the initial feature prototype.

[0087] In this embodiment, in step S3, as Figure 2As shown, perform inter-class information extraction processing on the image feature Z q to obtain new samples;

[0088] S31. Reconstruct the image feature Z q to obtain the reconstructed image feature Z q ';

[0089] Reconstruct the image feature Z q according to the following method to obtain the reconstructed image feature Z q ':

[0090] Transform the image feature Z q into the image feature Z q ' according to the following formula:

[0091]

[0092] where N q is the number of samples in the query set, with a size of N×m, N represents the number of query set categories, and m represents the number of samples in each category of the query set; C, H, and W are the number of channels, height, and width of the sample feature vector respectively; → is the transformation symbol; is the set of real numbers.

[0093] S32. Construct an inter-class information extractor (CCIE); the inter-class information extractor includes a 5×1 convolutional layer, a 1×1 convolutional layer, a Transformer layer, a LeakyRelu activation function, and an average pooling function;

[0094] Input the reconstructed image feature Z q ' into the inter-class information extractor and perform the following operations, as Figure 4 shown:

[0095] Perform 5×1 convolution to obtain the fused feature map Z' q-f ;

[0096] Process the fused feature map Z' q-f through the LeakyReLu activation function and adjust the image size to obtain the size-adjusted feature map Z' q-s ;

[0097] Perform 1×1 convolution on the feature map Z' q-s to obtain the fused feature map

[0098] Use the Transformer layer for self-attention calculation to further enhance the feature map to obtain the enhanced feature Z' q-e ;

[0099] The feature Z' enhanced by processing through the LeakyReLu activation function and global average pooling q-e , and new samples are output. Among them, each new sample can form a combined set Q'.

[0100] Similarly, in combination with the actual working conditions, different inter-class information extraction methods can be adopted according to the degree of data distribution deviation. In the face of a large data distribution deviation, the extraction principle shown in Figure 4 (a) can be used for inter-class information extraction; in the face of a small data distribution deviation, the extraction principle shown in Figure 4 (b) can be used for inter-class information extraction.

[0101] In this embodiment, both the first classifier and the second classifier adopt existing cosine classifiers;

[0102] The loss function L is determined according to the following formula reg :

[0103]

[0104] where p(y q′ =k'|x q′ ) represents the probability that the q'-th new sample x q′ belongs to the k'-th class of samples, and y q′ represents the predicted label of the q'-th new sample; Y q′ represents the true label of the q'-th new sample; N' is the number of classes of new samples; m is the number of new samples.

[0105] The loss function L is determined according to the following formula meta :

[0106]

[0107] where P(y q =k|x q ) represents the probability that the q-th pooled feature image x q belongs to the k-th class of samples, and y q represents the predicted label of the q-th pooled feature image; Y q represents the true label of the q-th pooled feature image; N is the number of classes of the pooled feature images; M is the number of the pooled feature images. Among them, the pooled feature images are the query set embeddings Q_Eembeddings as shown in Figure 2 .

[0108] The above loss function L reg and the loss function L meta corresponding probabilities both adopt the following same calculation principle:

[0109]

[0110] The cosine similarity function φ(·) calculates the distance between the sample x (corresponding to the new sample or the pooled feature image) and each class prototype c i and selects the class with the highest probability score as the predicted class of the sample and assigns a label to it.

[0111] Among them, for the new samples generated by extracting inter-class information, soft labels are used for cross-entropy calculation. Since the new samples contain information of various samples, the soft label values are (0.2, 0.2, 0.2, 0.2, 0.2), indicating that the probabilities of belonging to 5 classes are all 0.2, and their sum is 1.

[0112] The loss function L is determined according to the following formula:

[0113] L = L meta + λL reg ;

[0114] where λ is the loss balance factor, and λ can take the value of 0.1.

[0115] The constructed loss function L is used to optimize the network parameters of the image classification model, so as to obtain a more mature image classification model.

[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A small-sample image classification method based on prototype complementation, characterized in that: It includes the following steps: S1. Extract the image features Z of the support set s and the image features Z of the query set q ; S2. Perform in-class information extraction processing on the image feature Z s to obtain the feature prototype c final , which specifically includes: S21. Reconstruct the image feature Z s to obtain the reconstructed image feature Z s ′; S22. Construct an intra-class information extractor; the intra-class information extractor includes a 3×3 convolutional layer, a 1×1 convolutional layer, a Transformer layer, a LeakyRelu activation function, a max pooling function, and an average pooling function; Input the reconstructed image feature Z s ' into the intra-class information extractor for the following operations: Perform 3×3 convolution to obtain a fused feature map; After passing the fused feature map through the LeakyReLu activation function and the max pooling operation, the size of the feature map is reduced to the size of the normal sample feature map; Perform 1×1 convolution to obtain a second fused feature map; use the Transformer layer to perform self-attention calculation to further enhance the second fused feature map to obtain enhanced features; Process the enhanced features through the LeakyReLu activation function and global average pooling, and output class description information; S23. Use the sum of the category description information and the initial feature prototype as the feature prototype c final ; S3. Extract the inter-class information from the image feature Z q to obtain a new sample; perform average pooling on the image feature Z q to obtain the pooled feature image; Among them, for the image feature Z q perform inter-class information extraction processing to obtain new samples, specifically including: S31. Reconstruct the image feature Z q to obtain the reconstructed image feature Z q ′; S32. Construct an inter-class information extractor; the inter-class information extractor includes a 5×1 convolutional layer, a 1×1 convolutional layer, a Transformer layer, a LeakyRelu activation function, and an average pooling function; Input the reconstructed image feature Z q into the inter-class information extractor and perform the following operations: Perform a 5×1 convolution to obtain the fused feature map Z′ q-f ; The fused feature map Z' q-f is processed by the LeakyReLu activation function and the image size is adjusted to obtain the feature map Z' with adjusted size q-s ; Perform a 1×1 convolution on the feature map Z′ q-s to obtain the fused feature map Perform self-attention calculation using the Transformer layer on the feature map to further enhance it and obtain the enhanced feature Z' q-e ; The feature Z' enhanced by processing through the LeakyReLu activation function and global average pooling q-e , output a new sample; S4. Input the feature prototype c final into the first classifier along with the new sample, and calculate the loss function L reg ; Input the feature prototype c final and the pooled feature image into the second classifier, and calculate the loss function L meta ; S5. Use the loss function L reg and the loss function L meta , construct the loss function L of the image classification model, input the image to be tested into the image classification model, and output the classification result of the image to be tested.

2. The small-sample image classification method based on prototype complementation according to claim 1, wherein: Reconstruct the image feature Z according to the following method s to obtain the reconstructed image feature Z s ': Transform the image feature Z according to the following formula s into the image feature Z s-1 : Among them, N S is the number of samples in the support set, with a size of N×K. N represents the number of support set categories, and K represents the number of samples in each category of the support set; C, H, and W are the number of channels, height, and width of the sample feature vector, respectively; → is the transformation symbol; is the set of real numbers; Transform the image feature Z according to the following formula s-1 into the image feature Z s ': Among them, is the floor symbol.

3. The small-sample image classification method based on prototype complementation according to claim 1, wherein: Determine the initial feature prototype according to the following method: The image features Z of the support set s are subjected to global average pooling, and then the mean value of the pooled features for each category is calculated, and the obtained mean value is used as the initial feature prototype.

4. The small-sample image classification method based on prototype complementation according to claim 1, characterized in that: Reconstruct the image feature Z according to the following method q to obtain the reconstructed image feature Z q ': Transform the image feature Z according to the following formula q into the image feature Z q ': Among them, N q is the number of query set samples, with a size of N×m. N represents the number of query set categories, and m represents the number of samples in each category of the query set; C, H, and W are the number of channels, height, and width of the sample feature vector respectively; → is the transformation symbol; is the set of real numbers.

5. The small-sample image classification method based on prototype complementation according to claim 1, wherein: Determine the loss function L according to the following formula reg : Among them, p(y q′ = k′|x q′ ) represents the probability that the q'-th new sample x q′ belongs to the k'-th class of samples, and y q′ represents the predicted label of the q'-th new sample; Y q′ represents the true label of the q'-th new sample; N' is the number of classes of new samples; m is the number of new samples.

6. The small-sample image classification method based on prototype complementation according to claim 1, characterized in that: Determine the loss function L according to the following formula meta : Among them, P(y q = k|x q ) represents the probability that the q-th pooled feature image x q belongs to the k-th class of samples, and y q represents the predicted label of the q-th pooled feature image; Y q represents the true label of the q-th pooled feature image; N is the number of classes of the pooled feature images; M is the number of pooled feature images.

7. The small-sample image classification method based on prototype complementation according to claim 1, wherein: Determine the loss function L according to the following formula: L = L meta + λL reg ; where λ is the loss balance factor.