A small-sample learning image classification method based on meta-features
Through the small sample learning method based on meta features, combined with local and global feature measurement modules, the problem of low overfitting and classification accuracy in small sample learning is solved, and efficient image classification under a small number of samples is achieved.
Patent Information
- Application Number
- CN202210919656.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-02
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-08-02
AI Technical Summary
Existing deep learning methods are prone to overfitting in small sample learning scenarios, and shallow network models have low prediction accuracy in coarse and fine-grained tasks, making it difficult to effectively use a small number of samples for classification.
Using a small sample learning method based on meta features, meta features are extracted through a two-layer convolutional network, and combined with the measurement module of local features and global features, local and global features are measured using cosine similarity and Euclidean distance, and information is fused to improve the accuracy of feature representation.
In the small sample learning scenario, the classification accuracy is significantly improved, and the robustness and feasibility on different data sets are demonstrated, and the feature extraction and classification effects are optimized.
Smart Images

Figure CN115272688B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of small-sample learning image classification, and relates to a small-sample learning image classification method based on meta-features. Background Art
[0002] The development of deep learning enables computers to achieve good results in solving many artificial intelligence-related tasks, such as image recognition, speech recognition, and machine translation. However, all kinds of deep learning methods require a large amount of available labeled data for training various models. Therefore, in some fields where data is scarce, deep learning cannot solve the corresponding tasks. Specifically, many current methods are based on the condition of having a large amount of training data, but such models are powerless when faced with tasks with only a small amount of data. On the contrary, humans can quickly recognize and learn a new category with only a small number of samples or no samples. For example, people can recognize an object they have never seen before or hear a description of it and then recognize it the next time they see it. Inspired by humans' ability to learn well even with only a small number of samples, we want machines to have this ability, which helps reduce data collection work and brings good news to some fields where it is difficult to collect data.
[0003] In the small-sample learning scenario, the limitation of the sample size will lead to serious overfitting of the deep network model, but it is not easy for the shallow network model to extract good sample features. At the same time, in tasks with fine and coarse granularities, the prediction results of fine-grained tasks often have low accuracy. In order to enable the shallow network model to extract better sample features and be able to handle tasks with fine and coarse granularities in small-sample learning, we hope to make full use of the existing small number of samples and use multiple classification models to train an abstract meta-feature from two dimensions of local features and global features, obtain a more general feature representation, and further obtain image features that are easier to classify. Summary of the Invention
[0004] Based on existing research, the present invention proposes a meta-feature extraction network. The purpose of this network is to extract a meta-feature, and then, based on this meta-feature, use two different methods to measure the similarity between samples. One method pays more attention to local features, and the other method pays more attention to global features. Then, through the meta-feature extraction network, the information of local features and global features is included to further improve the accuracy of feature representation.
[0005] Technical Solution of the Present Invention:
[0006] A small-sample learning image classification method based on meta-features. The small-sample learning image classification method is based on a single extracted meta-feature and uses two methods to measure the similarity between samples, obtaining local features and global features respectively. Then, through a meta-feature extraction network, the information of the local features and global features is fused to improve the accuracy of feature representation. The specific steps are as follows:
[0007] Step 1: Small-sample learning task construction stage
[0008] First, from the small-sample learning image dataset, a training set and a test set are proportionally divided, and for the corresponding small-sample learning tasks, a large number of tasks are randomly sampled and constructed. Each task contains a support set and a query set. Each small-sample learning classification task includes N categories, with K samples in each category. A two-layer convolutional network M σ is used to extract features from the picture x to obtain the meta-feature m, where m = M σ (x);
[0009] Step 2: Image meta-feature extraction and sample similarity measurement stage
[0010] A two-layer convolutional network is used for meta-feature extraction, and then two modules of local feature measurement and global feature measurement are used to constrain and optimize the meta-feature. The samples are classified from these two different perspectives, and the information of the two modules is fused.
[0011] The meta-feature m is processed through a local feature processing network and a global feature processing network respectively
[0012] 2.1 Use a local feature processing network f to process the meta-feature m to obtain the processed local feature representation m f = f θ (m) ∈ R c×w×d , where c is the number of channels after convolution, and w and d are the length and width of the feature matrix in each channel; the set composed of the local feature descriptors of the local feature representation m f is F refers to the number of local feature descriptors that a sample can finally obtain; in each small-sample learning classification task, there are a total of KF local feature descriptors for one category in the support set; for the query set samples to be classified the local feature representation is obtained. Using the k-nearest neighbor method, the similarity between the query set sample and the support set category o is measured,
[0013]
[0014] Cosine similarity is used to measure the distance between local feature descriptors, Among all the local feature descriptors representing class o, the j-th of the k local feature descriptors that are most similar to ; After calculating the similarities between the query set samples and all classes in the support set, the soft classification prediction label for the query set samples is obtained All samples in the query set and the true label y of the samples use the cross-entropy loss,
[0015]
[0016] where y j represents the true label of the j-th query set, and 1[y j =o] means taking 1 when y j =o, otherwise taking 0;
[0017] 2.2 Use a global feature processing network g to process the meta-feature m to obtain the processed global feature representation m g = g δ (m); For each class, calculate the global feature representation representing the class according to the global feature representation of each sample, and finally take the average; For class o, its global feature representation is
[0018]
[0019] where S o represents the sample set of class o, and |S o | represents the number of samples in class o, and (x i , y i ) represents the feature vector and label of the sample;
[0020] Then use the Euclidean distance function d(·) to obtain the distance between two samples in the embedding space. The distance scores of the query set samples to class i form a distance score vector d g , that is
[0021]
[0022] where, after c i represents the global feature representation of class i, normalize all the obtained distances, and obtain the probability P that the query set sample belongs to a certain class o,
[0023]
[0024] Use the cross-entropy loss function, that is
[0025]
[0026] Step 3: Model Training Phase
[0027] Use cosine similarity to calculate the similarity between the query set images and the support set images in each task. During training, according to the existing labels of the images, use the stochastic gradient descent strategy to train the meta-feature extraction network and the networks of the two metric modules to obtain the optimal model parameters.
[0028] According to the loss functions of the local feature representation and the global feature representation, use the weighted summation method to obtain the final loss function, that is
[0029] L = αL f + βL g
[0030] Both α and β are positive hyperparameters used to adjust the weight ratios of the two feature metric parts; the final predicted label
[0031] Step 4: Classification Task Evaluation Phase
[0032] In the state where the query set has no labels, calculate the similarity between the query set samples and the support set samples, and classify the query set samples into the most similar category. Then fuse the classification accuracies of the two modules and statistically calculate the classification accuracy under a large number of tasks.
[0033] The beneficial effects of the present invention. The small-sample image classification method based on meta-features can, through this new method, extract better image features using only a small number of pictures in the small-sample learning scenario. Because in the classification process, the images of different categories may vary greatly or little. For example, the spotted dog and the Shiba Inu are very similar, and more attention needs to be paid to local features, while the spotted dog and the bicycle are very different, and more attention needs to be paid to global features. Therefore, the present invention uses a meta-feature extraction network to extract the meta-feature representation, and uses a local metric module and a global metric model to constrain and optimize the representation effect of the meta-feature, so as to obtain a better feature representation. At the same time, compared with the recently proposed shallow network model, the model in this paper has improved the final classification accuracy in the small-sample learning image classification problem and demonstrated robustness on different small-sample learning data sets, providing a new feature extraction method for small-sample learning image classification. Brief Description of the Drawings
[0034] Figure 1 is the overall framework diagram of the method of the present invention.
[0035] Figure 2 is the structural diagram of the meta-feature extraction network in the method of the present invention.
[0036] Figure 3 is the structural diagram of the local feature processing network in the method of the present invention.
[0037] Figure 4 It is the structural diagram of the global feature processing network in the method of the present invention. Detailed implementation manners
[0038] The following elaborates on the specific implementation manners of the present invention to further illustrate the starting point of the present invention and the corresponding technical solutions.
[0039] The present invention is a small-sample learning image classification method based on meta-features. The main purpose is to enable the model to extract more distinguishable features from a small number of samples for measuring the similarity between images to achieve the purpose of classification. The overall framework diagram of this method is as Figure 1 shown, which is specifically divided into the following three steps:
[0040] (1) Feature representation extraction and measurement module
[0041] First, the small-sample learning image datasets miniImageNet and tieredImage are used. According to the small-sample learning task N-way, K-shot, the dataset is divided, that is, there are N categories in each task, and each category has K pictures. These image sets are called the support set; and some pictures to be classified, called the query set. The purpose of the task is to accurately classify the pictures to be classified into N categories.
[0042] As Figure 1 shown, the task is a 5-way, K-shot task, and the image samples used are three-channel color pictures with a pixel size of 3×84×84. We need to use a two-layer convolutional network M σ , and the specific structure of the network is as Figure 2 shown. The input of the entire network is the small-sample learning image sample x, and the output is the meta-feature m = M σ (x).
[0043] After the image sample passes through a meta-feature extraction network, a relatively abstract meta-feature m will be obtained. In order to enable the meta-feature to obtain better information in both the local feature and global feature dimensions, this method uses two different ways to measure the similarity between samples after obtaining the meta-feature, referring to the local feature descriptor of the deep nearest neighbor neural network and the prototype calculation of the prototype network respectively. This process is called meta-feature dual measurement.
[0044] First, for the part measured using the local feature descriptor, a local feature processing network f is used to process the meta-feature to obtain the processed feature representation m f = f θ (m) ∈ R c×w×d, where c is the number of channels after convolution, w and d are the length and width of the feature matrix in each channel, and the structure of the processing network is as Figure 3 shown. The entire network consists of two convolutional networks with 64 channels, similar to the meta-feature extraction network, but without the max pooling layer. To make the local information of the sample images easier to measure, we need to obtain as many local feature descriptors as possible during the feature processing. Therefore, the max pooling layer is not used to reduce the dimension of the feature representation, so that more local feature descriptors can be obtained finally. And the local feature descriptors can be considered as obtained by transforming the dimension of the feature representation m f i.e., m f = [h f1 , …, h fF ∈ R c×F (F = wd). That is, for the feature representation m f of a sample, the set composed of its local feature descriptors is F refers to the number of local feature descriptors that a sample can finally obtain. For the N-way K-shot task, there are a total of NF local feature descriptors for one category in the support set. For the query set samples to be classified finally, the feature representation can be obtained. We use the k-nearest neighbor method to measure the similarity between the query set samples and the support set category o, and the specific calculation is as shown in the formula,
[0045]
[0046] where the cosine similarity is used to measure the distance between local feature descriptors, represents the j-th of the k local feature descriptors that are most similar to among all the local feature descriptors of category o. After calculating the similarity between the query set sample and all categories in the support set, the soft classification prediction label for the query set sample is obtained. For all samples in the query set, finally, the cross-entropy loss is used with the true label y of the sample,
[0047]
[0048] where y j represents the true label of the j-th query set, and 1[y j = o] means taking 1 if y j = o, otherwise taking 0.
[0049] Then comes the global feature measurement part. A global feature processing network g is used to process the meta-feature m to obtain the processed feature representation m g = gδ (m), like the processing network and the meta-feature extraction network, is composed of two convolutional networks with 64 channels, including a max-pooling layer for dimensionality reduction, as Figure 4 shown. The Euclidean distance is used to measure the global features. For global features, the more crucial the information contained, the better, that is, the smaller the dimension of the feature representation, the better. If the dimension of the feature representation is large, it may contain a lot of unimportant interference information, thus having a certain impact on the similarity measurement.
[0050] For the few-shot learning task of N-way K-shot, there are N categories in the support set S, and each category contains K samples. For each category, it is necessary to calculate the feature representation representing this category based on the feature representations of each sample, and finally the average method is used to obtain it. For category o, its feature representation is
[0051]
[0052] where S o represents the sample set of category o, |S o | represents the number of samples in category o, and (x i , y i ) represents the feature vector and label of the sample. Then the Euclidean distance function d(·) is used to calculate the distance between two samples in the embedding space, and the distance scores of the query set sample x to each category form a distance score vector d g , that is
[0053] d gi = d(g δ (M σ (x)), c i )
[0054] After that, all the obtained distances are normalized. Since it is desired that the distance from the query set sample to the correct category is as small as possible, a negative sign needs to be added before the distance score during normalization. Finally, the probability P that the sample x belongs to a certain category o is obtained,
[0055]
[0056] Finally, the cross-entropy loss function is used, that is
[0057]
[0058] After obtaining the loss functions of the two parts, the present paper uses the weighted summation method to obtain the final loss function, that is
[0059] L = αL f + βL g
[0060] Both α and β are positive hyperparameters used to adjust the weight ratios of the two feature measurement parts.
[0061] (2) Classification module for performing classification tasks
[0062] During the training process, we used a meta-feature extraction network to obtain meta-feature representations, and then used two different measurement methods to predict the labels of the query set samples. In the loss part, we used weight hyperparameters to balance the ratios of the two measurement parts. In the final label prediction part, the purpose of using meta-features is to increase a little information with global feature measurement and obtain more information in the local feature measurement part. Therefore, we used the similarity measurement of the local feature measurement part as the final measurement criterion, discarded the final distance score of the global feature measurement, and only used the loss part. So, the final predicted label where is the predicted label calculated in the local feature measurement part, and the predicted label in the test process is obtained in the same way. Finally, the performance of the model is evaluated through classification accuracy.
[0063] (3) Experimental results
[0064] For the publicly available few-shot image classification dataset, this method is used for true label estimation to evaluate the effectiveness of this method.
[0065] The information of the dataset used is as follows:
[0066] 1) miniImageNet dataset. It is a classic few-shot learning image classification dataset, which includes 100 categories, 600 images for each category, and each image is an 84×84 color image. According to previous few-shot learning works, we divided the dataset into a training set, a validation set, and a test set, with 64, 16, and 20 categories respectively for each dataset.
[0067] 2) tieredImageNet dataset. This dataset is a larger data set with 608 categories and a total of 779,165 images. Each image is also an 84×84 color image. We divided the dataset into a training set, a validation set, and a test set, with 20 major categories (351 sub-categories), 6 major categories (97 sub-categories), and 8 major categories (160 sub-categories) respectively for each dataset.
[0068] We compared our method with previous excellent methods, namely: Matching Network (MN), Prototypical Network (PN), Model-agnostic meta-learning (MAML), Relation Network (RN), Deep Nearest Neighbor Nwural Network (DN4). Both our model and these methods used a shallow convolutional network with 64 convolutional kernels. We trained the models of all methods in the same experimental environment and then tested them in the same environment. Each test ran for 600 iterations. After 10 tests, the average of the 10 results was taken to obtain the final classification accuracy.
[0069] Tables 1 and 2 respectively show the final test classification accuracies of all methods for four different few-shot learning tasks on the miniImageNet dataset and the tieredImageNet dataset. For each task, the optimal accuracy is shown in bold. It can be seen that except for the 5-way 1-shot task on tieredImageNet, the final results of our method are the best. At the same time, our method can achieve the optimal performance on different datasets, indicating the feasibility and robustness of the present invention in few-shot learning image classification tasks from the experimental results.
[0070]
[0071] Table 1
[0072]
[0073] Table 2
[0074] Throughout the process, weight parameters α and β were involved. To investigate the importance of the two modules, this paper conducted experiments on these two parameters using four few-shot learning tasks on the miniImageNet dataset and the tieredImageNet dataset. The experimental results are shown in Tables 3 and 4.
[0075]
[0076]
[0077] Table 3
[0078]
[0079] Table 4
[0080] In the experiment, we gradually controlled the proportion of the two modules proportionally. At the same time, the original intention of our proposed few-shot learning method based on meta-features is to use the global feature metric method to optimize the local feature metric, and the local feature metric is also used as the standard in the final classification. Therefore, in the proportion experiment, the proportion of the global feature metric was not allowed to exceed that of the local feature metric. On the miniImageNet dataset, as can be seen from the experimental data in Table 3, when α and β are 0.8 and 0.2 respectively, the classification accuracy is the highest in the two few-shot learning tasks of 5-way; when α and β are 0.6 and 0.4 respectively, the classification accuracy is the highest in the two few-shot learning tasks of 10-way. On the tieredImageNet dataset, as can be seen from the experimental data in Table 4.5, when α and β are 0.6 and 0.4 respectively, the classification accuracy is the highest in the two few-shot learning tasks of 5-way; when α and β are 0.7 and 0.3 respectively, the classification accuracy is the highest in the two few-shot learning tasks of 10-way. These experiments can prove that through the meta-feature method, the global feature metric can optimize the local feature metric, thereby obtaining better results. It is proved that through the meta-feature method, the global feature metric can optimize the local feature metric, thereby obtaining better results.
[0081] The above are the specific embodiments of the present invention and the technical principles applied. If changes are made according to the concept of the present invention and the functions and effects generated therefrom do not exceed the spirit covered by the specification and the drawings, they should still fall within the protection scope of the present invention.
Claims
1. A small-sample learning image classification method based on meta-features, characterized in that The small-sample learning image classification method is based on a meta-feature extracted, uses two methods to measure the similarity between samples, and obtains local features and global features respectively; then, through a meta-feature extraction network, the information of local features and global features is fused to improve the accuracy of feature representation; the specific steps are as follows: Step 1: From the small-sample learning image dataset, divide the training set and the test set proportionally, and for the corresponding small-sample learning tasks, randomly sample and construct a large number of tasks, each task containing a support set and a query set; each small-sample learning classification task includes N categories, with K samples in each category; use a two-layer convolutional network M σ , extract features from the picture x to obtain the meta-feature m, m = M σ (x); Step 2: The meta-feature m is processed through a local feature processing network and a global feature processing network respectively 2.1 Process the meta - feature m using a local feature processing network f to obtain the processed local feature representation m f = f θ (m) ∈ R c×w×d , where c is the number of channels after convolution, and w and d are the length and width of the feature matrix in each channel; the set composed of the local feature descriptors of the local feature representation m f is F refers to the number of local feature descriptors that a sample can finally obtain; in each small - sample learning classification task, there are a total of KF local feature descriptors for one category in the support set; for the query - set samples to be classified obtain the local feature representation Use the k - nearest neighbor method to measure the similarity between the query - set sample and the support - set category o Cosine similarity is used to measure the distance between local feature descriptors, among all local feature descriptors representing class o, the j-th one among the k local feature descriptors that are most similar to After calculating the similarities between the query set samples and all classes in the support set, the soft classification prediction label for the query set samples is obtained For all samples in the query set and the true label y of the samples, cross-entropy loss is used, where y j represents the true label of the j-th query set, 1[y j = o] means that when y j = o, take 1, otherwise take 0; 2.2 Use a global feature processing network g to process the meta-feature m to obtain the processed global feature representation m g = g δ (m); for each category, calculate the global feature representation representing the category according to the global feature representation of each sample, and finally take the average; for category o, its global feature representation is Among them, S o represents the sample set of class o, and |S o | represents the number of samples in class o. (x i , y i ) represents the feature vector and label of the sample; Then, the Euclidean distance function d(·) is used to obtain the distance between two samples in the embedding space, and the distance scores of the query set samples to class i form a distance score vector d g , that is Among them, c i After representing the global feature representation of class i, perform a normalization operation on all the obtained distances to obtain the probability P that a query set sample belongs to a certain class o. Use the cross-entropy loss function, that is Step 3: According to the loss functions of local feature representation and global feature representation, use the weighted summation method to obtain the final loss function, that is L = αL f + βL g Both α and β are positive hyperparameters used to adjust the weight ratios of the two feature measurement parts; the final predicted label
Citation Information
Patent Citations
Hyperspectral image classification method based on self-paced learning double-flow multi-scale dense connection network
CN112733659A
Apparatus and method for training classification model and apparatus for classifying with classification model
EP3699813A1