Method for Applying Meta-Learning Based on Transfer Learning and Attention Mechanism to Few-Shot Image Classification
Through DenseNet network pre-training and attention mechanism correction features, combined with transfer learning and parameter adjustment in meta-learning stage, the problems of model overfitting and catastrophic forgetting in small sample picture classification are solved, achieving higher classification accuracy.
Patent Information
- Application Number
- CN202111615640.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-27
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2041-12-27
AI Technical Summary
The prior art requires a large amount of data to achieve high accuracy in small sample picture classification, and traditional meta-learning methods ignore the attention mechanism, leading to the problems of model overfitting and catastrophic forgetting.
DenseNet network is used for pre-training, parameters are frozen, and feature correction is performed in combination with attention mechanism. Through parameter adjustment in transfer learning and meta-learning stages, network parameter updates are reduced, prior knowledge is used to train large-scale data, and scaling and translation parameters are introduced for feature weighting, and valueless features are eliminated.
Improve the accuracy of small sample picture classification, avoid model overfitting and catastrophic forgetting, and improve classification accuracy under new tasks.
Smart Images

Figure CN114492581B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep learning image classification, and particularly relates to a method for applying transfer learning and attention mechanism meta-learning to few-shot image classification. Background Art
[0002] Deep learning has achieved great success in many fields. For example, the results obtained in object detection, image classification, semantic segmentation, etc. can exceed those of humans. However, they usually require a large amount of data to achieve relatively high accuracy, and the cost of collecting and annotating a large amount of data is also very high. Humans can generalize the concept of "lion" from an illustration in a book. Therefore, enabling machines to "generalize" the concept of a certain object from a small number of samples has attracted the attention of a large number of researchers. Learning from a small amount of data is a challenge faced by machine vision. In recent years, meta-learning has shown good performance in improving few-shot learning in machine vision.
[0003] Meta-learning is "learning to learn". Different from traditional machine learning methods, traditional machine learning methods use fixed learning algorithms to solve given tasks from scratch. The purpose of meta-learning is to improve the learning algorithm itself, obtain experience in multiple learning tasks, usually covering the distribution of related tasks, and use this experience to improve future learning performance. Meta-learning is a task-level learning method aimed at accumulating experience by learning multiple tasks, while the base learner focuses on modeling the data distribution of a single task. A representative in this regard is model-agnostic meta-learning (MAML), which learns to search for the optimal initialization state to quickly adapt to new tasks of the base learner. Its task-agnostic property makes it possible to be extended to few-shot supervised learning and unsupervised reinforcement learning. However, this method has limitations. Each task is usually modeled by a low-complexity base learner (such as a shallow neural network) to avoid model overfitting, thus unable to use deeper and more powerful network architectures. And existing meta-learning methods usually ignore the existence of the attention mechanism, which has been proven to be important in the process of cognition and learning.
[0004] As the research on deep networks continues to deepen, a new problem has emerged: as information about the input or gradients passes through many layers, it may disappear or "wash out" by the time it reaches the end or start of the network. In recent years, many methods have been proposed to solve this problem. For example, ResNets connect from one layer to the next through skip connections, but the contributions of many layers are small in practice and can be randomly discarded during training. This makes the state of ResNets similar to an unfolded recurrent neural network, but its number of parameters is large because each layer has its own weights. In this paper, DenseNet is used as a feature extractor. DenseNet clearly distinguishes the information added to the network from the information retained. DenseNet layers are very narrow (e.g., 12 filters per layer), adding only a small group of feature maps to the "collective knowledge" of the network and keeping the remaining feature maps unchanged - the final classifier makes decisions based on all the feature maps in the network. Another major advantage of DenseNet is its improved information flow and gradients across the entire network, which makes them easy to train. Each layer has direct access to the gradients from the loss function and the original input signal, resulting in an implicit deep supervision.
[0005] In recent years, attention mechanisms have also been widely applied in computer vision systems and machine translation. The attention mechanism in neural networks is a resource allocation scheme that allocates computational resources to more important tasks while solving the problem of information overload when computational power is limited. In neural network learning, generally speaking, the more parameters a model has, the stronger its expressive power and the larger the amount of information it stores, but this will bring the problem of information overload. Then, by introducing the attention mechanism, focusing on the information that is more crucial for the current task among numerous input information, reducing the attention to other information, and even filtering out irrelevant information, the problem of information overload can be solved, and the efficiency and accuracy of task processing can be improved. So, what is the attention mechanism? When we look at a scene, we must be looking at a certain place in the scene, and when our vision moves, the attention also moves with the movement of our eyes. That is to say, when a person pays attention to a certain scene, the attention distribution in each space within the scene is inconsistent. Therefore, we can draw on the attention mechanism of the human brain and only select some key information inputs for processing to improve the efficiency of the neural network. Summary of the Invention
[0006] For the above-mentioned problems, in order to solve the problem of too few training samples, the present invention uses the DenseNet network for pre-training to extract features. After the pre-training, the DenseNet network parameters are frozen. The network parameters trained with a large amount of data can ensure that they have good initialization, and DenseNet will use multiple feature reuse to improve the utilization rate of fewer features. The attention mechanism is used to perform channel weighting on the features extracted by the feature extractor. The corrected features can retain valuable features and eliminate worthless features. In the meta-learning stage, displacement and deviation are used to perform meta-learning of parameters, which reduces the parameters of the network and avoids the problem of catastrophic forgetting.
[0007] A method for classifying small sample images based on transfer learning and attention mechanism meta-learning, comprising the following steps:
[0008] (1) Obtaining data: Read the pre-trained images in the dataset. The images are divided by task. Images of different tasks are in different folders. Images are read according to the task distribution.
[0009] (2) Construction of the transfer learning and attention mechanism meta-learning network framework: including a fixed feature extractor and different category output layers due to the different number of classification tasks in the pre-training and meta-learning stages;
[0010] (2.1) The model framework of the pre-training stage includes: using the DenseNet network as a feature extractor to extract the features of the input image, followed by an average pooling layer to reduce the dimension of the features extracted by the DenseNet network, remove redundant information, flatten the pooled features, followed by a fully connected layer, and finally a category output layer determined according to the classification task;
[0011] (2.2) The model framework of the meta-learning stage includes: using the DenseNet network as a feature extractor to extract the features of the input image, and inputting the extracted image features into the attention mechanism module. The attention mechanism uses channel attention, and performs global average pooling on the feature map of each channel to obtain the attention weighted value. This weighted value is then applied to the original feature map to weight the values of each channel. The weighted features are flattened, followed by a fully connected layer, and finally a category output layer different from that in the training stage;
[0012] (3) training the model framework of the pre-training phase and the model framework of the meta-learning phase constructed in step (2);
[0013] (3.1) Initialize the pre-trained network parameters, input the training data set into the pre-trained network framework to optimize the parameters of the pre-trained network framework, learn the network parameter weights W and biases b in the convolutional layer, reduce the feature distribution between the same tasks through the cross-entropy loss function, and finally use softmax to calculate the category with the highest probability of belonging to the sample, which is the predicted category in the pre-training stage of the picture;
[0014] (3.2) Update the parameters of the pre-trained network;
[0015] (3.3) Repeat steps (3.1) and (3.2) until the number of network iterations reaches the preset number of iterations, and take the network parameters Θ(W, b) of the iteration with the best accuracy;
[0016] (3.4) The pre-training ends, and the network parameters Θ(W, b) are fixed and no longer updated;
[0017] (3.5) Initialize the meta-learning network parameters, input the test data into the meta-learning network framework. At this time, the weights W and biases b used in the convolutional layer of the feature extractor in the network are the parameters with the best iteration accuracy in the pre-training stage, and introduce two new parameters: scaling and translation;
[0018] (3.6) Update the meta-learning network parameters;
[0019] (3.7) Repeat steps (3.5) and (3.6) until the number of network iterations reaches the preset number of iterations, and take the network parameters of the iteration with the best accuracy;
[0020] (3.8) After the meta-learning stage, use the validation data set to verify the network model, and the final classification accuracy output by the network is the final model evaluation accuracy.
[0021] Further, the step (1) includes the following steps:
[0022] (1.1) To increase the difficulty of classification, the size of the pictures in the data set is 84×84. Resize the picture size to 40×40, and then randomly cut pictures of 36×36 size; each picture is converted to three RGB channels and converted into a three-dimensional matrix of c×h×w, where h and w are the height and width of the image respectively, and c is the number of channels;
[0023] (1.2) Convert the training pictures into four-dimensional matrix data of n S ×c×h×w, where n S represents the number of training samples of task T; randomly select un-trained pictures in the same task as validation data and convert them into four-dimensional matrix data of n T ×c×h×w, where n T represents the number of data samples used for validation in the same task;
[0024] (1.3) The picture category is encoded in one-hot. If there are N categories of pictures in total, the label of the first category is represented as [1, 0, 0, ..., 0] 1×N , and the label of the second category is represented as [0, 1, 0, ..., 0] 1×N , …, the label of the Nth category of pictures is represented as [0, 0, 0, ..., 1] 1×N .
[0025] Furthermore, the step (2) includes the following steps:
[0026] Training of the feature extractor: The DenseNet network is pre-trained with the training dataset to obtain the DenseNet densely connected network model parameters. The input of the densely connected network is a matrix of batch_size×c×h×w, and the size of batch_size depends on the computer memory;
[0027] In the pre-training stage: The densely connected network is used to extract features. The extracted features have 342 channels. Then, they are flattened. To reduce the dimension after flattening, global pooling is performed on the features of each channel to obtain a 342×1×1 vector. After flattening, a fully connected layer is connected. The number of neurons in the fully connected layer is fixed at 600. The last layer is the output layer, and the softmax function is used to calculate the probability of each category to which the picture belongs;
[0028] In the meta-learning stage: The densely connected network is used for feature extraction. The extracted features have 342 channels. The extracted features are fed into the attention mechanism module. The shape of the feature map is (128, 10, 10, 342), where 128, 10, 10, and 342 are the batch-size, width, height, and number of channels of the feature map respectively; Then, they are flattened. To reduce the dimension after flattening, global pooling is performed on the features of each channel to obtain a 342×1×1 vector. After flattening, a fully connected layer is connected. The number of neurons in the fully connected layer is fixed at 600. The last layer is the output layer, and the softmax function is used to calculate the probability of each category to which the picture belongs.
[0029] Furthermore, the step (3.2) includes the following steps:
[0030] Initializing network parameters, pre-training part: Input the pictures of the training dataset into the DenseNet densely connected network. In the DenseNet densely connected network, there is a stage of feature reuse. Therefore, the lth layer of the network framework receives the feature maps of all previous layers, x0…x l-1 , input:
[0031] x l =Hl ([x0,x1,...,x l-1 ) (1)
[0032] where [x0,x1,...,x l-1 represents the concatenation of all the feature maps generated in layers 0 to l-1, and H l represents the tensor formed after concatenation;
[0033] For ease of implementation, multiple inputs are concatenated into a tensor; at this stage, data from other datasets or domain adaptation is not considered, and pre-training is performed on the available few-shot learning benchmark data; specifically, for a specific few-shot dataset, all class data D is merged for pre-training.
[0034] First, a feature extractor Θ and an auxiliary classifier θ are randomly initialized, and then they are optimized by gradient descent,
[0035]
[0036] where L is defined as the empirical loss, α is the learning rate, and the learning rate α is set to 0.01,
[0037]
[0038] At this stage, the feature extractor Θ is learned; the parameters it learns will be frozen in the subsequent meta-training and meta-testing stages; the learned auxiliary classifier θ will be discarded because the subsequent few-shot tasks have different classification objectives.
[0039] Furthermore, the two new parameters introduced in step (3.5): scaling and translation, denoted as and perform weighted scaling on the weights W in the feature extractor, perform weighted translation on the biases b; the image features extracted by the feature extractor are input into the attention mechanism layer, channel weighting is performed on the extracted features, and then the weighted features are flattened, fully connected, and finally the class with the highest probability of the sample is calculated using softmax as the predicted class in the image meta-learning stage.
[0040] Furthermore, step (3.6) includes the following steps:
[0041] For a given task T, the current base learner, that is, the classifier θ′, is optimized by gradient descent using the loss of the training data in task T:
[0042]
[0043] where β is the learning rate, and the learning rate β is set to 0.01, Indicates performing a gradient operation on the subsequent expression, which is the empirical loss of the training task T; corresponds to a classifier that only works in the current task;
[0044] Initialize Then optimize it using the test loss of the test images in task T:
[0045]
[0046] Let be the learning rate, and set the learning rate to 0.0001, Indicates performing a gradient operation on the subsequent meta-learning network parameters and is the empirical loss of the test task T;
[0047]
[0048] In this step of update, the same learning rate in Equation (5) will be used to update θ′ At this time, θ′ in Equation (6) is the final trained base learner obtained by training on the test training images in Equation (5).
[0049] Furthermore, the meta-learning network parameters will be updated relying on W and b fixed in the training phase in step (3.4), and the specific steps are as follows:
[0050] For the already trained feature extractor Θ, for the l-th layer containing K neurons, there are K pairs of parameters, namely the weights and biases, denoted as {(W i,k , b i,k )}; after training, K pairs of scalars are obtained Assume M is the input, and apply to the weights and biases through Equation (7):
[0051]
[0052] where ⊙ represents element-wise multiplication.
[0053] Compared with the prior art, the beneficial effects of the present invention are:
[0054] The present invention uses the method of meta - learning to deal with the problem of few - shot classification. For few - shot classification, the number of samples is small, and the available features are relatively few. The DenseNet densely - connected network is used as a feature extractor. Feature reuse in the network framework makes full use of the feature maps extracted after each step of convolution, and fully utilizes the limited features. At the same time, in the pre - training stage, by training network parameters with a large - scale data set of pictures, better network parameters and prior knowledge can be obtained. In the meta - learning stage, parameters of scaling and translation are introduced, and these two parameters are updated through the training set, reducing the number of parameters to be trained in the network. Meanwhile, prior knowledge is fully utilized, and the problem of catastrophic forgetting will not occur. Convolution for feature extraction is very important. The introduction of the attention mechanism can correct features. After correction, valuable features are retained and valueless features are removed, improving the classification accuracy of the model and enhancing the utilization rate of few - shot classification features, so as to more accurately predict the category of pictures by the model in the case of new tasks. Brief Description of the Drawings
[0055] Figure 1 It is the structure diagram of the transfer learning and attention - mechanism - based meta - learning network of the present invention.
[0056] Figure 2 It is the rising curve of the cross - validation experiment accuracy on the miniImageNet data set of the present invention. Detailed Embodiment
[0057] The present invention will be further described below in conjunction with the accompanying drawings and specific experiments.
[0058] It is generally recognized that the number of samples for few - shot classification is small, and the characteristics of meta - learning can well solve this problem. Specifically, the number of samples used for training is small, and using a conventional model will lead to the problem of overfitting. Since the present invention uses a large - scale network framework, it is obviously unrealistic to use the method of fine - tuning all parameters. The fundamental motivation is to use the method of meta - learning to reduce the number of updated parameters of the model framework for few - shot classification.
[0059] A method for solving few - shot image classification based on transfer learning and attention - mechanism - based meta - learning of the present invention includes the following steps:
[0060] A method for applying transfer learning and attention - mechanism - based meta - learning to few - shot image classification includes the following steps:
[0061] (1) Obtain data: Read the pictures in the pre - training of the data set. The pictures are divided by task, and pictures of different tasks are in different folders, and the pictures are read according to the task distribution.
[0062] (1.1) To increase the difficulty of classification, the images in the dataset are resized from 84×84 to 40×40, and then randomly cropped into images of size 36×36; each image is converted into an RGB three-channel and transformed into a three-dimensional matrix of c×h×w, where h and w are the height and width of the image respectively, and c is the number of channels;
[0063] (1.2) The training images are converted into four-dimensional matrix data of n S ×c×h×w, where n S represents the number of training samples for task T; in the same task, randomly selected un-trained images are used as validation data and converted into four-dimensional matrix data of n T ×c×h×w, where n T represents the number of data samples used for validation in the same task;
[0064] (1.3) The image categories are encoded using one-hot encoding. If there are N categories of images, the label of the first category is [1, 0, 0,..., 0] 1×N , the label of the second category is [0, 1, 0,..., 0] 1×N , …, the label of the Nth category of images is [0, 0, 0,..., 1] 1×N .
[0065] (2) Construction of the transfer learning and attention mechanism meta-learning network framework: including a fixed feature extractor and different category output layers adopted due to the different numbers of classification tasks in the pre-training and meta-learning stages;
[0066] Training of the feature extractor: The DenseNet network is pre-trained using the training dataset to obtain the model parameters of the DenseNet densely connected network. The input of the densely connected network is a matrix of batch_size×c×h×w, and the size of batch_size depends on the computer memory;
[0067] In the pre-training stage: The densely connected network is used to extract features. The extracted features have 342 channels, and then they are flattened. To reduce the dimension after flattening, global pooling is performed on the features of each channel to obtain a vector of 342×1×1. After flattening, a fully connected layer is connected. The number of neurons in the fully connected layer is fixed at 600, and the last layer is the output layer, and the softmax function is used to calculate the probability of each category to which the image belongs;
[0068] The meta - learning stage: A densely connected network is used for feature extraction. The extracted features have 342 channels. The extracted features are fed into the attention mechanism module. The shape of the feature map is (128, 10, 10, 342), where 128, 10, 10, and 342 are the batch - size, width, height, and number of channels of the feature map respectively. The setting of the attention mechanism can be understood as follows: Use some networks to calculate a weight, operate this weight with the feature map, and change this feature map to obtain a feature map with enhanced attention. Convolution for feature extraction is very important. The attention mechanism can correct the features. The corrected features can retain valuable features and eliminate worthless features. Then flattening is performed. To reduce the dimension after flattening, global pooling is performed on the features of each channel to obtain a 342×1×1 vector. After flattening, a fully - connected layer is connected. The number of neurons in the fully - connected layer is fixed at 600. The last layer is the output layer, and the softmax function is used to calculate the probability of each category to which the image belongs.
[0069] (3) Train the model frameworks of the pre - training stage and the meta - learning stage built in step (2);
[0070] (3.1) Initialize the parameters of the pre - trained network. Input the training data set into the pre - trained network framework to optimize the parameters of the pre - trained network framework, learn the network parameter weights W and biases b in the convolutional layer, reduce the feature distribution between the same tasks through the cross - entropy loss function, and finally use softmax to calculate the category with the highest probability for the sample, which is the predicted category in the pre - training stage of the image;
[0071] (3.2) Update the parameters of the pre - trained network;
[0072] Initialize the network parameters. Pre - training part: Input the training data set images into the DenseNet densely connected network. In the DenseNet densely connected network, there is a stage of feature reuse. Therefore, the l - th layer of the network framework receives the feature maps of all previous layers, x0…x l-1 , Input:
[0073] x l =H l ([x0,x1,...,x l-1 ) (1)
[0074] where [x0,x1,...,x l-1 represents the concatenation of all feature maps generated in the 0 - th layer to the (l - 1) - th layer, and H l represents the tensor formed after concatenation;
[0075] For ease of implementation, multiple inputs are concatenated into a tensor; at this stage, data from other datasets or domain adaptation is not considered, and pre-training is performed on the available few-shot learning benchmark data; specifically, for a specific few-shot dataset, all class data D is combined for pre-training. For example, for miniImageNet, there are a total of 64 classes in the training split of the dataset D, with each class containing 600 samples, which are used to pre-train a 64-class classifier.
[0076] First, a feature extractor Θ and an auxiliary classifier θ are randomly initialized and then optimized via gradient descent.
[0077]
[0078] where L is defined as the empirical loss and α is the learning rate, and the learning rate α is set to 0.01.
[0079]
[0080] At this stage, the feature extractor Θ is learned; the parameters it learns will be frozen in the subsequent meta-training and meta-testing stages; the learned auxiliary classifier θ will be discarded because the subsequent few-shot tasks have different classification objectives.
[0081] (3.3) Repeat steps (3.1) and (3.2) until the number of network iterations reaches the preset number of iterations, and take the network parameters Θ(W, b) of the iteration with the best accuracy.
[0082] (3.4) The pre-training is completed, and the network parameters Θ(W, b) are fixed and no longer updated.
[0083] (3.5) Initialize the meta-learning network parameters, input the test data into the meta-learning network framework. At this time, the weights W and biases b used in the convolutional layer of the feature extractor are the parameters with the best iteration accuracy in the pre-training stage, and two new parameters are introduced: scaling and translation.
[0084] The two new parameters: scaling and translation, are denoted as and The weights W in the feature extractor are weighted and scaled, and the biases b are weighted and translated; the image features extracted by the feature extractor are input into the attention mechanism layer, the extracted features are channel-weighted, then the weighted features are flattened and fully connected, and finally the class with the highest probability of the sample is calculated using softmax, which is the predicted class in the meta-learning stage of the image.
[0085] (3.6) Update the meta-learning network parameters.
[0086] The step (3.6) includes the following steps:
[0087] For a given task T, optimize the current base learner, that is, the classifier θ′, by using the loss of the training data in task T through gradient descent:
[0088]
[0089] where β is the learning rate, and the learning rate β is set to 0.01. denotes performing a gradient operation on the subsequent expression. is the empirical loss of training task T. corresponds to the classifier that only works in the current task.
[0090] This is different from formula (2). Here, the feature extractor Θ is not updated. It should be noted that the classifier here is different from the large-scale auxiliary classifier θ in the previous stage, that is, in formula (2); this classifier is less than the large-scale classifier and classifies sample pictures in a new few-shot scenario; corresponding to the classifier that only works in the current task and optimized for the previous task Initialization;
[0091] Initialization Then optimize it by using the test loss of the test pictures in task T:
[0092]
[0093] is the learning rate, and set the learning rate to 0.0001. denotes performing a gradient operation on the subsequent meta-learning network parameters and is the empirical loss of test task T.
[0094]
[0095] In this step update, the same learning rate in formula (5) will be used to update θ′. At this time, θ′ in formula (6) is finally trained from the base learner after training on the test training pictures in formula (5).
[0096] The meta-learning network parameters will be updated relying on the fixed W and b in step (3.4) during the training phase. The specific steps are as follows:
[0097] For the already trained feature extractor Θ, there are K pairs of parameters for the l-th layer containing K neurons, which are the weights and biases respectively, denoted as {(Wi,k , b i,k )}; After training, K pairs of scalars are obtained Assume M is the input, and Apply it to the weights and biases through formula (7):
[0098]
[0099] where ⊙ represents element-wise multiplication.
[0100] (3.7) Repeat steps (3.5) and (3.6) until the number of network iterations reaches the preset number of iterations, and take the network parameters of the iteration with the best accuracy;
[0101] (3.8) After the meta-learning stage, use the validation dataset to verify the network model, and the classification accuracy finally output by the network is the final model evaluation accuracy.
[0102] The present invention can be further illustrated by the following experiments:
[0103] To verify the effectiveness of the present invention, experiments were conducted on the Omniglo, miniImageNet, and FC100 datasets respectively.
[0104] To demonstrate the multi-task nature of meta-learning, the dataset is divided into a training set, a validation set, and a test set.
[0105] Since Omniglot is a much simpler dataset than MiniImagenet, existing meta-learning methods can easily achieve an accuracy of over 95% on most test tasks generated on Omniglot. Therefore, we only test the TML method on Omniglot. Similar to the experiment on Miniimagenet, we also train the meta-learner on 200,000 randomly generated tasks and set the learning rate to 0.001. The experimental results are shown in Table 1. It can be seen that the proposed method TML achieves relatively advanced performance in few-shot image classification tasks.
[0106] miniimagenet was proposed by Vinyalset for few-shot learning evaluation. Due to the use of ImageNet images, it has a high complexity, but compared with running on the complete ImageNet dataset, it requires fewer resources and infrastructure. There are a total of 100 categories, and each category has 600 84×84 color picture samples. These 100 categories are divided into 64, 16, and 20 categories for sampling tasks of meta-training, meta-validation, and meta-testing respectively, and related work is carried out.
[0107] Fewshot-CIFAR100 (FC100) is based on the currently popular object classification dataset CIFAR100. It provides a more challenging scenario with lower image resolution and a more challenging meta-training / test split (separated according to object superclasses). It contains 100 object classes, with 600 sample images of 32×32 for each class. These 100 classes belong to 20 superclasses. The meta-training data comes from 60 classes belonging to 12 superclasses. The meta-validation and meta-test sets contain 20 classes, belonging to 4 superclasses respectively. These partitions conform to the superclasses, thus minimizing the information overlap between the training, validation, and test tasks.
[0108] Train a large-scale deep neural network model on all training data points and stop training after 100 iterations. We use the same task sampling method as in the related work. Specifically, 1) consider 5-way classification; 2) sample tasks with 1-shot or 5-shot for 5 classes, including 1 or 5 samples for training and 15 (uniform) samples for testing. A total of 8k tasks are sampled for meta-training, and 600 random tasks are sampled for meta-validation and meta-test respectively.
[0109] Table 1 Experimental Accuracy on Omniglot Dataset
[0110]
[0111] Table 2: Experimental Accuracy on FC100 Dataset
[0112]
[0113] Table 3: Experimental Accuracy on miniImageNet Dataset
[0114]
[0115]
[0116] Table 4: Cross-Experiment Results
[0117]
[0118] The method of the present invention learns prior knowledge from a large amount of training data. In the case of only using a small amount of labeled training data, it can help the deep neural network converge faster and at the same time reduce the possibility of network overfitting. This method uses a DenseNet network as a feature extractor. The difficulty of the few-shot classification task is the small number of samples. The feature extractor network adopted by this method uses the method of feature reuse to make full use of the limited pictures. In the pre-training stage, a dense network is used to train a large amount of data to train the weights and biases of the network. The features finally extracted by the feature extractor are flattened, followed by a fully connected layer and a classification layer. At this time, 64-class classification is performed. Since the amount of data in the pre-training stage is relatively large, the network parameters trained are also relatively good. After the pre-training is completed, the trained weights and biases are fixed, and the subsequent classifier is modified for the next meta-learning stage. In the meta-learning stage, the prior knowledge learned in the pre-training stage is used to scale the weights and translate the biases in the network for meta-learning. Only these two parameters are updated, and the weights and biases are not updated. Training with a large amount of data provides a good initialization for the weights of the deep network, enabling meta-learning to converge quickly under fewer tasks. These operations keep the weights of the trained deep network unchanged, thus avoiding the problem of catastrophic forgetting, and thereby improving the classification accuracy of the image dataset.
[0119] It should be understood that the present invention is not limited to the above specific examples. On the basis of the present invention, those skilled in the art can make equivalent deformations or substitutions without departing from the spirit of the present invention. These equivalent variations or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for small-sample image classification based on transfer learning and attention mechanism meta-learning, characterized in that, The steps include the following: (1) Obtain data: Read the pictures in the pre-training of the dataset. The pictures are divided by tasks, and the pictures of different tasks are in different folders. Read the pictures according to the task distribution; (2) Build the transfer learning and attention mechanism meta-learning network framework: It includes a fixed feature extractor and different category output layers adopted due to the different numbers of classification tasks in the pre-training and meta-learning stages; (2.1) The model framework in the pre-training stage includes: Use the DenseNet network as the feature extractor to extract the features of the input pictures, followed by an average pooling layer to reduce the dimension of the features extracted by the DenseNet network, remove redundant information, flatten the pooled features, followed by a fully connected layer, and finally a category output layer determined according to the classification task; (2.2) The model framework in the meta-learning stage includes: Use the DenseNet network as the feature extractor to extract the features of the input pictures, input the extracted picture features into the attention mechanism module. The attention mechanism uses channel attention, perform global average pooling on the feature map of each channel to obtain the attention weighting value, then apply this weighting value to the original feature map, weight the values of each channel, flatten the weighted features, followed by a fully connected layer, and finally a category output layer different from the training stage; (3) Train the model frameworks in the pre-training stage and the meta-learning stage built in step (2); (3.1) Initialize the pre-training network parameters, input the training dataset into the pre-training network framework to optimize the parameters of the pre-training network framework, learn the network parameter weights W and biases b in the convolutional layer, reduce the feature distribution between the same tasks through the cross-entropy loss function, and finally use softmax to calculate the category with the highest probability of the sample belonging to as the predicted category in the picture pre-training stage; (3.2) Update the parameters of the pre-training network; (3.3) Repeat steps (3.1) and (3.2) until the network iteration times reach the preset iteration times, and take the network parameters Θ(W, b) of the iteration times with the best accuracy; (3.4) After the pre-training ends, the network parameters Θ(W, b) are fixed and no longer updated; (3.5) Initialize the meta-learning network parameters, input the test data into the meta-learning network framework. At this time, the weights W and biases b used in the convolutional layer of the feature extractor in the network are the parameters with the best iteration accuracy in the pre-training stage, and introduce two new parameters: scaling and translation; (3.6) Update the meta-learning network parameters; (3.7) Repeat steps (3.5) and (3.6) until the network iteration times reach the preset iteration times, and take the network parameters of the iteration times with the best accuracy; (3.8) After the meta-learning stage ends, use the validation dataset to validate the network model, and the final classification accuracy output by the network is the final model evaluation accuracy.
2. The method for applying meta-learning based on transfer learning and attention mechanism to few-shot image classification according to claim 1, wherein, The said step (1) includes the following steps: (1.1) To increase the difficulty of classification, the images in the dataset are sized 84×84, resized to 40×40, and then randomly cropped into images of size 36×36; each image is converted to three RGB channels and transformed into a three-dimensional matrix of c×h×w, where h and w are the height and width of the image respectively, and c is the number of channels; (1.2) Training images are converted into four-dimensional matrix data of n S ×c×h×w, where n S represents the number of training samples for task T; randomly selected un-trained images in the same task are used as validation data and converted into four-dimensional matrix data of n T ×c×h×w, where n T represents the number of data samples used for validation in the same task; (1.3) The image categories are encoded using one-hot encoding. If there are N categories of images, the label for the first category is [1, 0, 0, ..., 0] 1×N , and the label for the second category is [0, 1, 0, ..., 0] 1×N , …, the label for the Nth category of images is [0, 0, 0, ..., 1] 1×N .
3. The method for applying meta - learning based on transfer learning and attention mechanism to few - shot image classification according to claim 1, characterized in that, Step (2) includes the following steps: Training of the feature extractor: The DenseNet network is pre-trained using the training dataset to obtain the DenseNet densely connected network model parameters, where the input to the densely connected network is a matrix of batch_size×c×h×w, and the size of batch_size depends on the computer memory; In the pre-training stage: The densely connected network is used to extract features. The extracted features have 342 channels and are then flattened. To reduce the dimension after flattening, global pooling is performed on the features of each channel to obtain a vector of 342×1×1. After flattening, a fully connected layer is connected. The number of neurons in the fully connected layer is fixed at 600. The last layer is the output layer, and the softmax function is used to calculate the probability of each category to which the image belongs; In the meta-learning stage: The densely connected network is used for feature extraction. The extracted features have 342 channels. The extracted features are fed into the attention mechanism module. The shape of the feature map is (128, 10, 10, 342), where 128, 10, 10, and 342 are the batch-size, width, height, and number of channels of the feature map respectively; then it is flattened. To reduce the dimension after flattening, global pooling is performed on the features of each channel to obtain a vector of 342×1×1. After flattening, a fully connected layer is connected. The number of neurons in the fully connected layer is fixed at 600. The last layer is the output layer, and the softmax function is used to calculate the probability of each category to which the image belongs.
4. A method for applying meta - learning based on transfer learning and attention mechanism to few - shot image classification according to claim 1, characterized in that, Step (3.2) includes the following steps: Initialize network parameters, pre-training part: Input the training dataset images into the DenseNet densely connected network. In the DenseNet densely connected network, there will be a stage of feature reuse. Therefore, the l-th layer of the network framework receives the feature maps of all previous layers, x0…x l-1 , Input: x l = H l ([x0,x1,...,x l-1 ) (1) where [x0, x1,..., x l-1 represents the concatenation of all feature maps generated in the 0th to l-1th layers, and H l represents the tensor formed after concatenation; For ease of implementation, multiple inputs are concatenated into a tensor; at this stage, data or domain adaptation from other datasets is not considered, and pre-training is performed on the available few-shot learning benchmark data; specifically, for a specific few-shot dataset, all class data D is merged for pre-training; First, a feature extractor Θ and an auxiliary classifier θ are randomly initialized, and then they are optimized by gradient descent, where L is defined as the empirical loss, α is the learning rate, and the learning rate α is set to 0.01, At this stage, the feature extractor Θ is learned; the parameters it learns will be frozen in the subsequent meta-training and meta-testing stages; the learned auxiliary classifier θ will be discarded because the subsequent few-shot tasks have different classification targets.
5. A method for applying meta - learning based on transfer learning and attention mechanism to few - shot image classification according to claim 1, characterized in that, The two new parameters introduced in step (3.5): scaling and translation, denoted as and perform weighted scaling on the weights W in the feature extractor, perform weighted translation on the bias b; input the image features extracted by the feature extractor into the attention mechanism layer, perform channel weighting on the extracted features, then flatten and fully connect the weighted features, and finally use softmax to calculate the class with the highest probability of the sample belonging to, which is the predicted class in the meta-learning stage of the image.
6. The method for applying meta - learning based on transfer learning and attention mechanism to few - shot image classification according to claim 5, wherein, Step (3.6) includes the following steps: For a given task T, the current base learner, that is, the classifier θ′, is optimized by gradient descent using the loss of the training data in task T: where β is the learning rate, and the learning rate β is set to 0.01, denotes the gradient operation on the subsequent expression, L T(tr) is the empirical loss of the training task T; corresponds to the classifier that only works in the current task; Initialization Then, optimize it using the loss of the test image in task T for damage testing: is the learning rate, set the learning rate to be 0.0001, indicating the gradient operation on the subsequent meta-learning network parameters is performed, is the empirical loss of the test task T; In this step update, the same learning rate in Equation (5) will be used to update θ'. At this time, θ' in Equation (6) is finally trained from the base learner trained on the test training images in Equation (5).
7. A method for applying meta - learning based on transfer learning and attention mechanism to few - shot image classification according to claim 6, characterized in that, The meta-learning network parameters It will be updated relying on the fixed W and b in step (3.4) during the training phase. The specific steps are as follows: The trained feature extractor Θ has K pairs of parameters for the l-th layer containing K neurons, namely the weights and biases, denoted as {(W i,k , b i,k )); After training, K pairs of scalars are obtained Suppose M is the input, and is applied to the weights and biases through Equation (7): where ⊙ represents element-wise multiplication.