Few-shot Learning Method and Device Based on Context Memory and Fine-grained Calibration
By adopting a method based on context memory and fine-grained calibration in small sample learning, we learn relationship embedding from the perspective of global and local features, the problems of insufficient generalization ability and difficulty in extracting subtle differences in the existing technology are solved, and higher accuracy and better generalization ability are achieved.
Patent Information
- Application Number
- CN202011128422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2040-10-20
AI Technical Summary
Existing small sample learning methods have difficulty providing strong generalization capabilities when dealing with multiple categories of data sets and lack the ability to extract subtle differences from a local perspective.
A small sample learning method based on context memory and fine-grained calibration is adopted to learn discriminant relationship embeddings from the global feature perspective through a class-sensitive context memory network, and predict the global query-to-class similarity; then, the rough relationship embedding predicted by the global stage is supplemented and improved from the local feature perspective to obtain a more accurate local query-to-class similarity evaluation.
It significantly reduces classification errors, improves accuracy and generalization capabilities in the new tasks, and can more effectively integrate the global and local feature information of the sample.
Smart Images

Figure CN112308123B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a few-shot learning method, in particular to a few-shot learning method based on context memory and fine-grained calibration, and also relates to a corresponding few-shot learning device, belonging to the technical field of machine learning. Background Art
[0002] Data is an important resource in the field of machine learning. How to train a model in the case of lack of data? Few-shot learning is an effective solution. Few-shot learning refers to the study of how to extract effective concepts from one or a few limited samples under the condition of sparse samples (each class may have only one or a few limited samples), so that the model can quickly adapt to these new unseen classes.
[0003] In recent years, people have successively proposed a variety of few-shot learning methods. These methods can be roughly divided into few-shot learning methods based on optimization, few-shot learning methods based on generation, and few-shot learning methods based on metric. In the few-shot learning method based on optimization, first train the required meta-learner on a series of learning tasks and continuously optimize it to obtain the best performance on the target task distribution (including potential unseen tasks). In the few-shot learning method based on generation, try to use the generative meta-learner to augment the few-shot data or learn the classification weights for predicting new classes. In the few-shot learning method based on metric, classification is performed by comparing the support and query sample features in the shared feature space, and good results have been obtained. Early studies in 2016 introduced the scenario training mechanism into few-shot learning, encoded each support sample in the context of the entire support set using bidirectional LSTM (Long Short-Term Memory Network), and matched the query sample with the support sample through the attention mechanism.
[0004] Most of the existing few-shot learning methods focus on abstracting global information from each sample. This is consistent with the human cognitive mechanism. The human cognitive mechanism usually first makes a rough recognition from a global perspective. However, in practice, when it is difficult to distinguish objects from a global perspective, humans will further resort to local detailed features. To make up for the deficiencies of the existing methods, some recent studies have proposed to extract subtle differences from a local perspective and achieved good results on fine-grained datasets. However, by only focusing on global or local features, the existing methods still cannot provide strong generalization ability on datasets with multiple classes.
[0005] In the Chinese patent application with the application number 201910600332.6, the National University of Defense Technology proposed a few-shot machine learning method for multimodal data. Aiming at the typical application of identifying and classifying multimodal data under the condition of few samples, this method adopts the methods of multimodal data encoding, hierarchical pooling, and relational network learning. Under the support of a small number of labeled samples, it can quickly identify and classify new category data, achieving a recognition accuracy superior to that of several current typical algorithms, and having good representation ability and generalization ability. Summary of the Invention
[0006] The primary technical problem to be solved by the present invention is to provide a few-shot learning method based on context memory and fine-grained calibration.
[0007] Another technical problem to be solved by the present invention is to provide a few-shot learning device based on context memory and fine-grained calibration.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] According to the first aspect of the embodiments of the present invention, there is provided a few-shot learning method based on context memory and fine-grained calibration, including the following steps:
[0010] S1, randomly sample a predetermined task from the dataset;
[0011] S2, load the parameters of the multi-layer convolutional neural network and the class-sensitive context memory network, and use the feature extraction network to extract the local features and global features of each sample in the task;
[0012] S3, perform hidden state update through the state update mechanism of the recurrent neural network;
[0013] S4, use the bidirectional update mechanism to obtain the final relational embedding;
[0014] S5, learn discriminative relational embeddings from the perspective of global features through the class-sensitive context memory network, and then predict the global query-to-class similarity;
[0015] S6, calculate the prediction reliability τ;
[0016] S7, if τ > τ0, take the class with the highest global query-to-class similarity as the final classification result; otherwise, further calculate the local query-to-class similarity to obtain the final classification result.
[0017] Preferably, in the step S1, the training task or the test task is obtained by using the random sampling method from the training set or the test set;
[0018] In the training task or the test task, a support set is included. and a query set The support set mentioned above includes N different categories, and each category includes K samples with known labels; the query set includes T samples with unknown labels; where K, N, and T are all positive integers.
[0019] Preferably, in step S2, a multi-layer convolutional neural network is used to extract features from each sample in the sampled task, obtaining three-dimensional local features of each sample, and then the local features are converted into one-dimensional global features through a fully connected layer or global pooling.
[0020] Preferably, the multi-layer convolutional neural network is WideResNet or ResNet.
[0021] Preferably, step S3 includes the following sub-steps:
[0022] S31, modify the information flow of the recurrent neural network and add residual connections;
[0023] S32, consider the historical information of the previous two iterations, linearly superimpose the two, and use the result as the hidden state variable before the update of the recurrent neural network, and update the hidden state through the state update mechanism of the recurrent neural network.
[0024] Preferably, step S4 includes the following sub-steps:
[0025] Use a bidirectional update mechanism to further learn context information, concatenate the hidden states output from two opposite directions, and use the result as the final relation embedding.
[0026] Preferably, in step S5, the calculation formula for the global query-to-class similarity is:
[0027]
[0028] where is the similarity score learned according to each relation embedding.
[0029] Preferably, in step S6, the two categories with the highest similarity are selected from the calculation results of the global query-to-class similarity and The prediction reliability τ of the two is calculated through the following formula:
[0030]
[0031] Preferably, in step S7, τ is compared with a pre-set reliability threshold τ0. If the value of τ is greater than the set reliability threshold τ0, the class with the highest prediction probability is regarded as the final result; otherwise, the classes and are determined to be difficult-to-distinguish classes, and fine-grained calibration is performed to obtain a more accurate local query-to-class similarity.
[0032] According to a second aspect of the embodiments of the present invention, a few-shot learning device based on context memory and fine-grained calibration is provided, including a processor and a memory. The processor reads a computer program in the memory and is configured to perform the following operations:
[0033] S1, randomly sample a predetermined task from a dataset;
[0034] S2, load the parameters of a multi-layer convolutional neural network and a class-sensitive context memory network, and use a feature extraction network to extract local features and global features of each sample in the task;
[0035] S3, perform hidden state update through the state update mechanism of a recurrent neural network;
[0036] S4, use a bidirectional update mechanism to obtain a final relationship embedding;
[0037] S5, learn a discriminative relationship embedding from the perspective of global features through a class-sensitive context memory network, and then predict the global query-to-class similarity;
[0038] S6, calculate the prediction reliability τ;
[0039] S7, if τ > τ0, the class with the highest global query-to-class similarity is used as the final classification result; otherwise, the local query-to-class similarity is further calculated to obtain the final classification result.
[0040] Compared with the prior art, the present invention mimics the cognitive mechanism of humans in few-shot learning, combines a learning process from "coarse" to "fine", first learns discriminative relational embeddings from the perspective of global features through a class-sensitive context memory network, and predicts the global query-to-class similarity; then, from the perspective of local features, supplements and improves the coarse relational embeddings predicted in the global stage to obtain a more accurate evaluation result of the local query-to-class similarity. The present invention integrates the global feature information and local feature information of the samples, and can significantly reduce the classification error. When dealing with a completely new task, the present invention has higher accuracy and better generalization ability. Brief Description of the Drawings
[0041] Figure 1 It is a flowchart of the few-shot learning method based on context memory and fine-grained calibration provided by the present invention;
[0042] Figure 2 It is a schematic diagram of the few-shot learning device based on context memory and fine-grained calibration provided by the present invention. Detailed Embodiments
[0043] The technical content of the present invention will be described in detail below with reference to the drawings and specific embodiments.
[0044] Currently, deep neural networks have very important applications in many aspects such as image recognition, speech recognition, and natural language processing. However, the models of deep neural networks usually have millions of parameters and need to be supervised and trained with a large amount of labeled data to obtain a relatively good effect. In practice, it is often difficult to provide sufficient labeled data for the models of deep neural networks.
[0045] Therefore, the embodiments of the present invention provide a few-shot learning method based on context memory and fine-grained calibration. The inspiration for this method comes from the "coarse" to "fine" learning process in human visual perception. By simulating the cognitive mechanism of humans in few-shot learning, first learn discriminative relational embeddings from the perspective of global features through a class-sensitive context memory network, and predict the global query-to-class (query to class) similarity; then, through a fine-grained calibration module, supplement and improve the coarse relational embeddings predicted in the global stage from the perspective of local features to obtain a more accurate local query-to-class (query to class) similarity evaluation. Next, in combination with Figure 1 A detailed description will be given.
[0046] As Figure 1 shown, the few-shot learning method provided by the embodiments of the present invention mainly includes the following steps:
[0047] S1, randomly sample a predetermined task from the dataset X;
[0048] S2, load the parameters of the multi-layer convolutional neural network and the class-sensitive context memory network, and use the feature extraction network to extract the local features and global features of each sample in the task;
[0049] S3, perform hidden state update through the state update mechanism of the recurrent neural network;
[0050] S4, use the bidirectional update mechanism to obtain the final relationship embedding;
[0051] S5, learn discriminative relationship embeddings from the perspective of global features through the class-sensitive context memory network, and then predict the global query-to-class similarity;
[0052] S6, calculate the prediction reliability τ;
[0053] S7, if τ > τ0, take the class with the highest global query-to-class similarity as the final classification result; otherwise, further calculate the local query-to-class similarity to obtain the final classification result.
[0054] In an embodiment of the present invention, step S1 specifically includes the following sub-steps: obtain a training task or a test task Γ from the training set or the test set by means of random sampling. In this task, a support set and a query set The support set contains N different classes, and each class contains K samples with known labels. The query set contains T samples with unknown labels. Among them, K, N, and T are all positive integers. For example, in the training set or the test set, task sampling is performed in two experimental settings of 5-way 1-shot and 5-way 5-shot and trained separately. In the task, one sample is sampled from each class as the query sample, and a total of 10,000 tasks are sampled for subsequent testing.
[0055] In one embodiment of the present invention, step S2 specifically includes the following sub-steps: Using a multi-layer convolutional neural network, such as WideResNet or ResNet, to extract features from each sample in the sampled task (the corresponding multi-layer convolutional neural network is called the feature extraction network), obtaining three-dimensional local features of each sample, and then converting the local features into one-dimensional global features through a fully connected layer or global pooling. After obtaining the global features, further use a class-sensitive context memory network to predict the global query-to-class similarity, and finally obtain a more accurate query-to-class similarity evaluation through a local fine-grained calibration module.
[0056] It should be noted that both WideResNet and ResNet in the above embodiments are typical multi-layer convolutional neural networks. Among them, WideResNet was published by Sergey Zagoruyko in 2016, aiming to improve ResNet from the perspective of increasing the network width, thereby improving performance and training speed. In one embodiment of the present invention, the WideResNet used in the training process is the WRN-28-10 architecture, and the parameters are pre-trained on the entire training set through cross-entropy classification loss, and the parameters remain unchanged during the training process of the few-shot learning model.
[0057] In one embodiment of the present invention, step S3 specifically includes the following sub-steps:
[0058] S31, modify the information flow of the Recurrent Neural Network (RNN for short) and add residual connections;
[0059] S32, update the hidden state through the state update mechanism of the recurrent neural network.
[0060] It should be noted that the above recurrent neural network can be specifically implemented by GRU (Gate Recurrent Unit) or LSTM (Long-Short Term Memory) to solve problems such as long-term memory and gradients in backpropagation. Compared with LSTM, GRU uses fewer parameters than LSTM, but can also achieve a function equivalent to LSTM. Considering the computing power of the hardware and the time cost, GRU is preferably used in the embodiments of the present invention. GRU provides a recurrent gating mechanism, where the reset gate mainly determines how much of the past information needs to be deleted, and the update gate determines how much of the past information (from previous time steps) needs to be memorized and passed to the future.
[0061] In a preferred embodiment using GRU, the above step S31 specifically includes the following operations:
[0062] During the update process of the hidden state of the GRU, a residual connection is added to further consider the historical information of the previous two iterations, and the two are linearly superimposed and used as the hidden state variable before the GRU update. Through this operation, catastrophic forgetting can be avoided and the context information of the entire category can be learned better.
[0063] The above step S32 specifically includes the following operations:
[0064] When learning the relationship between the query sample and the nth category, the global feature of the query sample is used to initialize the hidden state of the GRU. In the kth iteration, the global feature of the kth support set sample from this category is input into the modified GRU to obtain the updated hidden state:
[0065]
[0066]
[0067]
[0068]
[0069] Among them, represents the updated hidden state, is the hidden state of the previous two iterations. When k = 1, W z ,U z ,W r ,U r ,W h ,U h are all learnable parameters, σφ are the sigmoid activation function and the tanh activation function respectively, and represent the update gate and the reset gate in the GRU respectively.
[0070] In an embodiment of the present invention, step S4 specifically includes the following sub-steps:
[0071] Use a bidirectional update mechanism to further learn context information, and concatenate the hidden states output from two opposite directions as the final relationship embedding:
[0072]
[0073] Therefore, for the nth category, we can obtain a set of relationship embeddings
[0074] In an embodiment of the present invention, step S5 specifically includes the following sub-steps:
[0075] Query samples and classes The calculation formula for the global query-to-class similarity is as follows:
[0076]
[0077] where is the similarity score learned based on each relationship embedding.
[0078] Then calculate the cross-entropy loss:
[0079]
[0080] where
[0081]
[0082] In an embodiment of the present invention, step S6 specifically includes the following sub-steps:
[0083] Calculate the prediction reliability from the calculation result of the global query-to-class similarity, and determine whether fine-grained calibration is required?
[0084] Specifically, select the two classes with the highest similarity from the calculation result of the global query-to-class similarity and Calculate the prediction reliability τ of the two through the following formula:
[0085]
[0086] In an embodiment of the present invention, step S7 specifically includes the following sub-steps:
[0087] Compare τ with a pre-set reliability threshold τ0. If the value of τ is greater than the set reliability threshold τ0, then consider the class with the highest prediction probability as the final result. Otherwise, consider the classes and as difficult to distinguish classes, perform fine-grained calibration, and obtain a more accurate local query-to-class similarity.
[0088] Specifically, split the local features output for each sample to obtain a set of local feature blocks [p1,…,p M , where p j is the jth local feature block. For each feature block p j of the query sample, find the L local feature blocks with the closest distance in the local feature block space of the support samples of class , and then calculate the query sample and class Local query-to-class similarity:
[0089]
[0090] To implement the few-shot learning method based on context memory and fine-grained calibration provided by the present invention, the present invention also provides a few-shot learning device based on context memory and fine-grained calibration. As Figure 2 shown, the device includes a memory 21 and a processor 22, and may further include a communication component, a sensor component, a power supply component, a multimedia component, and an input / output interface according to actual needs. Among them, the memory, the communication component, the sensor component, the power supply component, the multimedia component, and the input / output interface are all connected to the processor 22. The above-mentioned memory 21 may be a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, etc., and the processor may be a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a digital signal processing (DSP) chip, etc. Other communication components, sensor components, power supply components, multimedia components, etc. can all be implemented by general components in existing smart phones, and will not be specifically described here.
[0091] On the other hand, in the above-mentioned few-shot learning device based on context memory and fine-grained calibration, the processor 22 reads the computer program in the memory 21 and is used to perform the following operations:
[0092] S1, randomly sample a predetermined task from the data set X;
[0093] S2, load the parameters of the multi-layer convolutional neural network and the class-sensitive context memory network, and use the feature extraction network to extract the local features and global features of each sample in the task;
[0094] S3, perform hidden state update through the state update mechanism of the recurrent neural network;
[0095] S4, use the bidirectional update mechanism to obtain the final relationship embedding;
[0096] S5, learn the discriminative relationship embedding from the perspective of global features through the class-sensitive context memory network, and then predict the global query-to-class similarity;
[0097] S6, calculate the prediction reliability τ;
[0098] S7. If τ > τ0, use the class with the highest global query-to-class similarity as the final classification result; otherwise, further calculate the local query-to-class similarity to obtain the final classification result. During the training process of the few-shot learning method and device provided by the present invention, steps S1 to S5 are iteratively executed, and the parameters of the class-sensitive context memory network are updated through the cross-entropy loss l. In one embodiment of the present invention, the Adam optimizer can be used during training, with the initial learning rate set to 0.001, halved every 15,000 generations on the miniImageNet dataset and every 30,000 generations on the tieredImageNet dataset. The input image size is set to 84×84, 20 images are trained in each batch, and gradient decay is added during training, with the gradient decay rate set to 0.000001. Keep training until the loss function converges, save the parameters of each layer of the neural network with the best performance on the validation set, and complete the training of the class-sensitive context memory network.
[0099] During the testing process of the few-shot learning method and device provided by the present invention, steps S1 to S7 are sequentially executed to obtain the prediction results for the categories to which each image in the query set belongs. In one embodiment of the present invention, the prediction reliability threshold τ0 for fine-grained calibration during testing can be set to 1.5, and the number of nearest neighbors L is set to 3. During the execution process, the cross-entropy loss l is no longer updated.
[0100] Compared with the prior art, the present invention mimics the human cognitive mechanism in few-shot learning, combines a learning process from "coarse" to "fine", first learns discriminative relational embeddings from the perspective of global features through a class-sensitive context memory network and predicts the global query-to-class similarity; then, supplements and improves the coarse relational embeddings predicted in the global stage from the perspective of local features to obtain a more accurate evaluation result of the local query-to-class similarity. The present invention integrates the global feature information and local feature information of the samples, can significantly reduce the classification error, and has higher accuracy and better generalization ability when dealing with new tasks.
[0101] The above has described in detail the few-shot learning method and device based on context memory and fine-grained calibration provided by the present invention. For those of ordinary skill in the art, any obvious changes made without departing from the essential content of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.
[0102] The above has described in detail the few-shot learning method and device based on context memory and fine-grained calibration provided by the present invention. For those of ordinary skill in the art, any obvious changes made without departing from the essential content of the present invention will constitute an infringement of the patent right of the present invention and will bear corresponding legal responsibilities.
Claims
1. A few-shot learning method based on context memory and fine-grained calibration, characterized in that It includes the following steps: S1. Randomly sample a predetermined task from the dataset; S2. Load the parameters of the multi-layer convolutional neural network and the class-sensitive context memory network, and use the feature extraction network to extract the local features and global features of each sample in the task; S3. Modify the information flow of the recurrent neural network and add residual connections; consider the historical information of the previous two iterations, linearly superimpose the two, and use the result as the hidden state variable before the update of the recurrent neural network, and update the hidden state through the state update mechanism of the recurrent neural network; S4. Use the bidirectional update mechanism to obtain the final relationship embedding; S5. Learn the discriminative relationship embedding from the perspective of global features through the class-sensitive context memory network, and then predict the global query-to-class similarity; S6. Calculate the prediction reliability τ; S7, if τ > τ0, use the class with the highest global query-to-class similarity as the final classification result; otherwise, further calculate the local query-to-class similarity to obtain the final classification result. 2. The few-shot learning method according to claim 1, characterized in that: In step S1, the training task or the test task is obtained by using the random sampling method from the training set or the test set; In the training task or the test task, a support set is included. and a query set The support set includes N different categories, and each category contains K samples with known labels; the query set contains T samples with unknown labels; where K, N, and T are all positive integers.
3. The few-shot learning method according to claim 1, characterized in that: In step S2, use the multi-layer convolutional neural network to extract the features of each sample in the sampled task to obtain the three-dimensional local features of each sample, and then convert the local features into one-dimensional global features through the fully connected layer or global pooling.
4. The few-shot learning method according to claim 3, characterized in that: The multi-layer convolutional neural network is WideResNet or ResNet.
5. The few-shot learning method according to claim 1, characterized in that Step S4 includes the following sub-steps: Use the bidirectional update mechanism to further learn the context information, concatenate the hidden states output from two opposite directions, and use the result as the final relationship embedding.
6. The few-shot learning method according to claim 1, characterized in that: In step S5, the calculation formula of the global query-to-class similarity is: where is the similarity score learned according to each relational embedding.
7. The few-shot learning method according to claim 1, characterized in that: In the step S6, select the two classes with the highest similarity from the calculation results of the global query-to-class similarity and Calculate the prediction reliability τ of the two through the following formula:
8. The few-shot learning method according to claim 1, characterized in that: In the step S7, τ is compared with a preset reliability threshold τ0. If the value of τ is greater than the set reliability threshold τ0, the category with the highest prediction probability is regarded as the final result; otherwise, the categories and are determined to be difficult-to-distinguish categories, and fine-grained calibration is performed to obtain a more accurate local query-to-class similarity.
9. A few-shot learning device based on context memory and fine-grained calibration, characterized in that It includes a processor and a memory. The processor reads the computer program in the memory and is used to perform the following operations: S1. Randomly sample a predetermined task from the dataset; S2. Load the parameters of the multi-layer convolutional neural network and the class-sensitive context memory network, and use the feature extraction network to extract the local features and global features of each sample in the task; S3. Modify the information flow of the recurrent neural network and add residual connections; consider the historical information of the previous two iterations, linearly superimpose the two, and use the result as the hidden state variable before the update of the recurrent neural network, and update the hidden state through the state update mechanism of the recurrent neural network; S4. Use the bidirectional update mechanism to obtain the final relationship embedding; S5. Learn the discriminative relationship embedding from the perspective of global features through the class-sensitive context memory network, and then predict the global query-to-class similarity; S6. Calculate the prediction reliability τ; S7, if τ > τ0, use the class with the highest global query-to-class similarity as the final classification result; otherwise, further calculate the local query-to-class similarity to obtain the final classification result.
Citation Information
Patent Citations
Multi-modal data-oriented small sample machine learning method and system, and medium
CN110363239A
Small sample learning method based on multi-scale metric learning
CN111639679A