Small sample image classification method and system based on feature pyramid and feature fusion
By constructing a feature pyramid relationship network model and combining feature fusion method, the problem of insufficient accuracy of the existing small sample image classification method is solved, and high-accuracy image classification under small sample data is achieved.
Patent Information
- Application Number
- CN202210733595.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The existing small sample image classification methods have shortcomings in terms of accuracy, especially the failure to effectively utilize the similarity and measurement of the middle layer of the model network, resulting in a low recognition accuracy.
A small sample image classification method based on feature pyramids and feature fusion is adopted to improve the accuracy of image classification by constructing a multi-layer neural network feature pyramid relationship network model, including feature extraction modules, relationship modules and feature fusion modules.
Through the feature pyramid relationship network model, the accuracy of small sample image classification is significantly improved, and because the model itself is small, it can quickly obtain detection results and has a high accuracy rate.
Smart Images

Figure CN115272692B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of small sample learning and meta-learning, and in particular to a small sample image classification method and system based on feature pyramid and feature fusion. Background Art
[0002] Deep neural network models often require a large number of training samples with labeled data to achieve better training results. In reality, however, sample labels often require a lot of manpower and material resources, or in some cases, there are very few sample data that can be used. At this time, if a small number of samples are directly used for training, overfitting problems will occur. Small sample learning is created to solve this type of problem.
[0003] The basic model of small sample learning is defined as p = C (f (x | θ) | w), where the feature extractor can be represented as f, the classifier can be represented as C, x represents the input image to be recognized, θ represents the parameters of the feature extractor f, w represents the parameters of the classifier C, and p represents the predicted result of the model output. In the process of small sample learning, due to the small number of samples, direct training will lead to overfitting of the model parameters θ and w, and the accuracy of the target task will decrease.
[0004] We define D as a training set of similar previous tasks for which a large amount of data is available base , the small sample learning dataset containing the target detection task is defined as D novel . Place the model in D base The model is trained on D novel Training is performed on the original image to obtain new model parameters θ1 and w1, and the original parameters are updated. The updated new model p=C(f(x|θ1)|w1) can complete the image classification task more accurately.
[0005] Focusing on the core issue of small number of samples, the existing small sample learning strategies are mainly solved through methods based on data enhancement, metric learning, model and parameter optimization. Metric learning is also called similarity learning. The task goal of metric learning is to learn a pairwise similarity metric S(·,·), where similar samples have higher similarity scores and dissimilar samples have lower similarity scores. S can be either a distance metric that does not need to be learned or a neural network that can be learned. The similarity score of its output can be used to query the classification of test set samples. However, the existing small sample learning based on metric learning only focuses on the similarity and distance metrics between the final outputs of the model, but not the similarity and metrics of the intermediate layers of the model network. The recognition accuracy is low, which affects the final classification accuracy of the model. Summary of the invention
[0006] The purpose of the present invention is to overcome the problem of low accuracy in small sample image classification in the prior art, and to provide a small sample image classification method based on feature pyramid and feature fusion, which improves the accuracy in small sample image classification through a new feature pyramid relationship network and a new feature fusion method.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A small sample image classification method based on feature pyramid and feature fusion includes the following steps:
[0009] S1. Construct a feature pyramid relational network model of a multi-layer neural network, where each layer of the neural network includes a feature extraction module, a relational module, and a feature fusion module;
[0010] S2. Obtain and expand the data set, and divide the expanded data set into a training set, a validation set, and a test set;
[0011] S3. Use C-way K-shot method to train the feature pyramid relational network model, and sample the support set and query set from the training set for each training;
[0012] S4. Input the support set image and the query set image, the feature extraction module extracts the features of the image, and outputs the feature vector of the image, and the feature fusion module fuses the feature vectors of the support set image and the query set image;
[0013] S5. The fused feature vector is input into the relationship module, and the relationship module outputs the similarity score of the support set image and the query set image. The similarity scores output by all relationship modules are processed to obtain the final similarity score;
[0014] S6. Calculate the loss of the feature pyramid relationship network model, update the parameters of the feature pyramid relationship network model, and repeat the iterative training until the error value of the loss tends to be stable;
[0015] S7. Save the trained feature pyramid relational network model, and use the feature pyramid relational network model for small sample image classification test.
[0016] Furthermore, the feature fusion module includes feature fusion items, which are:
[0017] C′(F S ,F Q )=Concate(F S ,F Q ,Mul(F S ,F Q ))
[0018] In the formula, FS The feature vector representing the query set image, F Q represents the feature vector of the support set image, Concate(·,·) represents the concatenation operation on the feature channel, and Mul(·,·) operation represents multiplying the corresponding elements of the feature map by position.
[0019] Furthermore, in step S6, the similarity scores of a group of images are regarded as a regression task, and the mean square error MSE function is used as the loss function of each layer of the neural network. The mean square error MSE function is:
[0020] MSE(r,y S ,y Q )=(r-1(y S = =y Q )) 2
[0021] In the formula, r represents the similarity score output by each layer of the neural network, y S represents the label of the support set image, y Q Represents the labels of the query set images.
[0022] Further, in step S6, the loss of the feature pyramid relationship network model is calculated using a loss function, and the loss function is:
[0023]
[0024] In the formula, r l Represents the similarity score of the output of the l-th layer neural network, y S represents the label of the support set image, y Q represents the label of the query set image, MSE represents the mean square error function, and n represents the number of layers of the neural network.
[0025] Furthermore, in step S2, the acquired data set is expanded by rotating the data set at angles of 90 degrees, 180 degrees, and 270 degrees.
[0026] Furthermore, in the feature pyramid relationship network model, the activation function of the last fully connected layer of the relationship module uses the Sigmoid function, and all other activation functions use the ReLU function.
[0027] A small sample image classification system based on feature pyramid and feature fusion, comprising:
[0028] A feature extraction module, used to extract features of an input image;
[0029] A feature fusion module is used to fuse the features of the input image;
[0030] The relation module is used to determine the similarity between the input support set image features and the query set image features.
[0031] Furthermore, the feature extraction module includes four convolution blocks and two 2*2 maximum pooling layers, and the convolution block, the maximum pooling layer, the convolution block, the maximum pooling layer, the convolution block, and the convolution block are connected in sequence.
[0032] Furthermore, the relationship module includes two convolution blocks, two 2*2 maximum pooling layers, a ReLU fully connected layer and a Sigmoid fully connected layer. The convolution block, the maximum pooling layer, the convolution block, the maximum pooling layer, the ReLU fully connected layer, and the Sigmoid fully connected layer are connected in sequence.
[0033] Furthermore, the convolution block includes a convolution layer, a Batch Norm layer and a ReLU activation function layer. The convolution kernel size of the convolution layer is 3*3 and the number of output channels is 64.
[0034] Compared with the prior art, the present invention improves the accuracy of small sample image classification by constructing a feature pyramid relational network (FPRN) model. Moreover, since the feature pyramid relational network (FPRN) model itself is small in size, the detection results can still be quickly obtained through the feature pyramid relational network (FPRN) model with high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 Schematic diagram of the structure of the feature pyramid relationship network FPRN.
[0036] Figure 2 This is a structural diagram of the feature extraction module.
[0037] Figure 3 It is a structural diagram of the relationship module.
[0038] Figure 4 It is a structural diagram of the convolution block Conv Block. DETAILED DESCRIPTION
[0039] The small sample image classification method and system based on feature pyramid and feature fusion of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0040] The present invention discloses a small sample image classification method based on feature pyramid and feature fusion. The small sample image classification method based on feature pyramid and feature fusion comprises the following steps:
[0041] S1. Construct a feature pyramid relational network model of a multi-layer neural network, where each layer of the neural network includes a feature extraction module, a relation module, and a feature fusion module.
[0042] S2. Obtain and expand the data set, and divide the expanded data set into a training set, a validation set, and a test set.
[0043] S3. The feature pyramid relational network model is trained using the C-way K-shot method, and the support set and query set are sampled from the training set for each training.
[0044] S4. Input the support set image and the query set image, the feature extraction module extracts the features of the image and outputs the feature vector of the image, and the feature fusion module fuses the feature vectors of the support set image and the query set image.
[0045] S5. The fused feature vector is input into the relationship module, and the relationship module outputs the similarity score between the support set image and the query set image. The similarity scores output by all relationship modules are processed to obtain the final similarity score.
[0046] S6. Calculate the loss of the feature pyramid relational network model, update the parameters of the feature pyramid relational network model, and repeat the iterative training until the error value of the loss tends to be stable.
[0047] S7. Save the trained feature pyramid relational network model, and use the feature pyramid relational network model for small sample image classification test.
[0048] See also Figure 1 The present invention also discloses a small sample image classification system based on feature pyramid and feature fusion. The small sample image classification system based on feature pyramid and feature fusion includes a feature extraction module, a feature fusion module and a relationship module. The feature extraction module is used to extract features of an input image, the feature fusion module is used to fuse features of the input image, and the relationship module is used to determine the similarity between input support set image features and query set image features.
[0049] Specifically, in a neural network, the deeper the neural network level, the larger the receptive field, and the more attention is paid to the overall features of the image. The shallower the neural network level, the smaller the receptive field, and the more attention is paid to the local features of the image. For example, when classifying animals, a deep neural network can distinguish the characteristics of specific animal species, and a shallow neural network can extract hair features, background texture features, etc. Therefore, the features of the shallow neural network can also be used to help distinguish animal species. Based on this, the present invention proposes a Feature Pyramid Relation Network (FPRN) model.
[0050] The feature pyramid relational network model has a multi-layer neural network, each layer of which includes a feature extraction module (FEM), a feature fusion module and a relation module (RM). The feature extraction module is used to extract the features of the input image, the feature fusion module is used to fuse the features of the input image, and the relation module is used to determine the similarity between the input support set image features and the query set image features.
[0051] like Figure 2 As shown, the feature extraction module includes four convolution blocks and two 2*2 maximum pooling layers. In the feature extraction module, the convolution block, the maximum pooling layer, the convolution block, the maximum pooling layer, the convolution block, and the convolution block are connected in sequence, and each convolution block outputs a group of two feature vectors for feature fusion.
[0052] like Figure 3 As shown in the figure, the relationship module includes two convolution blocks, two 2*2 maximum pooling layers, a ReLU fully connected layer and a Sigmoid fully connected layer. In the relationship module, the convolution block, the maximum pooling layer, the convolution block, the maximum pooling layer, the ReLU fully connected layer and the Sigmoid fully connected layer are connected in sequence. The input of the relationship module is the features fused by the feature fusion module, and its output is a similarity score, which is used to judge the similarity between the input support set image features and the query set image features.
[0053] like Figure 4 As shown in the figure, the convolution block consists of a convolution layer with a convolution kernel size of 3*3 and 64 output channels, a Batch Norm layer, and a ReLU activation function layer.
[0054] In the feature pyramid relationship network model, except for the activation function of the last fully connected layer of the relationship module, which is the Sigmoid function, all other activation functions use the ReLU function. The final output of the relationship module uses the Sigmoid function because the present invention expects to output a similarity score between 0 and 1.
[0055] The feature fusion module includes feature fusion items, which are:
[0056] C′(F S ,F Q )=Concate(F S ,F Q ,Mul(F S ,F Q ))
[0057] In the formula, F S represents the feature vector generated by the convolutional block of the query set image, F QIt represents the feature vector generated by the convolution block of the support set image, Concate(·,·) represents the concatenation operation on the feature channel, and Mul(·,·) operation represents multiplying the corresponding elements of the feature map by position.
[0058] The acquired data set is expanded by rotating it at angles of 90, 180, and 270 degrees.
[0059] For C-way 1-shot problems, the feature map of the support set extracted by the feature extraction module can be directly fused with the feature map of the query set. For C-way K-shot (K>1) problems, the feature maps extracted by the feature extraction module from multiple images of the support set need to be added element by element at corresponding positions, and then fused with the feature map of the query set, and then put into the relationship module to calculate the similarity score.
[0060] During the training phase, C categories are randomly selected from the training set, and K samples are selected from each category as the support set of the Feature Pyramid Relational Network (FPRN) model; a batch of samples are then selected from the remaining data in these C categories as the query set. The Feature Pyramid Relational Network (FPRN) model hopes to learn the ability to distinguish these C categories from these C*K samples during training, and different categories will be sampled in each round of training.
[0061] The feature vectors obtained at different network depths of the support set image and the query set image are fused together through the feature fusion term. In the feature fusion term proposed in the present invention, Mul(F S ,F Q ) can be introduced to S and F Q The region of interest is more prominent, which is more conducive to the feature pyramid relationship network to judge the similarity.
[0062] The relation module outputs a similarity score between 0 and 1, and the similarity scores output by the relation modules of all neural networks are weighted averaged to obtain the final similarity score of the Feature Pyramid Relational Network (FPRN) model.
[0063] The similarity scores of the support set images and the query set images are regarded as regression tasks, and the mean square error MSE function is used as the loss function of each layer of the neural network. The mean square error MSE function is:
[0064] MSE(r,y s ,y Q )=(r-1(y S = =y Q )) 2
[0065] In the formula, r represents the similarity score output by each layer of the neural network, y S represents the label of the support set image, y Q represents the label of the query set image. When the labels are the same, (y S = =y Q ) has a value of 1. When the labels are different, (y S = =y Q ) has a value of 0.
[0066] In the Feature Pyramid Relational Network (FPRN) model, the relational module of each layer of the neural network outputs a similarity score, so the total loss function of the Feature Pyramid Relational Network (FPRN) model is:
[0067]
[0068] In the formula, r l Represents the similarity score of the output of layer l, y S represents the label of the support set image, y Q represents the label of the query set image, and MSE represents the mean squared error function.
[0069] The loss of the Feature Pyramid Relational Network (FPRN) model is calculated through the loss function, and the parameters of the Feature Pyramid Relational Network (FPRN) model are updated by back propagation. The iterative training is repeated until the error value of the loss calculated by the loss function tends to be stable.
[0070] The trained feature pyramid relationship network model is saved and used for small sample image classification test. The small sample image classification method based on feature pyramid and feature fusion proposed in the present invention has achieved good detection results on two public data sets.
[0071] The Omniglot dataset contains 1623 character classes in 50 different languages, and each character class contains 20 samples written by different people. During the training process, 20-way 1-shot training consists of 1 support set image and 10 query set images per category, and 20-way 5-shot training consists of 5 support set images and 5 query set images per category. During the testing process, the present invention randomly samples 1000 times in the test set to evaluate the classification results of the feature pyramid relational network (FPRN) model, where 1-shot samples 1 test set image each time, and 5-shot samples 5 test set images each time.
[0072] The miniImagenet dataset consists of 60,000 color images in 100 categories, each of which contains 600 samples. The present invention uses 64 categories for training, 16 categories for verification, and 20 categories for testing. On the miniImagenet dataset, the present invention adopts the settings of 5-way 1-shot and 5-way 5-shot. During the training process, each training of 5-way 1-shot consists of 1 support set image and 15 query set images per category, and each training of 5-way 5-shot consists of 5 support set images and 10 query set images per category. During the test process, the present invention randomly samples 600 times in the test to evaluate the classification results of the feature pyramid relational network (FPRN) model, wherein 15 test set images are sampled each time in the settings of 5-way 1-shot and 5-way 5-shot.
[0073] The present invention compares the results of the feature pyramid relational network (FPRN) model with other popular image classification methods based on small sample learning models of metric learning. The models used for comparison mainly include Siamese Network, Prototype Network, Matching Network and Relation Network. The comparison results of the feature pyramid relational network model (FPRN) and these model benchmarks on the Omniglo dataset are shown in Table 1.
[0074] Table 1 Experimental results of Omniglot dataset
[0075]
[0076] The comparison results of the Feature Pyramid Relational Network (FPRN) model with the Siamese Network, Prototype Network, Matching Networks and Relation Network model benchmarks on the miniImagenet dataset are shown in Table 2.
[0077] Table 2 Experimental results of miniImagenet dataset
[0078]
[0079] As shown in Table 1 and Table 2, experimental data show that the feature pyramid relational network (FPRN) model proposed in the present invention has achieved the highest judgment accuracy in various experiments. On the Ominiglot dataset, the feature pyramid relational network (FPRN) model proposed in the present invention can achieve a classification accuracy of 98.3% in the 20-way 1-shot setting, and a classification accuracy of 99.2% in the 20-way 5-shot setting. On the miniImagenet dataset, the feature pyramid relational network (FPRN) model proposed in the present invention can achieve a classification accuracy of 50.2% in the 5-way 1-shot setting, and a classification accuracy of 66.7% in the 5-way 5-shot setting.
[0080] The present invention compares the detection speed of the relational network model and the feature pyramid relational network (FPRN) model in the 5-way 1-shot setting on the miniImagenet dataset. The graphics card used in the experiment of the present invention is an NVIDIA Quadro P2000 graphics card, the discrimination speed of using the relational network is 17.1fps, the discrimination speed of using the feature pyramid relational network model (FPRN) is 16.3fps, and the feature pyramid relational network (FPRN) model is 4.7% slower than the relational network model in discrimination speed. In this experimental setting, the detection accuracy of the feature pyramid relational network (FPRN) model is 50.2%, and the detection accuracy of the relational network model is 47.3%. The absolute value of the detection accuracy of the feature pyramid relational network (FPRN) model is 2.9% higher than that of the relational network model, and the accuracy is increased by 6.1% in percentage calculation. The feature pyramid relational network (FPRN) model obtains a 6.1% percentage accuracy increase under the condition of sacrificing 4.7% of the detection time.
[0081] In summary, the present invention improves the accuracy of small sample image classification by constructing a feature pyramid relational network (FPRN) model. Since the feature pyramid relational network (FPRN) model itself is small in size, the detection results can still be quickly obtained through the feature pyramid relational network (FPRN) model with high accuracy.
[0082] The above description is a detailed description of the preferred feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. All equivalent changes or modified changes completed under the technical spirit disclosed by the present invention should fall within the patent scope covered by the present invention.
Claims
1. A small sample image classification method based on feature pyramid and feature fusion, characterized in that: The following steps are involved: S1. Construct a feature pyramid relational network model of a multi-layer neural network, where each layer of the neural network includes a feature extraction module, a relational module, and a feature fusion module; S2. Obtain and expand the data set, and divide the expanded data set into a training set, a validation set, and a test set; S3. Use C-way K-shot method to train the feature pyramid relational network model, and sample the support set and query set from the training set for each training; S4. Input the support set image and the query set image, the feature extraction module extracts the features of the image, and outputs the feature vector of the image, and the feature fusion module fuses the feature vectors of the support set image and the query set image; S5. The fused feature vector is input into the relationship module, and the relationship module outputs the similarity score of the support set image and the query set image. The similarity scores output by all relationship modules are processed to obtain the final similarity score; S6. Calculate the loss of the feature pyramid relationship network model, update the parameters of the feature pyramid relationship network model, and repeat the iterative training until the error value of the loss tends to be stable; S7. Save the trained feature pyramid relationship network model, and use the feature pyramid relationship network model for small sample image classification test; The feature fusion module includes feature fusion items, which are: C′(F S ,F Q )=Concate(F S ,F Q ,Mul(F S ,F Q )) In the formula, F S The feature vector representing the query set image, F Q represents the feature vector of the support set image, Concate(·,·) represents the concatenation operation on the feature channel, and Mul(·,·) operation represents multiplying the corresponding elements of the feature map by position; In step S6, the similarity scores of a group of images are regarded as a regression task, and the mean square error MSE function is used as the loss function of each layer of the neural network. The mean square error MSE function is: In the formula, r represents the similarity score output by each layer of the neural network, y S represents the label of the support set image, y Q Represents the label of the query set image; when the labels are the same, (y S = =y Q ) has a value of 1. When the labels are different, (y S = =y Q ) has a value of 0; In step S6, the loss function is used to calculate the loss of the feature pyramid relationship network model, and the loss function is: In the formula, r l Represents the similarity score of the output of the l-th layer neural network, y S denotes the label of the support set image, y Q represents the label of the query set image, MSE represents the mean square error function, and n represents the number of layers of the neural network.
2. The small sample image classification method based on feature pyramid and feature fusion according to claim 1, characterized in that: In step S2, the acquired data set is expanded by rotating the data set at angles of 90 degrees, 180 degrees, and 270 degrees.
3. The small sample image classification method based on feature pyramid and feature fusion according to claim 1, characterized in that: In the feature pyramid relationship network model, the activation function of the last fully connected layer of the relationship module uses the Sigmoid function, and all other activation functions use the ReLU function.
4. A small sample image classification system based on feature pyramid and feature fusion, applying the small sample image classification method based on feature pyramid and feature fusion according to any one of claims 1 to 3, characterized in that: include: A feature extraction module, used to extract features of an input image; A feature fusion module is used to fuse the features of the input image; The relation module is used to determine the similarity between the input support set image features and the query set image features.
5. The small sample image classification system based on feature pyramid and feature fusion as claimed in claim 4, characterized in that: The feature extraction module includes four convolution blocks and two 2*2 maximum pooling layers. The convolution block, maximum pooling layer, convolution block, maximum pooling layer, convolution block, and convolution block are connected in sequence.
6. The small sample image classification system based on feature pyramid and feature fusion as claimed in claim 4, characterized in that: The relationship module includes two convolution blocks, two 2*2 maximum pooling layers, a ReLU fully connected layer and a Sigmoid fully connected layer. The convolution block, the maximum pooling layer, the convolution block, the maximum pooling layer, the ReLU fully connected layer, and the Sigmoid fully connected layer are connected in sequence.
7. The small sample image classification system based on feature pyramid and feature fusion according to claim 5 or 6, characterized in that: The convolution block includes a convolution layer, a Batch Norm layer, and a ReLU activation function layer. The convolution kernel size of the convolution layer is 3*3, and the number of output channels is 64.
Citation Information
Patent Citations
Attention mechanism relationship comparison network model method based on small sample learning
CN110020682A
Small sample commodity image classification method, a device, equipment and storage medium
CN113780335A