A Few-Shot Image Recognition Method Combining Contrastive Learning

By dividing the data set into basic classes and few-sample classes, and fine-tuning the MOCO model using a comparative learning framework, the problems of poor recognition of few-sample data and high preprocessing cost are solved, and high precision and low-cost image recognition of few-sample images are achieved.

CN114022754BActive Publication Date: 2025-07-04SUN YAT SEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111365133.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-17
Publication Date
2025-07-04
Estimated Expiration
2041-11-17

AI Technical Summary

Technical Problem

The prior art has poor effect in identifying small sample data, and the preprocessing work is too high, making it difficult to effectively utilize prior knowledge.

Method used

The data set is divided into basic classes and small sample classes, the MOCO model is trained using unlabeled data, and further fine-tuned through the comparative learning framework, combining the comparative loss function and the classification loss function to optimize the model.

Benefits of technology

It improves the accuracy of data recognition for few samples, reduces the cost of preprocessing, and realizes simple, practical and efficient image recognition for few samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114022754B_ABST
    Figure CN114022754B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot image recognition method combined with contrastive learning, which relates to the technical field of machine learning. The technical solution is as follows: A few-shot image recognition method combined with contrastive learning includes the following steps: S1. Given a labeled data set and an unlabeled data set, divide the labeled data set into a base class data set and a few-shot class data set according to categories; S2. Train a MOCO model through the unlabeled data set and the base class data set to obtain a pre-trained MOCO model; S3. Further fine-tune and train the pre-trained MOCO model through a contrastive learning framework; S4. Classify and recognize the targets in the few-shot class data set through the MOCO model, solving the problems of poor recognition effect of few-shot data and too high preprocessing work cost in the prior art, and providing a few-shot image recognition method combined with contrastive learning, which has the characteristics of simplicity, practicality, high accuracy and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and more specifically, to a few-shot image recognition method combining contrast learning. Background Art

[0002] Early machine learning has achieved many results in application scenarios with sufficient data volume, but it often fails to achieve the expected results in tasks with scarce data. Few-shot learning, a deep learning method, is inspired by the human way of understanding the world and things. It is found that usually by using prior knowledge, accurate conclusions can be quickly obtained for unknown new tasks. Few-shot learning methods make the learning of tasks feasible by combining the available supervised information in experience with some prior knowledge. In order to imitate this efficient learning method, researchers have successively proposed many efficient, practical and effective deep learning models in the field of few-shot learning. The existing typical methods are mainly supervised learning methods, which can be roughly classified into the following categories according to data, model, and algorithm: using prior knowledge data as supervision conditions such as meta-learning methods, using prior knowledge data for comparison and constraint such as metric learning methods, and generative models, multi-task combination methods, etc. In addition, there are also methods to improve the overall effect starting from data, such as obtaining more samples from the training set, unlabeled data, and similar task data.

[0003] Classic methods rely on deep convolutional networks and require a large number of training samples. Even the above-mentioned few-shot learning algorithms, which can learn new categories after training on basic categories, mainly focus on single-task scenarios rather than more general multi-label scenarios or transfer applications of different tasks. For example, a Chinese patent discloses a self-supervised image classification method based on contrast learning. By extracting features from views, calculating the loss through unsupervised contrast, an unsupervised classification model is obtained, and part of the unlabeled data is manually labeled as the training and validation set; then, as a pre-trained model, it is fine-tuned according to the training and validation set. However, the existing methods either have high workload requirements for pre-training and pre-processing, or the generative model method is difficult to deduce and has high computational costs, etc., and do not make good use of the ability to efficiently utilize prior knowledge. When combining multiple tasks, new tasks can be used to add constraints to the original results, but the process of model training needs to be re-executed, resulting in additional time costs. When the data or tasks are complex, overfitting is likely to occur in the training tasks. The method of adding metric learning mostly uses the embedding space, which can solve the overfitting problem to a certain extent. However, the task scenario for obtaining prior knowledge needs to be consistent with the few-shot task, and the effect is often poor when the test set deviates too much from the training set. Therefore, it is basically only used for supervised learning and has certain requirements for data distribution. Summary of the Invention

[0004] The present invention aims to solve the problems of poor recognition effect on few-shot data and high preprocessing cost in the prior art, and provides a few-shot image recognition method combined with contrastive learning, which has the characteristics of simplicity, practicality, high accuracy and low cost.

[0005] To achieve the above object of the present invention, the following technical solutions are adopted:

[0006] A few-shot image recognition method combined with contrastive learning, comprising the following steps:

[0007] S1. Given a labeled dataset and an unlabeled dataset, divide the labeled dataset into a base class dataset and a few-shot class dataset according to categories;

[0008] S2. Train a MOCO model through the unlabeled dataset and the base class dataset to obtain a pre-trained MOCO model;

[0009] S3. Further fine-tune and train the pre-trained MOCO model through a contrastive learning framework;

[0010] S4. Classify and recognize the targets in the few-shot class dataset through the MOCO model.

[0011] Preferably, the base class dataset and the few-shot class dataset in S1 do not intersect, that is where C base is the base class dataset, C novel is the few-shot class dataset, and each few-shot class data has n training samples, n < 10.

[0012] Furthermore, in step S2, the MOCO model is a self-supervised model, including an encoder module, a multi-layer perceptron module, and a queue module; the encoder module processes the tensor corresponding to the image data to obtain data such as a feature matrix, and constructs a "key" in the unsupervised learning dictionary data to retrieve the corresponding data; the multi-layer perceptron module is used to obtain image features; the queue module completes the storage and maintenance of the dictionary data by setting queue rules.

[0013] Even further, step S2 specifically includes the following steps:

[0014] S201. Collect images (Img) in the unlabeled dataset;

[0015] S202. Define the encoder as encoder f in the encoder module, and the corresponding parameter is ω;

[0016] S203. Encode the image Img with the encoder:

[0017] E = f(Img),

[0018] where \(E\in\mathbb{R}\) n ;

[0019] S204. After obtaining the encoding, through the projection layer MLP of one of the multi-layer perceptrons:

[0020] E' = MLP(E);

[0021] S205. The queue module adds the obtained E' after each round of training to a queue after this round of training, maintains it according to the "first in first out" rule to ensure the real-time nature of the queue, and uses the encoding saved in the queue as negative samples for training in each round of training.

[0022] Furthermore, step S3 specifically includes the following steps:

[0023] S301. Use contrastive learning to augment the targets in the few-shot class dataset;

[0024] S302. Map the first encoding of the augmented targets through the projection layer MLP of the multi-layer perceptron to obtain the second encoding, and construct the positive sample class and the negative sample class;

[0025] S303. Substitute the negative sample class and the positive sample class into the contrastive loss function, and further adjust the few-shot class model through the contrastive loss function;

[0026] S304. Classify and identify the augmented targets through a classifier;

[0027] S305. Substitute the result of the classification and identification into the classification loss function to further adjust the pre-trained MOCO model.

[0028] Furthermore, for step S301, the specific steps are: for each class \(c\in C\) base \(\cup C\) novel , randomly sample two images Img1, Img2, and perform augmentation on the two images respectively through data augmentation to obtain four images belonging to class c Pair them up in pairs to obtain 6 groups of positive sample pairs for supervision in contrastive learning.

[0029] Furthermore, for step S302, the specific steps are:

[0030] St01: For each image pair (Img q , Img k ) corresponding to each image Img, define two encoders f q , f k , with corresponding parameters \(\omega\) q , \(\omega\) k , where fk is a momentum copy of the slowly updated f q and does not participate in the direct gradient update, that is:

[0031] ω k ← mω k + (1 - m)ω q

[0032] where m ∈ [0, 1) is the momentum parameter;

[0033] St02: Encode the images Img q and Img k separately using two encoders:

[0034] E q = f q (Img q ), E k = f k (Img k )

[0035] E q , E k ∈ R n , obtaining the first encoding Then perform MLP mapping through a multi - layer perceptron:

[0036] E′ q = MLP q (E q ), E k ′ = MLP k (E k )

[0037] obtaining the second encoding Construct positive sample classes by taking data from the few - shot class dataset,

[0038] and queue the data belonging to the base class to construct negative sample classes with the few - shot classes.

[0039] Furthermore, in step S303, the contrast loss function is:

[0040]

[0041] where P + is the inner product between positive sample pairs, P - is the inner product between negative sample classes, and t is the temperature parameter. The larger t is, the flatter the distribution is.

[0042] Furthermore, in step S304, classify and identify the image pairs , specifically:

[0043]

[0044] Among them, f c is a classifier, and q(Img) is the prediction result of the classifier.

[0045] Furthermore, the classification loss function is as follows:

[0046]

[0047] Among them, x represents the data of all categories, which includes the base class and the few-shot class; p(x) is the label corresponding to x.

[0048] The beneficial effects of the present invention are as follows:

[0049] By dividing the given data into labeled data and unlabeled data, and dividing the labeled data set into a base class data set and a few-shot class data set according to types, the data is classified in advance, which is convenient for subsequent training and classification. First, the unlabeled data set and the base class data set are used to train the MOCO model to achieve simple pre-training, and then the pre-trained MOCO model is further fine-tuned through a contrastive learning framework to improve the accuracy of the MOCO model. Finally, the trained MOCO model is used to classify and identify the targets in the few-shot class data set. Thus, the present invention well solves the problems of poor recognition effect of few-shot data and high preprocessing work cost in the prior art. The present invention has the characteristics of being simple and practical, high in accuracy, and low in cost. Description of the Drawings

[0050] Figure 1 is a schematic flow chart of the few-shot image recognition method.

[0051] Figure 2 is a schematic diagram of the process of pre-training the MOCO model and further fine-tuning by combining contrastive learning in the few-shot image recognition method.

[0052] Figure 3 is a schematic diagram of the process of expanding the targets in the few-shot class data set through contrastive learning in the few-shot image recognition method.

[0053] Figure 4 is a schematic diagram of the change in the accuracy rate after adjustment of the few-shot image recognition method.

[0054] Figure 5 is a schematic diagram of the change in the classification loss function after adjustment of the few-shot image recognition method. Detailed Embodiments

[0055] The present invention will be described in detail below with reference to the drawings and specific embodiments.

[0056] Example 1

[0057] As Figure 1 shown, a few-shot image recognition method combining contrastive learning includes the following steps:

[0058] S1. Given a labeled dataset and an unlabeled dataset, divide the labeled dataset into a base class dataset and a few-shot class dataset by category;

[0059] S2. Train a MOCO model with the unlabeled dataset and the base class dataset to obtain a pre-trained MOCO model;

[0060] S3. Further fine-tune and train the pre-trained MOCO model through a contrastive learning framework;

[0061] S4. Classify and recognize the targets in the few-shot class dataset through the MOCO model.

[0062] The base class dataset and the few-shot class dataset in S1 do not intersect, that is where C base is the base class dataset, C novel is the few-shot class dataset, and each few-shot class data has n training samples, where n < 10.

[0063] In step S2, the MOCO model is a self-supervised model, including an encoder module, a multi-layer perceptron module, and a queue module; the encoder module processes the tensor corresponding to the image data to obtain data such as a feature matrix, and constructs the "key" in the unsupervised learning dictionary data to retrieve the corresponding data; the multi-layer perceptron module is used to obtain image features; the queue module completes the storage and maintenance of the dictionary data by setting queue rules.

[0064] Step S2 specifically includes the following steps:

[0065] S201. Collect images (Img) in the unlabeled dataset;

[0066] S202. Define the encoder as encoder f in the encoder module, and the corresponding parameter is ω;

[0067] S203. Encode the image Img with the encoder:

[0068] E = f(Img),

[0069] where E ∈ R n ;

[0070] S204. After obtaining the encoding, pass through a projection layer MLP of the multi-layer perceptron:

[0071] E′ = MLP(E);

[0072] S205. The described queue module adds the E′ obtained in each round of training to a queue after this round of training, maintains it according to the "first in, first out" rule to ensure the real-time nature of the queue, and uses the encoded data saved in the queue as negative samples for training in each round of training.

[0073] As Figure 2 and Figure 3 shown, step S3 specifically includes the following steps:

[0074] S301. Use contrastive learning to augment the targets in the few-shot class dataset;

[0075] S302. Map the first encoding of the augmented targets through the multi-layer perceptron projection layer MLP to obtain the second encoding, and construct the positive sample class and the negative sample class;

[0076] S303. Substitute the described negative sample class and positive sample class into the contrastive loss function, and further adjust the few-shot class model through the contrastive loss function;

[0077] S304. Classify and identify the augmented targets through a classifier;

[0078] S305. Substitute the result of the classification and identification into the classification loss function to further adjust the pre-trained MOCO model.

[0079] For step S301, the specific steps are: For each class c ∈ C base ∪C novel , randomly sample two images Img1, Img2, and augment each of the two images through data augmentation to obtain four images belonging to class c Pair them up in pairs to obtain 6 groups of positive sample pairs for supervision in contrastive learning.

[0080] For step S302, the specific steps are:

[0081] St01: For each image pair (Img q , Img k ) corresponding to each image Img, define two encoders f q , f k , with corresponding parameters ω q , ω k , where f k is a momentum copy of f q that is slowly updated and does not participate in direct gradient updates, that is:

[0082] ωk ←mω k +(1 - m)ω q

[0083] where m ∈ [0, 1) is the momentum parameter;

[0084] St02: For the image Img q and Img k are encoded by two encoders respectively:

[0085] E q = f q (Img q ), E k = f k (Img k )

[0086] E q , E k ∈ R n , obtaining the first encoding Then, perform MLP mapping through a multi - layer perceptron:

[0087] E' q = MLP q (E q ), E k ' = MLP k (E k )

[0088] obtaining the second encoding Take the data in the few - shot class dataset to construct the positive sample class,

[0089] and enqueue the data belonging to the base class to construct the negative sample class with the few - shot class.

[0090] In step S303, the contrast loss function is:

[0091]

[0092] where P + is the inner product between positive sample class pairs, P - is the inner product between negative sample classes, and t is the temperature parameter. The larger t is, the flatter the distribution is.

[0093] In step S304, classify and identify the image pair specifically as:

[0094]

[0095] where, f c is the classifier, and q(Img) is the prediction result of the classifier.

[0096] The classification loss function is as follows:

[0097]

[0098] Among them, x represents the data of all categories, which includes the base class and the few-shot class; p(x) is the label corresponding to x.

[0099] Compared with the prior art, the model structure provided by the present invention can more effectively obtain the features to be extracted for the few-shot class, and better make up for the limitation of the data volume in this task; First, the unlabeled dataset and the base class dataset are used to train the MOCO model to achieve simple pre-training, and then the pre-trained MOCO model is further fine-tuned through a contrastive learning framework. On the basis of data augmentation in the MOCO framework, a contrastive learning method is added to perform data augmentation on two image samples respectively to improve the accuracy of the MOCO model. Finally, the trained MOCO model is used to classify and identify the targets in the few-shot class dataset. This method reduces the requirements for sample data and the dependence on data distribution in few-shot learning to a certain extent. Therefore, the present invention solves well the problems of poor recognition effect of few-shot data and high preprocessing work cost in the prior art, and provides a few-shot image recognition method combined with contrastive learning, which has the characteristics of simplicity, high accuracy and low cost.

[0100] Embodiment 2

[0101] As Figure 1 shown, a few-shot image recognition method combined with contrastive learning includes the following steps:

[0102] S1. Given a labeled dataset and an unlabeled dataset, divide the labeled dataset into a base class dataset and a few-shot class dataset according to categories;

[0103] S2. Train the MOCO model through the unlabeled dataset and the base class dataset to obtain a pre-trained MOCO model;

[0104] S3. Further fine-tune the pre-trained MOCO model through a contrastive learning framework;

[0105] S4. Classify and identify the targets in the few-shot class dataset through the MOCO model.

[0106] In S1, the base class dataset and the few-shot class dataset do not intersect, that is where C base is the base class dataset, C novel is the few-shot class dataset, and each few-shot class data has n training samples, where n < 10.

[0107] In step S2, the MOCO model is a self-supervised model, including an encoder module, a multi-layer perceptron module, and a queue module. The encoder module processes the tensor corresponding to the image data to obtain data such as a feature matrix, and constructs the "keys" in the unsupervised learning dictionary data to retrieve the corresponding data. The multi-layer perceptron module is used to obtain image features. The queue module completes the storage and maintenance of the dictionary data by setting queue rules.

[0108] Step S2 specifically includes the following steps:

[0109] S201. Collect the images (Img) in the unlabeled dataset;

[0110] S202. Define the encoder as encoder f in the encoder module, and the corresponding parameter is ω;

[0111] S203. Encode the image Img with the encoder:

[0112] E = f(Img),

[0113] where E ∈ R n ;

[0114] S204. After obtaining the encoding, pass it through the projection layer MLP of a multi-layer perceptron:

[0115] E' = MLP(E);

[0116] The queue module adds the E' obtained in each round of training to a queue after this round of training, and maintains it according to the "first in, first out" rule to ensure the real-time nature of the queue, and uses the encoding saved in the queue as negative samples for training in each round of training.

[0117] As Figure 2 and Figure 3 shown, step S3 specifically includes the following steps:

[0118] S301. Use contrastive learning to expand the targets in the few-shot class dataset;

[0119] S302. Map the first encoding of the expanded target through the projection layer MLP of the multi-layer perceptron to obtain the second encoding, and construct positive and negative sample classes;

[0120] S303. Substitute the negative and positive sample classes into the contrastive loss function, and further adjust the few-shot class model through the contrastive loss function;

[0121] S304. Classify and identify the expanded targets through a classifier;

[0122] S305. Substitute the result of classification recognition into the classification loss function to further adjust the pre-trained MOCO model.

[0123] As Figure 4 and Figure 5 shown, in this embodiment, the accuracy of the adjusted MOCO model has increased significantly, while the classification loss has decreased significantly.

[0124] Step S301, the specific steps are: for each category c ∈ C base ∪C novel , randomly sample two images Img1, Img2, and augment the two images respectively through data augmentation to obtain four images belonging to category c Pair them in pairs to obtain 6 groups of positive sample pairs for supervision in contrastive learning.

[0125] Step S302, the specific steps are:

[0126] St01: For each image pair (Img q , Img k ) corresponding to each image Img, define two encoders f q , f k , and the corresponding parameters are ω q , ω k , where f k is a momentum copy of f q that is slowly updated and does not participate in direct gradient update, that is:

[0127] ω k ← mω k + (1 - m)ω q

[0128] where m ∈ [0, 1) is the momentum parameter;

[0129] St02: Encode the image Img q and Img k respectively with two encoders:

[0130] E q = f q (Img q ), E k = f k (Img k )

[0131] E q , E k ∈ R n , and obtain the first encoding Then perform MLP mapping through a multi-layer perceptron:

[0132] E′ q = MLP q (Eq), E k ′ = MLP k (E k )

[0133] Obtain the second encoding Take the data in the few-shot class dataset to construct the positive sample class, and put the data belonging to the base class into the queue to construct the negative sample class with the few-shot class.

[0134] In step S303, the contrast loss function is as follows:

[0135]

[0136] where P + is the inner product between positive sample class pairs, P - is the inner product between negative sample classes, and t is the temperature parameter. The larger t is, the flatter the distribution is.

[0137] In step S304, classify and identify the image pair specifically as follows:

[0138]

[0139] where f c is the classifier, and q(Img) is the prediction result of the classifier.

[0140] The classification loss function is as follows:

[0141]

[0142] where x represents the data of all classes, including the base class and the few-shot class; p(x) is the label corresponding to x.

[0143] Compared with the prior art, the model structure provided by the present invention can more effectively obtain the features to be extracted for small-sample classes, and better make up for the limitation of the data volume in this task; First, the unlabeled dataset and the base-class dataset are used to train the MOCO model to achieve simple pre-training, and then the pre-trained MOCO model is further fine-tuned through a contrastive learning framework. On the basis of data augmentation in the MOCO framework, a contrastive learning method is added to perform data augmentation on two image samples respectively, improving the accuracy of the MOCO model. Finally, the trained MOCO model is used to classify and identify the targets in the few-shot class dataset. This method reduces the requirements for sample data and the dependence on data distribution in few-shot learning to a certain extent. Therefore, the present invention well solves the problems of poor recognition effect of few-shot data and high preprocessing work cost in the prior art, and provides a few-shot image recognition method combined with contrastive learning, which has the characteristics of simplicity, high accuracy and low cost.

[0144] Embodiment 3

[0145] As Figure 1 shown, a few-shot image recognition method combined with contrastive learning includes the following steps:

[0146] S1. Given a labeled dataset and an unlabeled dataset, divide the labeled dataset into a base-class dataset and a few-shot class dataset according to categories;

[0147] S2. Train a MOCO model with the unlabeled dataset and the base-class dataset to obtain a pre-trained MOCO model;

[0148] S3. Further fine-tune the pre-trained MOCO model through a contrastive learning framework;

[0149] S4. Classify and identify the targets in the few-shot class dataset through the MOCO model.

[0150] The base-class dataset and the few-shot class dataset in S1 do not intersect, that is where C base is the base-class dataset, C novel is the few-shot class dataset, and each few-shot class data has n training samples, where n < 10.

[0151] In step S2, the MOCO model is a self-supervised model, including an encoder module, a multi-layer perceptron module, and a queue module; the encoder module processes the tensor corresponding to the image data to obtain data such as a feature matrix, and constructs the "keys" in the unsupervised learning dictionary data to retrieve the corresponding data; the multi-layer perceptron module is used to obtain image features; the queue module completes the storage and maintenance of the dictionary data by setting queue rules.

[0152] Step S2 specifically includes the following steps:

[0153] S201. Collect the images (Img) in the unlabeled dataset;

[0154] S202. Define the encoder as encoder f in the encoder module, and the corresponding parameter is ω;

[0155] S203. Encode the image Img with the encoder:

[0156] E = f(Img),

[0157] where E ∈ R n ;

[0158] S204. After obtaining the encoding, pass it through the projection layer MLP of a multi-layer perceptron:

[0159] E' = MLP(E);

[0160] The queue module adds the E' obtained in each round of training to a queue after this round of training, and maintains it according to the "first in, first out" rule to ensure the real-time nature of the queue, and uses the encoding saved in the queue as negative samples for training in each round of training.

[0161] As Figure 2 and Figure 3 shown, step S3 specifically includes the following steps:

[0162] S301. Use contrastive learning to expand the targets in the few-shot class dataset;

[0163] S302. Map the first encoding of the expanded target through the projection layer MLP of the multi-layer perceptron to obtain the second encoding, and construct positive and negative sample classes;

[0164] S303. Substitute the negative and positive sample classes into the contrastive loss function, and further adjust the few-shot class model through the contrastive loss function;

[0165] S304. Classify and identify the expanded targets through a classifier;

[0166] S305. Substitute the result of classification recognition into the classification loss function to further adjust the pre-trained MOCO model.

[0167] Step S301, the specific steps are as follows: For each category c ∈ C base ∪C novel , randomly sample two images Img1 and Img2, and augment the two images respectively through data augmentation to obtain four images belonging to category c Pair them up in pairs to obtain 6 groups of positive sample pairs for supervision in contrastive learning.

[0168] Step S302, the specific steps are as follows:

[0169] St01: For each image pair (Img q , Img k ) corresponding to each image Img, define two encoders f q , f k , and the corresponding parameters are ω q , ω k , where f k is a momentum copy of f q that is slowly updated and does not participate in direct gradient update, that is:

[0170] ω k ←mω k +(1 - m)ω q

[0171] where m ∈ [0, 1) is the momentum parameter;

[0172] St02: Encode the image Img q and Img k respectively with the two encoders:

[0173] E q =f q (Img q ), E k =f k (Img k )

[0174] E q , E k ∈R n , to obtain the first encoding Then perform MLP mapping through a multi-layer perceptron:

[0175] E′ q =MLP q (E q ), E k ′=MLP k (E k)

[0176] Obtain the second encoding Construct positive sample classes by taking data from the few-shot sample dataset,

[0177] and those belonging to the base class are enqueued to construct negative sample classes with the few-shot classes.

[0178] In step S303, the contrast loss function is as follows:

[0179]

[0180] where P + is the inner product between positive sample class pairs, P - is the inner product between negative sample classes, and t is the temperature parameter. The larger t is, the flatter the distribution is.

[0181] In step S304, the image pairs are classified and recognized, specifically as follows:

[0182]

[0183] where f c is the classifier, and q(Img) is the prediction result of the classifier.

[0184] The classification loss function is as follows:

[0185]

[0186] where x represents data of all classes, including the base class and the few-shot classes; p(x) is the label corresponding to x.

[0187] As shown in the following table, compared with the prior art, the model structure provided by the present invention can efficiently identify and classify the minority sample data of one sample and five samples, wherein the recognition ratio of the minority sample class of the five samples in each data set exceeds 60%, and the features to be extracted of the minority sample class are more effectively obtained, which better compensates for the limitation of the amount of data in this task; firstly, the unlabeled data set and the basic class data set are used to train the MOCO model to achieve simple pre-training, and then the pre-trained MOCO model is further fine-tuned and trained through the contrastive learning framework, and the contrastive learning method is added on the basis of the data augmentation of the MOCO framework, and the data augmentation of the two image samples is achieved separately to improve the accuracy of the MOCO model, and finally the trained MOCO model is used to classify and identify the targets in the minority sample class data set. This method reduces the requirements of the minority sample learning for the sample data and the dependence on the data distribution to a certain extent, so that the present invention solves the problems of the prior art in poor recognition effect of minority sample data and high cost of pre-processing work, and proposes a minority sample image recognition method combined with contrastive learning, which is simple and practical, high in accuracy and low in cost.

[0188]

[0189] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, and are not intended to limit the implementation methods of the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. A few-shot image recognition method combined with contrastive learning, characterized in that: Including the following steps: S1. Given a labeled dataset and an unlabeled dataset, divide the labeled dataset into a base class dataset and a few-shot class dataset according to categories; S2. Train a MOCO model with the unlabeled dataset and the base class dataset to obtain a pre-trained MOCO model; S201. Collect the images in the unlabeled dataset ; S202. Define the encoder in the encoder module as , and the corresponding parameter is ; S203. Encode the image using an encoder: , Among them, ; S204. After obtaining the encoding, pass it through a projection layer MLP of a multi-layer perceptron: ; S205. The queue module adds the obtained in each round of training to a queue after this round of training, maintains it according to the "first in, first out" rule, ensures the real-time nature of the queue, and uses the encoding saved in the queue as negative samples for training in each round of training;​​ S3. Further fine-tune and train the pre-trained MOCO model through a contrastive learning framework; S301. Use contrastive learning to augment the targets in the few-shot class dataset; S302. Map the first encoding of the augmented targets through a multi-layer perceptron projection layer MLP to obtain a second encoding, and construct positive sample classes and negative sample classes; S303. Substitute the negative sample classes and positive sample classes into a contrastive loss function, and further adjust the few-shot class model through the contrastive loss function; S304. Classify and identify the augmented targets through a classifier; S305. Substitute the results of the classification and identification into a classification loss function to further adjust the pre-trained MOCO model; S4. Classify and identify the targets in the few-shot class dataset through the MOCO model.

2. The few-shot image recognition method combining contrastive learning according to claim 1, wherein: The basic class dataset and the few-shot class dataset described in S1 do not intersect , where is the basic class dataset, is the few-shot class dataset, and each of the few-shot class data has training samples, .

3. The few-shot image recognition method combining contrastive learning according to claim 1, wherein: In step S2, the MOCO model is a self-supervised model, including an encoder module, a multi-layer perceptron, and a queue module; the encoder module processes the tensor corresponding to the image data to obtain feature matrix data and constructs the "keys" in the unsupervised learning dictionary data to retrieve the corresponding data; the multi-layer perceptron is used to obtain image features; The queue module completes the storage and maintenance of the dictionary data by setting queue rules.

4. The few-shot image recognition method combined with contrastive learning according to claim 1, wherein: Step S301, the specific steps are as follows: For each category , randomly sample two images , and augment the two images respectively through data augmentation to obtain four images belonging to the category . Pair them in pairs to obtain 6 pairs of positive samples for supervision in contrastive learning. ​ 5. The few-shot image recognition method combining contrastive learning according to claim 4, wherein: Step S302, the specific steps are: St01: For each image The corresponding image pair , define two encoders , and the corresponding parameters are respectively , where is a momentum copy that is slowly updated and does not participate in direct gradient updates: wherein is the momentum parameter; St02: For the image and are respectively encoded by two encoders: Obtain the first encoding ; then perform MLP mapping through a multi-layer perceptron: Obtain the second encoding , take the data in the few-shot class dataset to construct the positive sample class, and put the belonging to the base class into the queue to construct the negative sample class with the few-shot class.

6. The few-shot image recognition method combined with contrastive learning according to claim 5, characterized in that: In step S303, the contrastive loss function is: where is the inner product between positive sample pairs, is the inner product between negative sample classes, and t is the temperature parameter. The larger t is, the flatter the distribution is.

7. The few-shot image recognition method combined with contrastive learning according to claim 6, wherein: In step S304, the image pair mentioned above is classified and recognized, specifically as follows: Among them, is a classifier, is the prediction result of the classifier.

8. The few-shot image recognition method combined with contrastive learning according to claim 7, characterized in that: The classification loss function is: Among them, represents data of all categories, which includes base classes and few-shot classes; is the corresponding label.