Small sample image classification method based on adaptive effective semantic expansion
By adopting an adaptive effective semantic expansion method based on triple network architecture in small sample image classification, combined with strong and weak enhancement joint learning and cluster allocation, the problems of insufficient data volume and insufficient structural variation in small sample learning are solved, and efficient unsupervised small sample image classification is achieved.
Patent Information
- Application Number
- CN202510129536.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-30
AI Technical Summary
The existing deep clustering method for contrast learning is not sufficiently mined for finite data in small sample learning, especially when constructing positive sample pairs, there are problems of insufficient variability and 'false positive'.
A small sample image classification method based on adaptive effective semantic augmentation is adopted, and an effective semantic extraction framework based on joint enhancement of triple network architecture is constructed, combined with strong and weak enhancement joint learning and cluster allocation, unsupervised model training is carried out. The method includes key semantic positioning, adaptive cleaning steps, and nearest neighbor semantic mining to improve the generalization capability of the model.
This method can adaptively perform effective sample enhancement, alleviate the problem of insufficient data volume of small sample tasks, and reduce the probability of "false positive" in the construction process of positive sample pairs, realizing the effective expansion of unsupervised small sample image classification.
Smart Images

Figure CN120070967A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a small sample image classification method based on adaptive effective semantic expansion, and belongs to the field of deep learning. Background Art
[0002] Few-shot learning is a machine learning method that aims to enable models to learn and infer when data is scarce. In traditional machine learning, models usually require a large amount of labeled data for training in order to perform well at test time. However, in many practical application scenarios, obtaining a large number of labeled samples may be difficult or uneconomical. Few-shot learning reduces the reliance on a large number of training samples by adopting different strategies. Typical techniques include meta-learning, transfer learning, metric learning, and data augmentation.
[0003] Unsupervised learning extracts representations from data without additional labels and has broad application prospects. Recently, contrastive learning with instance discrimination as a proxy task has become the dominant method for unsupervised representation learning. It constructs positive sample pairs through data augmentation. Common augmentation methods include random cropping, rotation, coloring, etc. These methods usually enhance samples of the same category to construct positive sample pairs with various data augmentations and optimize the contrast loss. The purpose is to bring enhancements with similar instances closer in the embedding space and away from enhancements with low similarity. After optimization, the constructed proxy task can learn useful representations without manual annotation.
[0004] The data enhancement method is a method that expands sample data in various ways to increase the number of training samples. This is a direct and effective method in small sample learning. The basic idea of this method is to generate new and diverse training samples by performing a series of transformations and expansions on the existing sample data. These transformations can be simple geometric operations such as rotation and flipping, or complex pixel-level operations such as color conversion and noise addition. Through these transformations, more data scenarios can be simulated without increasing the actual data collection cost, allowing the model to mine more implicit semantic information during the training process.
[0005] In current contrastive learning deep clustering methods, some weak enhancements are usually used to form positive sample pairs. This method is not sufficient to mine limited data in small sample learning. Summary of the invention
[0006] This application proposes a small sample image classification method based on adaptive effective semantic expansion, which belongs to an implementation of contrastive learning and deep clustering in small sample learning for feature enhancement in the field of deep learning, and specifically involves a pseudo-supervised clustering method based on meta-features. Combined with the construction of proxy tasks, a combination of multiple enhancements is applied to improve the generalization ability of the model.
[0007] To solve the above technical problems, the technical solution adopted by the present invention is a small-sample image classification method based on adaptive effective semantic augmentation, including the following steps:
[0008] 1) Adopt key semantic localization to adaptively enhance the feature information of new categories;
[0009] 2) Perform secondary optimization on semantic features;
[0010] 3) Introduce an adaptive cleaning step to eliminate the noise generated during enhancement;
[0011] In step 1), construct an effective semantic extraction framework based on the joint enhancement of a triple network architecture;
[0012] Execute a strategy of one-way strong enhancement and two-way weak enhancement on the original data set, and input it into contrastive learning at the instance level and the clustering level; perform model training in a purely unsupervised manner;
[0013] In step 2), perform nearest neighbor semantic mining;
[0014] Use Conaug(·) to perform two-way weak augmentation on the original samples, and the two generated related samples are denoted as And
[0015] The network will enhance the samples and and convert them into feature representations and
[0016] In step 3), for each input image, the first-way and second-way weak enhancements form a positive sample pair, and the third-way strong enhancement and the second-way weak enhancement samples form another positive sample pair;
[0017] Form negative sample pairs between the enhanced samples from different input images;
[0018] In step 3), the optimization of the loss function consists of an instance-level loss and a clustering-level loss.
[0019] Optimized, for the above small-sample image classification method based on adaptive effective semantic augmentation, in step 1), the backbone network ResNet of the constructed triple network joint model inputs the data set into the triple network joint model for pre-training, and the training data set is an image data set.
[0020] Optimized, for the above small-sample image classification method based on adaptive effective semantic augmentation, in step 1), the triple network performs one type of strong enhancement and two types of weak enhancements on each input image among the given N image samples, denoted as xi , with N strongly augmented and 2·N weakly augmented samples;
[0021] Instance-level projection Transform into where i ∈ [1, N] and j ∈ {1, 2, 3};
[0022] Cluster-level projection Transform into
[0023] Construct two types of feature matrices via two projectors respectively;
[0024] Perform unsupervised network training by simultaneously optimizing the instance-level contrast loss and the cluster-level contrast loss.
[0025] Optimized, in the above small-sample image classification method based on adaptive effective semantic augmentation, in step 2), introduce the support set S for nearest neighbor mining;
[0026] From the Obtain the nearest neighbor information and feed into the instance-level projection and the cluster-level projection to obtain the nearest neighbor features and
[0027] Optimized, in the above small-sample image classification method based on adaptive effective semantic augmentation, in step 3), use the cosine similarity function to measure the similarity between sample pairs, and the cosine similarity function is defined as: where u and v represent two feature vectors, and s(·, ) is the cosine similarity function used to calculate the distance between two vectors.
[0028] Optimized, in the above small-sample image classification method based on adaptive effective semantic augmentation, in step 3), optimize the consistency of the contrast pairs constructed by two augmentations while minimizing the consistency of negative sample pairs; the instance-level contrast loss of the augmented sample pairs is:
[0029]
[0030] where, i ∈ [1, N], j ∈ {1, 2, 3}, τ R is the temperature coefficient.
[0031] Optimized, the above-mentioned few-shot image classification method based on adaptive effective semantic augmentation calculates all the losses on the triple network architecture, expands the instance-level contrast loss from one pair to multiple pairs, and identifies all positive sample pairs;
[0032] All the losses on the triple network architecture are expressed as
[0033] Based on reconstructing instance-level positive sample pairs, the first and second weak augmentations are used to form positive sample pairs, and the third strong augmentation and its second weak augmentation samples form another positive sample pair strategy. The instance-level contrast loss is defined as L ins = L ins (t 1 , t 2 ) + L ins (t 2 , t 3 )
[0034] where t 1 , t 1 and t 3 are three-way data augmentations.
[0035] Optimized, the above-mentioned few-shot image classification method based on adaptive effective semantic augmentation defines the cluster-level contrast loss as:
[0036]
[0037] where i ∈ [1, M], j ∈ {1, 2, 3}, τ C is the temperature coefficient.
[0038] Optimized, the above-mentioned few-shot image classification method based on adaptive effective semantic augmentation, after traversing all clusters, the cluster-level contrast loss is further expressed as:
[0039]
[0040] where H(Y) is the entropy function of the cluster assignment probability;
[0041]
[0042] where m is the sample number, and P(·) is to avoid the trivial solution that most samples are assigned to the same cluster;
[0043] The cluster-level contrast loss is defined as:
[0044] L clu = L clu (t 1 , t 2 ) + L clu(t 2 ,t 3 )
[0045] The total loss function consists of an instance-level contrastive loss and a clustering-level contrastive loss, i.e.:
[0046] L = L ins + L clu
[0047] For the optimized few-shot image classification method based on adaptive effective semantic augmentation, in step 1), the backbone network ResNet of the constructed triple network joint model is used to input the dataset into the triple network joint model for pre-training, and the training datasets are all image datasets;
[0048] In step 1), in the contrastive learning at the instance level and the clustering level, the model is trained in a purely unsupervised manner;
[0049] In step 2), the implementation of nearest neighbor semantic mining is to perform two-way weak augmentation on the original samples by Conaug(·) to generate two related samples, and introduce the obtain the nearest neighbor information
[0050] In step 2), the of nearest neighbor semantic mining is fed into the instance-level projection and the clustering-level projection to obtain the nearest neighbor features and The beneficial effects of this application are as follows:
[0051] The technical solution of this application simultaneously utilizes strong and weak augmentation joint learning representations and clustering assignments in a network with multiple augmented views. It can adaptively perform effective sample augmentation, alleviating the problems of insufficient data volume in few-shot tasks and "false positives" in the construction process of positive sample pairs in complex proxy tasks.
[0052] The technical solution of this application solves three aspects of problems: First, it takes key semantic localization to adaptively augment the feature information of new categories; second, it introduces an adaptive cleaning step to eliminate the noise generated during augmentation; finally, it realizes effective augmentation for unsupervised few-shot image classification through an effective semantic extraction framework jointly augmented by a triple network architecture.
[0053] The technical solution of this application proposes an effective semantic extraction framework based on a triple network architecture joint enhancement, and simultaneously utilizes strong and weak enhancements to jointly learn representations and cluster assignments in a network with multiple enhanced views. A model with a triple shared network is adopted, which combines one strong enhancement and two weak enhancements into contrastive learning at the instance level and cluster level, and the model can be trained in a purely unsupervised manner.
[0054] The technical solution of this application, by adding a strong data enhancement module, alleviates the problem of insufficient variability of positive sample pairs caused by the expansion of the same image; adopts key semantic positioning to adaptively generate more suitable blocks to mine more fine-grained semantic information between instances; and introduces the nearest neighbor mining of the support set into the comparative clustering framework, thereby fully mining the scarce data information in small sample tasks. The adaptive and effective data enhancement method not only fully expands the data and alleviates the problem of insufficient data volume in small sample tasks, but also reduces the probability of "false positives" in the construction process of positive sample pairs in complex proxy tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 This is a schematic diagram of the principle of the small sample image classification method based on adaptive effective semantic expansion of the present application;
[0056] Figure 2 Schematic diagram of the principle of the nearest neighbor semantic mining module in the embodiment. DETAILED DESCRIPTION
[0057] The technical features of the present invention are further described below in conjunction with specific embodiments.
[0058] The present invention provides a small sample image classification method based on adaptive effective semantic expansion, comprising:
[0059] Construct an effective semantic extraction framework based on the joint enhancement of a triple network architecture, which can simultaneously utilize strong and weak enhancements to jointly learn representations and clustering assignments in a network with multiple enhanced views. Incorporate one-way strong enhancement and two-way weak enhancement into contrastive learning at the instance level and the clustering level, and train the model in a purely unsupervised manner. Specifically, first, add a strong data augmentation module to alleviate the problem of insufficient variability in positive sample pairs caused by augmenting the same image. Second, use key semantic localization to adaptively generate more appropriate blocks and mine finer-grained semantic information between instances. Finally, by introducing nearest neighbor mining of the support set into the contrastive clustering framework, scarce data information in few-shot tasks is fully exploited. Generally speaking, this adaptive method of effective data augmentation not only fully expands the data and alleviates the problem of insufficient data volume in few-shot tasks, but also reduces the probability of "false positives" in the construction of positive sample pairs for complex proxy tasks. This method is a further exploration of imitating the human learning method and is of great significance.
[0060] Specifically, the few-shot image classification method based on adaptive effective semantic augmentation in this application includes the following steps:
[0061] 1) Construct an effective semantic extraction framework based on the joint enhancement of a triple network architecture, which jointly learns representations and clustering assignments using strong and weak enhancements simultaneously in a network with multiple enhanced views. Execute a strategy of one-way strong augmentation and two-way weak augmentation on the original dataset and input it into contrastive learning at the instance level and the clustering level, and the model can be trained in a purely unsupervised manner. The triple network performs one type of strong augmentation and two types of weak augmentations on each input image from a given N image samples, denoted as x i , with N strong augmentations and 2·N weak augmentation samples.
[0062] Instance-level projection Project to where i ∈ [1, N] and j ∈ {1, 2, 3}. Clustering-level projection Project to Construct two types of feature matrices via two projectors respectively. Thereafter, unsupervised network training can be performed by simultaneously optimizing the instance-level contrast loss and the clustering-level contrast loss. In Figure 1 , the model framework consists of three parts: a triple architecture network architecture with strong and weak enhancements, nearest neighbor representation mining, instance-level projection, and clustering-level projection.
[0063] 2) The implementation of nearest neighbor semantic mining is to first perform two-way weak augmentation on the original samples using Conaug(·) to generate two related samples, denoted as And After that, the network will transform the augmented samples and into feature representations respectively and
[0064] In addition, a support set S is introduced for nearest neighbor mining. It obtains the nearest neighbor information from the support set S and feeds it into the instance-level projection and the cluster-level projection to obtain the nearest neighbor features and
[0065] In Figure 2 the implementation of nearest neighbor semantic mining, the original samples are first subjected to two-way weak augmentation with Conaug(·) to generate two related samples, denoted as and After that, the network will transform the augmented samples and into feature representations respectively and In addition, a support set S is introduced for nearest neighbor mining. It obtains the nearest neighbor information from the support set S and feeds it into the instance-level projection and the cluster-level projection to obtain the nearest neighbor features and
[0066] 3) The loss function consists of an instance-level loss and a cluster-level loss. For each input image, the first-way and second-way weak augmentations form a positive sample pair, and the third-way strong augmentation and the second-way weak augmentation sample form another positive sample pair. Meanwhile, negative pairs are formed between the augmented samples from different input images. The purpose of defining the above positive and negative sample pairs is to maximize the consistency of the positive sample pairs while increasing the distance of the negative pairs, so as to obtain good representations.
[0067] To measure the similarity of instance pairs, the cosine similarity function is used to measure the similarity between sample pairs, which is defined as:
[0068]
[0069] where u and v represent two feature vectors.
[0070] To optimize the two augmentations (e.g., augmentation t 1 and augmentation t 2) Consistency of the constructed contrast pairs while minimizing the consistency of negative sample pairs. The instance-level contrast loss of the enhanced sample pairs is:
[0071] For \(i\in[1,N]\), \(j\in\{1,2,3\}\), \(\tau\) R is the temperature coefficient.
[0072] To expand the instance-level contrast loss from one pair to multiple pairs and identify all positive sample pairs, we calculate all losses on the triple network architecture.
[0073]
[0074] Since the reconstructed instance-level positive sample pairs are formed by the first and second weak augmentations to form positive sample pairs, and the third strong augmentation and its second weak augmentation samples form another positive sample pair strategy. Therefore, the instance-level contrast loss is defined as:
[0075] \(L\) ins \(=\) \(L\) ins \((t\) 1 , \(t\) 2 ) \(+\) \(L\) ins \((t\) 2 , \(t\) 3 ) #
[0076] where \(t\) 1 , \(t\) 1 and \(t\) 3 are triple data augmentations.
[0077] Similarly, the cluster-level contrast loss can be defined as:
[0078]
[0079] where, for \(i\in[1,M]\), \(j\in\{1,2,3\}\), \(\tau\) C is the temperature coefficient.
[0080] After traversing all clusters, the cluster-level contrast loss can be further expressed as:
[0081]
[0082] where \(H(Y)\) is the entropy function of the cluster assignment probability:
[0083]
[0084] where \(P(\cdot)\) is to avoid the trivial solution that most samples are assigned to the same cluster. Therefore, the cluster-level contrast loss is defined as:
[0085] \(L\)clu = L clu (t 1 , t 2 ) + L clu (t 2 , t 3 )
[0086] The total loss function consists of instance-level contrastive loss and cluster-level contrastive loss, that is:
[0087] L = L ins + L clu .
[0088] The above-mentioned few-shot image classification method based on adaptive effective semantic augmentation solves three aspects of problems: First, key semantic localization is adopted to adaptively enhance the feature information of new categories; second, an adaptive cleaning step is introduced to eliminate the noise generated during enhancement; finally, an effective semantic extraction framework jointly enhanced by a triple network architecture realizes effective augmentation of unsupervised few-shot image classification.
[0089] In step 1), the backbone network ResNet of the constructed triple network joint model is used to pre-train the dataset by inputting it into the triple network joint model, and the training datasets are all image datasets.
[0090] In step 1), in the contrastive learning at the instance level and the cluster level, the model can be trained in a purely unsupervised manner. In Figure 1 , the model framework consists of a triple architecture network architecture with strong and weak augmentations, nearest neighbor representation mining, instance-level projection, and cluster-level projection.
[0091] In step 2), the implementation of nearest neighbor semantic mining is to perform two-way weak augmentation on the original samples by Conaug(·) to generate two related samples, and introduce the to obtain the nearest neighbor information
[0092] In step 2), the of nearest neighbor semantic mining is fed into the instance-level projection and the cluster-level projection to obtain the nearest neighbor features and In Figure 2 , the implementation of nearest neighbor semantic mining is to first perform two-way weak augmentation on the original samples by Conaug(·) to generate two related samples, denoted as And After that, the network will enhance the samples and respectively convert them into feature representations and In addition, a support set S is introduced for nearest neighbor mining. It obtains nearest neighbor information and feeds it into the instance-level projection and the cluster-level projection to obtain nearest neighbor features and
[0093] In step 3), the optimization of the loss function consists of the instance-level loss and the cluster-level loss.
[0094] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions, or substitutions made by those of ordinary skill in the art within the scope of the essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. A small sample image classification method based on adaptive effective semantic expansion, characterized by: The following steps are involved: 1) Adopt key semantic positioning to adaptively enhance the feature information of new categories; 2) Perform secondary optimization on semantic features; 3) Introducing an adaptive cleaning step to remove noise generated during enhancement; In step 1), an effective semantic extraction framework based on the joint enhancement of the triple network architecture is constructed; The original dataset is subjected to one-way strong enhancement and two-way weak enhancement strategies and then input into instance-level and cluster-level contrastive learning. Model training in a purely unsupervised manner; In step 2), nearest neighbor semantic mining is performed; The original sample is weakly amplified in two ways using Conaug (·), and the two related samples generated are expressed as and The network will enhance the sample and Convert to feature representation and In step 3), for each input image, the first path and the second path of weak enhancement form a positive sample pair, and the third path of strong enhancement and the second path of weak enhancement samples form another positive sample pair; Form negative sample pairs between augmented samples from different input images; In step 3), the optimization of the loss function consists of instance-level loss and cluster-level loss.
2. The small sample image classification method based on adaptive effective semantic expansion according to claim 1 is characterized by: In step 1), a backbone network ResNet of the triple network joint model is constructed, and a data set is input into the triple network joint model for pre-training, and the training data set is an image data set.
3. The small sample image classification method based on adaptive effective semantic expansion according to claim 1 is characterized by: In step 1), the triple network performs one type of strong enhancement and two types of weak enhancement on each input image given N image samples, denoted as x i , with N strong enhancements and 2·N weak enhancement samples; Instance-level projection Will Transformed to i∈[1,N] and j∈{1,2,3}; Cluster-level projection Will Transformed to Two types of feature matrices are constructed through two projectors respectively; Unsupervised network training is performed by simultaneously optimizing both instance-level contrastive loss and cluster-level contrastive loss.
4. The small sample image classification method based on adaptive effective semantic expansion according to claim 1, characterized in that: In step 2), the support set S is introduced to perform nearest neighbor mining; From the support set S Get nearest neighbor information and will Feeding to instance-level projection and cluster-level projection To obtain the nearest neighbor features and 5. The small sample image classification method based on adaptive effective semantic expansion according to claim 1, characterized in that: In step 3), the cosine similarity function is used to measure the similarity between sample pairs. The cosine similarity function is defined as: where u and v represent two feature vectors, and s(·,·) is the cosine similarity function, which is used to calculate the distance between two vectors.
6. The small sample image classification method based on adaptive effective semantic expansion according to claim 1, characterized in that: In step 3), Optimize the consistency of the contrast pairs constructed by the two enhancements, while minimizing the consistency of the negative sample pairs; the instance-level contrast loss of the enhanced sample pairs is: Among them, i∈[1,N], j∈{1,2,3}, τ R is the temperature coefficient.
7. The small sample image classification method based on adaptive effective semantic expansion according to claim 1 is characterized by: By calculating all losses on the triple network architecture, we expand the instance-level contrastive loss from one pair to multiple pairs to identify all positive sample pairs; All losses on the triplet network architecture are expressed as Based on the strategy of reconstructing instance-level positive sample pairs, the first and second paths are weakly enhanced to form a positive sample pair, and the third path is strongly enhanced and its second path is weakly enhanced to form another positive sample pair. The instance-level contrast loss is defined as L ins =L ins (t1,t2)+L ins (t2,t3) Among them, t1, t1 and t3 are three-way data enhancement.
8. The small sample image classification method based on adaptive effective semantic expansion according to claim 1, characterized in that: The cluster-level contrastive loss is defined as: Among them, i∈[1,M], j∈{1,2,3}, τ C is the temperature coefficient.
9. The small sample image classification method based on adaptive effective semantic expansion according to claim 1, characterized in that: After traversing all clusters, the cluster-level contrastive loss is further expressed as: Among them, H(Y) is the entropy function of cluster assignment probability; Where m is the sample number and P(·) is a trivial solution that avoids most samples being assigned to the same cluster; The cluster-level contrastive loss is defined as: <h2 style=";text-align:left;direction:ltr">L<h2 style=";text-align:left;direction:ltr"> clu <h2 style=";text-align:left;direction:ltr"> =L<h2 style=";text-align:left;direction:ltr"> clu <h2 style=";text-align:left;direction:ltr"> (t1,t2)+L<h2 style=";text-align:left;direction:ltr"> clu <h2 style=";text-align:left;direction:ltr"> (t2,t3) The total loss function consists of instance-level contrast loss and cluster-level contrast loss, namely: L=L ins +L clu 10. The small sample image classification method based on adaptive effective semantic expansion according to claim 1, characterized in that: In step 1), the backbone network ResNet of the constructed triple network joint model is input into the triple network joint model for pre-training, and the training data sets are all image data sets; In step 1), the model is trained in a purely unsupervised manner in contrastive learning at the instance level and cluster level; In step 2), the implementation of nearest neighbor semantic mining is to perform two-way weak amplification on the original sample by Conaug (·), generate two related samples, and introduce the support set S Get nearest neighbor information In step 2), the nearest neighbor semantic mining Feeding to instance-level projection and cluster-level projection Get nearest neighbor features and