A Small-Sample Image Classification Method Based on Similarity Feature Fusion

By modeling the similarity between support and basic samples and categories in small sample learning, fusion features are generated and classifiers are trained, the feature deviation problem caused by differences between basic categories and support categories is solved, which improves classification accuracy and reduces training complexity and cost.

CN115965818BActive Publication Date: 2025-07-29UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310032701.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-10
Publication Date
2025-07-29
Estimated Expiration
2043-01-10

AI Technical Summary

Technical Problem

In the existing small sample learning methods, the difference between the basic category and the support category leads to feature representation deviations, affecting the classifier's response ability, and traditional data generation methods introduce noise to reduce classification accuracy.

Method used

By directly modeling the similarity between samples and basic samples, support categories and basic categories, the pre-trained CNN and word embedding models extract features, combine text and sample similarity relationships, generate fusion features and train classifiers to reduce deviation and noise, and improve classification accuracy.

Benefits of technology

It effectively reduces the semantic bias introduced due to category differences, improves the accuracy of small sample image classification, simplifies the training process, and reduces the computational cost.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965818B_ABST
    Figure CN115965818B_ABST
Patent Text Reader

Abstract

The present invention discloses a small-sample image classification method based on similarity feature fusion, comprising the following steps: Step 1: Extract features from the input image; Step 2: Extract the similarity relationship at the text end; Step 3: Extract the similarity relationship between samples; Step 4: Feature fusion based on text similarity; Step 5: Feature fusion based on sample similarity; Step 6: Multi-stage feature fusion; Step 7: Model training and testing. Based on the similarity between samples and categories, the present invention can improve the diversity of small-sample image features and perfect the category representation of small-sample images by performing feature fusion on the features of input small-sample images and the natural image features of the basic categories, thereby helping the classifier improve its response ability to small-sample images and enhancing the accuracy of small-sample image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image classification, and specifically relates to a small-sample image classification method based on similarity feature fusion. Background Art

[0002] In recent years, convolutional neural networks (CNNs) have demonstrated powerful performance in a large number of visual tasks including image classification and segmentation. However, they rely on large-scale labeled data for training, and the labeling of large-scale data requires a huge amount of human and material costs, which limits their application scenarios. To solve this problem, the task of few-shot learning (FSL) has been proposed. Its aim is to complete the classification of test samples through a limited number of training samples.

[0003] Currently, in the few-shot learning (FSL) task, a pre-training method is often adopted. It directly extracts the sample features of the support classes using a feature extractor (Backbone) pre-trained on the base classes, and trains a classifier using the features of the support samples. Training a robust feature extractor (Backbone) can effectively improve the performance of the few-shot learning (FSL) model. However, designing, training, and validating a feature extractor from scratch is time-consuming and expensive. Moreover, since the base classes and the support classes do not intersect, the feature extractor (Backbone) pre-trained on the base classes tends to focus more on the texture and structure information of the base class samples it has learned, resulting in its neglect of the details of the support samples, and there is a problem of weak classification performance.

[0004] To solve the above problem of insufficient classification performance on a small number of support samples, a data generation-based approach generates more new samples based on the current support samples to assist the optimization process of the classifier. However, it ignores the differences between the base classes and the support classes, and introduces additional noise during the data generation process, which will instead mislead the classifier.

[0005] Based on the above analysis, currently, how to reduce the deviation between the feature representations introduced by the differences between the base classes and the support classes, and between the base samples and the support samples, so as to improve the response ability of the classifier to the support classes, is an urgent problem to be solved in few-shot learning. Summary of the Invention

[0006] The present invention is proposed to solve the above-mentioned deficiencies of the prior art, and provides a small-sample image classification method based on similarity feature fusion. By directly modeling the similarity between the support samples and the base samples, and between the support classes and the base classes, the accuracy of small-sample image classification can be improved.

[0007] To achieve the above object of the present invention, the following technical solutions are adopted:

[0008] The method for classifying small - sample images based on similarity feature fusion of the present invention is characterized by the following steps:

[0009] Step 1: Feature extraction of the input image:

[0010] Step 1.1: Obtain a natural image set and input it into a pre - trained CNN model for feature extraction to obtain the feature representation of the natural image and its basic category set, denoted as where, represents the feature representation of the i - th natural image, and d represents the dimension of the feature representation, represents the basic category to which the i - th natural image belongs, and C base represents the basic category set of the natural image set, |C base | represents the number of basic categories of the natural image set, N base represents the number of natural images in each basic category;

[0011] Step 1.2: Obtain another image sample set and input it into the pre - trained CNN model for feature extraction to obtain the feature representation of the image sample and its support category set, denoted as where, represents the feature representation of the j - th image sample, and represents the support category to which the j - th image sample belongs, and C novel represents the support category set of the image sample, and satisfies C novel ∩C base =φ, |C novel | represents the number of support categories of the image sample, N novel represents the number of image samples in each support category;

[0012] Step 2: Extraction of text - end similarity relationships:

[0013] Step 2.1: Use a pre - trained word - embedding model to extract the vector representation of the text information of each basic category in the basic category set C base where, represents the vector representation of the text information of the k - th basic category, and t represents the dimension of the vector representation;

[0014] Step 2.2: Use the pre - trained word - embedding model to extract the vector representation of the text information of each support category in the support category set C novel where, where, The vector representation of the text information of the s-th support category

[0015] Step 2.3: Calculate the vector representation of the text information of the s-th support category using Equation (1) and the vector representation of the text information of the i-th base category the distance between them and use it as the text similarity relationship between the s-th support category and a base category, so as to obtain the text similarity relationship vector between the s-th support category and all base categories

[0016]

[0017] In Equation (1), denotes the vector inner product of and and respectively denote the L2 norms of and

[0018] Step 3: Extract the similarity relationship between samples:

[0019] Calculate the feature representation of the j-th image sample using Equation (2) and the feature representation of the i-th natural image the distance between them and use it as the similarity relationship between the j-th image sample and a natural image, so as to obtain the sample similarity relationship vector between the j-th image sample and all natural images

[0020]

[0021] In Equation (2), denotes the vector inner product of and and respectively denote the L2 norms of and

[0022] Step 4: Feature fusion based on text similarity and generate the fused feature

[0023] Step 5: Feature fusion based on sample similarity and generate the fused feature

[0024] Step 6: Multi-stage feature fusion and generate the fused feature

[0025] Step 7: Model Training and Testing:

[0026] Step 7.1. According to the feature extraction module, extract the feature representations of the images from the basic sample set and the support set. The similarity feature fusion module consists of feature fusion based on text similarity, feature fusion based on sample similarity, and multi-stage feature fusion. Perform feature fusion on the support samples according to the selection of the feature fusion method to obtain the fused samples

[0027] Step 7.2. Use Equation (3) to construct the loss function L;

[0028]

[0029] In Equation (3), L CE represents the cross-entropy loss, Γ represents the classifier, and λ is the harmonic factor during feature fusion; represents the category of the support samples, and is consistent with the category of the fused samples ;

[0030] Step 7.3. Use the gradient descent algorithm to train the classifier Γ, and calculate the loss function L to update the parameters of the classifier Γ. When the number of training iterations reaches the set number, stop training to obtain the trained classifier Γ * , which is used to predict the category of new image samples.

[0031] The small-sample image classification method based on similarity feature fusion according to the present invention is also characterized in that the step 4 includes:

[0032] Step 4.1. Denote the vector representation of the text information corresponding to the support category of the feature representation of the j-th image sample in V novel as and extract the text similarity relationship R base with all the basic categories in the basic category set C T (j);

[0033] Step 4.2. Select β basic category sets corresponding to the β closest distances from the text similarity relationship R of the feature representation T of the j-th image sample, and use the feature representations of all the natural images in the β basic category sets as the text-end alternative set where, represents the feature representation of the r-th natural image in the text-end alternative set D textual and is used as the alternative feature representation;

[0034] Step 4.3: Generate a random vector V on the text side T ∈R d , and the text-side random vector V T Obey 0-1 uniform distribution V T ~U(0,1), define the hyperparameter α, and α∈[0,1], according to the random vector V T With the hyperparameter α, use formula (4) to construct the text end mask vector M T ∈R d ;

[0035]

[0036] In formula (4), v Tt Represents the random vector V on the text side T The tth random value in m Tt Indicates M T The tth mask value in ;

[0037] Step 4.4: Based on the candidate feature representation And the text-side mask vector M T , use formula (5) to represent the features of the j-th image sample Perform feature fusion to generate fused features

[0038]

[0039] In formula (5), represents the vector inner product, and λ is the harmonic factor of random sampling from the Beta(2,2) distribution.

[0040] The step 5 comprises:

[0041] Step 5.1: Feature representation of the jth image sample extract With the base category set D base The similarity relationship R between samples of the feature representation of all natural images in I (j);

[0042] Step 5.2: From the current sample The similarity relationship R between samples I (j) Select the feature representations of the natural images with the closest distance to each other as the sample end candidate set D instance ,and in, Denotes the sample end candidate set D instance The feature representation of the rth natural image in is used as an alternative feature representation;

[0043] Step 5.3: Generate a random vector V at the sample end I ∈R d , V I follows a 0-1 uniform distribution V I ~U(0,1). Define a hyperparameter α, and α ∈ [0,1]. According to the random vector V I and the hyperparameter α, use Equation (6) to construct a mask vector M at the sample end I ∈R d ;

[0044]

[0045] In Equation (6), v Ik represents the k-th random value in the random vector V at the sample end I ; m Ik represents the k-th mask value in M I ;

[0046] Step 5.4: According to the alternative feature representation and the mask vector M at the sample end T , use Equation (7) to perform feature fusion on the feature representation of the j-th image sample to generate a fused feature

[0047]

[0048] In Equation (7), represents the inner product of vectors, and λ is a harmonic factor randomly sampled from a Beta(2,2) distribution.

[0049] The said Step 6 includes:

[0050] Step 6.1: For the feature representation V novel of the j-th image sample, the vector representation corresponding to the text information of its support category is denoted as Extract the text similarity relationship R base with all basic categories in the basic category set C T (j), and extract the sample similarity relationship R base with the feature representations of all natural images in the basic category set D I (j);

[0051] Step 6.2: From the text similarity relationship R of the feature representation TSelect a set of β base categories corresponding to the β closest distances in (j), and use the feature representations of all natural images in the β sets of base categories as the text - end alternative set Among them, represents the text - end alternative set D textual The feature representation of the r - th natural image in

[0052] Step 6.3: From the text alternative set D textual Select γ base - image samples with the closest distances as the alternative set D I according to the sample - similarity relationship R candidate , and where x f candidate represents the feature representation of the f - th natural image in the alternative set D candidate and is used as the alternative feature representation for feature fusion;

[0053] Step 6.4: Generate a random vector V, where V ∈ R d , and the random vector V follows a 0 - 1 uniform distribution V ∼ U(0,1). Define a hyperparameter α, and α ∈ [0,1]. According to the random vector V and the hyperparameter α, use Equation (8) to construct the sample - end mask vector M, where M ∈ R d ;

[0054]

[0055] Step 6.5: According to the alternative feature representation and the mask vector M, use Equation (9) to perform feature fusion on the feature representation of the j - th image sample to generate the fused feature

[0056]

[0057] In Equation (9), represents the vector inner product, and λ is a harmonic factor randomly sampled from the Beta(2,2) distribution.

[0058] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program that supports the processor to execute the few - shot image classification method, and the processor is configured to execute the program stored in the memory.

[0059] A computer - readable storage medium according to the present invention, characterized in that a computer program is stored on the computer - readable storage medium, and when the computer program is run by a processor, it executes the steps of the few - shot image classification method.

[0060] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0061] 1. The present invention designs a small-sample image classification method based on similarity feature fusion. By directly modeling the similarity between support samples and base samples, and between support classes and base classes, it solves the problems of information loss and insufficient attention to support feature details when using a feature extractor pre-trained on base classes to extract the features of support class samples.

[0062] 2. The present invention simultaneously utilizes the similarity between base classes and support classes, and between base samples and support samples to generate new samples with stronger discriminability, representativeness, and expressive ability. Compared with traditional data generation-based methods, it reduces the bias and noise introduced in the data generation process, fully considers the differences between base classes and support classes, better assists the training of the classifier, and improves the classification accuracy of the small-sample classification method.

[0063] 3. The present invention directly trains the classifier by generating support features. Compared with the traditional scheme based on training a feature extractor, it is simpler and more efficient, greatly reducing the complex time cost and expensive computational cost brought by training the feature extractor. At the same time, it makes up for the semantic bias introduced by class differences and improves the classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a flowchart of the small-sample image classification method based on similarity feature fusion of the present invention;

[0065] Figure 2 is a schematic diagram of extracting the similarity relationship between samples of the present invention;

[0066] Figure 3 is a schematic diagram of extracting the similarity relationship at the text end of the present invention;

[0067] Figure 4 is a schematic diagram of the feature fusion method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] In this embodiment, a small-sample classification method based on similarity feature fusion directly models the similarity between support samples and base samples, and between support classes and base classes, and generates new samples based on the similarity to improve the description of support samples and assist the optimization process of the classifier, thereby reducing the semantic bias introduced by class differences and improving the accuracy of the small-sample image classification method. Specifically, as Figure 1 shown, it is carried out according to the following steps:

[0069] Step 1: Extract features from the input image:

[0070] Before performing similarity relation extraction, first convert the image samples from the natural image set and another image set into feature representations through a CNN model pre-trained on the natural image set.

[0071] Step 1.1: Obtain the natural image set and input it into the pre-trained CNN model for feature extraction to obtain the feature representation of the natural image and its basic category set, denoted as Among them, represents the feature representation of the i-th natural image, and d represents the dimension of the feature representation, represents the basic category to which the i-th natural image belongs, and C base represents the basic category set of the natural image set, |C base | represents the number of basic categories in the natural image set, N base represents the number of natural images in each basic category;

[0072] Step 1.2: Obtain another image sample set and input it into the pre-trained CNN model for feature extraction to obtain the feature representation of the image sample and its support category set, denoted as Among them, represents the feature representation of the j-th image sample, and represents the support category to which the j-th image sample belongs, and C novel represents the support category set of the image sample, and satisfies C novel ∩C base =φ, |C novel | represents the number of support categories of the image sample, N novel represents the number of image samples in each support category;

[0073] Step 2: Text-side similarity relation extraction:

[0074] To achieve feature fusion based on category text similarity, it is necessary to extract the similarity relations between the text information of each support category and the text information of all basic categories. First, convert the semantic labels of the basic categories and support categories into vector representation forms through a pre-trained word embedding method, and then calculate the Cosine distance between the vector representation of the support category and the vector representation of each basic category as the text-side similarity relation.

[0075] Step 2.1: Use the pre-trained word embedding model to extract the vector representation of the text information of each basic category in the basic category set C base in Among them, represents the vector representation of the text information of the k-th basic category, t represents the dimension of the vector representation;

[0076] Step 2.2: Use the pre-trained word embedding model to extract the support category set C novel The vector representation of the text information of each support category in in, The vector representation of the text information of the s-th support category,

[0077] Step 2.3: Use formula (1) to calculate the vector representation of the text information of the s-th support category and the vector representation of the i-th basic category text information The distance between And as the text-side similarity relationship between the s-th support category and a basic category, the text-side similarity relationship vector between the s-th support category and all basic categories is obtained.

[0078]

[0079] In formula (1), express and The vector inner product of and Respectively and L2 paradigm;

[0080] Step 3: Extract similarity relationships between samples:

[0081] In order to achieve similarity feature fusion between samples, it is necessary to extract the similarity relationship between each image sample of the support category and all natural image samples. For each image sample of the support category, the Cosine distance between its feature representation and the feature representation of all natural image samples is calculated as the similarity relationship between the samples.

[0082] Step 3.1: Use formula (2) to calculate the feature representation of the jth image sample and the feature representation of the i-th natural image The distance between And as the similarity relationship between the jth image sample and a natural image, the sample similarity relationship vector between the jth image sample and all natural images is obtained

[0083]

[0084] In formula (2), express and The vector inner product, and respectively represent and the L2 norm of;

[0085] Step 4: Feature fusion based on text similarity:

[0086] Step 4.1. Denote the vector representation of the text information corresponding to the support category in V for the feature representation of the j-th image sample as novel and extract and the text similarity relationship R base with all the base categories in the base category set C T (j);

[0087] Step 4.2. As Figure 2 shown, select the set of β base categories corresponding to the β closest distances from the text similarity relationship R of the feature representation of the j-th image sample T (j), and use the feature representations of all the natural images in the set of β base categories as the text-side alternative set where, denotes the feature representation of the r-th natural image in the text-side alternative set D textual and use it as the alternative feature representation;

[0088] Step 4.3. Generate a text-side random vector V T ∈R d , and the text-side random vector V T obeys a 0-1 uniform distribution V T ~U(0,1). Define a hyperparameter α, and α ∈ [0,1]. In this example, α = 0.7. According to the random vector V T and the hyperparameter α, use Equation (3) to construct a text-side mask vector M T ∈R d ;

[0089]

[0090] In Equation (3), v Tt represents the t-th random value in the text-side random vector V T ; m Tt represents the t-th mask value in M T ;

[0091] Step 4.4. According to the alternative feature representation and the text-side mask vector M T , use Equation (4) to process the feature representation of the j-th image samplePerform feature fusion to generate the fused features

[0092]

[0093] In Equation (4), denotes the inner product of vectors, and λ is the harmonic factor randomly sampled from the Beta(2, 2) distribution;

[0094] Step 5: Feature fusion based on sample similarity:

[0095] Step 5.1. For the feature representation of the j-th image sample Extract and the sample similarity relationship R base with the feature representations of all natural images in the basic category set D I (j);

[0096] Step 5.2. As Figure 3 shown, select the feature representations of γ natural images with the closest distances from the sample similarity relationship R of the current sample as the sample-side alternative set D I (j), where instance where, where, denotes the feature representation of the r-th natural image in the sample-side alternative set D instance and serves as the alternative feature representation. In this example, γ = 512.

[0097] Step 5.3. Generate the sample-side random vector V I ∈ R d , V I follows a 0-1 uniform distribution V I ~ U(0, 1). Define the hyperparameter α, and α ∈ [0, 1]. In this example, α = 0.7. According to the random vector V I and the hyperparameter α, use Equation (5) to construct the sample-side mask vector M I ∈ R d ;

[0098]

[0099] In Equation (5), v Ik denotes the k-th random value in the sample-side random vector V I ; m Ik denotes the k-th mask value in M I ;

[0100] Step 5.4. According to the alternative feature representation and the sample-side mask vector M T, use Equation (6) to represent the features of the j-th image sample Perform feature fusion to generate the fused features

[0101]

[0102] In Equation (6), represents the inner product of vectors, and λ is the harmonic factor randomly sampled from the Beta(2, 2) distribution;

[0103] Step 6: Multi-stage feature fusion:

[0104] Step 6.1. For the feature representation V novel The vector representation corresponding to the text information of the support class in is denoted as Extract The text similarity relationship R with all basic classes in the basic class set C base (j), extract T (j), the sample similarity relationship R between the feature representations of all natural images in the basic class set D and the basic class set D base (j); I (j);

[0105] Step 6.2. Select β basic class sets corresponding to the β closest distances from the text similarity relationship R (j) of the j-th image sample's feature representation, and use the feature representations of all natural images in the β basic class sets as the text-end alternative set T where represents the feature representation of the r-th natural image in the text-end alternative set D and is used as the alternative feature representation. In this example, β = 2; textual

[0106] Step 6.3. Select γ basic image samples with the closest distances from the text alternative set D textual according to the sample similarity relationship R I (s) as the alternative set D candidate , and where x f candidate represents the feature representation of the f-th natural image in the alternative set D candidate and is used as the alternative feature representation for feature fusion. In this example, γ = 512;

[0107] Step 6.4. As Figure 4 shown, generate a random vector V, where V ∈ R d ​, and the random vector V follows a 0-1 uniform distribution V ∼ U(0, 1). Define the hyperparameter α, and α ∈ [0, 1]. In this example, α = 0.7. According to the random vector V and the hyperparameter α, use Equation (7) to construct the sample-side mask vector M, where M ∈ R d ;

[0108]

[0109] Step 6.5: According to the alternative feature representation and the mask vector M, use Equation (8) to perform feature fusion on the feature representation of the j-th image sample to generate the fused feature

[0110]

[0111] In Equation (8), represents the vector inner product, and λ is the harmonic factor randomly sampled from the Beta(2, 2) distribution;

[0112] Step 7: Model training and testing:

[0113] Step 7.1: According to the feature extraction module, extract the feature representations of the images from the basic sample set and the support set. The similarity feature fusion module consists of feature fusion based on text similarity, feature fusion based on sample similarity, and multi-stage feature fusion. Perform feature fusion on the support samples according to the selection of the feature fusion method to obtain the fused samples

[0114] Step 7.2: Use Equation (9) to construct the loss function L;

[0115]

[0116] In Equation (9), L CE represents the cross-entropy loss, Γ represents the classifier, and λ is the harmonic factor during feature fusion; represents the category of the support sample and is consistent with the category of the fused sample ;

[0117] Step 7.3: Use the gradient descent algorithm to train the classifier Γ and calculate the loss function L to update the parameters of the classifier Γ. When the number of training iterations reaches the set number, stop training to obtain the trained classifier Γ * , which is used to predict the category of new image samples.

[0118] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above-mentioned few-shot classification method, and the processor is configured to execute the program stored in the memory.

[0119] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above-mentioned few-shot classification method.

Claims

1. A small-sample image classification method based on similarity feature fusion, characterized in that Follow these steps: Step 1: Feature extraction of input image: Step 1.

1. Obtain a natural image set and input it into a pre-trained CNN model for feature extraction to obtain the feature representation of the natural image and its basic category set, denoted as where represents the feature representation of the i-th natural image, and d represents the dimension of the feature representation, represents the basic category to which the i-th natural image belongs, and C base represents the basic category set of the natural image set, |C base | represents the number of basic categories of the natural image set, N base represents the number of natural images in each basic category; Step 1.2: Obtain another set of image samples and input them into the pre-trained CNN model for feature extraction to obtain the feature representation of the image samples and their supporting category set, which is recorded as in, represents the feature representation of the j-th image sample, and represents the support category to which the jth image sample belongs, and C novel Represents the support category set of image samples, and satisfies C novel ∩C base =φ,|C novel | represents the number of supported categories of image samples, N novel Indicates the number of image samples in each support category; Step 2: Extract similarity relationships on the text side: Step 2.1: Use a pre-trained word embedding model to extract the vector representations of the text information of each basic category in the basic category set C base among them where represents the vector representation of the text information of the k-th basic category and t represents the dimension of the vector representation Step 2.2: Use the pre-trained word embedding model to extract the vector representations of the text information of each support category in the support category set C novel among them wherein represents the vector representation of the text information of the s-th support category Step 2.3: Use formula (1) to calculate the vector representation of the text information of the s-th support category and the vector representation of the i-th basic category text information The distance between And as the text-side similarity relationship between the s-th support category and a basic category, the text-side similarity relationship vector between the s-th support category and all basic categories is obtained. In formula (1), represents the vector inner product of and respectively represent the L2 norms of Step 3: Extract similarity relationships between samples: Use formula (2) to calculate the feature representation of the jth image sample and the feature representation of the i-th natural image The distance between And as the similarity relationship between the jth image sample and a natural image, the sample similarity relationship vector between the jth image sample and all natural images is obtained In formula (2), represents the vector inner product of and respectively represent the L2 norms of Step 4: Feature fusion based on text similarity and generate fused features Step 5: Feature fusion based on sample similarity and generate fused features Step 6: Multi-stage feature fusion and generate fused features Step 7: Model training and testing: Step 7.1, based on the feature extraction module, extract the feature representation of the image from the basic sample set and the support set, and form a similarity feature fusion module by the feature fusion based on text similarity, the feature fusion based on sample similarity and the multi-stage feature fusion, and extract the feature representation of the image from the support sample set. Perform feature fusion according to the feature fusion method selection to obtain the fused sample Step 7.2: Use formula (3) to construct the loss function L; In formula (3), L CE represents the cross entropy loss, Γ represents the classifier, and λ is the harmonic factor during feature fusion; Indicates the category of the supporting sample, and is consistent with the fused sample The categories are consistent; Step 7.3: Train the classifier Γ using the gradient descent algorithm, calculate the loss function L to update the parameters of the classifier Γ, and stop training when the number of training iterations reaches the set number of times to obtain the trained classifier Γ * , which is used to predict the class of new image samples.

2. The small-sample image classification method based on similarity feature fusion according to claim 1, wherein The step 4 comprises: Step 4.1: Represent the feature of the j-th image sample In V novel The vector representation of the text information corresponding to the support category is recorded as and extract With the base category set C base The text similarity relationship R of all basic categories in T (j); Step 4.

2. Select, from the text similarity relationship R of the j-th image sample T , the set of β basic categories corresponding to the β closest distances, and use the feature representations of all the natural images in the β sets of basic categories as the text-end alternative set . of the text similarity relationship R T (j) to select the set of β basic categories corresponding to the β closest distances, and use the feature representations of all the natural images in the β sets of basic categories as the text-end alternative set Among them, represents the feature representation of the r-th natural image in the text-end alternative set D textual and serves as the alternative feature representation; textual in the text-end alternative set D and serves as the alternative feature representation; Step 4.3: Generate a random vector V on the text side T ∈R d , and the text-side random vector V T Obey 0-1 uniform distribution V T ~U(0,1), define the hyperparameter α, and α∈[0,1], according to the random vector V T With the hyperparameter α, use formula (4) to construct the text end mask vector M T ∈R d ; In formula (4), v Tt represents the t-th random value in the text-side random vector V T ; m Tt represents the t-th masking value in M T ; Step 4.4: Based on the candidate feature representation And the text-side mask vector M T , use formula (5) to represent the features of the j-th image sample Perform feature fusion to generate fused features In formula (5), represents the inner product of vectors, and λ is a harmonic factor randomly sampled from the Beta(2, 2) distribution.

3. The small-sample image classification method based on similarity feature fusion according to claim 2, wherein The step 5 comprises: Step 5.1: Feature representation of the jth image sample extract With the base category set D base The similarity relationship R between samples of the feature representation of all natural images in I (j); Step 5.

2. Select, from the sample - to - sample similarity relationship R of the current sample I (j), the feature representations of γ natural images with the closest distances as the sample - side alternative set D instance , and where represents the feature representation of the r - th natural image in the sample - side alternative set D instance and serves as the alternative feature representation; Step 5.3: Generate the sample - side random vector V I ∈R d ,V I obeys the 0 - 1 uniform distribution V I ~U(0,1), define the hyperparameter α, and α ∈ [0,1]. According to the random vector V I and the hyperparameter α, use Equation (6) to construct the sample - side mask vector M I ∈R d ; In formula (6), v Ik Represents the random vector V at the sample end I The kth random value in m Ik Indicates M I The kth mask value in ; Step 5.4, according to the alternative feature representation and the sample end mask vector M T , use Equation (7) to perform feature fusion on the feature representation of the j-th image sample to generate the fused feature In formula (7), represents the inner product of vectors, and λ is the harmonic factor randomly sampled from the Beta(2, 2) distribution.

4. The small-sample image classification method based on similarity feature fusion according to claim 3, wherein The step 6 comprises: Step 6.1: Feature representation of the jth image sample V novel The vector representation of the text information corresponding to its support category is recorded as extract With the base category set C base The text similarity relationship R of all basic categories in T (j), and extract With the base category set D base The sample similarity relationship R of the feature representation of all natural images in I (j); Step 6.

2. Select, from the text similarity relationship R of the j-th image sample, a set of β base categories corresponding to the β closest distances, and use the feature representations of all natural images in the β sets of base categories as the text-end alternative set T (j), where denotes the feature representation of the r-th natural image in the text-end alternative set D ; textual in the text-end alternative set D Step 6.3: From the text selection set D textual Based on the similarity relationship R between samples I (s) Select γ nearest basic image samples as the candidate set D candidate ,and Among them, x f candidate represents the candidate set D candidate The feature representation of the f-th natural image in is used as an alternative feature representation for feature fusion; Step 6.4: Generate a random vector V, where V ∈ R d , and the random vector V follows a 0-1 uniform distribution V ∼ U(0,1). Define a hyperparameter α, and α ∈ [0,1]. According to the random vector V and the hyperparameter α, use Equation (8) to construct the sample-side mask vector M, where M ∈ R d ; Step 6.5, according to the alternative feature representation and the mask vector M, use Equation (9) to perform feature fusion on the feature representation of the j-th image sample to generate the fused feature In Equation (9), represents the inner product of vectors, and λ is a harmonic factor randomly sampled from a Beta(2, 2) distribution.

5. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor to execute the small sample image classification method according to any one of claims 1 to 4, and the processor is configured to execute the program stored in the memory.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the small sample image classification method according to any one of claims 1 to 4 are executed.

Citation Information

Patent Citations

  • Training method and device for word embedding model

    CN109308354A

  • Method, device and system for identifying text in image

    CN113128494A