Small sample image classification method based on feature set credibility inference
By employing a dual attention mechanism and multi-scale feature aggregation, combined with linear regression and set metrics, the problem of insufficient feature representation in small sample image classification is solved, thereby improving the model's classification accuracy and robustness.
Patent Information
- Application Number
- CN202510899621.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-11-25
AI Technical Summary
Existing deep learning methods struggle to classify effectively with small sample sizes. Traditional single-feature representations are prone to introducing bias and lack the ability to adaptively extract multiple feature embeddings.
We employ a lightweight network model based on a dual attention mechanism, combined with multi-scale feature aggregation and linear regression, and improve feature representation capability and classification accuracy through feature credibility inference and set measurement methods.
It improves the model's classification accuracy and robustness with small sample sizes, enhances its ability to identify complex scenes and fine-grained differences, suppresses noise interference, and improves the model's discriminative ability.
Smart Images

Figure CN121010794A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to a small sample image classification method, in particular to a small sample image classification method based on feature set credibility inference. BACKGROUND
[0002] Deep learning-based image classification methods have achieved remarkable results, but these methods usually rely on a large amount of labeled data for training. In the case of scarce training data, traditional deep learning methods often have difficulty achieving ideal results. Unlike this, humans can learn efficiently from limited samples, which is the ability that deep learning methods lack. Therefore, few-shot learning (FSL) has emerged to address how to effectively classify new objects with very few samples. This method focuses on improving learning and generalization ability in the case of extremely limited training samples. In practical applications, obtaining large-scale labeled data is both expensive and time-consuming, so few-shot learning provides a feasible solution to achieving high performance under sample constraints.
[0003] Currently, research in this field can be mainly divided into two categories: gradient-based methods and metric-based methods. Gradient-based methods optimize through gradient descent, enabling the model to adapt more quickly and accurately to new tasks and data, thereby enhancing the model's generalization ability and performance. This method aims to extract general knowledge from diverse few-shot tasks to facilitate rapid adaptation to new tasks. Unlike this, the main goal of metric learning is to accurately measure the similarity or distance between samples in the feature space, optimizing classification results by minimizing intra-class distance and maximizing inter-class distance. Traditional metric learning methods usually include two core components: a feature extractor and a classifier. In the prototype network, each class is represented by a prototype vector, which is the mean of all sample feature vectors in this class. The feature extractor extracts a separate feature vector for each sample. In the inference stage, the prototype network identifies the class prototype closest to the test sample by calculating the Euclidean distance or cosine similarity between the test sample and each class prototype. To improve the representation ability of the model, researchers mainly focus on two aspects of optimization: one is to design a more efficient network architecture to extract more accurate prototype features; the other is to develop a more accurate similarity evaluation method to ensure accurate measurement between different images. Single feature representation helps to unify the description of input images, making the model more robust, especially when facing noise or deformation. By extracting key features rather than complex multi-layer information, the model can maintain high stability. However, relying solely on a single feature representation for measurement can introduce bias. Therefore, there is an urgent need for a method that can adaptively extract a set of feature embedding to represent images, further improving model performance. SUMMARY
[0004] The application aims to provide a small sample image classification method based on feature set credibility inference, which extracts multiple feature embeddings through a double attention mechanism, adjusts the weight of the features by combining linear regression and prior knowledge, designs a set measurement-based scheme to solve the problem of how to measure the similarity of feature embeddings between the support set and the query set, and accurately captures the subtle differences between different categories of images to improve image recognition capability.
[0005] Technical scheme: The small sample image classification method based on feature set credibility inference comprises the following steps:
[0006] (1) Prepare a small sample image dataset, including miniImageNet, tieredImageNet and CUB dataset, divide the dataset into a support set and a query set, and generate meta-training and meta-testing tasks;
[0007] (2) Embed a double attention mechanism mapper in the feature extraction model, construct a lightweight network model based on the double attention mechanism, and perform image feature extraction;
[0008] (3) Input the image data into the network model, use a multi-scale feature aggregation technology to enhance the feature representation capability;
[0009] (4) Apply linear regression to the image data of the support set and the query set, infer the credibility of the features in the dataset, assign weights to the features, and adjust the features using semantic features after adding the weights;
[0010] (5) When classifying images, use a set measurement method to judge the similarity of the support set and the query set;
[0011] (6) Train different features of the same image through global self-supervised contrastive loss, and verify and evaluate the classification effect of the model on the test set.
[0012] Preferably, step (1) comprises predefining the miniImageNet, tieredImageNet and CUB dataset, and dividing the support set and the query set through random sampling technology to construct an N-way K-shot meta-training task; the model is trained on the support set and the classification performance is tested on the query set.
[0013] Preferably, the multi-scale feature aggregation adjusts the spatial size of the low-level features through an average pooling operation, introduces a convolution layer to learn the weight w of the low-level features f, and adjusts the low-level features f and the weight w through the low-level features f and the weight w. l l l l The matrix multiplication and activation function processing are carried out, the aggregation of low-level features is realized, the low-level features f l and high-level features f h are added, the fusion of low-level features to high-level features is completed;
[0014] f h =σ(w l (f l )·f l )+f h ,
[0015] wherein f h is a high-level feature, f l is a low-level feature, and σ is an activation function.
[0016] Preferably, the step (4) uses a linear regression method to evaluate the reliability of the final output feature set, identifies the features that most contribute to the final classification result, and quantifies the importance of these features in the classification decision through the form of weights. Prior knowledge is introduced as a weight, and the prior knowledge is used to evaluate the feature embedding to determine its possibility of belonging to a certain category. At the same time, according to these prior information, the contribution degree of each feature embedding in the feature set to the classification task is adjusted.
[0017] Preferably, the lightweight network model based on the dual attention mechanism includes a channel attention module, which calculates a channel attention map a c from , reshapes into R p×c×n , multiplies with its transpose matrix, and finally applies a softmax layer to obtain the attention map a c ∈R p×c×n .
[0018]
[0019] In the formula, is a feature vector obtained by inputting an image into a feature extractor, is a scale factor.
[0020] Preferably, the lightweight network model based on the dual attention mechanism includes a semantic information feature fusion module, which improves the visual features by calculating the semantic features of the entire category, constructs more information-rich positive and negative samples in the small sample classification task, constructs the original knowledge for all categories, and generates the semantic information of the category. Let A={a i}i∈(0-F) represent the set of class components / attributes, F represent the number of attributes, and R represent the association matrix of attributes and classes. If the attribute a i is related to the class k, then otherwise GloVe averages the word embeddings of the semantic embeddings of all classes and attributes:
[0021]
[0022] Based on a pre-trained feature extractor and raw knowledge, base class prototypes and attribute semantic features are extracted. Then, visual and semantic features are fused using addition operations.
[0023] h m =h m +H
[0024] Among them, h m Let H represent the feature set obtained after passing through the dual attention module, and let H represent the semantic features of each category extracted by wordNet.
[0025] Preferably, the set measurement method includes predicting a label for each feature in the feature set, and in order to measure the confidence of the predicted label relative to unlabeled data, inferring the confidence of the feature by regressing each instance from the feature to the label space:
[0026]
[0027] The goal of this formula is to minimize the feature set h by adjusting the parameter γ. m The difference between the predicted value Y and the squared Frobenius norm is used to measure the difference, where h m Let represent the set of features extracted by the feature extractor and the dual attention mapper, γ represent the parameters of the linear model used to combine these features, and R(γ) represent the regularization term applied to control the complexity of the model. The ultimate goal is to rank the features in the set according to a specific criterion.
[0028] Preferably, as the parameter λ changes, γ gradually becomes sparse until all elements disappear. By determining the λ value corresponding to the disappearance of γ, the features in the feature set are sorted by confidence.
[0029] Preferably, the set metric method uses a negative cosine similarity function to calculate the minimum sum of distances between a set of feature embeddings of the query set and the centroids of the support set:
[0030]
[0031] Here, the set-based metric is denoted as d. set =(X q S n ), quantifying the distance between two sets of features, X q S represents the query set. nThis represents the prototype of the support set.
[0032] Preferably, the formula for calculating the global self-supervised comparison loss is:
[0033]
[0034] Where · represents the inner product after L2 regularization, and τ1 is a scalar temperature parameter. It is an indicator function, positive sample z i From the same sample x i It is extracted through a dual attention mapper.
[0035] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0036] (1) This invention introduces a lightweight network structure with a dual attention mechanism, which effectively extracts local and global features of images, enhances the semantic information representation of images, and maintains low computational resource consumption. By using a multi-scale aggregation strategy, it integrates feature representations at different levels, enabling the model to have richer feature perception capabilities and enhancing its ability to recognize complex scenes and fine-grained differences.
[0037] (2) This invention proposes a feature credibility evaluation method based on linear regression and sparse constraints, which can effectively screen and sort the final feature set, suppress noise or interfering features, and improve classification and discrimination capabilities. By integrating external knowledge or domain experience to adaptively adjust feature weights, the model becomes more dependent on key features, further improving classification accuracy and robustness under small sample conditions.
[0038] (3) Based on the improved measurement method, the present invention uses the negative cosine similarity function for set measurement, which better measures the semantic distance between the support set and the query set and improves the model’s ability to distinguish between category boundaries. Attached Figure Description
[0039] Figure 1 This is a flowchart of the method described in this invention.
[0040] Figure 2 This is a flowchart of the dual attention mapper of the present invention.
[0041] Figure 3 This is a flowchart illustrating the workflow for inferring the credibility of a feature set in this invention.
[0042] Figure 4 This is a graph showing the accuracy and loss during the training of the model in this invention. Detailed Implementation
[0043] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0044] This invention provides a few-sample image classification method based on feature set confidence inference, including:
[0045] (1) Prepare a small sample image dataset, including miniImageNet, tieredImageNet and CUB datasets, divide the dataset into support set and query set, and generate meta-training and meta-testing tasks;
[0046] (2) Embed a dual attention mechanism in the feature extraction model to construct a lightweight network model based on the dual attention mechanism for image feature extraction;
[0047] (3) Input image data into the network model and use multi-scale feature aggregation technology to enhance feature representation capability;
[0048] (4) Apply linear regression to the image data of the support set and query set to infer the credibility of features in the dataset, assign weights to the features, and adjust the weighted features using semantic features.
[0049] (5) When classifying images, set measurement method is used to determine the similarity between support set and query set;
[0050] (6) The model is trained on different features of the same image by global self-supervised contrastive loss, and the classification effect of the model on the test set is verified and evaluated.
[0051] Step (1) involves predefining the miniImageNet, tieredImageNet, and CUB datasets, and dividing them into support and query sets using random sampling techniques. An N-way K-shot meta-training task is constructed, where K instances are randomly sampled from N classes as the support set, and Q instances are sampled from the remaining instances as the query set. The model is trained on the support set and its classification performance is tested on the query set, with the partitioning ensuring a uniform distribution of data categories.
[0052] In the specific implementation process, X train X val and X test These represent the training set, validation set, and test set, respectively, with the corresponding labels represented as Y. train Y val and Y test The set as a whole is represented by D. train =X train ,Y train D val =X val ,Y val and D test =X test ,Y testC train C val and C test , representing the categories in the training, validation, and test sets. Specifically, each N-way K-shot set E consists of a support set and a query set. First, starting from C... base Randomly sample N classes for meta-training (or from C) novel (Random sampling is used for meta-testing), where K instances are randomly sampled from each class to obtain support set S = {xi, yi}N*Ki = 1. Then, Q instances are sampled from each selected class to obtain the query set Q = {x i ,y i}N*Q i =1, where yi∈{1,2,...,N}, The fragmented training process classifies the samples in Q into the categories corresponding to the samples in S.
[0053] Step (2) constructs a lightweight network model based on the dual attention mechanism by embedding a dual attention mechanism mapper in the feature extractor to extract richer feature representations. Using a set of feature embeddings to represent each image allows the network model to aggregate richer and more useful features from different views of the image to capture image detail information.
[0054] Step (3) uses spatial multi-scale feature aggregation technology to enable the dual attention mapper to capture richer and more diverse feature representations, aggregates high-level features with low-level features, and inputs the aggregated features into the dual attention network for processing.
[0055] Step (4) proposes a feature credibility evaluation mechanism based on linear regression to further improve the discrimination performance in small sample classification tasks. This mechanism aims to identify features that positively contribute to classification decisions from the final extracted feature set, while suppressing potential interfering information, thereby achieving effective feature screening and optimization.
[0056] Specifically, the model first extracts features from the input samples and enhances their semantic expressiveness through an attention mechanism, resulting in a final feature set. For this set, a linear modeling process is introduced to simulate the mapping relationship between features and labels, thereby evaluating the reliability of each feature in the prediction task. This evaluation process is conducted using supervised learning, fitting the relationship between the feature space and the label space through regression modeling. The higher the fit between a feature and the label, the greater its corresponding regression weight, indicating a more significant contribution to classification; conversely, features with poor fit are considered redundant or distracting. To improve the robustness and discriminative power of feature selection, sparsity constraints are introduced during model training. By progressively applying a sparsity mechanism, the weights of some features are made to approach zero, thereby achieving feature-level filtering and ranking. This process not only helps reduce the interference of noisy features but also enhances the model's dependence on key features. During sparsification, the importance of features in the feature set can be ranked according to the order of feature weight decay, thus evaluating their reliability.
[0057] This invention introduces prior knowledge as auxiliary information to further calibrate feature credibility. This prior information can come from domain experience, external knowledge bases, or pre-trained models, and is used to measure the correlation between a feature and a specific category. By fusing prior information with regression results, the model can more accurately determine whether a feature has classification value and assign normalized adaptive weights to each feature accordingly. These weights are used to adjust the actual contribution ratio of each feature in subsequent classification processes, thereby enhancing the influence of positive features on the classification results and suppressing the misleading effects of negative features.
[0058] The set metric method in step (5) uses the negative cosine similarity function to calculate the sum of the minimum distances between the feature embeddings in the query set and the centroids of the support set.
[0059] Step (6) evaluates the similarity between different feature embeddings within the same image feature set using a global self-supervised contrastive loss. This aims to maximize the similarity of feature embeddings within the same set while minimizing the similarity between different image feature sets, effectively improving the model's recognition ability. The model's generalization ability in few-shot classification tasks is then trained, and its classification performance on the test set is validated and evaluated.
[0060] The lightweight network model based on the dual attention mechanism includes a multi-scale feature aggregation algorithm and a mapper algorithm, such as... Figure 1 As shown. To enable the dual attention mapper to effectively capture rich and diverse feature representations, an efficient spatial multi-stage feature aggregation module is introduced, such as... Figure 2As shown, the network consists of multiple stages, each corresponding to a convolutional module. The first stage serves as the initial module of the entire network, primarily used to extract basic low-level features from the input image. This stage relies solely on the input image itself and does not involve any feature fusion operations. Starting from the second stage, each module embeds two types of feature maps: the low-level feature map from the previous stage of the current module. High-level feature mapping of the current module Where C, W, and H represent the number of channels, height, and width of the feature map, respectively.
[0061] In this process, multi-scale feature aggregation adjusts the spatial size of low-level features through average pooling, enabling them to be effectively aggregated with high-level features. To achieve adaptive aggregation of low-level features, convolutional layers are introduced to learn the low-level feature f. l weight w l Next, by analyzing the low-level features f l and weight w l Matrix multiplication and activation function processing are performed to aggregate low-level features. Finally, the low-level features f are processed... l and advanced features f h Matrix addition is performed to fuse low-level features into high-level features.
[0062] f h =σ(w l (f l )·f l )+f h ,
[0063] Among them, f h It is a high-level feature, f l These are low-level features, and σ is the activation function. Next, the obtained features are input into a dual attention mapper to extract richer features.
[0064] In the channel attention module, directly from In the computation of channel attention mapping a c Specifically, will Remodeling into R p×c×n ,Will Multiply by its transpose matrix, and finally apply a softmax layer to obtain the attention map a. c ∈R p×c×n :
[0065]
[0066] In the formula, It is the feature vector obtained by the feature extractor from the input image. It is a scaling factor.
[0067] In the self-attention mechanism, The input is fed into a convolutional layer, resulting in three parameterized elements: and Use two parameterized elements and Calculation Note Figure a m :
[0068]
[0069] In the formula, a m yes Self-attention score, It is a scaling factor.
[0070] Finally, the output is obtained using self-attention and channel attention:
[0071]
[0072] Where, E∈R p×c×w×h This represents the final mapping obtained through a shallow dual attention mechanism. The feature vector h is calculated by averaging the p-th dimension of the patch. m .
[0073] The semantic information feature fusion module improves visual features by calculating the semantic features of the entire category, thereby constructing more information-rich positive and negative samples in few-shot classification tasks. It constructs original knowledge for all classes. Knowledge refers to the attribute features that a class should have. By generating semantic information for the categories in this way, the model can more effectively capture key information. In WordNet, let A = {a} i}i∈(0-F). Represents the set of class components / attributes. F represents the number of attributes, and R represents the association matrix between attributes and classes. If attribute a i Related to class k, then otherwise =0. GloVe averages the word embeddings of the semantic embeddings of all classes and attributes:
[0074]
[0075] Where H represents the set of semantic embedding vectors for categories and attributes, obtained using the GloVe word embedding model. k This represents the embedding vector of the k-th category. Let |C| represent the embedding vector of the i-th attribute, with a total of |C| for all categories (base and novel). base |+|C novel | , totaling F attributes.
[0076] Based on a pre-trained feature extractor and existing knowledge, base class prototypes and attribute semantic features are extracted. Visual and semantic features are then fused using addition.
[0077] h m =h m +H
[0078] Where h m Let H represent the feature set obtained after passing through the dual attention module, and let H represent the semantic features of each category extracted by wordNet.
[0079] To address the issue of the contribution of extracted feature sets to classification tasks, a linear regression method is used to infer the reliability of the feature set. This allows for the identification of more positive features for classification tasks with small sample sizes. The goal is to increase the contribution of positive features to the classification task and reduce the interference of negative features. A set metric is used to predict labels for each feature in the feature set. To measure the reliability of the predicted labels relative to unlabeled data, a linear model assumption is introduced. The reliability of a feature is inferred by regressing each instance from the feature space to the label space.
[0080]
[0081] The goal of this formula is to minimize the feature set h by adjusting the parameter γ. m The difference between the predicted value Y and the actual value Y is measured using the squared Frobenius norm. In this formula, h m Let represent the set of features extracted by the feature extractor and the dual attention mapper, γ represent the parameters of the linear model used to combine these features, and R(γ) represent the regularization term applied to control the complexity of the model. The ultimate goal is to rank the features in the set according to a specific criterion.
[0082] As the parameter λ changes, γ gradually becomes sparse until all elements disappear. By determining the λ value corresponding to the disappearance of γ, the features in the feature set are ranked by confidence. In classification tasks, each feature contributes differently, and some features may even provide misleading information. Therefore, after ranking the feature set by confidence, normalization is performed to assign adaptive weights to each feature to more effectively utilize valuable information.
[0083] The query feature set class is inferred from the support set of each class by comparing the query feature set with each corresponding instance. The set-based metric is denoted as d. set =(X q S n This quantifies the distance between two sets of features. Where X... q S represents the query set.n This represents the prototype of the support set. (Using h) m Let (x) represent the set of feature embeddings extracted from the query set. This represents the prototype array extracted from the support set. The negative cosine similarity function is used to compute the minimum sum of distances between a set of feature embeddings of the query set and the centroids of the support set.
[0084]
[0085] Self-supervised contrastive loss aims to enhance the similarity between different feature embeddings within the same image feature set, while reducing the similarity of feature embeddings across different images. Specifically, from the meta-training set D... train A batch of samples were randomly selected from the middle And a set of features is extracted using a dual attention mapper. here, x represents i The extracted set of feature embeddings is considered a positive sample pair. Define f. φ A feature extractor with learnable parameter φ, used to extract samples Convert to feature map Global features are obtained after passing through a Global Average Pooling (GAP) layer. The projection head proj(·) is implemented using an MLP with one hidden layer to generate projection vectors. Then, calculate the global self-supervised contrastive loss:
[0086]
[0087] Where · represents the inner product after L2 regularization, and τ1 is a scalar temperature parameter. It is an indicator function. Here, the positive sample z i From the same sample x i It is extracted through a dual attention mapper.
[0088] The method proposed in this invention is validated below. miniImageNet is a subset of ImageNet, containing 100 categories with 600 instances per category. The dataset is divided into a training set (64 categories), a validation set (16 categories), and a test set (20 categories). tieredImageNet, also from ImageNet, contains 779,165 images from 608 categories. The dataset is divided into a training set (351 categories), a validation set (97 categories), and a test set (160 categories). Fewshot-CIFAR100 (FC-100) is a subset of CIFAR-100, typically using 60 categories for training and 20 categories for validation and testing. CUB is a fine-grained classification dataset covering 200 bird species, with 100 categories used for basic classification, 50 categories for evaluation, and the remaining 50 categories for identifying novel categories.
[0089] This invention uses ResNet12 as the basic model architecture and is implemented strictly according to the Meta-Baseline specifications. Model parameters are initialized using the He-normal method. For optimization, this invention uses the stochastic gradient descent (SGD) algorithm with an initial learning rate of 0.1. In experiments on the miniImageNet dataset, the learning rate was iteratively adjusted at episodes 12,000, 14,000, and 16,000. In experiments on the tieredImageNet dataset, the learning rate was halved after every 24,000 episodes. All experiments were evaluated over 2000 episodes. During training, each batch contained 4 episodes as training samples.
[0090] To evaluate the performance of this patented model, it was compared with a variety of methods, including ProtoNet, DMF, SetFeat, DeepBDC, ELMOS, and CORL. This includes both classic few-shot learning (FSL) methods and models that have previously reported state-of-the-art results in the field.
[0091] Table 1 compares the miniImageNet and tieredImageNet datasets, with ResNet-12 as the backbone network in both benchmark methods. Table 2 compares the CUB dataset.
[0092]
[0093]
[0094] Table 1
[0095]
[0096] Table 2
[0097] This invention extracts multiple feature embeddings through a dual attention mechanism, combines linear regression with prior knowledge to adjust feature weights, and designs a scheme based on set metric, which enables the model to more accurately capture subtle differences between different categories and improve image recognition capabilities.
Claims
1. A few-sample image classification method based on feature set confidence inference, characterized in that, include: (1) Prepare a small sample image dataset, including miniImageNet, tieredImageNet and CUB datasets, divide the dataset into support set and query set, and generate meta-training and meta-testing tasks; (2) Embed a dual attention mechanism mapper in the feature extraction model to construct a lightweight network model based on the dual attention mechanism for image feature extraction; (3) Input image data into the network model and use multi-scale feature aggregation technology to enhance feature representation capability; (4) Apply linear regression to the image data of the support set and query set to infer the credibility of features in the dataset, assign weights to the features, and adjust the weighted features using semantic features. (5) When classifying images, set measurement method is used to determine the similarity between support set and query set; (6) The model is trained on different features of the same image by global self-supervised contrastive loss, and the classification effect of the model on the test set is verified and evaluated.
2. The few-sample image classification method based on feature set confidence inference according to claim 1, characterized in that, Step (1) includes predefining the miniImageNet, tieredImageNet and CUB datasets, and dividing the support set and query set by random sampling techniques to construct the N-way K-shot meta-training task; The model is trained on the support set and its classification performance is tested on the query set.
3. The few-sample image classification method based on feature set confidence inference according to claim 1, characterized in that, The multi-scale feature aggregation adjusts the spatial size of low-level features through average pooling and introduces convolutional layers to learn low-level features f. l weight w l By analyzing low-level features f l and weight w l Matrix multiplication and activation function processing are performed to aggregate low-level features. This is achieved by processing the low-level features f. l and advanced features f h Perform matrix addition to complete the fusion of low-level features into high-level features; f h =σ(w l (f l )·f l )+f h , Among them, f h It is a high-level feature, f l σ is a low-level feature, and σ is the activation function.
4. The few-sample image classification method based on feature set confidence inference according to claim 1, characterized in that, In step (4), the reliability of the final output feature set is evaluated using a linear regression method. The features that contribute most to the final classification result are identified, and their importance in the classification decision is quantified using weights. Prior knowledge is introduced as weights, and the feature embeddings are evaluated using this prior knowledge to determine their probability of belonging to a certain category. Simultaneously, the contribution of each feature embedding in the feature set to the classification task is adjusted based on this prior information.
5. A few-sample image classification method based on feature set credibility inference according to claim 1, characterized in that, The lightweight network model based on a dual attention mechanism includes a channel attention module, from... In the computation of channel attention mapping a c ,Will Remodeling into R p×c×n ,Will Multiply by its transpose matrix, and finally apply a softmax layer to obtain the attention map a. c ∈R p×c×n : In the formula, It is the feature vector obtained by the feature extractor from the input image. It is a scaling factor.
6. The few-sample image classification method based on feature set confidence inference according to claim 1, characterized in that, The lightweight network model based on the dual attention mechanism includes a semantic information feature fusion module. By calculating the semantic features of the entire category, it improves visual features, constructs more information-rich positive and negative samples in few-sample classification tasks, builds original knowledge for all classes, and generates semantic information for the categories. Let A = {a} i Let}i∈(0-F), representing the set of class components / attributes, where F represents the number of attributes, and R represents the association matrix between attributes and classes. If attribute a i Related to class k, then otherwise GloVe averages the word embeddings of the semantic embeddings of all classes and attributes: Based on a pre-trained feature extractor and raw knowledge, base class prototypes and attribute semantic features are extracted. Then, visual and semantic features are fused using addition operations. h m =h m +H Among them, h m Let H represent the feature set obtained after passing through the dual attention module, and let H represent the semantic features of each category extracted by wordNet.
7. A few-sample image classification method based on feature set credibility inference according to claim 1, characterized in that, The set metric method includes predicting a label for each feature in the feature set. To measure the confidence of the predicted label relative to unlabeled data, the confidence of the feature is inferred by regressing each instance from the feature to the label space. The goal of this formula is to minimize the feature set h by adjusting the parameter γ. m The difference between the predicted value Y and the squared Frobenius norm is used to measure the difference, where h m Let represent the set of features extracted by the feature extractor and the dual attention mapper, γ represent the parameters of the linear model used to combine these features, and R(γ) represent the regularization term applied to control the complexity of the model. The ultimate goal is to rank the features in the set according to a specific criterion.
8. A few-sample image classification method based on feature set confidence inference according to claim 7, characterized in that, As the parameter θ changes, γ will gradually become sparsified until all elements disappear. By determining the λ value corresponding to the disappearance of γ, the features in the feature set are sorted by confidence.
9. A few-sample image classification method based on feature set confidence inference according to claim 1, characterized in that, The set metric method uses a negative cosine similarity function to calculate the minimum sum of distances between a set of feature embeddings of the query set and the centroids of the support set: Here, the set-based metric is denoted as d. set =(X q S n ), quantifying the distance between two sets of features, X q S represents the query set. n This represents the prototype of the support set.
10. A few-sample image classification method based on feature set confidence inference according to claim 1, characterized in that, The formula for calculating the global self-supervised comparison loss is as follows: Where · represents the inner product after L2 regularization, and τ1 is a scalar temperature parameter. It is an indicator function, positive sample z i From the same sample x i It is extracted through a dual attention mapper.
Citation Information
Cited By
Small sample image classification method based on parallel expert structure
CN122156833A
Small sample image classification method based on parallel expert structure
CN122156833B