A Fine-Grained Image Classification Method Based on Feature Fusion and Semantic Enhancement

Through multi-level attention mechanism and semantic enhancement technology, combined with cross-entropy loss function optimization, the problems of insufficient feature extraction and low classification accuracy in the existing fine-grained image classification methods are solved, and higher classification accuracy and feature region positioning accuracy are achieved.

CN118799646BActive Publication Date: 2025-07-29SICHUAN DIGITAL ECONOMY RESEARCH INSTITUTE (YIBIN) +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411084301.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-08
Publication Date
2025-07-29
Estimated Expiration
2044-08-08

AI Technical Summary

Technical Problem

The existing fine-grained image classification methods have shortcomings in feature extraction and feature fusion, and it is difficult to accurately locate discriminant feature regions, especially in scenarios where subcategory similarity is high and background information is complex.

Method used

Using multi-level attention mechanism and semantic enhancement technology, global and local features are extracted through the ViT model, combined with cross attention mechanism and cross entropy loss function optimization, key tokens are selected and quadratic blocking is performed to enhance the discriminant characteristics.

Benefits of technology

The classification accuracy in complex backgrounds and high similarity subcategory scenarios is improved, and the positioning ability and feature extraction of discriminant feature regions are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118799646B_ABST
    Figure CN118799646B_ABST
Patent Text Reader

Abstract

The present invention discloses a fine-grained image classification method based on feature fusion and semantic enhancement. The method includes the following steps: First, use the Vision Transformer (ViT) model for feature extraction, divide the input image into non-overlapping patches, convert them into embedding vectors through linear projection, and input them into the Transformer encoder to generate global features. Then, combine multi-level attention fusion with semantic information, extract the attention weights in each layer of the Transformer, and combine with the semantic embeddings generated by the pre-trained language model to calculate the importance scores of each token and select the key tokens. Next, perform secondary chunking and projection on the key tokens to re-select the secondary key tokens. Through the cross-attention mechanism, fuse the global features and local features to generate fused features. Finally, combine the fused features with the global classification features, input them into the classifier for classification, and generate classification outputs. Through multi-level attention fusion, semantic enhancement, and key token selection, the present invention realizes the accurate localization of discriminative feature regions of fine-grained images, enhances the discriminability of features, and improves the classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fine-grained image classification. Specifically, it relates to a fine-grained image classification method based on feature fusion and semantic enhancement. The aim is to improve the deficiencies of existing fine-grained image classification methods in accurately locating discriminative feature regions and extracting discriminative features through strategies such as multi-level attention mechanisms and semantic enhancement, thereby improving the classification accuracy of the model in scenarios with high sub-category similarity and complex image background information. Background Art

[0002] Fine-grained image classification is an important research direction in computer vision tasks, aiming to classify and identify images of similar sub-categories under a certain large category. Fine-grained image recognition has a wide range of applications in fields such as biodiversity detection, intelligent retail systems, intelligent transportation, and military target recognition.

[0003] Existing fine-grained image recognition technologies mainly rely on deep learning methods and have deficiencies in feature extraction and feature fusion.

[0004] Chinese Patent Application CN 118135290A proposes a fine-grained image classification method based on Transformer. The Transformer backbone network is used to divide the image into small patches, and then each token is scored through an information discarding module, and the low-score tokens are filtered according to the scores. Then, a feature selection module is used to select strong discriminative regions and splice them to obtain an aggregated feature. This method can filter out noise and irrelevant tokens to a certain extent and then select discriminative feature regions. However, it is difficult for a single-stage Vision Transformer (ViT) model to fully extract the discriminative fine-grained features in the image. The information discarding module may discard some key tokens containing discriminative information; and only using a single cross-entropy loss function to optimize the model results in insufficient discriminability of the extracted features and difficulty in distinguishing the subtle differences between sub-categories.

[0005] Chinese Patent Application CN 118262180A proposes a fine-grained image classification method based on foreground-background multi-level division and feature fusion. This method extracts features through image segmentation based on the attention intensity distribution value to obtain a foreground image and multi-level background images, and performs two weighted fusions to improve the classification accuracy. However, when the single-level attention mechanism assigns weights to all tokens, it is difficult to utilize the more discriminative features and fully extract the fine-grained information of the image; the single-level feature weighting method is difficult to fully fuse the fine-grained features and global features of the image, reducing the classification accuracy.

[0006] Chinese Patent CN 114676776 A proposes a fine-grained image classification method based on Transformer. Each head of the multi-head attention mechanism is used to select the token feature with the highest degree of association with the classification token as the global feature, and then a local region containing semantic components is cropped from the input image to obtain the local feature, and the two features are fused. However, there are problems of feature duplication and feature loss in the selection of global features through multi-head selection; it is impossible to flexibly adapt to different images by selecting image patch tokens with a high degree of association with the classification token according to an empirical threshold; adopting a single center loss function lacks direct optimization of the classification boundary, and it is difficult to find an accurate classification boundary in scenarios where there are only slight differences between categories or feature space overlap. Summary of the Invention

[0007] To solve the above problems, the present invention provides a fine-grained image classification method based on feature fusion and semantic enhancement. The fine-grained features of the image are extracted layer by layer through a multi-level attention mechanism to accurately locate the discriminative feature region; a semantic enhancement technology and a key token selection mechanism are introduced, the global feature and the fine-grained feature with strong discrimination are respectively extracted based on the ViT model, and the global feature and the fine-grained feature are fused by using the cross-attention mechanism to more comprehensively capture the details and global information of the image; the model is optimized based on the cross-entropy loss and the contrast loss, which is beneficial to extracting discriminative features and improving the classification accuracy.

[0008] A fine-grained image classification method based on feature fusion and semantic enhancement includes the following steps:

[0009] S1. Feature extraction of the ViT model: The input image I is divided into non-overlapping patches and converted into an embedding vector E through linear projection i , and input into the Transformer encoder to generate the global feature E global .

[0010] S2. Multi-level attention fusion and key token selection combining semantic information: Extract the attention weight a in each layer of the Transformer l , and combine the semantic embedding emb generated by the pre-trained language model to calculate the importance score s of each token i , and select the key token denoted as z key .

[0011] S3. Semantic enhancement and refined key token selection: Perform secondary partitioning and projection on z in S2 key to generate a new embedding vector E' i,j , re-input it into the Transformer encoder, and select the secondary key token denoted as z' again key .

[0012] S4. Cross - attention fusion: Through the cross - attention mechanism, fuse the global feature Z global and the local feature z′ key to generate the fused feature Z fused .

[0013] S5. Feature fusion and classification: Combine the fused feature Z fused with the global classification feature E global and input it into the classifier for classification to generate the classification output y′.

[0014] S6. Loss function design and training: Design the cross - entropy loss L cross and the contrastive loss L con , and optimize the model parameters through backpropagation.

[0015] Furthermore, the ViT model feature extraction in S1 includes:

[0016] S101. Image chunking and linear projection: Divide the input image I into several non - overlapping patches (i.e., tokens) of size 16×16, and then convert each patch P i into an embedding vector E i as shown in Equation (1):

[0017] E i = LinearProjection(P i ) (1)

[0018] where E i represents the embedding vector of the i - th patch.

[0019] S102. Embedding sequence combination: Combine all the embedding vectors into a sequence E = [E1, E2,..., E i , and add a classification token embedding vector E0: E′ = [E0, E1, E2,..., E i

[0020] S103. Transformer encoder: Input the embedding vector sequence E′ into a multi - layer Transformer encoder, and generate the query matrix Q, the key matrix K, and the value matrix V through linear transformation. Each layer of the Transformer encoder consists of a multi - head self - attention mechanism and a feed - forward neural network. The calculation process of the attention mechanism is shown in Equation (2):

[0021]

[0022] where softmax represents the softmax activation function d​k Denotes the dimension of the key matrix; A l Denotes the attention weight matrix of the l-th layer.

[0023] S104. Global features:

[0024] A1. In the Transformer encoders of all layers, the classification token E0 will absorb the information from other tokens layer by layer and update to form a vector E containing global features global .

[0025] A2. Perform a weighted sum of the attention weight matrices of all layers to obtain the global feature representation Z fused by the attention mechanism globa , as shown in Equation (3):

[0026]

[0027] where, V l Denotes the value matrix of the l-th layer; A l Denotes the attention weight matrix of the image I at the l-th layer; L denotes the total number of layers of the Transformer; Z global Denotes the global feature representation.

[0028] Furthermore, the multi-level attention fusion and key token selection for combining semantic information in S2 include:

[0029] S201. Extract and fuse attention weights: Extract the attention weight A of each token from each layer of the Transformer in S1 l . Perform matrix multiplication fusion on the attention weights of all layers to generate the fused global attention, as shown in Equation (4):

[0030]

[0031] where, a fimal Denotes the fused global attention, Denotes the attention weight of each token at the first layer.

[0032] S202. Normalization processing: Perform normalization processing on the attention weights of each token in the global attention, as shown in Equation (5):

[0033]

[0034] where, a final,i Denotes the attention weight of the i-th token, ∑ j a final,j Denotes the sum of the attention weights of all tokens in the global attention; Represents the attention weight of the i-th token after normalization.

[0035] S203. Generate semantic embeddings, fuse semantic information to calculate importance scores: Use a pre-trained language model to generate semantic embeddings emb for each token, then perform an element-wise multiplication (Hadamard product) of the normalized attention weights and the generated semantic embeddings, and perform normalization processing to generate enhanced attention weights, as shown in Equation (6):

[0036]

[0037] where, ⊙ represents element-wise multiplication; emb i represents the semantic embedding vector of the i-th token; a enhanced,i represents the enhanced attention weight of the i-th token.

[0038] Use the enhanced attention weight a enhanced,i to calculate the importance score of each token, as shown in Equation (7):

[0039] a i = a enhanced,i · emb i (7)

[0040] where, s i represents the importance score of the i-th token, and select several tokens with the highest scores as key tokens denoted as z key .

[0041] 4. Further, the semantic enhancement and refinement key token selection in S3 includes:

[0042] S301. Perform secondary chunking and linear projection on key tokens: Further chunk the z key selected in S2, and divide each key token into patches of size 16×16 again. Perform linear projection on these new-chunked patches to generate new embedding vectors, as shown in Equation (8):

[0043] E′ i,j = LinearProjection(P i,j ) (8)

[0044] where, P i,j represents the j-th new chunk of the i-th key token; E′ i,j represents the embedding vector of the j-th new chunk of the i-th key token.

[0045] S302. Embedded Sequence Combination: Combine these embedded vectors into a new sequence, add the classification token E0, and input it into the Transformer encoder for processing, as shown in Equation (9):

[0046] E″ i =[E0, E′ i,1 , E′ i,2 ,…, E′ i,j (9)

[0047] S303. Secondary Selection of Key Tokens: The calculation method is the same as that in S2. Calculate the secondary attention weights and select the secondary key tokens denoted as z′ key .

[0048] Furthermore, the cross-attention fusion in S4 includes:

[0049] S401. Introduction of Cross-Attention Mechanism: After the secondary key token selection, introduce the cross-attention mechanism to fuse the global features and local features. Calculate the cross-attention as shown in Equation (10):

[0050]

[0051] Among them, Q represents the query matrix of the global feature Z global ; K represents the key matrix of the secondary key token feature z′ key ; d kj represents the dimension of the key matrix; A cross represents the cross-attention weight matrix.

[0052] S402. Apply Cross-Attention Weights to Generate Fusion Features as shown in Equation (11):

[0053] Z fused =A cross V (11)

[0054] Among them, V represents the value matrix of the secondary key token feature z′ key ; Z fused represents the fused feature matrix.

[0055] Furthermore, the feature fusion and classification in S5 include:

[0056] S501. Final Feature Fusion: Fuse the cross-attention fusion feature Z fused with the global classification feature vector E global obtained in S1, as shown in Equation (12):

[0057] Z final =[E global , Zfused (12)

[0058] Among them, Z final represents the finally fused feature vector.

[0059] S502. Classification output: Input the finally fused feature vector Z that contains the classification token and the cross-attention fused features final into the classifier for classification, as shown in Equation (13):

[0060] y′ = Classifier(Z final ) (13)

[0061] Among them, Classifier represents the classifier; y′ represents the classification output of the model.

[0062] Furthermore, the loss function design and training in S6 include:

[0063] A1. Calculate the cross-entropy loss: After each forward propagation ends, calculate the cross-entropy loss for the classification task, as shown in Equation (14):

[0064]

[0065] Among them, y n represents the true label of the nth sample; p n represents the predicted probability of the nth sample; L cross represents the cross-entropy loss.

[0066] A2. Calculate the contrastive loss: To further extract discriminative features, maximize the similarity between samples of the same class, and at the same time minimize the similarity between samples of different classes. After each forward propagation ends, calculate the contrastive loss, as shown in Equation (15):

[0067]

[0068] Among them, B represents the batch size; α represents the hyperparameter that controls the similarity threshold; y n and y m represent the labels of the nth sample and the mth sample; Sim(z n , z m ) represents the similarity between the nth and the mth samples; is the Kronecker delta function, which takes the value 1 when y n = y m and 0 otherwise; L con represents the contrastive loss.

[0069] A3. Combine the loss function: The final loss function L FIt is the cross - entropy loss \(L\) cross and the contrastive loss \(L\) con which is the weighted sum as shown in Equation (16):

[0070] \(L\) F \(=\) \(L\) cross \(+\beta L\) con (16)

[0071] where \(\beta\) represents the weight of the contrastive loss.

[0072] A4. Backpropagation and parameter update:

[0073] After each forward propagation, calculate the total loss function \(L\) F and perform backpropagation to minimize the total loss as shown in Equation (17):

[0074]

[0075] where \(\theta\) represents the model parameters; \(\eta\) represents the learning rate; represents the gradient of the loss function with respect to the model parameters.

[0076] Compared with the prior art, the advantages of the present invention are as follows:

[0077] 1. The present invention introduces semantic embeddings generated by a pre - trained language model, enhancing the discriminability of features. Through multi - level attention fusion and the selection of key tokens that combine semantic information, the present invention improves the ability to locate discriminative feature regions. In addition, through the secondary chunking and selection of key tokens, the present invention further refines the feature extraction process, captures more discriminative fine - grained features, and improves the classification accuracy of the model in scenarios such as complex backgrounds and highly similar sub - categories.

[0078] 2. The present invention fuses global features and local features through a cross - attention mechanism, enhancing the discriminability of features and improving the classification accuracy.

[0079] 3. The present invention combines cross - entropy loss and contrastive loss. Compared with the method of using a single cross - entropy loss function, the introduction of contrastive learning is beneficial to enhancing the discriminability of features, alleviating the problem of difficult to distinguish subtle differences between categories, and improving the classification accuracy. Brief Description of the Drawings

[0080] Figure 1 is the flowchart of the method of the present invention;

[0081] Figure 2 is the network structure diagram of the present invention. Detailed Embodiments

[0082] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only specific embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0083] The present invention captures discriminative image features through key tokens, introduces semantic information, and uses a two-stage ViT model. In the first stage, key tokens with global features are captured, and in the second stage, based on the key tokens, the blocks are refined again to capture key tokens with more discriminative local features, alleviating the deficiencies of inaccurate regional localization and incomplete feature extraction of discriminative features of fine-grained image subcategories. This paper applies computer vision technology to the fine-grained image classification task, proposes a fine-grained image classification method based on feature fusion and semantic enhancement, and provides a new idea for the fine-grained image classification problem. Specifically:

[0084] As Figure 1 、 Figure 2 shown, a fine-grained image classification method based on feature fusion and semantic enhancement includes the following steps:

[0085] S1. Feature extraction of the ViT model: The input image I is divided into non-overlapping patches and converted into an embedding vector E through linear projection i , and then input into the Transformer encoder to generate the global feature E global ;

[0086] The step S1 includes:

[0087] S101. Image patching and linear projection:

[0088] The input image I is divided into a number of non-overlapping patches (i.e., tokens) of size 16×16, and then each patch P i is converted into an embedding vector E i as shown in Equation (1):

[0089] E i = LinearProjection(P i ) (1)

[0090] where E i represents the embedding vector of the i-th patch.

[0091] S102. Embedding sequence combination:

[0092] Combine all the embedded vectors into a sequence E = [E1, E2,..., E i , and add a classification token embedded vector E0:

[0093] E′ = [E0, E1, E2,..., E i

[0094] S103. Transformer Encoder:

[0095] Input the embedded vector sequence E′ into a multi-layer Transformer encoder, and generate a query matrix Q, a key matrix K, and a value matrix V through linear transformation. Each layer of the Transformer encoder consists of a multi-head self-attention mechanism and a feed-forward neural network. The calculation process of the attention mechanism is shown in Equation (2):

[0096]

[0097] where softmax represents the softmax activation function d k represents the dimension of the key matrix; A l represents the attention weight matrix of the l-th layer.

[0098] S104. Global Feature Obtaining:

[0099] A1. In the Transformer encoders of all layers, the classification token E0 will absorb the information from other tokens layer by layer and be updated to form a vector E global .

[0100] A2. Perform a weighted sum of all layers of attention weight matrices to obtain the global feature representation Z globa , Z globa The calculation formula is shown in Equation (3):

[0101]

[0102] where V l represents the value matrix of the l-th layer; A l represents the attention weight matrix of the l-th layer; L represents the total number of layers of the Transformer; Z global represents the global feature representation.

[0103] S2. Multi-level Attention Fusion and Key Token Selection Combining Semantic Information: Extract the attention weights A l in each layer of the Transformer, and combine the semantic embedding emb generated by the pre-trained language model to calculate the importance score s i ​, select the key token and denote it as z key ;

[0104] The step S2 includes:

[0105] S201. Extract and fuse attention weights: Extract the attention weight A of each token from each layer of Transformer in S1 l . Perform matrix multiplication fusion on the attention weights of all layers to generate the fused global attention, a final The calculation process is shown in Equation (4):

[0106]

[0107] where a final represents the fused global attention, represents the attention weight of each token in the first layer.

[0108] S202. Normalization processing:

[0109] Perform normalization processing on the attention weights of each token in the global attention, The calculation formula is shown in Equation (5):

[0110]

[0111] where a final,i represents the attention weight of the i-th token, ∑ j a final,j represents the sum of the attention weights of all tokens in the global attention; represents the attention weight of the i-th token after normalization.

[0112] S203. Generate semantic embeddings and fuse semantic information to calculate importance scores:

[0113] Use the pre-trained language model to generate semantic embeddings emb for each token, then perform element-wise multiplication (Hadamard product) on the normalized attention weights and the generated semantic embeddings, and perform normalization processing to generate enhanced attention weights, as shown in Equation (6):

[0114]

[0115] where ⊙ represents element-wise multiplication; emb i represents the semantic embedding vector of the i-th token; a enhanced,i represents the enhanced attention weight of the i-th token.

[0116] Use the enhanced attention weight aenhanced,i Calculate the importance score, s, of each token i The calculation process is shown in Equation (7):

[0117] s i = a enhanced,i · emb i (7)

[0118] Among them, s i represents the importance score of the i-th token. Select several tokens with the highest scores as the key tokens, denoted as z key .

[0119] S3. Semantic enhancement and refined key token selection: Perform secondary chunking and projection on z key in Step 2 to generate a new embedding vector E′ i,j , re-enter it into the Transformer encoder, and select the secondary key tokens again, denoted as z′ key ;

[0120] The said Step S3 includes:

[0121] S301. Perform secondary chunking and linear projection on the key tokens:

[0122] Perform further chunking on z key selected in S2. Each key token is divided into patches of size 4×4 again. Perform linear projection on these new chunks of patches to generate new embedding vectors, as shown in Equation (8):

[0123] E′ i,j = LinearProjection(P i,j ) (8)

[0124] Among them, P i,j represents the j-th new chunk of the i-th key token; E′ i,j represents the embedding vector of the j-th new chunk of the i-th key token.

[0125] S302. Embedding sequence combination:

[0126] Combine these embedding vectors into a new sequence, add the classification token E0, and input it into the Transformer encoder for processing, as shown in Equation (9):

[0127] E″ i = [E0, E′ i,1 , E′ i,2 , …, E′ i,j (9)

[0128] S303. Secondary selection of key tokens:

[0129] The calculation method is the same as that in S2. Calculate the secondary attention weight, and select the secondary key token denoted as z′ key .

[0130] S4. Cross-attention fusion: Through the cross-attention mechanism, fuse the global feature Z global and the local feature z′ key to generate the fused feature Z fused ;

[0131] The step S4 includes:

[0132] S401. Introduce the cross-attention mechanism:

[0133] After the secondary key token selection, introduce the cross-attention mechanism to fuse the global feature and the local feature, and calculate the cross-attention as shown in Equation (10):

[0134]

[0135] where Q represents the query matrix of the global feature Z global ; K represents the key matrix of the secondary key token feature z′ key ; d k represents the dimension of the key matrix; A cross represents the cross-attention weight matrix.

[0136] S402. Apply the cross-attention weight to generate the fused feature as shown in Equation (11):

[0137] Z fused = A cross V (11)

[0138] where V represents the value matrix of the secondary key token feature z′ key , and Z fused represents the fused feature matrix.

[0139] S5. Feature fusion and classification: Combine the fused feature Z fused with the global classification feature E global and input it into the classifier for classification to generate the classification output y′;

[0140] The step S5 includes:

[0141] S501. Final feature fusion:

[0142] Combine the cross-attention fusion feature Z fusedFuse with the global classification feature vector E obtained in S1 global as shown in Equation (12):

[0143] Z final = [E global , Z fused (12)

[0144] where Z final represents the finally fused feature vector.

[0145] S502. Classification output:

[0146] Input the finally fused feature vector Z containing the classification token and the cross-attention fusion feature final into the classifier for classification, as shown in Equation (13):

[0147] y' = Classifier(Z final ) (13)

[0148] where Classifier represents the classifier; y' represents the classification output of the model.

[0149] S6. Loss function design and training: Design the cross-entropy loss L cross and the contrastive loss L con , and optimize the model parameters through backpropagation.

[0150] The step S6 includes:

[0151] A1. Calculate the cross-entropy loss:

[0152] After each forward propagation ends, calculate the cross-entropy loss of the classification task, as shown in Equation (14):

[0153]

[0154] where y n represents the true label of the nth sample; p n represents the predicted probability of the nth sample; L cross represents the cross-entropy loss.

[0155] A2. Calculate the contrastive loss:

[0156] To further extract discriminative features, maximize the similarity between samples of the same class, and at the same time minimize the similarity between samples of different classes. After each forward propagation ends, calculate the contrastive loss as shown in Equation (15):

[0157]

[0158] Among them, B represents the batch size; α represents the hyperparameter that controls the similarity threshold; y n and y m represent the labels of the nth sample and the mth sample; Sim(z n , z m ) represents the similarity between the nth and the mth samples; is the Kronecker delta function, which takes the value 1 when y n = y m and 0 otherwise; L con represents the contrastive loss.

[0159] A3. Combined loss function:

[0160] The final loss function L F is the weighted sum of the cross-entropy loss L cross and the contrastive loss L con , as shown in Equation (16):

[0161] L F = L cross + βL con (16)

[0162] Among them, β represents the weight of the contrastive loss.

[0163] A4. Backpropagation and parameter update:

[0164] After each forward propagation, calculate the total loss function L F and perform backpropagation to minimize the total loss, as shown in Equation (17):

[0165]

[0166] Among them, θ represents the model parameters; η represents the learning rate; represents the gradient of the loss function with respect to the model parameters.

[0167] Although the above-described illustrative specific embodiments of the present invention have been described to facilitate understanding of the present invention by those skilled in the art, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

Claims

1. A fine-grained image classification method based on feature fusion and semantic enhancement, characterized in that, It includes the following steps: S1. ViT model feature extraction: The input image I is segmented into non-overlapping patches and converted into an embedding vector E through linear projection i , which is input into the Transformer encoder to generate the global feature E global ; S2. Multi-level attention fusion and key token selection combining semantic information: Extract the attention weights A in each layer of the Transformer l and combine with the semantic embedding emb generated by the pre-trained language model to calculate the importance score s of each token i select the key tokens and denote them as z key ; S3. Semantic enhancement and refined key token selection: Perform secondary chunking and projection on z in S2 key to generate a new embedding vector E' i,j , re-enter it into the Transformer encoder, and select the secondary key token again, denoted as z' key ; S4. Cross-attention fusion: Through the cross-attention mechanism, the global feature Z global and the local feature z′ key are fused to generate the fused feature Z fused ; S5. Feature fusion and classification: Fuse the feature Z fused with the global classification feature E global and input them into a classifier for classification to generate a classification output y'; S6. Loss function design and training: Design the cross-entropy loss L cross and the contrastive loss L con , and optimize the model parameters through backpropagation; The ViT model feature extraction in step S1 includes: S101. Image chunking and linear projection: The input image I is divided into a number of non-overlapping patches of size 16×16, and then each patch P i is converted into an embedding vector E i as shown in Equation (1): E i = LinearProjection(P i ) (1) Among them, E i represents the embedding vector of the i-th patch; S102. Embedding sequence combination: Combine all the embedding vectors into a sequence E = [E1, E2,..., E i , and add a classification token embedding vector E0: E' = [E0, E1, E2,..., E i ; S103. Transformer encoder: Input the embedding vector sequence E′ into a multi-layer Transformer encoder, and generate a query matrix Q, a key matrix K, and a value matrix V through linear transformation. Each layer of the Transformer encoder consists of a multi-head self-attention mechanism and a feed-forward neural network; the calculation process of the attention mechanism is shown in Equation (2): Among them, softmax represents the softmax activation function d k represents the dimension of the key matrix; A l represents the attention weight matrix of the l-th layer; S104. Global feature: A1. In the Transformer encoders of all layers, the classification token E0 will absorb information from other tokens layer by layer and update to form a vector E containing global features global ; A2. Weight-sum the attention weight matrices of all layers to obtain the global feature representation Z fused by the attention mechanism global As shown in Equation (3): Among them, V l represents the value matrix of the l-th layer; A l represents the attention weight matrix of the image I at the l-th layer; L represents the total number of layers of the Transformer; Z global represents the global feature representation; The multi-level attention fusion and key token selection combining semantic information in step S2 includes: S201. Extract and fuse attention weights: Extract the attention weight A of each token from each layer of the Transformer in S1 l ; Perform matrix multiplication fusion on the attention weights of all layers to generate the fused global attention as shown in Equation (4): Among them, a final represents the fused global attention, represents the attention weight of each token at the l-th layer; S202. Normalization processing: Normalize the attention weights of each token in the global attention, as shown in Equation (5): Among them, a final,i represents the attention weight of the i-th token, and ∑ j a final,j represents the sum of the attention weights of all tokens in the global attention; represents the normalized attention weight of the i-th token; S203. Generate semantic embeddings and fuse semantic information to calculate importance scores: Use a pre-trained language model to generate semantic embeddings emb for each token, then multiply the normalized attention weights element-wise with the generated semantic embeddings, and perform normalization processing to generate enhanced attention weights, as shown in Equation (6): where, ⊙ represents element-wise multiplication; emb i represents the semantic embedding vector of the i-th token; a enhanced,i represents the attention weight of the enhanced i-th token; Use the enhanced attention weight a enhanced,i Calculate the importance score for each token as shown in Equation (7): s i = a enhanced,i ·emb i (7) Among them, s i represents the importance score of the i-th token, and several tokens with the highest scores are selected as key tokens, denoted as z key .

2. The fine-grained image classification method based on feature fusion and semantic enhancement according to claim 1, characterized in that, The semantic enhancement and refined key token selection in S3 includes: S301. Perform secondary chunking and linear projection on the key tokens: Select the z selected in S2 key Perform further chunking, splitting each key token into patches of size 4×4 again; perform linear projection on these newly chunked patches to generate new embedding vectors, as shown in Equation (8): E′ i,j = LinearProjection(P i,j ) (8) where, P i,j represents the j-th new chunk of the i-th key token; E' i,j represents the embedding vector of the j-th new chunk of the i-th key token; S302. Embedding sequence combination: Combine these embedding vectors into a new sequence, add a classification token E0, and input it into the Transformer encoder for processing, as shown in Equation (9): E″ i = [E0, E′ i,1 , E′ i,2 , …, E′ i,j (9) S303. Secondary selection of key tokens: The calculation method is the same as S2. Calculate the secondary attention weight and select the secondary key token, denoted as z'. key。 3. A fine-grained image classification method based on feature fusion and semantic enhancement according to claim 2, characterized in that, The cross-attention fusion in S4 includes: S401. Introduce the cross-attention mechanism: After the secondary key token selection, introduce the cross-attention mechanism to fuse global features and local features, and calculate the cross-attention as shown in Equation (10): where Q represents the query matrix of the global feature Z global ; K represents the key matrix of the secondary key token feature z′ key ; d k represents the dimension of the key matrix; A cross represents the cross-attention weight matrix; S402. Apply the cross-attention weights to generate fused features as shown in Equation (11): Z fused = A cross V (11) Among them, V represents the value matrix of the secondary key token feature z′ key , Z fused represents the fused feature matrix.

4. The fine-grained image classification method based on feature fusion and semantic enhancement according to claim 3, characterized in that The feature fusion and classification in S5 includes: S501. Final feature fusion: Fuse the cross-attention fusion feature Z fused with the global classification feature vector E obtained in S1 global for fusion as shown in Equation (12): Z final = [E global , Z fused (12) Among them, Z final represents the finally fused feature vector; S502. Classification output: The final fused feature vector Z containing the classification token and the cross-attention fusion feature final is input into a classifier for classification, as shown in Equation (13): y′=Classifier(Z final ) (13) Among them, Classifier represents the classifier; y′ represents the classification output of the model.

5. The fine-grained image classification method based on feature fusion and semantic enhancement according to claim 4, wherein The loss function design and training in S6 includes: A1. Calculate the cross-entropy loss: After each forward propagation ends, calculate the cross-entropy loss of the classification task, as shown in Equation (14): where y n represents the true label of the nth sample; p n represents the predicted probability of the nth sample; L cross represents the cross-entropy loss; A2. Calculate the contrastive loss: To further extract discriminative features, maximize the similarity between samples of the same category, and minimize the similarity between samples of different categories. After each forward propagation ends, calculate the contrastive loss, as shown in Equation (15): Among them, B represents the batch size; α represents the hyperparameter that controls the similarity threshold; y n and y m represent the labels of the nth sample and the mth sample; Sim(z n , z m ) represents the similarity between the nth and the mth samples; is the Kronecker delta function, which takes the value 1 when y n = y m and 0 otherwise; L con represents the contrastive loss; A3. Combine the loss functions: The final loss function L F is the cross-entropy loss L cross and the contrastive loss L con which is a weighted sum as shown in Equation (16): L F = L cross + βL con (16) Among them, β represents the weight of the contrastive loss; A4. Backpropagation and parameter update: After each forward propagation, calculate the total loss function L F And perform backpropagation to minimize the total loss, as shown in Equation (17): Among them, θ represents the model parameters; η represents the learning rate; represents the gradient of the loss function with respect to the model parameters.

Citation Information

Patent Citations

  • Transform-based fine-grained image classification method

    CN114676776A

  • Fine-grained image classification method based on foreground and background multistage division and feature fusion

    CN118262180A

  • Transform-based fine-grained image classification method

    CN118135290A

  • Weak supervision image semantic segmentation method based on sample mixing and contrast learning

    CN118154884A