A Fine-Grained Classification Method Based on Multi-Granularity Interaction and Feature Recombination Network
By using the Swin-Transformer backbone network and self-attention mechanism for local positioning in the fine-grained visual classification task, and combining multi-grained feature interaction and feature recombination network, the problem of capturing subtle differences and overfitting in fine-grained visual classification is solved, and a more efficient fine-grained image classification is achieved.
Patent Information
- Application Number
- CN202310867063.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-14
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-07-14
AI Technical Summary
In the fine-grained visual classification task, it is difficult for the prior art to effectively capture subtle differences between objects with high visual similarity, and there is also the problem of overfitting.
Fine-grained global image features are extracted using a backbone network based on Swin-Transformer, and local image positioning is guided through self-attention weights, combining multi-grained feature interaction and feature recombination networks to enhance granularity perceived features and regional feature descriptions.
It realizes efficient discriminant local area positioning for fine-grained images, combines global and local feature representation learning, improves fine-grained visual classification performance and reduces the risk of overfitting.
Smart Images

Figure CN116883748B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a fine-grained classification method based on a multi-granularity interaction and feature recombination network. Background Art
[0002] Fine-grained visual classification (FGVC) has received extensive attention in computer vision research. Different from traditional visual classification tasks, the target categories of fine-grained visual classification are more detailed, aiming to distinguish sub-categories with similar visual appearances (such as classifying sub-categories of objects such as birds, cars, airplanes, etc.). As one of the basic capabilities of the visual system, FGVC has become the basis for various visual applications in real-world scenarios. Due to the visual similarity between fine-grained sub-categories, they can only be distinguished by capturing subtle visual differences. At the same time, there are also huge visual differences such as lighting, background, pose, and occlusion between objects of the same sub-category, further increasing the complexity of this task. Therefore, compared with traditional visual classification tasks, fine-grained visual classification is a unique and more challenging problem.
[0003] In recent years, with the rapid development of deep convolutional neural networks, FGVC has achieved great development. Due to the powerful feature extraction ability of convolutional neural networks, existing fine-grained visual classification research mainly focuses on using convolutional neural networks (CNNs) for feature extraction to learn subtle differences. Some FGVC research works obtain discriminative components for learning by localizing components. However, annotating based on object bounding boxes and component annotations requires professional knowledge. Limited by this, methods for automatically mining discriminative features through attention and other methods to achieve localization only using image-level labels have received more attention recently. Another part of the methods learn subtle differences by encoding discriminative features. However, such methods ignore the perception of context relationships to construct a feature representation for global information description, and to a certain extent, some key information is lost. Benefiting from the application of the self-attention mechanism, Transformer has shown superior performance in various recent visual tasks. Self-attention models a global dependence relationship by calculating the connections between different positions. Compared with convolutional neural networks, Transformer has stronger global information representation ability and can construct the context relationship of the visual space to obtain a richer feature representation. Therefore, in the FGVC task, Transformer can be an effective alternative to convolutional neural networks. However, the structure that focuses on global representation will cause a decline in the local information extraction ability and it is difficult to capture the subtle differences between objects with high visual similarity. At the same time, the data volume of each sub-category in FGVC is small and there are large variations. Constructing complex feature representations will lead to the classification model overfitting to certain specific feature patterns. Therefore, effective solutions are urgently needed. Summary of the Invention
[0004] The present invention proposes a fine-grained classification method based on a multi-granularity interaction and feature recombination network, which can accurately and effectively perform fine-grained classification.
[0005] The present invention adopts the following technical solutions.
[0006] A fine-grained classification method based on a multi-granularity interaction and feature recombination network includes the following steps:
[0007] Step S1: Extract the features of the fine-grained global image through a backbone network based on Swin-Transformer, then guide local image localization through self-attention weights, and extract local features in the form of weight sharing;
[0008] Step S2: Enhance the granularity-aware features by embedding a multi-granularity feature enhancement module, and at the same time combine cross-attention feature interaction to further enrich the region-level feature description;
[0009] Step S3: Use dynamic class-level center representation to guide high-difference channel recombination and exchange to retain potential class-invariant features and explore diverse combinations of feature patterns;
[0010] Step S4: Perform iterative training according to the specified training parameters, update the model parameters by optimizing the combined loss, continuously save the optimal model according to the validation accuracy, and obtain the combined prediction result using the final model.
[0011] Step S1 specifically includes the following steps;
[0012] Step S11: Calculate the patch importance factor; specifically: input the fine-grained global image X global to the Swin-T backbone network, and use the feature embedding F=(F (1) ,..., F (patch_N) ) T , w′ patch_i represents the weight of each patch, which is obtained by taking the mean of F (patch_i) along the embedding size D. Then use Softmax to normalize W′ p =(w′ 1 ,..., w′ patch_N ) T to obtain the patch importance factor W p =(w 1 ,..., w patch_N ) T , and the specific calculation formula is: w′ patch_i =meann(F (patch_i) )
[0013] where mean(·) represents calculating the element-wise mean of the input, patch_i ∈ [1, patch_N], and patch_N represents the number of patches;
[0014]
[0015] Step S12: Calculate the attention localization map; specifically: Obtain the attention weight matrix Att through matrix multiplication of the attention weights of each layer in the last stage of the Swin-T backbone network in a recursive manner; then, using the attention weight matrix Att = (Att 1 ,..., Att patch_N ) T , re-weight and aggregate the attention weights assigned to each patch through the patch importance factor W p to obtain the attention localization map G. The specific calculation process is as follows
[0016]
[0017] where layer_L represents the corresponding layer number of this stage, and Att layer_i represents the attention weight matrix calculated for the layer_layer_i of this stage.
[0018]
[0019] where Att patch_j represents the vector of each patch along the embedding size D dimension, and w patch_j represents the weight assigned to the patch_j-th patch;
[0020] Step S13: Calculate the mask matrix. Further calculate the patch mask matrix Mask(pos_x, pos_y) through the attention weight matrix Att(pos_x, pos_y). The specific calculation method is as follows
[0021]
[0022] where pos_x, pos_y represent the positions corresponding to each patch in the plane, represents the result after the element-wise mean operation of the localization map G;
[0023] Step S14: Using the above-computed mask matrix Mask(pos_x, pos_y), the local region that the Transformer network pays the most attention to for the input fine-grained image is obtained by extracting the largest connected component; guided by the self-attention mechanism, this region contains less noise interference and a higher proportion of discriminative information, which helps to improve the performance of fine-grained visual classification; then, the relevant region is further located and enlarged to obtain the local localization image X local , together with the fine-grained global image X gobal , are input into the network, and the parameters of the entire network are shared for model training.
[0024] Step S2 specifically includes the following steps;
[0025] Step S21: Using the Swin-Transformer backbone network, that is, the fine-grained global image X extracted by the Swin-T backbone network global , the multi-granularity enhanced feature global_F of the global image is further calculated E , and the specific calculation method is as follows
[0026] global_F E = DE(Up(global_F layer_l )) + DE(global_F layer_l - 1)
[0027] where global_F l represents the feature embedding output by the l-th layer of the Swin-T backbone network, layer_l ∈ [1, 4]. DE(·) represents the detail enhancement layer, and Up(·) represents the upsampling operation. The one-dimensional convolution is used to upsample the feature embedding to align the feature embedding dimensions of the previous stage;
[0028] The detail enhancement layer takes the features output by the specified stage of the Swin-T backbone network as input. First, the 1×1 convolutional layer is used to integrate the original embedding representation, and the proposed local information supplement layer is further added to it. Finally, the 1×1 convolutional layer is used again to integrate the preliminarily enhanced feature embedding; the local information supplement is implemented in the form of a residual, and the specific calculation is as follows
[0029] global_F′ layer_l = Conv(global_F layer_l ; conv_θ 1 ),
[0030] global_F″ layer_l = conv(σ(global′ layer_l + LS(global_F′ layer_l; ls_θ)); conv_θ 2 )
[0031] Among them, LS(·; θ) represents the local information supplement layer with parameters ls_θ, which is implemented using a 3×3 depth convolution. σ(·) represents the GELU function, Conv(·) represents the 1×1 convolution layer, and conv_θ 1 and conv_θ 2 respectively represent the corresponding parameters of two 1×1 convolution layers;
[0032] Step S22: Input the local positioning image X local into the Swin-T backbone network with shared weights, and calculate the multi-granularity enhanced feature local_F of the local image in the same way as in step S21 E ;
[0033] Step S23: Calculate the query, key, and value matrices. The original multi-granularity feature interaction module feature F Multi ; respectively perform different linear transformations on the multi-granularity enhanced feature global_F of the global image E and the multi-granularity enhanced feature local_F of the local image E to calculate the corresponding Query 1 , Key 2 , Value 2 , which respectively correspond to the query, key, and value matrices. The specific calculation methods are as follows
[0034]
[0035]
[0036]
[0037] Among them, respectively represent the corresponding linear transformation weight matrices;
[0038] Step S24: Calculate the original multi-granularity feature interaction module feature; specifically: perform cross-attention calculation on the query matrix Query 1、 key matrix Key 2 and value matrix Value 2 calculated in step S23. Through the inner product operation of the corresponding elements in Query 1 and Key 2 , calculate the weights for each value vector in the value matrix, and map them to the final output through normalization. The specific calculation methods are as follows
[0039]
[0040] Among them, represents the scaling factor, and the Softmax function is used to normalize the attention weight matrix. The output result of this function is the feature F of the original multi-granularity feature interaction module Multi .
[0041] Step S3 specifically includes the following steps;
[0042] Step S31: Calculate the channel average spatial representation; specifically: generate class centers in a parameter-free manner as fine-grained class-level feature representations, and iteratively update them during the training process. This generation method ensures the overall feature representation of training samples of the same class, and at the same time does not require additional learnable parameters to optimize to fit the data distribution of fine-grained training samples, reducing the sensitivity of the network to the data distribution; specifically, define the class center representation of the i-th class in the t-th epoch as To generate a robust feature center representation for each category from the sample feature space, calculate the mean of all training samples of the same category. The specific calculation is as follows
[0043]
[0044] where epoch_t represents the t-th epoch in the model training process, represents the j-th sample of the i-th class in the t-th epoch, and K class_i represents the total number of training samples of the i-th class;
[0045] For the class feature centers obtained in each iteration process calculated above, further calculate the channel average spatial representation of the i-th class in the t-th epoch Since the model can easily obtain the class feature center representation of training samples in each parameter iteration process, therefore, the channel average spatial representation can be updated in a simple way during each parameter iteration process. The specific calculation process is as follows
[0046]
[0047] where represents the j-th channel feature of the i-th class in the t-th epoch; D represents the total number of channels of each patch;
[0048] Step S32: Calculate the weight score of each channel of the i-th class of the i-th sample The specific calculation method is as follows
[0049]
[0050] where channel_N represents the number of channels, is the average spatial representation of the channels of the i-th class calculated for the (t-1)-th epoch (i.e., the previous round). It should be noted that at the start of training, it is randomly initialized with Gaussian parameters, and the class centers are initialized as random points in the feature space;
[0051] Step S33: Further calculate the channels with high differences relative to the class centers. Specifically, by calculating the channel mask of the current i-th sample filter the channel features with high differences, and the calculation method is as follows
[0052]
[0053] where channel_d represents the channel dimension index of the feature, the index starts from 0, and channel_d ∈ [0, D - 1]; m pp (·) represents the value of the (pp × D)-th element obtained by sorting in descending order, pp is the channel selection probability parameter, pp ∈ [0, 1], which determines the proportion of feature channel recombination;
[0054] Step S34: Randomly select another training sample during training to recombine the channel features of the current training sample. By obtaining more feature combinations, the learning of the classifier for fine-grained object features is stabilized, so as to retain the relatively stable class-invariant channel feature representations and exchange and confuse the high-difference channel features; the recombined features of the i-th sample in the training batch are denoted as The specific calculation process is as follows
[0055]
[0056] where, M sample_i represents the channel mask calculated for the i-th sample, and F sample_j represents the j-th training sample in the same batch during training;
[0057] Step S35: Update the class center representation after the end of the current training epoch
[0058] Step S4 specifically includes the following steps;
[0059] Step S41: Calculate the overall loss function L of the network total,The training losses of the fine-grained global image and the local localization image obtained by the self-attention guided localization module are respectively denoted as L raw、 L local , and the loss of the feature recombination enhancement strategy during training is denoted as L re , and the specific calculation formula is as follows
[0060] L raw = f cls (Cls(F Global ), gt)
[0061] L local = f cls (Cls(F Local ), gt)
[0062]
[0063] L total = L raw + L local + L re
[0064] where F Global represents the features of the fine-grained global image extracted by the Swin-T backbone network, F Local represents the features of the local localization image extracted by the Swin-T backbone network, Cls(·) represents the classifier, gt represents the ground-truth label of the input image, and f cls (·) represents the cross-entropy function; represents the features after the output of the multi-granularity feature interaction module is recombined for the times_r-th time, and times_R represents the total number of times the feature recombination enhancement strategy is applied during training;
[0065] Step S42: Perform iterative training according to the specified parameters, and continuously update the gradient for iterative training according to the overall loss L total calculated in Step S41;
[0066] Step S43: Calculate the predicted label through the fine-grained global image features F Gobal , local localization image features F Local , and original multi-granularity feature interaction module features F Multi extracted by the model, and its calculation method is as follows
[0067]
[0068] Step S44: During the training process, model validation is performed at a certain iteration interval according to the validation interval flag, and the optimal model is continuously saved. When the iteration number reaches the preset maximum iteration number threshold, the training process ends, and the optimal fine-grained prediction accuracy of the current specified dataset is returned.
[0069] In step S4, when training the model, the training image data used is labeled by category, and there is no need to manually label the corresponding annotation boxes when locating the discriminative local regions of the image data.
[0070] Compared with the prior art, the present invention has the following beneficial effects:
[0071] 1. A fine-grained classification method based on a multi-granularity interaction and feature recombination network constructed by the present invention can accurately and effectively perform efficient discriminative local region localization on fine-grained images, and jointly perform global and local feature representation learning to improve the fine-grained visual classification performance.
[0072] 2. The present invention constructs a self-attention guided localization method, which uses self-attention weight aggregation to guide the adaptive selection of relatively important regions of the image, and learns discriminative region-level feature representations.
[0073] 3. The present invention constructs a multi-granularity feature interaction learning method, which enhances the granularity-aware features by embedding a multi-granularity feature enhancement module, and at the same time combines cross-attention feature interaction to further enrich the region-level feature description, so as to promote the learning of spatial context recognition clues.
[0074] The present invention constructs a feature recombination enhancement method, which uses dynamic class-level center representations to guide the recombination and exchange of high-difference channels, so as to retain potential class-invariant features and explore diverse combinations of feature patterns, thereby effectively improving the robustness of the FGVC model. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] The following further details the present invention in conjunction with the drawings and specific embodiments:
[0076] Att Figure 1 is a schematic diagram of the principle of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0077] The following further describes the present invention in conjunction with the drawings and embodiments.
[0078] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0079] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0080] As shown in the figure, a fine-grained classification method based on a multi-granularity interaction and feature recombination network includes the following steps:
[0081] Step S1: Extract the features of the fine-grained global image through a backbone network based on Swin-Transformer, then guide the local image localization through the self-attention weights, and extract the local features in the form of weight sharing;
[0082] Step S2: Enhance the granularity-aware features by embedding a multi-granularity feature enhancement module, and at the same time combine cross-attention feature interaction to further enrich the region-level feature description;
[0083] Step S3: Use the dynamic class-level center representation to guide the high-difference channel recombination and exchange, so as to retain the potential class-invariant features and explore diverse combinations of feature patterns;
[0084] Step S4: Perform iterative training according to the specified training parameters, update the model parameters by optimizing the combined loss, continuously save the optimal model according to the validation accuracy, and use the combined prediction results obtained by the final model.
[0085] Step S1 specifically includes the following steps;
[0086] Step S11: Calculate the patch importance factor; specifically: Input the fine-grained global image X global to the Swin-T backbone network, and use the feature embedding F=(F (1) ,..., F (patch_N) ) T , w′ patch_i represents the weight of each patch, which is obtained by taking the mean of F (patch_i) along the embedding size D. Then use Softmax to normalize W′ p =(w′ 1 ,..., w′ patch_N ) T to obtain the patch importance factor W p =(w 1 ,..., w patch_N ), and the specific calculation formula is: w′ T , and the specific calculation formula is: w′patch_i = mean(F (patch_i) )
[0087] where mean(·) represents calculating the element-wise mean of the input, patch_i ∈ [1, patch_N], and patch_N represents the number of patches;
[0088]
[0089] Step S12: Calculate the attention localization map; specifically: Obtain the attention weight matrix Att by using matrix multiplication on the attention weights of each layer in the last stage of the Swin-T backbone network in a recursive manner; then, re-weight and aggregate the attention weights assigned to each patch through the patch importance factor W 1 ,..., Att patch_N ) T , and re-weight and aggregate the attention weights assigned to each patch through the patch importance factor W p to obtain the attention localization map G. The specific calculation process is as follows
[0090]
[0091] where layer_L represents the corresponding layer number of this stage, and Att layer_i represents the attention weight matrix calculated for the layer_i-th layer of this stage.
[0092]
[0093] where Att patch_j represents the vector of each patch along the embedding size D dimension, and w patch_j represents the weight assigned to the patch_j-th patch;
[0094] Step S13: Calculate the mask matrix. Further calculate the patch mask matrix Mask(pos_x, pos_y) through the attention weight matrix Att(pos_x, pos_y). The specific calculation method is as follows
[0095]
[0096] where pos_x, pos_y represent the positions corresponding to each patch in the plane, represents the result after the element-wise mean operation of the localization map G;
[0097] Step S14: Using the calculated mask matrix Mask(pos_x, pos_y), the local region that the Transformer network pays the most attention to for the input fine-grained image is obtained by extracting the largest connected component; guided by the self-attention mechanism, this region contains less noise interference and a higher proportion of discriminative information, which helps to improve the performance of fine-grained visual classification; then, the relevant region is further located and enlarged to obtain the local localization image X local , together with the fine-grained global image X gobal , are input into the network together, and the parameters of the entire network are shared for model training.
[0098] Step S2 specifically includes the following steps;
[0099] Step S21: Using the Swin-Transformer backbone network, that is, the fine-grained global image X extracted by the Swin-T backbone network global , the multi-granularity enhanced feature global_F of the global image is further calculated E , and the specific calculation method is as follows
[0100] global_F E =DE(Up(global_F layer_l )) + DE(global_F layer_l-1 )
[0101] Among them, global_F l represents the feature embedding output by the l-th layer of the Swin-T backbone network, layer_l ∈ [1, 4]. DE(·) represents the detail enhancement layer, and Up(·) represents the upsampling operation. The one-dimensional convolution is used to perform the upsampling operation on the feature embedding to align the feature embedding dimensions of the previous stage;
[0102] The detail enhancement layer takes the features output by the specified stage of the Swin-T backbone network as input. First, the 1×1 convolutional layer is used to integrate the original embedding representation, and the proposed local information supplement layer is further added to it. Finally, the 1×1 convolutional layer is used again to integrate the preliminarily enhanced feature embedding; the local information supplement is implemented in the form of a residual, and the specific calculation is as follows
[0103] global_F′ layer_l =Conv(globaLF layer_l ; conv_θ 1 ),
[0104] global_F″ layer_l =conv(σ(global′ layer_l +LS(global_F′ layer_l; ls_θ)); conv_θ 2 )
[0105] Among them, LS(·; θ) represents the local information supplementation layer with parameters ls_θ, which is implemented using a 3×3 depth convolution. σ(·) represents the GELU function, Conv(·) represents the 1×1 convolution layer, and conv_θ 1 and conv_θ 2 respectively represent the corresponding parameters of two 1×1 convolution layers;
[0106] Step S22: Input the local positioning image X local into the Swin-T backbone network with weight sharing, and calculate the multi-granularity enhanced feature local_F of the local image in the same way as in step S21 E ;
[0107] Step S23: Calculate the query, key, and value matrices. The original multi-granularity feature interaction module feature F Multi ; respectively perform different linear transformations on the multi-granularity enhanced feature global_F of the global image E and the multi-granularity enhanced feature local_F of the local image E to calculate the corresponding Query 1, Key 2 , Value 2 , which respectively correspond to the query, key, and value matrices, and the specific calculation methods are as follows
[0108]
[0109]
[0110]
[0111] Among them, respectively represent the corresponding linear transformation weight matrices;
[0112] Step S24: Calculate the original multi-granularity feature interaction module feature; specifically: perform cross-attention calculation on the query matrix Query 1 , key matrix Key 2 and value matrix Value 2 calculated in step S23. Through the inner product operation of the corresponding elements in Query 1 and Key 2 , calculate the weights for each value vector in the value matrix, and map them to the final output through normalization. The specific calculation methods are as follows
[0113]
[0114] Among them, represents the scaling factor, and the Softmax function is used to normalize the attention weight matrix. The output result of this function is the feature F of the original multi-granularity feature interaction module Multi .
[0115] Step S3 specifically includes the following steps;
[0116] Step S31: Calculate the channel-averaged spatial representation; specifically: generate class centers in a parameter-free manner as fine-grained class-level feature representations, and iteratively update them during the training process. This generation method ensures the overall feature representation of training samples of the same class, and at the same time does not require additional learnable parameters for optimization to fit the data distribution of fine-grained training samples, reducing the sensitivity of the network to the data distribution; specifically, define the class center representation of the i-th class in the t-th epoch as To generate a robust feature center representation for each category from the sample feature space, calculate the mean of all training samples of the same category, and the specific calculation is as follows
[0117]
[0118] where epoch_t represents the t-th epoch in the model training process, represents the j-th sample of the i-th class in the t-th epoch, and K class_i represents the total number of training samples of the i-th class;
[0119] For the class feature centers obtained in each iteration process calculated above, further calculate the channel-averaged spatial representation of the i-th class in the t-th epoch Since the model can easily obtain the class feature center representation of training samples in each parameter iteration process, therefore, the channel-averaged spatial representation can be updated in a simple way during each parameter iteration process, and the specific calculation process is as follows
[0120]
[0121] where represents the j-th channel feature of the i-th class in the t-th epoch; D represents the total number of channels of each patch;
[0122] Step S32: Calculate the weight score of each channel of the i-th class of the i-th sample The specific calculation method is as follows
[0123]
[0124] where channel_N represents the number of channels, is the average spatial representation of the channels of the i-th class calculated for the (t-1)-th epoch (i.e., the previous epoch). It should be noted that at the start of training, it is randomly initialized with Gaussian parameters, and the class centers are initialized as random points in the feature space;
[0125] Step S33: Further calculate the channels with high differences relative to the class centers. Specifically, by calculating the channel mask of the current sample_i to screen the channel features with high differences, and the calculation method is as follows
[0126]
[0127] where channel_d represents the channel dimension index of the feature, the index starts from 0, and channel_d ∈ [0, D - 1]; m pp (·) represents the value of the pp×D-th element obtained by sorting in descending order, pp is the channel selection probability parameter, pp ∈ [0, 1], which determines the proportion of feature channel recombination;
[0128] Step S34: Randomly select another training sample during training to recombine the channel features of the current training sample, and by obtaining more feature combinations, stabilize the learning of the classifier for fine-grained object features, so as to retain the relatively stable class-invariant channel feature representations and exchange and confuse the high-difference channel features; the recombined features of the sample_i in the training batch are denoted as The specific calculation process is as follows
[0129]
[0130] where, M sample_i represents the channel mask calculated for the sample_i, and F sample_j represents the sample_j-th training sample in the same batch during training;
[0131] Step S35: Update the class center representation after the end of the current training epoch
[0132] Step S4 specifically includes the following steps;
[0133] Step S41: Calculate the overall loss function L of the network total,The training losses of the fine-grained global image and the local localization image obtained by the self-attention guided localization module are respectively denoted as L raw and L local , and the loss of the feature recombination enhancement strategy during training is denoted as L re , and the specific calculation formula is as follows
[0134] L raw = f cls (Cls(F Global ), gt)
[0135] L local = f cls (Cls(F Local ), gt)
[0136]
[0137] L total = L raw + L local + L re
[0138] where F Global represents the features of the fine-grained global image extracted by the Swin-T backbone network, F Local represents the features of the local localization image extracted by the Swin-T backbone network, Cls(·) represents the classifier, gt represents the ground-truth label of the input image, and f cls (·) represents the cross-entropy function; represents the features after the r-th recombination of the output of the multi-granularity feature interaction module, and times_R represents the total number of times the feature recombination enhancement strategy is applied during training;
[0139] Step S42: Perform iterative training according to the specified parameters, and continuously update the gradient for iterative training according to the overall loss L total calculated in step S41;
[0140] Step S43: Calculate the predicted label through the fine-grained global image features F Gobal , local localization image features F Local , and original multi-granularity feature interaction module features F Multi extracted by the model, and its calculation method is as follows
[0141]
[0142] Step S44: During the training process, model validation is performed at a certain iteration interval according to the validation interval flag, and the optimal model is continuously saved. When the iteration number reaches the preset maximum iteration number threshold, the training process ends, and the optimal fine-grained prediction accuracy of the current specified dataset is returned.
[0143] In step S4, when training the model, the training image data used is labeled by category, and there is no need for manual annotation of the corresponding annotation box when locating the discriminative local region of the image data.
[0144] Specifically, in this embodiment, only category annotation is used, and a series of additional manual annotations such as annotation boxes are not required. In view of the uniqueness of the fine-grained visual classification task, the present invention proposes a new Transformer-based fine-grained visual classification framework. This framework guides the location of discriminative local regions through self-attention, and improves the model's ability to distinguish fine-grained subcategories through the joint learning of local features and global features. On this basis, the present invention learns more robust multi-granularity feature representations of global and local interactions through dynamic feature reconstruction, effectively improving the fine-grained image classification accuracy.
[0145] The above are only the preferred embodiments of the present invention, and all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope of the present invention.
Claims
1. A fine-grained classification method based on a multi-granularity interaction and feature recombination network, characterized in that: It includes the following steps: Step S1: Extract the features of the fine-grained global image through a backbone network based on Swin-Transformer, then guide local image localization through self-attention weights, and extract local features in the form of weight sharing; Step S2: Enhance the granularity-aware features by embedding a multi-granularity feature enhancement module, and at the same time combine cross-attention feature interaction to further enrich the regional-level feature description; Step S3: Use dynamic class-level center representation to guide high-difference channel recombination and exchange, to retain potential class-invariant features, and explore diverse feature pattern combinations; Step S4: Perform iterative training according to the specified training parameters, update the model parameters by optimizing the combined loss, continuously save the optimal model according to the validation accuracy, and obtain the combined prediction result using the final model; Specifically, step S3 includes the following steps; Step S31: Calculate the channel average spatial representation; specifically: generate class centers in a parameter-free manner as fine-grained class-level feature representations, and iteratively update them during the training process. This generation method ensures the overall feature representation of training samples of the same class, and at the same time does not require additional learnable parameters for optimization to fit the data distribution of fine-grained training samples, reducing the sensitivity of the network to the data distribution; specifically, define the class center representation of the \(i\)-th class in the \(t\)-th epoch as To generate a robust feature center representation for each class from the sample feature space, calculate the mean of all training samples of the same class, and the specific calculation is as follows Among them, epoch_t represents the epoch_t-th epoch in the model training process, represents the sample_j-th sample of the class_i-th class in the epoch_t-th epoch, K class_i represents the total number of training samples of the class_i-th class; For the class feature centers in each iteration calculated above, further calculate the channel-averaged spatial representation of the \(class_i\)-th class in the \(epoch_t\)-th epoch. Since the model can easily obtain the class feature center representation of the training samples in each parameter iteration process, therefore, the channel-averaged spatial representation can be updated in a simple way during each parameter iteration. The specific calculation process is as follows Among them represents the feature of the j-th channel of the i-th class in the epoch_t-th epoch; D represents the total number of channels of each patch; Step S32: Calculate the weight scores of each channel for each of the sample_i categories in the class_i-th category The specific calculation method is as follows where channel_N represents the number of channels, is the channel average spatial representation of the i-th class calculated for the (t-1)-th epoch. It should be noted that at the beginning of training, it is randomly initialized with Gaussian parameters, and the class centers are initialized as random points in the feature space; Step S33: Further calculate the high-difference channels relative to the category center; specifically, calculate the channel mask of the current sample_i sample Filter the high-difference channel features, and the calculation method is as follows where channel_d represents the channel dimension index of the feature, with the index starting from 0 and channel_d ∈ [0, D - 1]; m pp (·) represents the value of the pp × D-th element obtained by sorting in descending order. pp is the channel selection probability parameter, pp ∈ [0, 1], which determines the proportion of feature channel recombination; Step S34: During the training process, randomly select another training sample to recombine the channel features of the current training sample. By obtaining more feature combinations, the learning of the classifier for fine-grained object features is stabilized, so as to retain the relatively stable class-invariant channel feature representations and exchange and confuse the highly different channel features. The recombined features of the \(sample_i\)-th sample in the training batch are denoted as The specific calculation process is as follows Among them, M sample_i represents the channel mask calculated from the Msample_i-th sample, and F sample_j represents the sample_j-th training sample in the same batch during the training process; Step S35: Update the class center representation after the current training epoch ends 2. A fine-grained classification method based on a multi-granularity interaction and feature recombination network according to claim 1, characterized in that: Specifically, step S1 includes the following steps; Step S11: Calculate the patch importance factor; specifically: Input the fine-grained global image X global to the Swin-T backbone network, and use the feature embedding F = (F (1) ,..., F (patch_N) ) T output by the last stage of the Swin-T backbone network. Here, w′ patch_i represents the weight of each patch, which is obtained by taking the mean of F (patch_i) along the embedding size D; then use Softmax to normalize W′ p =(w′ 1 ,..., w′ patch_N ) T to obtain the patch importance factor W p =(w 1 ,..., w patch_N ) T . The specific calculation formula is: w′ patch_i =mean(F (patch_i) ) where mean(·) represents taking the element-wise mean of the input, patch_i ∈ [1, patch_N], and patch_N represents the number of patches; Step S12: Calculate the attention localization map; Specifically: the attention weight matrix Att is obtained by using matrix multiplication on the attention weights of each layer in the last stage of the Swin-T backbone network in a recursive manner; then, using the attention weight matrix Att = (Att 1 ,..., Att patch_N ), T , the attention weights assigned to each patch are re-weighted and aggregated through the patch importance factor W p to obtain the attention localization map G. The specific calculation process is as follows Among them, layer_L represents the layer corresponding to this stage, and Att layer_i represents the attention weight matrix calculated for the layer_i layer at this stage; Among them, Att patch_j represents the vector of each patch along the embedding size D dimension, and w patch_j represents the weight assigned to the j-th patch; Step S13: Calculate the mask matrix, and further calculate the patch mask matrix Mask(pos_x, pos_y) through the attention weight matrix Att(pos_x, pos_y). The specific calculation method is as follows Among them, pos_x and pos_y represent the positions corresponding to each patch under the plane, which represents the result after the element mean operation on the positioning map G; Step S14: Using the obtained mask matrix Mask(pos_x, pos_y) calculated above, the local area that the Transformer network pays the most attention to for the input fine-grained image is obtained by extracting the largest connected component; guided by the self-attention mechanism, this area contains less noise interference and a higher proportion of discriminative information, which helps to improve the performance of fine-grained visual classification; then, the relevant area is further located and enlarged to obtain the local localization image X local , which is input into the network together with the fine-grained global image X gobal , and the parameters of the entire network are shared for model training.
3. A fine-grained classification method based on a multi-granularity interaction and feature recombination network according to claim 1, characterized in that: Specifically, step S2 includes the following steps; Step S21: Use the Swin-Transformer backbone network, i.e., the fine-grained global image X extracted by the Swin-T backbone network, to further calculate the multi-granularity enhanced feature global_F of the global image global The specific calculation method is as follows E global_F E =DE(Up(global_F layer_l )) + DE(global_F layer_l-1 ) Among them, global_F l represents the feature embedding output by the l-th layer of the Swin-T backbone network, where layer_l ∈ [1, 4]; DE(·) represents the detail enhancement layer, and Up(·) represents the upsampling operation. A one-dimensional convolution is used to perform the upsampling operation on the feature embedding to align the feature embedding dimensions of the previous stage; The detail enhancement layer takes the features output by the specified stage of the Swin-T backbone network as input. First, it integrates the original embedding representation through a 1×1 convolutional layer, and further adds the proposed local information supplement layer therein. Finally, it uses the 1×1 convolutional layer again to integrate the preliminarily enhanced feature embedding; The local information supplement is implemented in the form of a residual, and the specific calculation is as follows global_F′ layer_l = Conv(global_F layer_l ; conv_θ 1 ), global_F″ layer_l = Conv(σ(global′ layer_l + LS(global - F′ layer_l ; ls_θ)); conv_θ 2 ) where LS(·; θ) represents the local information supplementation layer with parameter ls_θ, which is implemented using a 3×3 depth convolution; σ(·) represents the GELU function, Conv(·) represents the 1×1 convolution layer, and conv_θ 1 and conv_θ 2 represent the corresponding parameters of two 1×1 convolution layers respectively; Step S22: Input the local positioning image X local into the Swin-T backbone network with shared weights, and calculate the multi-granularity enhanced feature local_F of the local image in the same way as in step S21 E ; Step S23: Calculate the query, key, and value matrices; the feature F of the original multi-granularity feature interaction module Multi ; respectively perform different linear transformations on the multi-granularity enhanced feature global_F of the global image E and the multi-granularity enhanced feature local_F of the local image E to calculate the corresponding Query 1 , Key 2 , Value 2 , which respectively correspond to the query, key, and value matrices, and the specific calculation methods are as follows Among them, respectively represent the corresponding linear transformation weight matrices; Step S24: Calculate the features of the original multi-granularity feature interaction module; specifically: the query matrix Query calculated in step S23 1 , the key matrix Key 2 and the value matrix Value 2 perform cross-attention calculation; through the inner product operation of the corresponding elements in Query 1 and Key 2 calculate the weights for each value vector in the value matrix and map them to the final output through normalization. The specific calculation method is as follows Among them, represents the scaling factor, and the Softmax function is used to normalize the attention weight matrix; the output result of this function is the feature F of the original multi-granularity feature interaction module Multi .
4. A fine-grained classification method based on a multi-granularity interaction and feature recombination network according to claim 1, characterized in that: Step S4 Specifically includes the following steps; Step S41: Calculate the overall network loss function L total , and define the training losses of the fine-grained global image and the local localization image obtained by the self-attention guided localization module as L raw and L local respectively. The loss of the feature recombination enhancement strategy during training is expressed as L re . The specific calculation formula is as follows L raw = f cls (Cls(F Global ), gt) L local = f cls (Cls(F Local ), gt) L total = L raw + L local + L re Among them, F Global represents the features of the fine-grained global image extracted by the Swin-T backbone network, and F Local represents the features of the local localization image extracted by the Swin-T backbone network. Cls(·) represents the classifier, gt represents the ground-truth label of the input image, and f cls (·) represents the cross-entropy function; represents the features after the output of the multi-granularity feature interaction module is recombined for the times_r-th time, and times_R represents the total number of times the feature recombination enhancement strategy is applied during the training process; Step S42: Perform iterative training according to the specified parameters, and calculate the overall loss L according to Step S41 total Continuously update the gradient for iterative training; Step S43: Calculate the predicted label The fine-grained global image feature F extracted by the model Gobal , the local localization image feature F Local , the original multi-granularity feature interaction module feature F Multi , and its calculation method is as follows Step S44: During the training process, perform model validation at a certain iteration interval according to the validation interval flag, and continuously save the optimal model. When the iteration number reaches the preset maximum iteration number threshold, the training process ends, and the optimal fine-grained prediction accuracy of the current specified dataset is returned.
5. A fine-grained classification method based on a multi-granularity interaction and feature recombination network according to claim 1, characterized in that: When training the model in step S4, the training image data is labeled by category, and there is no need for manual annotation of the corresponding annotation boxes when locating the discriminative local regions of the image data.
Citation Information
Patent Citations
Monitoring video multi-granularity marking method based on generalized multi-labeling learning
CN107133569A
Multi-granularity information fusion fine-granularity image classification method and system
CN114299343A