Low-Resolution Fine-Grained Image Classification Method Based on Component Attention Correction and Detail Enhancement

By adopting component attention correction and detail enhancement methods in low-resolution fine-grained image classification, the problems of feature distribution differences and sparse image features in low-resolution image classification tasks are solved, and higher classification accuracy and local fine-grained feature characterization capabilities are achieved.

CN119600338BActive Publication Date: 2025-06-17DALIAN UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411639987.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-06-17
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

In practical application scenarios, low-resolution fine-grained image classification task is difficult to effectively identify the existing technology due to feature distribution differences and sparse image features.

Method used

Using a method based on component attention correction and detail enhancement, a variety of discriminant component features are mined through the component attention correction module, and the high-level semantic features of the low-resolution image block are mapped to the underlying texture details of the high-resolution image block through dynamic reconstruction losses.

Benefits of technology

The classification accuracy of low-resolution fine-grained images is improved, and the local fine-grained feature characterization ability is enhanced, and the negative impact of resolution reduction and sparse image features on classification performance is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600338B_ABST
    Figure CN119600338B_ABST
Patent Text Reader

Abstract

A low-resolution fine-grained image classification method based on component attention correction and detail enhancement, belonging to the field of low-resolution fine-grained image recognition. This method first mines diverse and discriminative local representations from low-resolution images; uses learnable class embedding vectors to encode the global fine-grained feature representations of low-resolution image patch embeddings; designs an attention correction loss to adjust the attention activation intensity of each component query vector in the intermediate encoding layer to each image patch embedding. By forcing the high-level semantic features of low-resolution image patches to be mapped to the underlying texture details of the corresponding high-resolution image patches, the extracted local fine-grained feature representations and global fine-grained feature representations are input into a feature fusion layer for information interaction and fusion, and the fused feature representations are input into a classification layer for class prediction. The present invention can accurately perceive and associate local fine-grained features in target components and improve the classification accuracy of low-resolution fine-grained images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of low-resolution fine-grained image recognition, and relates to a low-resolution fine-grained image classification method based on component attention correction and detail enhancement. Background Art

[0002] Fine-grained image classification aims to distinguish different sub-categories under the same coarse-grained category. These sub-categories usually have highly similar appearance features and very tiny fine-grained differences. Traditional fine-grained image classification methods mainly focus on locating key local regions from high-resolution input images for extracting rich fine-grained features of the target. Thanks to the powerful feature encoding ability of Transformer, some research works design an image patch sampling mechanism to select image patches containing key fine-grained features and filter out discriminative-less image patch features. These methods have been experimentally verified on multiple high-resolution fine-grained datasets and achieved high recognition accuracy, greatly promoting the development of the fine-grained image recognition field.

[0003] However, in practical application scenarios such as remote monitoring and automatic retail, due to the over-long shooting distance resulting in a small overall area of the target in the image, it is not easy to collect high-resolution image data. Therefore, it is inevitable to face the problem of low-resolution fine-grained image classification. Due to the large difference in feature distribution between high-resolution data and low-resolution data, when a fine-grained recognition model trained on high-resolution images is directly applied to low-resolution images, unreliable prediction results often occur. Although by reducing the resolution of training images, the problem of inconsistent feature distribution between high-resolution training images and low-resolution test images can be alleviated. However, it is still difficult to overcome the negative impact of resolution reduction and sparse image features on classification performance only by training the model with low-resolution data. To solve this problem, an intuitive solution is to use image super-resolution methods to reconstruct the texture details of low-resolution images. Some methods integrate the super-resolution model into the recognition model to construct an end-to-end network structure, and the downstream recognition task guides the restoration task of fine-grained details.

[0004] In addition to the methods based on image super-resolution, another paradigm for solving the low-resolution image classification task is to use the features of high-resolution data to guide and constrain the feature learning process of low-resolution data. Some methods design various knowledge distillation strategies using the feature information at different levels in high-resolution data and transfer them to the feature learning process of low-resolution data. In addition, the methods based on the encoder-decoder structure constrain the encoder to extract resolution-invariant features from low-resolution images by forcing the high-level semantic features of low-resolution images to reconstruct the texture details of the corresponding high-resolution images.

[0005] Although the above methods improve the ability to distinguish low-resolution fine-grained targets to a certain extent, most of them tend to identify low-resolution image targets by learning global feature representations consistent with high-resolution images, and still lack the ability to learn local discriminative features of low-resolution images. Summary of the Invention

[0006] In view of the above problems existing in the prior art, the present invention provides a low-resolution fine-grained image classification method based on component attention correction and detail enhancement, which mines diverse and discriminative component features through a component attention correction module; designs an attention correction loss to adjust the attention activation intensity of each component query vector in the intermediate encoding layer to the feature of each image patch, reduces the attention of the component query vector to the non-discriminative region in the low-resolution image, and enhances the sensitivity of the component query vector to fine-grained features; by applying a dynamic reconstruction loss to the low-resolution image patch, forcing the high-level semantic features of the low-resolution image patch to be mapped to the underlying texture details of the corresponding high-resolution image patch, further enhancing the local fine-grained feature representation ability, thereby improving the classification accuracy of low-resolution fine-grained images.

[0007] The technical solution adopted to achieve the above purpose is as follows:

[0008] A low-resolution fine-grained image classification method based on component attention correction and detail enhancement, the low-resolution fine-grained image classification method includes the following steps:

[0009] Step 1: Divide the low-resolution image and the high-resolution image into non-overlapping image patch sequences respectively, and linearly transform and map these two groups of image patch sequences into high-dimensional image patch features; specifically:

[0010] For a given low-resolution image I LR , divide it into a sequence of N image patches of size 16×16, denoted as Then input it into the linear transformation layer to map it into D-dimensional low-resolution image patch features, denoted as {v 1 , v 2 ,..., v N}, where D is the dimension of the low-resolution image patch features.

[0011] For the high-resolution image I HR during the training stage, perform the above operations in the same way, divide it into a sequence of N image patches of size 16×16, denoted as Then input it into the linear transformation layer to map it into D-dimensional high-resolution image patch features, denoted as {u 1 , u 2 ,..., u N}, where D is the dimension of the high-resolution image patch features.

[0012] Step 2: Construct a component attention correction module, set a group of learnable component query vectors and a learnable class embedding, and use them together with the low-resolution image patch features and high-resolution image patch features generated in Step 1 as the input of the component attention correction module, to perceive and associate the local and global fine-grained features of the targets in the low-resolution image and high-resolution image, and output the encoded low-resolution image component features, the global fine-grained feature representation of the low-resolution image, the high-resolution image component features, and the global fine-grained feature representation of the high-resolution image; specifically:

[0013] The feature extractor of the component attention correction module uses ViT-B_16, which contains L encoding layers, and the parameters of the low-resolution image patch feature extractor and the high-resolution image patch feature extractor are shared. The input of the feature extractor is Q learnable component query vectors a class embedding v cls 、N low-resolution image patch features and N high-resolution image features.

[0014] In the feature extractor, the component query vectors perform feature perception and interaction with the low-resolution image patch features to calculate the attention scores of each component query vector on all low-resolution image patch features (, where represents the attention score of the i-th component query vector on the j-th image patch in the l-th encoding layer of the feature extractor, i ∈ {1,..., Q}, j ∈ {1,..., N}, l ∈ {1,..., L}); use these attention scores to re-weight and sum the low-resolution image patch features respectively, aggregate the fine-grained features of different parts of the image target, and output Q low-resolution image component features and N low-resolution image patch features {v (L),1 , v (L),2 ,..., v (L),N}. In the feature extractor, the class embedding calculates its attention map E with all low-resolution image patch features l , aggregates all low-resolution image patch features, and extracts and outputs the global fine-grained feature representation of the low-resolution image

[0015] Meanwhile, the component query vectors perform feature perception and interaction with all high-resolution image patch features, encode the fine-grained features of different parts of the high-resolution image target, and output Q high-resolution image component features and N high-resolution image patch features {u (L),1 , u (L),2 ,..., u (L),NIn the feature extractor, the category embedding extracts and outputs the global fine-grained feature representation of the high-resolution image by aggregating all high-resolution image patch features.

[0016] Step 3: Use the classification loss function to classify the low-resolution image component features output in step 2 and high-resolution image component features The discriminability is supervised. The orthogonal loss function is used to supervise the query vectors of different components in step 2 to focus on different regions of the target. The contrast loss function is used to further increase the distance between the global fine-grained feature representations of low-resolution images of different categories in step 2, and to reduce the distance between the global fine-grained feature representations of low-resolution images of the same category; specifically:

[0017] The classification loss function The calculation formula is as follows:

[0018]

[0019] Where C[·] represents the classification head; y i Indicates that the image belongs to the i-th category in the dataset, and C is the total number of categories; represents the low-resolution image component features output from step 2, represents the high-resolution image component feature output by step 2; j represents the index of the component query vector, and Q represents the number of learnable component query vectors.

[0020] The orthogonal loss function L o The calculation formula is as follows:

[0021]

[0022] in, and They represent the i-th and j-th low-resolution image component query vectors output by the last encoding layer respectively. and They represent the i-th and j-th high-resolution image component query vectors output by the last encoding layer, respectively. ||·||2 represents the L2 norm.

[0023] The contrast loss function The calculation is as follows:

[0024]

[0025] in, Represents the global fine-grained feature representation of a low-resolution image with class label yi; Indicates that the category label is y jGlobal fine-grained feature representation of the low-resolution image; Sim(·) represents a function for calculating the similarity of the features of two samples; α is used to control the distance between the features of different categories.

[0026] Step 4: According to the attention scores of each component query vector on the image patch features in Step 2 Extract the indices and attention scores of the top K low-resolution image patch features corresponding to each component query vector (here, the attention scores of the top K image patch features corresponding to the i-th component query vector are denoted as ), and construct an attention correction loss function to correct the attention activation scores of each component query vector on the low-resolution image patch features;

[0027] The attention correction loss function L ac is calculated as follows:

[0028]

[0029] s.t. A[d] = {A j [d] | A i [d] ≥ A j [d], d ∈ D i,j}},

[0030] where A[d] represents the attention score to be optimized of the component query vector on the d-th image patch, A i [d] represents the attention score of the i-th component query vector on the d-th image patch; A j [d] represents the attention score of the j-th component query vector on the d-th image patch; d represents the image patch index; D i,j represents the index set of the repeated image patches among the top K image patches selected by the i-th and j-th component query vectors, and the calculation method is as follows:

[0031] D i,j = H(P i , P j ),

[0032] where, represents the set of the top K image patch features attended to by the i-th component query vector, and H(P i , P j ) is used to calculate the indices of the repeated image patch features attended to by the i-th and j-th component query vectors.

[0033] During the training process, the number of repeated image patch features searched by different component query vectors will gradually approach 0. Therefore, the value of |D i,j | varies between [0, K].

[0034] Step 5: Superimpose the attention scores of each part query vector in Step 4 on the top K low-resolution image patch features to obtain the key region attention map of all part query vectors on the image patch features. Then, extract and fuse the attention maps of all encoding layers of the feature extractor in the part attention correction module described in Step 2, and fuse the fused attention map with the key region attention map to obtain the dynamic attention map; specifically:

[0035] The dynamic attention map A mod is calculated as follows:

[0036]

[0037] where M i represents the binary mask of the i-th part query vector. The mask values at the corresponding positions of the top K image patch features with the highest attention scores on the image patch features are 1, and the values at the remaining positions are 0. β is a threshold used to further adjust the attention scores of each part query vector on the image patch features. E i represents the attention map obtained by multiplying the attention maps E c output from all encoding layers in Step 2. ⊙ represents the Hadamard product. l The first term on the right side of the above formula is used to extract and adjust the attention scores of the image patch features noticed by all part query vectors, and the second term is used to extract the activation scores of the discriminative regions noticed in the multiplied attention map.

[0038] The first term on the right side of the above formula is used to extract and adjust the attention scores of the image patch features noticed by all part query vectors, and the second term is used to extract the activation scores of the discriminative regions noticed in the multiplied attention map.

[0039] Step 6: Introduce a detail enhancement module. The input of this detail enhancement module is the low-resolution image patch features, which are mapped to the texture details of the high-resolution image patches, and the dynamic attention map A mod obtained in Step 5 is used to construct the dynamic reconstruction loss function of the low-resolution image, which is used to constrain the fine-grained feature representation of the low-resolution image patches to approximate that of the high-resolution image patches during the training process; specifically:

[0040] The detail enhancement module is stacked by L' decoding layers, and its input is the N low-resolution image patch features {v (L),1 , v (L),2 ,..., v (L),N} output in Step 2, and the output is the reconstructed high-resolution image patch

[0041] The dynamic reconstruction loss function of the low-resolution image is calculated as follows:

[0042]

[0043] Among them, represents the j-th high-resolution image patch; represents the j-th high-resolution image patch reconstructed from the features of the j-th low-resolution image patch.

[0044] In order to further improve the ability of the detail enhancement module to reconstruct the texture details of key regions, the present invention uses the high-resolution image patch features {u (L),1, u (L),2 ,..., u (L),N} output in step 2 as the input of the detail enhancement module to reconstruct the corresponding high-resolution image patch The dynamic reconstruction loss function of the high-resolution image is calculated as follows:

[0045]

[0046] Among them, represents the high-resolution image patch reconstructed from the high-level semantic features of the high-resolution image patch. The total dynamic reconstruction loss is calculated as follows:

[0047]

[0048] It should be noted that in the test phase, the detail enhancement module is removed, and the low-resolution image does not need to pass through this module.

[0049] Step 7: Introduce a feature fusion layer in the component attention correction module, and reset a learnable class embedding v' cls , concatenate it with the Q low-resolution image component features output in step 2, and input them into the feature fusion layer for information interaction and fusion to output the fused feature representation; introduce a classification layer for predicting the category of the low-resolution image, the input of the classifier is the fused feature representation, and the output is the predicted category. Specifically:

[0050] Take the Q part features extracted in step 2 and a newly set learnable class embedding v' cls as the input of the feature fusion layer, perform information interaction and fusion in the feature fusion layer, and output the fused feature representation Input it into the classification layer and use the classification supervision loss function for category prediction.

[0051] The classification supervision loss function L c is calculated as follows:

[0052]

[0053] Among them, C represents the number of categories in the dataset, yi represents the class label of the image; C[·] represents the classification layer, represents the probability that the classification layer predicts the input as the i-th class.

[0054] Step 8: Input the low-resolution fine-grained test image into the part attention correction module trained in Steps 1 to 7, and output the predicted class of the low-resolution fine-grained image.

[0055] It should be noted that in the test stage, the model only needs the low-resolution image as the input. After prediction by the part attention correction module, the high-resolution image and the detail enhancement module are only used in the training stage.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] (1) The present invention designs a part attention correction module with attention correction loss. By correcting the attention activation scores of each part query vector in the intermediate encoding layer for the features of each image patch, it accurately perceives and correlates the local fine-grained features in the target part.

[0058] (2) On the other hand, the present invention introduces a detail enhancement module with dynamic reconstruction loss into the detail enhancement module, maps the high-level semantic features of the low-resolution image patch to the corresponding high-resolution texture details, and makes the fine-grained feature representations of the low-resolution image patch and the high-resolution image patch consistent.

[0059] (3) The present invention can fully exploit the local fine-grained visual differences of the target in the low-resolution image and improve the classification accuracy of the low-resolution fine-grained image. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 is a schematic diagram of a low-resolution fine-grained image classification method based on part attention correction and detail enhancement of the present invention.

[0061] Figure 2 is a flowchart of a low-resolution fine-grained image classification method based on part attention correction and detail enhancement of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0062] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0063] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0064] As Figure 1 shown, the present invention provides a low-resolution fine-grained image classification method based on component attention correction and detail enhancement, comprising the following steps:

[0065] Step 1: Divide the low-resolution image and the high-resolution image into non-overlapping sequences of image patches respectively, and linearly transform and map these two sets of image patch sequences into high-dimensional image patch features; specifically:

[0066] For a given low-resolution image I LR , divide it into a sequence of N image patches of size 16×16, and denote it as Then input it into the linear transformation layer to map it into D-dimensional low-resolution image patch features, denoted as {v 1 , v 2 ,..., v N}, where D is the dimension of the low-resolution image patch features.

[0067] For the high-resolution image I HR in the training stage, perform the above operations in the same way, divide it into a sequence of N image patches of size 16×16, and denote it as Then input it into the linear transformation layer to map it into D-dimensional high-resolution image patch features, denoted as {u 1 , u 2 ,..., u N}, where D is the dimension of the high-resolution image patch features.

[0068] Step 2: Construct a component attention correction module, set a group of learnable component query vectors and a learnable class embedding, and use them together with the low-resolution image patch features and high-resolution image patch features generated in Step 1 as the input of the component attention correction module, for perceiving and associating the local and global fine-grained features of the targets in the low-resolution image and high-resolution image, and outputting the encoded low-resolution image component features, the global fine-grained feature representation of the low-resolution image, the high-resolution image component features, and the global fine-grained feature representation of the high-resolution image; specifically:

[0069] The feature extractor of the component attention correction module uses ViT-B_16, which contains L encoding layers, and the parameters of the low-resolution image patch feature extractor and the high-resolution image patch feature extractor are shared. As Figure 1 (a) shows, the input of the feature extractor is Q learnable component query vectors a class embedding v cls , N low-resolution image patch features, and N high-resolution image patch features.

[0070] In the feature extractor, the component query vectors perform feature perception and interaction with the low-resolution image patch features to calculate the attention scores of each component query vector on all low-resolution image patch features (, where represents the attention score of the i-th component query vector on the j-th image patch in the l-th encoding layer of the feature extractor, i ∈ {1,..., Q}, j ∈ {1,..., N}, l ∈ {1,..., L}); use these attention scores to re-weight and sum the low-resolution image patch features respectively, aggregate the fine-grained features of different parts of the image target, and output Q low-resolution image component features and N low-resolution image patch features {v (L),1 , v (L),2 ,..., v (L),N}}. In the feature extractor, the class embedding calculates its attention map E l with all low-resolution image patch features, aggregates all low-resolution image patch features, and extracts and outputs the global fine-grained feature representation of the low-resolution image

[0071] Meanwhile, the component query vectors perform feature perception and interaction with all high-resolution image patch features, encode the fine-grained features of different parts of the image target, and output Q high-resolution image component features and N high-resolution image patch features {u (L),1 , u (L),2 ,..., u (L),NIn the feature extractor, the category embedding extracts and outputs the global fine-grained feature representation of the high-resolution image by aggregating all high-resolution image patch features.

[0072] Step 3: Use the classification loss function to classify the low-resolution image component features output in step 2 and high-resolution image component features The discriminability is supervised. The orthogonal loss function is used to supervise the query vectors of different components in step 2 to focus on different regions of the target. The contrast loss function is used to further increase the distance between the global fine-grained feature representations of low-resolution images of different categories in step 2, and to reduce the distance between the global fine-grained feature representations of low-resolution images of the same category; specifically:

[0073] The classification loss function The calculation formula is as follows:

[0074]

[0075] Where C[·] represents the classification head; y i Indicates that the image belongs to the i-th category in the dataset, and C is the total number of categories; represents the low-resolution image component features output from step 2, represents the high-resolution image component feature output by step 2; j represents the index of the component query vector, and Q represents the number of learnable component query vectors.

[0076] The orthogonal loss function L o The calculation formula is as follows:

[0077]

[0078] in, and They represent the i-th and j-th low-resolution image component query vectors output by the last encoding layer respectively. and They represent the i-th and j-th high-resolution image component query vectors output by the last encoding layer, respectively. ||·||2 represents the L2 norm.

[0079] The contrast loss function The calculation is as follows:

[0080]

[0081] in, Represents the global fine-grained feature representation of a low-resolution image with class label yi; Indicates that the category label is y jGlobal fine-grained feature representation of the low-resolution image; Sim(·) represents a function for calculating the similarity between the features of two samples; α is used to control the distance between the features of different categories.

[0082] Step 4: According to the attention scores of each component query vector in the image patch features in Step 2 Extract the indices and attention scores of the top K low-resolution image patch features corresponding to each component query vector (here, the attention scores of the top K image patch features corresponding to the i-th component query vector are denoted as ), and construct an attention correction loss function to correct the attention activation scores of each component query vector for the low-resolution image patch features;

[0083] The attention correction loss function L ac is calculated as follows:

[0084]

[0085] s.t. A[d] = {A j [d] | A i [d] ≥ A j [d], d ∈ D i,j}},

[0086] where A[d] represents the attention score to be optimized by the component query vector on the d-th image patch, A i [d] represents the attention score of the i-th component query vector on the d-th image patch; A j [d] represents the attention score of the j-th component query vector on the d-th image patch; d represents the image patch index; D i,j represents the index set of the overlapping image patches among the top K image patches selected by the i-th and j-th component query vectors, and the calculation method is as follows:

[0087] D i,j = H(P i , P j ),

[0088] where, represents the set of the top K image patch features attended to by the i-th component query vector, and H(P i , P j ) is used to calculate the indices of the overlapping image patch features attended to by the i-th and j-th component query vectors. During the training process, the number of overlapping image patch features searched by different component query vectors will gradually approach 0. Therefore, the value of |D i,j | varies between [0, K].

[0089] Step 5: Superimpose the attention scores of each part query vector in Step 4 on the top K low-resolution image patch features to obtain the key region attention map of all part query vectors on the image patch features. Then, extract and fuse the attention maps of all encoding layers of the feature extractor in the part attention correction module described in Step 2, and fuse the fused attention map with the key region attention map to obtain the dynamic attention map; specifically:

[0090] As Figure 1 (b) shows, the dynamic attention map A mod is calculated as follows:

[0091]

[0092] where M i represents the binary mask of the i-th part query vector. The mask values at the corresponding positions of the top K image patch features with the highest attention scores on the image patch features are 1, and the values at the remaining positions are 0. β is a threshold used to further adjust the attention scores of each part query vector on the image patch features. E i represents the attention map obtained by multiplying the attention maps E c output from all encoding layers in Step 2. ⊙ represents the Hadamard product. l The first term on the right side of the above formula is used to extract and adjust the attention scores of the image patch features noticed by all part query vectors, and the second term is used to extract the activation scores of the discriminative regions noticed in the multiplied attention map.

[0093] The first term on the right side of the above formula is used to extract and adjust the attention scores of the image patch features noticed by all part query vectors, and the second term is used to extract the activation scores of the discriminative regions noticed in the multiplied attention map.

[0094] Step 6: Introduce a detail enhancement module. The input of this detail enhancement module is the low-resolution image patch features, which are mapped to the texture details of high-resolution image patches, and the dynamic attention map A mod obtained in Step 5 is used to construct a dynamic reconstruction loss function for the low-resolution image, which is used to constrain the fine-grained feature representation of the low-resolution image patch to approximate that of the high-resolution image patch during the training process; specifically:

[0095] The detail enhancement module is stacked by L' decoding layers. Its input is the N low-resolution image patch features {v (L),1 , v (L),2 ,..., v (L),N} output in Step 2, and the output is the reconstructed high-resolution image patch

[0096] The dynamic reconstruction loss function of the low-resolution image is calculated as follows:

[0097]

[0098] Among them, represents the j-th high-resolution image patch; represents the j-th high-resolution image patch reconstructed from the features of the j-th low-resolution image patch.

[0099] In order to further improve the ability of the detail enhancement module to reconstruct the texture details of key regions, the present invention uses the high-resolution image patch features {u (L),1 , u (L ), ,2 ,..., u (L),N} output in step 2 as the input of the detail enhancement module to reconstruct the corresponding high-resolution image patch The dynamic reconstruction loss function of the high-resolution image is calculated as follows:

[0100]

[0101] Among them, represents the high-resolution image patch reconstructed from the high-level semantic features of the high-resolution image patch. The total dynamic reconstruction loss is calculated as follows:

[0102]

[0103] It should be noted that in the test phase, the detail enhancement module is removed, and the low-resolution image does not need to pass through this module.

[0104] Step 7: Introduce a feature fusion layer in the component attention correction module, and reset a learnable class embedding v' cls , concatenate it with the Q low-resolution image component features output in step 2, and input them into the feature fusion layer for information interaction and fusion to output the fused feature representation; introduce a classification layer for predicting the category of the low-resolution image, the input of the classifier is the fused feature representation, and the output is the predicted category. Specifically:

[0105] Take the Q part features extracted in step 2 and the newly set learnable class embedding v' cls as the input of the feature fusion layer, perform information interaction and fusion in the feature fusion layer, and output the fused feature representation Input it into the classification layer, and use the classification supervision loss function for category prediction.

[0106] The classification supervision loss function L c is calculated as follows:

[0107]

[0108] Among them, C represents the number of categories in the dataset, and y i represents the category label of the image; C[·] represents the classification layer, indicating the probability that the classification layer predicts the input as the i-th category.

[0109] Step 8: Input the low-resolution fine-grained test image into the part attention correction module trained in Steps 1 to 7, and output the predicted category of the low-resolution fine-grained image.

[0110] It should be noted that in the test phase, the model only needs a low-resolution image as input. After prediction by the part attention correction module, the high-resolution image and the detail enhancement module are only used in the training phase.

[0111] According to the above steps, the present invention trains and tests the proposed low-resolution fine-grained image classification method based on part attention correction and detail enhancement. The low-resolution fine-grained dataset used is generated by bilinear downsampling of the high-resolution fine-grained benchmark datasets CUB-200-2011, Nabirds, Standford-Dogs, Standford-Cars, and FGVC-Aircraft, and is labeled as CUB-S, Nabird-S, Dog-S, Car-S, and Aircraft-S respectively. The sizes of the generated low-resolution images and high-resolution images are 28×28 and 224×224 respectively. In the test phase, the present method only uses the low-resolution image as the model input.

[0112] The evaluation metric adopted by this method is accuracy, that is, the ratio of the number of samples whose categories are correctly predicted in all test data to the number of samples in all test data. The higher the accuracy, the better the recognition result. The following experimental results show that the method proposed by the present invention has better classification accuracy. The present invention corrects the attention activation scores of each image patch feature by the part query vector, guides the part query vector to discover more discriminative fine-grained clues, and enhances the local fine-grained feature representation ability by introducing the dynamic reconstruction loss function constraint, thereby improving the classification accuracy of low-resolution fine-grained images.

[0113] Table 1 Comparison results between the present invention and existing traditional methods

[0114] Method CUB-S Nabird-S Dog-S Cars-S Aircraft-S Baseline network 76.26 68.13 74.27 69.08 73.17 TransFG 76.75 68.91 73.62 70.50 68.91 SIM-Trans 76.61 68.92 72.69 71.69 69.92 The present invention 80.47 72.54 76.70 75.15 74.52

[0115] The above-described embodiments only represent the implementation manners of the present invention, but should not be construed as limiting the scope of the present invention patent. It should be noted that for those skilled in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A low-resolution fine-grained image classification method based on component attention correction and detail enhancement, characterized in that: The low-resolution fine-grained image classification method comprises the following steps: Step 1: Divide the low-resolution image and the high-resolution image into non-overlapping image block sequences, and perform linear transformation on the two sets of image block sequences and map them into high-dimensional image block features; Step 2: Build a component attention correction module, set a set of learnable component query vectors and a learnable category embedding, and use them together with the low-resolution image block features and high-resolution image block features generated in step 1 as the input of the component attention correction module, which is used to perceive and associate the local and global fine-grained features of the targets in the low-resolution image and the high-resolution image, and output the encoded low-resolution image component features, the global fine-grained feature representation of the low-resolution image, the high-resolution image component features, and the global fine-grained feature representation of the high-resolution image; Step 3: Use the classification loss function to classify the low-resolution image component features output in step 2 and high-resolution image component features The discriminability of the image is supervised; the supervision of the orthogonal loss function is used to make the query vectors of different components in step 2 focus on different areas of the target; the contrast loss function is used to expand the distance of the global fine-grained feature representation of low-resolution images of different categories in step 2, and to reduce the distance of the global fine-grained feature representation of low-resolution images of the same category; Step 4: Based on the attention score of each component query vector on the image patch feature in step 2 Extract the index and attention score of the first K low-resolution image block features corresponding to each component query vector, and record the attention score of the first K image block features corresponding to the i-th component query vector as Construct an attention correction loss function to correct the attention activation score of each component query vector to the low-resolution image patch feature; Step 5: Superimpose the attention scores of each component query vector on the first K low-resolution image block features in step 4 to obtain the key area attention map of all component query vectors on the image block features; then extract the attention maps of all encoding layers of the feature extractor in the component attention correction module in step 2 and fuse them, and fuse the fused attention map with the key area attention map to obtain the dynamic attention map A. mod ; Step 6: Introduce a detail enhancement module, which takes the low-resolution image block features as input, maps them to the texture details of the high-resolution image blocks, and uses the dynamic attention map A obtained in step 5 mod A dynamic reconstruction loss function of a low-resolution image is constructed to constrain the fine-grained feature representation of a low-resolution image block to approach the fine-grained feature representation of a high-resolution image block during training; Step 7: Introduce a feature fusion layer in the component attention correction module and reset a learnable category embedding v′ cls , and combine it with the Q low-resolution image component features output from step 2 The images are connected in series and input into the feature fusion layer for information interaction and fusion, and the fused feature representation is output; a classification layer is introduced to predict the category of the low-resolution image. The input of the classifier is the fused feature representation, and the output is the predicted category; Step 8: Input the low-resolution fine-grained test image into the component attention correction module trained in steps 1 to 7, and output the predicted category of the low-resolution fine-grained image.

2. According to claim 1, a low-resolution fine-grained image classification method based on component attention correction and detail enhancement is characterized in that: The step 1 is specifically as follows: For a given low-resolution image I LR , divide it into N image block sequences of the same size, which are recorded as It is then input into the linear transformation layer and mapped into a D-dimensional low-resolution image block feature, denoted as {v 1 ,v 2 ,...,v N }, where D is the dimension of the low-resolution image patch feature; For the high-resolution image I in the training phase HR , perform the above operation and divide it into N image block sequences of the same size, which are recorded as It is then input into the linear transformation layer and mapped into a D-dimensional high-resolution image block feature, denoted as {u 1 ,u 2 ,...,u N }, where D is the dimension of the high-resolution image patch features.

3. The low-resolution fine-grained image classification method based on component attention correction and detail enhancement according to claim 2 is characterized in that: The step 2 is specifically as follows: The feature extractor of the component attention correction module adopts ViT-B_16, which includes L encoding layers, and the parameters of the low-resolution image block feature and the high-resolution image block feature extractor are shared; the input of the feature extractor is Q learnable component query vectors A category embedding v cls , N low-resolution image block features and N high-resolution image features; In the feature extractor, the component query vector performs feature perception and interaction with the low-resolution image patch features, and calculates the attention score of each component query vector on all low-resolution image patch features. in represents the attention score of the i-th component query vector on the j-th image block in the l-th encoding layer of the feature extractor, i∈{1,…,Q}, j∈{1,…,N}, l∈{1,…,L}); the attention score is used to re-weight and sum the low-resolution image block features, aggregate the fine-grained features of different parts of the image target, and output Q low-resolution image component features and N low-resolution image patch features {v (L),1 ,v (L),2 ,...,v (L),N }; In the feature extractor, the category embedding is obtained by computing the attention map E between it and all low-resolution image patch features l , aggregate all low-resolution image patch features, extract and output the global fine-grained feature representation of the low-resolution image At the same time, the component query vector encodes the fine-grained features of different parts of the high-resolution image target by performing feature perception and interaction with all high-resolution image block features, and outputs Q high-resolution image component features and N high-resolution image patch features {u (L),1 ,u (L),2 ,...,u (L),N In the feature extractor, the category embedding extracts and outputs the global fine-grained feature representation of the high-resolution image by aggregating all high-resolution image patch features.

4. The low-resolution fine-grained image classification method based on component attention correction and detail enhancement according to claim 3 is characterized in that: The step 3 is specifically as follows: The classification loss function The calculation formula is as follows: Where C[·] represents the classification head; y i Indicates that the image belongs to the i-th category in the dataset, and C is the total number of categories; represents the low-resolution image component features output from step 2, represents the high-resolution image component feature outputted in step 2; j represents the index of the component query vector, and Q represents the number of learnable component query vectors; The orthogonal loss function L o The calculation formula is as follows: in, and Respectively represent the i-th and j-th low-resolution image component query vectors output by the last encoding layer; and They represent the i-th and j-th high-resolution image component query vectors output by the last encoding layer respectively; ||·||2 represents the L2 norm; The contrast loss function The calculation is as follows: in, Indicates that the category label is y i Global fine-grained feature representation of low-resolution images; Indicates that the category label is y j The global fine-grained feature representation of the low-resolution image; Sim(·) represents the function of calculating the similarity of two sample features; α is used to control the distance between features of different categories.

5. The low-resolution fine-grained image classification method based on component attention correction and detail enhancement according to claim 4 is characterized in that: The step 4 is specifically as follows: The attention correction loss function L ac The calculation is as follows: s.t.A[d]={A j [d]∣A i [d]≥A j [d],d∈D i,j }, Among them, A[d] represents the attention score that needs to be optimized on the d-th image block of the part query vector, A i [d] represents the attention score of the i-th component query vector on the d-th image patch; A j [d] represents the attention score of the j-th component query vector on the d-th image patch; d represents the image patch index; D i,j represents the index set of repeated image blocks in the first K image blocks selected by the i-th and j-th component query vectors, and is calculated as follows: D i,j =H(P i ,P j ), in, represents the set of the first K image block features that the i-th component query vector focuses on, H(P i ,P j ) is used to calculate the index of the i-th and j-th component query vectors focusing on the repeated image block features; During the training process, the number of repeated image block features searched by different component query vectors will gradually approach 0, so |D i,j The value of | varies between [0,K].

6. The low-resolution fine-grained image classification method based on component attention correction and detail enhancement according to claim 5, characterized in that: The step 5 is specifically as follows: The dynamic attention map A mod The calculation formula is as follows: Among them, M i Represents the binary mask of the i-th component query vector, and the K image patch features with the highest attention scores on the image patch features are in M i The mask value of the corresponding position in is 1, and the values ​​of the other positions are 0. β is a threshold used to further adjust the attention score of each component query vector on the image block feature. c represents the attention map E in all encoding layers output by step 2 l The attention map obtained by cumulative multiplication; ⊙ represents the Hadamard product.

7. The low-resolution fine-grained image classification method based on component attention correction and detail enhancement according to claim 6 is characterized in that: The step 6 is specifically as follows: The detail enhancement module is composed of L' decoding layers stacked together, and its input is the N low-resolution image block features {v (L),1 ,v (L),2 ,...,v (L),N }, the output is the reconstructed high-resolution image patch The dynamic reconstruction loss function of the low-resolution image The calculation is as follows: in, represents the jth high-resolution image block; represents the jth high-resolution image block reconstructed from the jth low-resolution image block features; Using the high-resolution image block features {u (L),1 ,u (L),2 ,...,u (L),N }As the input of the detail enhancement module, the corresponding high-resolution image block is reconstructed The dynamic reconstruction loss function of high-resolution images is calculated as follows: in, represents the high-resolution image patch reconstructed by the high-level semantic features of the high-resolution image patch; the total loss of dynamic reconstruction is calculated as follows:

8. The low-resolution fine-grained image classification method based on component attention correction and detail enhancement according to claim 6, characterized in that: The step 7 is specifically as follows: The Q part features extracted in step 2 A learnable category embedding v′ with the new setting cls As the input of the feature fusion layer, information interaction and fusion are performed in the feature fusion layer, and the fused feature representation is output Input it into the classification layer and use the classification supervision loss function to make category prediction; The classification supervision loss function L c The calculation is as follows: Among them, C represents the number of categories in the data set, y i represents the category label of the image; C[·] represents the classification layer, Represents the probability that the classification layer predicts the input as the i-th class.

Citation Information

Patent Citations

  • Pedestrian re-identification method and device based on mixed attention decoupling re-identification network

    CN116524533A

  • Low-resolution image fine-grained recognition method based on easy-to-confusion feature mining and decoupling

    CN118736273A