An interpretable fine-grained image recognition method, device, equipment and medium
Through the dual-branch structure framework and feature fusion method, the problem of inaccurate black box characteristics and feature extraction in fine-grained image classification of deep learning models is solved, and the interpretability and accuracy improvement in high-risk fields is achieved.
Patent Information
- Application Number
- CN202510147874.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-02-11
AI Technical Summary
Deep learning models have black box characteristics in fine-grained image classification, which is difficult to apply in high-risk fields such as medical decision-making and autonomous driving. The feature extraction of existing prototype learning methods is not accurate enough and it is easy to lose detailed features.
The dual-branch structure framework is adopted, combining feature fusion and prototype learning, and the global and local features of the image are captured through the Transformer architecture, and important features are selected using the self-attention mechanism and feature selection fusion module, and weighted calculations are performed on the full connection layer to generate image recognition results.
The classification accuracy and interpretability of the model are improved, making it more suitable for application in high-risk areas, and enhancing the ability to capture local details of the image.
Smart Images

Figure CN119625436B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fine-grained image recognition, and particularly relates to an interpretable fine-grained image recognition method, device, equipment and medium. Background Art
[0002] In the field of artificial intelligence, fine-grained image recognition is a challenging task, which requires algorithms to be able to identify different sub-categories within the same large category, and these sub-categories are extremely similar in visual features. With the development of deep learning technology, especially the application of convolutional neural networks (CNNs) and vision transformers (ViTs), the accuracy of fine-grained image classification has been significantly improved. However, the black-box nature of deep learning models limits their application in high-risk fields, such as medical decision-making, autonomous driving, and toxic biological identification, because in these fields, the decision-making process of the models needs to have a high degree of interpretability and credibility.
[0003] Among them, the goal of the field of computer vision is to simulate and expand the functions of the human visual system, enabling computers to extract useful information from images and videos. And the fine-grained image classification task is a core problem in computer vision, which requires algorithms to be able to identify sub-categories with subtle differences within the same large category. Although deep learning models have made progress in this field, due to the opacity of their internal information processing process, it is difficult for humans to understand and interpret the decision-making process of the models, which to a certain extent limits their application in high-risk fields.
[0004] To solve this problem, pre-explainable algorithms and post-explainable algorithms have been proposed on the market currently. Pre-explainable algorithms are designed to make the model structure itself interpretable, or the inference process of the model is understandable to humans, such as linear / logistic regression models and decision tree models. Post-explainable algorithms, on the other hand, are for trained models, and design explainable algorithms to explain the behavior logic and decision basis of the models. However, these methods either sacrifice the accuracy of the models or the explanations provided are not completely credible.
[0005] The method based on prototype learning is a research direction proposed in recent years. It learns the prototype parts representing the typical features of each class for each class, and identifies the class by calculating the similarity between the features of the input picture and the prototype. This method is inspired by the human visual system and can provide a certain degree of interpretability. However, existing methods based on prototype learning usually use convolutional neural networks as the backbone network for feature extraction, which results in the extracted features being prone to losing detailed features, and the prototype learning is not accurate enough, and is prone to the problem of concentrating on the background.
[0006] In view of this, the present application is proposed. Summary of the Invention
[0007] The present invention provides an interpretable fine-grained image recognition method, device, equipment and medium, which can at least partially improve the above problems.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] An interpretable fine-grained image recognition method, which includes:
[0010] Obtain the image x to be recognized, input the image x into the feature extraction layer for feature extraction preprocessing, and generate an output vector;
[0011] Use the class encoding vector of the output vector as the global feature, input it into the global branch of the prototype learning layer for calculation, and obtain the similarity score between the global feature and the global prototype;
[0012] Obtain N local features output by the feature extraction layer, and select K preset foreground information-related features for local prototype learning to obtain the similarity score between the local feature and the local prototype;
[0013] Use a fully connected layer to perform probability calculation processing on the similarity score between the global feature and the global prototype and the similarity score between the local feature and the local prototype, obtain the final classification probability, and generate an image recognition result based on the classification probability.
[0014] The present invention also provides an interpretable fine-grained image recognition device, which includes:
[0015] A feature extraction unit, configured to obtain the image x to be recognized, input the image x into the feature extraction layer for feature extraction preprocessing, and generate an output vector;
[0016] A global branch unit, configured to use the class encoding vector of the output vector as the global feature, input it into the global branch of the prototype learning layer for calculation, and obtain the similarity score between the global feature and the global prototype;
[0017] A local branch unit, configured to obtain N local features output by the feature extraction layer, and select K preset foreground information-related features for local prototype learning to obtain the similarity score between the local feature and the local prototype;
[0018] A fully connected unit, configured to use a fully connected layer to perform probability calculation processing on the similarity score between the global feature and the global prototype and the similarity score between the local feature and the local prototype, obtain the final classification probability, and generate an image recognition result based on the classification probability.
[0019] The present invention also provides an interpretable fine-grained image recognition device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the interpretable fine-grained image recognition method described in any one of the above is implemented.
[0020] The present invention also provides a readable storage medium, which stores a computer program that can be executed by the processor of the device where the storage medium is located to implement the interpretable fine-grained image recognition method described in any one of the above.
[0021] In summary, the interpretable fine-grained image recognition method aims to improve the accuracy and interpretability of deep learning models in fine-grained image classification tasks, especially in high-risk fields such as medical decision-making and autonomous driving. By designing a dual-branch structure framework, this method effectively combines feature fusion and prototype learning to achieve an in-depth understanding of the global and local features of images.
[0022] Specifically, in the feature extraction layer, this method uses the Transformer architecture and the self-attention mechanism to capture the global features and local detail information in the image. In particular, a feature selection and fusion module is introduced, which can select important features from each layer of the Transformer and fuse these features to enhance the model's ability to recognize foreground features while reducing the interference of background features. The prototype learning layer is the core of this method, which includes a global branch and a local branch, responsible for learning the global prototype and local prototype of the image respectively. The global branch focuses on the overall information of the image, while the local branch focuses on the details and local features of the image. This dual-branch structure enables the model to understand the image from different scales, improving the classification accuracy and the interpretability of the model. Finally, the fully connected layer weights and sums the similarity scores obtained from the global branch and the local branch to obtain the final classification probability.
[0023] Generally speaking, the fine-grained image recognition method not only improves the classification accuracy of the model, but also enhances the interpretability of the model, making it more suitable for applications in fields with high requirements for model interpretability. By combining prototype learning and feature fusion, this method provides a new solution for the field of fine-grained image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a schematic flowchart of the interpretable fine-grained image recognition method provided by the first embodiment of the present invention;
[0025] Figure 2 is a schematic overall structure diagram of the interpretable fine-grained image recognition method provided by the first embodiment of the present invention;
[0026] Figure 3It is a schematic structural diagram of the feature selection and fusion module of the interpretable fine-grained image recognition method provided by an embodiment of the present invention;
[0027] Figure 4 It is a schematic diagram of the visualization results of local prototypes before and after using the ARWS module provided by an embodiment of the present invention;
[0028] Figure 5 It is a schematic diagram of the visualization results of the global prototype provided by an embodiment of the present invention;
[0029] Figure 6 It is a reasoning process diagram of the interpretable fine-grained image recognition method provided by an embodiment of the present invention;
[0030] Figure 7 It is a module schematic diagram of the interpretable fine-grained image recognition device provided by the second embodiment of the present invention. Detailed implementation manners
[0031] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] Refer to Figure 1 、 Figure 2 、 Figure 6 As shown, the first embodiment of the present invention discloses an interpretable fine-grained image recognition method, which can be executed by an interpretable fine-grained image recognition device (hereinafter referred to as the recognition device), and particularly, by one or more processors in the recognition device, to implement the following method:
[0033] Please refer to Figure 3 、 Figure 4 , S1, obtain the image x to be recognized, input the image x into the feature extraction layer for feature extraction preprocessing, and generate an output vector;
[0034] Specifically, step S1 includes: obtaining the image x to be recognized, , where the size of the image x is , H is the image height, W is the image width, C is the image channel number, is a matrix;
[0035] Cut the picture x into N patches of the same size, and tile the patches into a sequence image with a length of , , , where the size of the patch is , is the resolution of each image block, is the final number of generated blocks;
[0036] Perform a linear projection process on the sequence of images to generate a feature sequence;
[0037] Add a class encoding vector to the feature sequence as the encoded feature of image x;
[0038] Add a position encoding vector to connect the two-dimensional position information of image x to the one-dimensional feature sequence, obtaining a visual vector , , where is the first image sequence, is the second image sequence, is the Nth image sequence, E is the identity matrix, is the product of the first image sequence and the identity matrix, is the product of the second image sequence and the identity matrix, is the product of the Nth image sequence and the identity matrix;
[0039] Input the visual vector into a module with L stacked Transformer encoders to obtain the output vector of the lth layer Transformer encoder, , , where each Transformer encoder includes a multi-head self-attention and a multi-layer perceptron, LN represents the normalization operation, MSA represents the multi-head self-attention, and MLP represents the multi-layer perceptron, is the output vector of the (l - 1)th layer Transformer encoder;
[0040] Use the class encoding vector of the output vector as the global feature and input it into the global branch, and use the fused feature of the output vector as the local feature and input it into the local branch.
[0041] Preferably, the calculation steps of the results of the multi-head self-attention of the lth layer Transformer encoder are specifically:
[0042] Perform a linear transformation process on the visual sequence input to the lth layer Transformer encoder to obtain three identical visual sequences, which are respectively called the query vector q, the key vector k, and the value vector v;
[0043] Attention weights based on the h-th self-attention head , the similarity between the query vector q and all key vectors k is calculated using the softmax function, and each value vector v is multiplied by its corresponding attention weight to obtain a weighted value vector. The sum of all weighted value vectors is used as the output vector. The calculation formula is: , , , where, is the linear projection matrix of the h-th self-attention head, W is the linear projection matrix, is the length of the h-th self-attention head, T is the transpose of the matrix, and A is the self-attention weight map;
[0044] The results of H self-attention heads are calculated in parallel, and the sum is calculated to obtain the final multi-head self-attention result .
[0045] Preferably, the calculation steps of the fused feature are as follows:
[0046] Obtain the self-attention weight map of the h-th attention head of the l-th layer Transformer encoder , , , where, is the attention value between the i-th visual sequence and the j-th visual sequence, is the attention value between the class encoding vector and all visual sequences;
[0047] Based on the self-attention weight map , calculate the attention scoring standard value matrix of the l-th layer Transformer encoder. Its recursive formula is: , , where, is the identity matrix, and H is the number of attention heads;
[0048] Let be the attention scoring standard value between the i-th visual sequence and the class encoding vector. Sort in ascending order of i from 1 to N, and select the first visual sequences with the highest attention scoring standard values as the important visual sequences of each layer, denoted as: , where, is the i-th important visual sequence in the l-th layer Transformer encoder, is the number of selected important visual sequences;
[0049] Obtain the fused feature , Denote the features selected by the ARWS module from the (L-1)-th layer.
[0050] In this embodiment, the interpretable fine-grained image recognition method mainly consists of three parts: a feature extraction layer, a prototype learning layer, and a fully connected layer. The feature extraction layer uses the ViT structure because ViT can capture complex patterns and long-range dependencies in images through its self-attention mechanism, thus effectively extracting global features and local detail information, providing a rich information basis for subsequent prototype learning and classification decisions. The feature selection and fusion module designed by this method is added to the feature extraction layer to select important features in each layer of ViT and finally fuse the selected features.
[0051] Specifically, in this embodiment, the feature extraction module adopts the Transformer architecture. First, obtain the image x to be recognized, whose size is , that is, the height, width, and number of color channels. The image x is sliced into N patches of the same size, and the size of each patch is , forming a sequence image. Then, perform a linear projection operation on the sequence image, that is, the Patch Embedding operation, to generate a feature sequence . At the same time, add a class encoding vector and a position encoding vector to the feature sequence to form a visual vector, effectively integrating the two-dimensional position information into the one-dimensional feature sequence.
[0052] Specifically, the class encoding vector (class token) is regarded as the encoded feature of the input image by the model, which can represent the learned global feature and is used for subsequent global prototype learning. Different from convolutional neural networks, a position encoding vector (Position Embedding) also needs to be added to add the two-dimensional position information of the image to the one-dimensional sequence, enabling the model to utilize the prior information of the image. Finally, the input image x is encoded into a series of visual tokens.
[0053] The visual vectors are then input into a module that stacks L Transformer encoders, each encoder containing multiple self-attention heads and a multi-layer perceptron. Among them, multi-head self-attention is the core of the Transformer architecture. In each layer of the Transformer encoder, the visual sequence undergoes linear transformation processing to generate query vector q, key vector k, and value vector v. Based on the self-attention weights, the similarity between q and k is calculated using the softmax function, and v is multiplied by the corresponding attention weights, which can emphasize the information that is more important for the current element, obtaining a weighted value vector. The sum of all weighted value vectors is used as the output vector. This process calculates the results of H self-attention heads in parallel and then performs a summation calculation to obtain the final multi-head self-attention result. This process can be regarded as the degree to which each element in the input sequence allocates attention when considering the information of all other elements.
[0054] In the calculation of the fused features, the self-attention weight map A of the h-th attention head of the l-th layer of the Transformer encoder is obtained. Based on this weight map, the attention scoring criterion value matrix is calculated and recursively updated until convergence. Then, the top M visual sequences with the highest attention scoring criterion values are selected as the important visual sequences for each layer. These sequences represent important visual information and are used for subsequent feature fusion. Among them, the feature selection and fusion module is named Attention Rollout WeightSelection (ARWS). A new attention scoring criterion, attention rollout, is introduced in this module. This value is calculated using the attention scores built into ViT and does not introduce additional parameters, enabling better selection of discriminative local features. The ARWS module is used in each layer of the network, which can better utilize the low-level, middle-level, and high-level information of the network.
[0055] The role of the ARWS module is to select the top K important features from each Transformer Layer. The way to select important features is to introduce the attention rollout value. An important structure for feature learning in the Transformer architecture is the multi-head attention mechanism. In this way, the features selected by the ARWS module from the L-1 layer are fused to form global features and local features. The global features are identified by the class encoding vector, while the local features are identified by the fused features. This method not only improves the accuracy of image recognition but also enhances the model's ability to capture local details of the image, especially when dealing with complex and diverse image data.
[0056] Among them, attention weights are generally considered to be corresponding to feature importance. However, recent research has shown that attention weights do not necessarily correspond to the importance of input features. Another important assumption is that the original attention lacks recognizability at higher levels. Therefore, the attention rollout value is introduced in the ARWS module to select the most important features in each layer. For the ViT model, the attention rollout matrix of the l (l ≥ 1) layer is recursively defined as: , . The initial attention rollout value . The class token is usually used for the final classification and aggregates the most information in the image. Therefore, the token with a higher attention score to the class token is considered to be more helpful for classification, which means that the token is more discriminative and more important. Similarly, it can be concluded that the token with a higher attention rollout value to the class token is more important.
[0057] Please refer to Figure 5 , S2, use the class encoding vector of the output vector as the global feature and input it into the global branch of the prototype learning layer for calculation to obtain the similarity score between the global feature and the global prototype;
[0058] Specifically, step S2 includes: predefined learnable global prototypes , and learnable local prototypes ;
[0059] Obtain the global feature output by the feature extraction layer, calculate the similarity score between the global feature and each global prototype. The calculation formula is: , where is the i-th global prototype, is the global feature and the similarity score between the i-th global prototype, is the smoothing term.
[0060] In this embodiment, the prototype learning layer includes two branches, one is the global branch and the other is the local branch, which are respectively used to learn the global prototype and the local prototype. First, learnable global prototypes and A learnable local prototype, which is learned by the model during the training process and can represent the core features of different categories. These prototypes represent the typical visual features of their respective categories and can ultimately be visualized as part of the input image.
[0061] After obtaining the global feature vector from the feature extraction layer, calculate the similarity score between the global feature vector and each predefined global prototype. This calculation is completed through a specific calculation formula to quantify the similarity between the global feature and each global prototype.
[0062] S3, obtain the local features output by the N feature extraction layers, and select K preset foreground information-related features for local prototype learning to obtain the similarity score between the local features and the local prototypes;
[0063] Specifically, step S3 includes: extracting the self-attention weight map , and selecting the top K visual vectors with the highest attention scores;
[0064] Set a binary mask , and modify the attention score of the i-th position encoding vector to the category encoding vector in the formula of the softmax function as: , to retain the visual sequence saved using the foreground preservation mask, where, is the j-th binary mask, is the k-th binary mask;
[0065] Perform calculation processing on the visual sequence saved using the foreground preservation mask and each local prototype, calculate its similarity score, and after max pooling, obtain the similarity score of the local prototype. The calculation formula is: , where, represents the top K visual vectors with the highest attention scores selected, represents a single visual vector selected from the visual vector each time, is the i-th local prototype, is the similarity score between the local feature and the i-th local prototype.
[0066] In this embodiment, among the local features, the features related to the foreground information are the key to classification and can better represent the prototype of a class. Therefore, in the local branch, first, K features related to the foreground information are selected from the N local features output from the feature extraction layer for learning the local prototype. To avoid introducing new parameters, the module endeavors to discover the key local regions from the existing self-attention weight maps of the Transformer, and thus selects the local features beneficial to the classification result for learning the local prototype. The ViT network adds a class token, which aggregates the global features of the image and can be used as the feature representation of the entire image for classification. The importance of the foreground information of the network for network classification is self-evident. Therefore, the local features more important to the class token are selected from the self-attention weights of the class token, and the selected features can basically represent the foreground information of the network.
[0067] Specifically, in this embodiment, this step focuses on obtaining N local features from the local features output by the feature extraction layer, and selecting K features related to the foreground information from them for local prototype learning to obtain the similarity scores between the local features and the local prototypes. This process first involves extracting the self-attention weight map and identifying the top K visual vectors with the highest attention scores, which are considered the most representative and can capture the key foreground information in the image. To further optimize this selection process, this method sets a binary mask, which is used to modify the softmax function to retain the visual sequences related to the foreground information and with higher attention scores. Specifically, the j-th binary mask and the k-th binary mask are used to adjust the attention score formula, enabling the model to focus more on the foreground information, thereby improving the accuracy and relevance of the local features.
[0068] Next, the visual sequences saved by the foreground-preserving mask are used to perform calculation processing with each local prototype to calculate their similarity scores. This calculation is completed through a specific formula, which involves quantifying the similarity between each visual vector and the local prototype. After the max pooling operation, the similarity scores of the local prototypes are obtained. This step helps to extract the most representative information from multiple local features and enhances the model's ability to recognize local details.
[0069] S4. Use the fully connected layer to perform probability calculation processing on the similarity scores between the global features and the global prototype and the similarity scores between the local features and the local prototype to obtain the final classification probability, and generate an image recognition result based on the classification probability.
[0070] Specifically, step S4 includes: inputting the similarity scores between the global features and the global prototype and the similarity scores between the local features and the local prototype into a fully connected layer respectively, and using the softmax function as the activation function for calculation to obtain the global branch probability and the local branch probability , and the calculation formula is: , , where , , , are all parameters of the fully connected layer;
[0071] Performing weighted sum processing on the global branch probability and the local branch probability to obtain the final classification probability, and the formula is: , where is the weighting coefficient of the global branch, is the weighting coefficient of the local branch.
[0072] In this embodiment, a fully connected layer is used to perform probability calculation processing on the similarity scores between the global features and the global prototype and the similarity scores between the local features and the local prototype to obtain the final classification probability, and an image recognition result is generated based on this probability. This step is a key link in the entire image recognition process, which fuses local and global information to improve the accuracy and robustness of recognition.
[0073] Specifically, the similarity scores between the global features and the global prototype and the similarity scores between the local features and the local prototype are respectively input into a fully connected layer. In the fully connected layer, the softmax function is used as the activation function for calculation to obtain the global branch probability and the local branch probability. This calculation involves the parameters of the fully connected layer, including weights and biases, which are optimized during the training process to best represent the relationships between classes. The softmax function can convert any real-valued vector into a probability distribution, making each element's value between 0 and 1, and the sum of all elements is 1, which makes the softmax function very suitable for multi-classification problems. Immediately afterwards, weighted sum processing is performed on the global branch probability and the local branch probability to obtain the final classification probability. This step involves the weighting coefficient of the global branch and the weighting coefficient of the local branch
[0074] Specifically, in this embodiment, ProtoPNet is used as the basic model for interpretable fine-grained image classification, and the effectiveness of this method is verified through a publicly available fine-grained image dataset. For the convenience of description, this method denotes the feature selection and fusion module as ARWS and the global branch in the prototype learning layer as GB. The experimental results of comparing the ProtoPNet model, the ProtoPNet+GB model, and the final model of this method, ProtoPNet+GB+ARWS, are used to verify the effectiveness of this method, as shown in Table 1. Among them, the evaluation index of model accuracy is accuracy. The proportion of the number of images correctly recognized by the model to the total number of images. The results show that compared with the initial model, the accuracy of this method has increased by 5.77%, 5.75%, and 3.12% on the CUB, Dogs, and Cars datasets, respectively.
[0075] Table 1 Recognition accuracy results of each method
[0076]
[0077] In summary, the interpretable fine-grained image recognition realizes the full-process automated processing from image feature extraction to classification probability calculation through deep learning technology and the Transformer architecture. In the feature extraction stage, the input image is first sliced into multiple patches and linearly projected to generate a feature sequence. Subsequently, a class encoding vector and a position encoding vector are added to the feature sequence to form visual vectors, which are input into a stack of multiple Transformer encoder modules to obtain a deep feature representation. Each Transformer encoder contains a multi-head self-attention mechanism and a multi-layer perceptron, which can effectively capture the global and local features of the image. Secondly, in the prototype learning stage, the global branch and the local branch are used to process the global features and local features respectively. By calculating the similarity scores between these features and the preset global prototypes and local prototypes, the key information of the image is further refined. This process utilizes the self-attention weight map and the binary mask technology to ensure that the model focuses on the most important foreground information in the image. Finally, in the classification probability calculation stage, the fully connected layer is used to process the global and local similarity scores, which are converted into a probability distribution through the softmax function and combined with weighted sum processing to obtain the final classification probability. This method not only improves the classification accuracy but also enhances the model's ability to recognize local details of images, especially when dealing with complex and diverse image data.
[0078] Compared with the prior art, the interpretable fine-grained image recognition method has the following advantages:
[0079] The training process of the model is improved by adopting a dual-branch structure, which learns the global prototype and local prototype of the category respectively, rather than only generating the local prototype. The two prototypes add higher interpretability to the network and can use the global branch to guide the training of the local branch, further improving the accuracy of the network. A feature selection and fusion module is added to the feature extraction layer, enabling the network to select more important features for prototype learning, further enhancing the network's representation ability, and making the prototype learning more focused on the foreground rather than the background through this module. A new method for calculating feature importance evaluation is used, which is more effective than the traditional method of using attention scores to evaluate feature importance. Briefly speaking, the interpretable fine-grained image recognition method improves the accuracy and robustness of image recognition, enhances the generalization ability of the model, and enables it to maintain a high recognition accuracy when facing new and unseen data. In addition, this method also improves the flexibility and adaptability of the model because it can balance the importance of global and local features according to the requirements of different application scenarios. Finally, this method provides a new solution for the field of image recognition, with broad application prospects and practical value.
[0080] Please refer to Figure 7 , the second embodiment of the present invention provides an interpretable fine-grained image recognition device, which includes:
[0081] A feature extraction unit 201, configured to obtain an image x to be recognized, input the image x into a feature extraction layer for feature extraction preprocessing, and generate an output vector;
[0082] A global branch unit 202, configured to use the category encoding vector of the output vector as global features, input them into the global branch of the prototype learning layer for calculation, and obtain the similarity score between the global features and the global prototype;
[0083] A local branch unit 203, configured to obtain N local features output by the feature extraction layer, and select K preset foreground information-related features for local prototype learning to obtain the similarity score between the local features and the local prototype;
[0084] A fully connected unit 204, configured to use a fully connected layer to perform probability calculation processing on the similarity score between the global features and the global prototype and the similarity score between the local features and the local prototype, obtain the final classification probability, and generate an image recognition result based on the classification probability.
[0085] The third embodiment of the present invention provides an interpretable fine-grained image recognition device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the interpretable fine-grained image recognition method described in any one of the above.
[0086] A fourth embodiment of the present invention provides a readable storage medium storing a computer program that can be executed by a processor of a device where the storage medium is located to implement the interpretable fine-grained image recognition method described in any one of the above.
[0087] Exemplarily, each of the above-mentioned devices and each process step can be implemented by a computer program, which can be divided into one or more units. The one or more units are stored in the memory and executed by the processor to complete the present invention.
[0088] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0089] The memory can be used to store the computer program and / or module. The processor realizes various functions of the present invention by running or executing the computer program and / or module stored in the memory and calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash device, or other volatile solid-state storage devices.
[0090] Among them, if the unit integrated in the electronic device or printer is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0091] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative work.
[0092] The above is the preferred embodiment of the present invention. It should be pointed out that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. An interpretable fine-grained image recognition method, characterized in that: include: Obtain an image x to be identified, input the image x into a feature extraction layer for feature extraction preprocessing, and generate an output vector; The category encoding vector of the output vector is used as the global feature and input into the global branch of the prototype learning layer for calculation to obtain the similarity score between the global feature and the global prototype; Obtaining N local features output by the feature extraction layer, and selecting K preset foreground information related features for local prototype learning, to obtain similarity scores between the local features and the local prototype; Using a fully connected layer, a probability calculation process is performed on the similarity scores between the global feature and the global prototype and the similarity scores between the local feature and the local prototype to obtain a final classification probability, and based on the classification probability, an image recognition result is generated; Get the image x to be identified, input the image x into the feature extraction layer for feature extraction preprocessing, and generate an output vector, specifically: Get the image x to be recognized, , where the size of image x is , H is the image height, W is the image width, C is the number of image channels, is a matrix; Divide the image x into N patches of equal size and tile the patches to a length of Sequence images of , , where the patch size is , is the resolution of each image block, is the number of blocks finally generated; For the sequence of images Perform linear projection processing to generate a feature sequence; Add the category encoding vector to the feature sequence , as the encoding feature of image x; Add a positional encoding vector , connect the two-dimensional position information of image x to the one-dimensional feature sequence to obtain the visual vector , ,in, is the first image sequence, is the second image sequence, is the Nth image sequence, E is the unit matrix, is the product of the first image sequence and the identity matrix, is the product of the second image sequence and the identity matrix, is the product of the Nth image sequence and the identity matrix; The visual vector Input to the module of stacking L Transformer encoders to get the output vector of the lth layer Transformer encoder , , , where each Transformer encoder includes one or more self-attention heads and a multi-layer perceptron, LN represents the normalization operation, MSA represents multi-head self-attention, and MLP represents a multi-layer perceptron. is the output vector of the l-1th layer Transformer encoder; The output vector The category encoding vector of , input to the global branch, and output vector Fusion features As local features, input to the local branch.
2. The interpretable fine-grained image recognition method according to claim 1, characterized in that: The calculation steps of multiple self-attention head results of the l-th layer Transformer encoder are as follows: The visual sequence input to the l-th layer Transformer encoder Perform linear transformation to obtain three identical visual sequences, which are called query vector q, key vector k, and value vector v; Attention weight based on the h-th self-attention head , use the softmax function to calculate the similarity between the query vector q and all key vectors k, and multiply each value vector v by its corresponding attention weight to obtain a weighted value vector, and take the sum of all weighted value vectors as the output vector. The calculation formula is: , , ,in, is the linear projection matrix of the h-th self-attention head, W is the linear projection matrix, is the length of the h-th self-attention head, T is the transpose of the matrix, and A is the self-attention weight map; Calculate the results of H self-attention heads in parallel, add them up, and get the final multi-head self-attention result .
3. The interpretable fine-grained image recognition method according to claim 2, characterized in that: Fusion Features The calculation steps are as follows: Get the self-attention weight map of the hth attention head of the lth layer Transformer encoder , , ,in, is the attention value between the i-th visual sequence and the j-th visual sequence, The attention value between the category encoding vector and all visual sequences; Based on self-attention weight map , calculate the attention score standard value matrix of the l-th layer Transformer encoder, and its recursive formula is: , ,in, is the identity matrix, H is the number of attention heads; set up is the standard value of the attention score of the i-th visual sequence and the category encoding vector, and the order of i from 1 to N is Before sorting, filtering The visual sequence with the highest attention score standard value is taken as the important visual sequence of each layer, which is expressed as: ,in, is the i-th important visual sequence in the l-th layer Transformer encoder, is the number of important visual sequences selected; Get the fused features , Represents the features selected by the ARWS module from the L-1th layer.
4. The interpretable fine-grained image recognition method according to claim 3, characterized in that: The category encoding vector of the output vector is used as the global feature and input into the global branch of the prototype learning layer for calculation to obtain the similarity score between the global feature and the global prototype, which is: Predefined A learnable global prototype ,and Learnable local prototypes ; Get the global features output by the feature extraction layer , calculate the global feature The similarity score with each global prototype is calculated as: ,in, is the ith global prototype, Global feature The similarity score with the i-th global prototype, is the smoothing term.
5. The interpretable fine-grained image recognition method according to claim 4, characterized in that: Obtain N local features output by the feature extraction layer, and select K preset foreground information related features for local prototype learning to obtain the similarity score between the local features and the local prototype, specifically: Extracted from the attention weight map , and select the first K visual vectors with the highest attention scores; Set a binary mask , the attention score of the i-th position encoding vector to the category encoding vector The softmax function in the formula is modified as follows: , to preserve the vision sequence saved using the foreground save mask, where is the j-th binary mask, is the kth binary mask; The visual sequence saved using the foreground preservation mask is processed with each local prototype to calculate its similarity score, and after maximum pooling, the similarity score of the local prototype is obtained. The calculation formula is: ,in, represents the K selected visual vectors with the highest attention scores, Indicates that each time from the visual vector A single visual vector selected from is the i-th local prototype, is the similarity score between the local feature and the i-th local prototype.
6. The interpretable fine-grained image recognition method according to claim 5, characterized in that: The similarity scores between the global feature and the global prototype and the similarity scores between the local feature and the local prototype are subjected to probability calculation using a fully connected layer to obtain the final classification probability, specifically: The similarity scores between the global feature and the global prototype and the similarity scores between the local feature and the local prototype are respectively input into the fully connected layer, and the softmax function is used as the activation function for calculation to obtain the global branch probability and local branching probability , the calculation formula is: , ,in, , , , These are the parameters of the fully connected layer; For the global branch probability and local branching probability Perform weighted sum processing to obtain the final classification probability, the formula is: ,in, is the weight coefficient of the global branch, is the weight coefficient of the local branch.
7. An interpretable fine-grained image recognition device, characterized in that: include: A feature extraction unit is used to obtain an image x to be identified, input the image x into a feature extraction layer for feature extraction preprocessing, and generate an output vector; The global branch unit is used to input the category encoding vector of the output vector as a global feature into the global branch of the prototype learning layer for calculation to obtain a similarity score between the global feature and the global prototype; A local branch unit, used to obtain the local features output by the N feature extraction layers, and select K preset foreground information related features for local prototype learning to obtain a similarity score between the local features and the local prototype; A fully connected unit, used to perform probability calculation processing on the similarity scores between the global feature and the global prototype and the similarity scores between the local feature and the local prototype using a fully connected layer to obtain a final classification probability, and generate an image recognition result based on the classification probability; Get the image x to be identified, input the image x into the feature extraction layer for feature extraction preprocessing, and generate an output vector, specifically: Get the image x to be recognized, , where the size of image x is , H is the image height, W is the image width, C is the number of image channels, is a matrix; Divide the image x into N patches of equal size and tile the patches to a length of Sequence images , , where the patch size is , is the resolution of each image block, is the number of blocks finally generated; For the sequence of images Perform linear projection processing to generate a feature sequence; Add the category encoding vector to the feature sequence , as the encoding feature of image x; Add a positional encoding vector , connect the two-dimensional position information of image x to the one-dimensional feature sequence to obtain the visual vector , ,in, is the first image sequence, is the second image sequence, is the Nth image sequence, E is the unit matrix, is the product of the first image sequence and the identity matrix, is the product of the second image sequence and the identity matrix, is the product of the Nth image sequence and the identity matrix; The visual vector Input to the module of stacking L Transformer encoders to get the output vector of the lth layer Transformer encoder , , , where each Transformer encoder includes one or more self-attention heads and a multi-layer perceptron, LN represents the normalization operation, MSA represents multi-head self-attention, and MLP represents a multi-layer perceptron. is the output vector of the l-1th layer Transformer encoder; The output vector The category encoding vector of , input to the global branch, and output vector Fusion features As local features, input to the local branch.
8. An interpretable fine-grained image recognition device, characterized in that: The invention comprises a processor, a memory and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the interpretable fine-grained image recognition method as described in any one of claims 1 to 6 is implemented.
9. A readable storage medium, characterized in that: A computer program is stored, and the computer program can be executed by a processor of the device where the storage medium is located to implement the interpretable fine-grained image recognition method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Explanatable image recognition method based on visual Transform and prototype learning
CN119295818A