A fine-grained image recognition and classification method, system, storage medium and device
By adopting a Transformer-based visual model in fine-grained image classification, combining feature pyramid network, weakly supervised sampling, neural decision forest and trusted multi-view classification module, the model's shortcomings in identifying differentiated areas in the image are solved, and higher classification accuracy and interpretability are achieved.
Patent Information
- Application Number
- CN202510251909.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The prior art is difficult to effectively identify the most distinguishing areas in the image in fine-grained image classification, resulting in low classification accuracy and lack of interpretability of the model.
Using a visual model based on Transformer, multi-scale features are fused through the feature pyramid network module, the weakly supervised sampling module screens out beneficial features, the neural decision forest module conducts decision integration, and the trusted multi-view classification module is used to integrate evidence to generate the final classification results.
The accuracy of fine-grained image classification is improved, allowing the model to discover discriminant areas faster, taking into account the interpretability and training efficiency of the model.
Smart Images

Figure CN119762894B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly relates to a fine-grained image recognition and classification method, system, storage medium and device. Background Art
[0002] Fine-grained visual classification (FGVC) is a highly challenging task, whose main goal is to discover and classify objects based on discriminative regions and features in images. This task is even more difficult when distinguishing between classes with extremely similar appearances. Due to the high inter-class similarity and low intra-class difference in most images, as well as problems such as inconsistent environments and poses, both humans and machines face bottlenecks in recognition.
[0003] In the past, convolutional neural networks (CNNs) have achieved remarkable results in this field. However, with the introduction of the Vision Transformer (ViT) model, ViT has outperformed CNNs in multiple vision tasks and has become the new cutting-edge technology. In recent years, the application of ViT in FGVC tasks has gradually increased. For example, some models use the attention mechanism to extract local features of images and optimize their discriminative ability. However, whether it is a CNN-based or ViT-based model, its essence is a black-box learning mechanism, and there are problems with the interpretability of the model.
[0004] In the fine-grained image classification task, the model needs to identify the image category from tiny visual differences. This type of task is highly challenging because in fine-grained images, the visual features of different classes often have subtle differences, and the presence of background and irrelevant regions may also interfere with the model's focus on key parts, affecting the classification effect. To improve the performance of the model in these tasks, the key lies in helping the model find the most discriminative parts or "key patches" in the image. The current mainstream methods include part-based localization, attention-based methods, and feature-channel-based methods. These methods optimize the model's ability to focus on discriminative regions in different aspects and can usually be combined to improve the classification effect, but these methods are usually not interpretable, and people do not know why the model pays attention to the so-called discriminative regions.
[0005] In the fusion methods for fine-grained image classification, there are feature-level fusion and decision-level fusion. Feature-level fusion is to integrate the features of different layers or different models during the feature extraction process, such as concatenation, weighting, or fusion through an attention mechanism, enabling the model to better capture local details and global semantic information and enhancing the recognition ability for tiny differences. Decision-level fusion is to synthesize the results of multiple models or multiple inferences during the classification decision stage, and obtain the final prediction result through voting, averaging, or weighting, etc., thereby improving the robustness and classification accuracy of the overall system. For decision-level fusion, direct fusion is too simple. If appropriate fusion weights or strategies are not found, it may lead to unsatisfactory fusion effects and even reduce the classification accuracy, especially in fine-grained classification tasks, and the fusion means lack interpretability. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a fine-grained image recognition and classification method, system, storage medium, and device, aiming to improve the recognition and classification accuracy of fine-grained images and make fine-grained image classification more accurate.
[0007] The first aspect of the present invention provides a fine-grained image recognition and classification method, and the method includes:
[0008] Obtain the image dataset corresponding to the pre-collected fine-grained images, and import the image dataset into the vision model based on Transformer;
[0009] Import the feature maps of each stage into the feature pyramid network module of the vision model to obtain the first target feature map that fuses shallow-layer information;
[0010] Import the first target feature map into the weakly supervised sampling module of the vision model, screen out the target features beneficial to classification, and stack the target features into a target vector and a second target feature map;
[0011] Import the second target feature map and the initial prediction value without passing through the activation function into the neural decision forest module of the vision model, and obtain the target prediction value corresponding to each granularity and store it in the output sequence;
[0012] Import the output sequence into the credible multi-view classification module of the vision model, and fuse the prediction value distributions of multiple opinions to obtain the final fused classification result.
[0013] According to one aspect of the above technical solution, the step of importing the feature maps of each stage into the feature pyramid network module of the vision model to obtain the first target feature map that fuses shallow-layer information includes:
[0014] Import each stage feature map in the image dataset into the feature pyramid network module of the visual model;
[0015] Through the feature pyramid network module, layer-by-layer fuse the deep features containing semantic information with the shallow features containing spatial information to obtain the first target feature map.
[0016] According to one aspect of the above technical solution, the steps of importing the first target feature map into the weakly supervised sampling module of the visual model, screening out the target features beneficial to classification, and stacking the target features into a target vector and a second target feature map include:
[0017] Import the first target feature map into the weakly supervised sampling module of the visual model, and perform class prediction on the pixel values, i.e., feature points, of the first target feature map through the fully connected layer in the weakly supervised sampling module;
[0018] Perform multi-class probability conversion processing on the prediction results. If the highest prediction probability of the feature point is greater than the preset threshold, determine that the feature point is a target feature helpful for classification;
[0019] Based on the target features, perform feature stacking to output a target vector and a second target feature map.
[0020] According to one aspect of the above technical solution, the method further includes:
[0021] If the highest prediction probability of the feature point is less than the preset threshold, regard it as an invalid feature point with a contribution to fine-grained classification less than the preset value, and filter and remove the invalid feature points.
[0022] According to one aspect of the above technical solution, the steps of importing the second target feature map and the initial prediction value without passing through the activation function into the neural decision forest module of the visual model, and obtaining the target prediction value corresponding to each granularity and storing it in the output sequence include:
[0023] The neural decision forest module consists of multiple neural decision trees. Each neural decision tree consists of multiple nodes, and each node performs hierarchical routing on the input features through a hierarchical structure, so that the visual model captures the fine-grained information in the second target feature map.
[0024] According to one aspect of the above technical solution, the steps of importing the output sequence into the credible multi-view classification module of the visual model, and fusing the prediction value distributions of multiple opinions to obtain the final fused classification result include:
[0025] Import the output sequence into the credible multi-view classification module of the visual model;
[0026] In the trusted multi-view classification module, the category probability distribution is modeled by a variational Dirichlet distribution, and Dempster-Shafer is used for evidence integration to fuse the predicted value distributions of multiple opinions and obtain the finally fused classification result.
[0027] The second aspect of the present invention is to provide a fine-grained image recognition and classification system, which is applied to the method described in the above technical solution. The system includes:
[0028] An image data import module, configured to obtain an image data set corresponding to pre-collected fine-grained images and import the image data set into a vision model based on Transformer;
[0029] A shallow information fusion module, configured to import the feature maps of each stage into the feature pyramid network module of the vision model to obtain a first target feature map that fuses shallow information;
[0030] A pixel feature screening module, configured to import the first target feature map into the weakly supervised sampling module of the vision model, screen out target features beneficial to classification, and stack the target features into a target vector and a second target feature map;
[0031] A decision-making module, configured to import the second target feature map and the initial predicted value without passing through an activation function into the neural decision forest module of the vision model, respectively obtain the target predicted value corresponding to each granularity and store it in the output sequence;
[0032] A classification module, configured to import the output sequence into the trusted multi-view classification module of the vision model, fuse the predicted value distributions of multiple opinions, and obtain the finally fused classification result.
[0033] According to one aspect of the above technical solution, the shallow information fusion module is specifically configured to:
[0034] Import the feature maps of each stage in the image data set into the feature pyramid network module of the vision model;
[0035] Through the feature pyramid network module, layer-by-layer fusion of deep features containing semantic information and shallow features containing spatial information is performed to obtain a first target feature map.
[0036] The third aspect of the present invention is to provide a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above technical solution is implemented.
[0037] A fourth aspect of the present invention is to provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the above technical solution is implemented.
[0038] Compared with the prior art, the fine-grained image recognition and classification method, system, storage medium, and device shown in the present invention have the following beneficial effects:
[0039] The present invention proposes a fine-grained image recognition and classification method for fine-grained image classification. While taking into account the interpretability of the model, multi-granularity information is used to improve the training efficiency of the visual model, guide the model to discover discriminative regions faster, and thus improve the model accuracy. Moreover, a vision model based on transformer is used as a feature extractor, and feature maps with multi-granularity information having different receptive fields output at different levels are extracted. After passing through a feature pyramid, key feature screening is finally performed on each layer of granularity feature maps, and the multi-granularity information containing key features beneficial for classification is sent into an interpretable prototype decision tree forest. Finally, the decision forest uses an interpretable dynamic decision fusion algorithm to form a prediction result with prediction robustness to form the final classification, thereby being able to improve the recognition and classification accuracy of fine-grained images and making the classification more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The above and / or additional aspects and advantages of the present invention will become apparent and easier to understand from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0041] Figure 1 is a schematic flowchart of a fine-grained image recognition and classification method in an embodiment of the present invention;
[0042] Figure 2 is a structural block diagram of a fine-grained image recognition and classification system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, features, and advantages of the present invention more obvious and understandable, the following provides a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings. Several embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.
[0044] It should be noted that when an element is referred to as "fixed to" another element, it can be directly on the other element or there may also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used in this article are only for the purpose of illustration.
[0045] Unless otherwise defined, all technical and scientific terms used in this article have the same meaning as those commonly understood by those skilled in the technical field to which this invention belongs. The terms used in the description of this invention in this article are only for the purpose of describing specific embodiments and are not intended to limit this invention. The term "and / or" used in this article includes any and all combinations of one or more of the related listed items.
[0046] Embodiment 1
[0047] Please refer to Figure 1 , the first embodiment of the present invention provides a fine-grained image recognition and classification method, and the method includes steps S10 - step S50:
[0048] Step S10, obtain the image data set corresponding to the pre-collected fine-grained image, and import the image data set into the vision model based on Transformer.
[0049] First of all, it should be noted that fine-grained images refer to images with subtle differences visually, and these differences are mainly reflected between different sub-categories under the same large category. Due to the very small visual differences between sub-categories, it is difficult to classify fine-grained images, and high-resolution images and complex algorithms are required to achieve accurate classification. For example, distinguishing different breeds of birds, car models, dog breeds, etc.
[0050] In this embodiment, the fine-grained images of the target object will be collected first, and then an image data set will be constructed based on the fine-grained images collected at different time stages, and the image data set will be imported into the vision model based on Transformer, and the vision model based on Transformer is the Swin Transformer model.
[0051] Among them, the above-mentioned vision model includes a feature pyramid network module, a weakly supervised sampling module, a neural decision forest module, and a reliable multi-view classification module.
[0052] Step S20, import the feature maps of each stage into the feature pyramid network module of the vision model to obtain the first target feature map that fuses shallow-layer information.
[0053] In this embodiment, the step of importing each stage feature map into the feature pyramid network module of the visual model to obtain the first target feature map that fuses shallow layer information includes:
[0054] Import each stage feature map in the image dataset into the feature pyramid network module of the visual model;
[0055] Through the feature pyramid network module, layer-by-layer fusion of deep features containing semantic information and shallow features containing spatial information is performed to obtain the first target feature map.
[0056] First of all, it should be noted that the vision model based on Transformer has four stages, which are equivalent to four feature extraction steps, respectively used to extract features from images. The feature maps extracted from each stage are sent to the feature pyramid for processing to fuse multi-level feature map information, so as to obtain the first target feature map that fuses shallow layer information.
[0057] Among them, the feature pyramid network module, that is, the FPN module (Feature Pyramid Network, FPN), is a module for multi-scale feature fusion and is commonly used in image processing tasks to enhance the model's recognition ability for objects of different scales.
[0058] Specifically, the FPN module constructs a bottom-up feature pyramid to layer-by-layer fuse deep features including semantic information and shallow features including spatial information.
[0059] In specific implementation, the FPN module extracts feature maps from different levels of the backbone network of the visual model, first performs dimensionality reduction projection to make the feature maps corresponding to different levels have the same feature dimension, then aligns the feature maps of different resolutions through upsampling operations, and performs layer-by-layer feature addition to obtain the first target feature map that fuses shallow layer information.
[0060] Adopting the feature pyramid network module shown in this embodiment can effectively integrate multi-scale information in the image, enabling the vision model to retain key global semantic information and capture subtle local features when processing fine-grained tasks, thereby improving the accuracy and robustness of overall classification or detection.
[0061] In step S30, import the first target feature map into the weakly supervised sampling module of the visual model, screen out the target features beneficial to classification, and stack the target features into a target vector and a second target feature map.
[0062] In this embodiment, the step of importing the first target feature map into the weakly supervised sampling module of the visual model, screening out target features beneficial to classification, and stacking the target features into a target vector and a second target feature map includes:
[0063] Import the first target feature map into the weakly supervised sampling module of the visual model, and perform class prediction on the pixel values, i.e., feature points, of the first target feature map through the fully connected layer in the weakly supervised sampling module;
[0064] Perform multi-class probability conversion processing on the prediction result. If the highest prediction probability of the feature point is greater than the preset threshold, determine that the feature point is a target feature helpful for classification;
[0065] Perform feature stacking based on the target features, and output a target vector and a second target feature map.
[0066] The method further includes:
[0067] If the highest prediction probability of the feature point is less than the preset threshold, it is regarded as an invalid feature point with a contribution to fine-grained classification less than the preset value, and the invalid feature point is filtered out.
[0068] In this embodiment, the step of importing the first target feature map into the weakly supervised sampling module of the visual model, screening out target features beneficial to classification, and stacking the target features into a target vector and a second target feature map is specifically implemented by using a weakly supervised sampling module (Weakly Supervised Sampling, WSS). In the WSS module, the pixel value of the first target feature map is called a feature point. This WSS module adopts a simple and efficient design, and performs class prediction on each feature point through a fully connected layer.
[0069] Specifically, after performing multi-class probability conversion processing (softmax processing) on the prediction result, if the highest prediction probability of the feature point exceeds a certain preset threshold, the feature point is considered to be a valid feature helpful for classification, that is, the above-mentioned target feature, and will be retained and used for subsequent feature fusion steps.
[0070] On the contrary, if the highest prediction probability of the feature point does not reach the preset threshold, it is regarded as a feature point with a small contribution to fine-grained classification, determined as an invalid feature point, and then filtered out.
[0071] By using the above threshold-based feature point screening method, the WSS module can automatically screen out the most discriminative feature points for fine-grained image classification, effectively reducing the computational burden of the visual model, and helping the visual model to better focus on the discriminative key parts in fine-grained images, thereby improving the classification accuracy.
[0072] Step S40: Import the second target feature map and the initial prediction value without being processed by an activation function into the neural decision forest module of the visual model, and obtain the target prediction value corresponding to each granularity and store it in the output sequence.
[0073] In this embodiment, a neural decision forest composed of multiple neural decision trees is first constructed to enhance the overall robustness of the model. Each neural decision tree consists of multiple nodes, and the nodes perform hierarchical routing on the input features through a hierarchical structure, enabling the visual model to capture the fine-grained information in the second target feature map. Each neural decision tree in the neural decision forest independently learns different feature distributions and discrimination rules, thereby improving the model's adaptability to complex data and classification accuracy as a whole. By integrating the decisions of multiple neural decision trees, the neural decision forest can effectively balance the bias and variance of a single tree model, improve the resistance to noise and sample bias, and finally achieve robust and accurate classification of the input image.
[0074] Step S50: Import the output sequence into the trusted multi-view classification module of the visual model, and fuse the predicted value distributions of multiple opinions to obtain the final fused classification result.
[0075] In this embodiment, the step of importing the output sequence into the trusted multi-view classification module of the visual model and fusing the predicted value distributions of multiple opinions to obtain the final fused classification result includes:
[0076] Import the output sequence into the trusted multi-view classification module of the visual model;
[0077] In the trusted multi-view classification module, model the category probability distribution through a variational Dirichlet distribution, and use Dempster-Shafer for evidence integration to fuse the predicted value distributions of multiple opinions and obtain the final fused classification result.
[0078] Among them, in this embodiment, a trusted multi-view classification module (Trusted Multi-View Classification, TMC) is used to effectively integrate multi-granularity information and improve the reliability and credibility of decisions.
[0079] In this embodiment, different from traditional feature or output level fusion methods, the TMC module fuses evidence information from each perspective at the evidence level, so as to achieve stable and more reasonable uncertainty estimation for multi-view integration. In this embodiment, the categorical probability distribution is modeled by variational Dirichlet distribution, and Dempster-Shafer theory is used for evidence integration to ensure the optimization of multi-view information in terms of interpretability and credibility.
[0080] Specifically, the TMC module combines the uncertainties of each perspective and adaptively assigns perspective weights during the visual model training process, making the fused decision more reliable, more robust, and having theoretically guaranteed credibility. In addition, the TMC module not only does not require additional calculations or network modifications, but also provides evidence-based integration for the classification of each granularity through subjective uncertainty, thus significantly improving the classification accuracy and interpretability.
[0081] In summary, compared with the prior art, the beneficial effects of adopting the fine-grained image recognition and classification method shown in this embodiment are as follows:
[0082] In this embodiment, a fine-grained image recognition and classification method is proposed for fine-grained image classification. While taking into account the interpretability of the model, multi-granularity information is used to improve the training efficiency of the visual model and guide the model to discover discriminative regions faster, thereby improving the model accuracy. Moreover, a vision model based on transformer is used as a feature extractor, and feature maps with multi-granularity information having different receptive fields of different hierarchical outputs are extracted. After passing through a feature pyramid, finally, key feature screening is performed on each layer of granularity feature maps, and the multi-granularity information containing key features beneficial for classification is sent into an interpretable prototype decision tree forest. Finally, the decision forest uses an interpretable dynamic decision fusion algorithm to form a prediction result with prediction robustness for final classification, so as to be able to improve the recognition and classification accuracy of fine-grained images and make the classification more accurate.
[0083] Embodiment Two
[0084] Please refer to Figure 2 , the second embodiment of the present invention provides a fine-grained image recognition and classification system, and the system includes: an image data import module 10, a shallow information fusion module 20, a pixel feature screening module 30, a decision module 40, and a classification module 50.
[0085] The image data import module 10 is used to obtain an image data set corresponding to pre-collected fine-grained images and import the image data set into a vision model based on Transformer.
[0086] First of all, it should be noted that fine-grained images refer to images with subtle visual differences, which are mainly reflected among different sub-categories under the same major category. Due to the very small visual differences between sub-categories, fine-grained image classification is difficult and requires high-resolution images and complex algorithms to achieve accurate classification. For example, distinguishing different breeds of birds, car models, dog breeds, etc.
[0087] In this embodiment, first, fine-grained images of the target object will be collected, and then an image dataset will be constructed based on the fine-grained images collected at different time stages. The image dataset will be imported into a vision model based on Transformer, and the vision model based on Transformer is the Swin Transformer model.
[0088] Among them, the above-mentioned vision model includes a Feature Pyramid Network module, a Weak Supervision Sampling module, a Neural Decision Forest module, and a Trusted Multi-View Classification module.
[0089] The shallow information fusion module 20 is used to import the feature maps of each stage into the Feature Pyramid Network module of the vision model to obtain a first target feature map that fuses shallow information.
[0090] In this embodiment, the shallow information fusion module 20 is specifically used for:
[0091] Import the feature maps of each stage in the image dataset into the Feature Pyramid Network module of the vision model;
[0092] Through the Feature Pyramid Network module, layer-by-layer fusion of deep features containing semantic information and shallow features containing spatial information is performed to obtain a first target feature map.
[0093] Among them, the Feature Pyramid Network module, that is, the FPN module (Feature Pyramid Network, FPN), is a module for multi-scale feature fusion and is commonly used in image processing tasks to enhance the model's ability to recognize objects at different scales.
[0094] Specifically, the FPN module constructs a bottom-up feature pyramid to perform layer-by-layer fusion of deep features containing semantic information and shallow features containing spatial information.
[0095] In specific implementation, the FPN module extracts feature maps from different levels of the backbone network of the vision model, first performs dimensionality reduction projection to make the feature maps corresponding to different levels have the same feature dimension, then aligns the feature maps with different resolutions through upsampling operations, and performs layer-by-layer feature addition to obtain a first target feature map that fuses shallow information.
[0096] By adopting the feature pyramid network module shown in this embodiment, multi-scale information in the image can be effectively integrated, enabling the visual model to retain key global semantic information and capture subtle local features when processing fine-grained tasks, thereby improving the accuracy and robustness of overall classification or detection.
[0097] The pixel feature screening module 30 is used to import the first target feature map into the weakly supervised sampling module of the visual model, screen out the target features beneficial to classification, and stack the target features into a target vector and a second target feature map.
[0098] In this embodiment, the pixel feature screening module 30 is specifically used for:
[0099] Import the first target feature map into the weakly supervised sampling module of the visual model, and perform class prediction on the pixel values, i.e., feature points, of the first target feature map through the fully connected layer in the weakly supervised sampling module;
[0100] Perform multi-class probability conversion processing on the prediction results. If the highest prediction probability of the feature point is greater than the preset threshold, determine that the feature point is a target feature helpful for classification;
[0101] Stack the features based on the target features, and output a target vector and a second target feature map.
[0102] In this embodiment, the pixel feature screening module 30 is further used for:
[0103] If the highest prediction probability of the feature point is less than the preset threshold, it is regarded as an invalid feature point with a contribution to fine-grained classification less than the preset value, and the invalid feature point is filtered out.
[0104] In this embodiment, during the use of the weakly supervised sampling module (Weakly Supervised Sampling, WSS), the pixel value of the first target feature map is called a feature point. This WSS module adopts a simple and efficient design, and performs class prediction on each feature point through a fully connected layer.
[0105] Specifically, after performing multi-class probability conversion processing (softmax processing) on the prediction results, if the highest prediction probability of the feature point exceeds a certain preset threshold, the feature point is considered an effective feature helpful for classification, that is, the above-mentioned target feature, and will be retained and used for subsequent feature fusion steps.
[0106] Conversely, if the highest prediction probability of the feature point does not reach the preset threshold, it is regarded as a feature point with a small contribution to fine-grained classification, determined as an invalid feature point, and then filtered out.
[0107] By adopting the above-mentioned threshold-based feature point screening method, the WSS module can automatically screen out the most discriminative feature points for fine-grained image classification, effectively reducing the computational burden of the visual model and helping the visual model better focus on the discriminative key parts in the fine-grained image, thereby improving the classification accuracy.
[0108] The decision-making module 40 is used to import the second target feature map and the initial prediction value without passing through the activation function into the neural decision forest module of the visual model, respectively obtain the target prediction value corresponding to each granularity and store it in the output sequence.
[0109] In this embodiment, first, a neural decision forest composed of multiple neural decision trees is constructed to enhance the overall robustness of the model. Each neural decision tree consists of multiple nodes, and the nodes perform hierarchical routing on the input features through a hierarchical structure, enabling the visual model to capture the fine-grained information in the second target feature map. Each neural decision tree in the neural decision forest independently learns different feature distributions and discrimination rules, thereby improving the model's adaptability to complex data and classification accuracy as a whole. By integrating the decisions of multiple neural decision trees, the neural decision forest can effectively balance the bias and variance of a single tree model, improve the resistance to noise and sample bias, and finally achieve robust and accurate classification of the input image.
[0110] The classification module 50 is used to import the output sequence into the trusted multi-view classification module of the visual model, fuse the predicted value distributions of multiple opinions, and obtain the final fused classification result.
[0111] In this embodiment, the classification module 50 is specifically used for:
[0112] Import the output sequence into the trusted multi-view classification module of the visual model;
[0113] In the trusted multi-view classification module, the class probability distribution is modeled by a variational Dirichlet distribution, and Dempster-Shafer is used for evidence integration to fuse the predicted value distributions of multiple opinions and obtain the final fused classification result.
[0114] Among them, in this embodiment, a trusted multi-view classification module (Trusted Multi-View Classification, TMC) is used to effectively integrate multi-granularity information and improve the reliability and credibility of decision-making.
[0115] In this embodiment, different from traditional feature or output-level fusion methods, the TMC module fuses the evidence information from each perspective at the evidence level, so as to achieve a stable and more reasonable uncertainty estimation for multi-view integration. In this embodiment, the categorical probability distribution is modeled by a variational Dirichlet distribution, and the Dempster-Shafer theory is used for evidence integration to ensure the optimization of multi-view information in terms of interpretability and credibility.
[0116] Specifically, the TMC module combines the uncertainties of each perspective and adaptively assigns perspective weights during the visual model training process, making the fused decision more reliable, more robust, and having theoretically guaranteed credibility. In addition, the TMC module not only does not require additional calculations or network modifications, but also provides evidence-based integration for the classification of each granularity through subjective uncertainty, thus significantly improving the classification accuracy and interpretability.
[0117] In summary, compared with the prior art, the beneficial effects of adopting the fine-grained image recognition and classification system shown in this embodiment are as follows:
[0118] In this embodiment, a fine-grained image recognition and classification system is proposed for fine-grained image classification. While taking into account the interpretability of the model, multi-granularity information is used to improve the training efficiency of the visual model, guiding the model to discover discriminative regions faster, thereby improving the model accuracy. Moreover, a vision model based on a transformer is used as a feature extractor, and feature maps with multi-granularity information having different receptive fields at different hierarchical outputs are extracted. After passing through a feature pyramid, finally, key feature screening is performed on each layer of granularity feature maps, and the multi-granularity information containing key features beneficial for classification is sent into an interpretable prototype decision tree forest. Finally, the decision forest uses an interpretable dynamic decision fusion algorithm to form a prediction result with prediction robustness for final classification, so as to be able to improve the recognition and classification accuracy of fine-grained images and make the classification more accurate.
[0119] Embodiment III
[0120] The third embodiment of the present invention provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described in the above embodiment is implemented.
[0121] Embodiment IV
[0122] The fourth embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method described in the above embodiment is implemented.
[0123] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0124] The above-described embodiments merely represent several implementation manners of the present invention. The descriptions thereof are relatively specific and detailed, but should not be construed as a limitation to the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A fine-grained image recognition and classification method, characterized in that: The method comprises: Obtain an image dataset corresponding to a pre-collected fine-grained image, and import the image dataset into a Transformer-based visual model; Importing each stage feature map into the feature pyramid network module of the visual model to obtain a first target feature map that integrates shallow information; Importing the first target feature map into the weakly supervised sampling module of the visual model, screening out target features that are useful for classification, and stacking them into a target vector and a second target feature map based on the target features; Importing the second target feature map and the initial prediction value that has not been processed by the activation function into the neural decision forest module of the visual model, respectively obtaining the target prediction value corresponding to each granularity and storing them in the output sequence; The output sequence is imported into the credible multi-view classification module of the visual model, and the predicted value distributions of multiple opinions are fused to obtain a final fused classification result.
2. The fine-grained image recognition and classification method according to claim 1, characterized in that: The step of importing each stage feature map into the feature pyramid network module of the visual model to obtain a first target feature map that integrates shallow information includes: Importing the feature map of each stage in the image data set into the feature pyramid network module of the visual model; The deep features containing semantic information and the shallow features containing spatial information are fused layer by layer through the feature pyramid network module to obtain a first target feature map.
3. The fine-grained image recognition and classification method according to claim 1, characterized in that: The step of importing the first target feature map into the weakly supervised sampling module of the visual model, screening out target features that are useful for classification, and stacking the target features into a target vector and a second target feature map comprises: Importing the first target feature map into the weakly supervised sampling module of the visual model, and performing category prediction on the pixel values, i.e., feature points, of the first target feature map through the fully connected layer in the weakly supervised sampling module; Performing multi-classification probability conversion processing on the prediction results, if the highest prediction probability of the feature point is greater than a preset threshold, determining the feature point as a target feature that is helpful for classification; Feature stacking is performed based on the target features, and a target vector and a second target feature map are output.
4. The fine-grained image recognition and classification method according to claim 3 is characterized in that: The method further comprises: If the highest prediction probability of the feature point is less than the preset threshold, it is regarded as an invalid feature point whose contribution to fine-grained classification is less than a preset value, and the invalid feature point is filtered out.
5. The fine-grained image recognition and classification method according to claim 1, characterized in that: The step of importing the second target feature map and the initial prediction value not processed by the activation function into the neural decision forest module of the visual model, obtaining the target prediction value corresponding to each granularity and storing them in the output sequence includes: The neural decision forest module is composed of multiple neural decision trees, each of which is composed of multiple nodes, and each node hierarchically routes input features through a hierarchical structure so that the visual model captures fine-grained information in the second target feature map.
6. The fine-grained image recognition and classification method according to claim 1, characterized in that: The step of importing the output sequence into the trusted multi-view classification module of the visual model, fusing the predicted value distributions of multiple opinions, and obtaining the final fused classification result comprises: importing the output sequence into a trusted multi-view classification module of the vision model; In the credible multi-view classification module, the category probability distribution is modeled by variational Dirichlet distribution, and Dempster-Shafer is used to integrate evidence to fuse the predicted value distributions of multiple opinions to obtain the final fused classification result.
7. A fine-grained image recognition and classification system, characterized in that: The method applied to any one of claims 1 to 6, wherein the system comprises: An image data import module, used to obtain an image data set corresponding to a pre-collected fine-grained image, and import the image data set into a Transformer-based visual model; A shallow information fusion module, used for importing the feature map of each stage into the feature pyramid network module of the visual model to obtain a first target feature map that fuses the shallow information; A pixel feature screening module, used for importing the first target feature map into the weakly supervised sampling module of the visual model, screening out target features that are useful for classification, and stacking the target features into a target vector and a second target feature map based on the target features; A decision module, used for importing the second target feature map and the initial prediction value that has not been processed by the activation function into the neural decision forest module of the visual model, respectively obtaining the target prediction value corresponding to each granularity and storing it in an output sequence; The classification module is used to import the output sequence into the trusted multi-view classification module of the visual model, fuse the predicted value distributions of multiple opinions, and obtain a final fused classification result.
8. The fine-grained image recognition and classification system according to claim 7, characterized in that: The shallow information fusion module is specifically used for: Importing the feature map of each stage in the image data set into the feature pyramid network module of the visual model; The deep features containing semantic information and the shallow features containing spatial information are fused layer by layer through the feature pyramid network module to obtain a first target feature map.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Transform-based fine-grained image classification method
CN118135290A
Visual intention understanding method and system based on uncertainty cross-granularity evidence feature fusion network, and storage medium
CN118196592A