Image classification method and device based on gradient guidance, equipment and storage medium
By employing a gradient-guided image classification method, which utilizes gradient attention to generate and transform heatmaps, and combines a fully connected layer optimization model, the problem of insufficient interpretability in existing image classification technologies is solved, achieving higher accuracy and reliability. This method is applicable to image classification in the medical and financial fields.
Patent Information
- Application Number
- CN202511066177.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing image classification methods suffer from insufficient interpretability in high-risk domains. The generated heatmaps deviate from the actual discriminant regions, the weight updates of fully connected layers lack spatial importance guidance, and the methods fail to effectively utilize the global correlation of high-level semantic features, all of which affect the accuracy and reliability of classification.
This gradient-guided image classification method utilizes gradient attention to generate gradient-weighted class activation heatmaps, which are then converted into attention weights. These heatmaps dynamically modulate feature images and are combined with fully connected layers for classification. A composite loss function is used to optimize the model, enabling visualization and simultaneous optimization of the model's decision-making process.
It improves the accuracy and interpretability of image classification, can more accurately focus on lesion areas, enhances the model's diagnostic and identification accuracy in medical imaging and finance, and improves the model's performance in different application scenarios.
Smart Images

Figure CN120953677A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to gradient-guided image classification methods, apparatus, devices, and storage media. Background Technology
[0002] In recent years, deep learning has made significant progress in image classification tasks, but its "black box" nature limits its application in high-risk domains. To improve model interpretability, researchers have proposed various gradient visualization methods, such as Grad-CAM (Gradient-weighted Class Activation Mapping, a technique for visualizing convolutional neural networks), which generates heatmaps through backpropagation gradients to identify key regions. However, these methods are typically used only as post-processing tools and fail to directly integrate interpretability into the model's optimization objectives.
[0003] In existing technologies, some studies have attempted to combine classification and interpretation tasks through multi-task learning frameworks, or to utilize attention mechanisms to enhance feature selectivity. For example, some studies have proposed spatiotemporal attention networks for interpreting time-series data, but their design does not consider the gradient modulation problem of fully connected layers. Regarding optimization methods, inverse optimization techniques have been used to infer the decision logic of models, but have not yet been systematically applied to enhance the interpretability of image classification.
[0004] In summary, the current methods have three main limitations: (1) The separation of interpretability and classification objectives leads to a deviation between the generated heatmap and the actual discriminant region. For example, in the medical field, this deviation may lead to doctors misjudging the lesion area and affecting the accuracy of diagnosis; (2) As a key component of classification decision, the weight update of the fully connected layer lacks spatial importance guidance. For example, in the financial field, this may lead to the model being insensitive to key features in the image and affecting the accuracy of classification; (3) Existing gradient attention methods are mostly designed for convolutional layers and fail to effectively utilize the global correlation of high-level semantic features. For example, for medical images, this limitation may lead to the model performing poorly when recognizing complex lesion patterns. For images in the financial field, it may ignore key details in the image and affect the reliability of classification. Summary of the Invention
[0005] This invention provides a gradient-guided image classification method, apparatus, computer device, and storage medium, aiming to improve the accuracy and interpretability of image classification.
[0006] In a first aspect, embodiments of the present invention provide a gradient-guided image classification method, including:
[0007] Acquire an input image and extract features from the input image to obtain a feature image;
[0008] Gradient attention is used to generate a gradient-weighted class activation heatmap of the feature image;
[0009] The gradient-weighted activation heatmap is converted into attention weights;
[0010] The attention image is obtained by modulating the feature image using the attention weights;
[0011] The attention image is input into a classification network, and the classification network outputs the corresponding classification result to construct an image classification model.
[0012] The image classification model is used to perform image classification processing on the specified image to be classified.
[0013] Secondly, embodiments of the present invention provide a gradient-guided image classification device, comprising:
[0014] A feature extraction unit is used to acquire an input image and extract features from the input image to obtain a feature image;
[0015] A heatmap generation unit is used to generate a gradient-weighted class activation heatmap from the feature image using gradient attention.
[0016] A heatmap conversion unit is used to convert the gradient-weighted activation heatmap into attention weights.
[0017] An image modulation unit is used to modulate the feature image using the attention weights to obtain an attention image;
[0018] The model building unit is used to input the attention image into the classification network and have the classification network output the corresponding classification result, thereby constructing an image classification model;
[0019] The model classification unit is used to perform image classification processing on a specified image to be classified using the image classification model.
[0020] Thirdly, embodiments of the present invention provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the gradient-guided image classification method as described in the first aspect.
[0021] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the gradient-guided image classification method as described in the first aspect.
[0022] This invention provides a gradient-guided image classification method, apparatus, computer device, and storage medium. The method includes: acquiring an input image and extracting features from the input image to obtain a feature image; generating a gradient-weighted class activation heatmap from the feature image using gradient attention; converting the gradient-weighted class activation heatmap into attention weights; modulating the feature image using the attention weights to obtain an attention image; inputting the attention image into a classification network, and having the classification network output corresponding classification results to construct an image classification model; and using the image classification model to perform image classification processing on a specified image to be classified. This invention first acquires an input image and extracts features to obtain a corresponding feature image. Then, a gradient-weighted class activation heatmap is generated using gradient attention and converted into attention weights. Next, the feature image is modulated using the attention weights to obtain an attention image. Then, the attention image is input into a classification network, and corresponding classification results are output to construct an image classification model. Finally, the model is used to classify a specified image to be classified. This invention deeply integrates a gradient-guided attention mechanism into the classification network structure. By dynamically modulating feature weights, it achieves simultaneous visualization and optimization of the model's decision-making process, thereby improving the accuracy and interpretability of image classification. In practical applications, an end-to-end connection is used to ensure that gradient information flows throughout the entire network, guaranteeing collaborative optimization among components. Attached Figure Description
[0023] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 A flowchart illustrating a gradient-guided image classification method provided in an embodiment of the present invention;
[0025] Figure 2 This is a schematic diagram of a sub-process of a gradient-guided image classification method provided in an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of another sub-process of a gradient-guided image classification method provided in an embodiment of the present invention;
[0027] Figure 4 A schematic block diagram of a gradient-guided image classification device provided in an embodiment of the present invention;
[0028] Figure 5 A schematic block diagram of a gradient-guided image classification device provided in an embodiment of the present invention;
[0029] Figure 6 Another schematic block diagram of a gradient-guided image classification device provided in an embodiment of the present invention;
[0030] Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0032] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0033] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0034] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0035] The gradient-guided image classification method provided in this invention can be applied in client-server interaction environments, where the client communicates with the server via a network. The server acquires an input image and extracts features from it to obtain a feature image; it then generates a gradient-weighted class activation heatmap using gradient attention; the heatmap is converted into attention weights; these weights are used to modulate the feature image to obtain an attention image; the attention image is input to a classification network, which outputs the corresponding classification result, thereby constructing an image classification model. The client sends an image to be classified, and the server uses the image classification model to perform image classification processing on the specified image. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. Specific embodiments of the invention are described in detail below.
[0036] Please see below. Figure 1 This invention provides a gradient-guided image classification method, specifically including steps S101 to S106.
[0037] Step S101: Obtain the input image and extract features from the input image to obtain a feature image;
[0038] Step S102: Generate a gradient-weighted class activation heatmap of the feature image using gradient attention;
[0039] Step S103: Convert the gradient-weighted class activation heatmap into attention weights;
[0040] Step S104: Modulate the feature image using the attention weights to obtain an attention image;
[0041] Step S105: Input the attention image into the classification network, and have the classification network output the corresponding classification result to construct an image classification model;
[0042] Step S106: Perform image classification processing on the specified image to be classified using the image classification model.
[0043] In this embodiment, the input image is first acquired and its features are extracted to obtain the corresponding feature image. Next, a gradient-weighted class activation heatmap is generated using gradient attention and converted into attention weights. Then, the feature image is modulated using these attention weights to obtain an attention image. This attention image is then input into a classification network, which outputs the corresponding classification result, thus constructing an image classification model. Finally, this model is used to classify the specified image to be classified.
[0044] This embodiment deeply integrates a gradient-guided attention mechanism into the classification network structure. By dynamically modulating feature weights, it achieves simultaneous visualization and optimization of the model's decision-making process, thereby improving the accuracy and interpretability of image classification. In practical applications, an end-to-end connection is adopted to ensure that gradient information flows throughout the entire network, guaranteeing collaborative optimization among components.
[0045] Specifically, the gradient-guided image classification method provided in this embodiment is applicable to the medical field. When processing medical images, it can more accurately focus on lesion areas, thereby improving the accuracy and reliability of diagnosis. For example, when identifying nodules in lung CT images, the image classification method provided in this embodiment can highlight the nodule area, assisting doctors in making accurate diagnoses. Simultaneously, the gradient-guided image classification method provided in this embodiment is applicable to the financial field. For example, when identifying key information in financial instruments, such as signatures, amounts, or dates, this method can accurately emphasize these key features, improving the accuracy and efficiency of classification. Furthermore, this method is also applicable to the security monitoring field. By accurately classifying people or objects in surveillance videos, such as identifying abnormal behavior or specific individuals, it can effectively improve public safety levels. Therefore, by combining gradient guidance and attention mechanisms, this embodiment not only enhances the interpretability of image classification but also significantly improves the model's performance in different application scenarios.
[0046] In one embodiment, step S101 includes:
[0047] The input image is processed using a standard CNN architecture to extract multi-level image features to obtain the feature image.
[0048] This embodiment uses a convolutional neural network to extract multi-level image features from the input image. For example, it employs convolutional kernels of different scales to extract local and global features, enhancing the richness and robustness of feature representation. The feature image retains key information from the input image, providing a solid foundation for subsequent steps. Through this step, the input image is transformed into a representation in a high-dimensional feature space, facilitating further processing and analysis by subsequent modules. In a specific embodiment, the standard CNN architecture can be a ResNet50 neural network, using ResNet50 to extract deep features from the input image. This provides strong support for subsequent gradient attention generation and attention weighting. Furthermore, the residual connection structure of ResNet50 helps alleviate the vanishing gradient problem in deep neural networks, enabling the network to be trained and optimized more effectively.
[0049] In one embodiment, such as Figure 2As shown, step S102 includes steps S201 to S202.
[0050] Step S201: Calculate the gradient weight of each channel of the feature image according to the following formula:
[0051]
[0052] Where, α c Let H represent the gradient weights, W represent the height of the feature image, c represent the channel class of the feature image, and y represent the gradient weights. c Let i represent the predicted score for category c, and i and j represent the i-th and j-th scores, respectively.
[0053] Step S202: Perform channel summation on the feature image using the gradient weights according to the following formula to obtain the gradient-weighted class activation heatmap:
[0054]
[0055] Where M represents the gradient-weighted activation heatmap, ReLU represents the ReLU activation function, and Fc represents the feature map of the c-th channel.
[0056] In this embodiment, the input image is set as... After multi-level feature extraction, the feature image is obtained. Therefore, when calculating the gradient weight of each channel for the target category c, it can be calculated according to the following formula:
[0057]
[0058] Based on the obtained gradient weights, the heatmap The result is obtained by channel-weighted summation followed by ReLU activation:
[0059]
[0060] This heatmap details the contribution of each region to the classification decision at different spatial locations. Specifically, each point in the heatmap corresponds to a specific region in the image, and the intensity of the point's color visually indicates the region's importance in the classification process. Darker points mean that the region plays a more crucial role in the final classification decision, and the features it contains have a greater impact on the classification model's judgment. Conversely, lighter points indicate that the region's contribution to the classification decision is relatively smaller, and its impact on the model's final judgment is more limited. This visualization method clearly identifies which regions are key factors in the classification decision, thereby enabling a better understanding and optimization of the model's performance.
[0061] This embodiment calculates a Grad-CAM heatmap (i.e., a gradient-weighted class activation heatmap) and obtains the gradient weights of each channel through backpropagation, effectively revealing the contribution of different regions in the image to the classification results. This step not only enhances the interpretability of the model but also provides crucial information for the subsequent generation of attention weights. By accurately locating important feature regions in the image, the model can focus more on these key details, thereby reducing unnecessary information interference and further improving the accuracy and robustness of classification.
[0062] In one embodiment, step S103 includes:
[0063] The gradient-weighted activation heatmap is converted into attention weights using a projection function according to the following formula:
[0064] φ(M)=Softmax(Linear(Flatten(M)));
[0065] Where φ(M) represents the attention weight, M represents the gradient-weighted class activation heatmap, Softmax represents the Softmax activation function, latten represents the flattening operation, and Linear represents the learnable linear transformation layer.
[0066] This embodiment converts the heatmap M into attention weights using a projection function φ. The Flatten operation flattens M into an HW-dimensional vector, and the Linear layer is a learnable linear transformation layer used to map the flattened vector to the attention weight space. The Softmax activation function ensures that the sum of all attention weights is 1, thus making the weight at each position represent relative importance. By converting the two-dimensional gradient-weighted activation heatmap into attention weights for the feature vector, this embodiment effectively transforms the gradient-weighted activation heatmap into spatial attention weights, providing crucial guidance for subsequent feature modulation.
[0067] In one embodiment, such as Figure 3 As shown, step S104 includes steps S301 to S303.
[0068] Step S301: Perform spatial average pooling on the feature image to obtain an initial feature vector;
[0069] Step S302: Multiply the attention weights and the initial feature vector element-wise according to the following formula to obtain the attention feature vector:
[0070] x attn =x⊙φ(M);
[0071] Where, x attnLet represent the attention feature vector, x represent the initial feature vector, and φ(M) represent the attention weights;
[0072] Step S303: Modulate the feature image using the attention feature vector to obtain the attention image.
[0073] This embodiment first performs spatial average pooling on the feature image to obtain an initial feature vector. Then, the initial feature vector is modulated using the obtained attention weights to obtain a modulated attention feature vector. Subsequently, based on the extracted attention feature vector, the feature image undergoes detailed modulation processing. This modulation process aims to generate an attention image specifically for input into the classification network. In this way, the classification network can dynamically adjust itself according to the spatial importance of each feature when receiving input features. This dynamic adjustment mechanism effectively highlights the features of key regions in the image, significantly enhancing the contribution of these key features in the classification decision process, thereby improving the overall classification accuracy and robustness.
[0074] In one embodiment, step S104 includes:
[0075] The attention image is classified using a fully connected layer to obtain the corresponding classification result.
[0076] This embodiment employs a fully connected layer to achieve the final classification prediction, thus fully utilizing key information in the attention image and improving classification accuracy. By using a fully connected layer, this embodiment comprehensively considers the global features of the attention image, making the classification results more reliable. Furthermore, by combining gradient-guided attention mechanisms with fully connected layers, this embodiment also achieves model interpretability, making the classification decision process more transparent and helping users understand and trust the model's output. In practical implementation, the parameters of the fully connected layer can be optimized using the backpropagation algorithm to ensure that the overall performance of the model reaches its optimal level.
[0077] In other embodiments, other classifiers, such as support vector machines or decision trees, can be used to classify the attention images. However, using fully connected layers as classifiers generally yields better classification results because they can comprehensively consider the global features of the attention images and learn more complex classification boundaries. Furthermore, fully connected layers can be combined with gradient-guided attention mechanisms to improve model interpretability, which is important for users in practical applications.
[0078] In one embodiment, the gradient-guided image classification method further includes:
[0079] A composite loss function is constructed by combining cross-entropy loss and gradient attention loss according to the following formula:
[0080]
[0081] in, Represents the loss function. Represents cross-entropy loss, Let M represent gradient attention loss, M represent gradient-weighted class activation heatmap, and B represent the true discrimination region of the input image;
[0082] The composite loss function is used to optimize the parameters of the image classification model.
[0083] This embodiment employs a composite loss function with two components to train and optimize the image classification model. Cross-entropy loss measures the difference between the model's predictions and the true labels, ensuring accurate classification. Gradient attention loss constrains the consistency between the gradient-weighted class activation heatmap and the true discriminant regions of the input image, thereby improving the model's interpretability. Through backpropagation, the model parameters are continuously optimized, causing the loss function to gradually converge to its minimum, resulting in a high-performance image classification model. Furthermore, the joint optimization of the composite loss function allows the image classification model to simultaneously learn accurate classification boundaries and attention distributions consistent with human cognition. In practice, λ is used as a hyperparameter to balance the relative importance of cross-entropy loss and gradient attention loss; therefore, by adjusting the value of λ, a good trade-off between classification accuracy and interpretability can be achieved.
[0084] In summary, compared to traditional methods, the gradient-guided image classification method provided in this invention has the following three main characteristics: First, it innovatively advances the gradient heatmap generation process from the post-processing stage to the model training stage; second, it designs a differentiable heatmap-attention conversion mechanism, enabling gradient information to directly guide the parameter updates of fully connected layers; and finally, by introducing a composite loss function, it achieves joint optimization of classification accuracy and interpretability. This design not only ensures the high performance of the model but also allows its decision-making basis to be presented in an intuitive heatmap format, greatly improving the ease of understanding and use.
[0085] Figure 4 A schematic block diagram of a gradient-guided image classification device 400 provided in an embodiment of the present invention, the device 400 comprising:
[0086] The feature extraction unit 401 is used to acquire an input image and extract features from the input image to obtain a feature image;
[0087] Heatmap generation unit 402 is used to generate a gradient-weighted class activation heatmap from the feature image using gradient attention;
[0088] Heatmap conversion unit 403 is used to convert the gradient-weighted activation heatmap into attention weights;
[0089] Image modulation unit 404 is used to modulate the feature image using the attention weights to obtain an attention image;
[0090] The model building unit 405 is used to input the attention image into the classification network and have the classification network output the corresponding classification result to build an image classification model.
[0091] The model classification unit 406 is used to perform image classification processing on a specified image to be classified using the image classification model.
[0092] In this embodiment, the input image is first acquired and its features are extracted to obtain the corresponding feature image. Next, a gradient-weighted class activation heatmap is generated using gradient attention and converted into attention weights. Then, the feature image is modulated using these attention weights to obtain an attention image. This attention image is then input into a classification network, which outputs the corresponding classification result, thus constructing an image classification model. Finally, this model is used to classify the specified image to be classified.
[0093] This embodiment deeply integrates a gradient-guided attention mechanism into the classification network structure. By dynamically modulating feature weights, it achieves simultaneous visualization and optimization of the model's decision-making process, thereby improving the accuracy and interpretability of image classification. In practical applications, an end-to-end connection is adopted to ensure that gradient information flows throughout the entire network, guaranteeing collaborative optimization among components.
[0094] Specifically, the gradient-guided image classification method provided in this embodiment is applicable to the medical field. When processing medical images, it can more accurately focus on lesion areas, thereby improving the accuracy and reliability of diagnosis. For example, when identifying nodules in lung CT images, the image classification method provided in this embodiment can highlight the nodule area, assisting doctors in making accurate diagnoses. Furthermore, in the financial field, the gradient-guided image classification method provided in this embodiment can also effectively identify key features in images, such as facial features and key information in documents, thereby improving the accuracy of identity verification and document review, and enhancing the accuracy and robustness of classification. Therefore, by combining gradient guidance and attention mechanisms, this embodiment not only enhances the interpretability of image classification but also significantly improves the model's performance in different application scenarios.
[0095] In one embodiment, the feature extraction unit 401 includes:
[0096] A multi-level extraction unit is used to extract multi-level image features from the input image using a standard CNN architecture to obtain the feature image.
[0097] This embodiment uses a convolutional neural network to extract multi-level image features from the input image. For example, it employs convolutional kernels of different scales to extract local and global features, enhancing the richness and robustness of feature representation. The feature image retains key information from the input image, providing a solid foundation for subsequent steps. Through this step, the input image is transformed into a representation in a high-dimensional feature space, facilitating further processing and analysis by subsequent modules. In a specific embodiment, the standard CNN architecture can be a ResNet50 neural network, using ResNet50 to extract deep features from the input image. This provides strong support for subsequent gradient attention generation and attention weighting. Furthermore, the residual connection structure of ResNet50 helps alleviate the vanishing gradient problem in deep neural networks, enabling the network to be trained and optimized more effectively.
[0098] In one embodiment, such as Figure 5 As shown, the heatmap generation unit 402 includes:
[0099] Weight calculation unit 501 is used to calculate the gradient weight of each channel of the feature image according to the following formula:
[0100]
[0101] Where, α c Let H represent the gradient weights, W represent the height of the feature image, c represent the channel class of the feature image, and y represent the gradient weights. c Let i represent the predicted score for category c, and i and j represent the i-th and j-th scores, respectively.
[0102] The channel summation unit 502 is used to perform channel summation processing on the feature image using the gradient weights according to the following formula to obtain the gradient-weighted class activation heatmap:
[0103]
[0104] Where M represents the gradient-weighted activation heatmap, ReLU represents the ReLU activation function, and Fc represents the feature map of the c-th channel.
[0105] In this embodiment, the input image is set as... After multi-level feature extraction, the feature image is obtained. Therefore, when calculating the gradient weight of each channel for the target category c, it can be calculated according to the following formula:
[0106]
[0107] Based on the obtained gradient weights, the heatmap The result is obtained by channel-weighted summation followed by ReLU activation:
[0108]
[0109] This heatmap details the contribution of each region to the classification decision at different spatial locations. Specifically, each point in the heatmap corresponds to a specific region in the image, and the intensity of the point's color visually indicates the region's importance in the classification process. Darker points mean that the region plays a more crucial role in the final classification decision, and the features it contains have a greater impact on the classification model's judgment. Conversely, lighter points indicate that the region's contribution to the classification decision is relatively smaller, and its impact on the model's final judgment is more limited. This visualization method clearly identifies which regions are key factors in the classification decision, thereby enabling a better understanding and optimization of the model's performance.
[0110] This embodiment calculates a Grad-CAM heatmap (i.e., a gradient-weighted class activation heatmap) and obtains the gradient weights of each channel through backpropagation, effectively revealing the contribution of different regions in the image to the classification results. This step not only enhances the interpretability of the model but also provides crucial information for the subsequent generation of attention weights. By accurately locating important feature regions in the image, the model can focus more on these key details, thereby reducing unnecessary information interference and further improving the accuracy and robustness of classification.
[0111] In one embodiment, the heatmap conversion unit 403 includes:
[0112] The projection transformation unit is used to convert the gradient-weighted activation heatmap into attention weights using a projection function according to the following formula:
[0113] φ(M)=Softmax(Linear(Flatten(M)));
[0114] Where φ(M) represents the attention weight, M represents the gradient-weighted class activation heatmap, Softmax represents the Softmax activation function, latten represents the flattening operation, and Linear represents the learnable linear transformation layer.
[0115] This embodiment converts the heatmap M into attention weights using a projection function φ. The Flatten operation flattens M into an HW-dimensional vector, and the Linear layer is a learnable linear transformation layer used to map the flattened vector to the attention weight space. The Softmax activation function ensures that the sum of all attention weights is 1, thus making the weight at each position represent relative importance. By converting the two-dimensional gradient-weighted activation heatmap into attention weights for the feature vector, this embodiment effectively transforms the gradient-weighted activation heatmap into spatial attention weights, providing crucial guidance for subsequent feature modulation.
[0116] In one embodiment, such as Figure 6 As shown, the image modulation unit 404 includes:
[0117] The average pooling unit 601 is used to perform spatial average pooling on the feature image to obtain an initial feature vector.
[0118] Vector calculation unit 602 is used to perform element-wise multiplication of the attention weights and the initial feature vector according to the following formula to obtain the attention feature vector:
[0119] x attn =x⊙φ(M);
[0120] Where, x attn Let represent the attention feature vector, x represent the initial feature vector, and φ(M) represent the attention weights;
[0121] The vector modulation unit 603 is used to modulate the feature image using the attention feature vector to obtain the attention image.
[0122] This embodiment first performs spatial average pooling on the feature image to obtain an initial feature vector. Then, the initial feature vector is modulated using the obtained attention weights to obtain a modulated attention feature vector. Subsequently, based on the extracted attention feature vector, the feature image undergoes detailed modulation processing. This modulation process aims to generate an attention image specifically for input into the classification network. In this way, the classification network can dynamically adjust itself according to the spatial importance of each feature when receiving input features. This dynamic adjustment mechanism effectively highlights the features of key regions in the image, significantly enhancing the contribution of these key features in the classification decision process, thereby improving the overall classification accuracy and robustness.
[0123] In one embodiment, the model building unit 405 includes:
[0124] The image classification unit is used to classify the attention image using a fully connected layer to obtain the corresponding classification result.
[0125] This embodiment employs a fully connected layer to achieve the final classification prediction, thus fully utilizing key information in the attention image and improving classification accuracy. By using a fully connected layer, this embodiment comprehensively considers the global features of the attention image, making the classification results more reliable. Furthermore, by combining gradient-guided attention mechanisms with fully connected layers, this embodiment also achieves model interpretability, making the classification decision process more transparent and helping users understand and trust the model's output. In practical implementation, the parameters of the fully connected layer can be optimized using the backpropagation algorithm to ensure that the overall performance of the model reaches its optimal level.
[0126] In other embodiments, other classifiers, such as support vector machines or decision trees, can be used to classify the attention images. However, using fully connected layers as classifiers generally yields better classification results because they can comprehensively consider the global features of the attention images and learn more complex classification boundaries. Furthermore, fully connected layers can be combined with gradient-guided attention mechanisms to improve model interpretability, which is important for users in practical applications.
[0127] In one embodiment, the gradient-guided image classification device 400 further includes:
[0128] The loss building unit is used to construct a composite loss function by combining cross-entropy loss and gradient attention loss according to the following formula:
[0129]
[0130] in, Represents the loss function. Represents cross-entropy loss, Let M represent gradient attention loss, M represent gradient-weighted class activation heatmap, and B represent the true discrimination region of the input image;
[0131] The model optimization unit is used to optimize the parameters of the image classification model using the composite loss function.
[0132] This embodiment employs a composite loss function with two components to train and optimize the image classification model. Cross-entropy loss measures the difference between the model's predictions and the true labels, ensuring accurate classification. Gradient attention loss constrains the consistency between the gradient-weighted class activation heatmap and the true discriminant regions of the input image, thereby improving the model's interpretability. Through backpropagation, the model parameters are continuously optimized, causing the loss function to gradually converge to its minimum, resulting in a high-performance image classification model. Furthermore, the joint optimization of the composite loss function allows the image classification model to simultaneously learn accurate classification boundaries and attention distributions consistent with human cognition. In practice, λ is used as a hyperparameter to balance the relative importance of cross-entropy loss and gradient attention loss; therefore, by adjusting the value of λ, a good trade-off between classification accuracy and interpretability can be achieved.
[0133] In summary, compared to traditional methods, the gradient-guided image classification method provided in this invention has the following three main characteristics: First, it innovatively advances the gradient heatmap generation process from the post-processing stage to the model training stage; second, it designs a differentiable heatmap-attention conversion mechanism, enabling gradient information to directly guide the parameter updates of fully connected layers; and finally, by introducing a composite loss function, it achieves joint optimization of classification accuracy and interpretability. This design not only ensures the high performance of the model but also allows its decision-making basis to be presented in an intuitive heatmap format, greatly improving the ease of understanding and use.
[0134] Please see Figure 7 , Figure 7 This is a schematic block diagram of a computer device provided in an embodiment of the present invention. The computer device is equipped with both wireless and wired communication capabilities.
[0135] The computer device includes a processor 702, a memory, and a network interface 705 connected via a system bus 701. The memory may include a non-volatile storage medium 703 and internal memory 704.
[0136] The non-volatile storage medium 703 may store an operating system 7031 and a computer program 7032. When the computer program 7032 is executed, it causes the processor 702 to execute a gradient-guided image classification method.
[0137] The processor 702 provides computing and control capabilities to support the operation of the entire computer device.
[0138] The internal memory 704 provides an environment for the execution of the computer program 7032 in the non-volatile storage medium 703. When the computer program 7032 is executed by the processor 702, the processor 702 can execute a gradient-guided image classification method.
[0139] This network interface 705 is used for network communication with other devices. Those skilled in the art will understand that... Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0140] The processor 702 is used to run a computer program 7032 stored in a memory to implement any embodiment of the gradient-guided image classification method described above.
[0141] It should be understood that, in this embodiment of the invention, the processor 702 may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0142] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0143] Acquire an input image and extract features from the input image to obtain a feature image;
[0144] Gradient attention is used to generate a gradient-weighted class activation heatmap of the feature image;
[0145] The gradient-weighted activation heatmap is converted into attention weights;
[0146] The attention image is obtained by modulating the feature image using the attention weights;
[0147] The attention image is input into a classification network, and the classification network outputs the corresponding classification result to construct an image classification model.
[0148] The image classification model is used to perform image classification processing on the specified image to be classified.
[0149] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed, can perform the steps provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0150] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0151] Acquire an input image and extract features from the input image to obtain a feature image;
[0152] Gradient attention is used to generate a gradient-weighted class activation heatmap of the feature image;
[0153] The gradient-weighted activation heatmap is converted into attention weights;
[0154] The attention image is obtained by modulating the feature image using the attention weights;
[0155] The attention image is input into a classification network, and the classification network outputs the corresponding classification result to construct an image classification model.
[0156] The image classification model is used to perform image classification processing on the specified image to be classified.
[0157] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0158] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0159] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0160] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A gradient-guided image classification method, characterized in that, include: Acquire an input image and extract features from the input image to obtain a feature image; Gradient attention is used to generate a gradient-weighted class activation heatmap of the feature image; The gradient-weighted activation heatmap is converted into attention weights; The attention image is obtained by modulating the feature image using the attention weights; The attention image is input into a classification network, and the classification network outputs the corresponding classification result to construct an image classification model. The image classification model is used to perform image classification processing on the specified image to be classified.
2. The gradient-guided image classification method according to claim 1, characterized in that, The process of acquiring an input image and extracting features from the input image to obtain a feature image includes: The input image is processed using a standard CNN architecture to extract multi-level image features to obtain the feature image.
3. The gradient-guided image classification method according to claim 1, characterized in that, The step of generating a gradient-weighted class activation heatmap of the feature image using gradient attention includes: The gradient weight of each channel of the feature image is calculated according to the following formula: Where, α c Let H represent the gradient weights, W represent the height of the feature image, c represent the channel class of the feature image, and y represent the gradient weights. c Let i represent the predicted score for category c, and i and j represent the i-th and j-th scores, respectively. The gradient-weighted class activation heatmap is obtained by performing channel summation on the feature image using the gradient weights according to the following formula: Where M represents the gradient-weighted activation heatmap, ReLU represents the ReLU activation function, and Fc represents the feature map of the c-th channel.
4. The gradient-guided image classification method according to claim 1, characterized in that, The step of converting the gradient-weighted activation heatmap into attention weights includes: The gradient-weighted activation heatmap is converted into attention weights using a projection function according to the following formula: φ(M)=Softmax(Linear(Flatten(M))); Where φ(M) represents the attention weight, M represents the gradient-weighted class activation heatmap, Softmax represents the Softmax activation function, latten represents the flattening operation, and Linear represents the learnable linear transformation layer.
5. The gradient-guided image classification method according to claim 1, characterized in that, The process of modulating the feature image using the attention weights to obtain an attention image includes: The feature image is subjected to spatial average pooling to obtain an initial feature vector; The attention feature vector is obtained by element-wise multiplying the attention weights and the initial feature vector according to the following formula: x attn =x⊙φ(M); Where, x attn Let represent the attention feature vector, x represent the initial feature vector, and φ(M) represent the attention weights; The attention image is obtained by modulating the feature image using the attention feature vector.
6. The gradient-guided image classification method according to claim 1, characterized in that, The step of inputting the attention image into a classification network and having the classification network output the corresponding classification result to construct an image classification model includes: The attention image is classified using a fully connected layer to obtain the corresponding classification result.
7. The gradient-guided image classification method according to claim 1, characterized in that, Also includes: A composite loss function is constructed by combining cross-entropy loss and gradient attention loss according to the following formula: in, Represents the loss function. Represents cross-entropy loss, Let M represent gradient attention loss, M represent gradient-weighted class activation heatmap, and B represent the true discrimination region of the input image; The composite loss function is used to optimize the parameters of the image classification model.
8. A gradient-guided image classification device, characterized in that, include: A feature extraction unit is used to acquire an input image and extract features from the input image to obtain a feature image; A heatmap generation unit is used to generate a gradient-weighted class activation heatmap from the feature image using gradient attention. A heatmap conversion unit is used to convert the gradient-weighted activation heatmap into attention weights. An image modulation unit is used to modulate the feature image using the attention weights to obtain an attention image; The model building unit is used to input the attention image into the classification network and have the classification network output the corresponding classification result, thereby constructing an image classification model; The model classification unit is used to perform image classification processing on a specified image to be classified using the image classification model.
9. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the gradient-guided image classification method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the gradient-guided image classification method as described in any one of claims 1 to 7.