An Image Classification Result Feature Visualization Method Based on the Upsampling Mechanism and Class Activation Mapping

By adopting upsampling mechanism and class activation mapping in convolutional neural networks, high-resolution significant graphs are generated, which solves the problem of the lack of interpretability of convolutional neural network models in picture classification tasks, and achieves higher feature visualization accuracy and model interpretability.

CN116468941BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310400157.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-05-27
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

The existing convolutional neural network models lack interpretability in picture classification tasks, making it difficult to explain the decisions made by the model, especially in the application of the security and controllable field.

Method used

A feature visualization method of image classification results based on upsampling mechanism and class activation mapping is adopted to obtain higher resolution feature maps through multi-size upsampling, and a significant map is generated using class activation mapping and mask operations to visualize model decisions.

Benefits of technology

It improves the resolution and accuracy of feature visualization, enhances the interpretability of the model, reduces the cost of perturbation calculation, and is suitable for areas such as image classification, weakly supervised positioning and semantic segmentation of image objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116468941B_ABST
    Figure CN116468941B_ABST
Patent Text Reader

Abstract

The present invention discloses an image classification result feature visualization method based on an upsampling mechanism and class activation mapping. This method uses the activation map and gradient matrix of the last layer of a convolutional neural network. By magnifying the original input image at multiple scales, activation maps and gradient matrices with different resolutions are obtained; then they are fused and weighted and added together to obtain an initial mask. The mask after multi-scale fusion has richer feature information; the normalized mask is directly multiplied point by point with the input image to perturb the input image; the perturbed input image is fed into the model to obtain the corresponding class probability scores of each mask as weights, and finally the weights and masks are linearly added and combined to obtain the feature visualization result. The present invention is applied to an image classification neural network containing convolutional layers, and can present a feature visualization effect with less noise, higher resolution, and more accurate feature localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of feature visualization of computer vision and deep neural networks, and in particular to an image classification result feature visualization algorithm based on an upsampling mechanism and class activation mapping. Background Art

[0002] The statements in this section only provide background information related to the present disclosure, and these statements may constitute prior art. In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art.

[0003] At present, convolutional neural networks have made revolutionary progress in tasks such as image classification, object detection, and semantic segmentation with their superior performance. However, how to explain the decisions made by neural networks has become a difficult problem under research. Since the current models based on convolutional neural networks have many hidden layers and a huge number of parameters, it is impossible for humans to track the specific process from neural network input to output, which makes them lack interpretability, limiting the application of such deep neural networks in fields that focus on safety and control. In recent years, interpretable research on convolutional neural networks has gradually emerged. One of the research methods is to use saliency maps to intuitively present the importance of the contribution of each pixel in the input image to the prediction results of the image classification neural network.

[0004] There are many types of techniques for feature visualization of convolutional neural network classification results. The back-propagation-based method formulates a special attribution method to start from the output end and back-propagate through all the hidden layers in the middle to the input end, thereby associating the output result with each pixel in the input and assigning a contribution value to each pixel. However, this type of method is not reliable enough and is not even sensitive to model parameters. The perturbation-based method regards the model as a black box, focusing only on the input and output. By covering part of the input image with different masks and observing the changes in the model output, it finds the most important area in the input image for the model. This type of method is computationally expensive and lacks explanation of the model. The class activation mapping-based method uses the feature map of the convolution layer and obtains a saliency map by finding different weighting methods to weight the extracted feature maps. For example, the patent application number CN202210693405.2 is named "Image Processing Method, Device, Storage Medium and Electronic Device". This type of method uses data inside the neural network and is category-sensitive, so it has good interpretability. However, such methods usually select the convolutional layer close to the output end that contains rich category semantic information as the source of the feature map, which also leads to low resolution of the feature map. The final saliency map is obtained through upsampling interpolation and cannot provide detailed feature information. Summary of the invention

[0005] In view of the above problems, the object of the present invention is to solve part of the problems in the prior art, or at least alleviate these problems.

[0006] A method for visualizing image classification result features based on upsampling mechanism and class activation mapping, comprising the following steps:

[0007] Upsample the original input image to multiple sizes to obtain several input images with different resolutions;

[0008] Input the input picture to the image classification model, save the picture whose maximum category index in the model discrimination picture is the specified category index, and perform forward propagation to obtain a feature map;

[0009] Extracting a feature map set corresponding to the feature map from a specified target convolutional layer;

[0010] Back-propagating the feature map for a specified target category index, and extracting a set of gradient matrices corresponding to the feature map from a specified target convolutional layer;

[0011] Upsampling the feature map set and gradient matrix set to the original input image resolution, and accumulating and averaging them in channel order to obtain a fused feature map and a fused gradient matrix;

[0012] The fused gradient matrix is ​​globally averaged and pooled as a weight;

[0013] Multiplying the weight and the fused feature map in channel order to obtain an initial mask;

[0014] The initial masks are grouped according to the channel adjacency principle and accumulated to obtain masks, and the masks are normalized;

[0015] The normalized mask is superimposed on the original input image to perturb it and obtain a perturbed image.

[0016] Input the perturbed image into an image classification model to obtain a probability score and a mask weight of a specified category index;

[0017] The mask weight and the normalized mask are linearly weighted and combined, and then normalized to obtain a saliency map.

[0018] Furthermore, the original input image is upsampled to multiple sizes, including the following steps:

[0019] Select the original input image and adjust it to have the same length and width;

[0020] The original input image is upsampled multiple times, and each upsampling increases the length and width by x pixels, until the last upsampling resolution reaches ζ max =(Hmax ,W max )stop;

[0021] The upsampling process follows Among them, 0 =(H 0 ,W 0 ) is the resolution of the initial input image, H 0 and W 0 is the height dimension of the original input image, represents the number of pixels of the image under this dimension, t is the iteration count, N is the maximum number of iterations, ζ max =(H max ,W max ) is the upper sampling limit resolution, H max and W max is the maximum height and width of the upsampled image, representing the number of pixels in the height and width dimensions respectively. The resolution of all upsampled images must be less than or equal to ζ max .

[0022] The back propagation follows the formula in, is the trained image classification convolutional neural network model, I t is the tth upsampled image, isI t The corresponding feature map set, is a collection of gradients.

[0023] Furthermore, the feature map set and the gradient matrix set are upsampled to the original input image resolution, and are accumulated in channel order and averaged, including the following steps:

[0024] The feature map set is upsampled and interpolated to the original input image resolution and accumulated by channel. The formula is: Represents the set of feature maps after accumulation after the tth iteration;

[0025] After accumulation, take the average of each element in each feature map, the formula is:

[0026] The gradient matrix set is upsampled and interpolated to the original input image resolution and accumulated by channel. The formula is: Represents the set of gradient matrices after accumulation after the tth iteration;

[0027] After accumulation, take the average of each element in each gradient matrix. The formula is:

[0028] Among them, the maximum effective number of iterations tmax =Nd, N is the number of input images generated after upsampling the original single input image, N is the maximum number of iterations, and d is the number of images discarded because the image classification model determines that the maximum category index of the image is not the target category index.

[0029] Preferably, the upsampling interpolation function is a bilinear interpolation function.

[0030] Furthermore, the normalized mask is superimposed on the original input image for perturbation to obtain a perturbed image, which includes the following steps:

[0031] Use Gaussian function to transform the original input image I 0 After Gaussian blurring, we get image I′ 0 ;

[0032] The perturbation image I is calculated by the original input image, the normalized mask, and the Gaussian blurred image. r , the formula is:

[0033] I r =I 0 ⊙M′ r +I′ 0 ⊙(1-M′ r )

[0034] The original input picture M′ r is the rth mask M r The normalized mask; ⊙ is the dot multiplication operation, that is, the corresponding pixel positions are directly multiplied.

[0035] Furthermore, the perturbed image is input into an image classification model to obtain a probability score and a mask weight of a specified category index, including the following steps:

[0036] Input all the perturbed images obtained into the image classification model to obtain the probability score corresponding to the specified category index c

[0037] The Gaussian blurred image I′ 0 Input into the image classification model to get the probability score corresponding to the specified category index c

[0038] Calculate the normalized mask M′ r The mask weight

[0039] Furthermore, the mask weight and the normalized mask are linearly weighted and combined, and a saliency map is obtained by normalization, including the following steps:

[0040] The mask weight α r and the normalized mask M′ r Linear weighted combination to obtain the preliminary saliency map S c =ReLU(∑ r a r M r ), where ReLU is the activation function;

[0041] The preliminary saliency map is normalized and converted into pseudo-color to obtain an intuitively displayed saliency map.

[0042] The normalization method is min-max normalization.

[0043] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for visualizing image classification result features based on an upsampling mechanism and class activation mapping.

[0044] The present invention has the following beneficial effects:

[0045] 1. By upsampling the input image at multiple scales and inputting it into the CNN model for forward propagation, higher-resolution feature maps can be extracted. These higher-resolution feature maps can contain richer feature information. This actually utilizes the algorithm principle of the convolutional network and is interpretable in terms of mechanism.

[0046] 2. The present invention uses group sum operation to compress the number of masks, thereby significantly reducing the computational cost of disturbance, and achieves good results between accuracy and computational cost; compared with other visualization methods, it presents a feature visualization effect with less noise, higher resolution, and more accurate feature positioning;

[0047] 3. The present invention can be used to perform feature visualization analysis on the prediction results of image classification convolutional neural networks and analyze the characteristic patterns of neural network learning, thereby helping relevant developers to debug; in addition, the present invention can also be used in weakly supervised positioning, image object semantic segmentation and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A general flowchart of the feature visualization method for image classification results based on upsampling mechanism and class activation mapping;

[0049] Figure 2 A detailed example diagram of the feature visualization method for image classification results based on upsampling mechanism and class activation mapping;

[0050] Figure 3This is a comparison chart of the visualization effects of the present invention, Grad-CAM, Grad-CAM++, XGrad-CAM, Score-CAM, Group-CAM and CAMERAS under test pictures. DETAILED DESCRIPTION

[0051] The present invention is further described below in conjunction with the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention rather than to limit the present invention. Without departing from the technical idea of ​​the present invention, various substitutions and changes can be made according to common technical knowledge and customary means in the field, which should all be included in the scope of the present invention.

[0052] Aiming at the problem of low resolution and poor fineness of the current convolutional neural network feature visualization algorithm, the present invention proposes an image classification result feature visualization algorithm based on upsampling mechanism and class activation mapping.

[0053] The parameters that need to be specified in the feature visualization process of the present invention include the trained convolutional neural network model The discriminant image to be input Category index c, number of groups B, upper sampling limit resolution ζ max =(H max ,W max ), Gaussian blur parameter kernel size ,sigma, maximum number of iterations N, specify convolution layer l, upsampling interpolation function

[0054] like Figure 1 As shown in 2, a method for visualizing image classification result features based on upsampling mechanism and class activation mapping is characterized by comprising the following steps:

[0055] Upsample the original input image to multiple sizes to obtain several input images with different resolutions;

[0056] Input the input picture to the image classification model, save the picture whose maximum category index in the model discrimination picture is the specified category index, and perform forward propagation to obtain a feature map;

[0057] Extracting a feature map set corresponding to the feature map from a specified target convolutional layer;

[0058] Back-propagating the feature map for a specified target category index, and extracting a set of gradient matrices corresponding to the feature map from a specified target convolutional layer;

[0059] Upsampling the feature map set and gradient matrix set to the original input image resolution, and accumulating and averaging them in channel order to obtain a fused feature map and a fused gradient matrix;

[0060] The fused gradient matrix is ​​globally averaged and pooled as a weight;

[0061] Multiplying the weight and the fused feature map in channel order to obtain an initial mask;

[0062] The initial masks are grouped according to the channel adjacency principle and accumulated to obtain masks, and the masks are normalized;

[0063] The normalized mask is superimposed on the original input image to perturb it and obtain a perturbed image.

[0064] Input the perturbed image into an image classification model to obtain a probability score and a mask weight of a specified category index;

[0065] The mask weight and the normalized mask are linearly weighted and combined, and then normalized to obtain a saliency map.

[0066] Upsampling the original input image to multiple sizes includes the following steps:

[0067] Select the original input image and adjust it to have the same length and width;

[0068] The original input image is upsampled multiple times, and each upsampling increases the length and width by x pixels, until the last upsampling resolution reaches ζ max =(H max ,W max )stop;

[0069] The upsampling process follows Among them, 0 =(H 0 ,W 0 ) is the resolution of the initial input image, H 0 and W 0 is the height dimension of the original input image, and represents the number of pixels in the image under this dimension. t is the iteration count, N is the maximum number of iterations, and ζ max =(H max ,W max ) is the upper sampling limit resolution, H max and W max is the maximum height and width of the upsampled image, representing the number of pixels in the height and width dimensions respectively. The resolution of all upsampled images must be less than or equal to ζ max .

[0070] Input all input images into the image classification model, specify the target category index c, discard the images whose maximum category index determined by the model is not the target category index, and retain the other images.

[0071] Select the target convolution layer and extract the feature map from the target convolution layer during the forward propagation process. Each image corresponds to a set of feature maps of one resolution.

[0072] Back propagation is performed for the target category index, and a set of gradient matrices is extracted from the target convolutional layer. Similarly, each image corresponds to a set of gradient matrices.

[0073] The back propagation follows the formula in, is the trained image classification convolutional neural network model, I t is the tth upsampled image, isI t The corresponding feature map set, is a collection of gradients.

[0074] The feature map set and gradient matrix set are upsampled to the original input image resolution, and are accumulated and averaged in channel order, including the following steps:

[0075] The feature map set is upsampled and interpolated to the original input image resolution and accumulated by channel. The formula is: Represents the set of feature maps after accumulation after the tth iteration;

[0076] After accumulation, take the average of each element in each feature map, the formula is:

[0077] The gradient matrix set is upsampled and interpolated to the original input image resolution and accumulated by channel. The formula is: Represents the set of gradient matrices after accumulation after the tth iteration;

[0078] After accumulation, take the average of each element in each gradient matrix. The formula is:

[0079] Among them, the maximum effective number of iterations t max =Nd, N is the number of input images generated after upsampling the original single input image, N is the maximum number of iterations, and d is the number of images discarded because the image classification model determines that the maximum category index of the image is not the target category index.

[0080] The up-sampling interpolation function is a bilinear interpolation function. In the selection of the up-sampling interpolation function, the present invention gives priority to the bilinear interpolation function.

[0081] The gradient matrix is ​​globally averaged pooled, that is, the average value of the gradient matrix of each channel is taken as the weight.

[0082] The weights and feature maps are processed in channel order to obtain the initial mask set.

[0083] The normalized mask is superimposed on the original input image to perform perturbation to obtain a perturbed image, including the following steps:

[0084] Use Gaussian function to transform the original input image I 0 After Gaussian blurring, we get image I′ 0 ;

[0085] The perturbation image I is calculated by the original input image, the normalized mask, and the Gaussian blurred image. r , the formula is:

[0086] I r =I 0 ⊙M′ r +I′ 0 ⊙(1-M′ r )

[0087] The original input picture M′ r is the rth mask M r The normalized mask; ⊙ is the dot multiplication operation, that is, the corresponding pixel positions are directly multiplied.

[0088] Input the perturbed image into the image classification model to obtain the probability score and mask weight of the specified category index, including the following steps:

[0089] Input all the perturbed images obtained into the image classification model to obtain the probability score corresponding to the specified category index c

[0090] The Gaussian blurred image I′ 0 Input into the image classification model to get the probability score corresponding to the specified category index c

[0091] Calculate the normalized mask M′ r The mask weight

[0092] The mask weight and the normalized mask are linearly weighted combined and normalized to obtain a saliency map, including the following steps:

[0093] The mask weight α r and the normalized mask M′ r Linear weighted combination to obtain the preliminary saliency map S c =ReLU(∑ r a r M r), where ReLU is the activation function;

[0094] The preliminary saliency map is normalized and converted into pseudo-color to obtain an intuitively displayed saliency map.

[0095] The normalization method is min-max normalization.

[0096] The image classification model can adopt the CNN model.

[0097] The present invention is achieved through the following specific steps:

[0098] (1) Upsample the input discriminant image N times to obtain N input images. The image resolution obtained by each upsampling is given by the following formula:

[0099]

[0100] where ζ 0 =(H 0 ,W 0 ) is the resolution of the initial input image and t is the iteration count.

[0101] (2) Input the N pictures obtained above into the model one by one, and obtain the output score corresponding to each picture through forward propagation This score is the output before the Softmax layer, and if the corresponding score of category c is the largest among all categories, it is recorded as a valid image. If not, it is an invalid image, and the number of invalid images is recorded as d. Get the corresponding feature map set of each image from the specified convolutional layer l

[0102] (3) For each valid image output in step (2), back propagate for category c to obtain the feature map set in (2) The corresponding gradient matrix set The back propagation algorithm is given by the following formula

[0103]

[0104] (4) Upsample all gradient matrix sets and feature map sets to the resolution of the input image and accumulate them by channel:

[0105]

[0106]

[0107] After accumulation, each element in each feature map and each gradient matrix is ​​divided by t max , where t max =Nd, that is, the total number of pictures minus the number of discarded pictures.

[0108]

[0109]

[0110] (5) Average gradient matrix Applying global average pooling, we get The set of weights corresponding to each feature map in

[0111]

[0112] Where i and j represent the row and column index numbers of each gradient matrix, respectively, and m and n are the length and width of the gradient matrix.

[0113] (6) Set the weights Each weight and feature map set in Each feature map is multiplied in channel order to obtain the initial mask set M:

[0114]

[0115] Where M has K masks, and the size of each mask is consistent with the original input image.

[0116] (7) The initial mask set M is divided into B groups according to the order of adjacent channels, with g masks in each group. The g masks in each group are directly added together to obtain a mask. Assume that the mask M obtained after addition in the rth group is r , its calculation formula is:

[0117]

[0118] Among them, M k is the kth mask in the mask set M. ReLU is an activation function that sets the value less than 0 in the calculation result to 0. This is to filter out meaningless mask values.

[0119] (8) Normalize the mask processed in step (7) and the rth mask M r After normalization, it is recorded as M′ r , the specific normalization method is min-max normalization.

[0120] (9) Use the Gaussian function to transform the input image I 0 Gaussian blur is performed to obtain I′ 0 .

[0121] (10) For the mask obtained in step (8), input image I 0 And the Gaussian blurred image I′ 0, the perturbation image I is obtained through the following calculation r :

[0122] I r =I 0 ⊙M′ r +I′ 0 ⊙(1-M′ r )

[0123] Among them, ⊙ is the dot multiplication operation, that is, the corresponding pixel positions are directly multiplied.

[0124] (11) Input all the perturbation images obtained in (10) into the model and obtain the probability score of the corresponding category c The Gaussian blurred image I′ 0 Input into the model and obtain the probability score of the corresponding category c

[0125] The mask M′ in (8) can be calculated by the following formula: r The weight α r :

[0126]

[0127] (12) Saliency Map S c You can get it by following the steps below:

[0128]

[0129] S c It also needs to be normalized to min-max and converted to pseudo-color before it can be used as a salient image for intuitive display. The present invention is not limited to a single pseudo-color conversion scheme, and other schemes can also be used.

[0130] The model of the present invention can be obtained by directly downloading a pre-trained model or training a model by itself, as long as the model uses a convolutional layer as the main network architecture. Here, the VGG19 model trained on the ImageNet dataset provided in Torchvision is used as the target model for illustration.

[0131] 1. Select the original input image. The original input image should be adjusted to have the same length and width. Upsample the original input image multiple times. Each upsampling increases the length and width by 100 pixels each time until the last upsampling reaches 1000 pixels in length and width.

[0132] 2. Input all the input images in 1 into the VGG19 model, specify the target category index c, discard the images whose maximum category index the model determines is not the target category index, and retain the other images.

[0133] 3. Select the target convolution layer, for example, select the last convolution layer "features.34" of VGG19. Extract the feature map in the forward propagation process from the target convolution layer. Each image corresponds to a set of feature maps of a certain resolution. For specific examples, see the attached Figure 2 The feature map collection in .

[0134] 4. Perform back propagation for the target category index and extract the gradient matrix set from the target convolution layer. Similarly, each image corresponds to a gradient matrix set. See the attached Figure 2 A collection of medium gradient matrices.

[0135] 5. After upsampling all feature map sets and gradient matrix sets to the original input image resolution, they are accumulated in channel order. See the attached Figure 2 The fused gradient matrix and fused feature map.

[0136] 6. Perform global average pooling on the gradient matrix, that is, take the average value of the gradient matrix of each channel as the weight. See the attached Figure 2 The weight vector

[0137] 7. The weights and feature maps are processed in channel order to obtain the initial mask set M.

[0138] 8. The initial masks are grouped according to the channel adjacency principle and accumulated to obtain the masks, and the masks are normalized by min-max.

[0139] 9. Gaussian blur the original input image to get I′ 0 Press I to convert the original input image, mask, and Gaussian blurred image. r =I 0 ⊙M′ r +I′ 0 ⊙(1-M′ r ) to calculate the disturbed image I r , see attached for details Figure 2 Example in .

[0140] 10. Input the perturbed image into the model to obtain the probability score of the target category index and calculate the weight α r .

[0141] 11. Linearly weighted combination of weights and masks and min-max normalization are performed to obtain the saliency map.

[0142] from Figure 3It can be seen that the visualization effect of the present invention under the test picture shows a feature visualization effect with less noise, higher resolution and more accurate feature positioning compared with other mainstream feature visualization algorithms Grad-CAM (gradient-based class activation map), Grad-CAM++ (generalized version of Grad-CAM), XGrad-CAM (axiom-based improved Grad-CAM), Score-CAM (probability score weighted class activation map), Group-CAM (grouping-based improved Score-CAM) and CAMERAS (resolution-enhanced and rationality-preserving class activation map).

[0143] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without creative work should fall within the scope of protection of the present invention.

[0144] The present invention belongs to a feature visualization method combining perturbation and class activation mapping, which utilizes the advantages of the above two methods, and can provide high-resolution and accurate saliency maps at the same time, and can perform feature visualization on any image classification model based on convolutional neural network, and has high resolution and accurate feature positioning capabilities. The present invention can extract higher-resolution feature maps by upsampling the input image in multiple sizes and inputting it into the CNN model for forward propagation. These higher-resolution feature maps can contain richer feature information, which actually utilizes the algorithm principle of the convolutional network and has interpretability from a mechanism perspective; the number of masks is compressed by grouping and taking operations, so that the computational cost of the perturbation can be significantly reduced, and good results are achieved between accuracy and computational cost. The present invention can be used to perform feature visualization analysis on the prediction results of the image classification convolutional neural network and analyze the feature patterns of the neural network learning, thereby helping relevant developers to debug. In addition, the present invention can also be used in the fields of weakly supervised positioning, image object semantic segmentation, etc.

[0145] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for visualizing image classification result features based on an upsampling mechanism and class activation mapping.

Claims

1. An image classification result feature visualization method based on an upsampling mechanism and class activation mapping, characterized in that, it includes the following steps: Upsample the original input image at multiple scales to obtain several input images with different resolutions; Input the input image into an image classification model, save the image with the largest class index in the model's discrimination of the image as the specified class index for forward propagation, and obtain a feature map; Extract the corresponding set of feature maps from the specified target convolutional layer for the feature map; Perform backpropagation on the feature map for the specified target class index, and extract the corresponding set of gradient matrices from the specified target convolutional layer for the feature map; Upsample the set of feature maps and the set of gradient matrices to the resolution of the original input image, accumulate them in the channel order, and take the average to obtain a fused feature map and a fused gradient matrix; Globally average pool the fused gradient matrix to obtain a weight; Multiply the weight and the fused feature map in the channel order to obtain an initial mask; Group and accumulate the initial mask according to the principle of adjacent channels to obtain a mask, and normalize the mask; Overlay the normalized mask on the original input image for perturbation to obtain a perturbed image; Input the perturbed image into the image classification model to obtain the probability score of the specified class index and the mask weight; Linearly weight and combine the mask weight and the normalized mask, and normalize to obtain a saliency map.

2. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 1, characterized in that, Upsampling the original input image at multiple scales includes the following steps: Select the original input image and adjust the original input image to have the same length and width; The original input image is upsampled multiple times. Each time during upsampling, the length and width each increase by x pixel sizes until the resolution of the last upsampling reaches ζ max =(H max ,W max ) Stop; The upsampling process follows where ζ 0 =(H 0 , W 0 ) is the resolution of the initial input image, t is the iteration count, N is the maximum number of iterations, and ζ max =(H max , W max ) is the upsampling upper limit resolution, H max and W max are the height and width dimensions of the final upsampled image, representing the number of pixel points in the height and width dimensions respectively. The resolution of all upsampled images must be less than or equal to ζ max .

3. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 1, characterized in that, The backpropagation follows the formula where is the trained convolutional neural network model for image classification, and I t is the picture after the t-th upsampling, is the set of feature maps corresponding to I t , and is the set of gradients.

4. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 1, characterized in that, Upsampling the set of feature maps and the set of gradient matrices to the resolution of the original input image, accumulating them in the channel order, and taking the average includes the following steps: Upsample and interpolate the set of feature maps to the resolution of the original input image, and accumulate them channel by channel. The formula is as follows: represents the set of feature maps after accumulation after the t-th iteration; After accumulation, take the average of each element in each feature map. The formula is: Upsample and interpolate the gradient matrix set to the resolution of the original input image, and accumulate it channel by channel. The formula is as follows: represents the gradient matrix set after accumulation after the t-th iteration; After accumulation, take the average of each element in each gradient matrix. The formula is: Among them, the maximum effective iteration number t max = N - d, where N is the number of input images generated after upsampling the original single input image, N is also the maximum iteration number, and d is the number of images discarded because the image classification model determines that the maximum class index of the image is not the target class index.

5. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 4, characterized in that, The upsampling interpolation function is a bilinear interpolation function.

6. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 1, characterized in that, Overlaying the normalized mask on the original input image for perturbation to obtain a perturbed image includes the following steps: Perform Gaussian blurring on the original input image I 0 to obtain the image I' 0 ; Calculate the perturbed image I from the original input image, the masked image after normalization, and the image after Gaussian blur r , and the formula is: I r = I 0 ⊙ M′ r + I′ 0 ⊙ (1 - M′ r ) Among them, the original input image M′ r is the r-th mask M r after normalization processing; ⊙ is the dot product operation, that is, directly multiplying the corresponding pixel positions.

7. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 1, characterized in that, Inputting the perturbed image into the image classification model to obtain the probability score of the specified class index and the mask weight includes the following steps: All the obtained perturbed images are input into an image classification model to obtain the probability score corresponding to the specified class index c Input the Gaussian-blurred image I′ 0 into the image classification model to obtain the probability score corresponding to the specified class index c The calculated mask M' after normalization r of the mask weight 8. The image classification result feature visualization method based on an upsampling mechanism and class activation mapping according to claim 1, characterized in that, Linearly weighting and combining the mask weights and the normalized masks, and obtaining a saliency map through normalization, including the following steps: Multiply the mask weight α r and the normalized mask M′ r by linear weighted combination to obtain the preliminary saliency map S c = ReLU(∑ r a r M r ); where ReLU is the activation function; Normalize the preliminary saliency map, and obtain a visually displayed saliency map after pseudo-color conversion.

9. The method for visualizing the features of the image classification result based on the upsampling mechanism and the class activation mapping according to claim 1 or 8, characterized in that the normalization method is min-max normalization.

10. A computer-readable storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method for visualizing the features of the image classification result based on the upsampling mechanism and the class activation mapping according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Image processing method and device, storage medium and electronic equipment

    CN115049886A