Class Activation Map Generation Method, Neural Network Interpretability Method, and Related Devices

By sorting the feature map set, calculating the residual area set and fraction set, and activating function processing, the generated class activation map can more effectively explain the decision-making process of the neural network, improving interpretability.

CN114359588BActive Publication Date: 2025-06-20SHENZHEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111468069.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-06-20
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

In the existing class activation map generation method, the weights between the overlapping areas and non-overlapping areas of the feature map are not much different, resulting in the reduced interpretable effect of the generated class activation map on the neural network.

Method used

By sorting the first feature map set, the second feature map set is obtained, and the residual area set and residual score set corresponding to the second feature map set are calculated. The residual area set and residual score set are activated by the activation function to obtain the class activation map.

Benefits of technology

The interpretability effect of class activation graphs on neural networks is improved, and the interpretability of image processing is enhanced by reducing the weight of residual areas and increasing the weight of the main areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359588B_ABST
    Figure CN114359588B_ABST
Patent Text Reader

Abstract

The present invention discloses a class activation map generation method, a neural network interpretability method and related devices, which are applied to the field of image processing. The class activation map generation method includes: obtaining an original image; processing the original image through the convolutional layer of a neural network to obtain a first feature map set, where the first feature map set is a set of multiple feature maps; reordering the first feature map set to obtain a second feature map set; calculating a residual region set and a residual score set corresponding to the second feature map set; activating the residual region set and the residual score set through an activation function to obtain a class activation map. By assigning weights to the feature maps, the residual score set corresponding to the residual region set is reduced, that is, the weight of the residual region is reduced, and at the same time, the weight of the main region corresponding to the residual region is increased, thereby effectively improving the interpretability effect of the class activation map on the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and particularly to a method for generating class activation maps, a method for interpreting neural networks, and related devices. Background Art

[0002] Neural networks have achieved great success in many fields. However, due to their end-to-end "black box" nature, the knowledge processing and storage mechanisms of the intermediate layers are masked, which allows researchers to only understand and optimize neural networks from a macroscopic perspective, such as convex optimization and receptive fields, to a certain extent affecting their application value. As the most effective way of interpretation technology, various neural network interpretability methods explain the information and features extracted by neural networks from different perspectives. The application of class activation maps is an interpretability method for convolutional neural networks and graph convolutional neural networks. However, in related technologies, the weights between the overlapping regions and non-overlapping regions of the feature maps are not very different, thus reducing the interpretability effect of the generated class activation maps on neural networks. Summary of the Invention

[0003] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.

[0004] Embodiments of the present application provide a method for generating class activation maps, a method for interpreting neural networks, and related devices, which perform weight assignment on feature maps to improve the interpretability effect of class activation maps on neural networks.

[0005] In a first aspect, the present invention provides a method for generating class activation maps, including:

[0006] Obtain an original image;

[0007] Process the original image through the convolutional layer of a neural network to obtain a first set of feature maps, where the first set of feature maps is a set of multiple feature maps;

[0008] Reorder the first set of feature maps according to the importance of the feature maps in the first set of feature maps for a predicted category, where the predicted category is a predetermined category, to obtain a second set of feature maps;

[0009] Calculate a set of residual regions and a set of residual scores corresponding to the second set of feature maps;

[0010] Activate the set of residual regions and the set of residual scores through an activation function to obtain a class activation map.

[0011] The method for generating a class activation map according to an embodiment of the first aspect of the present invention has at least the following beneficial effects: In the embodiment of the present application, the first feature map set is sorted to obtain a second feature map set, and the residual region set and the residual score set corresponding to the second feature map set are calculated, and the class activation map is obtained by activating the residual region set and the residual score set. The second feature map set generated in the embodiment of the present application is sorted according to the importance of the feature map for the predicted class, which is convenient for generating the residual region set of the neural network, so that the residual score set corresponding to the residual region set is reduced, that is, the weight of the residual region is reduced, and the weight of the main region corresponding to the residual region is increased. Compared with the existing neural network interpretability methods, the technical embodiment of the present application effectively improves the interpretability effect of the class activation map on the neural network.

[0012] According to some embodiments of the first aspect of the present invention, re-sorting the first feature map set according to the importance of the feature map in the first feature map set for the predicted class to obtain a second feature map set includes:

[0013] Inverting the first feature map set to obtain a third feature map set;

[0014] According to the first feature map set, the third feature map set and the predicted class, a first target residual score set is obtained, where the first target residual score set includes a plurality of target residual scores, and the number of the target residual scores is equal to the number of the feature maps in the first feature map set;

[0015] Sorting the target residual scores in the first target residual score set in descending order to obtain a second target residual score set;

[0016] Sorting the feature maps in the first feature map set according to the sequence of the target residuals in the second target residual score set to obtain a second feature map set.

[0017] According to some embodiments of the first aspect of the present invention, obtaining the first target residual score set according to the first feature map set, the third feature map set and the predicted class includes:

[0018] Obtaining the feature map in the first feature map set as the first feature map;

[0019] Obtaining the feature map in the third feature map set as the second feature map, where the position of the second feature map in the third feature map set is the same as the position of the first feature map in the first feature map set;

[0020] Passing the original image and the predicted class through the convolutional layer and the fully connected layer in the neural network to obtain a first probability, where the first probability is the probability that the first feature map is judged as the predicted class;

[0021] Perform a Hadamard product operation on the second feature map and the original image to obtain a third feature map;

[0022] Pass the third feature map and the predicted category through the convolutional layer and the fully connected layer to obtain a second probability, where the second probability is the probability that the third feature map is judged to be the predicted category;

[0023] Calculate the difference between the first probability and the second probability to obtain a target residual score;

[0024] Repeat to obtain the feature maps in the first feature map set and the third feature map set, calculate the corresponding target residual scores until no new feature maps can be detected in the first feature map set and the third feature map set, and use the set of multiple calculated target residual scores as the first target residual score set.

[0025] According to some embodiments of the first aspect of the present invention, calculating the residual region set and the residual score set corresponding to the second feature map set includes:

[0026] Obtain the feature map in the second feature map set as a fourth feature map;

[0027] Find the maximum value of the fourth feature map and the feature maps before it to obtain a first saliency map;

[0028] Calculate the residual region and the residual score corresponding to the first saliency map;

[0029] Repeat to obtain the feature maps in the second feature map set, calculate the corresponding residual regions and residual scores until no new feature maps can be detected in the second feature map set, and use the multiple calculated residual regions as the residual region set and the multiple residual scores as the residual score set.

[0030] According to some embodiments of the first aspect of the present invention, calculating the residual region and the residual score corresponding to the first saliency map includes:

[0031] Obtain the feature map at the previous level of the first saliency map as a second saliency map;

[0032] Calculate the difference between the first saliency map and the second saliency map, and use the obtained calculation result as the residual region;

[0033] Invert the first saliency map to obtain a third saliency map;

[0034] Perform a Hadamard product operation on the third saliency map and the original image to obtain a fourth saliency map;

[0035] Obtain a third probability by passing the fourth significant map and the predicted category through the convolutional layer and the fully connected layer, where the third probability is the probability that the fourth significant map is judged to be the predicted category;

[0036] Calculate the difference between the first probability and the third probability to obtain a first region score;

[0037] Calculate the region score corresponding to the second significant map as the second region score;

[0038] Calculate the difference between the first region score and the second region score to obtain a residual score.

[0039] According to some embodiments of the first aspect of the present invention, the activating the residual region set and the residual score set through an activation function to obtain a class activation map includes:

[0040] Obtain a first residual region in the residual region set and a first residual score in the residual score set;

[0041] Multiply the first residual region by the first residual score to obtain a fifth significant map;

[0042] Repeatedly obtain the residual regions in the residual region set and the residual scores in the residual score set, calculate the corresponding significant maps until no new residual regions and residual scores can be detected in the residual region set and the residual score set, and use the sum of the multiple calculated significant maps as a significant map set;

[0043] Activate the significant map set through an activation function to obtain a class activation map.

[0044] According to some embodiments of the first aspect of the present invention, the activation function is a rectified linear unit function.

[0045] According to some embodiments of the first aspect of the present invention, the neural network is a convolutional neural network or a spatio-temporal graph convolutional neural network.

[0046] In a second aspect, the present invention provides a neural network interpretability method, including:

[0047] Obtain an original image and a class activation map, where the class activation map is obtained by passing the original image through a neural network and calculating using the class activation map generation method according to any one of the first aspect;

[0048] Overlay the class activation map with the original image to obtain a target image for interpreting the neural network;

[0049] Output the target image.

[0050] Since the neural network interpretable method in the second aspect applies the class activation map generation method of any item in the first aspect, it has all the beneficial effects of the first aspect of the present invention.

[0051] According to some embodiments of the second aspect of the present invention, when the neural network is a convolutional neural network, the superimposing the class activation map and the original image to obtain a target image for interpreting the neural network includes:

[0052] Converting the class activation map into a pseudo-color image;

[0053] Superimposing the pseudo-color image and the original image by addition to obtain a target image.

[0054] According to some embodiments of the second aspect of the present invention, when the neural network is a spatio-temporal graph convolutional neural network, the superimposing the class activation map and the original image to obtain a target image for interpreting the neural network includes:

[0055] Obtaining a skeleton motion sequence, where the skeleton motion sequence is obtained by the spatio-temporal graph convolutional neural network based on the object to be measured;

[0056] Sampling the skeleton motion sequence frame by frame to obtain skeleton motion vectors;

[0057] Drawing the importance of the skeleton sequence according to the class activation map to obtain a target image, where the target image is used to interpret and display the weight changes of skeleton nodes between different frames.

[0058] In a third aspect, the present invention provides a class activation map generation device, including:

[0059] An image acquisition module for acquiring an original image;

[0060] An image processing module for reordering the first feature map set according to the importance of the feature maps in the first feature map set for the predicted category to obtain a second feature map set;

[0061] An image sorting module for reordering the first feature map set to obtain a second feature map set;

[0062] A region residual score module for calculating a residual region set and a residual score set corresponding to the second feature map set;

[0063] An activation function module for activating the residual region set and the residual score set through an activation function to obtain a class activation map.

[0064] Since the class activation map generation device of the third party can execute the class activation map generation method of any item in the first aspect and / or the neural network interpretability method of any item in the second aspect, it has all the beneficial effects of the first aspect of the present invention.

[0065] Fourthly, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it is like the class activation map generation method of any item in the first aspect and / or like the neural network interpretability method of any item in the second aspect.

[0066] Since the processor in the fourth aspect executes the computer program like the class activation map generation method of any item in the first aspect and / or like the neural network interpretability method of any item in the second aspect when executing the computer program, it has all the beneficial effects of the first aspect of the present invention.

[0067] Fifthly, the present invention provides a computer storage medium, including computer executable instructions stored thereon. The computer executable instructions are used for the class activation map generation method of any item in the first aspect and / or like the neural network interpretability method of any item in the second aspect.

[0068] Since the computer storage medium in the fifth aspect can execute the class activation map generation method of any item in the first aspect and / or like the neural network interpretability method of any item in the second aspect, it has all the beneficial effects of the first aspect of the present invention.

[0069] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings

[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0071] Figure 1 It is a schematic structural diagram of a class activation map generation device provided by an embodiment of the present application;

[0072] Figure 2 It is a main step diagram of a class activation map generation method provided by an embodiment of the present application;

[0073] Figure 3 It is a step diagram of image sorting of a class activation map generation method provided by an embodiment of the present application;

[0074] Figure 4 It is a flowchart for calculating the first target residual score set of the class activation map generation method provided by an embodiment of the present application;

[0075] Figure 5 It is a flowchart for calculating the residual region set and the residual score set of the class activation map generation method provided by an embodiment of the present application;

[0076] Figure 6 It is a flowchart for calculating the residual region and the residual score of the class activation map generation method provided by an embodiment of the present application;

[0077] Figure 7 It is a flowchart for the activation step of the activation function of the class activation map generation method provided by an embodiment of the present application;

[0078] Figure 8 It is a main flowchart of the neural network interpretability method provided by an embodiment of the present application;

[0079] Figure 9 It is a flowchart of the neural network interpretability method of the convolutional neural network provided by an embodiment of the present application;

[0080] Figure 10 It is a flowchart of the neural network interpretability method of the spatio-temporal graph convolutional neural network provided by an embodiment of the present application;

[0081] Figure 11 It is an overall framework diagram of the class activation map generation network provided by an embodiment of the present application. Detailed implementation manners

[0082] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the embodiments of the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the embodiments of the present application.

[0083] It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that in the flowchart. Terms such as "first" and "second" in the specification, claims, and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0084] It should also be understood that references to "one embodiment" or "some embodiments" etc. described in the specification of the embodiments of the present application mean that specific features, structures or characteristics described in connection with that embodiment are included in one or more embodiments of the embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0085] The efficient feature extraction ability of neural networks drives researchers to study their internal complex operating mechanisms. However, in existing interpretable methods, there is a lack of discussion on the overlapping regions of feature maps, and non-target regions may contain features related to the target category, resulting in class activation maps containing interference from irrelevant regions.

[0086] Based on this, the embodiments of the present application provide a method, an interpretable method, a device, a device and a storage medium for generating class activation maps. In view of the problem that non-target regions may contain features related to the target category, the embodiments of the present application propose the concept of target residual score, and measure the score of the activation region by observing the difference between the image classification probability and the classification probability of the non-target region; then, in view of the problem of high overlapping regions in the feature map, the concept of regional residual score is proposed, and the score of the feature map is measured by evaluating the score of the residual region. The embodiments of the present application will assign weights to the feature map from two perspectives of target residual and regional residual.

[0087] The following further elaborates on the embodiments of the present application in conjunction with the accompanying drawings.

[0088] As Figure 1 shown, Figure 1 is a schematic structural diagram of a class activation map generation device provided by an embodiment of the present application. In Figure 1 the example, the class activation map generation device includes an image acquisition module 100, an image sorting module 300, an image processing module 200, a regional residual score module 400 and an activation function module 500.

[0089] Among them, the image acquisition module 100 is communicatively connected to the image processing module 200 respectively, and the image acquisition module 100 is used to acquire the original image.

[0090] The image processing module 200 is respectively communicatively connected to the image acquisition module 100, the image sorting module 300, and the regional residual score module 400. The image processing module 200 is configured to process the original image through the convolutional layer of the neural network to obtain a first feature map set, where the first feature map set is a set of multiple feature maps.

[0091] The image sorting module 300 is respectively communicatively connected to the image processing module 200 and the regional residual score module 300. The image sorting module 300 is configured to re-sort the first feature map set according to the importance of the feature maps in the first feature map set for the prediction category to obtain a second feature map set, and the prediction category is a predetermined category.

[0092] The regional residual score module 400 is respectively communicatively connected to the image processing module 200, the image sorting module 300, and the activation function module 500. The regional residual score module 400 is configured to calculate the residual region set and the residual score set corresponding to the second feature map set.

[0093] The activation function module 500 is communicatively connected to the regional residual score module 400. The activation function module 500 is configured to activate the residual region set and the residual score set through an activation function and obtain a class activation map.

[0094] In this embodiment, the image acquisition module 100 sends the acquired original image to the image processing module 200; the image processing module 200 processes the original image through the convolutional layer of the neural network to obtain a first feature map set, and sends the first feature map set to the image sorting module; the image sorting module 300 re-sorts the first feature map set according to the importance of the feature maps in the first feature map set for the prediction category, and sends the obtained second feature map set to the regional residual score module 400; the regional residual score module 400 calculates the residual region set and the residual score set corresponding to the second feature map set according to the relevant data obtained from the image processing module 200 and the image sorting module 300, and sends the residual region set and the residual score set to the activation function module 500; the activation function module 500 activates the residual region set and the residual score set through the activation function to obtain a class activation map.

[0095] The device and application scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0096] Those skilled in the art can understand that Figure 1The device structure shown does not constitute a limitation on the embodiments of the present application. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0097] In Figure 1 the shown device structure, each module can separately call the class activation map generation program stored therein to execute the class activation map generation method.

[0098] Based on the above device, various embodiments of the class activation map generation method of the embodiments of the present application are proposed.

[0099] As Figure 2 shown, Figure 2 is the main step diagram of the class activation map generation method provided by an embodiment of the present application. The class activation map generation method includes but is not limited to the following steps:

[0100] Step S100, obtain the original image.

[0101] Step S200, process the original image through the convolutional layer of the neural network to obtain a first feature map set, where the first feature map set is a set of multiple feature maps.

[0102] In one embodiment, the multiple feature maps obtained by passing the original image through the convolutional layer of the neural network also need to be upsampled to generate the first feature map set.

[0103] In one embodiment, the neural network is a convolutional neural network or a spatio-temporal graph convolutional neural network.

[0104] Step S300, reorder the first feature map set according to the importance of the feature maps in the first feature map set for the predicted class, where the predicted class is a predetermined class.

[0105] It should be understood that the predicted class is the target predicted class of the neural network, which can be the class to be expected to be measured or the class specified by the user.

[0106] It should be understood that the second feature map set is the result of reordering the first feature map set, and the feature maps in the obtained second feature map set are arranged in descending order of weight, which is convenient for generating the main region and the residual region that the neural network focuses on later.

[0107] Step S400, calculate the residual region set and the residual score set corresponding to the second feature map set.

[0108] Step S500, activate the residual region set and the residual score set through an activation function to obtain a class activation map.

[0109] It should be understood that the weights of the allocated feature maps are such that the set of residual scores corresponding to the set of residual regions is reduced, that is, the weights of the residual regions are reduced, while the weights of the main regions corresponding to the residual regions are increased. The set of residual regions and the set of residual scores are activated to obtain a class activation map. Compared with the existing neural network interpretability methods, the embodiments of the present application effectively improve the interpretability effect of the class activation map on the neural network.

[0110] In addition, in one embodiment, referring to Figure 3 and Figure 11 , step S300 may include but is not limited to the following steps:

[0111] Step 310: Invert the first set of feature maps to obtain a third set of feature maps.

[0112] It should be understood that the order of the feature maps in the third set of feature maps is opposite to that in the first set of feature maps. For example, take the k-th channel of the first set of feature maps, denoted as A k , then taking the k-th channel of the third set of feature maps is represented as 1 - A k .

[0113] Step S320: Obtain a first set of target residual scores according to the first set of feature maps, the third set of feature maps, and the predicted category. The first set of target residual scores includes multiple target residual scores, and the number of target residual scores is equal to the number of feature maps in the first set of feature maps.

[0114] It should be understood that the first set of target residual scores is calculated based on the first set of feature maps, the third set of feature maps, and the predicted category. There is a corresponding target residual for each feature map in the first set of feature maps or the third set of feature maps. Therefore, the number of target residual scores in the first set of target residual scores is equal to the number of feature maps in the first set of feature maps.

[0115] Step S330: Sort the target residual scores in the first set of target residual scores in descending order to obtain a second set of target residual scores.

[0116] Step S340: Sort the feature maps in the first set of feature maps according to the sequence of the target residuals in the second set of target residual scores to obtain a second set of feature maps.

[0117] It should be understood that sorting the feature maps in the first set of feature maps according to the sequence of the target residuals in the second set of target residual scores makes the feature maps in the obtained second set of feature maps arranged in descending order of weights, which is convenient for generating the main regions and residual regions that the neural network focuses on in the later stage.

[0118] In addition, in one embodiment, referring to Figure 4 and Figure 11, step S320 may include but is not limited to the following steps:

[0119] Step S321, obtain the feature map in the first feature map set as the first feature map.

[0120] Step S322, obtain the feature map in the third feature map set as the second feature map, where the position of the second feature map in the third feature map set is the same as the position of the first feature map in the first feature map set.

[0121] In one embodiment, the position of the second feature map in the third feature map set is the same as the position of the first feature map in the first feature map set, that is, take the k-th channel of the first feature map set as the first feature map, denoted as A k , then the second feature map is denoted as 1 - A k .

[0122] Step S323, pass the original image and the predicted category through the convolutional layer and the fully connected layer in the neural network to obtain the first probability, where the first probability is the probability that the first feature map is judged as the predicted category.

[0123] It should be understood that passing the original image X and the predicted category c through the convolutional layer and the fully connected layer in the neural network to obtain the first probability f c (X), where f c (·) is the probability of the neural network predicting the category c.

[0124] Step S324, perform the Hadamard product operation on the second feature map and the original image to obtain the third feature map.

[0125] It should be understood that performing the Hadamard product operation on the second feature Figure 1 -A k and the original image X to obtain the third feature map, then the third feature map is denoted as (1 - A k )X.

[0126] Step S325, pass the third feature map and the predicted category through the convolutional layer and the fully connected layer to obtain the second probability, where the second probability is the probability that the third feature map is judged as the predicted category.

[0127] It should be understood that passing the third feature map (1 - A k )X and the predicted category c through the convolutional layer and the fully connected layer to obtain the second probability, then the second probability is denoted as f c [(1 - A k )X], where f c (·) is the probability of the neural network predicting the category c.

[0128] Step S326, calculate the difference between the first probability and the second probability to obtain the target residual score.

[0129] It should be understood that the difference between the first probability f c (X) and the second probability f c [(1 - A k )X] is obtained to get the target residual score, and the target residual score can be expressed as:

[0130]

[0131] where f c (X) is the classification probability of observing the original image, and f c [(1 - A k )X] is the classification probability of removing the target region.

[0132] Step S327: Repeatedly obtain the feature maps in the first feature map set and the third feature map set, calculate the corresponding target residual scores until no new feature maps can be detected in the first feature map set and the third feature map set, and use the set of multiple calculated target residual scores as the first target residual score set.

[0133] It should be understood that the first target residual score set is the score set for the activation regions of each channel of the first feature map set.

[0134] In addition, in one embodiment, referring to Figure 5 and Figure 11 , step S400 may include but is not limited to the following steps:

[0135] Step S410: Obtain the feature maps in the second feature map set as the fourth feature map.

[0136] Step S420: Find the maximum value of the fourth feature map and the feature maps before the fourth feature map to obtain the first saliency map.

[0137] It should be understood that after obtaining the scores for the activation regions of each channel of the feature map, starting from the most central region concerned by the neural network, the importance of the surrounding regions is gradually explored. Among them, the most central region concerned by the neural network is the first feature map in the second feature map set. When the fourth feature map set in the second feature map set is selected, the maximum value of the fourth feature map and the feature maps before the fourth feature map is found to obtain the first saliency map. Then the first saliency map is a new activation region generated by combining the central region and its surrounding regions.

[0138] It should be understood that the maximum value of the fourth feature map and the feature maps before the fourth feature map is found to obtain the first saliency map, and the first saliency map can be expressed as:

[0139]

[0140] Among them, p is a sequence obtained according to the second residual score set. is the descending sorting operator, ∪(·) is the union operator, and max(·) indicates that the value at each position in the output new matrix is the maximum value of the corresponding positions in the two input matrices.

[0141] Step S430: Calculate the residual region and residual score corresponding to the first saliency map.

[0142] Step S440: Repeatedly obtain the feature maps in the second feature map set, calculate the corresponding residual regions and residual scores until no new feature map can be detected in the second feature map set, and use the calculated multiple residual regions as the residual region set and the multiple residual scores as the residual score set.

[0143] It should be understood that the residual regions in the residual region set and the residual score set correspond to the residual scores, and the number of residual regions is equal to the number of residual scores.

[0144] In addition, in one embodiment, referring to Figure 6 and Figure 11 , step S430 may include but is not limited to the following steps:

[0145] Step S431: Obtain the upper-level feature map of the first saliency map as the second saliency map.

[0146] Step S432: Calculate the difference between the first saliency map and the second saliency map, and use the obtained calculation result as the residual region.

[0147] It should be understood that for the first saliency map and the second saliency map The residual region of the first saliency map can be expressed as:

[0148]

[0149] Among them, R represents the residual region. Since the first saliency map is the maximum value of the fourth feature map and the feature maps before the fourth feature map, then R p can be 0.

[0150] Step S433: Invert the first saliency map to obtain the third saliency map.

[0151] It should be understood that by inverting the first saliency map , the third saliency map is obtained, and the third saliency map can be expressed as

[0152] Step S434: Perform a Hadamard product operation on the third saliency map and the original image to obtain the fourth saliency map.

[0153] It should be understood that the third significant map is subjected to Hadamard product operation with the original image X to obtain a fourth significant map, and the fourth significant map can be expressed as

[0154] Step S435: Pass the fourth significant map and the predicted category through a convolutional layer and a fully connected layer to obtain a third probability, where the third probability is the probability that the fourth significant map is judged as the predicted category.

[0155] It should be understood that the fourth significant map is passed through a convolutional layer and a fully connected layer with the predicted category c to obtain a third probability, and the third probability can be expressed as

[0156] Step S436: Calculate the difference between the first probability and the third probability to obtain a first region score.

[0157] It should be understood that the difference between the first probability f c (X) and the third probability is calculated to obtain a first region score, and the first region score can be expressed as:

[0158]

[0159] It should be understood that the first region score can also be expressed as the result of evaluating the importance of the new activation region to the surrounding regions, where the first significant map is the new activation region.

[0160] Step S437: Calculate the region score corresponding to the second significant map as the second region score.

[0161] It should be understood that according to Step S331, Step S332, Step S333, Step S334, Step S335, and Step S336, the region score corresponding to the second significant map is calculated as the second region score

[0162] Step S438: Calculate the difference between the first region score and the second region score to obtain a residual score.

[0163] It should be understood that the difference between the first region score and the second region score is calculated to obtain a residual score, and the residual score can be expressed as

[0164]

[0165] Since the first significant map is the maximum value of the fourth feature map and the feature maps before the fourth feature map, the residual score It can be 0.

[0166] In addition, in one embodiment, with reference to Figure 7 and Figure 11 , step S500 may include but is not limited to the following steps:

[0167] Step S510: Obtain the first residual region in the residual region set and the first residual score in the residual score set.

[0168] Step S520: Multiply the first residual region by the first residual score to obtain the fifth saliency map.

[0169] Step S530: Repeatedly obtain the residual regions in the residual region set and the residual scores in the residual score set, calculate their corresponding saliency maps until no new residual regions and residual scores can be detected in the residual region set and the residual score set, and use the sum of the calculated multiple saliency maps as the saliency map set.

[0170] Step S540: Activate the saliency map set through an activation function to obtain the class activation map.

[0171] In one embodiment, the activation function is the rectified linear unit function.

[0172] It should be understood that for the residual region R in the residual region set p and the residual score in the residual score set The final class activation map can be expressed as:

[0173]

[0174] In one embodiment, the neural network is a convolutional neural network or a spatio-temporal graph convolutional neural network.

[0175] In the embodiment of the present application, the first feature map set is sorted by the target residual score to obtain the second feature map set, and the corresponding residual region set and residual score set of the second feature map set are calculated, and the residual region set and the residual score set are activated to obtain the class activation map. The second feature map set generated in the embodiment of the present application is sorted according to the importance of the feature map to the predicted category, which is convenient for generating the residual region set of the neural network. The embodiment of the present application assigns weights to the feature maps from two perspectives of the target residual and the regional residual, so that the residual score set corresponding to the residual region set is reduced, that is, the weight of the residual region is reduced, and the weight of the main region corresponding to the residual region is increased, thereby effectively improving the interpretability effect of the generated class activation map for the neural network.

[0176] In addition, as Figure 8 shown, an embodiment of the method for interpreting a neural network provided by the embodiment of the present application includes but is not limited to the following steps:

[0177] Step S610: Obtain the original image and the class activation map, where the class activation map is calculated by using the class activation map generation method of steps S100 to S400.

[0178] Step S620: Superimpose the class activation map on the original image to obtain the target image for interpreting the neural network.

[0179] Step S630: Output the target image.

[0180] In addition, in an embodiment, referring to Figure 9 , when the neural network is a convolutional neural network, step S520 may include but is not limited to the following steps:

[0181] Step S621: Convert the class activation map into a pseudo-color image.

[0182] It should be understood that since the class activation map is a grayscale image, it is necessary to convert the class activation map into a pseudo-color image based on OpenCV.

[0183] Step S622: Superimpose the pseudo-color image and the original image by addition to obtain the target image.

[0184] It should be understood that this neural network interpretable method can be applied to image classification.

[0185] In addition, in an embodiment, referring to Figure 10 , when the neural network is a spatio-temporal graph convolutional neural network, step S520 may include but is not limited to the following steps:

[0186] Step S623: Obtain the skeleton motion sequence, where the skeleton motion sequence is obtained by the spatio-temporal convolutional neural network according to the object to be measured.

[0187] It should be understood that the skeleton motion sequence is obtained, and the minimum coordinate of the skeleton motion sequence is set as the origin of the drawing coordinate, that is, the coordinate of each point in the sequence is subtracted by the minimum coordinate.

[0188] Step S624: Sample the skeleton motion sequence frame by frame to obtain the skeleton motion vector.

[0189] It should be understood that the skeleton motion sequence is sampled, and one sample is taken every 5 frames to obtain the skeleton motion vector.

[0190] Step S625: Draw the importance of the skeleton motion sequence according to the class activation map to obtain the target image, where the target image is used to interpret and display the weight change of the skeleton nodes between different frames.

[0191] It can be understood that matplotlib can be used for plotting. The skeleton nodes and the connections of the skeleton are plotted on a three-dimensional coordinate according to the coordinates and time sequence, and the action interval is 1.5 times the width of the previous frame's action. Then, the importance is plotted on the skeleton motion sequence according to the class activation map, and the transparency of the node color is used as an interpretable way of node importance. Among them, the target image is used for interpretable display of the weight changes of the skeleton nodes between different frames.

[0192] It should be noted that steps S623 to S625 in this embodiment and steps S621 and S622 in the above-described embodiment as Figure 9 shown are parallel technical solutions to each other.

[0193] In addition, an embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it performs the class activation map generation method according to any one of the first aspects and / or the neural network interpretability method according to any one of the second aspects.

[0194] The processor and the memory can be connected through a bus or other means.

[0195] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and their combinations.

[0196] The non-transitory software programs and instructions required to implement the class activation map generation method and / or the neural network interpretability method of the above embodiments are stored in the memory. When executed by the processor, they execute the class activation map generation method in the above embodiments. For example, they execute the method steps S100 to S500 described above, Figure 2 or execute the neural network interpretability method in the above embodiments. For example, they execute the method steps S610 to S630 described above. Figure 8

[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.​

[0198] In addition, an embodiment of the embodiments of the present application further provides a computer-readable storage medium storing computer-executable instructions, which when executed by a processor or a controller, enable the above-mentioned processor to execute the class activation map generation method in the above-mentioned embodiments. For example, execute the method steps S100 to S500 described above in Figure 2 ; the method steps S310 to S340 in Figure 3 ; the method steps S321 to S327 in Figure 4 ; the method steps S410 to S440 in Figure 5 ; the method steps S431 to S438 in Figure 6 ; the method steps S510 to S540 in Figure 7 . Again, when executed by a processor in the above-mentioned embodiments, it enables the above-mentioned processor to execute the neural network interpretability method in the above-mentioned embodiments. For example, execute the method steps S610 to S630 described above in Figure 8 ; the method steps S621 to S622 in Figure 9 ; the method steps S623 to S625 in Figure 10 .

[0199] Those of ordinary skill in the art can understand that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or a non-transitory medium) and a communication medium (or a transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cassette, tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium generally contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0200] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for generating class activation maps, characterized in that, including: obtaining an original image; processing the original image through a convolutional layer of a neural network to obtain a first feature map set, where the first feature map set is a set of multiple feature maps; reordering the first feature map set according to the importance of the feature maps in the first feature map set for a prediction category, to obtain a second feature map set, where the prediction category is a predetermined category; calculating a residual region set and a residual score set corresponding to the second feature map set; activating the residual region set and the residual score set through an activation function to obtain a class activation map; wherein, the calculating the residual region set and the residual score set corresponding to the second feature map set includes: obtaining the feature maps in the second feature map set as fourth feature maps; obtaining a maximum value of the fourth feature map and the feature maps before the fourth feature map to obtain a first saliency map; calculating a residual region and a residual score corresponding to the first saliency map; repeatedly obtaining the feature maps in the second feature map set, calculating corresponding residual regions and residual scores until no new feature maps can be detected in the second feature map set, and taking the calculated multiple residual regions as the residual region set and multiple residual scores as the residual score set; the calculating the residual region and the residual score corresponding to the first saliency map includes: obtaining the feature map at the upper level of the first saliency map as a second saliency map; calculating a difference between the first saliency map and the second saliency map, and taking the obtained calculation result as the residual region; inverting the first saliency map to obtain a third saliency map; performing a Hadamard product operation on the third saliency map and the original image to obtain a fourth saliency map; passing the fourth saliency map and the prediction category through the convolutional layer and the fully connected layer to obtain a third probability, where the third probability is the probability that the fourth saliency map is judged as the prediction category; calculating a difference between the first probability and the third probability to obtain a first region score; calculating a region score corresponding to the second saliency map as a second region score; calculating a difference between the first region score and the second region score to obtain a residual score.

2. The method for generating class activation maps according to claim 1, characterized in that, the reordering the first feature map set according to the importance of the feature maps in the first feature map set for a prediction category to obtain a second feature map set includes: inverting the first feature map set to obtain a third feature map set; obtaining a first target residual score set according to the first feature map set, the third feature map set and the prediction category, where the first target residual score set includes multiple target residual scores, and the number of the target residual scores is equal to the number of the feature maps in the first feature map set; sorting the target residual scores in the first target residual score set in descending order to obtain a second target residual score set; sorting the feature maps in the first feature map set according to the sequence of the target residuals in the second target residual score set to obtain a second feature map set.

3. The method for generating class activation maps according to claim 2, characterized in that, the obtaining the first target residual score set according to the first feature map set, the third feature map set and the prediction category includes: Obtain the feature maps in the first feature map set as the first feature maps; Obtain the feature maps in the third feature map set as the second feature maps, where the position of the second feature maps in the third feature map set is the same as the position of the first feature maps in the first feature map set; Pass the original image and the predicted category through the convolutional layer and the fully connected layer in the neural network to obtain a first probability, where the first probability is the probability that the first feature maps are judged as the predicted category; Perform a Hadamard product operation on the second feature maps and the original image to obtain third feature maps; Pass the third feature maps and the predicted category through the convolutional layer and the fully connected layer to obtain a second probability, where the second probability is the probability that the third feature maps are judged as the predicted category; Calculate the difference between the first probability and the second probability to obtain the target residual score; Repeat obtaining the feature maps in the first feature map set and the third feature map set, calculate the corresponding target residual scores until no new feature maps can be detected in the first feature map set and the third feature map set, and use the set of multiple calculated target residual scores as the first target residual score set.

4. The method for generating class activation maps according to claim 1, characterized in that, The step of activating the residual region set and the residual score set through an activation function to obtain a class activation map includes: Obtain a first residual region in the residual region set and a first residual score in the residual score set; Multiply the first residual region by the first residual score to obtain a fifth saliency map; Repeat obtaining the residual regions in the residual region set and the residual scores in the residual score set, calculate their corresponding saliency maps until no new residual regions and residual scores can be detected in the residual region set and the residual score set, and use the sum of the multiple calculated saliency maps as the saliency map set; Activate the saliency map set through an activation function to obtain a class activation map.

5. The method for generating a class activation map according to claim 4, wherein, The activation function is a rectified linear unit function.

6. The method for generating a class activation map according to any one of claims 1 to 5, wherein, The neural network is a convolutional neural network or a spatio-temporal graph convolutional neural network.

7. A method for interpreting a neural network, wherein, It includes: Obtain an original image and a class activation map, where the class activation map is obtained by passing the original image through a neural network and using the class activation map generation method according to any one of claims 1 to 6; Overlay the class activation map on the original image to obtain a target image for explaining the neural network; Output the target image.

8. The method for interpreting a neural network according to claim 7, wherein, When the neural network is a convolutional neural network, the step of overlaying the class activation map on the original image to obtain a target image for explaining the neural network includes: Convert the class activation map into a pseudo-color image; Overlay the pseudo-color image and the original image through addition to obtain a target image.

9. The method for interpreting a neural network according to claim 7, wherein, When the neural network is a spatio-temporal graph convolutional neural network, the step of overlaying the class activation map on the original image to obtain a target image for explaining the neural network includes: Obtain a skeleton sequence, where the skeleton sequence is obtained by the spatio-temporal graph convolutional neural network based on the object to be measured; Sample the skeleton sequence frame by frame to obtain a skeleton motion vector; Drawing the importance of the skeleton motion sequence according to the class activation map to obtain a target image, where the target image is used to interpretably display the weight changes of skeleton nodes between different frames.

10. A class activation map generation device, wherein, Including: An image acquisition module for acquiring an original image; An image processing module for processing the original image through the convolutional layer of a neural network to obtain a first feature map set; An image sorting module for re-sorting the first feature map set according to the importance of the feature maps in the first feature map set for predicting classes to obtain a second feature map set; A region residual score module for calculating a residual region set and a residual score set corresponding to the second feature map set; An activation function module for activating the residual region set and the residual score set through an activation function to obtain a class activation map; Wherein, calculating the residual region set and the residual score set corresponding to the second feature map set includes: Obtaining the feature maps in the second feature map set as the fourth feature map; Taking the maximum value of the fourth feature map and the feature maps before the fourth feature map to obtain a first saliency map; Calculating the residual region and the residual score corresponding to the first saliency map; Repeatedly obtaining the feature maps in the second feature map set, calculating the corresponding residual regions and residual scores until no new feature maps can be detected in the second feature map set, and taking the calculated multiple residual regions as the residual region set and the multiple residual scores as the residual score set; Calculating the residual region and the residual score corresponding to the first saliency map includes: Obtaining the feature map at the previous level of the first saliency map as the second saliency map; Calculating the difference between the first saliency map and the second saliency map, and taking the obtained calculation result as the residual region; Inverting the first saliency map to obtain a third saliency map; Performing a Hadamard product operation on the third saliency map and the original image to obtain a fourth saliency map; Passing the fourth saliency map and the predicted class through the convolutional layer and the fully connected layer to obtain a third probability, where the third probability is the probability that the fourth saliency map is judged as the predicted class; Calculating the difference between the first probability and the third probability to obtain a first region score; Calculating the region score corresponding to the second saliency map as the second region score; Calculating the difference between the first region score and the second region score to obtain the residual score.

11. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that when the processor executes the computer program, it performs the class activation map generation method according to any one of claims 1 to 6 and / or the neural network interpretability method according to any one of claims 7 to 9.

12. A computer storage medium, characterized in that, It includes computer-executable instructions stored, and the computer-executable instructions are used to execute the class activation map generation method according to any one of claims 1 to 6 and / or the neural network interpretability method according to any one of claims 7 to 9.