Convolutional neural network-oriented positive and negative excitation saliency map generation method and system

By writing positive and negative excitation extraction functions in convolutional neural networks and using double-stranded backpropagation, the problem of relying on gradients and high computational complexity in the prior art is solved, and efficient and reliable significant graph generation and network interpretation are achieved.

CN119990342APending Publication Date: 2025-05-13NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510042997.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When generating significant graphs of convolutional neural networks, the prior art relies on gradients, has high computational complexity, and it is difficult to effectively interpret negative gradients, resulting in inaccurate results.

Method used

Using a positive and negative excitation extraction function that does not depend on gradients, by writing a positive and negative excitation extraction function for each layer of a convolutional neural network, the positive and negative excitation of each layer is calculated, and a significant graph of the entire network is obtained using double-stranded backpropagation.

Benefits of technology

The significant graph generation of convolutional neural networks with low computational complexity and high reliability is realized, which can effectively explain the network's decision-making process and provide a more convincing local explanation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990342A_ABST
    Figure CN119990342A_ABST
Patent Text Reader

Abstract

The invention discloses a convolutional neural network-oriented positive and negative excitation saliency map generation method and system. The method specifically comprises the following steps of: firstly, carrying out forward propagation on a convolutional neural network; compiling corresponding positive and negative excitation extraction functions for a linear layer, a convolutional layer, an activation layer, a pooling layer and a standardization layer of the convolutional neural network; calculating positive and negative excitation of each layer of the neural network; and finally, obtaining a saliency map of the whole neural network by using double-chain back propagation. According to the method, the excitation extraction function is written for each layer of the convolutional neural network, the complete information of each layer of the convolutional neural network can be utilized without gradient, the positive and negative excitation of each layer is extracted, double-chain back propagation is introduced, the excitation is combined into the saliency map with higher reliability, and the saliency map is more reliable. And therefore, more credible neural network interpretation can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning interpretability, and in particular to a method for generating positive and negative excitation saliency maps for convolutional neural networks. Background Art

[0002] A saliency map is an image that highlights the areas that our eyes focus on first. The goal of a saliency map is to reflect the importance of pixels to the human visual system. In computer vision, saliency maps have become a widely used technique to illustrate the decision-making process of Convolutional Neural Networks (CNNs). As a local explanation method, saliency maps can detect which pixels or regions in a sample have a greater impact on the model's decision.

[0003] However, some early methods, such as the Guided BP method, suffer from the problem of lack of category sensitivity. Researchers have addressed this problem by generating perturbation-based saliency maps. For example, Zeiler and Fergus used fixed-size pixel blocks to continuously mask the image and measured the saliency value of each block by observing changes in the target class; Fong and Vedaldi used optimization to promote mask learning to generate meaningful perturbations. However, perturbation-based saliency map generation methods have considerable randomness, poor results, and high computational complexity. In order to overcome these limitations and better characterize the image saliency of a given category, Class Activation Mapping (CAM) was proposed. CAM assumes that the fully connected layers of the model contain high-dimensional semantic information, such as the object concept of the category, and the feature map generated by the convolutional layer contains spatial structural information. In order to make the saliency map category sensitive, CAM obtains a mapping relationship from the target class to a specific layer of the feature map, that is, a linear weight, and then linearly merges the feature map to generate a saliency map.

[0004] Many existing CAMs and other saliency map generation methods rely heavily on gradients. Gradient-based methods can only produce accurate weights when the model is sufficiently linear, which makes gradient-based methods less effective when facing more complex models. In addition, interpreting negative gradients is also a major challenge, which has prompted some researchers to use the Rectified Linear Unit (ReLU) function to filter out such gradients. However, this approach will lead to further information loss in the network, which will have a negative impact on the accuracy of the saliency map. Summary of the invention

[0005] The purpose of the present invention is to provide a convolutional neural network saliency map generation method that is independent of gradient, has low computational complexity and high reliability.

[0006] The technical solution to achieve the purpose of the present invention is: a method for generating positive and negative excitation saliency maps for convolutional neural networks, comprising the following steps:

[0007] Step 1: Forward propagation of convolutional neural network;

[0008] Step 2: Write corresponding positive and negative excitation extraction functions for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network;

[0009] Step 3: Calculate the positive and negative excitations for each layer of the convolutional neural network;

[0010] Step 4: Use double-chain back-propagation to obtain the saliency map of the entire convolutional neural network.

[0011] Furthermore, for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network described in step 2, corresponding positive and negative excitation extraction functions are written, as follows:

[0012] Step 2.1, write the positive and negative excitation extraction function of the linear layer of the convolutional neural network;

[0013] Step 2.2, write the positive and negative excitation extraction function of the convolutional layer of the convolutional neural network;

[0014] Step 2.3, write the positive and negative excitation extraction function of the convolutional neural network activation layer;

[0015] Step 2.4, write the positive and negative excitation extraction function of the convolutional neural network pooling layer;

[0016] Step 2.5: Write the positive and negative excitation extraction function of the convolutional neural network normalization layer.

[0017] Furthermore, the positive and negative excitation extraction function of the convolutional neural network linear layer described in step 2.1 is written as follows:

[0018] Step 2.1.1, parameter definition:

[0019] The expression for setting the linear layer is:

[0020] Y=w·X+b

[0021] Among them, X represents input, w represents weight, b represents bias, and Y represents output; set X∈R V , Y∈R W , b∈R W , w∈R W×V , then for fixed w and b, the part of the equation related to X is Y′=Yb;

[0022] Step 2.1.2, obtain the linear layer positive and negative excitation coefficients:

[0023] The elements of the positive excitation coefficient of Y′ are defined as follows:

[0024]

[0025] Where, i∈W, j∈V; 0 is used to fill the blank part of the positive excitation. Since 0 is used as the zero element of the multiplication group and the unit element of the addition group in the real number field, it will not affect the excitation calculation;

[0026] Negative excitation coefficient of Y′ Use the same method to determine;

[0027] Since any Y is subject to the fixed influence of bias b, the excitation of Y′ is essentially the excitation of Y.

[0028] Furthermore, the positive and negative excitation extraction function of the convolutional neural network convolutional layer described in step 2.2 is written as follows:

[0029] Step 2.2.1, parameter definition:

[0030] The regular expression for setting the convolutional layer is:

[0031] Y=X*f

[0032] Among them, X represents input, f represents convolution kernel, and Y represents output;

[0033] In image convolution, define X∈R W×H , f∈R i×j , indicating the stride parameter s of the sliding step of the convolution kernel f, then the output Y∈R ([(W-i) / s]+1)×([(H-i) / s]+1) , abbreviated as Where o1 = [(Wi) / s] + 1, o2 = [(Hi) / s] + 1;

[0034] Step 2.2.2, get the positive and negative excitation coefficients of the convolutional layer:

[0035] The linear form of the convolutional layer is:

[0036] Y trans =X trans ·f trans

[0037]

[0038] f trans ∈R i·j

[0039] Among them, Y trans and f trans are the vector expansions of Y and f, respectively, X trans is the complex expansion of X according to f;

[0040] X trans The relationship between and the elements of X is expressed as follows:

[0041]

[0042] Among them, k∈[0,i-1] and l∈[0,j-1], therefore, the original convolution is equivalent to a linear layer, and the excitation can be obtained in a similar way to the linear layer. The input expanded to a vector can be restored to its original shape by directly using deconvolution;

[0043] Step 2.2.3, the incentive in tensor form is as follows:

[0044] For any element Y k,l ∈Y, where k∈[0,o1-1], l∈[0,o2-1], is calculated only from a slice of X, and the calculation formula is:

[0045]

[0046] The size of the slice is equal to the size of f. Therefore, the excitation of each element in the output Y is the corresponding slice from the input X. The specific excitation coefficient is the linear sum of all corresponding slices. When the spatial information is explicitly reflected in the excitation, the corresponding excitation can be expressed as The values ​​of the positions that are the same as the slice position are equal, and the values ​​of other positions are 0.

[0047] Furthermore, the positive and negative excitation extraction function of the activation layer of the convolutional neural network described in step 2.3 is written as follows:

[0048] Step 2.3.1, parameter definition:

[0049] In convolutional neural networks, ReLU is a commonly used activation function, which is defined as:

[0050] ReLU(x)=max(x,0)

[0051] in, This function is linear when x>0 and masks negative values ​​in the input signal;

[0052] Step 2.3.2, obtain the positive and negative excitation coefficients of the activation layer:

[0053] The coefficients of ReLU(x) can be obtained by max(x, 0) / x, where x≠0; since ReLU has no negative excitation, signals with values ​​greater than zero are considered positive excitation signals, so the coefficients of all positive excitation signals are set to 1, and the coefficients of negative excitations are set to zero for subsequent calculations in this particular case.

[0054] Furthermore, the positive and negative excitation extraction function of the convolutional neural network pooling layer described in step 2.4 is written as follows:

[0055] Step 2.4.1, parameter definition:

[0056] Maximum pooling and average pooling are two common pooling methods. The expression of the maximum pooling function is set as:

[0057] MaxPool(X)=max(X)

[0058] in, This method represents the original entire input signal X by selecting the maximum value in a given region;

[0059] The expression of the average pooling function is set as:

[0060]

[0061] Step 2.4.2, get the positive and negative excitation coefficients of the maximum pooling layer:

[0062] Similar to ReLU, the maximum pooling has only positive excitation and zero excitation. The only positive excitation signal of MaxPool(X) corresponds to the element X i,j ∈X, satisfying X i,j =max(X), the positive excitation coefficient of this element is 1;

[0063] Step 2.4.3, get the positive and negative excitation coefficients of the average pooling layer:

[0064] The average pooling function is similar to a linear function, and all weights are set to (W×H) -1 , the bias is set to 0, and the activation is calculated using the same principle as the linear layer.

[0065] Furthermore, the positive and negative excitation extraction function of the normalized layer of the convolutional neural network described in step 2.5 is written as follows:

[0066] Step 2.5.1, parameter definition:

[0067] Batch normalization BN is a commonly used normalization technique. The expression of BN is set as:

[0068]

[0069] in μ is the mean of X, σ is the variance of X, ∈ is a small bias to avoid division by zero, represents the Hadamard product, is the learning weight, is the deviation of the output;

[0070] Step 2.5.2, obtain the normalized layer positive and negative excitation extraction function:

[0071] For a given X, μ and σ are fixed and can therefore be treated as fixed parameters in static analysis. The expression can be rewritten as a linear function, which for any element in the output BN(X) is:

[0072]

[0073] in, The BN layer obtains the excitation coefficients in the same way as the linear layer.

[0074] Furthermore, the double-chain back propagation described in step 4 is used to obtain the saliency map of the entire neural network, as follows:

[0075] Step 4.1, parameter definition:

[0076] In any network F, given layer L n The input and output are represented by O n-1 and O n Indicates that O n-1 is equal to X when n = 1, so the positive excitation coefficient can be obtained and negative incentive coefficient Since zero excitation does not contribute to the output, a total of 2N excitation coefficients can be obtained for all N layers in the network, with positive excitation signals being positive for positive excitations and negative for negative excitations;

[0077] Step 4.2, double chain back propagation:

[0078] The characteristic of the excitation signal is transitive, which can be expressed as:

[0079]

[0080] This transitivity is reflected in the double-chain structure, where the corresponding positive and negative excitation signals of the two layers are cross-multiplied and then summed to complete the transfer of excitation; since the excitation coefficients have the same transitivity as the excitation signals, they can also be transferred in this way;

[0081] Step 4.3: Get the saliency map of the entire network:

[0082] The output O from any layer is obtained by iterative method j To any layer output O i The positive and negative incentive coefficients are as follows:

[0083]

[0084] Where i>j; when i=j+1, When i = N and j = 0, the resulting excitation coefficient is represents the excitation coefficient of the entire network, is the obtained saliency map, where represents the coefficient that has a positive contribution to the result for the corresponding pixel in the image, while represents the opposite coefficient.

[0085] A positive and negative excitation saliency map generation system for convolutional neural networks, the system is used to implement the positive and negative excitation saliency map generation method for convolutional neural networks, the system includes a first module to a fourth module, wherein:

[0086] The first module implements forward propagation based on convolutional neural network;

[0087] The second module writes corresponding positive and negative excitation extraction functions for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network;

[0088] The third module calculates the positive and negative excitations of each layer of the convolutional neural network;

[0089] The fourth module uses double-chain back-propagation to obtain the saliency map of the entire convolutional neural network.

[0090] A mobile terminal comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for generating positive and negative excitation saliency maps for convolutional neural networks is implemented.

[0091] Compared with the prior art, the present invention has the following significant advantages: (1) the use of positive and negative excitation extraction functions can utilize the complete information of each layer of the convolutional neural network without gradient, extract the positive and negative excitations of each layer to form a local explanation; (2) the double-chain back-propagation process is used to calculate the excitation coefficients between any two layers in the network, and two positive and negative excitation function graphs are obtained, which effectively reduces the time complexity associated with calculating the complete excitation and makes it reach a level equivalent to that of forward back-propagation; (3) the double-chain back-propagation method is used to combine the excitations into a more reliable saliency map, thereby obtaining a more convincing convolutional neural network explanation. BRIEF DESCRIPTION OF THE DRAWINGS

[0092] Figure 1 It is a flow chart of a method for generating positive and negative excitation saliency maps for convolutional neural networks according to the present invention.

[0093] Figure 2 is a saliency map generated by the method of the present invention in an embodiment of the present invention. DETAILED DESCRIPTION

[0094] In order to better understand the present invention, the content of the present invention is further described below in conjunction with the accompanying drawings.

[0095] like Figure 1 As shown, the present invention provides a method for generating positive and negative excitation saliency maps for convolutional neural networks, comprising the following steps:

[0096] Step 1: Forward propagation of convolutional neural network;

[0097] Step 2: For the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network, write the corresponding positive and negative excitation extraction functions, as follows:

[0098] Step 2.1, write the positive and negative excitation extraction function of the linear layer of the convolutional neural network, as follows:

[0099] Step 2.1.1, parameter definition:

[0100] The expression for setting the linear layer is:

[0101] Y=w·X+b

[0102] Among them, X represents input, w represents weight, b represents bias, and Y represents output; set X∈R V , Y∈R W , b∈R W , w∈R W×V , then for fixed w and b, the part of the equation related to X is Y′=Yb;

[0103] Step 2.1.2, obtain the linear layer positive and negative excitation coefficients:

[0104] The elements of the positive excitation coefficient of Y′ are defined as follows:

[0105]

[0106] Where, i∈W, j∈V; 0 is used to fill the blank part of the positive excitation. Since 0 is used as the zero element of the multiplication group and the unit element of the addition group in the real number field, it will not affect the excitation calculation;

[0107] Negative excitation coefficient of Y′ Use the same method to determine;

[0108] Since any Y is subject to the fixed influence of bias b, the excitation of Y′ is essentially the excitation of Y.

[0109] Step 2.2: Write the positive and negative excitation extraction function of the convolutional neural network convolutional layer, as follows:

[0110] Step 2.2.1, parameter definition:

[0111] The regular expression for setting the convolutional layer is:

[0112] Y=X*f

[0113] Among them, X represents input, f represents convolution kernel, and Y represents output;

[0114] In image convolution, define X∈R W×H , f∈R i×j , indicating the stride parameter s of the sliding step of the convolution kernel f, then the output Y∈R ([(W-i) / s]+1)×([(H-i) / s]+1) , abbreviated as Where o1 = [(Wi) / s] + 1, o2 = [(Hi) / s] + 1;

[0115] Step 2.2.2, get the positive and negative excitation coefficients of the convolutional layer:

[0116] The linear form of the convolutional layer is:

[0117] Y trans =X trans ·f trans

[0118]

[0119] f trans ∈R i·j

[0120] Among them, Y trans and f trans are the vector expansions of Y and f, respectively, X trans is the complex expansion of X according to f;

[0121] X trans The relationship between and the elements of X is expressed as follows:

[0122]

[0123] Among them, k∈[0,i-1] and l∈[0,j-1], therefore, the original convolution is equivalent to a linear layer, and the excitation can be obtained in a similar way to the linear layer. The input expanded to a vector can be restored to its original shape by directly using deconvolution;

[0124] Step 2.2.3, the incentive in tensor form is as follows:

[0125] For any element Y k,l ∈Y, where k∈[0,o1-1], l∈[0,o2-1], is calculated only from a slice of X, and the calculation formula is:

[0126]

[0127] The size of the slice is equal to the size of f. Therefore, the excitation of each element in the output Y is the corresponding slice from the input X. The specific excitation coefficient is the linear sum of all corresponding slices. When the spatial information is explicitly reflected in the excitation, the corresponding excitation can be expressed as The values ​​of the positions that are the same as the slice position are equal, and the values ​​of other positions are 0.

[0128] Step 2.3, write the positive and negative excitation extraction function of the convolutional neural network activation layer, as follows:

[0129] Step 2.3.1, parameter definition:

[0130] In convolutional neural networks, ReLU is a commonly used activation function, which is defined as:

[0131] ReLU(x)=max(x,0)

[0132] in, The function is linear when x>0 and masks the negative values ​​in the input signal;

[0133] Step 2.3.2, obtain the positive and negative excitation coefficients of the activation layer:

[0134] The coefficients of ReLU(x) can be obtained by max(x, 0) / x, where x≠0; since ReLU has no negative excitation, signals with values ​​greater than zero are considered positive excitation signals, so the coefficients of all positive excitation signals are set to 1, and the coefficients of negative excitations are set to zero for subsequent calculations in this particular case.

[0135] Step 2.4: Write the positive and negative excitation extraction function of the convolutional neural network pooling layer, as follows:

[0136] Step 2.4.1, parameter definition:

[0137] Maximum pooling and average pooling are two common pooling methods. The expression of the maximum pooling function is set as:

[0138] MaxPool(X)=max(X)

[0139] in, This method represents the original entire input signal X by selecting the maximum value in a given region;

[0140] The expression of the average pooling function is set as:

[0141]

[0142] Step 2.4.2, get the positive and negative excitation coefficients of the maximum pooling layer:

[0143] Similar to ReLU, the maximum pooling has only positive excitation and zero excitation. The only positive excitation signal of MaxPool(X) corresponds to the element X i,j ∈X, satisfying X i,j =max(X), the positive excitation coefficient of this element is 1;

[0144] Step 2.4.3, get the positive and negative excitation coefficients of the average pooling layer:

[0145] The average pooling function is similar to a linear function, and all weights are set to (W×H) -1 , the bias is set to 0 and the excitation is calculated using the same principle as the linear layer.

[0146] Step 2.5: Write the positive and negative excitation extraction function of the convolutional neural network normalization layer, as follows:

[0147] Step 2.5.1, parameter definition:

[0148] Batch normalization BN is a commonly used normalization technique. The expression of BN is set as:

[0149]

[0150] in μ is the mean of X, σ is the variance of X, ∈ is a small bias to avoid division by zero, represents the Hadamard product, is the learning weight, is the deviation of the output;

[0151] Step 2.5.2, obtain the normalized layer positive and negative excitation extraction function:

[0152] For a given X, μ and σ are fixed and can therefore be treated as fixed parameters in static analysis. The expression can be rewritten as a linear function, which for any element in the output BN(X) is:

[0153]

[0154] in, The BN layer obtains the excitation coefficients in the same way as the linear layer.

[0155] Step 3: Calculate the positive and negative excitations for each layer of the convolutional neural network;

[0156] Step 4: Use double-chain back propagation to obtain the saliency map of the entire convolutional neural network, as follows:

[0157] Step 4.1, parameter definition:

[0158] In any network F, given layer L nThe input and output are represented by O n-1 and O n Indicates that O n-1 is equal to X when n = 1, so the positive excitation coefficient can be obtained and negative incentive coefficient Since zero excitation does not contribute to the output, a total of 2N excitation coefficients can be obtained for all N layers in the network, with positive excitation signals being positive for positive excitations and negative for negative excitations;

[0159] Step 4.2, double chain back propagation:

[0160] The characteristic of the excitation signal is transitive, which can be expressed as:

[0161]

[0162] This transitivity is reflected in the double-chain structure, where the corresponding positive and negative excitation signals of the two layers are cross-multiplied and then summed to complete the transfer of excitation; since the excitation coefficients have the same transitivity as the excitation signals, they can also be transferred in this way;

[0163] Step 4.3: Get the saliency map of the entire network:

[0164] The output O from any layer is obtained by iterative method j To any layer output O i The positive and negative incentive coefficients are as follows:

[0165]

[0166] Where i>j; when i=j+1, When i = N and j = 0, the resulting excitation coefficient is represents the excitation coefficient of the entire network,

[0167] is the obtained saliency map, where represents the coefficient that has a positive contribution to the result for the corresponding pixel in the image, while represents the opposite coefficient.

[0168] The present invention also provides a positive and negative excitation saliency map generation system for convolutional neural networks, which is used to implement the positive and negative excitation saliency map generation method for convolutional neural networks. The system includes a first module to a fourth module, wherein:

[0169] The first module implements forward propagation based on convolutional neural network;

[0170] The second module writes corresponding positive and negative excitation extraction functions for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network;

[0171] The third module calculates the positive and negative excitations of each layer of the convolutional neural network;

[0172] The fourth module uses double-chain back-propagation to obtain the saliency map of the entire convolutional neural network.

[0173] The present invention also provides a mobile terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for generating positive and negative excitation saliency maps for convolutional neural networks is implemented.

[0174] The present invention is further described in detail below with reference to specific embodiments.

[0175] Example

[0176] This embodiment adopts the VGG16 convolutional neural network and uses the positive and negative excitation saliency map generation method for convolutional neural networks of the present invention to generate a saliency map for the ImageNet1k data set. The specific steps are as follows:

[0177] Step 1: Forward propagation of convolutional neural network;

[0178] Convolutional neural network is a type of feedforward neural network with deep structure and convolutional calculation, and is one of the representative algorithms of deep learning. The specific steps of VGG16 convolutional neural network forward propagation are as follows:

[0179] (1a) The image input size is 224*224*3. After convolution layer 1, convolution layer 2, and maximum pooling layer, the output size is 112*112*64.

[0180] (1b) After convolution layer 3, convolution layer 4, and maximum pooling layer, the output size is 56*56*128;

[0181] (1c) After convolution layer 5, convolution layer 6, convolution layer 7, and maximum pooling layer, the output size is 28*28*256;

[0182] (1d) After convolution layer 8, convolution layer 9, convolution layer 10, and maximum pooling layer, the output size is 14*14*512;

[0183] (1e) After convolution layer 11, convolution layer 12, convolution layer 13, and maximum pooling layer, the output size is 7*7*512;

[0184] (1f) The output of the previous layer is expanded into a one-dimensional vector, and the output size is 25088;

[0185] (1g) After the fully connected layer 14, the output size is 4096;

[0186] (1h) After the fully connected layer 15, the output size is 4096;

[0187] (1i) After the fully connected layer 16, the output size is 1000;

[0188] Step 2: Write corresponding positive and negative excitation extraction functions for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network. Since VGG16 does not have a normalization layer, the specific steps are as follows:

[0189] (2a) For the convolutional layers 1-13 in (1a) to (1e), use the method described in step 2.2 to write the extraction function of the positive and negative excitations of the convolutional layers;

[0190] (2b) For the maximum pooling layer in (1a) to (1e), use the method described in step 2.4 to write the extraction function of the positive and negative excitation of the pooling layer;

[0191] (2c) For the fully connected layers 14-16 in (1g) to (1i), use the method described in step 2.1 to write the extraction function of positive and negative excitations of the linear layer;

[0192] (2d) For the activation layers in (1a) to (1i), the activation layers are located after the convolutional layers 1-13 and the fully connected layers 14-16. Therefore, VGG16 has a total of 16 activation layers. Use the method described in step 2.3 to write the extraction function of the positive and negative excitations of the activation layer;

[0193] Step 3: Calculate the positive and negative excitations for each layer of the neural network. The specific steps are as follows:

[0194] (3a) Obtain the input and output of each layer, where the input of a layer is the output of the previous layer, and the output of a layer will be the input of the next layer;

[0195] (3b) obtaining the positive and negative excitation coefficients of each layer according to the positive and negative excitation extraction functions in (2a) to (2d);

[0196] Step 4: Use double-chain back propagation to obtain the saliency map of the entire neural network. The specific steps are as follows:

[0197] (4a) The double-chain back propagation of positive and negative excitation coefficients of two adjacent layers is expressed as:

[0198]

[0199] (4b) Any layer output O j To any layer output O iThe double-chain back propagation of positive and negative excitation coefficients is expressed as:

[0200]

[0201] When i = N and j = 0, the resulting excitation coefficient is Represents the incentive coefficient of the entire network; It is the saliency map; Figure 2 These are the original pictures, positive excitation pictures, negative excitation pictures, and the combined pictures of positive excitation and negative excitation corresponding to the four pictures in this embodiment.

[0202] In summary, by writing an excitation extraction function for each layer of the convolutional neural network, the present invention can utilize the complete information of each layer of the convolutional neural network without gradient, extract the positive and negative excitations of each layer, and introduce double-chain back propagation to combine the excitations into a more reliable saliency map, thereby obtaining a more convincing neural network explanation.

[0203] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A method for generating positive and negative excitation saliency maps for convolutional neural networks, characterized in that: Here are the steps: Step 1: Forward propagation of convolutional neural network; Step 2: Write corresponding positive and negative excitation extraction functions for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network; Step 3: Calculate the positive and negative excitations for each layer of the convolutional neural network; Step 4: Use double-chain back-propagation to obtain the saliency map of the entire convolutional neural network.

2. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 1, characterized in that: For the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network described in step 2, write the corresponding positive and negative excitation extraction functions, as follows: Step 2.1, write the positive and negative excitation extraction function of the linear layer of the convolutional neural network; Step 2.2, write the positive and negative excitation extraction function of the convolutional layer of the convolutional neural network; Step 2.3, write the positive and negative excitation extraction function of the convolutional neural network activation layer; Step 2.4, write the positive and negative excitation extraction function of the convolutional neural network pooling layer; Step 2.5: Write the positive and negative excitation extraction function of the convolutional neural network normalization layer.

3. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 2, characterized in that: Step 2.1 describes the preparation of the convolutional neural network linear layer positive and negative excitation extraction function, as follows: Step 2.1.1, parameter definition: The expression for setting the linear layer is: Y=w·X+b Among them, X represents input, w represents weight, b represents bias, and Y represents output; Let X∈R V ,Y∈R W ,b∈R W ,w∈R W×V , then for fixed w and b, the part of the equation related to X is Y ′ =Yb; Step 2.1.2, obtain the linear layer positive and negative excitation coefficients: Y ′ The individual elements of the positive excitation coefficient are defined as follows: Where, i∈W, j∈V; 0 is used to fill the blank part of the positive excitation. Since 0 is used as the zero element of the multiplication group and the unit element of the addition group in the real number field, it will not affect the excitation calculation; Y ′ Negative incentive coefficient Use the same method to determine; Since any Y is subject to the fixed influence of bias b, Y ′ The incentive of is essentially the incentive of Y.

4. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 2, characterized in that: Step 2.2 describes the preparation of the convolutional neural network convolutional layer positive and negative excitation extraction function, as follows: Step 2.2.1, parameter definition: The regular expression for setting the convolutional layer is: Y=X*f Among them, X represents input, f represents convolution kernel, and Y represents output; In image convolution, define X∈R W×H ,f∈R i×j , indicating the stride parameter s of the sliding step of the convolution kernel f, then the output Y∈R ([(W-i) / s]+1)×([(H-i) / s]+1) , abbreviated as Where o1 = [(Wi) / s] + 1, o2 = [(Hi) / s] + 1; Step 2.2.2, get the positive and negative excitation coefficients of the convolutional layer: The linear form of the convolutional layer is: Y trans =X trans ·f trans f trans ∈R i·j Among them, Y trans and f trans are the vector expansions of Y and f, respectively, X trans is the complex expansion of X according to f; X trans The relationship between and the elements of X is expressed as follows: Among them, k∈[0,i-1] and l∈[0,j-1], so the original convolution is equivalent to a linear layer, which obtains excitation in the way of a linear layer, expands to the input of a vector, and restores it to its original shape by directly using deconvolution; Step 2.2.3, the incentive in tensor form is as follows: For any element Y k,l ∈Y, where k∈[0,o1-1],l∈[0,o2-1], is calculated only from a slice of X, and the calculation formula is: The size of the slice is equal to the size of f. Therefore, the excitation of each element in the output Y is the corresponding slice from multiple inputs X. The specific excitation coefficient is the linear sum of all corresponding slices. When the spatial information is explicitly reflected in the excitation, the corresponding excitation is expressed as The values ​​of the positions that are the same as the slice position are equal, and the values ​​of other positions are 0.

5. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 2, characterized in that: Step 2.3 describes the preparation of the convolutional neural network activation layer positive and negative excitation extraction function, as follows: Step 2.3.1, parameter definition: In convolutional neural networks, ReLU is a commonly used activation function, which is defined as: ReLU(x)=max(x,0) in, This function is linear when x>0 and masks negative values ​​in the input signal; Step 2.3.2, obtain the positive and negative excitation coefficients of the activation layer: The coefficients of ReLU(x) can be obtained by max(x,0) / x, where x≠0; since ReLU has no negative excitation, signals with values ​​greater than zero are considered positive excitation signals, so the coefficients of all positive excitation signals are set to 1, and the coefficients of negative excitations are set to zero for subsequent calculations in this particular case.

6. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 2, characterized in that: Step 2.4 describes the writing of the convolutional neural network pooling layer positive and negative excitation extraction function, as follows: Step 2.4.1, parameter definition: Maximum pooling and average pooling are two common pooling methods. The expression of the maximum pooling function is set as: MaxPool(X)=max(X) in, This method represents the original entire input signal X by selecting the maximum value in a given region; The expression of the average pooling function is set as: Step 2.4.2, get the positive and negative excitation coefficients of the maximum pooling layer: Similar to ReLU, the maximum pooling has only positive excitation and zero excitation. The only positive excitation signal of MaxPool(X) corresponds to the element X i,j ∈X, satisfying X i,j =max(X), the positive excitation coefficient of this element is 1; Step 2.4.3, get the positive and negative excitation coefficients of the average pooling layer: The average pooling function is similar to a linear function, and all weights are set to (W×H) -1 , the bias is set to 0, and the activation is calculated using the same principle as the linear layer.

7. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 2, characterized in that: Step 2.5 describes the preparation of the convolutional neural network normalization layer positive and negative excitation extraction function, as follows: Step 2.5.1, parameter definition: Batch normalization BN is a commonly used normalization technique. The expression of BN is set as: in μ is the mean of X, σ is the variance of X, ∈ is a small bias to avoid division by zero, represents the Hadamard product, is the learning weight, is the deviation of the output; Step 2.5.2, obtain the positive and negative excitation extraction function of the normalized layer: For a given X, μ and σ are fixed and thus are treated as fixed parameters in static analysis. Rewriting the expression as a linear function gives, for any element in the output BN(X), the expression is: in, The BN layer obtains the excitation coefficients in the same way as the linear layer.

8. The method for generating positive and negative excitation saliency maps for convolutional neural networks according to claim 1, characterized in that: Use double-chain back propagation as described in step 4 to obtain the saliency map of the entire convolutional neural network, as follows: Step 4.1, parameter definition: In any network F, given layer L n The input and output are represented by O n-1 and O n Indicates that O n-1 Equal to X when n = 1, so the positive excitation coefficient is obtained and negative incentive coefficient Since zero excitation does not contribute to the output, a total of 2N excitation coefficients are obtained for all N layers in the network, with positive excitation signals being positive for positive excitations and negative for negative excitations; Step 4.2, double chain back propagation: The characteristic of the excitation signal is transitive, which can be expressed as: This transitivity is reflected in the double-chain structure, where the corresponding positive and negative excitation signals of the two layers are cross-multiplied and then summed to complete the transmission of the excitation; since the excitation coefficient has the same transitivity as the excitation signal, it can also be transmitted in the manner shown in the above formula; Step 4.3: Get the saliency map of the entire network: The output O from any layer is obtained by iterative method j To any layer output O i The positive and negative incentive coefficients are as follows: Where i>j; when i=j+1, When i = N and j = 0, the resulting excitation coefficient is represents the excitation coefficient of the entire network, is the obtained saliency map, where represents the coefficient that has a positive contribution to the result for the corresponding pixel in the image, while represents the opposite coefficient.

9. A positive and negative excitation saliency map generation system for convolutional neural networks, characterized in that: The system is used to implement the method for generating positive and negative excitation saliency maps for convolutional neural networks according to any one of claims 1 to 8, and the system includes a first module to a fourth module, wherein: The first module implements forward propagation based on convolutional neural network; The second module writes corresponding positive and negative excitation extraction functions for the linear layer, convolution layer, activation layer, pooling layer, and normalization layer of the convolutional neural network; The third module calculates the positive and negative excitations of each layer of the convolutional neural network; The fourth module uses double-chain back-propagation to obtain the saliency map of the entire convolutional neural network.

10. A mobile terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for generating positive and negative excitation saliency maps for convolutional neural networks is implemented.