An Enhancement Method for Visualization Algorithm of Image Classification Neural Network Based on Sliding Window Mechanism
By introducing a sliding window mechanism into the feature visualization algorithm of deep neural networks, the problem of low resolution and high noise in the prior art feature visualization results is solved, and a more refined and smooth display of feature information is achieved.
Patent Information
- Application Number
- CN202310428053.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-04-20
AI Technical Summary
In the prior art, deep neural networks are difficult to explain in picture classification tasks, and the feature visualization results have low resolution and high noise, making it difficult to provide convincing evidence.
The image classification neural network visualization algorithm based on the sliding window mechanism is used to intercept the local area of the original input image through the sliding window, generate several window pictures, and upsample them to the resolution of the original input image, and input the model to obtain the probability score of the specified category as weight. Combined with the feature visualization algorithm to generate significant images, and perform low-pass filtering and normalization processing to obtain the final enhanced significant images.
The significant graph resolution and information density of feature visualization algorithm are improved, and the positioning ability of key features that are interested in the model in the input picture is enhanced, so that the generated significant graphs are more refined and smooth.
Smart Images

Figure CN116403050B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of feature visualization of deep neural networks, and specifically to an enhancement method for an image classification neural network visualization algorithm based on a sliding window mechanism. Background Art
[0002] The statements in this part merely provide background technical information related to the present disclosure, and these statements may constitute prior art. In the process of implementing the present invention, the inventors found that at least the following problems exist in the prior art.
[0003] Currently, deep neural networks are widely applied to image classification tasks. Regarding the characteristic that deep neural networks are difficult to interpret, although there are many algorithms for visualizing the decision-making mechanism inside deep neural networks, most of these algorithms can only provide rough and noisy visualization results, and it is difficult to provide convincing evidence in some scenarios.
[0004] Feature visualization methods based on class activation mapping are a popular type of method, which have received extensive research and application. This type of method usually has class discriminability, and it is based on the feature maps output by the convolutional layers inside the CNN. Therefore, its interpretation effect and visualization images have gained the trust of researchers. The quality of the heatmaps generated by this type of method depends very much on the position of the convolutional layer selected in the network. Usually, this type of method will select deep convolutional networks, such as the convolutional network closest to the output layer, because the visualization images have clear class discriminability due to the rich class information it contains. However, the resolution of deep feature maps is very low, and it cannot contain more detailed information in the class. Shallow feature maps have a higher resolution, but they are full of noise and lack class information. Decomposition-based algorithms have a solid theoretical basis, with deep Taylor decomposition as the basic theoretical framework. This type of method infers from the network output end to the input end, decomposing the decision made by the current layer of the network into the contribution of the previous layer of the network until it infers to the corresponding elements in the input. However, this type of method needs to formulate inference rules for different network types, and some of these methods lack class sensitivity, and the visualization images generated by them also have problems of low resolution and a lot of noise. Summary of the Invention
[0005] In view of the above problems, an object of the present invention is to solve a part of the problems in the prior art, or at least alleviate these problems.
[0006] An enhancement method for an image classification neural network visualization algorithm based on a sliding window mechanism includes:
[0007] Using a sliding window algorithm to intercept local regions in the original input image to obtain a number of window images and the coordinate positions of the window images in the original input image;
[0008] Upsample the window image to the resolution of the original input image;
[0009] Input the upsampled window image into the model Obtain the probability score of the specified class index c of the window image as the weight;
[0010] Input the image set of the upsampled window image, the specified class index c, and the model As parameters into the feature visualization algorithm to be enhanced to obtain the corresponding window image saliency map;
[0011] Input the original input image into the feature visualization algorithm to be enhanced to obtain the corresponding original saliency map;
[0012] Downsample the window image saliency map and multiply it by its corresponding weight, and then add it to the pixel values of the corresponding region of the original saliency map according to its coordinate position to obtain the initial enhanced saliency map.
[0013] The enhancement method of the visualization algorithm of the image classification neural network based on the sliding window mechanism further includes performing low-pass filtering on the initial enhanced saliency map and performing min-max normalization to obtain the final enhanced saliency map.
[0014] Furthermore, downsampling the window image saliency map and multiplying it by its weight, and then adding it to the pixel values of the corresponding region of the original saliency map according to its coordinate position to obtain the initial enhanced saliency map, including the following steps:
[0015] Perform interpolation processing on the window image saliency map to make its resolution the same as that of the window image;
[0016] Multiply each pixel value of the interpolated window image saliency map by its corresponding probability score, and then add it to the corresponding pixel value in the original saliency map S 0 According to the following formula:
[0017]
[0018] Among them, S′ c Is the initial enhanced saliency map obtained after addition, Is the set of window image saliency maps, the set of probability scores Is the upsampled window image Regarding the confidence level of the class index c, Is the upsampled and interpolated window image The saliency map of the specified class index c obtained after downsampling and interpolation processing input into the specified feature visualization algorithm, and the resolution of this saliency map is reduced to be the same as that of the sliding window image; the above k 1Indicates the coordinates of the sliding window in the original input image when the window image is intercepted for the first time; k n Indicates the coordinates of the sliding window in the original input image when the window image is intercepted for the nth time.
[0019] Before obtaining a number of window images by intercepting local regions in the original input image using the sliding window algorithm, it further includes specifying in advance the parameters required before performing the enhancement algorithm; the parameters include but are not limited to a pre-trained convolutional neural network model Original input image Specify the class index c, the interpolation function φ(.), the feature visualization algorithm f, the width w and height h of the sliding window, the number of pixels moved by the sliding window each time stride, and the starting position left start of the sliding window.
[0020] Further, a sliding window with a specified width and height of w and h is located at the position with coordinates start in the original input image, where h and w are respectively smaller than the height H 0 and width W 0 ; the starting position left start of the sliding window is a coordinate on a two-dimensional coordinate axis, the origin of the coordinate axis is the upper left vertex of the original input image, the positive direction of the X axis points to the upper right vertex, and the positive direction of the Y axis points to the lower left vertex; the Start coordinate of (0, 0) means that the upper left vertex of the sliding window coincides with the upper left vertex of the original input image, i.e., the coordinate axis.
[0021] Preferably, the interpolation function is a bilinear interpolation function.
[0022] Further, using the sliding window algorithm to intercept local regions in the original input image to obtain a number of window images, and the coordinate positions of the window images in the original input image, including the following steps:
[0023] The sliding window remains inside the original input image and slides at a fixed stride;
[0024] In each sliding process, the image region within the sliding window is intercepted and stored in the window image set ; among them,
[0025] During the interception process, the current coordinate position of the sliding window is recorded as the subscript of the intercepted image within the current window, that is, the image of the nth sliding window screenshot and its attached coordinate position.
[0026] Further, the sliding window remains inside the original input image and slides at a fixed stride, including the following steps:
[0027] The sliding window starts from the origin and moves forward by a stride of stride pixel sizes along the positive X-axis each time;
[0028] If the right boundary of the sliding window coincides with the right boundary of the original input image, it slides downward by a stride of stride pixel sizes and then slides leftward, always keeping the sliding window completely inside the original input image.
[0029] Model It includes an image classification neural network model; the image classification neural network model includes, but is not limited to, the architecture of the target neural network of the feature visualization algorithm.
[0030] A computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, it implements the steps of the enhanced method of the image classification neural network visualization algorithm based on the sliding window mechanism.
[0031] The present invention has the following beneficial effects:
[0032] 1. The present invention uses a sliding window to intercept local regions in the picture for feature visualization to generate a fine saliency map corresponding to the local region, thereby enhancing the saliency map of the current feature visualization algorithm, and making up for the problems of low resolution and insufficient fine positioning of target features in the current feature visualization algorithm;
[0033] 2. The present invention is a general feature visualization enhancement algorithm, which can enhance almost all current feature visualization algorithms, does not need to modify the internal calculation process of the current feature visualization algorithm, and at the same time, the present invention is not limited by the architecture of the image classification network;
[0034] 3. The present invention can be applied to the feature visualization algorithm, making the saliency map of the algorithm present more fine feature information, thereby verifying the rationality of the algorithm and analyzing the feature patterns learned by the neural network. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a general flowchart of the enhanced method of the image classification neural network visualization algorithm based on the sliding window mechanism;
[0036] Figure 2 It is a detailed example diagram of the enhanced method of the image classification neural network visualization algorithm based on the sliding window mechanism;
[0037] Figure 3This is an enhanced saliency map display for the feature visualization algorithms Grad-CAM, LRP, partial LRP, and Transformer LRP applied in the present invention based on the ViT model, as well as an enhanced saliency map display for the feature visualization algorithms Grad-CAM and Grad-CAM++ applied in the VGG19 model. The saliency map within the rectangular frame is the enhanced saliency map. Detailed implementation mode
[0038] The following further describes the present invention with reference to the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention rather than limit it. Without departing from the technical idea of the present invention, various substitutions and changes made according to the common general knowledge and conventional means in the art shall be included within the scope of the present invention.
[0039] Aiming at the problems of low resolution and fuzzy target feature localization in the current feature visualization algorithm, the present invention proposes an enhancement method for the visualization algorithm of an image classification neural network with a sliding window mechanism. As Figure 1 shown in Figure 2, an enhancement method for the visualization algorithm of an image classification neural network with a sliding window mechanism includes:
[0040] Using the sliding window algorithm to intercept local regions in the original input image to obtain a number of window images and the coordinate positions of the window images in the original input image;
[0041] Upsampling the window images to the resolution of the original input image;
[0042] Inputting the upsampled window images into the model Obtaining the probability score of the specified class index c of the window images as the weight;
[0043] Taking the set of upsampled window images, the specified class index c, and the model as parameters and inputting them into the feature visualization algorithm to be enhanced to obtain the corresponding saliency maps of the window images;
[0044] Inputting the original input image into the feature visualization algorithm to be enhanced to obtain the corresponding original saliency map;
[0045] Downsampling the saliency maps of the window images, multiplying them by their corresponding weights, and then adding them to the pixel values of the corresponding regions of the original saliency map according to their coordinate positions to obtain the initial enhanced saliency map.
[0046] The present invention uses a sliding window to intercept local regions in the original image for feature visualization to generate saliency maps to enhance the quality of the saliency map of the original input image. Therefore, the set of saliency maps of the obtained window images is used to optimize S 0 , using The coordinates attached to each significant map can accurately locate the area of the low-resolution significant map corresponding to the original input image, and also directly correspond to the original significant map S 0 The area to be enhanced in. It can enhance all current feature visualization algorithms, improve the resolution and information density, and better locate the key features of interest to the model in the input image. Enhance any current feature visualization algorithm to improve their ability to locate target features.
[0047] The enhancement method for the visualization algorithm of the image classification neural network based on the sliding window mechanism further includes performing low-pass filtering on the initial enhanced significant map and performing min-max normalization to obtain the final enhanced significant map. Since the present invention uses the sliding window mechanism and the stride of each slide is fixed, it results in the initial enhanced significant map S' obtained by enhancement c There will be a grid phenomenon, that is, the transition of significant values between different windows is not smooth. Therefore, the present invention also proposes to use low-pass filtering to process S' c For further processing, so that the finally generated significant map is smoother. The processing process of the low-pass filter function is as follows:
[0048] S″ c = Δ(S' c , H(X,Y))
[0049] where H(X,Y) is the transfer function of the ideal low-pass filter, and its definition is as follows:
[0050]
[0051] where (X,Y) represents the coordinates in the frequency domain, D(X,Y) represents any point in the frequency domain, and D 0 is the cut-off frequency. The formula shows that for the frequency components below the cut-off frequency D 0 , its output value is 1, that is, retained; for the frequency components above the cut-off frequency D 0 , its output value is 0, that is, removed. Therefore, the essence of the ideal low-pass filter is to perform a filtering operation on the frequency domain through the filter function H(X,Y) in the frequency domain. The filtered significant map S″ c After performing min-max normalization, the final enhanced significant map S c can be obtained: It can only be used as a significant map for intuitive display after pseudo-color conversion. The present invention is not limited to a single pseudo-color conversion scheme.
[0052] Downsample the window image significant map, multiply it by its weight, and then add the pixel values according to its coordinate position and the corresponding area of the original significant map to obtain the initial enhanced significant map, including the following steps:
[0053] Interpolate the significant map of the window image to make its resolution the same as that of the window image;
[0054] Multiply each pixel value of the interpolated significant map of the window image by its corresponding probability score, and then add it to the corresponding pixel value in the original significant map S according to its coordinate position 0 as follows:
[0055]
[0056] where S′ c is the initial enhanced significant map obtained after addition, is the set of significant maps of window images, and the set of probability scores is the upsampled window image is the confidence level of the upsampled window image about the class index c, is the significant map after downsampling interpolation about the class index c obtained by inputting the upsampled and interpolated window image into the specified feature visualization algorithm, and the resolution of this significant map is reduced to be the same as that of the sliding window image. The above k 1 represents the coordinates of the sliding window in the original input image when the window image is intercepted for the first time; k n represents the coordinates of the sliding window in the original input image when the window image is intercepted for the nth time.
[0057] Before using the sliding window algorithm to intercept local regions in the original input image to obtain several window images, it also includes pre-specifying the parameters required before executing the enhancement algorithm; the parameters include but are not limited to the pre-trained convolutional neural network model the original input image the specified class index c, the interpolation function φ(.), the feature visualization algorithm f, the width w and height h of the sliding window, the number of pixels moved by the sliding window each time stride, and the starting position left start of the sliding window.
[0058] The sliding window with a specified width and height of w and h is located at the position with coordinates start in the original input image, where h and w are respectively smaller than the height H 0 and width W 0 of the input image; the starting position left start of the sliding window is a coordinate on a two-dimensional coordinate axis, the origin of the coordinate axis is the upper left vertex of the original input image, the positive direction of the X-axis points to the upper right vertex, and the positive direction of the Y-axis points to the lower left vertex; The start coordinate of (0, 0) means that the upper left vertex of the sliding window coincides with the upper left vertex of the original input image, that is, the coordinate axes coincide.
[0059] The starting position of the sliding window, start, is set as a coordinate on a two-dimensional coordinate axis, which facilitates the calculation that h and w can be divisible by H 0 and W 0 evenly.
[0060] In the selection of the interpolation function, the present invention preferentially considers that the interpolation function is a bilinear interpolation function.
[0061] Using the sliding window algorithm to intercept local regions in the original input image to obtain several window images and the coordinate positions of the window images in the original input image, including the following steps:
[0062] The sliding window slides inside the original input image at a fixed stride;
[0063] During each sliding process, the image region within the sliding window is intercepted and stored in the window image set ; where
[0064] During the interception process, the current coordinate position of the sliding window is recorded as the subscript of the intercepted image within the current window, that is, the image and its attached coordinate position of the nth sliding window screenshot.
[0065] The sliding window slides inside the original input image at a fixed stride, including the following steps:
[0066] The sliding window starts from the origin and moves forward by stride pixel sizes along the positive X-axis each time;
[0067] If the right boundary of the sliding window coincides with the right boundary of the original input image, it slides down by stride pixel sizes and then slides left, always keeping the sliding window completely inside the original input image.
[0068] Using the above window sliding algorithm, all local regions in the input image can be upsampled to the size of the input image using a sliding window of a fixed size.
[0069] Model includes an image classification neural network model; the image classification neural network model includes, but is not limited to, the architecture of the target neural network of the feature visualization algorithm.
[0070] The present invention is specifically implemented through the following steps:
[0071] (1) Specify a sliding window with width and height w and h respectively at the position with coordinates start in the original input image, where h and w are respectively less than the height H 0 and width W 0 of the input image, and for the convenience of calculation, h and w can be divisible by H 0and W 0 is divisible. Start is a coordinate on a two - dimensional coordinate axis. The origin of this coordinate axis is the upper - left vertex of the original input image. The positive direction of the X - axis points to the upper - right vertex, and the positive direction of the Y - axis points to the lower - left vertex. The start coordinate of (0, 0) means that the upper - left vertex of the sliding window coincides with the upper - left vertex of the original input image, i.e., the coordinate axes coincide.
[0072] (2) Starting from the origin, the sliding window moves stride pixels in the positive X - axis direction each time. If the right - hand boundary of the sliding window coincides with the right - hand boundary of the image, it slides down stride pixels and then slides the window to the left, always keeping the window completely within the image. During each sliding process, the image area within the window is intercepted and stored in the window image set .
[0073]
[0074] During the interception process, the coordinate position of the current window is recorded as the subscript of the image intercepted within the current window, i.e., the image and its attached coordinate position of the n - th screenshot of the sliding window.
[0075] (3) For each single image in the window image set of the sliding window in (2) apply the up - sampling interpolation function φ(.) to each image to upsample it to the resolution of the original input image: where is the set of up - sampled window images, represents the up - sampled image of the window image intercepted by the sliding window for the n - th time in this set.
[0076] (4) Input the images in the set of up - sampled window images into the model to obtain the probability score of the specified class index c for each image. This score can reflect the confidence level of the specified class index c in this image. Put the probability scores of all images regarding the specified class index c into the probability score set :
[0077]
[0078] where represents the confidence level of the image regarding the class index c. represents obtaining the probability score of the model regarding the class index c.
[0079] (5) Select the feature visualization algorithm f to be enhanced. The feature visualization algorithm needs to specify the input image, the neural network model, and the class index. The set of window images in (3) Specify the class index c and the model As parameters, input them into the visualization algorithm f to obtain the set of window saliency maps where The corresponding window saliency map is This saliency map has been upsampled and has the same resolution as the original input image.
[0080] (6) Similar to (5), input the original input image I 0 into the feature visualization algorithm f, keep other algorithm parameters unchanged, and obtain the original saliency map S for the specified class index c 0 .
[0081] (7) Interpolate the set of window saliency maps in (5) so that the resolution of each saliency map is reduced to the same as the resolution of the cropped image in the sliding window, that is, the same as the resolution of the images in the set of window images . The interpolation process is as follows:
[0082]
[0083] where is the set of downsampled window saliency maps.
[0084] (8) Multiply each pixel value of the saliency map in by its corresponding probability score in , and add it to the corresponding pixel in S 0 according to the coordinate position. The addition process follows the following formula:
[0085]
[0086] where S′ c is the initial enhanced saliency map obtained after addition.
[0087] (9) Use low-pass filtering to further process S′ c to make the finally generated saliency map smoother. The low-pass filtering function processing process is as follows: S″ c = Δ(S′ c , H(X, Y)), where H(X, Y) is the transfer function of the ideal low-pass filter, and its definition is as follows:
[0088]
[0089] where D0 is the cut-off frequency.
[0090] (10) The filtered saliency map S″ in (9) c is subjected to min-max normalization to obtain the final enhanced saliency map S c :
[0091]
[0092] S c also needs to go through min-max normalization and pseudo-color conversion before it can be used as a saliency map for intuitive display.
[0093] The following combines Figure 2 and specific embodiments to further elaborate on the present invention in detail.
[0094] The implementation examples and implementation situations of the complete method according to the invention content of the present invention are as follows:
[0095] The embodiments are described as follows:
[0096] 1. The feature visualization algorithm to be enhanced in the present invention needs to be adapted to the used image classification model. The pre-trained image classification model here comes from Torchvision, which is a vision toolkit based on the machine learning framework pytorch. It mainly includes four parts: dataset, data transformation, utility functions, and pre-trained models. Here, the VGG19 model trained on the ImageNet dataset provided by Torchvision is used as the target model and the feature visualization algorithm Grad-CAM for illustration. Among them, the ImageNet dataset is a large-scale image recognition dataset, containing more than one million manually annotated images with labels. VGG19 is a classic convolutional neural network model, which contains 19 convolutional layers and 3 fully connected layers and can be used for tasks such as image classification, object recognition, and target detection. Grad-CAM is a visual interpretation technique that can be used to understand and visualize the decision-making process of deep learning models.
[0097] 2. Select the original input picture. The original input picture should be adjusted to have the same length and width. For example, both the length and width are 224 pixel sizes. Apply the sliding window algorithm to intercept the picture area in each window on the input picture. Here, the length and width of the sliding window can be set to 96 pixel sizes, and the stride can be set to 32 pixel sizes.
[0098] 3. Upsample all the window pictures in 2 to a resolution of 224×224, and then input them into the VGG19 model to obtain the probability scores of each window picture for the specified class as weights.
[0099] 4. Select the target convolutional layer. For example, select the last convolutional layer "features.34" of VGG19. Apply Grad-CAM to all window images to obtain the saliency map corresponding to each image. Here, the resolution of the saliency map is 224×224.
[0100] 5. Similar to step 4, apply Grad-CAM to the original input image to obtain the original saliency map corresponding to the original input image.
[0101] 6. Downsample all the saliency maps obtained in step 4 to the resolution of the sliding window, i.e., 96×96.
[0102] 7. Multiply the saliency maps in step 6 by the corresponding weights in step 3, and then add the pixel values of the corresponding regions in the original saliency map in step 5 to obtain the initial enhanced saliency map.
[0103] 8. Apply ideal low-pass filtering to the initial enhanced saliency map.
[0104] 9. Normalize the saliency map obtained in step 8 using min-max normalization to obtain the final enhanced saliency map.
[0105] Figure 3 This is the display of the enhanced saliency map for the feature visualization algorithms Grad-CAM, LRP, partial LRP, and Transformer LRP applied to the present invention based on the ViT model, as well as the display of the enhanced saliency map for the feature visualization algorithms Grad-CAM and Grad-CAM++ applied to the VGG19 model. The saliency map within the rectangular frame is the enhanced saliency map. Among them, ViT (Vision Transformer) is a new type of image classification model. Different from traditional convolutional neural networks (CNNs), it uses the self-attention mechanism to capture features in images. Grad-CAM is a feature visualization technique that can be used to understand and visualize the decision-making process of deep learning models. Grad-CAM++ is a generalized improved version of Grad-CAM. Layer-wise relevance propagation (LRP) is a method for interpreting the output of neural networks, which can reveal the contribution degree of each input feature in the model classification decision. Partial LRP is an improved version of LRP, and Transformer LRP is the best feature visualization technique currently applied to the ViT model.
[0106] Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art and related fields without creative efforts shall fall within the protection scope of the present invention.
[0107] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the enhanced method of the visualization algorithm of the image classification neural network based on the sliding window mechanism are implemented.
Claims
1. An enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism, characterized in that, comprising: Use the sliding window algorithm to intercept local regions in the original input image to obtain a number of window images and the coordinate positions of the window images in the original input image, including the following steps: The sliding window remains inside the original input image and slides at a fixed stride; During each sliding process, intercept the image region within the sliding window and store it in the window image set ; Among them, During the interception process, record the current coordinate position of the sliding window as the subscript of the intercepted image within the current window, that is, the image of the nth sliding window screenshot and its attached coordinate position; Upsampling the window image to the resolution of the original input image; Input the upsampled window image into the model Obtain the probability score of the specified class index c of the window image as the weight; Input the image set of the upsampled window image, the specified class index c, and the model as parameters into the feature visualization algorithm to be enhanced, and obtain the corresponding window image saliency map; Inputting the original input image into the feature visualization algorithm to be enhanced to obtain the corresponding original saliency map; Downsample the significant map of the window image, multiply it by its corresponding weight, and then add it to the pixel values of the corresponding region of the original significant map according to its coordinate position to obtain an initial enhanced significant map, including the following steps: perform interpolation processing on the significant map of the window image to make its resolution the same as that of the window image; multiply each pixel value of the interpolated significant map of the window image by its corresponding probability score, and then add it to the corresponding pixel value in the original significant map S 0 according to its coordinate position, following the formula: Among them, S′ c is the initial enhanced saliency map obtained after addition, is the set of saliency maps of window images and the set of probability scores is the window image after upsampling is the confidence level regarding the class index c, is the window image after upsampling interpolation is the saliency map regarding the class index c obtained by inputting into the specified feature visualization algorithm and undergoing downsampling interpolation processing. The resolution of this saliency map is reduced to be the same as that of the sliding window image; k 1 represents the coordinates of the sliding window in the original input image when the window image is intercepted for the first time; k n represents the coordinates of the sliding window in the original input image when the window image is intercepted for the nth time; Performing low-pass filtering on the initial enhanced saliency map and performing min-max normalization to obtain the final enhanced saliency map.
2. The enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism according to claim 1, characterized in that, Before obtaining a number of window images by intercepting local regions in the original input image using the sliding window algorithm, it further includes pre-specifying the parameters required before executing the enhancement algorithm; the parameters include but are not limited to a pre-trained convolutional neural network model Original input image Specify the class index c, interpolation function φ(.), feature visualization algorithm f, width w and height h of the sliding window, number of pixels moved by the sliding window each time stride, and starting position coordinates start of the sliding window.
3. The enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism according to claim 2, characterized in that, A sliding window with a specified width of w and height of h is located at the position with coordinates start in the original input image, where h and w are respectively smaller than the height H and width W of the input image. 0 and width W 0 ; The start position on the left side of the sliding window is a coordinate on a two-dimensional coordinate axis. The origin of the coordinate axis is the upper left vertex of the original input image. The positive direction of the X-axis points to the upper right vertex, and the positive direction of the Y-axis points to the lower left vertex; The start coordinate of (0, 0) means that the upper left vertex of the sliding window coincides with the upper left vertex of the original input image, i.e., the coordinate axis.
4. The enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism according to claim 2, characterized in that, The interpolation function is a bilinear interpolation function.
5. The enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism according to claim 1, characterized in that, The sliding window slides inside the original input image with a fixed stride, including the following steps: The sliding window starts from the origin and moves forward by stride pixel sizes along the positive X-axis each time; If the right boundary of the sliding window coincides with the right boundary of the original input image, it slides downward by stride pixel sizes and then slides leftward, always keeping the sliding window completely inside the original input image.
6. The enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism according to claim 1, characterized in that, Model including an image classification neural network model; the image classification neural network model includes, but is not limited to, the architecture of a target neural network for a feature visualization algorithm.
7. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the steps of the enhancement method for the visualization algorithm of an image classification neural network based on a sliding window mechanism according to any one of claims 1 to 6 are implemented.