Visual interpretability method and device for small target detection model and medium

By introducing feature image pixel-level gradient weighting and sequencing grouping processing in the CAM method, combining the integral method of global decision contribution, a high-fine-grained mask is generated, which solves the problem of insufficient combination of local sensitivity and global decision contribution in small object detection, and achieves more efficient visual interpretation and accurate target positioning.

CN120374948AActive Publication Date: 2025-07-25CHINESE PEOPLES LIBERATION ARMY ARMY ARTILLERY & AIR DEFENSE ACAD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510456207.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The existing CAM visual interpretation methods are not fully combined with local sensitivity and global decision-making contribution in small object detection, resulting in poor visual interpretation effect and timeliness, large calculation volume and high storage occupancy.

Method used

The deep feature map is obtained through CNN forward transfer, and pixel-level gradient weighting and sequencing grouping of feature images are processed to generate high-fine-grained local sensitivity masks, and global decision contributions are obtained through fuzzification masks and integral sampling, and finally, linear weighting is performed to generate visual interpretation maps.

Benefits of technology

The visual accuracy and algorithm speed of small object detection are improved, and multiple similar targets can be accurately positioned in complex backgrounds, reducing random noise, and having better visual continuity and small background interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374948A_ABST
    Figure CN120374948A_ABST
Patent Text Reader

Abstract

The invention discloses a visual interpretability method for a small target detection model, and the method comprises the steps: S1, obtaining a deep feature map through CNN forward transmission, and obtaining a corresponding gradient of the feature map through specific class target classification score back propagation; s2, carrying out feature image pixel-level gradient weighting, and obtaining a mask containing high-fine-granularity local sensitivity information based on sequencing grouping processing; s3, inputting the fuzzified mask occlusion image into the CNN, obtaining global decision contribution through forward transmission including N times of integral sampling, calculating to obtain contribution importance scores, and taking a mean value as a group importance score; and S4, performing linear weighting on the grouped masks to obtain a class activation graph, and weighting the masks by using the importance score to obtain a visual interpretation graph. According to the method, the importance scores of the grouped masks are calculated in parallel based on the global decision contribution, the masks are processed by adopting a fuzzy boundary and integration method, and the sensitivity of different inputs to the global decision contribution and the algorithm speed are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision data processing, and particularly relates to a small target visual interpretation method, device, and storage medium based on model weight correlation and integral score. Background Art

[0002] The visual interpretation method is a technology that analyzes the decision results of deep neural network models represented by convolutional neural networks through visualization means, and its goal is to help users understand the internal mechanism of the model and trust the model decision.

[0003] With the rapid development and improvement of the visual interpretation method, the existing technology generally adopts the visual interpretation method based on class activation mapping (CAM) for use. This CAM visual interpretation method forms a class activation map by weighted combination of the feature maps of the last convolutional layer, and uses the image pixel value size to represent the correlation degree of the input at the corresponding position to the output result.

[0004] Although the CAM visual interpretation method has become a hot direction in this field due to its intuitive image and good interpretability, the existing CAM method has problems of poor visual interpretation effect and timeliness caused by the lack of full and effective combination of local sensitivity and global decision contribution.

[0005] Therefore, this application specifically proposes a visual interpretability method for small target detection models to solve the above technical problems. Summary of the Invention

[0006] The main purpose of the present invention is to provide a visual interpretability method for small target detection models to solve the technical problems proposed in the background art.

[0007] The present invention adopts the following technical solutions to solve the above technical problems:

[0008] A visual interpretability method for small target detection models includes the following steps:

[0009] S1. Obtain a deep feature map through the forward propagation of the CNN, and obtain the corresponding gradient of the feature map through the backpropagation of the classification score of a specific class target;

[0010] S2. After the pixel-level gradient weighting of the feature map and the processing based on ordinal grouping, obtain a set of masks containing high-fine-grained local sensitivity information;

[0011] S3. Input the blurred mask occluded image into the CNN, obtain the global decision contribution through the forward propagation including N times of integral sampling, calculate and obtain the contribution importance score, and take the average as the group importance score;

[0012] S4. Linearly weight the grouped mask to obtain the class activation map, and the class activation map weights the mask using the importance score to obtain the visual interpretation map.

[0013] Preferably, the specific process of feature image pixel-level gradient weighting in step S2 includes:

[0014] S21. Let f be the image classification CNN model and θ be the model parameters. For the given input image I0, obtain the feature map weights through global average pooling After the forward inference of the network, the classification score formula y c is:

[0015]

[0016] where represents the feature map of the k-th channel, and i and j respectively represent the row index and column index in the feature map;

[0017] S22. Explicitly encode each feature map weight c in the classification score formula y The calculation formula of the feature map weight is:

[0018]

[0019] where is the weighting coefficient of the pixel gradient corresponding to class c and the feature map, and the calculation formula is:

[0020]

[0021] where, is the pixel gradient value of the k-th channel at the (i, j) position of feature map A;

[0022] S23. Weight each pixel gradient at the (i, j) position of feature map A to obtain K groups of weighted feature maps

[0023] Preferably, the generation process of obtaining the mask by ordinal grouping processing in step S2 includes:

[0024] Divide the K groups of weighted feature maps equidistantly into G groups in a fixed order of the generated weighted feature maps, and G < K. Stack each group of weighted feature activation maps to obtain G masks M for generating the class activation map l , and the calculation formula is:

[0025]

[0026] Normalize the obtained low-resolution mask and upsample it to the original image size by bilinear interpolation to obtain a smoother mask. The calculation formula is as follows:

[0027]

[0028] M’ l is the processed smooth mask.

[0029] Preferably, the specific calculation process for obtaining the global decision contribution in step S3 includes:

[0030] S31. Apply Gaussian blur to the occluded area to obtain a blurred mask occluded image I’. l , and the calculation expression is:

[0031]

[0032] Among them, is the input image after Gaussian blur, and ⊙ is the Hadamard product;

[0033] S32. Gradually accumulate the input image features with linear interpolation as the integration path, calculate the forward pass classification scores on the target class N times and take the average as the group mask importance score to obtain a more accurate global decision contribution corresponding to the input.

[0034] Preferably, the specific operation process for calculating the contribution importance score in step S3 includes:

[0035] Let the output f c (M”0) of the baseline image M’0 be zero. Take a pure black image or a noise image as the baseline image and execute a method without gradients. We have:

[0036]

[0037] Among them, is the importance score of the l-th group of masks, and N is the number of integral sampling;

[0038] Perform the forward pass of N masks in parallel.

[0039] Preferably, in step S4, the importance scores are used to weight the G groups of masks to obtain a more refined and accurate visual interpretation map. The calculation expression is:

[0040]

[0041] Among them is the grouped mask of the i-th group of baseline images.

[0042] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method.

[0043] In yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which, when executed by the processor, causes the processor to execute the steps of the above method.

[0044] As can be seen from the above technical solutions, the present invention provides a method for visual interpretability of a small target detection model. Compared with the prior art, the present invention has the following advantages:

[0045] 1. The present invention mines high-fine-grained local sensitivity information in the feature map through pixel-level gradient weighting, generates a weighted feature map with strong correlation with the decision result, and merges the weighted feature maps in an ordered grouping manner, so as to obtain a mask that contains local sensitivity information, has low redundancy, and strong decision correlation, thereby improving the visual accuracy of the target detection result and facilitating use.

[0046] 2. The present invention calculates the importance scores of the grouped masks in parallel based on the global decision contribution, and processes the masks by using the fuzzy boundary and integral methods, improving the sensitivity of different inputs to the global decision contribution and the algorithm speed.

[0047] 3. The present invention combines the non-gradient global contribution idea, the model operation result has less random noise, can generate a smoother class activation mapping graph, is more concentrated, accurate and complete in locating the target area, has better visual continuity, and is convenient for reflecting the model classification decision.

[0048] 4. The present invention can discover and locate multiple same-type targets in a complex background with little background interference; at the same time, it is more effective in activating small targets, and the positioning area is more accurate, concentrated and complete, with better visual effects.

[0049] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Of course, any product implementing the present invention does not necessarily need to achieve all the above advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The schematic diagram of the drawings forming a part of this application is used to provide a further understanding of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0051] Figure 1 is a schematic diagram of the LG-CAM method of the present invention;

[0052] Figure 2 Schematic diagram for identifying problems of the existing CAM method

[0053] Figure 3 Schematic diagram for comparing the channel-level weight distribution of the feature activation map of the present invention

[0054] Figure 4 Schematic diagram for comparing the visual interpretation diagrams of different methods of the present invention

[0055] Figure 5 Schematic diagram for identifying the distinguishability of different categories of the present invention

[0056] Figure 6 Schematic diagram for comparing the insertion and deletion metric curves of the present invention

[0057] Figure 7 Schematic diagram for the target localization effect of the present invention

[0058] Figure 8 Schematic diagram for comparing the localization accuracy under different thresholds of the present invention Detailed implementation manners

[0059] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] In the embodiment, refer in detail to Figures 1 to 8 .

[0061] In the prior art, the gradient-free CAM method represented by ScoreCAM uses multiple groups of feature maps as masks to multiply with the original image points to occlude the original image, and inputs the perturbed image into the neural network. The importance of each layer of feature maps is measured by the difference between the forward transfer classification confidence score on the target class and the baseline, reflecting the overall decision-making of the CNN, as shown in the formula:

[0062]

[0063] wherein, I0 is the given input image, ⊙ is the Hadamard product, I b is the baseline image, is the input feature map, is the feature map upsampled to the size of the input image, and f(·) is the calculation of the confidence score.

[0064] Although a relatively accurate class activation mapping diagram can be obtained in the way of global decision-making contribution, when obtaining the importance score, it is limited by the selection of the baseline image and the input interference image, lacks sensitivity to the global decision-making contribution of different input features, and ignores the local sensitivity information in the forward propagation. The interpretability performance needs to be improved. At the same time, the forward inference has a high computational cost and a large memory occupation, resulting in poor real-time performance of the algorithm and prone to problems such as Figure 2 the positioning deviation, artifacts, noise, etc. shown

[0065] Therefore, to solve the existing technical problems as Figure 2 shown, referring to Figure 1 , an interpretable method for small target detection model vision (LG-CAM) is proposed in an embodiment of the present invention, including the following steps:

[0066] S1. Obtain a deep feature map through the forward propagation of the CNN, and obtain the corresponding gradient of the feature map through the backpropagation of the classification score of a specific class target.

[0067] S2. Obtain a group of masks containing high-fine-grained local sensitivity information through the pixel-level gradient weighting of the feature map and based on the ordered grouping method.

[0068] Among them, the specific process of pixel-level gradient weighting of the feature map includes:

[0069] S21. Let f be the image classification CNN model and θ be the model parameters. For the given input image I0, obtain the feature map weight through global average pooling After the forward inference of the network, the classification score formula y c is:

[0070]

[0071] where represents the feature map of the kth channel, and i and j respectively represent the row index and column index in the feature map;

[0072] S22. Explicitly encode each feature map weight c in the classification score formula y , and the calculation formula of the feature map weight is:

[0073]

[0074] where is the weighting coefficient of the pixel gradient corresponding to class c and the feature map, and the calculation formula is:

[0075]

[0076] where, is the pixel gradient value of the K-th channel at the (i, j) position of feature map A;

[0077] The high-fine-grained local sensitivity in the feature map is mined through pixel-level gradient weighting, enabling better activation of multiple similar targets as a whole;

[0078] At this time, referring to Figure 3 as shown, it is the mean distribution of the feature map and the channel-level weight distributions corresponding to the two algorithms. Compared with the prior art, it is found that another traditional visual interpretation method, Grad-CAM, contains a large number of negative-direction activations. The non-negative weight activation values obtained by the method of this application have a higher correlation with the feature map, reflecting the ability of the gradient to correspondingly affect the key regions of the decision result and being able to better reflect the importance of the feature map.

[0079] S23. Therefore, at this time, each pixel gradient at the (i, j) position of feature map A is weighted to obtain k groups of high-fine-grained weighted feature maps of different decision-related regions There is a calculation formula:

[0080]

[0081] where k ∈ {1, 2,..., K}.

[0082] In addition, in a specific embodiment, it should be further noted that the generation process of obtaining the mask by ordinal grouping processing includes:

[0083] S24. Divide the K groups of weighted feature maps equidistantly into G groups in the fixed order of the generated weighted feature maps, and G < K. Stack the weighted feature activation maps of each group to obtain G masks M for generating class activation maps l , and the calculation formula is:

[0084]

[0085] where the hyperparameter G is determined by experiments to obtain the optimal value.

[0086] At this time, the grouping method adopts ordinal grouping by merging in the fixed order of the feature map. The algorithm can obtain better performance and consistency. In the model of this application, K = 512, G = 64. Through grouping and merging, the subsequent calculation amount can be reduced by 8 times.

[0087] S25. To obtain a smoother mask, normalize the obtained low-resolution mask and upsample it to the original image size through bilinear interpolation to obtain a smoother mask. The calculation formula is:

[0088]

[0089] M’l It is the processed smooth mask.

[0090] S3. Input the blurred mask occluded image into the CNN, obtain the global decision contribution through the forward pass including N - time integral sampling, calculate the obtained contribution importance score and take the mean as the group importance score.

[0091] Among them, the specific calculation process for obtaining the global decision contribution includes:

[0092] S31. Perform Gaussian blur processing on the occluded area, and after processing, obtain the blurred mask occluded image I’. l , and the calculation expression is:

[0093]

[0094] Among them, is the input image after Gaussian blur, and ⊙ is the Hadamard product;

[0095] S32. The attribution between the model prediction and the input features should conform to two basic axioms: sensitivity (if there are feature deviations in a certain dimension for different inputs and the prediction outputs are different, then the attribution of this dimension is not zero) and implementation invariance (functionally equivalent models have the same input attribution). Therefore, linearly interpolate as the integration path to gradually accumulate the input image features, calculate the forward pass classification scores on the target class N times and take the average as the group mask importance score, obtaining a more accurate global decision contribution corresponding to the input;

[0096] At the same time, various gradient - free methods take a pure - black picture or a noise picture as the baseline image. Therefore, at this time, the output f c (M”0) of the baseline image M’0 is also set to zero. Take a pure - black picture or a noise picture as the baseline image and execute the gradient - free method, and there is:

[0097]

[0098]

[0099] Among them, is the importance score of the l - th group of masks, and N is the number of integral sampling;

[0100] It can further avoid the interference of unreasonable baseline images and improve the accuracy of representing the global decision contribution of different input features;

[0101] At the same time, in order to improve the running speed, expand the accumulation operation and calculate the forward passes of N masks in parallel, effectively improving the algorithm speed.

[0102] S4. Linearly weight the grouped mask to obtain the class activation map, which weights the mask using the importance score to obtain the visual interpretation map. The calculation expression is:

[0103]

[0104] where is the grouped mask of the i-th group of baseline images.

[0105] In summary, by mining the high-fine-grained local sensitivity information in the feature map through pixel-level gradient weighting and generating a weighted feature map with strong correlation with the decision result, and adopting an ordinal grouping method to merge the weighted feature maps, a mask containing local sensitivity information, low redundancy, and strong decision relevance can be obtained, thereby improving the visual accuracy of the object detection result. At the same time, the importance score of the grouped mask is calculated in parallel based on the global decision contribution, and the mask is processed using the fuzzy boundary and integral methods, which improves the sensitivity of different inputs to the global decision contribution and the algorithm speed, making it convenient to use.

[0106] Furthermore, the following system programming code is used to illustrate this method:

[0107] LG-CAM algorithm

[0108] Input: Image I0, model f, class c, number of groups G, Gaussian blur parameter.

[0109] Output: Class activation map

[0110] 1 Initialize the class activation map Base input image

[0111] 2 Obtain the feature map A of the target layer k, calculate the pixel-level weighted gradient W by the calculation formula of the feature map weight c , and calculate the weighted feature map F by the calculation formula of the high-fine-grained weighted feature map c ; c ;

[0112] 3 Use the mask calculation formula for generating the class activation map and the bilinear interpolation upsampling calculation formula to merge the weighted feature maps into G groups and generate the upsampled grouped mask M' l ;

[0113] 4 while l = 1 and l ≤ G

[0114] 5 while i = 0 and i < N

[0115] 6 Calculate the forward propagation score f of this integration c (M' l );

[0116] 7 i = i + 1

[0117] 8 end

[0118] 9 Calculate the mask importance score using the calculation formula for the importance score contributed by the gradient-free method, and obtain the calculation formula from the visual interpretation diagram and superimpose it to

[0119] 10 l = l + 1

[0120] 11 end

[0121] 12 Return

[0122] Experimental Example

[0123] Based on the above embodiments, in a specific experimental demonstration process, the visual interpretability method of the small target detection model proposed in the above embodiments is also experimentally verified based on the ImageNet classification dataset.

[0124] (1) Establish a dataset and evaluation metrics

[0125] The ImageNet validation set contains a total of 50,000 images, including 1000 target types, and provides the ground truth bounding boxes corresponding to the target classes with the highest prediction scores. 2000 images are extracted from the ImageNet validation set, and visual interpretation diagrams are generated for the target classes with the highest confidence for each image. When using the ImageNet image dataset, the following preprocessing is usually adopted: resize the image size to (224×224×3), convert the pixel values to the range [0,1], and normalize using the mean vector [0.485, 0.456, 0.406] and the standard deviation vector [0.229, 0.224, 0.225]. To ensure fairness, the pre-trained model weights in the pytorch model library are uniformly used for verification.

[0126] The evaluation of the visual interpretability method is divided into qualitative evaluation and quantitative evaluation. Qualitative evaluation can qualitatively evaluate the visual interpretability method from the perspectives of visual continuity, category distinguishability, multi-target and small-target visualization of the saliency maps obtained by different methods. Quantitative evaluation analysis can be carried out from two aspects: confidence evaluation and localization evaluation.

[0127] Confidence evaluation adopts the method of image occlusion experiment to force the model to change its decision, and verify the importance of the salient regions in the class activation mapping diagram for the prediction confidence. The metrics include deletion (Del) and insertion (Ins) metrics. Considering the two metrics comprehensively, it can be calculated by the following formula:

[0128] I-D = AUC(Ins) - AUC(Del)

[0129] Among them, the smaller the deletion metric, the better, and the larger the insertion metric and the I-D index, the better.

[0130] The localization accuracy of the weakly supervised object localization experiment evaluation method is evaluated. The localization accuracy is measured by the loc1 and loc5 metrics. The intersection over union (IoU) threshold between the predicted bounding box and the ground truth bounding box is set at intervals of 0.1 and tested one by one in the range of (0, 1). The loc1 and loc5 object localization accuracies under different thresholds are obtained and denoted as mloc1 and mloc5. The higher the object localization accuracy, the better.

[0131] (2) Ablation experiments of the improved method are carried out based on the above dataset and evaluation metrics

[0132] The parameter selection in the improved method is tested by ablation experiments on the ImageNet dataset to determine the optimal model structure and achieve adaptability to different input images. Since the number of groups and the number of integral samplings are key parameters and have a great impact on the computational cost and accuracy, in this section, the number of groups that can contain the most efficient information is first determined through multiple groups of experiments to ensure the basic algorithm accuracy and algorithm efficiency; secondly, multiple groups of integral sampling numbers are selected to conduct accuracy tests on the basis of the model with a fixed number of groups to determine the reasonable number of integral samplings. Subsequently, the grouping method, Gaussian blur parameter, and negative gradient operation that are independent of the results are tested respectively. Finally, the above improved method is verified to prove the effectiveness and speed advantage of the improvement relative to the baseline model.

[0133] (a) Number of groups test

[0134] The number of groups G in the LG-CAM algorithm is tested with different values. In the experiment, the number of integral samplings is set to 4, the grouping method is set to ordinal grouping and the negative gradient operation is used. The test results are shown in the following table. As the number of groups increases, the computational cost increases, and the comprehensive evaluation index I-D of the confidence level continuously rises. The localization accuracy index first rises and then decreases, and the maximum mloc1 is obtained when the number of groups is 32 and 64. It shows that there is a certain redundancy of information in the CNN feature map, and too large a number of grouping groups only brings a small improvement in accuracy at a greater computational cost. Considering that the I-D index when G = 64 is significantly improved compared to when G = 32, and the time consumption of the two is similar, the number of groups is selected as 64.

[0135]

[0136]

[0137] (b) Number of integral samplings test

[0138] The number of integral samplings represents the number of intervals in which the input image features increase from 0 to all image features. The influence of different numbers of integral samplings on the performance of the LG-CAM algorithm was verified. In the experiment, the number of groups was set to 64, the grouping method was ordinal grouping and the negative gradient operation was used, and the test results are shown in the following table. As the number of integral samplings increases, the localization accuracy and confidence indicators first increase and then decrease. The mloc1 localization accuracy reaches the highest at N = 4 and N = 10 respectively, the I-D indicator reaches the highest at N = 2, and it is sub-optimal at N = 4. It shows that by gradually accumulating the input image features in the way of integral sampling, this method can combine the contributions of the input features in different dimensions to the interpretation results, supplement the global decision-making information, and thus improve the algorithm effect.

[0139] At the same time, as the number of samplings increases, the time consumption increases significantly. Therefore, the number of integral samplings is selected as 4.

[0140]

[0141] (c) Grouping method test

[0142] The influence of different grouping methods on the algorithm performance was tested. In the experiment, the number of groups was set to 64, and the number of integral samplings was set to 4. The test results are shown in the following table. Among them, based on the Kmeans clustering, hierarchical clustering, and spectral clustering methods, the feature map is unfolded into a one-dimensional feature vector for clustering grouping. Based on the Structural Similarity (SSIM) grouping method, the structural similarity matrix between the feature activation maps is calculated and then clustering grouping is performed. It should be noted that affected by the initial value of the clustering algorithm, the obtained class activation mapping diagrams are not consistent. The results of the clustering methods in the table were tested 5 times and the average value was taken. From the following table, the spectral clustering reaches the optimal in the mloc1 and mloc5 localization accuracy indicators, and the ordinal grouping method is sub-optimal. In the I-D indicator, the ordinal grouping method is the best and the time consumption is the lowest. The ordinal grouping method does not perform artificial clustering and will not produce inconsistent results or unsatisfactory clustering; this method performs equidistant grouping according to the fixed order of generating the weighted feature map by feature extraction, making full use of the information structure of the CNN features itself, and obtaining a set of feature masks with low information redundancy and strong decision-making relevance more simply and efficiently. Therefore, the ordinal grouping method is selected.

[0143]

[0144] (d) Gaussian blur parameter test

[0145] During the calculation of the global decision contribution of the grouped mask, Gaussian blur processing is used for the occluded area to determine a reasonable Gaussian kernel size, and there is:

[0146]

[0147]

[0148] Among them, is the input image after Gaussian blur, ⊙ is the Hadamard product, and I’ l is the image masked by the blurred mask.

[0149] The results are shown in the following table. As the Gaussian kernel size increases, the I-D index improves and reaches the highest when the Gaussian kernel size is 51. The influence of the Gaussian kernel size on the positioning accuracy and time consumption is not obvious. Therefore, considering comprehensively, the Gaussian kernel size of 51 (kernel standard deviation of 50) is selected.

[0150]

[0151] (e) Negative gradient elimination test

[0152] In the weighted formula of pixel gradient, the LG-CAM algorithm uses ReLU to filter the negative gradient to illustrate the influence of the negative gradient on the model result. The experimental results are shown in the following table. w / oReLU means that the LG-CAM algorithm does not use ReLU in the formula calculation. It can be seen that by eliminating the negative gradient, a 2% improvement can be achieved in both the positioning accuracy and confidence metrics.

[0153]

[0154] (f) Verification of the effectiveness of the improved method

[0155] To verify the influence of the improved method on the algorithm's interpretability, based on the GroupCAM algorithm, the proposed improved method is gradually introduced to verify its effectiveness. The test results are shown in the following table. By comparing the GroupCAM+pp algorithm with the GroupCAM algorithm, it can be seen that through the improved method of high-fine-grained local sensitivity weighting of the feature map, the algorithm has improvements of 0.76% and 1.01% in the mloc1 and mloc5 positioning accuracy metrics respectively. Through the improved method of calculating the global decision contribution of the grouped mask, the IS method has an improvement of 0.121% in the I-D confidence metric. At the same time, by processing the grouped mask in parallel in the code, the algorithm speed is improved. Generally speaking, the algorithm of this application has improvements of 1.01%, 1.28%, and 0.023% in the mloc1, mloc5, and I-D metrics respectively compared with the baseline algorithm, indicating the effectiveness of the improved method proposed in this application.

[0156] (3) Comparative experiments with other methods

[0157] The method LG-CAM of this application is qualitatively and quantitatively compared with 7 interpretable methods. The comparison algorithms include gradient-based Guided-BP, perturbation-based RISE, class activation map-based GradCAM, GradCAM++, ScoreCAM, ISCAM, and GroupCAM. The number of RISE masks generated is 4000, the number of groups of GroupCAM is 64, and the number of integral samplings of ISCAM is 10.

[0158] (a) Qualitative analysis

[0159] For four typical example figures, the corresponding visual interpretation figures are generated by different visual interpretation methods as Figure 4 .

[0160] Visually, the Guided-BP method reflects all image features such as textures and edges learned by the convolutional layer of the CNN model. However, this type of method modifies the gradient information during the calculation process and cannot truly reflect the model's decision-making. The perturbation-based RISE method can reflect the key parts in the image that affect the CNN decision-making, with good visual continuity, but the range is prone to divergence and it is difficult to focus on the target area. On the contrary, the GradCAM method is difficult to activate all relevant areas of the target, with poor visual continuity. In comparison, GradCAM++, ScoreCAM, and GroupCAM have certain improvements in the integrity of the target localization area and visual continuity. However, for targets with complex backgrounds such as Figure 4 Example (4), there may still be a large amount of background noise, making it difficult to reflect the CNN decision-making characteristics. Compared with the comparison algorithms, the LG-CAM method combines the idea of non-gradient global contribution, and the results have less random noise, can generate a smoother class activation map, is more concentrated, accurate, and complete in localizing the target area, has better visual continuity, and reflects the main basis for the model to make classification decisions.

[0161] From the visualization effects of multi-targets and small targets, Figure 4 In Examples (2) and (3), RISE and GradCAM tend to only capture a single target in the image and are difficult to activate the small target area. GradCAM++ can localize multiple targets. ScoreCAM and GroupCAM methods have more complete and concentrated target activation and can localize and detect small targets, but there are large artifacts and noise interferences when the background is complex. The proposed LG-CAM in this application can discover and localize multiple similar targets in complex backgrounds with little background interference. At the same time, it is more effective in activating small targets, and the localization area is more accurate, concentrated, and complete, with better visual effects.

[0162] From the perspective of category discrimination ability, Figure 5It shows that the LG-CAM method can well distinguish different categories of targets. For the two categories of targets, butterflies and flowers, in the figure, the confidence score of the VGG19 model for classifying the input image as a butterfly is 94.8%, and the confidence score for classifying it as a flower is 1%. Although the latter confidence score is much lower than the former, through the backpropagation of different category confidences to obtain different activation feature map gradient weights, the method of this application can still give class activation maps corresponding to different target regions respectively.

[0163] (b) Quantitative analysis

[0164] L1. Image occlusion experiment

[0165] Figure 6 Correspond to Figure 4 The insertion and deletion metric curves and AUC values obtained from the image occlusion experiment of the images in (the left is the insertion metric curve, and the right is the deletion metric curve). It can be seen that LG-CAM has a faster rising insertion metric curve and a larger AUC, and a faster falling deletion metric curve and a smaller AUC. Figure 4 The I-D indicators of Examples (1), (2), and (4) in reach the highest 90.11%, 36.28%, and 63.41% respectively. Figure 4 The I-D indicator of Example (3) in reaches the second highest 79.17%, indicating that LG-CAM has better interpretability and is in good agreement with the visual qualitative evaluation results.

[0166] The following table gives the average values of the image occlusion experiment indicators of LG-CAM and the comparison methods on the validation set. It can be seen that LG-CAM has the best performance in the Del indicator, is only 0.07 lower than ISCAM in the Ins indicator, and is only 0.07 lower than ISCAM in the I-D comprehensive indicator, exceeding GradCAM++ by 4.17%. The image occlusion experiment shows that the visual interpretation map provided by the LG-CAM method can correspond to better classification confidence and has stronger interpretability.

[0167]

[0168] L2. Weakly supervised object localization accuracy experiment

[0169] The weakly supervised object localization accuracy experiment test is carried out on the LG-CAM method and 6 visual interpretation methods. Figure 7 For the visual effects of object localization of different methods, where the red box is the ground truth box of the object, and the green box is the bounding box obtained by the algorithm. From the visual intuitive perspective, the object localization box obtained by the algorithm of this application is more fitted to the ground truth box of the object and the localization is more accurate.

[0170] To evaluate the positioning accuracy more comprehensively and objectively, the positioning accuracy at different thresholds was tested on the validation set, and the target positioning accuracy curves of loc1 and loc5 were obtained. Refer to Figure 8 .

[0171] The average value of the positioning accuracy at each threshold was taken to obtain the mloc1 and mloc5 positioning accuracies, and the running speed of the algorithm was counted. In the weakly supervised target localization experiment, compared with the comparative algorithm, the algorithm of this application can achieve higher positioning accuracy at each threshold, and the mloc1 and mloc5 accuracies reach the highest, exceeding the second-place GradCAM++ by 0.12% and 0.1% respectively. This shows that the algorithm has better positioning ability. At the same time, the algorithm of this application takes into account both the algorithm speed and performance. The time consumption for a single-frame image is 70ms, which is 76 times that of ISCAM. This shows that the algorithm of this application has better performance in comprehensive performance.

[0172] In terms of the calculation results, the method of this application combines the idea of non-gradient global contribution. The operation results of the model have less random noise, can generate a smoother class activation mapping graph, are more concentrated, accurate and complete in localizing the target area, have better visual continuity, and are convenient for reflecting the model classification decision.

[0173] Therefore, this method can also discover and localize multiple similar targets in a complex background with little background interference; at the same time, it is more effective in activating small targets, and the localization area is more accurate, concentrated and complete, with better visual effects.

[0174] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor causes the processor to execute the steps of the above method.

[0175] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, it causes the processor to execute the steps of the above method.

[0176] In yet another embodiment provided by this application, a computer program product containing instructions is also provided, which when running on a computer causes the computer to execute the visual interpretability method of any small target detection model in the above embodiments.

[0177] It can be understood that the system provided by the embodiments of the present invention corresponds to the method provided by the embodiments of the present invention. The explanations, examples and beneficial effects of related content can refer to the corresponding parts in the above method.

[0178] This application embodiment also provides an electronic device including a processor, a communication interface, a memory and a communication bus. Among them, the processor, the communication interface and the memory complete communication with each other through the communication bus.

[0179] A memory for storing computer programs;

[0180] A processor, when executing the programs stored in the memory, implements the above-mentioned method for visual interpretability of the small target detection model.

[0181] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc.

[0182] The communication interface is used for communication between the above electronic device and other devices.

[0183] The memory can include a Random Access Memory (RAM), and can also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory can also be at least one storage device located far from the aforementioned processor.

[0184] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0185] It should also be noted that the electronic device further includes a terminal device, which can also be referred to as a terminal, a user equipment, a mobile station, a mobile terminal, etc. The terminal device can be a mobile phone, a smart TV, a wearable device, a tablet computer, a computer with wireless transceiver function, a virtual reality terminal device, an augmented reality terminal device, a wireless terminal in industrial control, a wireless terminal in unmanned driving, a wireless terminal in remote surgery, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, etc. The embodiments of the present application do not limit the specific technologies and specific device forms adopted by the terminal device.

[0186] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive SSD), etc.

[0187] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

[0188] In addition, it should be noted that if there are directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture. If the specific posture changes, the directional indications will also change accordingly.

[0189] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the meaning of "and / or" appearing throughout the text includes three parallel scenarios. Taking "A and / or B" as an example, it includes scenario A, scenario B, or the scenario where A and B are satisfied simultaneously. In addition, in the embodiments of the present invention, "a plurality" means two or more. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

Claims

1. A method for visual interpretability of small target detection models, characterized in that, It includes the following steps: S1. Obtain a deep feature map through the forward pass of the CNN, and obtain the corresponding gradient of the feature map through the backpropagation of the classification scores of specific class targets; S2. After the pixel-level gradient weighting of the feature image and the processing based on ordinal grouping, a set of masks containing high-fine-grained local sensitivity information is obtained; S3. Input the blurred mask-occluded image into the CNN, obtain the global decision contribution through the forward pass including N times of integral sampling, calculate and obtain the contribution importance score and take the average as the group importance score; S4. Linearly weight the grouped masks to obtain the class activation map, and the class activation map weights the masks using the importance score to obtain the visual interpretation map.

2. The method for visual interpretability of the small target detection model according to claim 1, wherein, The specific process of pixel-level gradient weighting of the feature image in step S2 includes: S21. Let \(f\) be an image classification CNN model, \(\theta\) be the model parameters. For a given input image \(I_0\), the feature map weights are obtained through global average pooling. After forward inference of the network, the classification score formula \(y\) c is: Among them is represented as the k-th channel feature map, where i and j respectively represent the row index and column index in the feature map; S22. For the classification score formula y c For each feature map weight within perform explicit coding. The calculation formula for the feature map weight is as follows: Among them is the weighted coefficient of the pixel gradient corresponding to class c and the feature map, and the calculation formula is: Among them, is the pixel gradient value of the K-th channel at the (i, j) position of the feature map A; S23. Weight each pixel gradient of the feature map A at the position (i, j) to obtain K groups of weighted feature maps 3. The method for visual interpretability of the small target detection model according to claim 2, wherein The generation process of obtaining the mask through the ordinal grouping process in step S2 includes: Group K weighted feature maps Equidistantly divide them into G groups according to the fixed order of the generated weighted feature maps, where G < K, and stack the weighted feature activation maps of each group to obtain G masks M for generating class activation maps l , and the calculation formula is: Normalize the obtained low-resolution mask and upsample it to the original image size through bilinear interpolation to obtain a smoother mask. The calculation formula is: M' l is the processed smoothed mask.

4. The method for visual interpretability of a small target detection model according to claim 3, wherein, The specific calculation process of obtaining the global decision contribution in step S3 includes: S31. Apply Gaussian blur processing to the occluded area, and after processing, obtain the blurred masked occluded image I'. l The calculation expression is as follows: Among them, is the input image after Gaussian blur, and ⊙ is the Hadamard product; S32. Gradually accumulate the input image features with linear interpolation as the integration path, calculate the forward pass classification scores on the target class N times and average them as the group mask importance score to obtain the global decision contribution.

5. The method for visual interpretability of the small target detection model according to claim 4, characterized in that, The specific operation process of calculating the contribution importance score in step S3 includes: Set the output f of the baseline image M'0 c (M”0) to zero, use a pure black image or a noise image as the baseline image, and perform a gradient-free method. We have: Among them, is the importance score of the l-th group of masks, and N is the number of integral samplings; Perform the forward pass of N masks in parallel.

6. The method for visual interpretability of the small target detection model according to claim 5, characterized in that In step S4, the G-group masks are weighted using the importance score to obtain a more refined and accurate visual interpretation map. The calculation expression is: wherein is the grouping mask for the i-th group.

7. A computer-readable storage medium, characterized in that, There is a computer program stored. When the computer program is executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.

8. A computer device, characterized in that, It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Convolutional neural network feature visualization method and system based on pixel gradient weighting

    CN112906867A

  • Sheet surface defect image recognition processing method based on convolutional neural network

    CN113505865A

  • Bird classification method based on comparison hierarchical correlation propagation theory

    CN114724184A

  • Image classification interpretable method based on comprehensive class activation mapping

    CN114913378A

  • Memory controler operating method and memory device

    KR102814915B1