Saliency maps for deep learning models

By generating support and disrupting significance maps, explaining the decision-making process of deep learning models in image quality evaluation, solving the shortcomings in the generation of significance maps in the prior art, and improving the interpretability of the model and image quality improvement capabilities.

CN120283268APending Publication Date: 2025-07-08KONINKLIJKE PHILIPS NV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202380080538.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-11-13
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has difficulty generating significant maps meaningful in nonbinary classification and regression tasks, especially in medical use cases, which cannot effectively explain the decision-making process of deep learning models.

Method used

By generating support significance maps and perturbation significance maps, indicating areas in the image that support and perturbation metric scores, respectively, provides an explanation of the deep learning model, supporting areas indicate high-quality parts, disturbing areas indicate low-quality parts, and image quality can be improved through post-processing.

Benefits of technology

Improve interpretability of deep learning models, help users understand and improve image quality, and enhance trust in model decisions, especially in medical image analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283268A_ABST
    Figure CN120283268A_ABST
Patent Text Reader

Abstract

A method and system for providing a saliency map for a deep learning model. The method includes inputting an input image into the deep learning model, the deep learning model being trained to output a metric score from a plurality of metric scores for the input image; generating, from the deep learning model, a support saliency map for the input image corresponding to a first range of the metric scores for the image, thereby providing one or more support regions of the image indicative of the first range of metric scores; and further generating, from the deep learning model, a perturbation saliency map for the image corresponding to a second range of the metric scores for the image, thereby providing one or more perturbation regions of the image indicative of the second range of metric scores.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning models. In particular, the present invention relates to the field of saliency maps obtained from deep learning models. Background Art

[0002] Large deep learning models are generally considered black boxes. This can be dangerous, especially when used in a medical environment. In particular, when there is a limited amount of data to train a deep learning model, or when there are unobvious biases in the data, there is a significant risk that the trained model overfits some details or incidental attributes of the data.

[0003] Saliency methods such as guided backpropagation or class activation maps (CAMs) have been proposed to visualize the results of deep learning models beyond the pure classification output. Saliency maps achieve this by highlighting the regions within the input data that contribute the most to the final network output of the network.

[0004] For example, a saliency map can be obtained by looking at the output of the last convolutional layer in a deep learning model corresponding to the selected classification. The convolutional layer can include multiple maps, which can be combined (e.g., weighted combination) to generate a saliency map. Saliency maps are typically used when overlaid on the input image, as it shows which regions of the image the deep learning model is focusing on when determining the classification of the image.

[0005] When asking experts to annotate images (especially for soft features, such as the quality or suitability of a specific diagnostic task), saliency maps can show visual cues for the decisions of deep learning models. For example, when using simple categories such as "good / bad" (e.g., binary classification), saliency maps perform well.

[0006] For many categories or problems on a continuous scale, there is no ad-hoc method for generating saliency maps. Therefore, there is a need for an improved method for providing saliency maps. Methods for generating meaningful saliency maps in non-binary classification and regression are not intuitive and have been described in the context of medical use cases.

[0007] Zhang Jing et al.'s "Explainability for Regression CNN in Fetal Head Circumference Estimation from Ultrasound Images" (ECCV 2020) discloses the use of saliency maps to measure head circumference ultrasound images.

[0008] US2021 / 0174497A1 discloses a method for using feature maps to generate improved saliency maps. Summary of the Invention

[0009] The present invention is defined by the claims.

[0010] According to an example of an aspect of the present invention, a method for providing a saliency map for a deep learning model is provided. The method includes:

[0011] Inputting an input image into the deep learning model, the deep learning model being trained to output a metric score for the input image from a plurality of metric scores;

[0012] Generating, from the deep learning model, a support saliency map for the input image corresponding to a first range of the metric scores for the image, thereby providing one or more support regions of the image indicating the first range of the metric scores; and

[0013] Generating, from the deep learning model, a perturbed saliency map for the image corresponding to a different second range of the metric scores for the image, thereby providing one or more perturbed regions of the image indicating the second range of the metric scores.

[0014] The first range and the second range are different because they do not cover exactly the same range of possible metric scores output by the deep learning model. In an embodiment, the first range and the second range do not overlap.

[0015] When calculating the output metric score, the support region and the perturbed region provide information about the relative importance of regions to the deep learning algorithm. In particular, the support region (indicating the first metric range) indicates the image regions that cause the determination of a specific metric score for the image, while the perturbed region (indicating the second metric range) indicates the regions that cause the deep learning model to deviate from the output of the specific metric score.

[0016] When providing the user with the context of the nature of the metric score, this can help the user determine the suitability of the image. For example, when the metric score is a quality score indicating the image quality, the support region may indicate the high-quality parts in the image, while the perturbed region may indicate the low-quality parts.

[0017] Substantially, the support region and the perturbed region act as an explanation of how the deep learning model determines the metric score, where the explanation is not based on explicit rules but on the displayed regions. Such an explanation can increase trust in the deep learning model because it not only outputs the metric score but also provides an explanation of how the metric score is obtained.

[0018] Metric scores typically involve non-measurable and / or subjective gradings that cannot be simply measured and / or may not simply depend on who / what is providing the metric score.

[0019] The deep learning model can be a regression model.

[0020] Alternatively, the deep learning model can be a classification model.

[0021] The support saliency map can be generated by applying weights to the metric scores relative to the distance of the metric scores in the first range from the support values included in the first range, and the perturbation saliency map can be generated by applying weights to the metric scores relative to the distance of the metric scores in the second range from the perturbation values included in the second range.

[0022] When attempting to visualize the decisions of a neural network, single saliency maps have been used in the past. However, especially in this case of image quality grading, it is not clear how to generate these maps because image quality assessment is typically formulated as a regression task (i.e., predicting a continuous output or grading scale). Regression tasks do not have a conventional saliency, and it is not clear how to extend saliency map generation for regression tasks in image quality grading. Therefore, the present invention also provides an extension of saliency maps for regression by internally redefining the problem as a number of classification problems.

[0023] The metric score can be a quality score.

[0024] The method can further include: processing the image separately at the positions of the support region and / or the perturbation region.

[0025] For example, the support region can correspond to a high-quality region, while the perturbation region can correspond to a low-quality image. Thus, the low-quality region can be processed to improve the quality (e.g., adapt contrast, sharpness, etc.), such that the overall quality of the image is improved without affecting the (already high) quality of the support region.

[0026] The method can further include: switching from a first image processing scheme to a second image processing scheme to process the image in response to the metric score output by the deep learning model being lower than or exceeding a score threshold.

[0027] The method can further include: retraining or tuning the deep learning model using user-specific images and metric scores provided by the user for the user-specific images.

[0028] This enables people to visualize and understand the user's decisions via the saliency maps when annotating images.

[0029] The method can further include: displaying the support saliency map and the perturbation saliency map superimposed on the image.

[0030] The present invention also provides a computer program carrier including computer program code which, when run on a processing system, causes the processing system to execute all steps of the foregoing method.

[0031] For example, the computer program carrier can be a relatively long-term data storage solution (e.g., a hard disk drive, a solid state drive, etc.) or a relatively temporary data storage solution (e.g., a bitstream).

[0032] The present invention also provides a system for providing a saliency map for a deep learning model, the system including a processor configured to:

[0033] input an input image into the deep learning model, the deep learning model being trained to output a metric score for the input image from a plurality of metric scores;

[0034] generate a support saliency map for the input image corresponding to a first range of the metric scores for the image from the deep learning model, thereby providing one or more support regions of the image indicating the first range of the metric scores; and

[0035] generate a scrambled saliency map for the image corresponding to a different second range of the metric scores for the image from the deep learning model, thereby providing one or more scrambled regions of the image indicating the second range of the metric scores.

[0036] The deep learning model can be a regression model.

[0037] The processor can be configured to: generate the support saliency map by applying weights to the metric scores based on the distance of the metric scores in the first range from a support value included in the first range, and the processor can be configured to: generate the scrambled saliency map by applying weights to the metric scores based on the distance of the metric scores in the second range from a scrambled value included in the second range.

[0038] The metric score can be a quality score.

[0039] The processor can further be configured to: separately process the image at the locations of the support regions and / or the scrambled regions.

[0040] The processor can further be configured to: retrain or tune the deep learning model using user-specific images and metric scores provided by the user for the user-specific images.

[0041] These aspects and other aspects of the present invention will be apparent with reference to the (one or more) embodiments described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To better understand the present invention and to more clearly show how to implement the present invention, reference will now be made, by way of example only, to the accompanying drawings, in which:

[0043] Figure 1 Two echocardiogram images are shown with support regions and disruption regions superimposed thereon for image quality;

[0044] Figure 2 A flowchart for obtaining a support saliency map and a disruption saliency map is shown; and

[0045] Figure 3 An abdominal computed tomography (CT) scan is shown. DETAILED DESCRIPTION

[0046] The present invention will be described with reference to the accompanying drawings.

[0047] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, system and method, are intended for illustrative purposes only and are not intended to limit the scope of the present invention. These and other features, aspects and advantages of the apparatus, system and method of the present invention will be better understood from the following description, the appended claims and the drawings. It should be understood that the drawings are merely schematic and are not drawn to scale. It should also be understood that throughout these drawings, the same reference numerals are used to indicate the same or similar parts.

[0048] The present invention provides a method and system for providing a saliency map for a deep learning model. The method includes: inputting an input image into the deep learning model, the deep learning model being trained to output a metric score from a plurality of metric scores for the input image; generating, from the deep learning model, a support saliency map for the input image corresponding to a first range of the metric scores for the image, thereby providing one or more support regions of the image indicating the first range of the metric scores; and also generating, from the deep learning model, a disruption saliency map for the image corresponding to a second range of the metric scores for the image, thereby providing one or more disruption regions of the image indicating the second range of the metric scores.

[0049] Currently, interpretable artificial intelligence (AI) is a growing field, especially in medical applications. For example, further automation for ultrasound examinations is being sought. In this context, it is desirable to detect whether an ultrasound image is suitable for a higher level of automatic evaluation, such as extended measurement, volume measurement, strain measurement, or other measurements. Currently, ultrasound images approved by the user are fed into the automatic evaluation.

[0050] The advantage of using a support saliency map and a perturbed saliency map is that they can help solve common problems of neural networks; they are perceived as black boxes. In particular, when modeling human annotations that rely on knowledge and experience specific to a task, neural networks are often perceived as black boxes. Both saliency maps serve as an interpretive metric for the scores from a neural network (or other AI / deep learning solution), which are not extracted according to clear rules but show the blurred patterns and appearances of features within the image.

[0051] Figure 1 Two cardiac ultrasound images are shown that are superimposed with a support region 106 and a perturbed region 108 for image quality. Figure 1 a) shows a relatively low-quality ultrasound cardiac image 102. The ultrasound image 102 is input into a deep learning model that is trained to output an image quality score for the image 102. Additionally, two saliency maps for the image 102 are obtained from the deep learning model.

[0052] The first saliency map is a support saliency map that shows the support region 106 of the image 102 that the deep learning model uses to support a high-quality score. In other words, the support region shows the image area that the deep learning model considers to be of high quality. It can be seen that the only support region in the image 102 is a small part of the wall.

[0053] The second saliency map is a perturbed map that shows the perturbed region 108 of the image 102 that causes the deep learning model to move away from a high-quality score. In other words, the perturbed region shows the image area that the deep learning model considers to be of low quality. In this case, it seems that the deep learning model uses the wall near the vertex of the image 102 and strong speckle noise towards the left side of the image 102 (shown in the perturbed region 108) to determine the relatively low-quality rating / score for the image.

[0054] In contrast, Figure 1 b) shows a relatively high-quality ultrasound image 104. The image 104 has well-defined boundaries with good contrast and good sidewall visibility, as shown by the support region 106 and a single relatively small perturbed region 108. These regions may drive a relatively high image quality score for the image 104.

[0055] Therefore, in the case of image quality, compared to a single quality score obtained from a deep learning score, the support region 106 that supports the saliency map and the perturbation region 108 that perturbs the saliency map provide more detailed information about the quality of images 102 and 104.

[0056] Compared to existing "global" labels (e.g., good image quality or bad image quality), the solutions discussed herein provide visual regions (i.e., regions 106 and 108) that can be overlaid on images 102 and 104. Thus, after learning to extract recognizable external metrics, the deep learning model can provide a support region and a perturbation region to visualize which parts of the image drive the model's decision, thereby giving insights into the regions of the input image that support the output label and the regions that oppose the output label. For example, if the basis of the label is unknown, the visibility of the myocardium in this case can be inferred from the regions 106 mainly referred to by the "good" label and the regions 108 mainly referred to by the "bad" label.

[0057] In addition, the deep learning model can receive user feedback on why and where a certain metric is met or not met.

[0058] Two desirable use cases have been identified. The first use case involves understanding expert decisions and being able to visualize expert decisions on data for internal improvement of methods and parameters. The second use case is to provide an explanation of certain metrics of their data to (clinical) users by visualizing the reasoning and thus driving trust in the deep learning model for a specific task.

[0059] Figure 2 A flowchart for obtaining a support saliency map 208 and a perturbation saliency map 212 is shown.

[0060] Initially, a classification problem or a regression problem is formulated for the deep learning model 204. To this end, the input image 202 must be relevant to the target. The target can be a discrete label (classification) or a label representing a continuous grading (regression). In the latter case, the continuous label is converted into a classification problem by defining the label ranges for the support class and the perturbation class. In the case of a regression target or a multi-class classification problem, the intermediate labels can be weighted according to their distance from the most supported (or perturbed) class.

[0061] The relationship between the label and the image is learned by an appropriate deep learning model 204 (e.g., a convolutional neural network (CNN)). In the case of a regression target, this can also be done in a multi-target learning manner, where one prediction branch of the network is used for regression at the original scale, and the second branch is used for classifying the support class and the perturbation class. The parameters of this model 204 should be selected in a way that adapts to the complexity of the problem.

[0062] Then, the trained model 204 can calculate a saliency map 208 for the support class and a saliency map 212 for the disruption class (e.g., Class Activation Map (CAM)) for each input 202. In a post-processing step, the support regions in the support map 208 and the disruption regions in the disruption map 212 can be used to identify features in the input image 202 that cause the prediction class for the input image 202 to be consistent (or inconsistent), as Figure 1 shown.

[0063] The deep learning model 204 can have the ability to process external labels. For example, the external label can be a quality score or different non-measurable technical quantities (e.g., noise level, distortion degree, etc.) that are usually rated / annotated by human observers. The external label can also be one of the blur level, presence / absence of features, sharpness of the image, etc.

[0064] The model 204 is trained to reproduce the external label by outputting a corresponding metric score 214 via a regression task or a classification task.

[0065] The main task (i.e., outputting the metric score) can be divided into two extreme binary classification categories, which are weighted by the distance to the label (e.g., very good quality, very poor quality).

[0066] Generate a first saliency map 208. The first saliency map 208 derives regions within the input image 202 that support the first range 206 of the metric score and produces a first set (support) of output regions.

[0067] Also generate a second saliency map 212. The second saliency map 212 derives regions within the input image 202 that support the second range 210 of the metric score and produces a second set (disruption) of output regions.

[0068] The saliency map can include output regions of the input image 202 that support the external label in a supporting, disrupting, or neutral manner. This can give rise to four saliency indicators: positive saliency for the positive class (or class range) 206, negative saliency for the positive class (or class range) 206, positive saliency for the negative class (or class range) 210, and negative saliency for the negative class (or class range) 210. The support region can be the region in the support saliency map corresponding to the positive saliency indicator for the positive class 206, and the disruption region can be the region in the disruption saliency map corresponding to the positive saliency indicator for the negative class 210.

[0069] Thus, it is possible to visualize / render the support regions and disruption regions on the input image 202 or its schematic depiction. Optionally or alternatively, a post-processing adaptation step can be used to drive image enhancement of the detected support regions and / or disruption regions.

[0070] Note that the regions shown are not typically displayed as binary choices. Instead, for each pixel, a probability or intensity map for the image may be shown for each image. For example, the brightness / color of each pixel in a region can depend on the probability / corresponding value in the intensity map.

[0071] For example, when the metric score 214 is image quality, contrast enhancement can be applied to low-quality regions of the image, and / or an algorithm can be used to zoom into good-quality regions and / or crop poor-quality regions.

[0072] The output quality score can also be used to determine a threshold for switching between different post-processing schemes. For example, if the image quality drops below a certain threshold, a post-processing scheme that ensures a minimum quality in all image regions can be used.

[0073] As discussed, it is also possible to use a metric score other than image quality to derive a visual interpretation using the support regions and disruption regions. For example, the metric score can be an indication of whether the input image 202 can be used for measurement. Thus, the support regions can indicate regions of the input image 202 that indicate that the measurement will be successful, and the disruption regions can indicate regions that suggest that the measurement will not be successful.

[0074] Thus, it is possible to use a saliency map to detect whether an external measurement will be successful and identify regions in the input image 202 that may cause errors. This can be a live feature during image capture.

[0075] Users can also classify their own data in a simple scheme (e.g., like / dislike), and it is possible to retrain the deep learning model 204 based on the user-specific classification. In a first step, this can be used to visualize features on the data that support grading of newly acquired images.

[0076] In addition, even if the support / disruption regions are not shown to the user, they can be used to apply image quality improvements based on the user's preferences. When the deep learning model 204 is learned to identify regions that the user does not like in an image, those regions can be specifically processed to better match the user's like / dislike. Since the classification follows a simple scheme (e.g., like / dislike), it is easy to generate such a classification for the user.

[0077] In highly mobile applications (e.g., having only a tablet or smartphone as an interface with few control buttons), classification can be done in a simple scheme (such as thumb up / thumb down or left / right swipe) and users are allowed to customize the image enhancement algorithm based on their preferences.

[0078] It is also possible to provide a metric score 214 to the user. This can be achieved by a second quality score regression / classification head parallel to the saliency classification head in the deep learning model 204.

[0079] It should be understood that this method can also be applied to segmentation or landmark detection problems. Similarly, this method can be applied to modalities other than ultrasound imaging and can be applied to many scenarios.

[0080] For example, when analyzing expert annotations of medical images or visually understanding the meaning / effect of a certain metric in an image, this method can prove to be helpful.

[0081] For the user, this method can provide a visual representation of the deep learning model on the system. An example would be to guide successful measurements by indicating areas of insufficient image quality during inspection.

[0082] In addition, this method can also be integrated into other workflows, for example, by automatic processing steps based on an agreement with a certain label or by automatically selecting the best-matching input image for other measurements to achieve this.

[0083] In a standard setting, the deep learning model is usually pre-trained and then only applied to the acquired images. In an extended setting, the deep learning model on the system can also be re-trained or fine-tuned. In particular, users can rate their own images according to the perceived image quality (or some other non-measurable metric score). After re-training, the model will then identify those image regions that have a positive impact on the user's perception and those image regions that have a negative impact on the user's perception. Then, the image regions with negative impacts can be post-processed with increased attention and / or effort to better match the image quality with the expectations of individual users.

[0084] It should be understood that for many non-measurable / subjective metric scores (e.g., quality), how to post-process the images will generally depend on the user's preferences. Therefore, by re-training the deep learning model using the user's preferences, effective, personalized, and automated post-processing can be achieved for different users.

[0085] Figure 3 Abdominal computed tomography (CT) scans 302 and 304 are shown. Figure 3a) shows a first CT image 302 that is input into a trained deep learning model to determine quality (or some other subjective metric score). In this example, the support region 306 has been identified as "high quality" and the disrupted region 308 has been identified as "low quality". Thus, for Figure 3 the next CT image 304 shown in b), the CT scan parameters can be adapted, or the image can be post-processed such that the region in the next CT image 304 corresponding to the disrupted region 308 of the previous CT image 302 has improved quality. Of course, the previous CT image 302 can also be post-processed to improve the quality of the disrupted region 308 while maintaining the (already high) quality of the support region 306.

[0086] By studying the drawings, the disclosure, and the appended claims, those skilled in the art will be able to understand and implement variations of the disclosed embodiments when practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the words "a" or "an" do not exclude a plurality.

[0087] Functions implemented by a processor can be implemented by a single processor or by multiple individual processing units that can be considered to together constitute a "processor". In some cases, such processing units can be remote from each other and communicate with each other in a wired or wireless manner.

[0088] The fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used advantageously.

[0089] A computer program can be stored / distributed on a suitable medium (e.g., an optical storage medium or a solid-state medium supplied together with or as part of other hardware), but can also be distributed in other forms (e.g., distributed via the Internet or other wired or wireless telecommunication systems).

[0090] If the term "adapted to" is used in the claims or the specification, it should be noted that the term "adapted to" is intended to be equivalent to the term "configured to". If the term "arranged" is used in the claims or the specification, it should be noted that the term "arranged" is intended to be equivalent to the term "system", and vice versa.

[0091] Any reference signs in the claims shall not be construed as limiting the scope.

Claims

1. A method for providing a saliency map for a deep learning model (204), the method comprising: Inputting an input image (202) into the deep learning model, the deep learning model being trained to output a metric score (214) for the input image from a plurality of metric scores; Generating, from the deep learning model, a support saliency map (208) for the input image corresponding to a first range of the metric scores for the image, thereby providing one or more support regions (106) of the image indicative of the first range of the metric scores; And Generating, from the deep learning model, a perturbed saliency map (212) for the image corresponding to a different second range of the metric scores for the image, thereby providing one or more perturbed regions (108) of the image indicative of the second range of the metric scores.

2. The method according to claim 1, wherein, The deep learning model is a regression model.

3. The method according to claim 2, wherein: The support saliency map is generated by applying weights to the metric scores relative to the distance of the metric scores in the first range from a support value included in the first range, and The perturbed saliency map is generated by applying weights to the metric scores relative to the distance of the metric scores in the second range from a perturbation value included in the second range.

4. The method according to any one of claims 1 to 3, wherein The metric score is a quality score.

5. The method according to any one of claims 1 to 4, wherein Further comprising: Processing the image separately at the location of the support region and / or the perturbed region.

6. The method according to any one of claims 1 to 5, further comprising: Switching from a first image processing scheme to a second image processing scheme to process the image in response to the metric score output by the deep learning model being lower than or exceeding a score threshold.

7. The method according to any one of claims 1 to 6 further comprises: Retraining the deep learning model using user-specific images and metric scores provided by the user for the user-specific images.

8. The method according to any one of claims 1 to 7, further comprising: Displaying the support saliency map and the perturbed saliency map superimposed on the image.

9. A computer program carrier comprising computer program code which, when run on a processing system, causes the processing system to perform all the steps of the method according to any one of the preceding claims.

10. A system for providing a saliency map for a deep learning model (204), the system comprising a processor configured to: Input an input image (202) into the deep learning model, the deep learning model being trained to output a metric score (214) for the input image from a plurality of metric scores; Generate, from the deep learning model, a support saliency map (208) for the input image corresponding to a first range of the metric scores for the image, thereby providing one or more support regions (106) of the image indicative of the first range of the metric scores; And Generate a perturbed saliency map (212) for the image from the deep learning model corresponding to a different second range of the metric score for the image, thereby providing one or more perturbed regions (108) of the image indicative of the second range of the metric score.

11. The system according to claim 10, wherein, The deep learning model is a regression model.

12. The system according to claim 11, wherein: The processor is configured to: generate the support saliency map by applying weights to the metric score based on the distance of the metric score in the first range from a support value included in the first range, and The processor is configured to: generate the perturbed saliency map by applying weights to the metric score based on the distance of the metric score in the second range from a perturbed value included in the second range.

13. The system according to any one of claims 10 to 12, wherein, The metric score is a quality score.

14. The system according to any one of claims 10 to 13, wherein, The processor is further configured to: process the image separately at the positions of the support region and / or the perturbed region.

15. The system according to any one of claims 10 to 14, wherein, The processor is further configured to: retrain the deep learning model using user-specific images and metric scores provided by the user for the user-specific images.

Citation Information

Patent Citations

  • Saliency mapping by feature reduction and perturbation modeling in medical imaging

    US20210174497A1