Visual explanation of classification

By training generative models and generative adversarial networks to generate interpretive masks, the problem of difficult interpretation of AI classification decisions in medical imaging is solved, achieving low-noise and high-transparency interpretation effects, and improving the understanding and reliability of medical image classification.

CN116888639BActive Publication Date: 2026-01-23SIEMENS MEDICAL SOLUTIONS USA INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180091086.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-18
Publication Date
2026-01-23
Estimated Expiration
2041-01-18

AI Technical Summary

Technical Problem

Existing interpretable artificial intelligence systems struggle to effectively explain their classification decisions in medical imaging applications, especially in medical images with low sample sizes and similar samples, where noise and noise patterns make interpretation difficult to understand.

Method used

By training a generative model to generate new images that are similar to the input image but are classified as candidate categories by the classifier, generative adversarial networks (cGANs) and optimization algorithms are used to generate explanatory masks, reducing noise and improving the clarity of the explanation.

Benefits of technology

The generated explanation mask has low noise levels, providing a simple and clear explanation that helps to understand the classifier's bias and decision-making process, thus improving the transparency of medical image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116888639B_ABST
    Figure CN116888639B_ABST
Patent Text Reader

Abstract

A framework for visual explanations for classification. The framework trains (204) a generative model to generate new images that resemble an input image but are classified by a classifier as belonging to one or more alternative classes. At least one explanation mask can then be generated (206) by performing an optimization based on the current input image and the new images generated from the current input image by the trained generative model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to digital medical data processing, and more specifically to visualized explanations of classifications. BACKGROUND

[0002] In recent years, artificial intelligence (AI) systems have made great strides in accuracy on a variety of tasks and domains. However, these systems are inherently black boxes, trading off transparency for transactional accuracy: these algorithms cannot explain their decisions. The lack of transparency is problematic, particularly in the medical domain, where humans must be able to understand how decisions are made in order to trust AI systems. Greater transparency would enable human operators to know when to trust AI decisions, and when to override it.

[0003] Explainable AI (referred to in the literature as XAI) is an emerging field, and a number of techniques have been published on this topic. The goal of XAI is to provide important factors that led to a classification. These methods can be grouped into the following categories: (1) conformal; (2) based on saliency; and (3) based on attention.

[0004] Symbolic reasoning systems were developed in the 70-90s with built-in interpretability. However, these systems do not handle well non-classification tasks, such as the interpretation of images. The saliency-based methods require the classifier to have its output differentiable with respect to its input. Many methods have been proposed in the literature, such as guided backpropagation, Grad-CAM, integrated gradients, etc. See, for example, Springenberg, Jost Tobias, et al. "Striving for Simplicity: The AllConvolutional Net." CoRR abs / 1412.6806 (2015); Selvaraju, R. R., Cogswell, M., Das, A., et al. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. Int J Comput Vis 128, 336-359 (2020); and Sundararajan, M., Taly, A., and Yan, Q., "Axiomatic Attribution for Deep Networks," 2017, which are incorporated herein by reference. Saliency-based methods mainly look at the influence of the input by computing the derivative of the input with respect to the neural network (NN) output. A well-trained neural network projects its input onto a low-dimensional manifold and then classifies it. However, due to the noise inherently present in the images, the NN can not project the input exactly onto the manifold. The derivative of the input with respect to the output will amplify the noise and lead to uninterpretable noise patterns in the saliency map. This effect is amplified in medical imaging applications, which usually have a low number of training samples and relatively similar samples (i.e., more likely to fall outside the manifold).

[0005] Attention-based methods use trainable attention mechanisms added to the neural network to help locate relevant locations in the image. See, for example, K. Li, Z. Wu, K. Peng, J. Ernst, and Y. Fu, "Tell Me Where to Look: Guided Attention Inference Network," 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, 2018, pp. 9215-9223, which is incorporated herein by reference. Attention-based methods do not "explain" the classification, but rather point to relevant areas that need further explanation. Therefore, they are less suitable for medical applications. SUMMARY

[0006] A framework for visualized explanations of classifications is described herein. According to one aspect, the framework trains a generative model to generate new images that are similar to an input image but are classified by a classifier as belonging to one or more alternative classes. Then, at least one explanation mask can be generated by performing an optimization based on the current input image and the new images generated from the current input image by the trained generative model. BRIEF DESCRIPTION OF DRAWINGS

[0007] A more complete understanding of the present disclosure and the many attendant aspects thereof will readily be acquired as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings.

[0008] Figure 1 An exemplary system is shown;

[0009] Figure 2 An exemplary method of generating an explanation mask is shown;

[0010] Figure 3 An exemplary cGAN architecture is shown;

[0011] Figure 4A An exemplary optimization architecture is shown;

[0012] Figure 4B An exemplary process for generating multiple explanation masks is shown;

[0013] Figure 5 An exemplary comparison of results is shown;

[0014] Figure 6 Another exemplary comparison of results is shown;

[0015] Figure 7 Results generated by the present framework are shown; and

[0016] Figure 8 Additional results generated by the present framework are shown. DETAILED DESCRIPTION

[0017] In the following description, numerous specific details are set forth such as examples of specific components, devices, methods, etc., in order to provide a thorough understanding of implementations of the present framework. It will be apparent, however, to one skilled in the art that these specific details need not be used to practice the present framework. In other instances, well-known materials or methods have not been described in detail in order to avoid unnecessarily obscuring the present framework. While the present framework is susceptible to various modifications and alternative forms, specific embodiments thereof have been shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that there is no intent to limit the present framework to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present framework. Further, some method steps that are described in the course of an exemplary method can be performed in various orders unless otherwise specifically noted, or unless it can be clear from the context that the order of performance is specifically intended.

[0018] The term "x-ray image" as used herein can mean either a visible x-ray image (e.g., displayed on a video screen) or a digital representation of an x-ray image (e.g., a file corresponding to x-ray detector pixel output). The term "in-treatment x-ray image" as used herein can refer to an image captured at any point in time during the treatment delivery phase of an intervention or treatment procedure, which can include times when the radiation source is turned on or off. At times, for ease of description, CT imaging data (e.g., cone-beam CT imaging data) can be used herein as an example imaging modality. However, it will be understood that data from any type of imaging modality can also be used in various implementations, including but not limited to radiographs, MRI, PET (positron emission tomography), PET-CT, SPECT, SPECT-CT, MR-PET, 3D ultrasound images, etc.

[0019] Unless specifically stated otherwise, and as apparent from the following, it will be appreciated that throughout the present specification, discussions utilizing terms such as "segmenting", "generating", "registering", "determining", "aligning", "positioning", "processing", "computing", "selecting", "estimating", "detecting", "tracking", or the like, can refer to the action and processes of a computer system, or similar electronic computing device, that manipulates and transforms data represented as physical (e.g., electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system memories or registers or other such information storage, transmission or display devices. Embodiments of the methods described herein can be implemented using computer software. If written in a programming language conforming to a recognized standard, sequences of instructions can be compiled for execution on a variety of hardware platforms and for interface to a variety of operating systems. Furthermore, the implementations of the present framework are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the present framework.

[0020] As used herein, the term "image" refers to multidimensional data composed of discrete image elements (e.g., pixels of a 2D image and voxels of a 3D image). An image can be a medical image of a subject collected, for example, by computed tomography, magnetic resonance imaging, ultrasound, or any other medical imaging system known to those skilled in the art. Images can also be provided from non-medical settings, such as from remote sensing systems, electron microscopes, etc. While an image can be considered as originating from R... 3 Functions in R, or to R 3 This method is not limited to such images and can be applied to images of any dimension, such as 2D pictures or 3D volumes. For 2D or 3D images, the domain of the image is typically a 2D or 3D rectangular array, where each pixel or voxel can be addressed by referencing a set of 2 or 3 mutually orthogonal axes. The terms "digital" and "digitized" as used herein will, as appropriate, refer to images or volumes in digital or digitized formats acquired through a digital acquisition system or by conversion from analog images.

[0021] The term "pixel" used for image elements in 2D imaging and image display and the term "voxel" used for volumetric image elements in 3D imaging are interchangeable. It should be noted that a 3D volumetric image is itself synthesized from image data obtained as pixels on a 2D sensor array and displayed as a 2D image from a certain viewpoint. Therefore, 2D image processing and image analysis techniques can be applied to 3D volumetric image data. In the following description, techniques described as operating on pixels can alternatively be described as operating on 3D voxel data stored and represented in the form of 2D pixel data for display. Similarly, techniques operating on voxel data can also be described as operating on pixels. In the following description, the terms "new input image," "pseudo-image," "output image," and "new image" are interchangeable.

[0022] One aspect of this framework is to provide interpretations of anomalies detected by any classifier when the task is to make a normal or anomalous decision, by training a generative model. The generative model can be trained to produce new images that resemble the input image but are classified by the classifier as belonging to one or more candidate categories. The generative model constrains the interpretation to remove noise from the interpretation mask (or graph). Advantageously, the noise level in the generated interpretation mask is much lower than in existing methods, making the interpretation of the mask simple and straightforward. The trained generative model can produce new input images x′, which can be used to understand what the classifier considers to be the category with the highest probability. This is very useful for understanding classifier bias (e.g., would an expert reader looking at x′ make the same classification?) or from a designer's perspective, it can ensure that the classifier is adequately trained and mimics how experts perceive diseases. These and other features and advantages will be described in more detail in this paper.

[0023] Figure 1 This is a block diagram of an exemplary system 100. System 100 includes a computer system 101 for implementing the framework described herein. In some implementations, computer system 101 operates as a standalone device. In other implementations, computer system 101 may be connected (e.g., using a network) to other machines, such as imaging device 102 and workstation 103. In a networked deployment, computer system 101 may operate as a server (e.g., a thin client server), a cloud computing platform, a client user machine in a server-client user network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.

[0024] In some implementations, computer system 101 includes a processor or central processing unit (CPU) 104 coupled via an input-output interface 121 to one or more non-transitory computer-readable media 105 (e.g., computer storage devices or memories), a display device 110 (e.g., a monitor), and various input devices 111 (e.g., a mouse or keyboard). Computer system 101 may also include supporting circuitry such as caches, power supplies, clock circuits, and communication buses. Various other peripheral devices, such as additional data storage devices and printing devices, may also be connected to computer system 101.

[0025] This technology can be implemented in various forms of hardware, software, firmware, dedicated processors, or combinations thereof, as part of microinstruction code, or as part of an application or software product, or a combination thereof, which is executed by an operating system. In some implementations, the technology described herein is implemented as computer-readable program code tangibly embodied in a non-transitory computer-readable medium 105. Specifically, this technology can be implemented by an interpretation module 106 and a database 109. The interpretation module 106 may include a training unit 102 and an optimizer 103.

[0026] The non-transitory computer-readable medium 105 may include random access memory (RAM), read-only memory (ROM), floppy disk, flash memory, and other types of memory, or combinations thereof. The computer-readable program code is executed by the CPU 104 to process medical data retrieved from, for example, imaging device 102. Thus, the computer system 101 is a general-purpose computer system, which becomes a special-purpose computer system when the computer-readable program code is executed. The computer-readable program code is not intended to be limited to any particular programming language or its implementation. It will be understood that the teachings of the disclosure contained herein can be implemented using various programming languages ​​and their encodings.

[0027] The same or different computer-readable media 105 may be used to store databases (or datasets) 109 (e.g., medical images). This data may also be stored on external storage devices or other memories. External storage devices may be implemented using a database management system (DBMS) managed by the CPU 104 and residing on memory, such as hard disks, RAM, or removable media. External storage devices may be implemented on one or more additional computer systems. For example, external storage devices may include data warehouse systems residing on separate computer systems, cloud platforms or systems, picture archiving and communication systems (PACS), or any other hospital, medical institution, clinic, testing facility, pharmacy, or other medical patient record storage system.

[0028] Imaging device 102 acquires medical image data 120 associated with at least one patient. This medical image data 120 may be processed and stored in database 109. Imaging device 102 may be a radiation scanner (e.g., X-ray, MR, or CT scanner) and / or suitable peripheral devices (e.g., keyboard and display device) for acquiring, collecting, and / or storing such medical image data 120.

[0029] Workstation 103 may include a computer and suitable peripherals, such as a keyboard and display device, and may operate in conjunction with the overall system 100. For example, workstation 103 may communicate directly or indirectly with imaging device 102, such that medical image data acquired by imaging device 102 can be rendered at workstation 103 and viewed on a display device. Workstation 103 may also provide other types of medical data 122 for a given patient. Workstation 103 may include a graphical user interface to receive user input via input devices for inputting medical data 122 (e.g., keyboard, mouse, touchscreen, voice or video recognition interface, etc.).

[0030] It should also be understood that, since some of the system components and method steps depicted in the accompanying drawings can be implemented in software, the actual connections between system components (or processing steps) may vary depending on how this framework is programmed. Given the teachings provided herein, those skilled in the art will be able to conceive of these and similar implementations or configurations of this framework.

[0031] Figure 2 An exemplary method 200 for generating an interpretation mask is shown. It should be understood that the steps of method 200 can be performed in the order shown or a different order. Additional, different, or fewer steps may also be provided. Furthermore, method 200 can be used... Figure 1 System 101, different systems or combinations thereof are used to achieve this.

[0032] At 202, training unit 102 receives training input images and a classifier f. The training input images may be medical images acquired directly or indirectly using medical imaging techniques such as high-resolution computed tomography (HRCT), magnetic resonance (MR) imaging, computed tomography (CT), spiral CT, X-ray, angiography, positron emission tomography (PET), fluoroscopy, ultrasound, single-photon emission computed tomography (SPECT), or combinations thereof. The training input images may include normal and abnormal images used to assess one or more diseases. For example, training input images may include normal and abnormal dopamine transporter scan (DaTscan) SPECT images used to assess Parkinson's disease. As another example, training input images may include amyloid-positive and amyloid-negative PET images used to assess amyloidosis. Abnormal images contain at least one abnormality (e.g., abnormal accumulation of α-synuclein in brain cells, amyloid deposits, lesions), while normal images do not contain any abnormalities.

[0033] In some implementations, the classifier f is a binary classifier, trained to classify training input images as normal or abnormal images. It should be understood that in other implementations, the classifier f can also be a non-binary classifier. The classifier f takes input x and returns an output O representing the classification probabilities among N classes, where c is the class with the highest probability. The classifier f can be implemented using machine learning techniques, including but not limited to neural networks, decision trees, random forests, support vector machines, co-evolutionary neural networks, or combinations thereof.

[0034] In step 204, training unit 102 trains a generative model using the training input image to generate new, high-quality pseudo-images x′. The generative model is trained to generate new input images x′ that are similar to (or as close as possible to) the training input image x but are classified by the classifier as belonging to one or more candidate categories (i.e., one or more distinct categories from the corresponding training input image). A generative model is a type of statistical model that can generate new data instances. A generative model includes the distribution of the data itself and indicates the likelihood of a given example. In some implementations, the generative model is a conditional generative model, where the input image x is modulated to generate a corresponding output image x′. A generative model can be, for example, a deep generative model formed by combining a generative model and a deep neural network. Examples of deep generative models include, but are not limited to, variational autoencoders (VAEs), generative adversarial networks (GANs), and autoregressive models.

[0035] In one implementation, the generative model includes a Generative Adversarial Network (GAN). A GAN is a machine learning framework that includes two neural networks competing against each other in a minimax game—a generator G and a discriminator D. The generative model can also be a conditional GAN. A conditional GAN ​​(cGAN) is a conditional generative model that learns from data, where an input image x is modulated to generate a corresponding output image x′, which is used as input to the discriminator for training. See, for example, Isola, Phillip et al., “Image-to-Image Translation with Conditional Adversarial Networks,” IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2017 (5967-5976), which is incorporated herein by reference. The goal of the generator G is to generate pseudo-inputs x′ that are indistinguishable from the real input x. The goal of the discriminator D is to identify the real input from the pseudo-inputs generated by G. The generator G and discriminator D are optimized sequentially, and the min-max game ideally converges to the generator G, producing a high-quality new input image x′ that simulates the distribution of x, while the discriminator D cannot guess whether the new input image x′ is real or fake. Unlike unconditional GANs, both G and D in cGANs observe the input image x.

[0036] Figure 3An exemplary cGAN architecture 300 is shown. The cGAN includes a generator 302 and a discriminator 304. The generator 302 can be trained to take an input image x (306) and generate a new input image x′ (308) that has the constraint of being as close as possible to the input image x (306) but classified into an alternative category by the classifier f. For example, if the input image x (306) is classified as “abnormal” by the classifier, the generator 302 can be trained to generate a new image x′ (308) that is a “normal” image and is very similar to the input image x (306). As another example, if the input image x (306) is classified as “normal” by the classifier, the generator 302 can be trained to generate a new image x′ (308) that is a “abnormal” image that is very similar to the input image x (306). The discriminator 304 can be trained to identify the true input image (310) from the pseudo-input image x′ (308) generated by the generator 302.

[0037] The cGAN objective can be formulated as follows:

[0038] G * =arg min G max D Loss cGAN (G, D)------------(1)

[0039] Loss cGAN (G, D) = E x,c [log D(x, c)]+E x,c [log(1-D(G(x,c),c)]-----(2)

[0040] Where x is the observed input image, c is the class with the highest probability, G(x, c) is the new image x′ of G, D(x, c) is the output of D, and Loss cGAN (G, D) are the loss functions of G and D. G attempts to minimize the loss function Loss against its opponent D. cGAN (G, D), the opponent D tries to maximize it. Loss function. cGAN (G, D) is the expected value E x,c [log D(x, c)] and E x,c The sum of [log(1-D(G(x,c),c)], where x and c are sampled from possible images and categories, respectively.

[0041] In other implementations, the cGAN objective is formulated as follows:

[0042] G * =arg min G max D LosscGAN (G,D)+αL1(G)---------------------(3)

[0043] Loss cGAN (G, D) = E x,c [log D(x, c)]+E x,c [log(1-D(G(x,c),c)]-----(4) and L1(G)=||xG(x,c)||1-------(5),

[0044] Where x is the observed input image, c is the class with the highest probability, G(x, c) is the new image x′ of G, D(x, c) is the output of D, and Loss cGAN (G, D) are the loss functions of G and D, α is the parameter, and L1(G) is the distance between the observed input image x and the new image x′ generated by G (i.e., x′ = G(x, c)). In this case, the work of D remains unchanged, but G is trained to not only fool D, but also to approximate the ground truth output in an L1 sense. In other words, G is penalized if it generates a new image x′ that is dissimilar (or not similar to) the input image x.

[0045] In other implementations, the cGAN objective is formulated as follows:

[0046] G * =arg min G max D Loss cGAN (G,D)+αL1(G)+βL(f(G))-----------(6)

[0047] Loss cGAN (G, D) = E x,c [log D(x, c)]+E x,c [log(1-D(G(x,c),c)]-----(7), and L1(G)=||xG(x,c)||1-------(8),

[0048] Where x is the observed input image, c is the class with the highest probability, G(x, c) is the new image x′ of G, D(x, c) is the output of D, and Loss cGAN(G, D) is the loss function of G and D, α and β are parameters, f is the classification function (or classifier), L1(G) is the distance between the observed input image x and the new image x′ generated by G (i.e., x′ = G(x, c)), and L is the loss term that penalizes the generator G if it generates an image that the classifier identifies as belonging to the incorrect category. Exemplary values ​​for parameters α and β can be, for example, 0.0002 and (0.5, 0.999), respectively. In the objective function (6), the disease classifier is linked to the generator 302 via the loss term L, such that the generator 302 is penalized if it generates a new image that the classifier considers to belong to the incorrect category. For example, if G attempts to generate a "normal" image from an "abnormal" input image, and the disease classifier classifies the generated new image as "abnormal" (instead of normal), then L is given a non-zero value to penalize G. This loss term ensures that the generator G and the classifier are connected.

[0049] Training cGANs can be achieved using various techniques, such as conditional autoencoders, conditional variational autoencoders, and / or other GAN variants. Furthermore, trained cGANs can be optimized using, for example, the ADAM optimizer—Adaptive Gradient Descent (ADAM). See, for example, Kingma, DP and Ba, J. (2014), Adam: A Method for Stochastic Optimization, which is incorporated herein by reference. Other types of optimization algorithms can also be used.

[0050] The trained generator G can generate new input images x′, which can be used to understand what the classifier considers to be category c. This is very useful for understanding the classifier's bias (e.g., an expert reader looking at x′ would make the same classification) or, from a designer's perspective, to ensure that the classifier is adequately trained and mimics expert perception of the disease.

[0051] return Figure 2 At 206, optimizer 103 generates an interpretation mask by performing optimization based on the current input image x and a new image x′ generated from the current input image by a trained generative model. The current input image x can be acquired from the patient by, for example, imaging device 102 using the same modality (e.g., a SPECT or PET scanner) used to acquire training images. The interpretation mask can then be generated by performing optimization based on the current input image x and the new image x′ to reduce the classifier's classification probability to a predetermined value. The interpretation mask represents the voxels in the current input image x that need to be changed to alter the classifier's decision; therefore, these voxels are possible interpretations of the classifier's classification. Each value in the interpretation mask represents a mixing factor between x and x′.

[0052] Figure 4AAn exemplary optimization architecture 400 is shown. Based on the observed input image x (404) and class c, a new image x′ (408) is generated by G of the trained cGAN. A mask (401) represents the portion of the input image x (404) to be mixed with the new input image x′ (408). A classification function f (402) takes the combination of the input image x (404) mixed by the mask (401) and the new pseudo-image x′ (408) as input and returns an output O representing the classification probability in N classes, where c is the class with the highest probability. The gradient of the mask (401) with respect to the output O of the classifier is constrained to be a mixture of x and x′, and x′ is constructed by design to be similar to x, thereby limiting unrealistic noise sources and enhancing robustness to noise.

[0053] The optimization aims to find a smaller mask that reduces the classifier probability of class c to 1 / N, where N is the total number of classes determined by the classifier. The optimization problem can be formulated as follows:

[0054]

[0055]

[0056] Where x″=x⊙(1-Mask)+x′⊙Mask--------(10)

[0057] α (e.g., 0.05) represents a scaling factor that controls the sparsity of the interpretation mask (411), and ⊙ represents the element-wise multiplication of the two terms. The combined input x” represents the sum of the current input image x and the new image x′ mixed by the previous mask. Optimization is performed on the combined input x” to minimize the probability of class c of the classifier (402). Thus, by construction, the combined input x” is in the same domain as the input x (404), and can be interpreted as follows.

[0058] Optimization can be achieved using backpropagation by calculating the partial derivative of the classifier output O(c) with respect to the mask (401). See, for example, Le Cun Y. (1986), “Learning Process in an Asymmetric Threshold Network”. Disordered Systems and Biological Organization, NATO ASI Series (Series F: Computer and Systems Sciences), vol 20. Springer, Berlin, Heidelber, which is incorporated herein by reference. Optimization stops once the classifier f(x”)(c) reaches a predetermined probability (e.g., 1 / N). Reaching a probability of 0 may not be desirable, as this could introduce noise into the mask, and 1 / N is found to be a good trade-off between noise and interpretation. In some implementations, backpropagation is applied up to 200 times with a learning rate of 0.1.

[0059] This framework is scalable to support multiple modes, both normal and abnormal. This can be accomplished by sampling multiple new input images x′ from the generator G and aggregating the interpretive masks for each x′ generated by architecture 400. Multiple interpretive masks can be generated from a single current input image x. Figure 4B An exemplary process 410 for generating multiple interpretation masks is shown. Multiple different interpretation masks 412 can be generated by passing the same single input image x(414) multiple times (e.g., 100 times) through a trained generative model to generate multiple new input images x′, and then passing the multiple new input images x′ through architecture 400 to generate multiple interpretation masks 412. Optionally, multiple different interpretation masks 412 can be aggregated to improve the robustness of the interpretation. Aggregation can be performed, for example, by averaging or clustering the interpretation masks.

[0060] As shown in the exemplary procedure 410, clustering can be performed to generate multiple clusters 416a-b. It should be understood that although only two clusters (cluster 1 and cluster 2) are shown, any other number of clusters can be generated. Different clusters can represent, for example, different lesions or other abnormalities. Clustering algorithms can include, for example, density-based spatial clustering (DBSCAN) with noise or other suitable techniques. A representative interpretation mask (e.g., cluster centers) can be selected for each cluster and presented to the user. The size of the clusters can be used to rank the interpretation masks for the user or characterize the importance of the interpretations.

[0061] return Figure 2At 208, the interpretation module 106 presents an interpretation mask. The interpretation mask can be displayed at a graphical user interface, for example, on workstation 103. The interpretation mask provides a visual explanation of the classification (e.g., anomaly classification) generated by the classifier. The noise level in the interpretation mask is advantageously much lower than the noise level produced by existing methods, making the interpretation of the interpretation mask simpler and clearer. Additionally, a new input image x′ generated by the trained generative model can also be displayed at the graphical user interface. The new input image x′ can be used to understand what the classifier considers to be the category with the highest probability.

[0062] This framework is implemented in the context of Parkinson's disease. A classifier is trained on DaTscan images to classify them as normal or abnormal images. The classifier is trained using 1356 images, tested on 148 images, and achieves 97% accuracy on the test data.

[0063] Figure 5 An exemplary comparison of results obtained using different conventional algorithms and this framework is shown to explain the classifier's classification. Conventional algorithms include Grad-CAM, backpropagation, guided backpropagation, and integral gradient algorithms. Column 502 shows randomly selected anomalous input DaTscan images. Columns 504, 506, 508, and 510 show explanatory maps generated based on the input images in column 502 using standard algorithms. Column 512 shows the explanatory map generated by this framework. The explanatory maps generated by conventional methods exhibit extreme noise, making them difficult to interpret. In contrast, the explanatory maps generated by this framework exhibit very low noise, making interpretation simple and straightforward.

[0064] Figure 6 Another exemplary comparison of results obtained using different conventional algorithms and this framework is shown to explain the classifier's classification. Columns 604, 606, 608, and 610 show the interpretive maps generated using conventional algorithms based on the input image in column 602. Column 612 shows the interpretive map generated by this framework. Compared to this framework, conventional methods generate interpretive maps exhibiting more extreme noise, making them difficult to interpret.

[0065] Figure 7The results generated by this framework are shown. Column 702 shows the input DaTscan image x from the test data. Column 706 shows the explanatory mask generated by this framework. Column 704 overlays the input image x with the explanatory mask to enable better visualization of spatial patterns relative to the input image. The computed mask patterns are highly correlated with asymmetric or bilateral reduction in shell-nucleus uptake and are a reasonable explanation for these scans being classified as anomalous. Column 712 shows the new input image x′ trained on the cGAN, which transforms the anomalous input image 702 into a normal image that closely matches the input image 702. Finally, column 710 shows the combined input x”, which represents the sum of the current input image x and the new image x′ mixed by the mask. These images x” closely match the input image 702 and reduce the classifier's probability to below 50%.

[0066] Figure 8 Additional results generated by this framework are shown. Column 802 shows the input DaTscan image x from the test data. Column 806 shows the explanatory mask generated by this framework. Column 804 overlays the input image x with the explanatory mask to enable better visualization of spatial patterns relative to the input image. The computed mask patterns are highly correlated with asymmetric or bilateral reduction in shell-nucleus uptake and are a plausible explanation for these scans being classified as anomalous. Column 812 shows the new input image x′ trained on the cGAN, which transforms the anomalous input image 802 into a normal image that closely matches the input image 802. Finally, column 810 shows the combined input x”, which represents the sum of the current input image x and the new image x′ mixed by the mask. These images x” closely match the input image 802 and reduce the classifier's probability to below 50%.

[0067] While this framework has been described in detail with reference to exemplary embodiments, those skilled in the art will understand that various modifications and substitutions can be made therein without departing from the spirit and scope of the invention as set forth in the appended claims. For example, within the scope of this disclosure and the appended claims, elements and / or features of different exemplary embodiments may be combined with and / or substituted for each other.

Claims

1. One or more non-transitory computer-readable media embodying an instruction program, said instruction program being machine-executable to perform operations for interpreting mask generation, said operations including: Receive input image and classifier; A new image similar to the input image is generated by a trained generative model but is classified by the classifier as belonging to a different category than the input image. The generative model is a conditional generative model, in which the input image is adjusted to generate a corresponding output image. At least one explanatory mask is generated by performing optimization based on the input image and the new image, wherein the optimization reduces the classification probability of the classifier to a predetermined value, wherein the optimization is performed by finding a smaller mask that reduces the classification probability of class c of the classifier to the predetermined value, class c being the class with the highest classification probability, and wherein the optimization is performed on the sum of the mixture of the current input image and the new image generated from the current input image by the trained generative model; as well as The explanation mask is presented by displaying the explanation mask at the graphical user interface.

2. The one or more non-transitory computer-readable media of claim 1, wherein the generative model comprises a conditional generative adversarial network (cGAN).

3. The one or more non-transitory computer-readable media according to claim 1, wherein the operation further comprises: The generative model is trained by penalizing it with a loss term in response to the generative model generating a new image that the classifier considers to belong to an incorrect category.

4. A system for visual explanation of classification, comprising: Non-transitory memory devices used for storing computer-readable program code; as well as A processor communicating with the memory device, the processor operating together with the computer-readable program code to perform operations including: Receive the training input image and the classifier; A generative model is trained based on the training input image to generate a new image that is similar to the training input image but is classified by the classifier as belonging to one or more candidate categories. The generative model is a conditional generative model, in which the input image is adjusted to generate a corresponding output image. At least one explanatory mask is generated by performing optimization based on the current input image and a new image generated from the current input image by a trained generative model. This optimization is performed by finding a smaller mask that reduces the classification probability of class c (the class with the highest classification probability) of the classifier to a predetermined value. The optimization is performed on the sum of the mixture of the current input image and the new image generated from the current input image by the trained generative model. The explanation mask is presented by displaying the explanation mask at the graphical user interface.

5. The system of claim 4, wherein the classifier comprises a binary classifier trained to classify the training input image as a normal or abnormal image.

6. The system of claim 4, wherein the processor operates together with the computer-readable program code to train a generative model by training a generative adversarial network (GAN).

7. The system of claim 4, wherein the processor operates together with the computer-readable program code to train a generative model by training a conditional generative adversarial network (cGAN).

8. The system of claim 4, wherein the processor operates together with the computer-readable program code to train the generative model by penalizing the generative model with a loss term in response to the generative model generating a new image that the classifier considers to belong to an incorrect category.

9. The system of claim 4, wherein the processor operates together with the computer-readable program code to train the generative model by penalizing the generative model with a loss term in response to the generative model generating a new image dissimilar to the training input image.

10. The system of claim 4, wherein the processor operates together with the computer-readable program code to train the generative model in such a way as to generate a first new image that is "normal" and similar to the first input image in response to receiving a first input image classified as "abnormal" by the classifier.

11. The system of claim 4, wherein the processor operates together with the computer-readable program code to train the generative model in such a way as to generate a second new image that is "abnormal" and similar to the second input image in response to receiving a second input image classified as "normal" by the classifier.

12. The system of claim 4, wherein each value in the interpretation mask represents a mixing factor between the current input image and a new image generated from the current input image by a trained generative model.

13. The system of claim 4, wherein the processor operates together with the computer-readable program code to generate the at least one interpretation mask by generating a plurality of different interpretation masks from the current input image.

14. The system of claim 13, wherein the processor operates together with the computer-readable program code to aggregate the plurality of different interpretation masks.

15. The system of claim 14, wherein the processor operates together with the computer-readable program code to cluster the plurality of different interpretation masks by performing clustering techniques.

16. The system of claim 4, wherein the predetermined value includes 1 / N, where N is the total number of categories determined by the classifier.

17. A method for visual explanation of classification, comprising: Receive the training input image and the classifier; A generative model is trained based on the training input image to generate a new image that is similar to the training input image but is classified by the classifier as belonging to one or more candidate categories. The generative model is a conditional generative model, in which the input image is adjusted to generate a corresponding output image. At least one explanatory mask is generated by performing optimization based on the current input image and a new image generated from the current input image by a trained generative model, wherein the optimization is performed by finding a smaller mask that reduces the classification probability of class c of the classifier to a predetermined value, class c being the class with the highest classification probability, and wherein the optimization is performed on the sum of the mixture of the current input image and the new image generated from the current input image by the trained generative model; as well as The explanation mask is presented by displaying the explanation mask at the graphical user interface.

Citation Information

Patent Citations

  • Determining a perturbation mask for a classification model

    EP3739515A1