Explainable ai for medical diagnosis

The method aligns visual characteristics with human-interpretable criteria to explain medical diagnoses, addressing the lack of transparency in machine learning models and improving trust and accuracy in clinical applications.

WO2025222209A1PCT designated stage Publication Date: 2025-10-23RUTGERS THE STATE UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/025650
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-19
Filing Date
2025-04-21
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Machine learning models in medical diagnosis operate as black boxes, lacking transparency and trustworthiness due to the inability to explain their decision-making processes, leading to potential errors and resistance in clinical integration.

Method used

A method and system that utilize a structured explainability model to generate input data representations based on human-interpretable criteria, evaluating similarity and outputting classification information with associated explanation information derived from concept scores and criteria.

Benefits of technology

Provides transparent and interpretable medical diagnoses by aligning visual characteristics with domain knowledge, enhancing trust and accuracy in clinical settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025025650_23102025_PF_FP_ABST
    Figure US2025025650_23102025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems are provided for performing classification using a structured, explainable classification model. The methods and systems may utilize knowledge anchor embeddings (or concept tokens) based on representations of human interpretable criteria derived from textual domain knowledge. Methods and systems for developing such models are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

EXPLAINABLE Al FOR MEDICAL DIAGNOSISCROSS-REFERENCE TO RELATED APPLICATION(S)

[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 636,207 filed April 19, 2024, the content of which is hereby incorporated by reference in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH

[0002] This invention was made with government support under HL127661 awarded by the National Institutes of Health. The government has certain rights in the invention.TECHNICAL FIELD

[0003] Various embodiments and implementations described herein relate generally to systems and methods for classifying images. Some embodiments and implementations provide for classification of a medical image or images by diagnostic criteria in an explainable way.BACKGROUND

[0004] Advancements in machine learning have allowed for improved and automated classification of images, sensor output, and other datasets. These techniques have allowed for techniques that can be leveraged in settings such as biomedical data processing, such as to lower diagnostic costs associated with image analysis, and improve diagnostic accuracy.

[0005] However, despite many advancements in functionality of machine learning, machine learning options have largely continued to operate as black boxes. While they can achieve high performance in things like image classification, they continue to be measured and assessed simply by their accuracy when presented with validation data. Yet simply understanding their accuracy and empirically determining what kinds of inputs generate more or less confidence in classification, fails to provide a true explainable behind the decision-making process. This lack of transparency compromises the trust and validation by users like healthcare professionals, leading to potential errors and resistance of integrating Al-derived insights into clinical settings.

[0006] While certain aspects of conventional technologies have been discussed herein and in theattached appendices to facilitate disclosure of the invention, such discussions in no way disclaim these technical aspects, and it is contemplated that embodiments of the present disclosure may encompass one or more of the conventional technical aspects discussed herein.

[0007] The present invention may address one or more of the problems and deficiencies of the prior art discussed herein. However, it is contemplated that the invention may prove useful in addressing other problems and deficiencies in a number of technical areas. Therefore, any given claimed embodiment should not necessarily be construed as limited to addressing any of the particular problems or deficiencies discussed herein.SUMMARY

[0008] The following presents a simplified summary of one or more aspects of the present disclosure, to provide a basic understanding of such aspects. This summary is not an extensive overview of all contemplated features of the disclosure and is intended neither to identify key or critical elements of all aspects of the disclosure nor to delineate the scope of any or all aspects of the disclosure. Its purpose includes presenting some concepts of one or more aspects of the disclosure in a simplified form as a prelude to the more detailed description that is presented later.

[0009] In one aspect, the present disclosure provides various implementations, embodiments, and variants of methods for classifying input data. For example, a method may comprise: receiving input data for a given classification task; generating, using a structured explainability model, a set of input data representations of features of the input data, each input data representation corresponding to a respective concept token associated with the classification task, wherein the concept tokens are based on representations of human-interpretable criteria for the classification task; evaluating similarity of the input data representations to the concept tokens, to obtain a set of concept scores for the input data; processing the set of concept scores via a classification output layer to obtain classification information for the input data; and outputting the classification information to a user with associated explanation information, the explanation information derived from the set of concept scores and the human-interpretable criteria.

[0010] In another aspect, the present disclosure various implementations, embodiments, and configurations of systems for performing knowledge-based classification of data. For example, a system may comprise: one or more processors; a vision-language model trained to perform a task of interest; a memory in communication with the one or more processors, wherein the memorycontains instructions that, when executed, cause the one or more processors to: receive data for a given classification; determine modality information of the data; provide the data and the modality information to the vision-language model, wherein the task of interest corresponds to the given classification; evaluate visual characteristics of the data across a set of data concept tokens corresponding to criteria categorizations for the task of interest; assess alignment scores for pairings of data concept tokens of the set of data concept tokens and associated visual characteristics; determine classification information for the data based at least in part on the assessment of the alignment scores, the classification information responsive to the task of interest; and output the classification information to a user with associated explainable support for the classification information derived from the alignment scores and data concept token characteristics.

[0011] These and other aspects of the disclosure will become more fully understood upon a review of the drawings and the detailed description, which follows. Other aspects, features, and embodiments of the present disclosure will become apparent to those skilled in the art, upon reviewing the following description of specific, example embodiments of the present disclosure in conjunction with the accompanying figures. While features of the present disclosure may be discussed relative to certain embodiments and figures below, all embodiments of the present disclosure can include one or more of the advantageous features discussed herein. In other words, while one or more embodiments may be discussed as having certain advantageous features, one or more of such features may also be used in accordance with the various embodiments of the disclosure discussed herein. Similarly, while example embodiments may be discussed below as devices, systems, or methods embodiments it should be understood that such example embodiments can be implemented in various devices, systems, and methods.BRIEF DESCRIPTION OF THE DRAWINGS

[0012] FIG. 1 is a process flow diagram of an example method for performing an explainable, diagnostic knowledge-based classification of single or multi-modal data.

[0013] FIG. 2 is a process flow diagram of an example method for creating a classification model that utilizes textual diagnostic knowledge and vision-language analysis.

[0014] FIG. 3 is a block diagram of an example system implementation.

[0015] FIGS. 4A and 4B illustrate a comparison of deep learning model characteristics and humanknowledge-based classification criteria.

[0016] FIG. 5 is a conceptual block diagram depicting a functional data flow according to some embodiments.

[0017] FIGS. 6A and 6B illustrate a diagram and a medical image to conceptually illustrating relationships between diagnostic criteria, textual descriptions, and a diagnostic analysis.

[0018] FIG. 7 illustrates a process diagram for generating a phenotypic risk prediction according to some embodiments.

[0019] FIG. 8 is a process flow diagram of an example method for generating aspects of a classification model according to various embodiments herein.DETAILED DESCRIPTION

[0020] The detailed description set forth below, in connection with the appended drawings and the attached appendices, is intended as a description of various configurations and is not intended to represent the only configurations in which the subject matter described herein may be practiced. The detailed description includes specific details to provide a thorough understanding of various embodiments of the present disclosure. However, it will be apparent to those skilled in the art that the various features, concepts and embodiments described herein may be implemented and practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form to avoid obscuring such concepts. Various embodiments, implementations, advantages, configurations, and examples of the present disclosure can also be found in the attached appendices with supporting data from experiments performed by the inventors.

[0021] As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the content clearly dictates otherwise. As used in this specification and the appended claims, the term “or” is generally employed in its sense including “and / or” unless the context clearly dictates otherwise.

[0022] As described below, various methods, systems, and model architectures are disclosed herein, which allow for true insight into how and why outputs of ML tasks like classification were determined. The techniques and principles underlying these methods, systems, and models are applicable to multiple types of input data, and can be applied to ML models that are used in a variety of domains and implementations.

[0023] In some examples, the techniques and advantages described below may be implemented in systems and methods for classification of medical images in an explainable way. For example, a ML model may be trained to analyze a type (or types) of image (e.g., from a given modality) to detect or classify objects or regions of interest (e.g., lesions, inflammation, tumors, occlusions, etc.), and then specifically report to a user what are the features of those objects that lead to the classification result - in which the features may align with expert / standard diagnostic criteria.

[0024] In other examples, these techniques and advantages may be leveraged in implementations that support unbiased assessments of multi-factor classifications. For example, where a complex system (biological, mechanical, environmental, etc.) or process is monitored via multiple different sensors (e.g., temperature, vibration, acoustic, pressure, electrochemical, motion, optical, etc.), a ML model may be trained to classify behavior of the system or process according to multiple sensor readings, and provide specific explanations for which features of which sensor readings supported the classification result.

[0025]

[0026] Embodiments of such methods, systems, software, and circuits are described below. First, general flowcharts and diagrams are presented to describe certain steps and / or components that may comprise or be common to multiple embodiments. Then, specific discussion is provided in regard to particular implementations generated by the inventors, along with validation data that helps illustrate the improvements and advantages offered by those implementations as well as other embodiments described herein.Methods and Techniques

[0027] FIG. 1 is a flow diagram illustrating an example process 100 for classification of single or multi-modal data in an explainable fashion. As described below, a given implementation of such a process might omit some or all illustrated features / steps, may utilizes features / steps in a different order, and / or may not require some illustrated features to achieve certain advantages or improvements. It should be appreciated that various suitable hardware, sensors, and system architectures for carrying out the operations or features described below may perform process 100.

[0028] At step 102, the process 100 receives a single or multi-modal data for a given classification. In some examples, the classification may define a type of identification, predication, confidence level, diagnostic task, subject, modality, or the like. Information regarding what the data representsmay further be collected. For example, the additional information may define what the data represents, how the data was acquired, as well as provide a description of the type of data. For example, the data may be a medical image, and the additional information may indicate how the data was acquired from a given anatomy of interest from a patient, as well as a request or indication of a desired diagnostic task. For example, the data may be a medical image, such as an X-ray image, CT image, MRI image, microscopy image, ultrasound image, biopsy section, lab test result (e.g., lateral flow test), or any other medical image or test that is visually inspected to determine diagnostic criteria. The anatomy of interest information may indicate the portion of the patient that was imaged, such as a brain scan, abdominal scan, brain tissue biopsy, etc. In some examples, the received request for a diagnostic task may include a request for an analysis of whether the medical image demonstrates presence of an identified disease or health state (or category of the same), such as melanoma, tumorous cancer, lesion, trauma, etc. Based on the information received and requested diagnostic task, process 100 will select an appropriate classification model or generate one (e.g., according to a method such as shown in FIG. 2).

[0029] At step 104, the process 100 evaluates features (e.g., visual characteristics of the image) across a set of data concept tokens. For example, a trained vision-language model (VLM) may be used to process the image to identify visual characteristics of the image according to a set of criteria axes or data concept tokens.

[0030] At step 106, the process 100 compares the data concept token characteristics to textual knowledge of diagnostically-relevant criteria for the given diagnostic task. In some examples, the analysis of domain knowledge associated with the data concept tokens performed at step 106 depends on the specific classification task defined at step 102. For example, a model may align feature data of a given image (which may include output of a vision-language model) to encoded data concept tokens that were learned for each of a number of criteria axes. In other embodiments, a pre-trained VLM may be utilized to generate textual description (or other output) of an image in specific categories that correspond to diagnostically relevant attributes of the same type of data, such as a same type of medical image of the same type of anatomy for the same type of diagnostic task (e.g., brain cancer detection from CT images), then the output can be compared to learned concepts / tokens associated with positive identification of a given diagnostic attribute.

[0031] At step 108, the process 100 may assess alignment scores for each data token / textual knowledge pair. In some examples, the alignment scores may be determined using one or morecosine projections, which can extract features from the domain knowledge and project them onto the received single or multi-modal data. For example, alignment scores may be determined for each of a set of criteria axes defined by knowledge-derived diagnostic criteria (e.g., asymmetry as a criteria for positive correlation with skin cancer). In some examples, the knowledge may be obtained directed from a user (e.g., a medical professional), or may be generated automatically by an associated system. The knowledge may provide tokens that can be used to assess the single or multi-modal data received at step 102.

[0032] At step 110, the process 100 determines a classification of the single or multi-modal data, based on the alignment scores. A linear function may be performed to assess the alignment scores for each of the data token comparisons. A variety of functions can be used in terms of how to determine the ultimate classification, and the specific diagnostic task and anatomy may influence the appropriate function to use. For example, where the knowledge base from which the diagnostically-relevant criteria were derived indicates that all of the criteria are required to be present in order to make a diagnosis, then the function can simply assess whether each alignment score meets a given confidence threshold. In other examples, some criteria or pairs of criteria may be more heavily weighted.

[0033] At step 112, the process 100 outputs a result to a user (e.g., the user that requested the diagnosis). For example, the output may include the classification determined by the model (e.g., ‘cancer’ or ‘not cancer’), as well as a description of exactly which criteria supported that classification. For example, the output may indicate that there was a high degree of likelihood that a given image displayed a criteria from the domain knowledge, such as nodules in a given anatomy of interest (assuming the presence of nodules was a diagnostically relevant criteria on which data concept tokens were established) as the reason or one of multiple reasons the classification was made. As another example, a negative output classification (e.g., ‘no cancer’) might be accompanied by a textual statement that the model determined there was not likely a presence of lesions in a given tissue, and that the lack of such lesions is (based on current standard of care) determinative that the patient does not have the type of cancer that was being considered by the diagnostic model.

[0034] It should be understood that the specific output of process 100 could take many forms, and that the classification need not be a binary decision. Rather, a confidence score could be output instead of a strict classification, or an alert to request that a human expert examine a specific image(with optionally an indication of any diagnostic criteria that were inconclusive or close to being present). Alternatively, the output could include more substantial discussion of the criteria and explanation of their relevance so that the patient could also understand the output and risk factors.

[0035] Referring now to FIG. 2, a flow diagram is shown illustrating an example process 200 for creating a diagnostic model for a given diagnostic task. The process / model described above with respect to FIG. 1 may be generated using the process of FIG. 2.

[0036] At step 202, the process 200 may receive certain preliminary information, such as an identification of the diagnostic task that the model is being developed to perform (including, for example, the types of images that the model will evaluate, the anatomy or biostructures that will be shown in the images, and other user preferences (e.g., format of the output). The process 200 then obtains textual classification criteria for this task. In some embodiments, this may be done through prompting of an LLM, through querying human experts, or a combination of both. For example, a pathologist may be a user of process 200 and desires to generate a model that learns data concept tokens specific to the pathologist’s preferred diagnostic criteria for the types of images that are generally obtained or obtainable from the radiology / laboratory equipment available within the pathologist’s home healthcare institution.

[0037] In an option step 204, the process 200 may obtain a set of training images and associated ground truth labels. These images may simply be labeled according to the confirmed classification of disease state, or may also contain textual information describing why the classification was made (e.g., a radiologist’s annotations to a medical image which describe diagnostically-relevant features of the image).

[0038] At step 206, the process 200 then catalogs a set of visual characteristics based on the diagnostic domain knowledge obtained in step 202. In some embodiments, this may include developing a set of classes or characteristics that will be the basis of visual or data concept tokens. At step 208, a ground truth label is recorded for each of the classes.

[0039] At step 210, the criteria or classes are then embedded as knowledge anchors via the VLM text encoder. As described below, these anchors ensure that the domain knowledge obtained from the LLM or human expert will remain relevant (or even the primary or only) factors on which classification of the disease state is made.

[0040] At step 212, a visual or data concept token is designated for each of the criteria. At step 214, the visual or data concept tokens are updated by comparing to characteristic embeddings toincrease the similarity of the token to associated criteria, as well as to decrease similarity to nonassociated criteria.[00411 Referring now to FIG. 3, a block diagram is shown which illustrates relationships and data flow among possible users and computational resources in various practical applications. In FIG. 3, a system 300 may include a computational resource 310, and optionally a user interface 302, a remote client / institution resource 304 (e.g., an electronic medical record system of a healthcare institution), and a remote user interface 306. In some embodiments, the resources and computational devices may be connected via a network 308, such as a local / private network, an Internet connection, or other type of network connection.

[0042] In some embodiments, a user may include a healthcare professional 302 who has obtained or ordered a test to be performed on a patient from which a diagnosis will be made. The individual 302 may enter information into a portal associated with system 300, indicating the disease state for which the individual 302 is requesting a diagnostic as well as the nature of the test that has been or will be performed (e.g., a whole-body CT scan, a biopsy, etc.). The institution 304 with which the individual 302 is associated may then perform the test or receive results of the test, and communicate them to a computational device 310 for diagnostic processing.

[0043] Computational device 310 may be operated by institution 304, may be a remote / cloud resource controlled by institution 304, or may be run by a third party provider. Computational device 310 comprises a communications system 318 that allows it to receive medical images (e.g., in connection with a P ACS or EMR). It may also include a display 316 and inputs 320, for example if interaction with a human expert is desired. The computational device also includes a memory 314 and a processor 312.

[0044] In some embodiments, the computational resource may already have an appropriate diagnostic model stored in memory 314 that can be used for a given diagnostic requested by user 302. In other embodiments, a new explainable diagnostic vision / text model 322 can be quickly generated with minimal or no input by a human expert.

[0045] The computational device 310 may thus run either or both of process 100 and process 200. The output of computational device 310 after running process 100 may be sent back to the original requesting user 302, or to an alternative user or location 306, such as to a patient, a third party expert (e.g., to confirm a close classification), or elsewhere.

[0046] Aspects of the present invention are described above with reference to flowchartillustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.Example Methods and Embodiments

[0047] In one example, a simple yet effective framework, referred to as “Explicd” for ease of reference, is disclosed: towards Explainable language-informed criteria-based diagnosis. Explicd initiates its process by querying domain knowledge from either large language models (LLMs) or human experts to establish diagnostic criteria across various concept axes (e.g., color, shape, texture, or specific patterns of diseases). By leveraging a pretrained vision-language model, Explicd injects these criteria into the embedding space as knowledge anchors, thereby facilitating the learning of corresponding visual concepts within medical images. The final diagnostic outcome is determined based on the similarity scores between the encoded visual concepts and the textual criteria embeddings. Through extensive evaluation of five medical image classification benchmarks, Explicd has demonstrated its inherent explainability and extends to improve classification performance compared to traditional black-box models.

[0048] As illustrated in Fig. 4B, human experts make diagnoses by meticulously analyzing key image features from color, shape, size, or specific patterns regarding the disease symptom, staging, and outcome. In essence, human-expert decisions are grounded in a set of differential diagnostic criteria, enabling us to distinguish between various medical conditions with confidence and transparency.

[0049] Figure 4A illustrates current deep learning models that often function as black boxes, offering predictions without revealing insights into their decision-making processes. FIG. 4B is adepiction of the decision-making process by human experts for skin lesions, grounded in domainspecific knowledge and precise criteria, facilitates explainable diagnoses.[00501 Typically, medical image classification benchmarks provide only the images and corresponding labels, largely omitting the detailed diagnostic information. Embodiments herein addresses this gap by querying domain knowledge from either large language models (LLMs) like GPT-4 or directly from human experts to formulate comprehensive diagnostic criteria, including key visual aspects such as color, shape, texture, or specific patterns associated with the classification task. We catalog the characteristics of each class based on these criteria axes. Within Explicd, these criteria are embedded as knowledge anchors using a VLM’s text encoder. Moreover, we use a set of visual concept tokens to encode visual concepts from images along these criteria axes. An intermediate criteria anchor contrastive loss encourages high similarity between the encoded visual concepts and the corresponding positive knowledge anchors. Finally, a linear layer predicts the final diagnosis class by integrating the alignment scores from all criteria.

[0051] FIG. 5 is a conceptual overview of one embodiment of an Explicd-type framework: Domain knowledge is queried from LLMs or human experts across criteria axes. Explicd then aligns encoded visual concepts with textual knowledge anchors, facilitating the learning of visual concepts. The final diagnostic prediction is made based on the alignment scores between visual and textual concepts with a linear function. Embodiments of Explicd may be a simple yet effective framework for explainable language-informed, criteria-based diagnosis via vision-language models. Embodiments may utilize domain knowledge from LLMs or human experts, defining diagnostic criteria for each class across specific concept axes. Embodiments may include a visual concept learning module alongside a criteria anchor contrastive loss to align fine-grained visual features with diagnostic criteria. Explicd demonstrates superior interpretability and performance compared to traditional black-box models on five public benchmarks.

[0052] Example Methods

[0053] Domain knowledge query and diagnostic criteria formulation: Disease diagnosis usually centers around various criteria axes describing distinctive characteristics across clinical classes. Drawing from inspiration, we first query domain knowledge from LLMs or consult human experts and formulate them into textual diagnosis criteria. Consider a set of training image-label pairs D = {( ,y)}, where x is the image and y G y is a label from a set of N classes. Specifically, we categorize the diagnosis criteria along K disentangled criteria axes specified by languagedepending on the task {CJ^. For instance, in the case of skin lesions, the criteria axes include asymmetry, border, color, diameter, texture, pattern. Subsequently, we query detailed knowledge on the typical characteristics for each class along each criteria axis Ct= {c , ... , c.1'). where 1 < rii < N denotes the number of possible options within a particular criteria axis. Take the color of skin lesion as an example, potential options could range from ‘a mixture of black brown and blue’ for melanoma, ‘uniform pigmentation of tan or brown’ for melanocytic nevus, among others. Notably, the quantity of typical characteristics for each criteria axis, n,-, may be less than the number of classes N, as different classes might exhibit identical characteristics for certain criteria axes; e.g., various types of benign skin lesions could all present symmetry. Additionally, the ground truth label for each diagnostic criterion is recorded, i.e. for each criteria axis, the associated class and characteristic options are marked as positive, whereas all other combinations are considered negative.

[0054] Visual Concept Learning

[0055] After collecting the textual form diagnostic criteria, we aim to align the visual features with these textual human knowledge. In particular, we propose a lightweight visual concept learning module for aligning the fine-grained visual concepts and the nuanced textual criteria. Specifically, given a pretrained vision-language model with visual encoder V and text encoder T, we first encode the queried diagnostic criteria into criteria anchor embeddings, {e; =GniXd, and d is the dimension of the embedded token. These criteria embeddings act as a sparse representation of human knowledge, serving as anchors to facilitate the learning of visual concepts.

[0056] To capture visual concepts effectively, our visual concept learning module employs a set of K learnable visual concept tokens p 6 7? / <xri, with each token designated to represent one of the K criteria axes. For a given image x and its feature map V(x), the concept encoding process is formalized as follows: p = cross — attention(p,V(x), V(x)), (1)

[0057] where p is the query and V(x) serves as the key and value of the cross-attention layer. The visual concept tokens p interact with the image feature map, thereby encoding the relevant visual concepts associated with specific criteria axes into p G LlKxd.

[0058] The learning of visual concepts is facilitated by a contrastive loss. For each criteria axis, the aggregated visual concept token p, is compared against all characteristic embeddings et,calculating a similarity score. The criteria anchor contrastive loss is as follows:

[0059] where T denotes the temperature parameter that adjusts the softness of the softmax distribution and we use dot product as the similarity function. The criteria anchor contrastive loss aims to increase the similarity between the encoded visual concept token p(and the positive criteria anchor embeddings while decreasing its similarity with the embeddings of negative characteristics, ensuring a more discriminative learning of visual concepts along each diagnostic criteria axis.

[0060] Explainable classification:

[0061] The above knowledge anchor loss £anchor enables the alignment of encoded visual concepts with the corresponding characteristic options along each diagnostic criteria axis. Intuitively, the similarity scores between the encoded visual concept token p, and the diagnostic criteria anchor ezindicate the model’s assessments for each diagnostic criterion. Mirroring the approach of human experts, who make their final diagnosis on the evaluations across multiple criteria, we use a linear layer to make prediction of the final class by integrating the alignment scores from all K criteria axes. y — W (concat

[0062] where concat(, ) represents the concatenation operation and W is the weights in the linear layer that inherently reflect the significance of each diagnostic criterion’s contribution towards the overall class prediction.

[0063] During the training phase, we optimize a joint objective that includes both the criteria anchor contrastive loss with cross-entropy loss for the final classification:

[0064] The embeddings of textual criteria anchors are precomputed and stored, ensuring that the training and inference overhead introduced by the additional components are negligible.

[0065] Experiments

[0066] Experimental setup:

[0067] We evaluate our method on five publicly available medical image classification benchmarks, which cover a diverse range of medical targets and modalities.

[0068] Dataset: ISIC2018 contains 10,015 dermoscopic images with seven skin lesioncategories for skin cancer classification. NCT-CRC-HE(NCT) includes 100,000 patch-based histological images of human colorectal cancer for training and 7180 patches for validation, with nine tissue classes for classification. IDRiD consists of 516 retinal fundus images annotated with 5 severity level grading of diabetic retinopathy. BUSI dataset contains 780 ultrasound images of breast masses categorized into normal, benign, and malignant classes for breast cancer classification. MIMIC-CXR contains 377,100 chest X-ray images. Cardiomegaly (CM) and Edema are used for binary classification.

[0069] Baselines. We compare our method with several baselines, including (1) VLMs zero-shot: We apply the general VLM CLIP and biomedical VLMs BioViL and BiomedCLIP [?] in a zero-shot setting for classification; (2) Supervised black-box models: We fine-tune ImageNet-pretrained ResNet50 and ViT-Base on the classification benchmarks; and (3) LaBo: a state-of-the-art explainable model with concept bottleneck.

[0070] Implementation Details. We prompt https: / / platform.openai.com / docs / models / gpt-4-and-gpt-4-turboGPT-4 to query domain knowledge and diagnostic criteria. We use the official implementation and pretrained weights of https: / / huggingface.co / docs / transformers / model_doc / clipCLIPViT-Base, https: / / huggingface.co / microsoft / BiomedVLP-BioViL-TBioViL and https: / / huggingface.co / microsoft / BiomedCLIP- PubMedBERT_256vit_base_patchl6_224BiomedCLIP. Our Explicd and LaBo are implemented based on BioViL-specialized for MIMIC-CXR dataset, while using BiomedCLIP for all other datasets. The fine-tuning of Explicd involves optimizing visual encoder, visual concept learning module and the final linear layer with AdamW optimizer, while keeping the text encoder fixed. All experiments are conducted using PyTorch with Nvidia A6000 GPUs.

[0071] Table 1 : Performance comparison across five benchmarks. Balanced accuracy is reported for CM and edema in MIMIC-CXR due to class imbalance; accuracy is reported for the other datasets.

[0072] Results.

[0073] Our proposed Explicd model demonstrates superior performance compared to various baseline methods across five medical image classification benchmarks, as shown in Table 1. The zero-shot performance of VLMs, including CLIP, BioViL, and BiomedCLIP, is generally poor, with CLIP performing close to random guessing across all datasets. This indicates that CLIP’S visual-text alignment is not effective for complex medical diagnosis tasks as it is trained on general vision data. Although BioViL and BiomedCLIP perform much better on the MIMIC-CXR dataset (CM and Edema tasks), likely due to their pretraining data being largely based on chest X-ray radiology reports, their performance on other datasets remains much lower than supervised trained black-box models like ResNet50 and ViT-Base, suggesting their limited generalization ability.

[0074] Explicd effectively combines explainability with high-level classification performance, outperforming not only the explainable model LaBo but also black-box models across all datasets. LaBo’s lower accuracy compared to the black-box models can be attributed to its reliance on well-aligned vision-language models, highlighting the challenge of maintaining high accuracy while providing strong explainability. In contrast, Explicd’ s superiority is due to the introduction of human knowledge and visual concept learning, which provides additional supervision for fine-grained alignment between visual features and diagnostic criteria. This strategy not only enhances fine explainability but also brings improvement in overall classification performance.

[0075] Diagnostic interpretation:

[0076] A distinguishing design of Explicd is its appealing ability to interpret its decision-making process. Fig. 2 (a) shows the alignment scores measured with cosine similarity between the encoded visual concept tokens and the embeddings of diagnostic criteria for skin lesions. The width of the lines indicates the strength of similarity, with larger widths representing higher similarity scores. We can see that Explicd can accurately predict the characteristics of each criteria axis, such as the presence of asymmetry, border irregularity, color variegation, and large diameter, which are key features in the diagnostic criteria we queried for melanoma diagnosis. The highsimilarity scores between the visual concepts and the corresponding diagnostic criteria demonstrate that Explicd has learned to identify and align these important visual features, leading to a correct final diagnosis of melanoma.

[0077] Furthermore, we visualize the heatmap of the average visual concept tokens with the image feature map of cardiomegaly in a chest X-ray image in Fig. 2 (b). The brighter regions indicate higher similarity scores, suggesting that the model is focusing on these areas when making its prediction. Cardiomegaly is a medical condition characterized by an enlarged heart. In the heatmap, we can observe that Explicd correctly focuses its attention on the heart area, indicating that Explicd has learned to align human knowledge regarding the key visual features of cardiomegaly with the relevant visual concepts in the X-ray image. By providing the alignment scores on criteria and highlighting the most important regions contributing to its prediction, Explicd provides a transparent and interpretable decision-making process that can be easily understood and verified by medical experts.

[0078] FIG. 6A shows alignment scores measured using cosine similarity between the encoded visual concept tokens and diagnostic criteria along each axis for skin lesion classification. The width of the lines represents the strength of similarity, with wider lines indicating higher scores. FIG. 6B shows a heatmap visualization of the encoded visual concept tokens overlaid on the image feature maps for a case of cardiomegaly. Brighter regions indicate higher similarity scores, suggesting a stronger focus on these areas by the model.

[0079] To address the lack of transparency and interpretability in current deep learning models for medical image analysis, we proposed Explicd, a comprehensive framework that integrates diagnostic criteria queried from LLMs, aligning visual concepts towards explainable classification. Explicd offers a novel means to understanding diseases along human- understandable criteria axes. Our extensive experiments highlight Explicd’ s superior performance over both traditional black-box approaches and existing explainable models, setting a new standard in both accuracy and interpretability. The clarity of Explicd’ s decision-making process promises to bolster trust and facilitate the integration of Al in clinical diagnostics. Moving forward, we aim to expand the incorporation of broad human knowledge within our diagnostic criteria and to refine the hierarchical representations of visual concepts, allowing for a more nuanced exploration of disease diagnosis and management.

[0080] FIG. 7 illustrates a process diagram for generating a phenotypic risk predictionaccording to some embodiments. FIG. 7 is an overarching diagram meant to illustrate certain aspects described above by visualizing possible data flow and potential interactions among different resources. In some examples, the process diagram illustrated in FIG. 7 may be associated with the processes and systems described above with respect to FIGS. 1-3. At block 702, images are provided as inputs (initially these may be training inputs; later they may be inputs meant for deployment / runtime classification) to a classification model pipeline or framework that has been structured to provide explainable outputs as described herein. In some examples, the images may include annotated images, based on previously-obtained and analyzed data sets, as well as pre- RCT annotated images or video. The input images are then pre-processed at block 704. For example, the pre-processing may include ROI detection, image augmentation, and the generation of feature maps. In some examples, additional text input may be received at block 706. The text inputs can include annotations (e.g., infectious stages), EMR phenotypic variables, as well as answers to survey questions. In some examples, this additional text may be used for training of an associated classification model using criteria anchors with contrastive loss. Moreover, the pre- processed data may further be provided as an input to a transformer with one or more crossattention mechanisms, as shown in block 708. At block 708, text associated with characteristics such as asymmetry, borders, colors, diameters, or the like may be anchored as visual or data tokens. Moreover, textual concepts associated with a prompt of interest may be extracted. At step 710, an output may be generated using both the analysis performed at 708, alone or in combination with the processing of the text inputs received at step 706.

[0081] Though the frameworks shown in FIG. 7 (and described in the context of certainExamples) related to classification of certain image types, it should be understood that other data types can also be processed for structured, explainable classification by leveraging the concepts described herein. For example, beyond simply medical image data, other image types and other data formats can also be used as inputs to a novel classifier having a process and structure similar to that shown in FIGs. 1-3 and 7 and described in the Examples section. Methods for determining human-understandable criteria that can be used as “anchors” in a structured, explainable classification model can be tailored to different types of input formats (and combinations thereof), while preserving the ability for the model to achieve high accuracy in an explainable way.

[0082] For example, a process 800 as illustrated in FIG. 8 may be performed to determine one or more tokens or textual descriptions for use in a structured, explainable classification system,allowing the system to associate human-interpretable criteria (which, e.g., a human expert would rely upon to make a diagnosis or classification) with machine-learned representations. The “descriptors” (which may be words, short phrases, graphs, relative comparisons, or other features / criteria, etc.) can be used, as discussed herein, to have a model learn features present in training / input data (e.g., images, signals, sensor data, etc.) that align with “axes” of variation in the input data that are human-recognizable, probative, and extractable from the input data.

[0083] In some instances, a process 800 may be performed once for a given classification task (such as for a given data type, domain, training set, etc.), or may be performed multiple times for different tailored datasets and / or when new or different anchors are desired (e.g., if diagnostic protocols change, meaning other descriptors are available).

[0084] At block 802, process 800 may receive task information associated with a classification objective. The task information may include: a set of potential class labels, categories, or conditions that a trained classification model can use (e.g., possible outputs); attributes of input data (e.g., parameters or fields existing in the input data); domain information (e.g., descriptions of relevant protocols, source / expert materials, etc.), and the like. In some implementations, the task information may further include domain context (such as information about experts that currently perform the task, like qualifications; methods for discerning accuracy; tolerance to mis-predictions (e.g., tolerance to false positives vs false negatives for the domain); output format (e.g., whether to include confidence scores for multiple outputs, whether class labels are the only outputs or whether continuous values should be used); prior examples; or whether metadata relating to the input data and / or classification setting can or should be utilized.

[0085] At block 804, process 800 may optionally include an analysis and / or pre-processing of data associated with the classification task. The input data may be uni- or multi- modal, and may include, but is not limited to, any one or more of multiple types of data such as: image-style data, such as depth images / point clouds, video sequences, grayscale intensity maps (e.g., for noncolor images or intensity measurements), thermal images, etc.; sensor or instrument readout data such as spectrograms, electrical or electrochemical measurements, etc.; biomedical data, such as images, videos (e.g., colonoscopy), patient history, blood tests, urine tests, other sample-tests, cognitive or behavioral test results, etc.; and / or environmental sensor measurements such as motion, acceleration, and directional data, vibration, acoustic, and audio signals, and / or temperature, pressure, etc. Based on the identified data types and available attributes of the inputdata type (which can be gleaned from training data, sample data, and / or entered by a user), process 800 may determine one or more feature dimensions that are available in the data and potentially interpretable by a human observer. For example, if the data comprises optical / RGB images, the available dimensions may include color, shape, texture, presence / number of objects of interest, border characteristics of objects of interest, etc. As another example, if the data includes grayscale medical images (e.g., MRI images and / or other data), the available dimensions may include intensity (rather than color channels), contrast, and presence of various structures. If the data includes audio signals such as electrical signals recorded via a microphone, the available dimensions may include frequency content, amplitude, and temporal patterns.

[0086] At block 806, process 800 may generate a collection of potential textual descriptors that can be aligned with the identified feature dimensions. (In some cases, blocks 804 and 806 can be addressed simultaneously or in alternate order.) In some embodiments, the descriptors may be manually provided by a user or domain expert. In other embodiments, the descriptors may be extracted from curated domain-specific sources, such as diagnostic manuals, clinical protocols, or scientific articles, such as by natural language processing or LLM-analysis via retrieval augmented generation. In further embodiments, a language model may be used to generate descriptors by prompting it to return descriptive terms relevant to the task and constrained to dimensions that correspond to the input data attributes. In yet other embodiments, descriptors may be statistically derived from text associated with labeled data, such as symptom descriptions, diagnostic reports, or user-provided annotations. In such cases, process 800 may identify recurrent terms or clusters of semantically similar expressions that correlate with particular class labels or outcomes and align with available data features.

[0087] At block 808, the candidate descriptors may be filtered and normalized to remove redundancy and ensure consistency. Filtering operations may include removing irrelevant words or parts of speech, standardizing phrasing, converting plural forms to singular, and mapping synonymous terms to a conceptual or categorical grouping. In some implementations, clustering techniques may be used to group related descriptors under a shared label. In some cases, a human reviewer may curate or validate the filtered list.

[0088] At block 810, each descriptor may be proj ected into a shared embedding space used by the classification model. In some embodiments, various machine learning models, neural networks, transformer models, or the like may be utilized to convert the descriptors (whichcomprise human-understandable concepts) into representations (like a numerical vector) of a given format, so they can be compared to representations of data features from input data (like vectors extracted as features of an image and / or sensor output, including multiple modalities or a single modality) in a common and comparable way. For example, a pretrained vision-language model may be used to generate embeddings for both textual descriptors and image features. Process 800 may also determine similarity measurements between some or all of the potential candidate descriptors and corresponding regions or latent features of the sensor data. In further implementations, process 800 may also generate second or third order derivations of the input data, such as transform operations (e.g., variations on Fourier transform functions), domain shifting (e.g., mapping data as though it were an image, to leverage native functions and pretraining of vision-language models), calculating derivatives or integrations (e.g., of accelerometer or RF data), interpretations (e g., estimating sound from motion data), etc. and attempt to align the results with candidate descriptors as well.

[0089] At block 812, a final set of tokens may be selected from the candidate descriptors based on various criteria. The selected tokens may be those that achieve high alignment with features of the input data, provide broad coverage across class distinctions, and are found by human users to have easy interpretability within the domain. In some embodiments, a refinement step may be used to calibrate the embeddings of the selected tokens using domain-specific training data.

[0090] Once a set of tokens are selected, a classification model can be structured and trained according to the various examples and descriptions herein.

[0091] Accordingly, it can be seen that the systems, methods, and advantages described herein may be leveraged for use beyond medical images and related biomedical information.

[0092] In the foregoing specification, implementations of the disclosure have been described with reference to specific example implementations thereof. It will be evident that various modifications may be made thereto without departing from the broader spirit and scope of implementations of the disclosure as set forth in the following claims. The specification and drawings are, accordingly, to be regarded in an illustrative sense rather than a restrictive sense.

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A method for classifying input data, the method comprising: receiving input data for a given classification task; generating, using a structured explainability model, a set of input data representations of features of the input data, each input data representation corresponding to a respective concept token associated with the classification task, wherein the concept tokens are based on representations of human-interpretable criteria for the classification task; evaluating similarity of the input data representations to the concept tokens, to obtain a set of concept scores for the input data; processing the set of concept scores via a classification output layer to obtain classification information for the input data; and outputting the classification information to a user with associated explanation information, the explanation information derived from the set of concept scores and the human- interpretable criteria.

2. The method of claim 1, wherein the input data comprises more than one data modality.

3. The method of claim 1, further comprising: receiving a diagnostic request from a user; determining attribute information for the input data; and determining whether a library of structured explainability models contains a trained model comprising an output layer corresponding to the diagnostic request and a set of concept tokens based on representations of human-interpretable criteria extractable from the attribute information.

4. The method of claim 1, wherein the human-interpretable criteria comprise diagnostic criteria categorizations that were derived from domain knowledge obtained from alarge language model via an automatic query formulated based on an anatomy, medical image, and diagnosis of interest.

5. The method of claim 1, wherein the step of evaluating similarity of the input data representations to the concept tokens includes evaluating similarity of embeddings of the input data obtained from a pre-trained, cross-attentional transformer model in a shared embedding space with the concept tokens.

6. The method of claim 1, wherein the concept tokens comprise visual concept tokens derived by aligning visual embeddings of training data to diagnostic criteria embeddings obtained from diagnostic criteria from domain knowledge, to increase similarity of the visual embeddings to the associated criteria embeddings, and decrease similarity of the visual embeddings to non-associated criteria embeddings.

7. The method of claim 6, wherein each concept token corresponds to one diagnostic criteria , and wherein no concept tokens share a same diagnostic criteria.

8. The method of claim 1, wherein the similarity scores are determined based on cosine similarity between an input data representation and a concept token.

9. The method of claim 1, wherein the vision-language model was trained by: providing a plurality of textual classification criteria for a domain-specific evaluation task to a vision-language model; cataloging, via the vision-language model, a plurality of visual characteristics or classes for each criterion of the plurality of textual classification criteria; inputting ground truth labels for each visual characteristic or class to the vision-language model; embedding the plurality of visual characteristics or classes as a plurality of knowledge anchors via a text encoder; designating one of the concept tokens for each criterion of the plurality of textual classification criteria; anddecreasing or increasing alignment of concept tokens to visual characteristics.

10. A system for performing knowledge-based classification of data, the system comprising: one or more processors; a vision-language model trained to perform a task of interest; a memory in communication with the one or more processors, wherein the memory contains instructions that, when executed, cause the one or more processors to: receive data for a given classification; determine modality information of the data; provide the data and the modality information to the vision-language model, wherein the task of interest corresponds to the given classification; evaluate visual characteristics of the data across a set of data concept tokens corresponding to criteria categorizations for the task of interest; assess alignment scores for pairings of data concept tokens of the set of data concept tokens and associated visual characteristics; determine classification information for the data based at least in part on the assessment of the alignment scores, the classification information responsive to the task of interest; and output the classification information to a user with associated explainable support for the classification information derived from the alignment scores and data concept token characteristics.

11. The system of claim 10, wherein the data is single or multi-modal data.

12. The system of claim 10, wherein the instructions further cause the one or more processors to: receive a diagnostic request from a user; and determine whether a library of explainable visual models contains a corresponding explainable model matching the task of interest.

13. The system of claim 10, wherein the criteria categorizations are diagnostic criteria categorizations that were derived from diagnostic knowledge obtained from a large language model via an automatic query formulated based on an anatomy, medical image, and diagnosis of interest.

14. The system of claim 10, wherein the one or more processors provides the data to a pre-trained vision language model when evaluating visual characteristics.

15. The system of claim 10, wherein the data concept tokens are visual concept tokens that were updated during a model development process by comparing the visual concept tokens to characteristic embeddings to increase the similarity of the token to an associated diagnostic criteria categorization and decrease the similarity of the token to non-associated diagnostic criteria categorizations.

16. The system of claim 15, wherein each visual concept token corresponds to one diagnostic criteria categorization, and wherein no visual concept tokens share a same diagnostic criteria categorization.

17. The system of claim 10, wherein the alignment scores were determined using one or more cosine projections extracted from the textual knowledge.

18. The system of claim 10, wherein the vision-language model was trained by: providing a plurality of textual classification criteria for a domain-specific evaluation task to the vision-language model; cataloging, via the vision-language model, a plurality of visual characteristics or classes for each criterion in the plurality of textual classification criteria; inputting ground truth labels for each characteristic or class in the plurality of visual characteristics or classes to the vision-language model; embedding the plurality of visual characteristics or classes as a plurality of knowledge anchors via a text encoder;designating the data concept tokens for each criterion of the plurality of textual classification criteria; and decreasing or increasing a similarity of data concept tokens based on a comparison of each data concept token to the plurality of visual characteristics.

Citation Information

Patent Citations

  • Explainable deep learning camera-agnostic diagnosis of obstructive coronary artery disease

    US20230309940A1

  • Systems and methods for providing explainability of natural language processing

    US20250165712A1