Multimodal data collection and diagnosis of rare ocular diseases
The multimodal data approach using an attention-based graph neural network and a curated concept bank improves diagnostic accuracy for rare ocular diseases by integrating diverse imaging modalities and textual data, addressing the challenge of scarce patient data and misdiagnosis.
Patent Information
- Application Number
- PCT/US2024/059516
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-05
- Filing Date
- 2024-12-11
- Publication Date
- 2025-07-10
AI Technical Summary
The diagnosis of rare ocular diseases is hindered by the scarcity of patient data and the difficulty in distinguishing between similar eye diseases, often leading to misdiagnosis by primary healthcare providers.
A multimodal data approach using an attention-based graph neural network (AM-GNN) that processes image-text pairs from various imaging modalities, including Fundus, MRI, and OCT, to generate latent representations and predict disease labels with confidence scores, leveraging a concept bank curated through iterative prompting of large language models.
Enhances diagnostic accuracy for rare ocular diseases by integrating diverse data sources, improving the understanding of complex conditions and reducing misdiagnosis through robust representation and decision-making.
Smart Images

Figure IMGF000024_0001 
Figure IMGF000025_0001 
Figure IMGF000025_0002
Abstract
Description
MULTIMODAL DATA COLLECTION ANDDIAGNOSIS OF RARE OCULAR DISEASESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority to and the benefit of U.S. Provisional Application Number 63 / 618,093 filed on January 5, 2024, which is hereby incorporated by reference in its entirety for all purposes.BACKGROUND
[0002] The prevalence of eye diseases represents a global challenge, affecting individuals' well-being. A considerable portion of the world’s population suffers from various forms of vision impairment, and this number is expected to rise substantially in the coming years. Beyond personal impact, vision impairment may place a considerable economic burden on society, contributing to substantial global productivity losses. The growing number of affected individuals highlights a demand for effective interventions. In addition to the economic costs, the burden may extend to healthcare systems for providing access to eye care services. Furthermore, vision loss often leads to complications and exacerbates comorbid conditions such as depression, cardiovascular diseases, diabetes, and hypertension. Recognizing the importance of early detection, timely intervention may contribute to slowing disease progression and preserving vision. To effectively manage vision impairment, widespread access to efficient eye care services may be effective, while also addressing the associated burdens to individual or healthcare systems.
[0003] Additionally, by definition, rare diseases affect a small portion of the population. Consequently, the limited number of individuals with these conditions often leads to a scarcity of patient data, making it difficult to fully understand the diseases. Furthermore, data related to rare diseases is frequently scattered across various healthcare institutions, complicating efforts toconsolidate and analyze it effectively. This fragmentation of data sources can hinder the creation of complete datasets needed for accurate analysis and diagnosis. Additionally, machine-learning models perform effectively when trained on diverse and representative datasets by learning patterns and generalizing to new, unseen data. With insufficient data for rare diseases, these models may struggle to generalize, resulting in poor performance on new cases.
[0004] The initial diagnosis by primary health care provider (general practitioner, optometrist, or local hospital emergency department) may help the ophthalmic outcome of patient. There is a possibility of misdiagnosis of acute eye diseases by primary health care providers in clinical practice. Specifically, given the plethora of eye diseases with highly similar features, it may be difficult to distinguish between the different eye diseases. For instance, iritis (anterior uveitis), a relatively common cause of red eye, is often misdiagnosed as conjunctivitis. If anterior uveitis is mistaken as conjunctivitis, it could lead to pain and blurred vision. Another common instance is the confusion of ocular ischemic syndrome (OIS) with diabetic retinopathy (DR) and central retinal vein occlusion as the exam of OIS often involve dot and blot hemorrhages, including dilated veins and narrowed arteries, which are similar with DR. Therefore, for accurate diagnosis of ocular diseases, broader features may be effective.SUMMARY
[0005] Certain embodiments of the present disclosure relate to diagnosis of ocular diseases by leveraging an attention-based multimodal graphical neural network (AM-GNN) based on multimodal data associated with different imaging modalities. The multimodal data may include one or more image-text pairs, where each image of the image-text pair may depict a retinal scan of a subject that may correspond to one or more imaging modalities e.g. Fundus, magnetic resonance imaging (MRI), and optical computed tomography (OCT). By applying one or more machine-learning models, latent representations (interchangeably used herein with embedding vectors) corresponding to each image-text pair may be generated. In some examples, the one or more machine-learning models may process an image and its associated text separately, using distinct encoders such as vision transformers for image embeddings and BERT (bidirectional encoder representations from transformers) for text embeddings. Alternatively, the one or more machine-learning models may process each image-text pair jointly by mapping it into a sharedmultimodal space, for example, using a model similar to ViLBERT (vision-and-language BERT).
[0006] In some aspects of the present disclosure, a concept bank may be accessed for prediction of a set of related concepts that may be present in the one or more image-text pairs. The concept bank may include e.g. modality-specific concepts and disease-specific concepts that may contribute to the prediction of a disease label of a set of disease labels. The concepts may include terminologies, synonyms, attributes, symptoms, ontologies and features related to various imaging modalities and disease labels that may support accurate prediction of each of the set of disease labels. The concept bank may be curated by prompting a large language model (LLM) with disease-guided prompts and modality-guided prompts. The LLM may be iteratively queried based on the crafted prompts, generating responses that may include unique concepts specific to each imaging modality (or each disease label) as well as shared concepts spanning multiple modalities (or disease labels). This iterative prompting may assist in collecting diverse concepts in the context of ocular disease diagnosis.
[0007] The one or more latent representations associated with the one or more image-text pairs may be fed to the attention-based multimodal graph neural network (AM-GNN) that generates a multimodal graph based on the concepts from the concept bank. The AM-GNN may be configured to generate associations between a set of nodes that includes the one or more image-text pairs, and the concepts associated with each disease label of the set of disease labels. The associations may be identified, based on an attention mechanism, iteratively between each pair of nodes of the set of nodes. The attention mechanism may include determining an attention score between each pair of nodes of the set of nodes by applying a compatibility function that computes relevance between latent representations associated with each pair of nodes. A nonlinearity may be introduced to the computed attention scores (e.g., by a non-linear activation function such as ReLU (rectified linear unit) or leaky ReLU) that may be further normalized (e.g., by SoftMax or sigmoid) to generate a final attention score.
[0008] The AM-GNN may predict the set of related concepts that are present in the one or more image-text pairs based on the multimodal graph, generating an output that includes a predicted disease label. The output may further include a confidence score for the predicted disease label, or for each of the set of disease labels. The confidence score may represent adegree of certainty of the AM-GNN in its prediction, quantifying a likelihood that the predicted disease label accurately corresponds to the input multimodal data based on the learned associations with the concepts. Higher confidence scores indicate stronger associations between the one or more image-text pairs and the predicted disease label, while lower scores suggest a higher degree of uncertainty, potentially requiring further validation or additional analysis. Additionally, the output may include the set of related concepts present in the input multimodal data along with respective confidence scores, indicating a degree to which each concept is relevant to the input multimodal data.
[0009] For prompting the LLM, the modality-specific prompts and disease-specific prompts may be crafted in a natural language format. The concepts from the generated responses may be extracted by applying one or more natural language processing (NLP) techniques, such as named entity recognition (NER), part-of-speech tagging, and dependency parsing, to identify and categorize relevant terminologies and features that comprise concepts. Additionally, the extracted concepts may undergo further refinement and validation through domain-specific filtering e.g., by a domain expert to enable their relevance and accuracy in relation to the set of disease labels and imaging modalities. In some examples, the concepts may be organized into a structured format (e.g., by GNN, knowledge graphs, topic modeling or unsupervised clustering techniques), enabling efficient integration into the concept bank for use in training the AM- GNN.
[0010] In some aspects, a multimodal dataset for training the AM-GNN may be collected by a data curation technique, particularly useful for collecting data related to rare ocular diseases. The data curation technique may include retrieving content via a web crawling tool (e.g., Scrapy) by automatically navigating through one or more data sources (e.g., publicly available repositories, databases, websites, research centers etc.) using a set of initial seed universal resource locators (URLs). Once the content is retrieved, web scraping tools may be used to parse the content and extract specific data, such as image URLs or images, textual descriptions, and / or associated metadata that are relevant to the concepts from the concept bank. The collected image-text pairs may include noise or irrelevant image-text pairs, therefore, to enhance the quality of the collected dataset, clustering may be performed.
[0011] For clustering, the collected image-text pair may be fed into the one or more machine-learning models configured to generate the corresponding latent representations. These generated latent representations may be clustered into a set of clusters based on the associated inherent patterns or similarities. For each cluster, a cluster refinement process may be performed to select the relevant image-text pairs associated with the latent representations. The clusters that are not correctly associated with images or texts, associated modalities, and / or that do not represent the given set of disease labels may be dropped. The refined image-text pairs may be included in the multimodal dataset for training the AM-GNN. In some examples, crafted prompts or the extracted concepts from the generated responses may be used in the data curation technique for retrieving content.
[0012] In some embodiments, the image-text paired data related to the set of disease labels identifies ocular diseases (e.g., diabetic retinopathy, age-related macular degeneration), where the images may belong to various imaging modalities. For instance, for the diagnosis of ocular diseases, imaging modalities may include, but not limited to, Fundus, X-ray, positron emission tomography (PET), functional magnetic resonance imaging (fMRI), ultrasound, optical coherence tomography (OCT) and fluorescein angiography (FA).
[0013] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.
[0014] In some embodiments, a computer-program product tangibly embodied in a non- transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods or processes disclosed herein.
[0015] In some embodiments, a system is provided that includes one or more means to perform part or all of one or more methods or processes disclosed herein.
[0016] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms and expressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the inventionclaimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present disclosure is described in conjunction with the appended figures:
[0018] FIG. 1 is an illustrative example of ocular data from one or more imaging modalities.
[0019] FIG. 2 depicts an example workflow of multimodal data curation, building a multimodal dataset including image-text pairs from the one or more imaging modalities.
[0020] FIG. 3 illustrates an exemplary network for curation of a concept bank.
[0021] FIG. 4 shows an example illustration of a network for generating latent representations from the one or more imaging modalities in accordance with some aspects of the present disclosure.
[0022] FIG. 5 illustrates an example of an attention-based multimodal graph neural network (AM-GNN), where inputs and output are represented as graphs with nodes and edges.
[0023] FIG. 6 illustrates an example of an attention mechanism for the AM-GNN.
[0024] FIG. 7 illustrates an exemplary workflow to predict ocular diseases using one or more image-text pairs including retinal scan in accordance with some aspects of the present disclosure.
[0025] FIG. 8 is an example illustration of a computing system in which various embodiments of the present disclosure may be implemented.DETAILED DESCRIPTION
[0026] The present disclosure relates to ocular disease diagnostic techniques by leveraging an attention-based multimodal graph neural network (AM-GNN) that processes multimodal data from different imaging modalities. The multimodal data may include one or more image-textpairs, where each image of the one or more image-text pairs may depict a retinal scan of a subject corresponding to one or more imaging modalities such as Fundus, ultrasound imaging, X-ray etc. The multimodal data incorporates two aspects: (1) combination of visual information (images) with accompanying textual descriptions thus, comprising image-text pairs, and (2) integration of images from different imaging modalities, each capturing distinct aspects such as structural details, blood flow, and functional characteristics. Some medical conditions may not be fully characterized by a single imaging modality. For example, MRI may offer detailed soft tissue contrast, while PET scans can reveal metabolic activity. Utilizing this multimodal approach may enhance the understanding of complex conditions and underlying pathology by combining diverse perspectives. Thus, a deeper understanding of various diseases may be achieved that further improves diagnostic accuracy.
[0027] Although various aspects of present disclosure are described with respect to ocular diseases, it will be understood by those skilled in the art that the techniques disclosed herein are not limited to this specific application. Rather, the techniques and approaches outlined may be equally applicable to the diagnosis or prediction of a wide range of diseases across various medical domains, including but not limited to cardiovascular, neurological, gastrointestinal, and musculoskeletal diseases. The techniques, as disclosed here, including the use of multimodal data and a combination of machine-learning models including encoders for generation of latent representations, one or more NLP (natural language processing) techniques for concept extraction and the attention-based multimodal graph neural network (AM-GNN) for generating a multimodal graph, can be adapted to other disease categories. Thus, the scope of the present disclosure extends beyond ocular diseases to encompass any medical field where image-based analysis from various imaging modalities accompanied by textual descriptions can be leveraged for predictive modeling.
[0028] For example, in cardiovascular diseases, image-text pairs from modalities such as echocardiography, magnetic resonance imaging (MRI), or computed tomography (CT) scans can be utilized to assess heart function and detect conditions such as coronary artery disease, myocardial infarction, or heart failure. Similarly, in neurology, CT scans, positron emission tomography (PET) or functional MRI (fMRI) images, paired with textual data describing the observed brain structures or abnormalities, can aid in diagnosing conditions such as Alzheimer's disease, Parkinson's disease, or brain tumors. Furthermore, image-text pairs from pulmonaryimaging modalities, such as chest X-rays or CT scans, can assist in detecting lung cancer, pneumonia, or pulmonary embolism. The combination of multimodal data including image-text pairs from different imaging modalities may facilitate the extraction of features for accurate disease prediction, enabling improved diagnostic capabilities across multiple medical disciplines.
[0029] For addressing multimodal data including different imaging modalities and imagetext pairs, various approaches may be adopted. In some examples, the machine-learning models (or encoders) may process image-text pairs separately for each imaging modality to generate latent representations (interchangeably used herein with embedding vectors). Alternatively, a single machine-learning model may process image-text pairs from multiple imaging modalities to generate corresponding latent representations. Similarly, to deal with image-text modality, separate image and text encoders may be used, generating corresponding embedded vectors that may be concatenated for further processing. In some examples, joint embedding vectors may be leveraged, where encoders generate embeddings in a shared multimodal space, for example, vision-and-language pretraining (ViLBERT), contrastive learning or unified vision-and-language pretraining (UNITER). The embedding vectors associated with each image-text pair generated by one or more machine-learning models may be further processed by the AM-GNN.
[0030] For training one or more machine-learning models and AM-GNN, a multimodal dataset may be curated manually from different available data repositories. Alternatively, or additionally, a data curation technique may be performed to automatically collect relevant image-text pairs for the targeted diseases from identified sources such as websites, medical databases, research papers, healthcare repositories and the like. Th data curation may include data crawling that involves automatically navigating through the identified sources to gather relevant content, such as images and associated text, by following links across pages. Once the relevant content is found, data scraping may be used to extract specific image-text pairs from the webpages, parsing the HTML or API responses to collect the image URLs and their corresponding descriptive text.
[0031] Together, these techniques may enable the automated collection of large volumes of multimodal data from diverse online sources facilitating the creation of rich multimodal dataset(s). The data curation technique may particularly be helpful in collecting and predicting rare diseases, increasing the training data for a machine-learning model to achievegeneralizability for learning better features for the targeted diseases. It may also increase the sample distribution of diseases, which are less represented in a dataset, thereby resulting in reduced bias of the machine-learning model for diagnosing and assessing rare diseases.
[0032] The process of web crawling and web scraping to collect image-text pairs for targeted diseases may often yield diverse and potentially noisy data due to variations in webpage layouts, image quality, irrelevant images or textual content. To enhance the quality and reliability of the collected dataset, clustering analysis can be employed. For clustering, the collected image-text pair may be fed into one or more machine-learning (ML) models configured to generate the corresponding embedding vectors (latent representations). These generated latent representations may be clustered based on the associated inherent patterns. The quality of the clusters may be assessed in a refinement process where an expert in the field (e.g., ophthalmologist, clinician, researchers or equivalent) may verify the correctness of the collected image-text pairs for the specified diseases. The human annotators / experts can manually review the cluster based on the factors such as data quality, uniqueness, diversity, and relevance with the associated text. During refinement, the clusters that are not correctly associated with text, image and / or that do not represent the targeted diseases, and related information may be dropped, if found.
[0033] In some aspects of the present disclosure, a concept bank is curated by identifying the concepts relevant to the targeted diseases and various imaging modalities. In the context of ocular diseases, the concept bank may include specific eye conditions or disease names (e.g., glaucoma, amblyopia), associated symptoms (e.g., fuch’s endothelial dystrophy is an ocular disease with symptoms such as corneal swelling, blurry vision and discomfort), imaging modalities (Fundus, OCT, fluorescence angiography (FA)), anatomical structures (e.g., lens, retina, cornea, iris), synonyms, antonyms and any other relevant attributes or terminologies.
[0034] Creating the concept bank may involve a collaborative approach by one or more natural language processing (NLP) techniques and an expert in the field (e.g., ophthalmologist or equivalent) to automate and verify the prior knowledge of concepts for the targeted diseases. For identification of concepts, an iterative prompting technique may be leveraged for querying a large language model (LLM) such as generative pre-trained transformers (GPT). LLM may be queried multiple times following a series of prompts (e.g., gradually expanding to related concepts), including disease-guided, modality-guided and comparison prompts in a naturallanguage format to generate responses including concepts. The comparison prompts may be used to find similarity scores between various concepts, where the similarity may be based on semantic similarity or contextual similarity.
[0035] The concepts and terminologies from the generated responses can be extracted by the one or more NLP techniques. For example, named entity recognition (NER) techniques (e.g., rule-based NER or ML-based NER) may parse the responses, removing irrelevant data or stop words (e.g., of, the, etc.) and identify named entities, such as disease names, symptoms, and / or anatomical structures. Part-of- Speech (POS) tagging techniques may help classify words into their respective grammatical categories (e.g., nouns, verbs, adjectives), aiding in the identification of relevant features and relationships. Finally, dependency parsing techniques may analyze the syntactic structure of sentences, helping to understand how words and concepts are related. These NLP techniques may be leveraged for the extraction, categorization, and organization of concepts, leading to a clearer understanding of the terminologies within the generated response.
[0036] The collected information (i.e., extracted concepts and / or responses) can be validated by engaging domain experts in the process, such as ophthalmologists, clinicians, domain experts or researchers, to gather relevant concepts and terms to ocular diseases. The validations by the domain experts may enable the inclusion of clinically significant and accurate information. In some examples, the terminologies from the concept bank may be incorporated into prompts to guide the data crawling process more effectively. These concepts can be used as search terms or keywords, enabling the search and navigation of related webpages to retrieve relevant information.
[0037] For training various models (e.g., AM-GNN or ML models), the multimodal dataset may include image-text pairs as inputs along with tagged labels including associated concepts, similarity scores and targeted diseases. During training, the latent representations from the refined image-text pairs and concepts may be provided to the AM-GNN that learns associations between inputs and outputs, generating a multimodal graph. Generation of the multimodal graph may involve leveraging an attention mechanism to dynamically weigh the importance of information from neighboring nodes in the graph during message-passing iterations. This may allow the model to focus on salient features and enhance fusion across different modalities. Thedisclosed techniques may fuse inter- and intra- modalities relationships together with the concepts to diagnose the targeted set of diseases.
[0038] Nodes in the multimodal graph may represent instances from image-text pairs, concepts and targeted diseases, and edges may signify relationships between them. The edges in the multimodal graph may have edge features including edge weights that may be predefined based on similarity scores and / or additional edge information e.g., type of relationship such as synonyms, antonyms etc. In other instances, these relationships (edges) are established during the training phase, where the model may learn to associate certain features from the image text pairs to concepts that further relate to the targeted diseases by computing attention scores. In the multimodal graph, the concept nodes serve as intermediary nodes between input nodes (i.e., image-text pairs) and disease nodes. The concepts may be shared or unique to a targeted disease and may be mapped to one or more image-text pairs. While the targeted disease nodes may be mutually exclusive.
[0039] While message passing, information may be exchanged between nodes that are influenced by attention weights, capturing nuanced relationships within and across modalities by focusing on specific nodes. Attention parameters (weights and biases) for each node or edge may be learned during training. Attention scores for each pair of nodes determine how much importance should be given to the information from one node or edge when updating the representation of another. The scoring may be done using a compatibility function (e g., dot product, cosine similarity, or a learned linear layer). The attention scores may be normalized across all neighbors of a node or edges connected to a node to obtain attention weights. These attention weights can be used to compute a weighted sum of the representations of neighboring nodes or edges. The information aggregated from neighboring nodes may be used to update the representation of each node. This disclosed approach, combining powerful embedding vectors from individual modalities with attention-based graph fusion, may enable a robust representation of multimodal data for enhanced understanding and decision-making, making it a valuable innovation with applications across various fields of healthcare analysis.
[0040] FIG. 1 shows illustrative examples of ocular data 100 obtained from one or more imaging modalities, each providing insights into the structure and function of the eye. Different imaging techniques may capture distinct aspects, such as structural details, blood flow, tissuedensity, and functional characteristics of ocular tissues. By integrating data from multiple modalities, clinicians may gain a better understanding of ocular pathologies and their broader impact on subjects’ health. For instance, when diagnosing rare ocular diseases like retinitis pigmentosa, asteroid hyalosis, neuroretinitis, or cystoid macular edema, combining data from diverse sources may allow an accurate and thorough assessment. The imaging modalities may include, but are not limited to, e.g., Fundus imaging 102, fluorescein angiography (FA) 104, optical coherence tomography (OCT) 106, wavefront aberrometry 108, and ultrasound imaging 125.
[0041] Each of these technologies may provide different types of information, enabling clinicians to make more informed decisions. For example, Fundus imaging 102 may offer high- resolution images of the retina, optic disc, and macula, enabling the detection of common retinal conditions such as diabetic retinopathy, macular degeneration, and glaucoma. By capturing a wide-field view of the eye, it provides an overall assessment of retinal health and can reveal changes in blood vessels or the presence of lesions. Fluorescein angiography (FA) 104, which involves injecting fluorescent dye into the bloodstream, may provide dynamic images that highlight blood flow in the retina and choroid. FA may particularly be useful for identifying vascular abnormalities, such as retinal vein occlusions, diabetic macular edema, and age-related macular degeneration (AMD), by visualizing areas of leakage or non-perfusion that might not be evident with other imaging techniques. Optical coherence tomography (OCT) 106, on the other hand, offers detailed cross-sectional images of the retinal layers and the optic nerve head that assist clinicians to assess the structural integrity of the retina. OCT 106 may help diagnosing and monitoring conditions such as macular edema, glaucoma, and retinal detachment, providing information on retinal thickness and the presence of fluid buildup.
[0042] Similarly, wavefront aberrometry 108 measures the optical imperfections of eyes by analyzing the way light is refracted as it passes through the cornea and lens. This approach may contribute to the diagnosis of refractive errors, astigmatism, and higher-order aberrations that may affect visual acuity and quality. Ultrasound imaging 125 uses sound waves to generate realtime images of the internal ocular structures, including the vitreous body, retina, and optic nerve. This imaging technique may particularly be useful for evaluating conditions that involve the posterior segment, such as retinal detachment, vitreous hemorrhages, and tumors, when other imaging modalities may be limited by opacity or media clouding. Integration of diverse imagingmodalities may provide precise diagnosis, particularly in complex or rare medical cases. In addition to these, other imaging techniques including fMRI (functional magnetic resonance imaging), CT scans, and X-rays can further contribute by revealing additional anatomical and physiological features, such as orbital masses, structural abnormalities, or vascular conditions, providing a holistic view of the patient's ocular health.
[0043] FIG. 2 illustrates an example workflow of multimodal data curation 200, collecting and refining multimodal data related to ocular diseases. The multimodal data includes image-text pairs, which combine visual information from various medical imaging modalities, such as OCT 106, Fundus imaging 102, and / or fMRI with accompanying clinical text descriptions. By incorporating data from diverse imaging modalities, a better understanding and a more holistic view of ocular conditions may be observed. The combination of image-text pairs may enhance disease diagnosis by offering additional context, significantly improving the performance of machine-learning models, particularly deep learning models, which can leverage both visual and textual information. Medical images alone may often present ambiguous features, requiring further interpretation. The textual component of image-text pairs may help resolve such ambiguities, providing clarifying information that supports a more accurate and precise diagnosis.
[0044] Multimodal data including image-text pairs can be manually collected, additionally or alternatively, multimodal data crawling 204 can be employed to gather image-text pairs related to targeted ocular diseases from World Wide Web (WWW) 202. This process may involve the systematic extraction of relevant data from various online sources such as medical websites, research databases, healthcare repositories, clinical studies, and publicly available healthcare datasets. Web crawling tools, such as Scrapy or Selenium, automate the process of navigating through web pages, following hyperlinks, and discovering additional content. These crawlers may utilize a set of seed URLs, traverse through different pages, and retrieve the HTML content of each page. Once the content is retrieved, web scraping tools, such as Beautiful Soup, may be used to parse the HTML structure and extract specific data, such as image URLs or images, textual descriptions, and / or associated metadata.
[0045] For ocular diseases, the multimodal data curation 200 may focus on extracting imagetext pairs 208 by identifying and selecting data that includes one or more keywords, phrases, orsentences relevant to specific ocular conditions. For example, keywords such as "retina", "optic nerve", "glaucoma", "macular degeneration", and "diabetic retinopathy", along with related medical terminology, may be used to filter and collect the relevant content. These keywords can appear in various forms, including single words, stems (e.g., "retinal," "ocular"), or specific sentence structures that describe the condition, symptoms, or diagnosis. Additionally, the text accompanying medical images — such as descriptions, captions, and alt-text — can be parsed so that the image is accurately associated with its corresponding diagnosis or condition. The automated web scraping 206 may send HTTP requests to web servers, receiving the HTML response or content, and parsing it for the relevant content. The automatic multimodal data curation 200 may enable the creation of large datasets for training machine-learning models, thereby improving ocular disease diagnosis, classification and / or decision making.
[0046] The multimodal data curation 200 for collecting image-text pairs 208 for targeted diseases may often yield diverse and potentially noisy data due to variations in webpage layouts, image quality, irrelevant images or textual content. In some examples, clustering 212 may be employed as a post-processing step to enhance the quality and reliability of the collected imagetext pairs 208. For clustering 212, the collected image-text pairs 208 may be processed using various embedding techniques to generate corresponding latent representations 210, interchangeably used herein with embedding vectors. Each image-text pair 208 may be embedded either individually (i.e., separate networks for image and text) or jointly, depending on the approach chosen. For individual embeddings, images can be processed using convolutional neural networks (CNNs) or transformer-based networks (e g., vison transformers) to extract visual features, while the accompanying text can be encoded using pretrained language models such as BERT or RoBERTa to generate textual embeddings. BERT employs a bidirectional attention mechanism, that enables consideration of both left and right context words simultaneously, encapsulating nuanced semantic relationships.
[0047] Alternatively, for joint embeddings, both the image and its associated text may be encoded together in a shared space using multimodal models such as UNITER (universal imagetext representation learning) or ViLT (vision and language transformer) that jointly processes both visual and textual features. For each image-text pair 208, these joint embedding techniques may create a single, unified latent representation 210 that characterizes the content in numerical form and captures the relationship between the image and text. Based on the generatedembeddings or latent representations 210 (whether individual or joint), the image-text pairs 208 may be further refined.
[0048] One or more unsupervised machine-learning techniques, including clustering techniques such as k-means, principal component analysis (PC A), DBSCAN, Gaussian mixture models, hierarchical clustering and the like, may be applied to the latent representations 210 to refine valid image-text pairs 208 from noisy or irrelevant data. Clustering 212 may involve grouping latent representations 210 based on similarity or the associated inherent patterns. However, the initial clustering results may still include some ambiguity or errors due to overlapping or misclassified data points. The quality of the clusters may be assessed in a cluster refinement 216 to select the clusters of interest from the grouped patterns. Cluster refinement 216 may be performed by human annotators (or domain experts) 214 after observing the clusters, for example, an expert in the field (ophthalmologist, clinician, researchers or equivalent) to verify the correctness of the collected image-text pairs 208 for the specified eye diseases.
[0049] The human annotators 214 can manually review the clusters based on the factors such as data quality, uniqueness, diversity, and relevance with the associated text. Other methods of cluster refinement 216 may include deploying one or more automated or semi-automated techniques, using validation metrics (e.g., silhouette score) or external validation metrics (e.g., adjusted random index). During cluster refinement 216, the clusters that are not correctly associated with image or text and / or that do not represent the targeted ocular diseases may be dropped, if found. Once the cluster refinement 216 is completed, the refined image-text pairs 218 associated with verified clusters may be extracted for creating a multimodal dataset comprising image-text pairs associated with the one or more image modalities.
[0050] In some aspects of the present disclosure, a knowledgebase— concept bank 322 may be created by curating a collection of relevant concepts, terms, diseases, and attributes associated with target disease e.g., ophthalmology, cardiology, or neurology. The purpose of the concept bank 315 is to serve as a structured knowledge repository for professionals in the targeted domain. The creation of the concept bank 315 may encompass a collaborative effort between natural language processing (NLP) and an expert in the field, such as an ophthalmologist, to automate and validate prior knowledge related to various eye diseases. Involving domain experts may facilitate the identification of concepts and terms related to ocular diseases, enabling theinclusion of clinically significant and accurate information. For guiding the search process, prompts (or queries) in natural language can be designed by specifying the information of interest. Well-designed prompts increase relevancy in search results by focusing on specific aspects of targeted diseases, such as particular diseases, imaging modalities, or anatomical structures. Prompts can be designed to cover a broad range of ocular diseases, thus avoiding the search to be limited to specific conditions.
[0051] In some examples, a prompt strategy (e.g., a series of prompt templates combined with chain-of-thought reasoning) may be employed to extract relevant information from large language models (LLMs) 320. Additionally, LLMs can understand and generate responses for diverse expressions, synonyms, and variations of terms to enable accurate data retrieval for each concept in textual data. Thus, enabling coverage during searches or analyses of targeted disease- related information. For example, synonyms of glaucoma are ocular hypertension, elevated intraocular pressure (IOP), silent thief of sight, with variations such as primary open-angle glaucoma (POAG), angle-closure glaucoma, normal-tension glaucoma (NTG). Similarly, red eye may also be known as conjunctival injection, bloodshot eye and astigmatism may be referred to as irregular cornea, cylindrical error or distorted vision. Concepts within the concept bank 322 can be organized hierarchically or interconnected to establish relationships, such as parent-child relationships representing broader categories and subcategories. This hierarchical organization may facilitate a structured and systematic approach to concept representation. The incorporation of established terms and structures contributes to the completeness and interoperability of concept bank 322. Leveraging existing knowledge resources, including medical textbooks, ontologies, and standardized terminologies like SNOMED CT and ICD, may further enrich the knowledge encapsulated within the concept bank 322.
[0052] LLMs 320 can also be used to generate responses to queries or prompts that cover a broad range of ocular diseases and imaging modalities. By employing well-designed prompts, the natural language understanding capabilities of LLMs can be harnessed to effectively build the concept bank 322 related to ocular diseases from open data repositories. The LLM 320 such as generative pre-trained transformers (GPT), text-to-text transfer transformers (T5), and XLNet are primarily designed for natural language understanding and generation. By utilizing an iterative prompting 302, the concept bank 322 may be built that includes highly specific, accurate terminologies related to ocular diseases. The iterative nature of the prompting processmay further enable continuous refinement and improvement based on the obtained responses. The iterative prompting technique 302 may involve crafting modality -guided and / or disease- guided prompts.
[0053] Modality-guided prompts 306 may help in tailoring the search to particular imaging modalities, where the goal is to extract general concepts related for specific imaging modalities (e g., fundus photography, OCT, fluorescein angiography (FA), etc.). Modality-guided prompts 306 can be crafted to generate concepts corresponding to dominant features or attributes that appear through a specific imaging modality. As a non-limiting example, a modality -guided prompt may be crafted as, "Describe the common features visible in a fundus image." For the given modality-guided prompt, LLM response may include, "Common features comprise optic disc, macula, retinal vessels, and the fovea." The extracted concepts from the query response may comprise “optic disc”, “macula”, “retinal vessels”, and “fovea”. The LLM may be queried iteratively, and the generated concepts from each prompt may then be integrated into the unified concept bank 322, forming the general ocular features that can be mapped to various ocular disease categories as a many-to-one integration 314.
[0054] For example, the iterative query 312 (i.e., second prompt) for refining concepts based on the previous response may be crafted as, “What are different types of retinal vessels visible in a fundus image?” Similarly, the third prompt may be crafted as, “How these types differ from each other?” The LLM response for the second and third prompt may include, “In a fundus image, retinal vessels can be classified into arteries and veins. Arteries appear smaller and lighter in color, whereas veins are larger and darker. The artery -vein ratio is a critical indicator of vascular health. Additionally, arterioles, which are smaller branches of arteries, and venules, smaller branches of veins, can also be seen.” For LLM’s response, the extracted concepts may include, “arteries (smaller, lighter color)”, veins (larger, darker color)”, “Arterioles (smaller branches of arteries), venules (smaller branches of veins), and “artery -vein ratio”.
[0055] Furthermore, the LLM may be prompted with disease-guided prompts 304 to gather disease-related concepts. As a non-limiting example, a disease-guided prompt including an ocular condition may be crafted as, “What are the key features seen in the OCT of a patient with diabetic retinopathy?" In response, LLM may output, "The OCT of diabetic retinopathy often shows retinal thickening, microaneurysms, and intraretinal fluid accumulation." Similarly, initerative prompting technique 302, concepts may be extracted from the generated responses that are specific to individual eye diseases or imaging modality. Different ocular diseases may share common features such as retinal vessel changes, but each ocular disease might also have unique characteristics such as specific patterns of retinal degeneration, vascular leakage, or changes in the optic nerve. This may be achieved by prompting the LLM 320 for common and diseasespecific and / or modality-specific terminologies, enabling the identification of both shared concepts 308 and unique concepts 310 associated with the imaging modality and / or diseases.
[0056] For identifying shared or unique concepts, a follow-up disease-guided prompt after extracting general features for diabetic retinopathy may comprise, “What are the common OCT features seen in diabetic retinopathy and age-related macular degeneration (AMD)?" The LLM may identify shared features (or concepts) 308, such as macular edema and retinal thickening. Similarly, for distinguishing between conditions, disease-guided prompts 304 can help identify unique signs; for instance, asking, “What unique signs are visible in the FA of a patient with AMD?" This iterative process may assist in extraction of both general and specific features, progressively refining the knowledgebase (or concept bank 322). By leveraging these strategies — modality-guided prompts 306, disease-guided prompts 304, iterative querying 312, and identifying shared concepts 308 and unique concepts 310, each ocular disease may be characterized enabling curation of the concept bank 322 for improved diagnostic accuracy and decision-making.
[0057] Concept extraction 321 from the generated responses of the LLM, may involve identifying and extracting relevant terminologies. The process may begin with parsing the response and removing common stop words (e.g., 'a', 'the', 'of), which helps focus on meaningful content. The domain-specific terms may be then identified using a combination of natural language processing (NLP) techniques. For example, named entity recognition (NER) may help identifying proper nouns and specialized terms, such as medical conditions (e.g., 'glaucoma') or technical jargon (e.g., 'retinal imaging'). Part-of-speech tagging may further refine the extraction by categorizing words as nouns or key phrases, aiding in the identification of important terminologies within the sentence structure. To capture the syntactic relationships between words and phrases, dependency parsing may be applied, identifying multi-word expressions or compound terms, such as 'visual acuity' or 'ocular pressure.' Additionally, TF-IDF (term frequency-inverse document frequency) can highlight terms that appear frequently in the LLM’sresponses but are rare across a broader corpus, helping identify relevant and unique terms. Finally, word embeddings from models such as BERT or Word2Vec can be used to group semantically similar terms, such as 'eye' and 'ocular,' enabling a deeper understanding of related terminologies. Together, these techniques may provide a robust framework for extracting key concepts from the generated responses, so that the extracted terms are both relevant and specific to the domain of ocular diseases.
[0058] Once, a broad set of concepts has been generated from the modality-guided 306 and disease-guided prompts 304, concept comparison prompts 316 may be crafted in iterative prompting technique 302. The goal may be to compare the concepts (such as features or characteristics) extracted from different prompts (e.g., features of diabetic retinopathy vs. features of AMD) and assess how similar or distinct those concepts are. For example, the concepts extracted for both diseases (e.g., retinal thickening, microaneurysms, and intraretinal fluid for diabetic retinopathy, and macular edema, drusen, and geographic atrophy for AMD), may be compared for similarities and differences. This comparison may be based on semantic similarity (i.e., how similar the meanings of these terms are) or contextual similarity or overlap (i.e., how often these terms appear together or in similar contexts).
[0059] Once concepts from different diseases are compared, similarity scores 326 may be assigned to quantify the level of overlap. The similarity scores 326 may be computed based on concept comparison techniques 324 that may account for semantic or contextual similarity, leveraging pre-trained models such as word embeddings (e.g., Word2Vec, GloVe, or BERT) to compute how similar two words or sets of words are based on their context in a language. These embeddings map each word (or concept) into a high-dimensional vector space, where similar words or concepts are closer together. For example: "retinal thickening" vs. "macular edema" may have a moderate similarity as both refer to retinal changes related to fluid retention. Conversely, "microaneurysms" vs. "drusen" may have a lower similarity since both represent different types of retinal abnormalities. These similarities can be quantified using a similarity score 326 ranging from 0 to 1 or -1 to 1, such as cosine similarity that measures the cosine of the angle between two vectors in the high-dimensional space. The closer the cosine value is to 1, the more similar the terms are. Alternatively, domain experts 214 may assign similarity scores 326 manually based on their knowledge, or a rule-based system may be employed with predefinedmappings e.g., "macular edema" and "retinal thickening" might be assigned a similarity score of 0.7 because both involve swelling of the retina due to fluid buildup.
[0060] After comparing each concept, a final similarity score can be generated for each pair of disease features. These scores quantify how closely related the concepts are. For example, the similarity score for "Retinal thickening" (Diabetic Retinopathy) vs. "Macular edema" (AMD) may be assigned as 0.8, and "Microaneurysms" (Diabetic Retinopathy) vs. "Drusen" (AMD) may be 0.3. The similarity scores 326 may help identify shared concepts 308 (e.g., having high similarity) and unique concepts 310 (e.g., having low similarity), making it easier to understand how the diseases overlap and where they differ. In some examples, the terminologies or concepts from the concept bank 322 may be integrated into iterative prompting 302 to guide the data crawling process effectively.
[0061] FIG. 4 shows an example illustration of a network 400, generating latent representations (embedding vectors) from different imaging modalities in accordance with some aspects of the present disclosure. The network 400 may preprocess n clusters of refined imagetext paired data 218 associated with n image modalities and feed into a corresponding machinelearning model 402a-n to generate modality-specific latent vectors. The preprocessing may include e.g. scaling, resizing and / or normalizing image data of the image-text pairs. The machine-learning models 402a-n may be trained to generate meaningful and context-rich representations. Similar to the process of generating latent representations 210, these machinelearning models 402a-n (or encoders) may generate embedding vectors corresponding to each image-text pair either individually (i.e., separate encoders for image and text) or jointly, depending on the approach chosen.
[0062] For individual embedding vectors, images can be processed using pre-trained convolutional neural networks (CNNs) (e.g., ResNet, VGG, and Inception) or transformer-based networks (e.g., vision transformers) to extract visual features, while the accompanying text can be encoded using pretrained language models such as BERT or RoBERTa to generate textual embeddings. These pre-trained networks are typically trained on large-scale datasets and can be fine-tuned or used as feature generators for specific downstream tasks. To fine-tune or apply transfer learning to pre-trained networks, the weights of the model can be initialized with those from a pretrained architecture that has been trained on a large dataset. Subsequently, the modelcan be further trained on domain-specific data by adjusting the weights of the later layers while potentially freezing earlier layers to retain general feature extraction and generation capabilities.
[0063] Alternatively, for joint embedding vectors, both the image and its associated text may be encoded together in a shared space using multimodal models such as UNITER, contrastive learning or vision and language transformers such as ViLT or ViLBERT that jointly processes both visual and textual features. In some examples, for generating latent representations for each image-text pair modality, custom -trained networks (e.g., neural networks, encoder-decoder networks, or any other machine-learning models capable of generating encoded representations for subsequent processing) may also be used. It may be understood that the machine-learning models 402a-n to generate latent representations from illustrative examples of FIG. 2 and FIG. 4 may either be different or the same. Additionally, it may be appreciated that the specific architecture and training methodology for each machine-learning model 402a-n may be similar or different depending on the characteristics of the imaging modalities and how these interact with textual data to extract meaningful and context-aware representations.
[0064] Within the domain of ocular diseases, the concept bank 322 may encompass concepts related to the targeted eyes diseases specific eye conditions (e.g., myopia, conjunctivitis, cataract, diabetic retinopathy), imaging modalities (fundus, OCT, FA, ), anatomical structures (e.g., macula, retina, cornea, iris), disease names (e.g., amblyopia, ocular hypertension), and associated symptoms (e.g., Fuch’s endothelial dystrophy presenting symptoms such as corneal swelling, blurry vision, and discomfort), along with other pertinent attributes of the targeted diseases.
[0065] In some examples, the concepts from concept bank 322 may be used to tag input dataset comprising multimodal data (i.e. image-text pairs) from one or more imaging modalities. For example, in one training sample, each image-text pair of a set of image-text pairs 218a-n may be tagged with one or more concepts and an associated ocular disease of a set of ocular diseases. This input dataset may be fed into an attention-based multimodal graph neural network (AM-GNN) 406 to construct a multimodal ocular disease graph 502. Alternatively, a concept graph 410 may be pre-constructed, where each ocular disease and their associated concepts are represented by nodes. The concept graph 410 may be constructed by graph encoding techniques 408 (e.g., graph attention networks (GAT), graph neural networks (GNNs), k-nearest neighbors (KNN), or community detection algorithms) to structure the nodes (concepts, diseases) and edges(relationships) into a graph. The concept graph 410 along with the training dataset may be fed into AM-GNN 406 to construct multimodal ocular disease graph 502.
[0066] FIG. 5 illustrates an example of an attention-based multimodal graph neural network (AM-GNN) 406, where inputs (i.e., 404a-n and / or 410) and output (i.e., 502) are represented as graphs with nodes and edges. As described previously, each input to the AM-GNN 406 may include a set of image-text pairs 404a-n and associated concepts during training, all of which are also represented as individual nodes. Each node in the graph has an embedding vector (or a latent representation). Relationships between different nodes may be represented by edges (or weighted edges) that may be stored in a data structure e.g., an adjacency matrix or a linked list. Relationships may be defined based on similarity scores 326 indicating semantic or contextual associations modality-specific and / or disease-specific characteristics. Bi-directional edges may be established between image-text pairs nodes 404a-n and corresponding concept nodes that are further linked with disease nodes 504a-m (e.g., as illustrated in 410).
[0067] In some instances, edge features may also be included in the node, representing additional information about the relationships between nodes. For example, edge features may capture type of relationship between two or more nodes such as synonym, antonym or part-of (e.g., cornea is a part-of eye) and / or confidence score. In other instances, these relationships (edges) may be established dynamically during the training phase (e.g., when the input is nongraph or unstructured). For example, AM-GNN 406 may learn to create edges within input image-text pair nodes (Ii, I2, .. .In) 404a-n based on various concepts e.g., disease specification, modality relationships or other parameters (e.g., image similarities, or linking a Fundus image with an OCT image associated with the same subject and representing same disease) during training. Additionally, AM-GNN 406 may also learn to create edges between image-text pair nodes 404a-n and concept nodes 506a-p that act as intermediary nodes between image-text pair nodes 404a-n and disease nodes 504a-m. The edges may be determined dynamically using attention mechanism for constructing graph.
[0068] The AM-GNN 406 may learn to handle image-text pair nodes 404a-n, disease nodes 504a-m, and concept nodes 506a-p. Image-text pair nodes (e.g., 404a-n) may serve as the primary input, encapsulating visual features from ocular images and textual descriptions associated with the ocular images (or retinal scans). The final output nodes (i.e., 504a-m) mayrepresent the predicted diseases based on associated concept(s) that contribute to the prediction. These nodes are part of the final output layer of the network. Each disease node is mutually exclusive, and the goal is to predict the presence or absence of these diseases based on the multimodal input and associated concepts. Concept nodes 506a-p may capture higher-level semantic information related to diseases. These may include anatomical concepts, physiological attributes, or specific clinical features associated with diseases. Each concept may be mapped to one or more image-text pairs 404a-n, for example, concept Ciu 506a may be characterized by image-text pair node 404a (Ii) and 404b (b) and the respective edges (e.g., 508a and 508b) may signify the extent of contribution that each image-text pair may provide to comprehend the concept 506a (Ciu).
[0069] Similarly, concepts, for example, Cis, C2S and Cqs may be shared concepts within one or more targeted eye diseases, forming a community 510 in the multimodal graph 502. Here, the subscripts for concept nodes include the first character representing the associated disease, while the second character indicates whether the concept is unique to a specific disease or shared among various targeted diseases. There exist edges between different concepts within the same community 510, where the weight of the edge may represent the similarity between different concepts. However, there may not be any edges between the different diseases (Li) as these are mutually exclusive. The weight of the edge between concepts and image-text pair may correspond to how important the concept is for that particular eye disease. Thus, targeted diseases may be predicted as a result of weighted sum of the different concepts.
[0070] FIG. 6 illustrates an example of an attention mechanism for the attention-based multimodal graph neural network (AM-GNN) 406. FIG. 6 shows an example of three nodes that may be related, where the nodes may represent image-text pairs, concepts or diseases, as shown in an example tree 602. For processing two nodes i and j, the attention mechanism of AM-GNN may focus on how attention coefficients may be computed and used. Each node in the example graph 602 has an associated embedded vector e.g.,and hj, which may be first linearly transformed using a learnable weight matrix W to produce Hi — Whtand Hj — Whj, new representations in a higher-dimensional space. For two nodes, i and j, their transformed embedded vectors H and Hj may be concatenated ([Hi ||fl / ]), and a learnable attention weightvector wamay be applied to compute a raw attention score, e^, through a compatibility function. This score captures the relationship between the two nodes based on their features.
[0071] The raw score etj 604 may be passed through a non-linear activation function to introduce non-linearity, such as ReLU (rectified linear unit) or Leaky ReLU 606, and then normalized using a SoftMax function 608 or a sigmoid function. This normalization causes attention coefficientsto sum to 1 over all neighbors j of node i. The resulting attention coefficient aLj quantifies how much importance should be given to the information from one node or edge when updating the representation of another the importance e.g., node j in updating the representation of node i. Finally, during message aggregation, these coefficientsare used as weights to combine the latent representations of neighboring nodes, producing the updated feature for node i. Alternatively, attention scores may be computed by the compatibility function (e.g., dot product, cosine similarity, or a learned linear layer). For example, the raw score604 may directly be computed by the dot product of 6) and Q, representing their similarity in the transformed feature space.
[0072] AM-GNN iteratively updates node representations, repeating the above-mentioned process for all nodes, thereby dynamically learning how information flows across the graph structure. During message passing, information may be exchanged between nodes, influenced by attention weights computed through the attention mechanism 610. Th attention mechanism 610 may enable the AM-GNN model 406 to focus on specific nodes capturing the nuanced relationships within and across modalities. Attention parameters (weights and biases) or attention scores for each node or edge may be learned during training. In some instances, attention mechanisms use multiple heads, each with their own set of attention parameters. By having multiple heads, the model can attend to various types of relationships (e.g., spatial, temporal, contextual) in the data that may not be captured by a single attention mechanism. The outputs from different heads may be then concatenated or averaged to provide a more robust representation.
[0073] The multimodal graph 502 generated by AM-GNN model 406 may be used to derive outputs 611 including predicted disease(s) 612 and associated concepts 614. Once the AM-GNN 406 is trained, the model can be applied to various downstream tasks related to disease prediction or classification (e.g., type of a disease), presence or absence of a specific disease, prognosisprediction (predicting likelihood of a disease from a confidence score), link prediction (e.g., predicting missing or potential connection between nodes), disease association analysis or construct a knowledge graph that may represent the relationships between diseases, symptoms, treatments, or other relevant entities. This knowledge graph can serve as a structured representation of information learned by the AM-GNN model 406. In other words, the AM-GNN 406 may integrate information from image-text pairs and concept nodes to make predictions that are interpretable and clinically relevant.
[0074] FIG. 6 illustrates exemplary outputs that may be generated from the multimodal graph 502 given a set of input multimodal data. In some examples, output may further comprise one or more concepts present in the image-text pair and corresponding confidence scores, which indicate the degree to which each concept is relevant to the input. These concepts may be derived from both the image content and the associated textual information, capturing concepts from the concept bank (e.g., anatomical terms, clinical signs, and disease-specific indicators).Additionally, the output may include a predicted disease of one or more diseases, accompanied by a confidence score, indicating a likelihood of the disease being present based on the input. Higher confidence scores indicate a stronger belief in the presence of certain features or the likelihood of a specific disease. The confidence scores associated with the predicted disease 612a may also indicate the confidence of the model in its diagnostic prediction. In some aspects of the present disclosure, AM-GNN 406 generates multiple outputs corresponding to one or more image-text pairs input.
[0075] The confidence scores associated with each concept 614a may indicate the certainty of the model regarding the presence of that particular feature or concept. This output may also reflect the absence or presence of certain diseases, depending on how the model is trained to classify negative or non-disease cases. For example, a likelihood (or a probability value) less than a certain threshold may represent absence of the particular ocular disease. In addition to disease classification and / or diagnosis, the AM-GNN model can also generalize across multiple disease categories, scoring their likelihood or presence based on the distribution of relevant concepts, effectively capturing the full spectrum of disease possibilities or absence within the given input. During inference, the input (one or more image-text pairs) may be treated as a set of new nodes. The set of input nodes may be dynamically connected to the relevant nodes of the concept graph 410 using attention mechanism 610. The AM-GNN 406 may generate anassociation of the set of input nodes to concepts in the multimodal graph 502 and output the detected concepts and diagnosed diseases based on the identified associations.
[0076] FIG. 7 illustrates an exemplary workflow 700 to predict ocular diseases using one or more image-text pairs including retinal scan in accordance with some aspects of the present disclosure. The blocks in workflow 700 are illustrated in a specific order, while the order can be modified, for example, some blocks may be performed before others, and some blocks may be performed simultaneously. The blocks can be performed by hardware, software, or a combination thereof. For detecting a disease from a set of diseases, a process at block 702 may include receiving multimodal data including the one or more image-text pairs 218a-n from one or more imaging modalities (e.g., OCT, Fundus, ultrasound) by a computing system. An image of the one or more image-text pairs 218a-n depicts a retinal scan of a subject. At block 704, one or more machine-learning models 402a-n may generate a latent representation, a numeric vector, for each image-text pair. The one or more machine-learning models 402a-n may process the image and its associated text either separately (e.g., using separate encoders such as vision transformers for image embedding and BERT for text embedding), or jointly (e.g., using a model similar to ViLBERT that integrates both modalities in a multimodal space). Additionally, the one or more machine-learning models 402a-n may be trained separately for each imaging modality to generate corresponding latent representations 404a-n. In some examples, the same machinelearning model may process image-text pairs 218a-n from multiple imaging modalities to generate respective latent representations.
[0077] At block 706, a concept bank 322 may be accessed that includes concepts for each disease of the set of diseases. The concept bank 322 may be curated by prompting a large language model (LLM) 320 with disease-guided prompts 304, modality-guided prompts 306 and comparison prompts 316. LLM 320 may be iteratively queried, thereby generating responses including concepts that represent terminologies and features that are unique to each imaging modality (or each disease of the set of diseases) as well as shared concepts 308 spanning multiple modalities (or the set of diseases). These concepts may also include attributes, patterns, symptoms, synonyms, ontologies, and features associated with particular diseases. This iterative prompting may assist in collecting diverse and accurate concepts in the context of disease diagnosis.
[0078] The concepts from the generated responses may be extracted by applying one or more natural language processing (NLP) techniques, such as named entity recognition (NER), part-of- speech tagging, and dependency parsing, to identify and categorize relevant terminologies and features (indicating concepts). These concepts may be then organized into a structured format, enabling efficient integration into the concept bank for use in training subsequent models. Additionally, the extracted concepts may undergo further refinement and validation through domain-specific filtering e.g., by a domain expert 214 to enable their relevance and accuracy in relation to the target ocular diseases and imaging modalities.
[0079] The one or more latent representations 404a-n associated with one or more image-text pairs 218a-n may be fed to an attention-based multimodal graph neural network (GNN) 406 that generates a multimodal graph 502 based on a set of relevant concepts, at block 708. The AM- GNN 406 may be configured to generate associations between each image-text pair and the concepts associated with each disease label using an attention mechanism 610. At block 710, an output 611 may be generated that includes a predicted disease label 612 based on the set of relevant concepts that are determined in the multimodal graph 502.
[0080] In some examples, the multimodal graph may output a confidence score for the predicted disease label, or for each of the potential disease labels. The confidence score may represent the model's degree of certainty in its prediction, quantifying a likelihood that the predicted disease label accurately corresponds to the input data based on the learned relationships between the visual and textual features. Higher confidence scores indicate stronger associations between the image-text pair and the predicted disease label, while lower scores suggest a higher degree of uncertainty, potentially requiring further validation or additional analysis. Additionally, the output may include the set of concepts present in the input multimodal data along with respective confidence scores that may be derived from the multimodal graph.
[0081] FIG. 8 is an example illustration of a computing system 800, in which various embodiments of the present disclosure may be implemented. The functionality described herein can be performed, at least in part or a combination of one or more hardware or software logic components. For example, the techniques described above for predicting ocular disease for a subject using one or more image-text pairs 218a-n depicting retinal scans by leveraging a combination of one or more machine-learning models 402a-n and AM-GNN 406 can beimplemented in computer-executable instructions. The instructions can be executed by processing unit 812 that may be a combination of an arithmetic logic unit 814 that performs arithmetic and logical operations, and a control unit 816 that may help in execution of the instructions. The control unit 816 may direct and coordinate the operation of the processor with other parts of the computer by synchronizing data flow between different components of the processing unit 812. It may manage the flow of instructions and data between various components. The control unit 816 may decode an operation code (opcode) and may convert them into control signals to coordinate how data moves within the processing unit 812. The control unit 816 may regulate execution units such as the arithmetic logic unit 814 and the flow of data to primary storage 804 and secondary storage 806.
[0082] To provide additional context for various aspects thereof, FIG. 8 and the following description are intended to provide a brief, general description of the computing system in which the various aspects can be implemented. While the description above is in the general context of computer-executable instructions that can run on one or more computing systems, those skilled in the art will recognize that a novel implementation also can be realized in combination with other program modules and / or as a combination of hardware and software. The computing system for implementing various aspects includes a processing unit 812 having one or more processors (also referred to as microprocessors), a computer-readable storage medium (where the medium is any physical device or material on which data can be electronically and / or optically stored and retrieved) such as a data storage unit 802 (computer readable storage medium / media also include magnetic disks, optical disks, solid state drives, external memory systems, and flash memory drives), and a system bus. The data storage unit 802 may have a primary storage 804 and a secondary storage 806 as described here in. The primary storage 804 and the secondary storage 806 may differ in speed of access, connection with the computer’s processor and data retrieval speeds. Primary storage 804 may often be directly connected to the computer's processor, boasts rapid data retrieval speeds. In contrast, secondary storage 806 may be designed for long-term storage and may have slower access times.
[0083] The computing system 800 can include various microprocessors, such as singleprocessor, multi-processor, single-core, and multi-core units for processing and storage. Additionally, experts in the field recognize that the innovative system and methods can be applied to other computing configurations, including minicomputers, mainframe computers,personal computers (such as desktops, laptops, and tablet PCs), handheld computing devices, microprocessor-based consumer electronics, and similar systems. These systems can be interconnected with one or more associated devices.
[0084] In some aspects, the computing system 800 can include one of several computers employed in a datacenter and / or computing resources (hardware and / or software) in support of cloud computing services for portable and / or mobile computing systems such as wireless communications devices, cellular telephones, and other mobile-capable devices. Cloud computing services include, but are not limited to, infrastructure as a service (laaS), platform as a service, software as a service (SaaS), storage as a service (StaaS), data as a service (DaaS), security as a service and APIs (application program interfaces) as a service. In some instances, data storage unit 802 can include computer-readable storage (physical storage) medium such as a volatile memory (e.g. random-access memory (RAM) also termed as the primary storage 804) and a non-volatile memory (e.g., (ROM)). A basic input / output system (BIOS) can be stored in the non-volatile memory and includes the basic routines that facilitate the communication of data and signals between components within the computing system, such as during startup. The volatile memory also includes a high-speed RAM such as static RAM for caching data.
[0085] As an illustrative example (without limiting the scope), the data storage unit 802 may include programming modules. These modules can encompass client applications, web browsers, mid-tier applications, relational database management systems (RDBMS), and more. Additionally, the storage unit 802 holds program data and an operating system. The operating system can include (for example) Microsoft Windows®, Apple Macintosh®, or Linux. Furthermore, commercially available UNIX®-like operating systems (such as GNU / Linux variants and Google Chrome OS) and mobile operating systems (like iOS, Windows® Phone, Android OS, BlackBerry® OS, and Palm® OS) are part of this landscape. Notably, portions of the operating system, program modules, and program data can be cached in storage unit 802 — both volatile memory (e.g., RAM) and non-volatile memory (e.g., ROM). This flexibility allows the disclosed architecture to be implemented using a variety of commercially available operating systems or combinations thereof (including virtual machines).
[0086] In some other examples, the computing system 800 may have additional features or functionality. For example, the computing system 800 may also include additional data storagedevices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or tape. Computer-readable media may include, at least, two types of computer-readable media, namely computer storage media and communication media. Computer storage media may include volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storage of information, such as computer readable instructions, data structures, program modules, or other data.
[0087] The storage media of the computing system 800 may also include removable storage, and non-removable storage. EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store the targeted information and which computing system 800 can access are examples of computer storage media in addition to RAM and ROM. Additionally, the computer-readable media might have computer-executable instructions that the processing unit 812 can use to carry out the different tasks and / or operations mentioned in this article. In contrast, communication media may embody computer readable instructions, data structures, programming modules, or other data in a modulated data signal, such as a carrier wave, or other transmission mechanism.
[0088] One or more input devices, such as a keyboard, mouse, pen, voice input device, touch input device, etc., may also be included in the computing system 800. There might also be one or more output devices 810, including speakers, printers, displays, and so on. These devices are not covered in detail here because they are well known in the field. To establish communication, the computing system 800 may further have one or more network interfaces. This would enable the computing system 800 to communicate with other systems or devices, for example, over a network. Both wired and wireless networks could be a part of these networks. Here, the computing system 800 is one example of a suitable device or system and is not intended to suggest any limitation as to the scope of use or functionality of the various embodiments described.
[0089] Other well-known computer environments, configurations, and / or systems that may be appropriate for use with the embodiments include, but are not limited to, network PCs, mainframe computers, programmable consumer electronics, set top boxes, game consoles,programmable consumer electronics, personal computers, server computers, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, and / or the like. For instance, part or all the computing system 800 components could be put into use in a cloud computing environment, where resources and / or services are made available for user devices to consume on a selected basis via a computer network.
[0090] Further, while certain aspects have been described using a particular combination of hardware and software, it should be recognized that other combinations of hardware and software are also possible. Certain aspects may be implemented only in hardware, or only in software, or using combinations thereof. The various processes described herein can be implemented on the same processor or different processors in any combination.
[0091] Where devices, systems, components or modules are described as being configured to perform certain operations or functions, such configuration can be accomplished, for example, by designing electronic circuits to perform the operation, by programming programmable electronic circuits (such as microprocessors) to perform the operation, such as by executing computer instructions or code, or processors or cores programmed to execute code or instructions stored on a non-transitory memory medium, or any combination thereof. Processes can communicate using a variety of techniques including but not limited to conventional techniques for inter-process communications, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0092] Some embodiments of the present disclosure include a system including one or more data processors. In some embodiments, the system includes a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein. Some embodiments of the present disclosure include a computer-program product tangibly embodied in a non-transitory machine- readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods and / or part or all of one or more processes disclosed herein.
[0093] The terms and expressions which have been employed are used as terms of description and not of limitation, and there is no intention in the use of such terms andexpressions of excluding any equivalents of the features shown and described or portions thereof, but it is recognized that various modifications are possible within the scope of the invention claimed. Thus, it should be understood that although the present invention as claimed has been specifically disclosed by embodiments and optional features, modification and variation of the concepts herein disclosed may be resorted to by those skilled in the art, and that such modifications and variations are considered to be within the scope of this invention as defined by the appended claims.
[0094] The present description provides preferred exemplary embodiments only, and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the present description of the preferred exemplary embodiments will provide those skilled in the art with an enabling description for implementing various embodiments. It is understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope as set forth in the appended claims.
[0095] Specific details are given in the present description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments may be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: receiving one or more image-text pairs from one or more imaging modalities, wherein an image of the one or more image-text pairs depicts a retinal scan of a subject; generating one or more latent representations associated with one or more imagetext pairs by applying one or more machine-learning models; accessing a concept bank that is generated by prompting a large language model (LLM), wherein the concept bank includes concepts that contribute to a prediction of a disease label of a set of disease labels, wherein the concepts include terminologies that are related to the one or more imaging modalities and each of the set of disease labels; generating a multimodal graph based on the one or more latent representations by applying an attention-based multimodal graph neural network (AM-GNN) that is configured to predict a set of relevant concepts by identifying associations between a set of nodes including the one or more image-text pairs, the set of disease labels and the concepts using an attention mechanism, wherein the attention-based multimodal graph neural network is trained on a multimodal dataset collected via a data curation technique; and generating an output that includes a predicted disease label based on the set of relevant concepts using the multimodal graph.
2. The computer-implemented method of claim 1, wherein prompting the large language model further including: generating responses by iteratively querying the large language model with modality-guided prompts, disease-guided prompts and comparison prompts in a natural language format; and extracting, from the generated responses, concepts associated with the one or more imaging modalities that contribute to a prediction of each of the set of disease labels based on one or more natural language processing techniques.
3. The computer-implemented method of claim 1, wherein the data curation technique includes: retrieving content by navigating through one or more data sources using universal resource locators (URLs); and extracting image-text pairs associated with the one or more imaging modalities from the retrieved content by identifying the concepts related to the disease label of the set of disease labels.
4. The computer-implemented method of claim 3, further including: generating latent representations associated with the extracted image-text pairs by leveraging the one or more machine-learning models; clustering the latent representations based on associated patterns; performing cluster refinement, for each cluster of a set of clusters, by verifying correctness of the image-text pairs associated with the latent representations; and including the verified image-text pairs in the multimodal dataset to train the attention-based multimodal graph neural network.
5. The computer-implemented method of claim 1, wherein the associations are identified iteratively between each pair of nodes of the set of nodes based on the attention mechanism that includes: determining an attention score by applying a compatibility function that computes relevance between latent representations associated with each pair of nodes of the set of nodes.
6. The computer-implemented method of claim 1, wherein the set of disease labels identifies one or more ocular diseases.
7. The computer-implemented method of claim 1, wherein the one or more imaging modalities include: Fundus, optical coherence tomography (OCT), fluorescence angiography (FA), function magnetic resonance imaging (fMRI), X-ray, positron emission tomography (PET), and ultrasound imaging.
8. A system comprising: one or more data processors; and a non-transitory computer readable storage medium containing instruction which, when executed on the one or more data processors, cause the one or more data processors to perform a set of operations including: receiving one or more image-text pairs from one or more imaging modalities, wherein an image of the one or more image-text pairs depicts a retinal scan of a subject; generating one or more latent representations associated with one or more image-text pairs by applying one or more machine-learning models; accessing a concept bank that is generated by prompting a large language model (LLM), wherein the concept bank includes concepts that contribute to prediction of a disease label of a set of disease labels, wherein the concepts include terminologies that are related to the one or more imaging modalities and each of the set of disease labels; generating a multimodal graph based on the one or more latent representations by applying an attention-based multimodal graph neural network (GNN) that is configured to predict a set of relevant concepts by identifying associations between a set of nodes including the one or more image-text pairs, the set of disease labels and the concepts using an attention mechanism, wherein the attention-based multimodal graph neural network is trained on a multimodal dataset collected via a data curation technique; and generating an output that includes a predicted disease label based on the set of relevant concepts using the multimodal graph.
9. The system of claim 8, wherein prompting the large language model further including: generating responses by iteratively querying the large language model with modality-guided prompts, disease-guided prompts and comparison prompts in a natural language format; andextracting, from the generated responses, the concepts associated with the one or more imaging modalities that contribute to a prediction of each of the set of disease labels based on one or more natural language processing techniques.
10. The system of claim 8, wherein the data curation technique including: retrieving content by navigating through one or more data sources using universal resource locators (URLs); and extracting image-text pairs associated with the one or more imaging modalities from the retrieved content by identifying the set of concepts related to the disease label of the set of disease labels.
11. The system of claim 10, wherein the set of operations further including: generating latent representations associated with the extracted image-text pairs by leveraging the one or more machine-learning models; clustering the latent representations based on associated patterns; performing cluster refinement, for each cluster of a set of clusters, by verifying correctness of the image-text pairs associated with the latent representations; and including the verified image-text pairs in the multimodal dataset to train the attention-based multimodal graph neural network.
12. The system of claim 8, wherein the associations are identified iteratively between each pair of nodes of the set of nodes based on the attention mechanism that includes: determining an attention score by applying a compatibility function that computes relevance between latent representations associated with each pair of nodes of the set of nodes.
13. The system of claim 8, wherein the set of disease labels identifies one or more ocular diseases.
14. The system of claim 8, wherein the one or more imaging modalities include: Fundus, optical coherence tomography (OCT), fluorescence angiography (FA), functionmagnetic resonance imaging (fMRI), X-ray, positron emission tomography (PET), and ultrasound imaging.
15. A computer-program product tangibly embodied in a non-transitory machine readable storage medium, including instructions configured to cause one or more data processors to perform to perform a set of operations comprising: receiving one or more image-text pairs from one or more imaging modalities, wherein an image of the one or more image-text pairs depicts a retinal scan of a subject; generating one or more latent representations associated with one or more imagetext pairs by applying one or more machine-learning models; accessing a concept bank that is generated by prompting a large language model (LLM), wherein the concept bank includes concepts that contribute to prediction of a disease label of a set of disease labels, wherein the concepts include terminologies that are related to the one or more imaging modalities and each of the set of disease labels; generating a multimodal graph based on the one or more latent representations by applying an attention-based multimodal graph neural network (GNN) that is configured to predict a set of relevant concepts by identifying associations between a set of nodes including the one or more image-text pairs, the set of disease labels and the concepts using an attention mechanism, wherein the attention-based graph neural network is trained on a multimodal dataset collected via a data curation technique; and generating an output that includes a predicted disease label based on the set of relevant concepts using the multimodal graph.
16. The computer-program product of claim 15, wherein prompting the large language model further including: generating responses by iteratively querying the large language model with modality-guided prompts, disease-guided prompts and comparison prompts in a natural language format; and extracting, from the generated responses, the concepts associated with the one or more imaging modalities that contribute to a prediction of each of the set of disease labels based on one or more natural language processing techniques.
17. The computer-program product of claim 15, wherein the data curation technique includes: retrieving content by navigating through one or more data sources using universal resource locators (URLs); and extracting image-text pairs associated with the one or more imaging modalities from the retrieved content by identifying the set of concepts related to the disease label of the set of disease labels.
18. The computer-program product of claim 17, further including: generating latent representations associated with the extracted image-text pairs by leveraging the one or more machine-learning models; clustering the latent representations based on associated patterns; performing cluster refinement, for each cluster of a set of clusters, by verifying correctness of the image-text pairs associated with the latent representations; and including the verified image-text pairs in the multimodal dataset to train the attention-based multimodal graph neural network.
19. The computer-program product of claim 15, wherein the associations are identified iteratively between each pair of nodes of the set of nodes based on the attention mechanism that includes: determining an attention score by applying a compatibility function that computes relevance between latent representations associated with each pair of nodes of the set of nodes.
20. The computer-program product of claim 15, wherein the set of disease labels identifies one or more ocular diseases.
Citation Information
Cited By
Method for intelligently describing liver space-occupying lesion ultrasonic image content by using LLM
CN120543547A
A method for intelligently describing the ultrasound image content of liver space-occupying lesions using LLM
CN120543547B
Eye image recognition system, method and device and storage medium
CN121527832A
Full-slice image analysis method based on spatial constraint attention and context awareness
CN121724955A