Systems and methods for generating synthetic medical images
A unified medical image-text generative model generates high-quality synthetic images using clinician feedback, addressing data scarcity and improving AI performance in healthcare applications by enhancing existing datasets and predictive capabilities.
Patent Information
- Application Number
- PCT/US2025/050496
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-11
- Filing Date
- 2025-10-10
- Publication Date
- 2026-04-16
AI Technical Summary
The scarcity of high-quality medical imaging datasets hampers the integration of advanced AI technologies into healthcare applications, particularly for less common conditions and underrepresented populations, due to privacy concerns and data availability issues.
A unified medical image-text generative model is used to create synthetic medical images based on textual instructions, trained with a latent stable diffusion model and reinforced by clinician feedback, enabling the generation of high-quality images across various imaging modalities and organs, and enhancing existing datasets through data augmentation.
The model improves the performance of AI systems in diagnostics, report generation, and self-supervised learning by providing diverse and clinically relevant synthetic images, addressing data scarcity and enhancing predictive capabilities in medical imaging.
Smart Images

Figure US2025050496_16042026_PF_FP_ABST
Abstract
Description
Attorney Docket No. 71951-711.602SYSTEMS AND METHODS FOR GENERATING SYNTHETIC MEDICAL IMAGESCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit and priority of International Patent Application PCT / CN2024 / 124230 entitled “Methods for Generating Synthetic Medical Images and Improving Clinical Models,” filed October 11, 2024, the contents of which are incorporated herein by reference in the entirety for all purposes.FIELD
[0002] The present disclosure relates to methods and systems for medical image processing and the generation of synthetic medical images for use in artificial intelligence training, selfimprovement, and diagnostic of diseases (such as cancer).BACKGROUND
[0003] Availability of large healthcare datasets is pivotal for driving Al model development and clinical applications. However, privacy concerns pose significant ethical and legal issues, making it difficult to share such data. The scarcity of high-quality imaging datasets has hindered the integration of cutting-edge Al technologies into medicine and healthcare applications.BRIEF SUMMARY
[0004] In many clinical and research settings, the scarcity of high-quality medical imaging datasets has hampered the potential of Al clinical applications. This issue is particularly pronounced in less common conditions, underrepresented populations, and emerging imaging modalities, where the availability of diverse and comprehensive datasets is often inadequate. To address this challenge, in some embodiments provided herein are data augmentation and the use of synthetic images with generative foundation models.
[0005] In some embodiments, disclosed herein is a unified medical image-text generative model. In some embodiments, using a method described herein, synthetic medical images of various organs across various imaging modalities based on textual instructions can be generated using a unified medical image-text generative model. In some embodiments, high quality synthetic images are generated using a unified medical image-text generative model disclosed herein, and clinician evaluations and rigorous objective measurements are used to validate the quality of the synthetic images. In some embodiments, a unified medical image-Attorney Docket No. 71951-711.602 text generative model disclosed herein exhibits an enhanced generative capability when presented with previously unseen data domains and is used as a generalist medical Al (GMAI). In some embodiments, synthetic images generated using a unified medical imagetext generative model disclosed herein augment existing datasets, boosting performance across multiple medical applications such as diagnostics, report generation and selfsupervised learning.
[0006] In some embodiments, provided herein is a computer-implemented method, comprising training a latent stable diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, thereby generating a trained latent stable diffusion model. In some embodiments, the method further comprises providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text. In some embodiments, the method further comprises processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality.
[0007] In any of the preceding embodiments, during the training, the modality information and the corresponding textual description information can be separately encoded. In any of the preceding embodiments, during the training, the modality information and the corresponding textual description information can be encoded using a BERT tokenizer.
[0008] In any of the preceding embodiments, the training can comprise introducing a series of random Gaussian noises to the training image input progressively. In any of the preceding embodiments, the training can comprise denoising using a U-Net architecture with a crossattention mechanism. In any of the preceding embodiments, the cross-attention mechanism can comprise cross-attention mechanisms on both the modality information and the corresponding description information. In any of the preceding embodiments, the U-Net architecture can comprise multiple U-Net layers. In any of the preceding embodiments, the U-Net architecture can comprise one or more shallow U-Net layers with cross-attention between image embedding and modality information embeddings. In any of the preceding embodiments, the U-Net architecture can comprise one or more deep U-Net layers with cross-attention between image embedding and description information embeddings.
[0009] In any of the preceding embodiments, the processing can comprise an iterative denoising process with sequential cross-attention. In any of the preceding embodiments, theAttorney Docket No. 71951-711.602 processing can comprise incorporating modality-specific information followed by refining with detailed descriptive information. In any of the preceding embodiments, the trained latent stable diffusion model can learn to distinguish between general modality characteristics and specific image details in descriptive texts. In any of the preceding embodiments, the training can comprise using pre-trained U-Net parameters sourced from a general domain and fine- tuning with a plurality of distinct modalities of medical images, each paired with its corresponding textual description. In any of the preceding embodiments, the training can comprise using a classifier-free guidance scale and a noise scheduler involving a pseudo- numerical method for diffusion.
[0010] In any of the preceding embodiments, the method can comprise a reinforcement learning strategy using human ratings. In any of the preceding embodiments, the method can comprise passing the synthetic medical image to a classification module to generate a classification result for the synthetic medical image. In any of the preceding embodiments, the method can comprise feeding the synthetic medical image to the classification module which decides whether the synthetic medical image should pass based on the classification result. In any of the preceding embodiments, the method can comprise a closed-loop, sustainable self-evolution mechanism which directly integrates human expert clinical judgment into an iterative optimization process. In any of the preceding embodiments, the method can comprise presenting the synthetic medical image to a clinician through a structured scoring interface, and inviting the clinician to provide quantitative scores across multiple dimensions comprising anatomical accuracy and clinical diagnostic value of the synthetic medical image. In any of the preceding embodiments, the method can comprise reward model training, wherein the quantitative scores serve as training data to train a reward model which learns and simulates clinician preferences to provide scores for the synthetic medical image that correspond to its clinical quality. In any of the preceding embodiments, output of the trained reward model can be used as reinforcement learning signals to fine-tune the trained latent stable diffusion model, incentivizing the trained latent stable diffusion model to produce synthetic medical images that achieve higher clinical scores.
[0011] In any of the preceding embodiments, the method can comprise cross-domain knowledge transfer and collaborative enhancement learning. In any of the preceding embodiments, the method can comprise inputting new domain data from one or more new domains in the trained latent stable diffusion model. In any of the preceding embodiments, the new domain data can comprise: (i) an additional medical image of an image modality that is different from the plurality of different image modalities and / or of an additional organ thatAttorney Docket No. 71951-711.602 is different from the plurality of different organs, and (ii) text input comprising modality information of the additional medical image concatenated with its corresponding textual description information.
[0012] In any of the preceding embodiments, the training can comprise using multimodal data comprising vital signs, imaging data, omics data on the cellular / molecular level (e.g., Genomic Alterations: Mutations, copy number variations (CNVs), chromosomal rearrangements, scDNA-seq; Epigenetics: DNA methylation, histone modifications, scATAC-seq; Transcriptomics: RNA-seq, gene expression profiles, scRNA-seq; Proteomics: Protein expression, post-translational modifications; Metabolomics: Targeted Metabolomics, Untargeted Metabolomics, Lipidomics, Fluxomics; Spatial Multi -omics Integration: Spatial Transcriptomics, Spatial Metabolomics, Spatial Proteomics, Spatial Epigenomics), Pathological Data (e.g., Histopathology, Immunohistochemistry (IHC), Cytopathology), Laboratory Data (e.g., Blood Tests, Tumor Markers, Liquid Biopsies), and EHRs electronic health records (EHRs) (e.g., including clinical notes, Demographics, Medical History, Symptoms & Signs, Treatment Records), or any combination thereof. In any of the preceding embodiments, the training can comprise using tabular data.
[0013] In some embodiments, disclosed herein is a computer-implemented method, comprising: (a) training a latent stable diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, wherein the training comprises denoising using one or more shallow U-Net layers with cross-attention between image embedding and modality information embeddings and one or more deep U- Net layers with cross-attention between image embedding and description information embeddings, thereby generating a trained latent stable diffusion model; (b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; (c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality; (d) presenting the synthetic medical image to a clinician through a structured scoring interface, and (e) inviting the clinician to provide quantitative scores across multiple dimensions comprising anatomical accuracy and clinical diagnostic value of the synthetic medical image, wherein the quantitative scores serve as training data to train a reward model which learnsAttorney Docket No. 71951-711.602 and simulates clinician preferences to provide scores for the synthetic medical image that correspond to its clinical quality.
[0014] In some embodiments, disclosed herein is a computer-implemented method, comprising: (a) training a latent stable diffusion model using: (i) training image input comprising medical images I, in} and their corresponding textual descriptionsD =dn} across a plurality of different image modalities M ={m^ m2, m3, ••• , mk} and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input [mk; dn] comprising modality information of each medical image concatenated with the corresponding textual description information, thereby generating a trained latent stable diffusion model, wherein the training comprises (i) introducing a series of random Gaussian noises to the training image input progressively, following the equation q(it|it-i)=linear noise scheduler defined as / 3t= / 3min+ t ■ anc[ (ii) reversing diffusion using a U-Net architecturewith cross-attention mechanisms on both modality information and description information:is predicted by the output of the cross-attention between image embedding and either modality information embedding or description information embedding; (b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; and (c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality, wherein the processing comprises an iterative denoising process with sequential cross-attention comprising incorporating modality-specific information followed by refining with detailed descriptive information
[0015] In some embodiments, disclosed herein is a method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising providing a mixed dataset comprising (i) a real medical image of the subject, and (ii) the synthetic medical image for the subject generated by the computer-implemented method of any one of the embodiments disclosed herein, and using the mixed dataset to provide the medical diagnosis, prediction, and / or prognosis for the subject. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for a rare disease or an atypical case of a common disease. In any of the preceding embodiments, the medical diagnosis, prediction, and / or prognosis can comprise stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof. In any of the preceding embodiments, theAttorney Docket No. 71951-711.602 medical diagnosis, prediction, and / or prognosis can be for one or more cancers. In any of the preceding embodiments, the method can comprise training a medical diagnosis, prediction, and / or prognosis model with the mixed dataset as input. In some embodiments, the medical diagnosis, prediction, and / or prognosis model comprises a Transformer classification model.
[0016] In any of the preceding embodiments, the plurality of different image modalities can comprise optical clearance tomography (OCT), computed tomography (CT), X-ray, fundus photography, magnetic resonance imaging (MRI), ultrasound image, endoscopy image, positron emission tomography (PET), single photon emission computed tomography (SPECT), microscopy image, medical photography, elastography image, a thermogram image, or any combination thereof. In any of the preceding embodiments, the plurality of different organs can comprise eye, lung, breast, intestine, brain, kidney, pancreas, bladder, or any combination thereof.
[0017] In one aspect, the present application provides a computer-implemented system for generating synthetic medical images comprising: a processor and a computer readable storage medium encoded with a computer program that causes the processor to generate a synthetic medical image from a text description. In some embodiments, the medical images comprise one or more types of medical images. In some embodiments, the medical images is selected from a group consisting of optical clearance tomography (OCT), computed tomography (CT), X-ray, fundus photography, magnetic resonance imaging (MRI), ultrasound image, endoscopy image, positron emission tomography (PET), single photon emission computed tomography (SPECT), microscopy image, medical photography, elastography image, and a thermogram image. In some embodiments, the computer program comprises a neural network. In some embodiments, the neural network comprises an encoder and a decoder. In some embodiments, the neural network further comprises a diffusion model. In some embodiments, the neural network further comprises reinforcement learning with human feedback. In some embodiments, the generation of synthetic medical images comprise iterative denoising from random noise.
[0018] In another aspect, the present application provides a computer-implemented system for generating synthetic medical images comprising a processor and a computer readable storage medium encoded with a computer program that causes the processor to: Create a first training set comprising pairs of medical images and text descriptions, train a neural network in a first stage using the first training set, generate a first set of synthetic medical images using text descriptions as input into the neural network, create a second training set comprising pairs of the first set of synthetic medical images and text descriptions, train theAttorney Docket No. 71951-711.602 neural network in a second stage using the second training set, subjective evaluations, and objective evaluations, and generate a second set of synthetic medical images using text descriptions as input into the neural network.
[0019] In some embodiments, the present application provides a method of diagnosing a disease in an individual using a computer-implemented system that is trained using the synthetic medical images. In some embodiments, the disease is a cancer. In some embodiments, the disease is an atypical case of a common disease. In some embodiments, the disease is in the eye, lung, breast, intestine, brain, kidney, pancreas, or bladder. In some embodiments, the individual is a human. In some embodiments, the present application provides a method of detecting a genetic mutation in an individual using a computer- implemented system that is trained using the synthetic medical images. In some embodiments, the present application provides a method of generating reports of medical images using a computer-implemented system that is trained using the synthetic medical images. In some embodiments, the present application provides a method of self-supervised learning using a computer-implemented system that is trained using the synthetic medical images.
[0020] In some embodiments, disclosed herein is a system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method of any one of the embodiments disclosed herein.
[0021] In some embodiments, disclosed herein is a non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method of any one of the embodiments disclosed herein.
[0022] In some embodiments, disclosed herein is a system comprising: at least one hardware processor; non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; and instructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method of any one of the embodiments disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings illustrate certain features and advantages of this disclosure. These embodiments are not intended to limit the scope of the appended claims in any manner.
[0024] FIG. 1A shows a generative model’s (MINIM) development and deployment. The training was conducted with a latent stable diffusion model on paired images and reportsAttorney Docket No. 71951-711.602 from different modalities and organs. In deployment, MINIM generated high-quality synthetic images given diverse textual descriptions. The synthetic images were evaluated through subjective assessments from clinicians and multiple objective metrics.
[0025] FIG. IB shows a method of using the synthetic images as an additional training source for diagnosis, report generation and self-supervised learning. The synthetic images can assist clinical applications, e.g., in accurate prediction of mutations of HER2 in breast cancer MR images and EGFR in lung cancer CT images, and improve survival analysis.
[0026] FIG. 2A shows performance of MINIM and other text-to-image generative models’ ability to create an accurate synthetic medical image as evaluated by clinicians. Clinicians conducted three rounds of scoring on the generated synthetic images. A score of 1 denoted a low-quality image; a score of 2 indicated a high-quality image but one that was irrelevant to the report; and a score of 3 signified a high-quality image that aligned with the report. Based on the scoring results, the two-stage RL strategy model was trained and iteratively applied to MINIM.
[0027] FIG. 2B shows synthetic medical images and their text prompts generated by a generative model trained on input from clinicians. In the text-conditioned synthesis of medical images, samples were created by prompting a fine-tuned model using annotations made by ophthalmologists and radiologists.
[0028] FIG. 2C shows performance comparison of objective assessments from six objective metrics (FID, IS, MS-SSIM, CAS, IIR, ITR) using MINIM and other text-to-image generative models.
[0029] FIG. 2D shows a demonstration of synthetic medical images using different methods.
[0030] FIG. 3A shows an exemplary two-stage reinforcement learning strategy using human ratings. The first stage involves the generation of synthetic images from a generative model. The generated images are then passed to a ‘Receive Selector’, which operates as a classification module. The second stage introduces an active selection process. Again, synthetic images are fed into the system, but, this time, the ‘Receive Selector’ classification module decides which images to pass based on the classification results.
[0031] FIG. 3B shows six subjective results of the synthetic images using reinforcement learning on human feedback (RLFH) from clinicians’ ratings.
[0032] FIG. 4A shows performance improvement of MINIM after being trained on a distinct and new medical image type (e.g., MRI), by accommodating image-text paired data from new domains. Data is organized by image type.Attorney Docket No. 71951-711.602
[0033] FIG. 4B shows performance improvement of MINIM’ s generative ability by accommodating image-text paired data from new domains (e.g., MRI). Data is organized by assessment metric.
[0034] FIG. 4C shows a demonstration of the synthetic MRI generated using MINIM and other generative methods.
[0035] FIG. 5A shows a flowchart of breast MRI data collection and analysis for HER2 mutation detection by MINIM. Breast MR images were split into Tumor and Benign, each of which has three modalities, that is, Tl, Tic and T2. These images were sent to MINIM for generative model training and downstream HER2 mutation detection.
[0036] FIG. 5B shows examples of atypical cases of common diseases including retinal vein occlusions (RVOs), diabetic retinopathy (DR), and choroidal neovascularization (CNV) as well as the reports generated MINIM or a clinician. When compared to reports generated by selected report generation models and ophthalmologists, the models can identify more comprehensive and accurate information. Incorporating OCT image data synthesized by MINIM into the self-supervised learning of OCT image models can enhance the representational capabilities of pre-trained models.
[0037] FIG. 6A shows performance comparison by incorporating synthetic images from MINIM and other generative models into real data for diagnostic tasks.
[0038] FIG. 6B shows performance improvement of a diagnostic classifier for OCT images after integration of synthetic medical images into the training data. Synthetic images assist in ophthalmological diagnostic. For diagnostic labels with lower classification performance, MINIM can selectively generate synthetic OCT images and integrate them into the training data, thereby further refining the diagnostic model’s accuracy.
[0039] FIG. 7 A shows performance comparison using BLEU-2 score as a metric for report generation task by incorporating synthetic images from MINIM and other generative models into real data.
[0040] FIG. 7B shows performance comparison using CIDEr score as a metric for report generation task by incorporating synthetic images from MINIM and other generative models into real data.
[0041] FIG. 7C shows performance comparison using ROUGE -L score as a metric for report generation task by incorporating synthetic images from MINIM and other generative models into real data.
[0042] FIG. 8 shows performance comparison on the representations from self-supervised learning using MINIM and other generative methods.Attorney Docket No. 71951-711.602
[0043] FIG. 9A shows performance metrics of Al-based EGFR mutation classification from CT images using increasing ratio of synthetic CT images in the training datasets. The figure shows Al-based mutation prediction performance and its effect on 5 -year survival rates of patients with advanced lung cancer. Baseline and performance comparison on EGFR mutation predictions using different numbers of synthetic images in training datasets is shown.
[0044] FIG. 9B shows (i) a baseline survival curve in patients who underwent standard chemotherapy, (ii) five-year survival curves of true EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort A, (iii) five- year survival curves of Al identified EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort A, (iv) five-year survival curves of true EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort, and (v) five-year survival curves of Al identified the EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort B.
[0045] FIGS. 10A-10D show Al-based EGFR mutation prediction and its effect on progress! on -free survival (PFS) rates of advanced lung cancer patients.
[0046] FIG. 10A shows PFS curves of true EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in the cohort A.
[0047] FIG. 10B shows PFS curves of Al identified EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort A.
[0048] FIG. 10C shows PFS curves of true EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort B.
[0049] FIG. 10D shows PFS curves of Al identified EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort B.
[0050] FIGS. 11A-11D show Al-based EGFR mutation prediction and its effect on objective response rate (ORR) of advanced lung cancer patients.
[0051] FIG. 11A shows ORR of true EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in the cohort A.
[0052] FIG. 11B shows ORR of Al identified EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort A.
[0053] FIG. 11C shows ORR of true EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort B.Attorney Docket No. 71951-711.602
[0054] FIG. 11D shows ORR of Al identified EGFR sensitive mutations (underwent TKI therapy) versus insensitive mutations (underwent chemotherapy) in cohort B.
[0055] FIG. 12 shows Al-based HER2 mutation detection from MRI images using synthetic MRI images in the training dataset.DETAILED DESCRIPTION
[0056] Generative Al has enabled significant breakthroughs in medical imaging research, enhancing the analysis of medical records and images, and improving diagnostics and treatment planning. Recent efforts have demonstrated that generative Al can synthesize high- quality chest X-ray images, morphologically preserving 3D images of the brain, and 2D images of histopathology and dermatology, consequently enhancing medical image understanding and improving downstream analysis.
[0057] In some embodiments, Generative Adversarial Networks (GANs) are used for generating synthetic medical images and enhancing data availability for Al models. In some embodiments, pipelines built upon GANs that have incorporated segmentation techniques are used herein to further reduce the required amount of dataset of training, which is tailored for data-scarce scenarios, such as in cases of rare diseases. In some embodiments, within a medical imaging modality, generative Al is used to generate high-quality synthetic images to support medical imaging research.
[0058] However, traditional GAN-based image generation struggles with generating images of diverse dimensionalities and is thus often limited to a single imaging modality. The limited exploration of one modality and the lack of investigation into the relationships among different medical imaging modalities further hinder the full utilization of extensive multimodal medical image datasets and the development of universal medical generative models.
[0059] In some embodiments, disclosed herein are a self-improving generative foundation model for synthetic medical image generation and clinical applications. In some embodiments, described herein are systems, software, and methods of generating medical images using a unified medical image-text generative model for facilitating a self-supervised learning, Al-assisted interpretation and reporting of medical images, Al-assisted interpretation and diagnosis using medical images.
[0060] In some embodiments, a generative adversarial model or generative adversarial network is used, which comprises a generative function and a discriminative function, wherein the generative function creates synthetic data (e.g., synthetic medical images), andAttorney Docket No. 71951-711.602 the discriminative function distinguishes between synthetic and real data. By training the generative function and / or the discriminative function on the one hand the generative function is configured to create synthetic data which is incorrectly classified by the discriminative function as real, on the other hand the discriminative function is configured to distinguish between real data and synthetic data generated by the generative function. In some embodiments, a generative adversarial model can be interpreted as a zero-sum game. In some embodiments, the training of the generative function and / or of the discriminative function is based on the minimization of a cost function. In some embodiments, based on a set of training data synthetic data can be generated that has the same characteristics as the training data set. In some embodiments, the training of the generative adversarial model comprises supervised learning and / or unsupervised learning. In some embodiments, the training of the generative adversarial model can be based on data not being annotated (unsupervised learning), so that there is low effort in training a generative adversarial model.
[0061] In some embodiments, a self-improving generative foundation model disclosed herein integrates medical images paired with textual descriptions across various modalities and organs. In some embodiments, a self-improving generative foundation model disclosed herein integrates medical images paired with textual descriptions across optical coherence tomography (OCT), fundus, chest X-ray, chest computed tomography (CT), or any combination thereof. In some embodiments, one or more synthetic images for each organ and imaging modality are generated based on textual descriptions. In some embodiments, the performance of self-improving generative foundation model in generating synthetic images based on textual inputs is evaluated, and its performance was tested across different medical imaging scenarios against other generative models. In some embodiments, data from a new domain - e.g., magnetic resonance imaging (MRI) datasets for brain and breast imaging - are incorporated in the self-improving generative foundation model for continuous learning and integration of new medical knowledge.
[0062] In some embodiments, the systems, software, and methods disclosed herein can incorporate one or more of the processes or subsystems disclosed herein, including but not limited to, Al-assisted generation of synthetic medical images, clinician-assisted improvement of Al models, Al-assisted generation of reports for medical images, Al-assisted self-improvement, Al-assisted classification or detection of a disease (such as cancer), Al- assisted classification or detection of a genetic mutation. Algorithms that can be used in the processes or subsystems include various models such as computer vision and naturalAttorney Docket No. 71951-711.602 language processing algorithms disclosed herein. Accordingly, the present disclosure contemplates any combination of the systems or subsystems and methods disclosed herein.
[0063] In some embodiments, disclosed herein is a computer-implemented system for generating synthetic medical images comprising: a processor and a computer readable storage medium encoded with a computer program that causes the processor to generate a synthetic medical image from a text description. In some embodiments, the medical images comprise one or more types of medical images, such as X-rays, CT, PET, PET-CT, MRI, OCT, FDG- PET, perfusion imaging, Radiomics, and / or fundus images. In some embodiments, the medical image is selected from a group consisting of optical clearance tomography (OCT), computed tomography (CT), X-ray, fundus photography, magnetic resonance imaging (MRI), ultrasound image, endoscopy image, positron emission tomography (PET), single photon emission computed tomography (SPECT), microscopy image, medical photography, elastography image, and a thermogram image. In some embodiments, the computer program comprises a neural network.
[0064] In some embodiments, the neural network comprises an encoder and a decoder. In some embodiments, the neural network further comprises a diffusion model. In some embodiments, the neural network further comprises reinforcement learning with human feedback. In some embodiments, the generation of synthetic medical images comprise iterative denoising from random noise.
[0065] In some embodiments, disclosed herein is a computer-implemented system for generating synthetic medical images comprising a processor and a computer readable storage medium encoded with a computer program that causes the processor to: create a first training set comprising pairs of medical images and text descriptions, train a neural network in a first stage using the first training set, generate a first set of synthetic medical images using text descriptions as input into the neural network, create a second training set comprising pairs of the first set of synthetic medical images and text descriptions, train the neural network in a second stage using the second training set, subjective evaluations, and objective evaluations, and generate a second set of synthetic medical images using text descriptions as input into the neural network.
[0066] In some embodiments, disclosed herein is a method of diagnosing a disease in an individual using a computer-implemented system that is trained using the synthetic medical images. In some embodiments, the disease is a cancer. In some embodiments, the disease is an atypical case of a common disease. In some embodiments, the disease is in the eye, lung, breast, intestine, brain, kidney, pancreas, or bladder. In some embodiments, the individual is aAttorney Docket No. 71951-711.602 human. In some embodiments, disclosed herein is a method of detecting a genetic mutation in an individual using a computer-implemented system that is trained using the synthetic medical images. In some embodiments, disclosed herein is a method of generating reports of medical images using a computer-implemented system that is trained using the synthetic medical images. In some embodiments, disclosed herein is a method of self-supervised learning using a computer-implemented system that is trained using the synthetic medical images.
[0067] In some embodiments, provided herein is a method of generating a unified multidomain medical image generation foundation model capable of generating synthetic medical images, and a method of using the foundation model for various downstream applications, including diagnosis of a disease or condition, report generation, self-supervised learning, or any combination thereof. Also disclosed herein in some embodiments are methods of using synthetic medical images generated by the foundation model in clinical applications, including for instance, detection of one or more genetic variants such as a cancer mutation, detection of a gene status in a cancer, or a combination thereof.
[0068] In some embodiments, the foundation model is a unified foundation model architecture rather than a series of independent single-task models. In some embodiments, the foundation model is based on one or more latent diffusion models, with high flexibility and scalability of its input conditions. In some embodiments, the foundation model concatenates imaging modality types, including chest CT, chest X-ray, fundus color photography, fundus OCT, and corresponding descriptive text as unified conditional inputs. In some embodiments, through cross-attention mechanisms, the model understands and fuses instructions from different sources within a unified latent space, thereby achieving multi-organ (e.g., ophthalmology, chest, brain, breast, etc.) and multi-modal image generation capabilities within a single model. In some embodiments, the "train once, deploy everywhere" architecture serves as the foundation for building universal medical Al generation models, fundamentally distinguishing it from the fragmented "one model, one task" approach.
[0069] In some embodiments, the foundation model is a self-evolution system based on clinical expert feedback (e.g., RLHF). In some embodiments, the foundation model comprises a closed-loop, sustainable self-evolution mechanism. In some embodiments, the foundation model directly integrates human expert clinical judgment into the model's iterative optimization process. In some embodiments, the process comprises structured clinical feedback collection, where after the system generates synthetic images, they are presented to clinicians through a structured scoring interface, and inviting physicians to provideAttorney Docket No. 71951-711.602 quantitative scores across multiple dimensions, including anatomical accuracy and clinical diagnostic value. In some embodiments, the process comprises reward model training, where the collected expert scores serve as training data to train a reward model which aims to learn and simulate clinician preferences, capable of providing scores for a generated image that correspond to its clinical quality. In some embodiments, the process comprises reinforcement learning-based model fine-tuning, where once the reward model is trained, it serves as an automated "virtual expert," and the reward model's output is used as reinforcement learning signals, employing algorithms such as Proximal Policy Optimization (PPO) to fine-tune the generation model itself. This process incentivizes the generation model to produce images that achieve higher clinical scores. This mechanism creates a self-learning Al system where model optimization objectives shift from traditional pixel -level similarity metrics to the critically important clinical usability, enabling generated image quality to continuously and directionally approximate real-world clinical standards.
[0070] In some embodiments, the foundation model comprises cross-domain knowledge transfer and collaborative enhancement learning. In some embodiments, the foundation model comprises continuous learning and knowledge transfer capabilities. In some embodiments, the foundation model is a multi-modal foundation model which, when encountering data from one or more entirely new domains, not only learns to generate images in the new domain but also reciprocally enhances its capabilities in original domains. For example, when the foundation model completes learning chest CT data and subsequently introduces chest MRI data for training, the quality and diversity of generated chest CT images are further improved. This collaborative enhancement effect demonstrates that the foundation model is learning universal medical imaging principles that transcend specific modalities and organs, such as tissue texture, pathological morphological principles, and spatial relationships, ensuring fusion rather than forgetting of new and old knowledge, thereby achieving continuous cumulative growth of model capabilities.
[0071] In some embodiments, the foundation model comprises a unified multi-modal architecture with end-to-end generation paradigm. In some embodiments, the foundation model comprises a unified multi-modal foundation model architecture, achieving generation through end-to-end latent diffusion models. In some aspects, the foundation model disclosed herein does not generate intermediate atlases. In some aspects, the foundation model disclosed herein does not use a dual-pathway, stepwise process that relies on "anatomical modeling" modules to generate intermediate atlases. In some aspects, the foundation model disclosed herein directly guides one or more diffusion models to create images from latentAttorney Docket No. 71951-711.602 space noise by concatenating modality instructions (such as "CT") with descriptive text. In some aspects, the superiority of this approach lies in its liberation from discrete, rigid intermediate representations, achieving direct mapping from high-level semantic instructions to high-fidelity pixels, thereby generating more diverse images containing anatomical variations common in the real world.
[0072] In some embodiments, the foundation model comprises one or more diffusion models, completing the denoising generation process in compressed, continuous latent spaces. In some aspects, the foundation model disclosed herein does not rely on "anatomical atlases" as interpretable intermediate steps, which constitute a technical bottleneck — the accuracy and granularity of atlases directly determine the upper limit of final image quality. In some aspects, the foundation model disclosed herein operates within latent spaces, not only dramatically improving computational efficiency but, more importantly, avoiding information loss caused by discretized intermediate representations. In some aspects, the operation in latent spaces endows the system with the capability to capture and generate more detailed, unstructured pathological features (such as infiltrative boundaries of early tumors).
[0073] In some embodiments, the foundation model disclosed herein comprises closed-loop optimization, e.g., for clinical value alignment. In some embodiments, the foundation model comprises Reinforcement Learning from Human Feedback (RLHF) based on clinical expert feedback, constructing a dynamic, continuously optimizable closed-loop system. In some aspects, the foundation model disclosed herein does not merely pursue "structural consistency" with textual descriptions. In some aspects, the foundation model disclosed herein optimizes the clinical diagnostic value of generated images through reward model learning. In some aspects, the foundation model's holistic semantic understanding capability is dynamically evolving, continuously aligning textual instruction interpretation with clinicians' actual judgments, thereby generating synthetic data that not only "looks right" but also "works well."
[0074] In some embodiments, the foundation model disclosed herein is an instruction-based interactive model. In some embodiments, the foundation model disclosed herein does not rely on complex frontend NLP modules to forcibly convert text into Structured Language Sequences (SLS) through fixed processes. In some embodiments, the foundation model disclosed herein directly understands and executes natural language instructions composed of simple concatenations. In some embodiments, the model disclosed herein comprises autonomous learning and generalization capabilities as a foundation model, making theAttorney Docket No. 71951-711.602 system universal and scalable, capable of easily adapting to new tasks and instructions rather than being confined to predefined structured paradigms.
[0075] In some embodiments, the foundation model disclosed herein comprises a diffusion generation paradigm. In some embodiments, the foundation model disclosed herein comprises a diffusion model-based generator, creating new samples that conform to textual conditions from random noise. In some embodiments, the foundation model disclosed herein does not train a Transformer model or network with training image input. In some embodiments, the foundation model disclosed herein does not train a Transformer model or network with training text input. In some embodiments, the foundation model disclosed herein does not train a Transformer model or network with training image input or training text input.
[0076] In some embodiments, the foundation model disclosed herein does not use multimodal Transformers as translators. In some embodiments, the foundation model disclosed herein may be used to translate between different types of highly structured features. In some embodiments, the foundation model disclosed herein is not designed to translate between different types of highly structured features. In some embodiments, the foundation model disclosed herein comprises generative capabilities, enabling it not only to reproduce patterns in training data but also to generalize and combine medical knowledge, creating entirely new yet clinically reasonable image instances.
[0077] In some embodiments, the foundation model disclosed herein comprises a RLHF- calibrated guidance mechanism. In some embodiments, the foundation model disclosed herein comprises an approach in feature fusion and / or guidance that is different from a Transformer fusion mechanism. In some embodiments, the foundation model disclosed herein may but does not require optimization of consistency between image and text features. In some embodiments, the foundation model disclosed herein is not aimed to optimize consistency between image and text features. In some embodiments, the foundation model disclosed herein comprises a feature fusion and / or guidance mechanism based on hierarchical guidance through U-Net cross-attention, introduces one or more reward models trained by RLHF. In some embodiments, the reward models provide additional optimization signals for the generation process, reflecting clinical relevance as defined by expert feedback. In some embodiments, the generation objectives of the foundation model disclosed herein encompass not only data-level fidelity but also considerations of clinical utility.
[0078] In some embodiments, disclosed herein is a unified medical image-text generative model trained on paired medical images with textual descriptions across various modalitiesAttorney Docket No. 71951-711.602 and organs (OCT, fundus, chest X-ray, chest CT, brain MRI and breast MRI). In some embodiments, the generative model disclosed herein is used to generate high-quality synthetic images for each organ and imaging modality given textual descriptions. In some embodiments, the generated data comprising synthetic medical images can function as a supplementary resource for the development of Al models, particularly in scenarios where there is a lack of sufficient data due to privacy or cost. In some embodiments, provided herein is a method of training a model on both real and synthetic data for improved predictive capabilities compared to training only on real data, with this enhancement being very beneficial in the context of rare disease diagnoses, report generation and / or self-supervised learning. In some embodiments, provided herein is a method of using the model to enhance the detection of one or more genetic variations, such as HER2 mutations in the breast. In some embodiments, provided herein is a method of using the model to predict patient responses to a therapy, such as an EGFR mutation target therapy for improved 5-year survival rates. The evaluations of classification accuracy in lung cancer and breast cancer demonstrate the clinical relevance and advantages of the model disclosed herein.
[0079] For lung cancer, EGFR mutation classification plays a critical role in determining the suitability of TKI-targeted treatments, which substantially improves survival in patients with advanced non-small cell lung cancer. By evaluating the model’s performance in both three- class (wild, sensitive and resistant) and binary (sensitive versus resistant + wild) classification tasks, the model’s robustness in accurately predicting EGFR status is demonstrated, which is crucial for guiding targeted therapy decisions. In some embodiments, the Al-driven approach provides a non-invasive, efficient and scalable alternative to traditional biopsy -based EGFR mutation detection, ultimately improving patient outcomes and expanding access to precision medicine.
[0080] Breast cancer is the most common cancer in women worldwide, posing a considerable threat to women’s health. HER2 is a critical biomarker for determining the benefit of HER2- targeted therapies. Traditionally, only HER2 amplification or overexpression was used to predict an enhanced survival benefit from HER2 -targeted therapy, and the model disclosed herein addresses this limitation by offering a more comprehensive and nuanced approach to HER2 status classification. By leveraging both real and synthetic datasets, our model identifies subtle patterns associated with HER2 -positive tumors that may not be detectable through conventional methods. This enhanced ability to classify HER2 status more accurately could lead to better identification of patients who would benefit from HER2 -targetedAttorney Docket No. 71951-711.602 therapies, ultimately improving treatment outcomes and survival rates in patients with HER2- positive breast cancer.
[0081] Despite the advancements in GANs, challenges persist in generating high-quality medical images due to limited training data. In some embodiments, the generative model disclosed herein addresses this limitation by integrating both textual and imaging data, effectively expanding the input space and enhancing the quality of the generated images. Furthermore, concerns regarding the clinical relevancy of generated images are also mitigated by training the model with paired image and text inputs, ensuring that the outputs are closely aligned with the clinical contexts described in the accompanying text.
[0082] Another longstanding question in the field of generative Al has been the extent to which synthetic data could enhance predictive model performance. In some embodiments, the generative model disclosed herein can self-improve through RLHF and transfer learning, enabling continuous learning and addressing the challenge of adapting to new data domains. In some embodiments, the generative model disclosed herein can adapt and learn prospectively over time across various applications.
[0083] In some embodiments, the generative model disclosed herein incorporates data from a varied array of sources, ensuring a comprehensive representation of different genders and ethnic groups, ensuring not only the model’s robustness but also its applicability across a broader demographic spectrum.
[0084] In some embodiments, the generative model disclosed herein incorporates detailed textual information, including findings and clinical details, to enhance the model’s conditioning capabilities, allowing for a more nuanced and comprehensive analysis of medical images in conjunction with their accompanying textual reports. In some embodiments, the generative model disclosed herein is not constrained by the token length of a CLIP text encoder.
[0085] In some embodiments, the generative model disclosed herein aligns text and images, e.g., when dealing with longer prompts. In some embodiments, the generative model disclosed herein employs reward models that are trained based on human feedback, for the precision of text-image alignment, making the model effective in handling complex, lengthy prompts. In some embodiments, the generative model disclosed herein addresses the risk of the model overfitting. In some embodiments, the generative model disclosed herein incorporates meticulous strategies, such as integrating adversarial loss and adaptively managing the learning rate during fine-tuning, especially under the constraints of limitedAttorney Docket No. 71951-711.602 data. In some embodiments, the generative model disclosed herein incorporates continuous human oversight and feedback, ensuring that the model remains accurate and reliable.
[0086] In some embodiments, provided herein include a computer-implemented method for generating synthetic medical images, a method of using the generated synthetic medical images, a system, a computer-readable medium, and a computer program product for performing any one or more steps of a method disclosed herein. In some embodiments, a method disclosed herein comprises obtaining a dataset comprising medical information, wherein the medical information comprises (i) real medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) modality information of each medical image concatenated with the corresponding textual description information. In some embodiments, the medical information comprises information specifying an anatomical structure to be depicted in a synthetic medical image. In some embodiments, the medical information does not comprise information specifying an anatomical structure to be depicted in a synthetic medical image. In some embodiments, the medical information comprises imaging information specifying image formation characteristics of the synthetic medical image. In some embodiments, the medical information does not comprise imaging information specifying image formation characteristics of the synthetic medical image. In some embodiments, the medical information does not comprise anamnesis or physical examination descriptions. In some embodiments, the method does not comprise generating an anatomical map of the anatomical structure, e.g., by means of anatomy modelling using the anatomical description as an input. In some embodiments, the method does not comprise conditioning a generative model using an anatomical map. In some embodiments, the method does not comprise branching off and embedding part of the medical information in a generative model, and using part of the information to generate an anatomical map which is then used for conditioning the generative model.
[0087] In some embodiments, provided herein is a computer-implemented method, comprising training a diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, thereby generating a trained diffusion model. Diffusion models, also known as diffusion probabilistic or score-based generative models, are a class of generative models. In principle, they generate data by reversing aAttorney Docket No. 71951-711.602 diffusion process, where a diffusion process gradually adds noise to data. The model is trained to reverse the diffusion process by starting from and gradually removing noise to generate output. In some embodiments, a latent diffusion model is trained and / or used according to embodiments of the present disclosure. In some embodiments, latent diffusion models are advantageous where a generative model generates image output, while also being able to handle other types of data and being suitable for use in multimodal scenarios. In some embodiments, latent diffusion models use an embedding mechanism to efficiently handle memory -intense image information (e.g. where parts of the image space representation are redundant pieces of information and can be compressed).
[0088] In some embodiments, embedding comprises taking a piece of information in one representation and transforming it into another representation which is more advantageous for some purpose. The information used to train a model disclosed herein may be provided in the form of training image input and / or training text input. In order to carry out a number of processing steps on this information, the training text input may be converted into a numerical representation that compresses, but most importantly preserves the target information encoded in the text. To achieve this, the text may be split into fragments. These fragments are called tokens and the process is referred to as tokenization. The resulting stream of tokens may then be transformed into even more condensed representations by a learned or given transformation (e.g. a linear projection). These representations are called embeddings, as they embed the original text information into a different space, which can be seen as a more appropriate space. The space is often referred to as latent space and may be seen as mathematically corresponding to a manifold that particularly preserves the relevant and discards the irrelevant information. Beyond this dimensionality reduction, similar information is better clustered in that space and can be better told apart. Encoders which typically perform the transformation into the embedding space can be learnt based on the task, can be taken from a generic pre-trained encoder and possibly refined, or can be manually constructed, for example. An example of a generic text encoder is Contrastive Language-Image Pretraining, CLIP. In addition to text, images can also be embedded into different space, e.g., a more appropriate and condensed space, in a similar manner. As this yields a more compact representation of the relevant information (e.g., memory footprint), it allows for reducing computational overhead and is, thus, particularly useful for models such as latent diffusion models.
[0089] Processing steps in a model disclosed herein can attend to the embeddings to process and recombine their inputs in a manner that is influenced by the information. The above mayAttorney Docket No. 71951-711.602 be referred to as an attention mechanism. An example for this attention mechanism is multi - head cross attention. Multi-head cross attention may be used, wherein “cross” refers to attending to some other piece of information (e.g., medical images to text descriptions as opposed to “self attention” - e.g. where a medical image attends to itself). "Multi-head" refers to multiple of these attention blocks that co-exist at the same place but have different learnt parameters. In contrast to spatial convolutions, certain attention mechanisms better model long-range dependencies.
[0090] Certain embodiments described herein are described with respect to methods and systems utilizing trained models, as well as with respect to methods and systems for training models. Features, advantages or alternative embodiments herein can be assigned to the other claimed objects and vice versa. In other words, claims and embodiments for training models can be improved with features described or claimed in the context of utilizing trained models, and vice versa. In particular, datasets used in the methods and systems for utilizing trained models can have the same properties and features as the corresponding datasets used in the methods and systems for providing trained models, and the trained models provided by the respective methods and systems can be used in the methods and systems for utilizing the trained models.
[0091] In some embodiments, a trained model (or “trained function”) mimics cognitive functions that humans associate with other human minds. In particular, by training based on training data the model is able to adapt to new circumstances and to detect and extrapolate patterns.
[0092] In some embodiments, parameters of a model can be adapted by means of training, including supervised training, semi -supervised training, unsupervised training, reinforcement learning, active learning, representation learning (or “feature learning”), or any combination thereof. In some embodiments, the parameters of the models can be adapted iteratively by several steps of training. In some embodiments, within the training a certain cost function can be minimized. In some embodiments, within the training of a neural network the backpropagation algorithm can be used.
[0093] In some embodiments, described herein is a multimodal data integration method using latent space for cross-scale biomedical applications, including generating synthetic medical images from multimodal training data or input using latent space, and using the generated synthetic medical images for cross-scale biomedical applications. In some embodiments, high dimensional multimodal data, such as medical images (e.g., X-rays, CT, PET, PET-CT, MRI, OCT, FDG-PET, perfusion imaging, Radiomics, and / or fundus images), omics data on theAttorney Docket No. 71951-711.602 cellular / molecular level (e.g., Genomic Alterations: Mutations, copy number variations (CNVs), chromosomal rearrangements, scDNA-seq; Epigenetics: DNA methylation, histone modifications, scATAC-seq; Transcriptomics: RNA-seq, gene expression profiles, scRNA- seq; Proteomics: Protein expression, post -translational modifications; Metabolomics: Targeted Metabolomics, Untargeted Metabolomics, Lipidomics, Fluxomics; Spatial Multi- omics Integration: Spatial Transcriptomics, Spatial Metabolomics, Spatial Proteomics, Spatial Epigenomics), Pathological Data (e.g., Histopathology, Immunohistochemistry (IHC), Cytopathology), Laboratory Data (e.g., Blood Tests, Tumor Markers, Liquid Biopsies), and EHRs (e.g., including clinical notes, Demographics, Medical History, Symptoms & Signs, Treatment Records), are used to build a shared latent space comprising low dimensional data. Techniques such as variational autoencoder (VAE) and / or generative adversarial network (GAN) can be used in the building of the latent space. For instance, VAE is a type of neural network that learns to reproduce its input and map data to a latent space. In some embodiments, the latent space comprises an embedding of a set of items within a manifold in which items resembling each other are positioned closer to one another, based on that many high-dimensional data sets can lie along low-dimensional latent manifolds within that space. In some embodiments, the latent space is a geometric latent space. Using the latent space, the multimodal data can be integrated and used to impute data within a modality and / or synthesize data across modalities. In some embodiments, the multimodal data can be integrated using the latent space to predict aging, disease trajectory, and / or development. In some embodiments, the multimodal data can be integrated using the latent space for crossscale mapping (e.g., from the molecular level to the cellular level, tissue level, and / or organismal level, or vice versa), and / or for making mechanism connections.
[0094] In some embodiments, multi -omics data are processed using unsupervised pretraining learning and / or deep learning architectures, such as convolutional neural networks (CNNs), generative adversarial networks (GANs), transformers and sequence models, and / or diffusion models. In some embodiments, the processed data are further processed using multi-model diffusion, for instance, using early integration at the data level, intermediate integration at the representation level, and / or late integration at the model level.
[0095] In some embodiments, provided herein is a method of generating a trained model (e.g., LLM), comprising (i) training a model using longitudinal multimodal data of a plurality of subjects, wherein the model comprises an examination encoder, a temporal embedding, and task-specific decoder heads, wherein for each subject, the longitudinal multimodal data comprise longitudinal EHR data from a chronological sequence of clinical visits of theAttorney Docket No. 71951-711.602 subject, and (ii) adapting the subject-level longitudinal representations to distinct tasks using the task-specific decoder heads, wherein the distinct tasks comprise first occurrence disease diagnosis and future disease prediction, thereby generating the trained model.
[0096] In some embodiments, provided herein is a method of generating a trained model (e.g., LLM), comprising: pretraining a model using longitudinal multimodal data of a plurality of subjects, wherein the model comprises an examination encoder, a temporal embedding, and task-specific decoder heads, wherein for each subject, the longitudinal multimodal data comprise EHR data from a chronological sequence of clinical visits of the subject. In some aspects, the examination encoder generates a contextualized representation of subject data from each clinical visit. In some aspects, the temporal embedding captures temporal relationships between clinical visits from the output of the examination encoder to generate a subject-level longitudinal representation for each subject. In some embodiments, the method further comprises finetuning the pretrained model by adapting the subject-level longitudinal representations to distinct tasks using the task-specific decoder heads, wherein the distinct tasks comprise first occurrence disease diagnosis and future disease prediction, thereby generating the trained model.
[0097] In some embodiments, the plurality of subjects comprise cancer subjects and the longitudinal multimodal data comprise routine laboratory results, vital signs, routine imaging data, or any combination thereof. In some embodiments, the routine laboratory results comprise results of any one or more of the biomarkers. In some embodiments, the routine laboratory results comprise complete blood count (CBC), blood chemistry, coagulation, or any combination thereof. In some embodiments, the vital signs comprise heart rate, blood pressure, body mass index (BMI), or any combination thereof. In some embodiments, the routine imaging data comprise chest X-rays (CXRs). In some embodiments, the longitudinal multimodal data comprise alterations to the pulmonary vasculature, shifts in bone density, changes in soft tissue composition, systemic inflammation, hormonal dysregulation, hemodynamic shifts, early cachexia, or any combination thereof.
[0098] In some embodiments, the longitudinal multimodal data comprise categorical variables and / or continuous variables. In some embodiments, the continuous variables are discretized to preserve their distributional characteristics. In some embodiments, for each subject, the longitudinal multimodal data comprise tabular clinical variables and image data. In some embodiments, for each modality of subject data, the examination encoder simultaneously captures both a value distribution and a semantic meaning of the modality of subject data.Attorney Docket No. 71951-711.602
[0099] In some embodiments, the temporal embedding comprises a linear positional embedding to learn time-dependent patterns in the longitudinal EHR data. In some embodiments, the temporal embedding is autoregressive and comprises causal masking to ensure unidirectional information flow in the autoregressive process.
[0100] In some embodiments, each of the task-specific decoder heads comprises a separate pathway applying a projection layer followed by Rectified Linear Unit (ReLU) activation. In some embodiments, each of the task-specific decoder heads comprises causal masking to prevent information from future clinical visits from influencing prediction at a given clinical visit.
[0101] In some embodiments, the multimodal model further comprises a missingness discriminator that determines whether a value of a particular feature is missing or not. In some embodiments, the model further comprises a gradient reversal layer (GRL) between the missingness discriminator and the examination encoder, wherein the GRL inverts the gradient during backpropagation, compelling the examination encoder to produce a representation that is independent of the missingness status of the particular feature. In some embodiments, the model further comprises a cohort discriminator that identifies the cohort label of each subject and forces the examination encoder to suppress cohort-specific information. In some embodiments, the pretraining is a self-supervised pretraining.
[0102] In some embodiments, provided herein is a method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising receiving a natural -language prompt for obtaining the medical diagnosis, prediction, and / or prognosis and a set of data related to the subject, and generating the medical diagnosis, prediction, and / or prognosis by inputting the prompt and the set of data in the trained model disclosed herein.
[0103] In some embodiments, the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for a pan-cancer diagnosis, prediction, and / or prognosis. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for one or more specific cancers. In some embodiments, the medical diagnosis, prediction, and / or prognosis is for: prediction of an organ-specific age, predicting of female aging, prediction of a hospital -acquired infection, prediction of a dialysis time series, prediction of diabetes and complications, prediction of pregnancy and child outcome, prediction of heartAttorney Docket No. 71951-711.602 failure, prediction of myopia, prediction of vision in anti-VEGF treatment, or any combination thereof.
[0104] In some embodiments, the set of data related to the subject comprises longitudinal EHR data of the subject. In some embodiments, at least part of the set of data related to the subject is not collected in association with generating a medical diagnosis, prediction, and / or prognosis of an oncologic indication. In some embodiments, the set of data related to the subject comprises routine laboratory results, vital signs, and routine imaging data. In some embodiments, the set of data related to the subject comprises: results of any one or more of the biomarkers, e.g., complete blood count (CBC), blood chemistry, coagulation, or any combination thereof; heart rate, blood pressure, body mass index (BMI), or any combination thereof; and chest X-rays (CXRs).
[0105] In some embodiments, provided herein is a method of generating a trained model (e.g., LLM), comprising (i) training a model using longitudinal multimodal data of a plurality of subjects, wherein for each subject, the longitudinal multimodal data comprise data from a chronological sequence of clinical visits of the subject, wherein the model comprises modality-specific alignment encoders, a decoder for cross-modal latent space prediction, and an autoregressive module, and wherein the pretraining comprises: (i) adversarially aligning the longitudinal multimodal data using the modality-specific alignment encoders, wherein each modality-specific alignment encoder embeds data of the corresponding modality into a unified, low-dimensional latent space, and wherein a modality discriminator identifies embedding sources in the latent space, (ii) using the decoder to predict a representation of a target modality in the latent space, wherein data from one or more clinical visits of a subject are missing in the target modality, and wherein the prediction imputes the missing data based on the subject’s historical data of the target modality and the subject’s historical data of one or more source modalities, thereby generating modalityinvariant representations of the subjects in the latent space; and (ii) feeding the modalityinvariant representations into the autoregressive module to model each subject's historical trajectory, thereby generating the trained model.
[0106] In some embodiments, provided herein is a method of generating a trained model (e.g., LLM), comprising: pretraining a model using longitudinal multimodal data of a plurality of subjects, wherein for each subject, the longitudinal multimodal data comprise data from a chronological sequence of clinical visits of the subject, wherein the model comprises modality-specific alignment encoders, a decoder for cross-modal latent space prediction, and an autoregressive module, and wherein the pretraining comprises: (i)Attorney Docket No. 71951-711.602 adversarially aligning the longitudinal multimodal data using the modality-specific alignment encoders, wherein each modality-specific alignment encoder embeds data of the corresponding modality into a unified, low-dimensional latent space, and wherein a modality discriminator identifies embedding sources in the latent space, (ii) using the decoder to predict a representation of a target modality in the latent space, wherein data from one or more clinical visits of a subject are missing in the target modality, and wherein the prediction imputes the missing data based on the subject’s historical data of the target modality and the subject’s historical data of one or more source modalities, thereby generating modalityinvariant representations of the subjects in the latent space. In some embodiments, the method further comprises finetuning the pretrained model by feeding the modality -invariant representations into the autoregressive module to model each subject's historical trajectory, thereby generating the trained model.
[0107] In some embodiments, the multimodal data comprise laboratory tests, imaging scans, and / or omics profiles. In some embodiments, the method further comprises feeding the output of the autoregressive module into a task-specific prediction head for a target application.
[0108] In some embodiments, the target application is prediction of a systemic disease, prediction of a chronic disease, prediction of an ocular disease such as myopia, and / or identification of aging trajectories of the subjects. In some embodiments, the target application is prediction of a disease selected from the group consisting of atrial fibrillation, coronary artery disease, diabetes, hypertension, ischemic stroke, multiple sclerosis, osteoporosis, Parkinson’s disease, and rheumatoid arthritis.
[0109] In some embodiments, the longitudinal multimodal data comprise fundus images, routine laboratory test results. In some embodiments, the longitudinal multimodal data comprise electronic health record (EHR) data, vital signs, imaging data, or any combination thereof. In some embodiments, the routine laboratory results comprise complete blood count (CBC), blood chemistry, coagulation, or any combination thereof. In some embodiments, the vital signs comprise heart rate, blood pressure, body mass index (BMI), or any combination thereof. In some embodiments, the imaging data comprise chest X-rays (CXRs).
[0110] In some embodiments, provided herein is a method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising receiving a natural -language prompt for obtaining the medical diagnosis, prediction, and / or prognosis and a set of data related to the subject, and generating the medical diagnosis, prediction,Attorney Docket No. 71951-711.602 and / or prognosis by inputting the prompt and the set of data in the trained model disclosed herein. In some embodiments, the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof. In some embodiments, the disease comprises one or more cancers.
[0111] In some embodiments, provided herein is a system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method disclosed herein. In some embodiments, provided herein is a non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method disclosed herein. In some embodiments, provided herein is a system comprising: at least one hardware processor; non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; and instructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method disclosed herein.Definitions
[0112] As used herein, the singular forms “a,” “an,” and “the” include the plural reference unless the context clearly dictates otherwise.
[0113] Reference to “about” a value or parameter herein includes (and describes) variations that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X”.
[0114] It is understood that aspects and variations of the present disclosure include “consisting” and / or “consisting essentially of’ aspects and variations.
[0115] When a range of values is provided, it is to be understood that each intervening value between the upper and lower limit of that range, and any other stated or intervening value in that states range, is encompassed within the scope of the present disclosure. Where the stated range includes upper or lower limits, ranges excluding either of those included limits are also included in the present disclosure.
[0116] The section headings used herein are for organization purposes only and are not to be construed as limiting the subject matter described. The description is presented to enable one of ordinary skill in the art to make and use an invention in the present disclosure and is provided in the context of a patent application and its requirements. VariousAttorney Docket No. 71951-711.602 modifications to the described embodiments will be readily apparent to those persons skilled in the art and the generic principles herein may be applied to other embodiments. Thus, the present disclosure is not intended to be limited to the embodiment shown but is to be accorded the widest scope consistent with the principles and features described herein.
[0117] The figures illustrate processes according to various embodiments. In the exemplary processes, some blocks are, optionally, combined, the order of some blocks is, optionally, changed, and some blocks are, optionally, omitted. In some examples, additional steps may be performed in combination with the exemplary processes. Accordingly, the operations as illustrated (and described in greater detail below) are exemplary by nature and, as such, should not be viewed as limiting.
[0118] The disclosures of all publications, patents, and patent applications referred to herein are each hereby incorporated by reference in their entireties. To the extent that any reference incorporated by reference conflicts with the instant disclosure, the instant disclosure shall control.Illustrative Embodiments
[0119] The following is a list of non-limiting illustrative embodiments disclosed herein:
[0120] Illustrative Embodiment 1 : A computer-implemented method, comprising: (a) training a latent stable diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, thereby generating a trained latent stable diffusion model; (b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; and (c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality.
[0121] Illustrative Embodiment 2: The method of Illustrative Embodiment 1, wherein in (a), the modality information and the corresponding textual description information are separately encoded.
[0122] Illustrative Embodiment 3: The method of Illustrative Embodiment 1 or Illustrative Embodiment 2, wherein in (a), the modality information and / or the corresponding textual description information are encoded using a BERT tokenizer.Attorney Docket No. 71951-711.602
[0123] Illustrative Embodiment 4: The method of any one of Illustrative Embodiments 1-3, wherein the training in (a) comprises introducing a series of random Gaussian noises to the training image input progressively.
[0124] Illustrative Embodiment 5: The method of Illustrative Embodiment 4, wherein the training in (a) comprises denoising using a U-Net architecture with a cross-attention mechanism.
[0125] Illustrative Embodiment 6: The method of Illustrative Embodiment 5, wherein the cross-attention mechanism comprises cross-attention mechanisms on both the modality information and the corresponding description information.
[0126] Illustrative Embodiment 7: The method of Illustrative Embodiment 5 or Illustrative Embodiment 6, wherein the U-Net architecture comprises multiple U-Net layers.
[0127] Illustrative Embodiment 8: The method of Illustrative Embodiment 7, wherein the U-Net architecture comprises one or more shallow U-Net layers with cross-attention between image embedding and modality information embeddings.
[0128] Illustrative Embodiment 9: The method of Illustrative Embodiment 7 or Illustrative Embodiment 8, wherein the U-Net architecture comprises one or more deep U- Net layers with cross-attention between image embedding and description information embeddings.
[0129] Illustrative Embodiment 10: The method of any one of Illustrative Embodiments 1-9, wherein the processing in (c) comprises an iterative denoising process with sequential cross-attention.
[0130] Illustrative Embodiment 11 : The method of any one of Illustrative Embodiments 1-10, wherein the processing in (c) comprises incorporating modality-specific information followed by refining with detailed descriptive information.
[0131] Illustrative Embodiment 12: The method of any one of Illustrative Embodiments 1-11, wherein the trained latent stable diffusion model learns to distinguish between general modality characteristics and specific image details in descriptive texts.
[0132] Illustrative Embodiment 13: The method of any one of Illustrative Embodiments 1-12, wherein the training in (a) comprises using pre-trained U-Net parameters sourced from a general domain and fine-tuning with a plurality of distinct modalities of medical images, each paired with its corresponding textual description.
[0133] Illustrative Embodiment 14: The method of any one of Illustrative Embodiments 1-13, wherein the processing in (c) comprises using a classifier-free guidance scale and a noise scheduler involving a pseudo-numerical method for diffusion.Attorney Docket No. 71951-711.602
[0134] Illustrative Embodiment 15: The method of any one of Illustrative Embodiments 1-14, comprising a reinforcement learning strategy using human ratings.
[0135] Illustrative Embodiment 16: The method of Illustrative Embodiment 15, wherein the reinforcement learning strategy comprises passing the synthetic medical image to a classification module to generate a classification result for the synthetic medical image.
[0136] Illustrative Embodiment 17: The method of Illustrative Embodiment 16, wherein the reinforcement learning strategy further comprises feeding the synthetic medical image to the classification module which decides whether the synthetic medical image should pass based on the classification result.
[0137] Illustrative Embodiment 18: The method of any one of Illustrative Embodiments 15-17, wherein the reinforcement learning strategy comprises a closed-loop, sustainable self-evolution mechanism which directly integrates human expert clinical judgment into an iterative optimization process.
[0138] Illustrative Embodiment 19: The method of Illustrative Embodiment 18, wherein the iterative optimization process comprises presenting the synthetic medical image to a clinician through a structured scoring interface, inviting the clinician to provide quantitative scores across multiple dimensions comprising anatomical accuracy and clinical diagnostic value of the synthetic medical image.
[0139] Illustrative Embodiment 20: The method of Illustrative Embodiment 19, wherein the iterative optimization process further comprises reward model training, wherein the quantitative scores serve as training data to train a reward model which learns and simulates clinician preferences to provide scores for the synthetic medical image that correspond to its clinical quality.
[0140] Illustrative Embodiment 21 : The method of Illustrative Embodiment 20, wherein output of the trained reward model is used as reinforcement learning signals to finetune the trained latent stable diffusion model, incentivizing the trained latent stable diffusion model to produce synthetic medical images that achieve higher clinical scores.
[0141] Illustrative Embodiment 22: The method of any one of Illustrative Embodiments 1-21, wherein the training in (a) comprises using multimodal data comprising vital signs, imaging data, omics data on the cellular / molecular level (e.g., Genomic Alterations: Mutations, copy number variations (CNVs), chromosomal rearrangements, scDNA-seq; Epigenetics: DNA methylation, histone modifications, scATAC-seq;Transcriptomics: RNA-seq, gene expression profiles, scRNA-seq; Proteomics: Protein expression, post-translational modifications; Metabolomics: Targeted Metabolomics,Attorney Docket No. 71951-711.602Untargeted Metabolomics, Lipidomics, Fluxomics; Spatial Multi -omics Integration: Spatial Transcriptomics, Spatial Metabolomics, Spatial Proteomics, Spatial Epigenomics), Pathological Data (e.g., Histopathology, Immunohistochemistry (IHC), Cytopathology), Laboratory Data (e.g., Blood Tests, Tumor Markers, Liquid Biopsies), and EHRs electronic health records (EHRs) (e.g., including clinical notes, Demographics, Medical History, Symptoms & Signs, Treatment Records), or any combination thereof, optionally wherein the multimodal data comprise tabular data.
[0142] Illustrative Embodiment 23: The method of any one of Illustrative Embodiments 1-21, comprising cross-domain knowledge transfer and collaborative enhancement learning, optionally wherein the method comprises inputting new domain data from one or more new domains in the trained latent stable diffusion model.
[0143] Illustrative Embodiment 24: The method of Illustrative Embodiment 23, wherein the new domain data comprise: (i) an additional medical image of an image modality that is different from the plurality of different image modalities and / or of an additional organ that is different from the plurality of different organs, and (ii) text input comprising modality information of the additional medical image concatenated with its corresponding textual description information.
[0144] Illustrative Embodiment 25: A method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising providing a mixed dataset comprising (i) a real medical image of the subject, and (ii) the synthetic medical image generated by the computer-implemented method of any one of Illustrative Embodiments 1-24 for the subject, and using the mixed dataset to provide the medical diagnosis, prediction, and / or prognosis for the subject.
[0145] Illustrative Embodiment 26: The method of Illustrative Embodiment 25, wherein the medical diagnosis, prediction, and / or prognosis is for a rare disease or an atypical case of a common disease.
[0146] Illustrative Embodiment 27: The method of Illustrative Embodiment 25 or Illustrative Embodiment 26, wherein the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof.
[0147] Illustrative Embodiment 28: The method of any one of Illustrative Embodiments 25-27, wherein the medical diagnosis, prediction, and / or prognosis is for one or more cancers.Attorney Docket No. 71951-711.602
[0148] Illustrative Embodiment 29: The method of any one of Illustrative Embodiments 25-28, comprising training a medical diagnosis, prediction, and / or prognosis model with the mixed dataset as input.
[0149] Illustrative Embodiment 30: The method of Illustrative Embodiment 29, wherein the medical diagnosis, prediction, and / or prognosis model comprises a Transformer classification model.
[0150] Illustrative Embodiment 31 : The method of any one of Illustrative Embodiments 1-30, wherein the plurality of different image modalities comprises optical clearance tomography (OCT), computed tomography (CT), X-ray, fundus photography, magnetic resonance imaging (MRI), ultrasound image, endoscopy image, positron emission tomography (PET), single photon emission computed tomography (SPECT), microscopy image, medical photography, elastography image, a thermogram image, or any combination thereof.
[0151] Illustrative Embodiment 32: The method of any one of Illustrative Embodiments 1-31, wherein the plurality of different organs comprises eye, lung, breast, intestine, brain, kidney, pancreas, bladder, or any combination thereof.
[0152] Illustrative Embodiment 33: A computer-implemented method, comprising: (a) training a latent stable diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, wherein the training comprises denoising using one or more shallow U-Net layers with cross-attention between image embedding and modality information embeddings and one or more deep U-Net layers with cross-attention between image embedding and description information embeddings, thereby generating a trained latent stable diffusion model; (b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; (c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality; (d) presenting the synthetic medical image to a clinician through a structured scoring interface, and (e) inviting the clinician to provide quantitative scores across multiple dimensions comprising anatomical accuracy and clinical diagnostic value of the synthetic medical image, wherein the quantitative scores serve as training data to train a reward model which learns and simulates clinician preferences to provide scores for the synthetic medical image that correspond to its clinical quality.Attorney Docket No. 71951-711.602
[0153] Illustrative Embodiment 34: A computer-implemented method, comprising: (a) training a latent stable diffusion model using: (i) training image input comprising medical , in} and their corresponding textual descriptions D =cross a plurality of different image modalities M = {m^ m2, m3, ••• , mk} and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input [mk; dn] comprising modality information of each medical image concatenated with the corresponding textual description information, thereby generating a trained latent stable diffusion model, wherein the training comprises (i) introducing a series of random Gaussian noises to the training image input progressively, following the equation q(it|it-i)is a linear noise scheduler defined as / 3t= pmin+ t ■an(j iTispurenoise input, and(ii) reversing diffusion using a U-Net architecture with cross-attention mechanisms on both modality information and description information: p0(it-1|it, £,T) =t ) where m is predicted by the output of the cross-attention between image embedding and either modality information embedding or description information embedding; (b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; and (c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality, wherein the processing comprises an iterative denoising process with sequential cross-attention comprising incorporating modality-specific information followed by refining with detailed descriptive information.
[0154] Illustrative Embodiment 35: A system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method of any one of Illustrative Embodiments 1-34.
[0155] Illustrative Embodiment 36: A non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method of any one of Illustrative Embodiments 1-34.
[0156] Illustrative Embodiment 37: A system comprising: at least one hardware processor; non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; and instructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method of any one of Illustrative Embodiments 1-34.Attorney Docket No. 71951-711.602EXAMPLES
[0157] The examples provided herein are included for illustrative purposes only and are not intended to limit the scope of the present disclosure.Example 1: A Medical Image-Text Generative Model
[0158] This example shows the development and application of a versatile and generalizable medical image-text generative model. To develop the model, a range of medical imaging modalities, including OCT, fundus, chest X-ray, and chest CT (Table 1), paired with corresponding textual descriptions as training data, were used. The model in this example operates through two major phases: an initial developmental phase that integrates all medical images and text descriptions into a stable diffusion model, and a deployment phase, where synthetic images are generated based on textual inputs (FIG. 1A). To comprehensively assess the model’s capabilities, a multi-faceted evaluation was conducted: 1) the quality of synthetic images was evaluated using both objective metrics and subjective assessments by clinicians; 2) downstream applications of the synthetic images were explored in diagnostics, report generation, and self-supervised learning (FIG. IB); 3) self-improvement strategies for the model was implemented with reinforcement learning on human feedback (RLHF) and transfer learning (FIG. 1A); 4) the model was applied to clinical tasks in detecting mutation and survival analysis (FIGS. 9A-9B, FIGS. 10A-10D, FIGS. 11A-11D, and FIG. 12).
[0159] To test the quality of synthetic images generated by the model, a comprehensive set of assessments was conducted, including both subjective and objective assessments. Subjective evaluations by clinicians were followed by objective metrics, including Frechet Inception Distance (FID), Inception Score (IS), and Multi-Scale Structural Similarity Index (MS-SSIM) to measure fidelity, diversity, and quality. Additionally, the Classification Accuracy Score (CAS), zero-shot image-image retrieval (IIR), and image-text retrieval (ITR) methods were incorporated into the assessment to assess the practical utility and relevancy of synthetic images.
[0160] All evaluations were performed on OCT, fundus, chest X-ray, and chest CT.Table 1. Data characteristics of the cohorts used in this study.Attorney Docket No. 71951-711.602RD: Refractive Disorders; AS: Amblyopia and Strabismus; GL (Glaucoma), ON (Optic Neuropathy), DR: Diabetic Retinopathy); ME: Macular Edema; AMD: Age-related Macular Degeneration; IAOD: Inherited and Autoimmune Ocular Diseases.Dataset descriptions
[0161] OCT data and reports were collected from multiple medical institutes to build an extensive dataset shown in Table 1. Image-report pairs from China Eye Image Investigation were combined and the data were randomly split based on patient identifications to build the training and validation sets (11,438 data pairs). Data from an eye hospital were used for external validation (2,287 data pairs). In addition to OCT images with reports from public datasets, a high-quality dataset comprising 13,725 paired image-text OCT samples was assembled.Attorney Docket No. 71951-711.602
[0162] The retinal fundus images were collected were collected, as shown in Table 1, including a Northern China cohort and a Southern China cohort. A total of 59,420 image-text pairs participated in the study.
[0163] The chest CT and X-ray images were collected, as shown in Table 1. For each patient with his / her medical record, 10 CT images and 10 X-ray images were selected, resulting in 74,900 paired data for 3,745 patients.
[0164] The brain MRI dataset was collected from multiple centers, including 1960 patients from 16 hospitals. These patients collectively contributed a total of 30,800 tumor- annotated images, including 13,859 Tl-channel images and 16,941 T2-channel images.
[0165] The breast tumor imaging dataset encompassed 3,694 patients diagnosed with breast MR images. These patients provided a total of 34,506 annotated tumor images and 35,000 normal breast images, which included 16017 Tl-channel images, 6129 T2-channel images and 47360 contrast-enhanced Tl-channel images. Notably, 3,200 images contain information on the HER2 gene, which is specifically utilized to train the model to associate HER2 gene expression with MR imaging characteristics. The remaining images serve as an additional dataset to train the stable diffusion model, ensuring robust and reliable performance.MINIM's framework and implementation details
[0166] MINIM is a framework designed for the generation of medical images through a diffusion model, leveraging paired medical images I =, in} and their corresponding textual descriptions D =dn} across different medical image modalities M = {m^ m2, m3, ••• , mk}. To facilitate the accurate generation of specific modality images, the modality information mkwas concatenated with the description information dnto generate the text input [mk; dn] of MINIM. During the training phase, MINIM employs a BERT tokenizer EBERTto encode both the modality information and description information separately. The framework introduces a series of random Gaussian noises to the input image progressively, following the equation:where / 3tis a linear noise scheduler defined as / 3t=nin+ t ■ ^max^min^an(j iTispurenoise input. Then MINIM learns to reverse the diffusion process using a U-Net architecture with cross-attention mechanisms on both modality information and description information:Attorney Docket No. 71951-711.602 where the 1Q is predicted by the output of the cross-attention between input texts and images, which incorporates cross-attention mechanisms at different U-Net layers. In the shallow layers of the U-Net encoder and decoder, MINIM applies cross-attention between image embedding U shaiiow^St) and modality information embeddings EBERT(m), and in the deep layers of the U-Net encoder and decoder, MINIM applies cross-attention between image embedding and description information embeddings EBERT(d). The cross-attention function is defined as:Cross Attention Q, K, 7) = softmax( -j=~)V , where Q is derived from the U-Net features, and K and V are derived from the corresponding text embeddings (EBERT(m) or EBERT(d ).
[0167] During inference, the generation of images occurs through an iterative denoising process starting from random noise iT. it-i = Ee t’ t. ET) + otz, where z~ / V(0, 1) and <Jt2= / 3t. This sequential cross-attention approach allows MINIM to first incorporate modality-specific information and then refine the features with more detailed descriptive information. By concatenating modality and description information, the model can learn to distinguish between general modality characteristics and specific image details described in the text, potentially leading to more accurate and controllable medical image generation.
[0168] To optimize MINIM for generating medical images, pre-trained U-Net parameters sourced from the general domain were utilized and the model was fine-tuned with four distinct modalities of medical images, each paired with its corresponding textual description. The fine-tuning experiments were conducted on a high-performance computing cluster equipped with 8 NVIDIA Al 00 GPUs. The proficiency in generating medical images was optimized by adjusting hyperparameters, initializations, and epochs. All the input images are rescaled at a resolution of 512^512 pixels and processed in batches of 128. In this configuration, fine-tuning a model required approximately 10 hours for Ik training steps, and 4 days for 10k steps. The U-Net model weights for the SD pipeline (version 1.4) and CLIP (ViT-large) were obtained from HuggingFace. The “transformers” and “diffusers” libraries were utilized in the implementations. The built-in safety checker was disabled during the generation process due to its high false-positive rate in response to medical prompts. For synthesizing medical images, a classifier-free guidance scale set at 4.0 was employed and theAttorney Docket No. 71951-711.602 inference steps were set at 100, utilizing the default noise scheduler, which involves the pseudo-numerical methods for diffusion models. The best-performing model was selected as the generative model of choice (FIG. 1A).Performance evaluations
[0169] As detailed in Examples 2-5 below, on average, MINIM enhances performance by 12% for ophthalmic, 15% for chest, 13% for brain and 17% for breast-related tasks. Furthermore, MINIM is useful in the accurate prediction of Zffi7?2-positive breast cancer from MRI images. Using a large retrospective simulation analysis, it is demonstrated that MINIM accurately identified targeted therapy-sensitive EGFR mutations using lung cancer computed tomography images, which could potentially lead to improved 5 -year survival rates.
[0170] In the examples herein, MINIM is a unified medical image-text generative model that integrates medical images paired with textual descriptions across various modalities and organs, including optical coherence tomography (OCT), fundus, chest X-ray and chest computed tomography (CT). MINIM can generate high-quality synthetic images for each organ and imaging modality based on textual descriptions.
[0171] MINIM’s performance in generating synthetic images based on textual inputs was evaluated, and its performance was tested across different medical imaging scenarios against other recent generative models. The model’s adaptability was demonstrated by incorporating additional magnetic resonance imaging (MRI) datasets for brain and breast imaging, allowing the assessment of MINIM’s ability to continuously learn and integrate new medical knowledge.Example 2: Subjective and Objective Assessment of Synthetic ImagesSubjective assessment
[0172] Clinicians specific for corresponding image modalities were invited to assess whether synthesized images are clinically accurate and useful (FIG. 2A). Clinicians rated the synthetic images on a scale of 1 to 3; a score of 1 denoted a low-quality image, 2 indicated a high-quality image but irrelevant to the report, and 3 signified a high-quality image that aligned well with the report. Three rounds of ratings were conducted, and the results were summarized by calculating the percentage of images that scored 3. In the first round, MINIM achieved an average of 70.75% (73% for OCT, 75% for fundus, 61% for chest CT, and 74% for chest X-ray), while the best competing generative models (StyleGAN-T) achieved anAttorney Docket No. 71951-711.602 average of 65.75% (62% for OCT, 78% for fundus, 58% for chest CT, and 65% for chest X- ray). In the third round, after implementing reinforcement learning, MINIM's performance showed significant improvement, achieving an average of 89.25% (91% for OCT, 95% for fundus, 83% for chest CT, and 88% for chest X-ray) (FIG. 2A).Image grading from clinicians
[0173] OCT and fundus images went through a tiered grading system consisting of multiple layers of trained graders of increasing expertise for verification and correction of image labels. Each image imported into the database started with a label matching the most recent diagnosis of the patient and examined each paired report. The first tier of graders consisted of medical students and ophthalmology residents who had taken and passed an interpretation course review. This first tier of graders conducted initial quality control on real patient images and excluded images containing severe artifacts or significant image resolution reductions. The second tier of graders consisted of five ophthalmologists who had at least 10 years of retinal subspecialty practice and they independently graded each image and corresponding report. The presence or absence of images reflecting features of vitreous, retinal, RPE and choroid, and various retinal disease diagnoses including choroidal neovascularization (active or in the form of subretinal fibrosis), macular edema, drusen, RVOs, epiretinal membrane, and other pathologies visible on the OCT scan were recorded. Finally, a third tier of two senior independent retinal specialists, each with over 20 years of clinical retina experience, verified the true labels and corresponding reports for each image. To account for human error in grading, a validation subset of 993 scans was graded separately by two ophthalmologist graders, with disagreement in clinical labels arbitrated by a senior retinal specialist. For chest X-ray, chest CT, brain MRI and breast MR images, all modalities were initially screened for quality control by removing all low-quality or unreadable scans. Similar to previous ophthalmological image grading, the diagnoses here then were graded by three radiology expert physicians who have at least 10 years of practice experience. In order to account for any grading errors, the evaluation set was also checked by another expert.
[0174] In assessing Al-based EGFR mutation status predictions, two independent retrospective cohorts with advanced lung cancer patients (Stage III or Stage IV) were employed. Patients from a period up to 2014 were used for training. Consecutive patients from 2015-2020 were used for a five-year survival evaluation (Cohort A and Cohort B from different hospitals). Kaplan-Meier survival curves were used to compare survival ratesAttorney Docket No. 71951-711.602 between EGFR-sensitive mutation patients versus EGFA-insensitive mutation patients. All patients received standard chemotherapy plus TKI therapy if they had undergone a tumor tissue biopsy and EGFR gene sequencing to determine the EGFR mutation status. An evaluation committee consisting of three senior oncologists gave blinded and deidentified Cohort A and Cohort B patients to a computer scientist for Al -based EGFR mutation status prediction.Objective assessment
[0175] Objective metrics were also employed to quantitatively assess the fidelity, diversity, and overall quality of the generated images. These images should show robust diversity to encompass a wide range of pathological variations without overfitting to the training data, while also maintaining high fidelity to real images, as reflected in the low FID, low MS-SSIM, and high IS scores (FIG. 2B). FID was calculated for each synthetic-real image pair, while IS and MS-SSIM were computed independently among synthetic images. The average results of MINIM and other competitive text-to-image generative methods (FIG. 2C) were reported. The findings revealed that MINIM’s performance (FID: 65.3, IS: 5.7±0.42, MS-SSIM: 0.16±0.03 for OCT; FID: 32.7, IS: 6.7±0.43, MS-SSIM: 0.11±0.03 for fundus; FID: 110.4, IS: 5.7±0.37, MS-SSIM: 0.19±0.06 for chest X-ray; FID: 94.8, IS: 5.7±0.38, MS-SSIM: 0.17±0.04 for chest X-ray) (Table 2) outperformed other generative methods by a large margin. Samples of MINIM' s synthetic images along with other generative methods indicate that MINIM can produce more realistic images in all settings (FIG. 2D)
[0176] The classification accuracy score (CAS) is an emerging proxy for the evaluation of synthetic data. CAS measures how accurately models trained exclusively on synthetic samples can classify real test data. For each of the four modalities and organs, MINIM was used to generate synthetic images 10 times the real dataset for each class label and reported CAS on the test set using 1 / 10 random images of each real dataset. FIG. 2C shows MINIM outperformed other generative methods for top-1 accuracy across different datasets (79.09% for OCT, 86.16% for fundus, 79.42% for chest CT and 77.23% for chest X- ray).Attorney Docket No. 71951-711.602Table 2. Ablation study on quantitative assessment of the synthetic image using metrics
[0177] To further validate the quality of the synthetic images from MINIM, a zeroshot IIR method was employed. Using real images as queries, this evaluation aimed to retrieve candidate synthetic images that are closely related in category to the query image. Both the query and candidate images were input into a pretrained image encoder, specific to their imaging modality, to rank the candidate images based on feature similarity to the query image using cosine similarities. The number of synthetic images that share the same category was reported as queried real images among the top ten candidates, using Precision@k metrics where k = 10, for each query. FIG. 2C demonstrated the superior results (62.25% for OCT, 64.83% for fundus, 59.44% for chest CT, and 53.17% for chest X-ray) from MINIM to other generative models in each case.Attorney Docket No. 71951-711.602
[0178] To evaluate whether synthesized images correspond to clinically relevant and accurate texts, zero-shot ITR was performed. In this approach, a query image embedding is mapped into a textual embedding space to retrieve the most relevant textual impression corresponding to the image. Retrieval precision was calculated using the Precision@k with k = 10, where a point of accuracy is awarded for each textual impression that shares the same category as the query image. FIG. 2C demonstrated the superior results (43.41% for OCT, 49.93% for fundus, 37.53% for chest CT, and 42.04% for chest X-ray) from MINIM compared to other generative models.Comparative methods
[0179] MINIM was compared with the latest text-to-image generative models in the Al community, including Imagen, DALLE, GigaGAN and StyleGAN-T. Pre-trained models on their released checkpoints were fine-tuned on the dataset. Hyperparameters were employed following the instructions in their original papers. An ablation study was performed to select the optimal hyperparameters for MINIM. The quality of the synthetic images was assessed using the average of three metrics (FID, IS, and MS-SSIM) on all datasets (Table 2). The results demonstrate that, as the number of training steps increased, the quality of synthetic images improved substantially, with FID decreasing to 57.91, MS- SSIM improving to 0.18, and IS increasing to 5.89 at 20,000 training steps, which was the choice in further experiments. The impact of different model initializations and image resolutions was also examined. Randomized weight initialization of U-Net and CLIP models led to poorer performance, as evidenced by higher FID scores. In addition, generating lower- resolution images (256x256) resulted in worse generative quality. In summary, the optimal performance was achieved by the full MINIM model, initialized with pre-trained weights and trained on 512x512 image resolution for 20,000 steps. In summary, the optimal performance was achieved by the full MINIM model, initialized with pre-trained weights and trained on 512 x 512 image resolution for 20,000 steps. This configuration showed the best balance of the three metrics.Example 3: Self-Improving StrategiesReinforcement learning from clinicians ’ ratings
[0180] The RLHF strategy utilizes a synthetic-in-the-loop approach to iteratively enhance MINIM’ s performance, which involves leveraging clinicians’ ratings, or feedback, to create a closed loop of continuous learning. In this example, the two-stage RLHF strategyAttorney Docket No. 71951-711.602 began with a panel of clinicians who assessed the quality of synthetic images in relation to their corresponding prompt text reports (FIG. 3A). Clinician feedback was then used to train a reward model that mimics the clinicians’ ratings. This reward model was incorporated into MINIM, serving as an additional layer of supervision to guide the model in training and improve overall performance. How RLHF incorporation improved image generation was compared with four other data augmentation methods across six subjective metrics(FIG. 3B). Notably, all five methods, including MINIM, exhibited improvements over their original generative models, with the two-stage RLHF strategy showcasing the most significant enhancements (e.g., for OCT, FID down to 43.3 vs. 65.3; IS up to 8.7±0.37 vs. 5.7±0.42; MS-SSIM down to 0.11±0.06 vs. 0.16±0.03). These results underscore the value of integrating human feedback into the reinforcement learning process, significantly refining the generative model's performance across multiple metrics.Reinforcement learning implementation details
[0181] A two-stage reinforcement learning with human feedback (RLHF) framework that employs synthetic data was used to iteratively train an agent. The framework is designed to harness the advantages of synthetic images and radiologist ratings, allowing for the generation of high-quality data, which is especially beneficial in situations where real -world data is scarce, expensive, or risky to obtain. The first stage involves the generation of synthetic images from MINIM model. The generated images are then passed to a receive selector, which operates as a classification module. In this stage, the classification module is designed to let all instances pass through without selection. Subsequently, human raters evaluate the synthetic data, and their assessments are encoded as ratings from 1 to 3, where 1 stands for ‘The image quality is too low.’, 2 stands for ‘The image quality meets the standard, but it does not correspond with the report.’, and 3 stands for ‘The image quality meets the standard and is generally consistent with the report.’. These ratings serve to update the RL policy, thereby closing the loop of this stage. The RL policy TT is updated according to the following equation:where R(s, a) denotes the rating for the state-action pair (s, a) and a is the learning rate.
[0182] The second stage introduces an active selection process. Again, synthetic images are fed into the system, but this time the receive selector classification module decides which images to pass based on the classification results. Only the selected images areAttorney Docket No. 71951-711.602 forwarded to the human raters for evaluation. The non-selected instances are discarded. The ratings from the human raters update the RL policy, similar to the first stage but with a different learning rate. The selective filtering focuses the policy update on higher quality images, which is crucial for improving learning efficiency and effectiveness. The updated policy is represented by:where / 3 is a different learning rate from a, reflecting the different data distributions between the two stages. Pseiected xi') is the probability that the instance is selected by the 'Receiveis the probability of taking action in state s under the current policy TT, and is the value function under policy n for state s. In this formulation, the term R(s, a) — l^(s) represents the temporal difference error, a measure of how the outcome was compared to the agent's expectations.Transfer Learning
[0183] A major advantage of MINIM over previous specialized models is its ability to leverage image-text paired medical data from various organs and imaging modalities.Following the inclusion of breast and brain MRI datasets in a new round of development, the performance of previously trained modalities was evaluated (FIG. 4A). Improvements across six modalities were observed, with metrics such as FID down to 51.3 from 65.3, IS increasing from 5.7±0.42 to 6.7±0.48, and MS-SSIM improving from 0.14±0.03 to 0.16±0.03. These findings highlight MINIM’ s continuous learning potential, driven by the ongoing integration of new medical data.
[0184] The objective evaluation for the synthetic brain and breast MR images (FIG. 4B) showed that MINIM can generate high-quality synthetic images in these two domains when mixing these two-modality paired datasets with existing four-modality paired datasets. A comparison of synthetic images generated by MINIM is presented alongside those produced by other generative models, all based on the same textual instructions (FIG. 4C), where it was observed that MINIM can generate images with enhanced realism and fidelity. More examples conditioned on different descriptions are shown in FIG. 5A and FIG. 5B.Attorney Docket No. 71951-711.602Example 4: Downstream Applications Using Synthetic ImagesDiagnosis
[0185] Adding synthetic data to the training dataset significantly enhances medical diagnosis performance. Multi-class diagnosis models were trained using a Swin Transformer classifier with diagnostic labels for each modality. Results from models trained on real datasets were compared with those trained on mixed datasets consisting various proportions of synthetic images. The baseline top-1 classification accuracy using real data exclusively was 0.64 for OCT, 0.74 for fundus, 0.58 for Chest CT, and 0.59 for Chest X-ray. Notably, with a 1 : 1 ratio of real to synthetic data, top-1 classification accuracy improved to 0.73 for OCT, 0.81 for fundus, 0.65 for Chest CT, and 0.69 for Chest X-ray. By increasing the proportion of synthetic data, parallel improvements in top-1 classification accuracy were observed, ultimately reaching average accuracies of 0.93 for OCT, 0.95 for fundus, 0.79 for Chest CT, and 0.86 for Chest X-ray at a 20: 1 ratio of synthetic to real data (FIG. 6A). To further validate the effectiveness of synthetic data, OCT images for classes with poor diagnostic performance were specifically synthesized. A confusion map visualizing the model’s classification results on real data (FIG. 6B) revealed that the performance for four diagnoses (Macular Hole of Post- Surgery, Choroidal and Retinal Degeneration, Retinal Vasculitis, and Endophthalmitis) was significantly worse than others. 100 tailored OCT images for each of the four diagnoses were subsequently generated and ophthalmologists were invited to further train the model by selecting typical images. This iterative approach led to marked improvements in diagnostic effectiveness (Table 4), illustrating the capacity of synthetic data to augment training datasets and address specific diagnostic challenges.Furthermore, the model successfully diagnosed atypical cases of common diseases, including retinal vein occlusions (RVOs), choroidal neovascularization (CNV), and diabetic retinopathy (DR), demonstrating its ability to recognize a wide spectrum of disease presentations (FIG. 5B) (Table 3).Ablation study
[0186] Table 3 presents a comparison of the classification performance using different configurations in diagnosing various eye conditions using synthetic OCT images.Attorney Docket No. 71951-711.602Table 3: Ablation study quantitative assessment of the synthetic images in a multiclassification task.
[0187] In Table 3, multi-label classification performance on the real OCT images and 5k synthetic OCT images with different settings are reported.
[0188] The table highlights the improvement in accuracy as the number of training steps increases. Initially, the original SD model had a random accuracy of 0.51. After 10,000 training steps, the model achieved significant improvements in accuracy: 0.82 for CNV classification, 0.84 for DME, 0.88 for DRUSEN, and 0.91 for Normal cases. Furthermore, the table investigates the impact of different model initializations and image resolutions on the classification performance. Randomized weight initialization of U-Net and CLIP models, as well as lower resolution images (256x256), led to suboptimal results compared to the fullAttorney Docket No. 71951-711.602MINIM model. The full MINIM model, trained for 20,000 steps, exhibited an exceptional average accuracy of 0.94, closely matching its performance on real data. This demonstrates the robustness and reliability of the MINIM model in generating high-quality synthetic images suitable for accurate multi -class diagnosis.Table 4. Multi-class diagnosis performance using real OCT images and mixed data of real and synthetic OCT images.Attorney Docket No. 71951-711.602Report Generation
[0189] For report generation, text descriptions that appeared more than 10 times in the real dataset were curated, resulting in 60 image reports for each modality. Synthetic images were generated for each report, creating a controlled and more balanced synthetic dataset compared to the real dataset collection process. Using the CLIP+GPT-2 framework, a report generation model was trained for each modality and its outputs were evaluated against clinician-provided annotations using metrics such as BLEU-n, CIDEr, and Rouge-L scores (FIGS. 7A-7C). Consistent with the diagnostic findings, the combination of real and synthetic data (20: 1) improved performance, yielding ROUGE-L scores of 38.6 for OCT, 46.3 for fundus, 35.6 for Chest CT and 43.0 for Chest Xray, compared to scores of 26.4, 29.2, 21.3, and 22.6, respectively, for real data alone. FIG. 5B presents selected examples of the generated OCT reports, showcasing the model's ability to capture disease information that may be overlooked by ophthalmologists. Moreover, these examples illustrate the potential of synthetic data not only as an augmentation strategy but also as a tool for enriching and enhancing the completeness of medical imaging reports.Self-supervised learning
[0190] Self-supervised learning allows the model to derive meaningful representations from unlabeled data by pre-training on auxiliary targets, a method particularly beneficial in medical imaging, where unlabeled data is abundant, and labeling is costly. To assess the effectiveness of the model’s synthetic images in self-supervised pre-training, images from different modalities were augmented and Siamese network branches were employed to learn invariant features. The DenseNet-121 model was pre-trained on auxiliary tasks, which involved fine-tuning only the classification layer from 100% to 2000% of the available training data. The results indicate that models fine-tuned with both synthetic and real images consistently outperformed those fine-tuned exclusively on real images, with top-1 accuracy improving significantly, achieving 79.7 for OCT, 86.0 for fundus, 78.0 for Chest CT, and 76.1 for Chest X-ray, compared to the initial accuracies of 54.7, 63.1, 57.8, and 51.4, respectively (FIG. 8).Implementation details for downstream analysis
[0191] To evaluate the effectiveness of synthetic data in diagnosis, the category labels that appeared more than five times in the training set were selected. The Swin-Transf ormer model was then used to train a multi-class classification diagnostic model on theAttorney Docket No. 71951-711.602 corresponding images and labels. First, the synthetic images of an equal number to the training images were generated using fine-tuned MINIM for each text description in the training dataset, and they were mapped to the corresponding diagnosis of the text. These images were then incorporated as additional inputs into the Swin-Transformer classification model during its training phase. For the evaluation of the generated diagnosis results, Accuracy, Fl score, and AUC were used as evaluation metrics. In a further scaling-up experiment, the number of synthetic images generated varied from 100%, 200%, 500%, 1000%, and 2000%, with the results of each variation reported accordingly.
[0192] The experimental settings on report generation using synthetic data for each image modality were as follows. Text descriptions with a frequency exceeding ten were chosen, along with their corresponding images, forming real datasets. These text descriptions were subsequently utilized for generating synthetic images, which served as additional inputs for the CLIP+GPT2-based image captioning model during its training phase. Subsequently, an independent test dataset was employed to produce corresponding reports, which were compared against the ground truth. For the evaluation of the generated reports compared to the ground-truth reports, established natural language generation metrics were employed. These included BLEU with 1 to 4 grams, CIDEr, and ROUGE -L. In a further scaling-up experiment, the number of synthetic images generated varied from 100%, 200%, 500%, 1000%, and 2000%, with the results of each variation reported accordingly.
[0193] In the assessment of self-supervised learning methodologies utilizing synthetic data, a Siamese network architecture was employed. For each image modality, two distinct augmentations of a synthetic image were introduced into separate branches of the network. The self-supervised pre-trained models of DenseNet-121 underwent linear probing, with only a classification layer fine-tuned on 100%, 200%, 500%, 1000%, and 2000% of the training data while keeping other model weights fixed. Top-1 Classification accuracy was reported using a mixture of real data and different numbers of synthetic data and real data only.Example 5: Clinical Applications Using Synthetic ImagesEGFR mutation detection in lung cancer
[0194] Whether synthetic images can improve the top-1 accuracy of EGFR mutation classifications was investigated using chest CT scans (FIG. 9A). A Swin-Transformer was employed to classify real CT images into three EGFR mutation types: wild, sensitive, and resistant. Baseline performance using only real data achieved an average top-1 accuracy of 81.5%. By progressively incorporating synthetic data into the training set, substantialAttorney Docket No. 71951-711.602 improvements in classification accuracy were observed. Remarkably, a 1: 1 ratio of synthetic to real data for training increased classification accuracy to 91.2%, with further gains reaching 95.4% when the synthetic to real data ratio reached 5: 1. Additional increases in synthetic data beyond this point did not yield significant improvements. In a subsequent binary classification task, combining EGFR wild type and resistant type mutations, the inclusion of synthetic data similarly enhanced model performance using the area under the receiver operating characteristic curve (AUROC). AUROC improved from 74.2% with real data alone to 96.5% when synthetic data was expanded fivefold. These trends underscore the value of synthetic data in improving the predictive accuracy of Al models for EGFR mutation detection in lung cancer.
[0195] These results were further validated in a retrospective clinical study using two independent cohorts of patients. The model’s prediction of advanced lung cancer patients, based on EGFR mutations, that could benefit from TKI -targeted therapy, was evaluated. The baseline survival curve of patients who underwent standard chemotherapy (FIG. 9B) were first determined. Then using the model’s predictions, patients with FGFF-sensitive mutations (who received TKI therapy) were identified and their survival rates were compared to those with FGFF-resistant mutations (who received chemotherapy). In the first group of patients (cohort A), FGFF-sensitive mutations had a 29.5% higher 5-year overall survival (OS) (EGFR sensitive mutations, 53.4% [95%CI 47.5% - 60.1%]; EGFR resistant mutations, 23.9% [95%CI 19.3% - 29.5%]) and a longer median survival time (EGFR sensitive mutations, 54.7 months [95%CI 45.4 - NA]; EGFR resistant mutations, 18.8 months [95%CI 16.6-21.1]; FIG. 9B). The 5-year progression-free survival (PFS) and objective response rate (ORR) also improved from 12.2% to 41.6% (FIGS. 10A-10B) and 12.3% to 30.9% (FIGS. 11A-11B), respectively. Similar results were confirmed in the second patient group (cohort B), with a 14.7% increase in 5-year OS (FGFF-sensitive mutations, 41.0% [95%CI 32.2% - 52.2%]; FGFF-resistant mutations, 26.3% [95%CI 21.6% - 32.1%]) and longer median survival time (FGFF-sensitive mutations, 40.6 months [95%CI 36.4 - 63.2]; EGFR- resistant mutations, 18.1 months [95%CI 16.2 - 21.0]; FIG. 9B). Error rates of the model’s predictions for both cohort A (FGFF-sensitive mutations 0.12 (4 / 327), FGFF-resistant mutations 0.018 (6 / 321)) and cohort B (FGFF-sensitive mutations 0.028 (4 / 142), EGFR- resistant mutations 0.04 (6 / 316)) were low. The PFS and ORR also improved from 19.1% to 44.7% (FIGS. 10C-10D) and, 18.9% to 33.1% (FIGS. 11C-11D), respectively.Attorney Docket No. 71951-711.602Implementation details ofEGFR mutation type detection
[0196] 3 -class EGFA-mutation type classification and 2-class classification were performed using a baseline that takes as input all real data, and a series of compared models including all synthetic images of the equal number of the baseline and combinations of real data and different ratios of synthetic images from 10% to 500%. Swin-Transformer was used to achieve an end-to-end architecture that encourages the model to predict a class probability. Training of the models by back-propagation of errors was performed in batches of 32 images resized to pixels for 50 epochs with a learning rate of 1 x 10'4. The training was performed using the Adam optimizer with a weight decay of 1 x 10'5. Transformations of random horizontal and vertical flips were added to each batch during training as data augmentation in order to enable improved and generalized network learning. The models selected for evaluation on test sets were the models with the best validation loss on the training set.HER2 mutation detection in breast cancer
[0197] Given the complexity of HER2 status and its critical role in determining treatment outcomes in HER2 -targeted therapies, this example explores integrating synthetic images from the model could enhance the accuracy of HER2 status classification. A baseline that was trained exclusively on real datasets was first established, achieving a classification accuracy of around 79.2% for distinguishing between three classes (no tumor / tumor with FF / G / tumor without H R2) (FIG. 12). With the progressive addition of synthetic images with text descriptions at ratios of 2, 3, 5, and 10 times the proportion of real images, significant improvements in the classification accuracy were observed, peaking at 94.0% when using a synthetic to real images ratio of 10: 1 (FIG. 12).Implementation details ofHER2 mutation detection
[0198] A three-layer CNN was trained to classify breast MR images into three classes: no tumor, tumor with HER2 mutation and tumor without HER2 mutation. The model comprises three convolutional layers, each followed by a ReLU activation function and a max pooling layer. At the end of the model, there are three linear functions separated by ReLU activations. A basic classifier was first trained on a dataset composed of solely real data (3,200 images). Among these, there were 1,598 images without tumors, 687 images containing tumors with HER2 mutations, and 915 images containing tumors without HER2 mutations. Using the same architecture, a new classifier was then trained on a hybrid dataset consisting of both real and synthetic data. This dataset included all of the aforementioned realAttorney Docket No. 71951-711.602 data, as well as 10 times of synthetic images. Within synthetic images, there were 16,274 images without tumors, 7,582 images containing tumors with HER2 mutations, and 6,292 images containing tumors without HER2 mutations. During training, the Adam optimizer with a loss function of CrossEntropy was employed. The learning rate was set to a fixed value of 0.001 and training included a total of 20 epochs using a batch size of 128. The test results of these two classifiers are depicted in FIG. 12).
[0199] The present disclosure is not intended to be limited in scope to the particular disclosed embodiments, which are provided, for example, to illustrate various aspects of the present disclosure. Various modifications to the compositions and methods described will become apparent from the description and teachings herein. Such variations may be practiced without departing from the true scope and spirit of the disclosure and are intended to fall within the scope of the present disclosure.
Claims
1. Attorney Docket No. 71951-711.602CLAIMS1. A computer-implemented method, comprising:(a) training a latent stable diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, thereby generating a trained latent stable diffusion model;(b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; and(c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality.
2. The method of claim 1, wherein in (a), the modality information and the corresponding textual description information are separately encoded.
3. The method of claim 1 or claim 2, wherein in (a), the modality information and / or the corresponding textual description information are encoded using a BERT tokenizer.
4. The method of any one of claims 1-3, wherein the training in (a) comprises introducing a series of random Gaussian noises to the training image input progressively.
5. The method of claim 4, wherein the training in (a) comprises denoising using a U-Net architecture with a cross-attention mechanism.
6. The method of claim 5, wherein the cross-attention mechanism comprises crossattention mechanisms on both the modality information and the corresponding description information.
7. The method of claim 5 or claim 6, wherein the U-Net architecture comprises multiple U-Net layers.
8. The method of claim 7, wherein the U-Net architecture comprises one or more shallow U-Net layers with cross-attention between image embedding and modality information embeddings.Attorney Docket No. 71951-711.6029. The method of claim 7 or claim 8, wherein the U-Net architecture comprises one or more deep U-Net layers with cross-attention between image embedding and description information embeddings.
10. The method of any one of claims 1-9, wherein the processing in (c) comprises an iterative denoising process with sequential cross-attention.
11. The method of any one of claims 1-10, wherein the processing in (c) comprises incorporating modality-specific information followed by refining with detailed descriptive information.
12. The method of any one of claims 1-11, wherein the trained latent stable diffusion model learns to distinguish between general modality characteristics and specific image details in descriptive texts.
13. The method of any one of claims 1-12, wherein the training in (a) comprises using pre-trained U-Net parameters sourced from a general domain and fine-tuning with a plurality of distinct modalities of medical images, each paired with its corresponding textual description.
14. The method of any one of claims 1-13, wherein the processing in (c) comprises using a classifier-free guidance scale and a noise scheduler involving a pseudo-numerical method for diffusion.
15. The method of any one of claims 1-14, comprising a reinforcement learning strategy using human ratings.
16. The method of claim 15, wherein the reinforcement learning strategy comprises passing the synthetic medical image to a classification module to generate a classification result for the synthetic medical image.
17. The method of claim 16, wherein the reinforcement learning strategy further comprises feeding the synthetic medical image to the classification module which decides whether the synthetic medical image should pass based on the classification result.Attorney Docket No. 71951-711.60218. The method of any one of claims 15-17, wherein the reinforcement learning strategy comprises a closed-loop, sustainable self-evolution mechanism which directly integrates human expert clinical judgment into an iterative optimization process.
19. The method of claim 18, wherein the iterative optimization process comprises presenting the synthetic medical image to a clinician through a structured scoring interface, and inviting the clinician to provide quantitative scores across multiple dimensions comprising anatomical accuracy and clinical diagnostic value of the synthetic medical image.
20. The method of claim 19, wherein the iterative optimization process further comprises reward model training, wherein the quantitative scores serve as training data to train a reward model which learns and simulates clinician preferences to provide scores for the synthetic medical image that correspond to its clinical quality.
21. The method of claim 20, wherein output of the trained reward model is used as reinforcement learning signals to fine-tune the trained latent stable diffusion model, incentivizing the trained latent stable diffusion model to produce synthetic medical images that achieve higher clinical scores.
22. The method of any one of claims 1-21, wherein the training in (a) comprises using multimodal data comprising vital signs, imaging data, omics data on the cellular / molecular level (e.g., Genomic Alterations: Mutations, copy number variations (CNVs), chromosomal rearrangements, scDNA-seq; Epigenetics: DNA methylation, histone modifications, scATAC-seq; Transcriptomics: RNA-seq, gene expression profiles, scRNA-seq; Proteomics: Protein expression, post-translational modifications; Metabolomics: Targeted Metabolomics, Untargeted Metabolomics, Lipidomics, Fluxomics; Spatial Multi -omics Integration: Spatial Transcriptomics, Spatial Metabolomics, Spatial Proteomics, Spatial Epigenomics), Pathological Data (e.g., Histopathology, Immunohistochemistry (IHC), Cytopathology), Laboratory Data (e.g., Blood Tests, Tumor Markers, Liquid Biopsies), and EHRs electronic health records (EHRs) (e.g., including clinical notes, Demographics, Medical History, Symptoms & Signs, Treatment Records), or any combination thereof, optionally wherein the multimodal data comprise tabular data.
23. The method of any one of claims 1-22, comprising cross-domain knowledge transfer and collaborative enhancement learning, optionally wherein the method comprises inputting new domain data from one or more new domains in the trained latent stable diffusion model.Attorney Docket No. 71951-711.60224. The method of claim 23, wherein the new domain data comprise: (i) an additional medical image of an image modality that is different from the plurality of different image modalities and / or of an additional organ that is different from the plurality of different organs, and (ii) text input comprising modality information of the additional medical image concatenated with its corresponding textual description information.
25. A method of generating a medical diagnosis, prediction, and / or prognosis for a subject, the method comprising providing a mixed dataset comprising (i) a real medical image of the subject, and (ii) the synthetic medical image generated by the computer- implemented method of any one of claims 1-24 for the subject, and using the mixed dataset to provide the medical diagnosis, prediction, and / or prognosis for the subject.
26. The method of claim 25, wherein the medical diagnosis, prediction, and / or prognosis is for a rare disease or an atypical case of a common disease.
27. The method of claim 25 or claim 26, wherein the medical diagnosis, prediction, and / or prognosis comprises stratifying disease risk, predicting disease incidence, classifying disease types and / or stages, assessing prognosis, or any combination thereof.
28. The method of any one of claims 25-27, wherein the medical diagnosis, prediction, and / or prognosis is for one or more cancers.
29. The method of any one of claims 25-28, comprising training a medical diagnosis, prediction, and / or prognosis model with the mixed dataset as input.
30. The method of claim 29, wherein the medical diagnosis, prediction, and / or prognosis model comprises a Transformer classification model.
31. The method of any one of claims 1-30, wherein the plurality of different image modalities comprises optical clearance tomography (OCT), computed tomography (CT), X- ray, fundus photography, magnetic resonance imaging (MRI), ultrasound image, endoscopy image, positron emission tomography (PET), single photon emission computed tomography (SPECT), microscopy image, medical photography, elastography image, a thermogram image, or any combination thereof.Attorney Docket No. 71951-711.60232. The method of any one of claims 1-31, wherein the plurality of different organs comprises eye, lung, breast, intestine, brain, kidney, pancreas, bladder, or any combination thereof.
33. A computer-implemented method, comprising:(a) training a latent stable diffusion model using: (i) training image input comprising medical images across a plurality of different image modalities and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input comprising modality information of each medical image concatenated with the corresponding textual description information, wherein the training comprises denoising using one or more shallow U-Net layers with cross-attention between image embedding and modality information embeddings and one or more deep U-Net layers with cross-attention between image embedding and description information embeddings, thereby generating a trained latent stable diffusion model;(b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text;(c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality;(d) presenting the synthetic medical image to a clinician through a structured scoring interface, and(e) inviting the clinician to provide quantitative scores across multiple dimensions comprising anatomical accuracy and clinical diagnostic value of the synthetic medical image, wherein the quantitative scores serve as training data to train a reward model which learns and simulates clinician preferences to provide scores for the synthetic medical image that correspond to its clinical quality.
34. A computer-implemented method, comprising:(a) training a latent stable diffusion model using: (i) training image input comprising medical images I =, in} and their corresponding textual descriptions D =dn} across a plurality of different image modalities M = {m^ m2, m3, ••• , mk} and a plurality of different organs, wherein each medical image is paired with its corresponding textual description, and (ii) training text input [mk; dn] comprising modality information of each medical image concatenated with the corresponding textual descriptionAttorney Docket No. 71951-711.602 information, thereby generating a trained latent stable diffusion model, wherein the training comprises:(i) introducing a series of random Gaussian noises to the training image input progressively, following the equation q(it|it-i)where jtis a linear noise scheduler defined asan(j iTispurenoiseinput, and(ii) reversing diffusion using a U-Net architecture with cross-attention mechanisms on both modality information and description information:predicted by the output of the cross-attention between image embedding and either modality information embedding or description information embedding;(b) providing a textual input comprising information of a desired image modality concatenated with information of a descriptive text; and(c) processing the textual input by the trained latent stable diffusion model to output a synthetic medical image of the desired image modality, wherein the processing comprises an iterative denoising process with sequential cross-attention comprising incorporating modality-specific information followed by refining with detailed descriptive information35. A system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method of any one of claims 1-34.
36. A non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method of any one of claims 1-34.
37. A system comprising: at least one hardware processor; non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; and instructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method of any one of claims 1-34.