Methods and systems employing generative foundation model for medical use

A generative foundation model trained on specialized medical data enhances diagnostic accuracy for rare diseases and emergency conditions, addressing AI's limitations by integrating multimodal imaging and human feedback, improving clinical decision-making and reducing harmful outputs.

WO2025227118A1PCT designated stage Publication Date: 2025-10-30ANTINOUS TECHNOLOGY CO LTD +3

Patent Information

Application Number
PCT/US2025/026516
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-01-10
Filing Date
2025-04-25
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Artificial intelligence systems face challenges in diagnosing rare diseases and emergency conditions due to a lack of specialized medical data and reliance on unfiltered public information, leading to potential misdiagnosis and over-reliance on inaccurate data.

Method used

A generative foundation model is trained on extensive datasets including electronic health records, biomedical literature, and medical textbooks, with specialized fine-tuning for rare diseases and emergency conditions, incorporating multimodal imaging and human-machine evaluation to enhance diagnostic accuracy and safety.

Benefits of technology

The model achieves diagnostic accuracy comparable to experienced clinicians, reduces harmful outputs, and supports physicians with precise diagnostic support, improving patient care and reducing errors, especially for junior practitioners.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025026516_30102025_PF_FP_ABST
    Figure US2025026516_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein are methods and devices for generating a trained generative foundation model for medical diagnosis, as well as using the trained generative foundation model for the detection of diseases and disorders including ophthalmic and systemic diseases. System and non-transitory computer-readable storage medium configured for performing the methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND SYSTEMS EMPLOYING GENERATIVE FOUNDATION MODEL FOR MEDICAL USECROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to U.S. Provisional Patent Application No. 63 / 639,619 filed April 27, 2024, International Patent Application No. PCT / US2024 / 026697 filed April 27, 2024, and Australian Patent Application No. 2025200208 filed January 10, 2025, all of which applications are incorporated herein by reference herein as if set forth in full.BACKGROUND OF THE DISCLOSURE

[0002] Artificial intelligence makes strides in specialized diagnostics but faces challenges in complex clinical scenarios, such as rare disease diagnosis and emergency condition identification.SUMMARY OF THE DISCLOSURE

[0003] In some embodiments, disclosed herein is a method for generating a medical diagnosis for a patient, comprising: receiving a natural -language prompt for obtaining the medical diagnosis and a set of data related to the patient; and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage. In some embodiments, the medical diagnosis comprises one or more ICD-10 codes. In some embodiments, the medical diagnosis relates to one or more diseases. In some embodiments, the one or more diseases comprise an ophthalmic disease or a systemic disease. In some embodiments, the medical diagnosis relates to one or more emergency conditions. In some embodiments, the one or more emergency conditions are related to an ophthalmic emergency. In some embodiments, the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof. In some embodiments, the LLM is trained to minimize an auto-regressive loss in the first stage. In some embodiments, the LLM is trained to minimize the auto-regressive loss in the second stage.

[0004] In some embodiments, disclosed herein is a method for generating a medical report for a patient, comprising: receiving a natural -language prompt for obtaining the medical report and a set of image data related to the patient; and generating the medical report by inputting the prompt and the set of image data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage. In some embodiments, the set of image data comprises one or more ophthalmic images, one or more radiological images, or any combination thereof. In someembodiments, the one or more ophthalmic images comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or any combination thereof. In some embodiments, the one or more radiological images comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof. In some embodiments, the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof. In some embodiments, the LLM is trained to minimize an auto-regressive loss in the first stage. In some embodiments, the LLM is trained to minimize the auto-regressive loss in the second stage. In some embodiments, the LLM comprises one or more vision encoders. In some embodiments, the one or more vision encoders comprise a Swin Transformer and / or a Vision Transformer.

[0005] In some embodiments, disclosed herein is a system comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one or more of the methods disclosed herein.

[0006] In some embodiments, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform any one or more of the methods disclosed herein.

[0007] In some embodiments, disclosed herein is a generative foundation model trained over millions of health system-scale electronic health records along with web-scale medical text corpora to acquire knowledge of both medical practices and theories, and use of the generative model for rare disease diagnosis (including rare ophthalmic diseases and rare systemic diseases), emergency condition identification (including ophthalmic emergencies and systemic emergencies), complex disease solving (“diagnostic puzzles”), or generating multimodal medical imaging reports (including ophthalmic images and radiology images such as X-rays and CT scans), wherein the generative model involves the use of language data for pre-training, language data for supervised finetuning using a instruction tuning approach (e.g., QA pairs), and a human-machine hybrid evaluation strategy, wherein both the pre-training and supervised finetuning phases involve scaling to extend the context window. In some embodiments, the human-machine hybrid evaluation strategy involves language data for automated evaluations as well as evaluations by generalists and by different specialists (e.g., ophthalmologists and radiologists) of varying levels of experience.

[0008] In some embodiments, disclosed herein is a method for generating a medical diagnosis relating to one or more ophthalmic diseases or conditions for a patient, the method comprising: receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient; and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage, wherein the LLM is further trained using ophthalmic imaging data in the first stage and / or the second stage. In some embodiments, the ophthalmic imaging data comprises image data from different imaging modalities. In some embodiments, the LLM comprises at least two vision encoders for processing of the image data from different imaging modalities. In some embodiments, the at least two vision encoders comprise a Swin Transformer and a Vision Transformer. In some embodiments, the LLM is trained using a reduced number of visual tokens obtained by dimension reduction of adjacent visual tokens of the ophthalmic imaging data. In some embodiments, the LLM comprises one or more linear projectors to map visual representations of the reduced number of visual tokens to a language space for training of the LLM.

[0009] In some embodiments, the ophthalmic imaging data comprises one or more images selected from the group comprising: an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, and an indocyanine green angiography (ICGA) image.

[0010] In some embodiments, the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof.

[0011] In some embodiments, the LLM is trained using an extended context window in the first stage and / or the second stage. In some embodiments, the extended context window is provided by linear rotary position embedding (RoPE) scaling.

[0012] In some embodiments, the LLM is trained to minimize an auto-regressive loss of text corpora in the first stage and / or the second stage.

[0013] In some embodiments, the LLM is trained with a reduced number of parameters in the second stage by using Low-Rank Adaptation of Large Language Models (LoRA).

[0014] In some embodiments, the set of data related to the patient is a set of image data comprising one or more ophthalmic images, one or more radiological images, or any combination thereof. In some embodiments, the one or more ophthalmic images comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or anycombination thereof. In some embodiments, the one or more radiological images comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

[0015] In some embodiments, the method further comprises generating a medical report for the patient comprising the medical diagnosis generated using the LLM. In some embodiments, the method further comprises comparing the medical diagnosis generated using the LLM with a medical diagnosis from a clinician, and determining a medical diagnosis for the patient based on the comparison.

[0016] In some embodiments, the one or more ophthalmic diseases or conditions comprise a retinal disorder, a visual pathway disorder, keratitis, a corneal scar and / or opacity condition, iridocyclitis, age-related cataract, cataract, a choroid disorder, a retinal detachment condition, a retinal vascular occlusion, a retinal disorder, glaucoma, a vitreous body disorder, and a globe disorder.

[0017] In some embodiments, disclosed herein is a system, comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one or more of the methods disclosed herein.

[0018] In some embodiments, disclosed herein is a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform any one or more of the methods disclosed herein.

[0019] In an embodiment, a non-transitory computer-readable medium having instructions stored thereon is disclosed, wherein the instructions, when executed by a processor, cause the processor to perform the methods above.

[0020] In some embodiments, provided herein is a method of generating a trained large language model for medical diagnosis, comprising: (a) pretraining a large language model (LLM) having a context window size using textual corpora to generate a pretrained LLM, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, and medical books, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize auto-regressive loss; and (b) finetuning the pretrained LLM using a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, and a multimodal imaging dataset, thereby generating a trained LLM for medical diagnosis.

[0021] In some embodiments, provided herein is a method of generating a trained large language model for medical diagnosis, comprising: (a) pretraining a large language model (LLM) having a context window size using textual corpora to generate a pretrained LLM, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, medical books, or any combination thereof, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize auto-regressive loss; and (b) finetuning the pretrained LLM using a domain-specific dataset comprising: a medical questionanswering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, a multimodal imaging dataset, or any combination thereof, thereby generating a trained LLM for medical diagnosis.

[0022] In any of the preceding embodiments, the pretraining in (a) can comprise extending the context window size of the LLM to generate a pretrained LLM having an extended context window size. In any of the preceding embodiments, the finetuning in (b) can comprise extending the context window size of the pretrained LLM. In any of the preceding embodiments, extending the context window size of the LLM and / or extending the context window size of the pretrained LLM can comprise linear rope scaling. In some embodiments, the linear rope scaling comprises downscaling position indices in Rotary Position Embedding (RoPE).

[0023] In any of the preceding embodiments, the pretraining in (a) can comprise optimization to minimize auto-regressive loss. In any of the preceding embodiments, the finetuning in (b) can comprise optimization to minimize auto-regressive loss.

[0024] In any of the preceding embodiments, the pretrained LLM can be finetuned using the medical QA dataset, the rare disease EHR dataset, the emergency EHR dataset, and the multimodal imaging dataset. In any of the preceding embodiments, the finetuning in (b) can comprise using different image encoders each for processing image data from a different imaging modality. In any of the preceding embodiments, the finetuning in (b) can comprise using a first image encoder for processing 2D images and a second image encoder for processing 3D images. In any of the preceding embodiments, the finetuning in (b) can comprise using a Swin Transformer and a Vision Transformer. In any of the preceding embodiments, the finetuning in (b) can comprise using Low-Rank Adaptation of Large Language Models (LoRA) to reduce the number of trainable parameters. In any of the preceding embodiments, the finetuning in (b) can comprise using a reduced number of visual tokens obtained by dimension reduction of adjacent visual tokens of the multimodal imaging dataset. In some embodiments,the method comprises using one or more linear projectors to map visual representations of the reduced number of visual tokens to a language space.

[0025] In any of the preceding embodiments, the multimodal imaging dataset can comprise one or more ophthalmic images, one or more radiological images, or any combination thereof. In any of the preceding embodiments, the multimodal imaging dataset can comprise one or more images selected from the group comprising: an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, and an indocyanine green angiography (ICGA) image. In any of the preceding embodiments, the multimodal imaging dataset can comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

[0026] In some embodiments, provided herein is a method of generating a medical diagnosis for a patient, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient, and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model generated by: (a) pretraining a large language model (LLM) having a context window size using textual corpora to generate a pretrained LLM, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, and medical books, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize auto-regressive loss; and (b) finetuning the pretrained LLM using a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, and a multimodal imaging dataset.

[0027] In some embodiments, provided herein is a method of generating a medical diagnosis for a patient, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient, and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model generated by: (a) pretraining a large language model (LLM) having a context window size, comprising: extending the context window size of the LLM, pretraining the LLM using textual corpora to generate a pretrained LLM having an extended context window size, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, medical books, or any combination thereof, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize auto-regressive loss; and (b) finetuning the pretrained LLM using a domain-specific dataset comprising: a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergencyEHR dataset, a natural language dataset, a multimodal imaging dataset, or any combination thereof.

[0028] In any of the preceding embodiments, the set of data related to the patient can comprise a set of image data comprising one or more ophthalmic images, one or more radiological images, or any combination thereof. In some embodiments, the one or more ophthalmic images comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or any combination thereof. In any of the preceding embodiments, the one or more radiological images can comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

[0029] In any of the preceding embodiments, the method can further comprise generating a medical report for the patient comprising the medical diagnosis generated using the trained LLM. In any of the preceding embodiments, the method can further comprise comparing the medical diagnosis generated using the trained LLM with a medical diagnosis from a clinician; and determining a medical diagnosis for the patient based on the comparison.

[0030] In any of the preceding embodiments, the medical diagnosis can be for rare disease diagnosis, urgent care diagnosis, and / or complex disease diagnosis. In any of the preceding embodiments, the medical diagnosis can be for diagnosis of an ophthalmic disease or condition. In some embodiments, the ophthalmic disease or condition comprises a retinal disorder, a visual pathway disorder, keratitis, a corneal scar and / or opacity condition, iridocyclitis, age-related cataract, cataract, a choroid disorder, a retinal detachment condition, a retinal vascular occlusion, a retinal disorder, glaucoma, a vitreous body disorder, and / or a globe disorder.

[0031] In any of the preceding embodiments, the method can further comprise (c) collecting evaluations and feedback from one or more medical practitioners on outputs from the pretrained and finetuned LLM. In some embodiments, the method further comprises: (d) using a reward model and Proximal Policy Optimization (PPO) to train the pretrained and finetuned LLM to generate responses aligned with clinical practice.

[0032] In some embodiments, provided herein is a system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method of any one of the preceding embodiments. In some embodiments, provided herein is a non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method of any one of the preceding embodiments. In some embodiments, provided herein is a system comprising: at least one hardware processor; non- transitory computer-readable medium coupled to at least one hardware processor, optionallywherein the coupling is over a network; and instructions stored in the non-transitory computer- readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method of any one of the preceding embodiments.

[0033] In any of the preceding embodiments, the textual corpora can comprise unstructured data. In any of the preceding embodiments, the textual corpora can comprise both structured data and unstructured data. In any of the preceding embodiments, the medical QA dataset can comprise unstructured data. In any of the preceding embodiments, the medical QA dataset can comprise both structured data and unstructured data. In any of the preceding embodiments, the rare disease EHR dataset can comprise unstructured data. In any of the preceding embodiments, the rare disease EHR dataset can comprise both structured data and unstructured data. In any of the preceding embodiments, the emergency EHR dataset can comprise unstructured data. In any of the preceding embodiments, the emergency EHR dataset can comprise both structured data and unstructured data. In any of the preceding embodiments, the natural language dataset can comprise unstructured data. In any of the preceding embodiments, the natural language dataset can comprise both structured data and unstructured data. In any of the preceding embodiments, the multimodal imaging dataset can comprise unstructured data. In any of the preceding embodiments, the multimodal imaging dataset can comprise both structured data and unstructured data.BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The details of the disclosed embodiments, both as to structure and operation, may be gleaned in part by study of the accompanying drawings, in which like reference numerals refer to like parts, and in which:

[0035] FIG. 1A shows the development pipeline and dataset overview of an exemplary generative foundational model. EHRs, electronic health records; RD, rare disease; EC, emergency condition; CMAIC, China Medical Al Investigation Consortium; MIMIC-IV, Medical Information Mart for Intensive Care IV; QA, question answering; PubMedQA, PubMed question answering; MedQA, medical question answering; MedMCQA, medical multiple-choice question answering; CXR, chest X-ray; CT, computed tomography; TOAC, the OpenAssistant Conversations dataset (TOAC); CoTC, Chain-of-Thought Collection dataset; AJEM, American Journal of Emergency Medicine; JACEP, Journal of American College of Emergency Physicians; NAJMS, North American Journal of Medical Sciences; AJMG, American Journal of Medical Genetics; CC-CXRI, China Consortium of Chest X-ray Image Investigation; CC-CCII, China Consortium of Chest CT Image Investigation; MIMIC-CXR, Medical Information Mart for Intensive Care Chest X-ray; lU-Xray, Indiana University Chest X-ray Dataset.

[0036] FIG. IB illustrates the structural design of a 32-billion-parameter generative foundation model. The architecture comprises multiple MetaGP Blocks, each consisting of RMSNorm, Linear, Group Query Attention and FeedForward layers, with attention mechanisms facilitating multimodal data integration. The left panel depicts a single MetaGP Block, while the right panel shows the overall model flow, starting from an Embedding layer through N MetaGP Blocks, and ending with RMSNorm and Linear layers for output generation.

[0037] FIG. 2A provides an overview of the evaluation framework using clinical cases from local EHR data, public datasets, and radiologic imaging (CXR and CT). Evaluation metrics included manual scores for diagnostic accuracy, clinical relevance, and reasoning, alongside automated metrics such as Fl score, accuracy, recall, and precision.

[0038] FIG. 2B provides an example of a rare disease diagnosis with scoring criteria for diagnostic accuracy.

[0039] FIG. 2C shows comparison of MetaGP-generated reports with physician-generated reports for imaging diagnoses.

[0040] FIG. 3A shows comparison of average diagnostic performance for rare systemic diseases among MetaGP, GPT-4, and general practitioners with varying experience levels. Data are represented as mean.

[0041] FIG. 3B shows distribution of diagnostic scores (ranging from -2 to 2) for MetaGP, GPT-4, and individual practitioners in rare disease cases. Data are represented as mean.

[0042] FIG. 3C shows average accuracy, Fl scores, recall, and precision metrics of MetaGP, GPT-4, and BERT in diagnosing rare diseases. Data are represented as mean.

[0043] FIG. 4A shows comparison of average diagnostic performance for emergency conditions among MetaGP, GPT-4, and general practitioners with varying experience levels. Data are represented as mean.

[0044] FIG. 4B shows distribution of diagnostic scores (ranging from -2 to 2) for MetaGP, GPT-4, and individual practitioners in emergency condition cases. Data are represented as mean.

[0045] FIG. 4C shows average accuracy, Fl scores, recall, and precision metrics of MetaGP, GPT-4, and BERT in diagnosing emergency conditions. Data are represented as mean.

[0046] FIGS. 5A-5C compare diagnostic performance of MetaGP, GPT-4, and BERT in rare disease categories based on ICD-10 codes. MetaGP generally outperforms the other models, with statistical significance tested by the Mann-Whitney-Wilcoxon test. Data are represented as mean.

[0047] FIGS. 5D-5H compare diagnostic accuracy of MetaGP, GPT-4, and BERT in emergency condition categories, also based on ICD-10 codes. MetaGP consistently showshigher accuracy, especially in critical categories, with statistical significance confirmed by the Mann-Whitney-Wilcoxon test. Data are represented as mean.

[0048] FIG. 6A shows comparison of MetaGP, Med-Flamingo, and LLaVA-Med in generating CXR reports, evaluated using ROUGE-L and BertScore metrics on internal and external validation sets. MetaGP achieved significantly higher scores than the other models in both metrics, indicating superior report quality. Data are represented as mean. *** / ? < 0.001.

[0049] FIG. 6B shows performance comparison of MetaGP, Med-Flamingo, and LLaVA- Med in CT report generation, with evaluations conducted on internal and prospective validation sets. MetaGP consistently outperformed competing models, achieving higher ROUGE-L and BertScore values. Data are represented as mean. *** / ? < 0.001.

[0050] FIG. 6C shows rater assessments of CXR reports generated by MetaGP, physicians, and other models. Eight raters, including one senior and seven junior radiologists, evaluated the reports in a blinded setting to determine which reports were of higher quality or if they were equivalent. The proportions reflect preferences for physician-generated reports, Al-generated reports, or a tie in report quality.

[0051] FIG. 6D shows rater assessments of CT reports. The pie charts show the distribution of preferences among the raters, highlighting MetaGP’ s performance in comparison to physician-generated reports and other models.

[0052] FIG. 7A shows an exemplary schematic diagram of the development of EyeGPT, according to some embodiments of the disclosure.

[0053] FIG. 7B shows an exemplary composition of text data for pretraining, according to some embodiments of the disclosure.

[0054] FIG. 7C shows an exemplary composition of image sets from different modalities, according to some embodiments of the disclosure.

[0055] FIG. 7D shows exemplary diagnostic performance scores achieved when evaluated on real-world clinical cases, according to some embodiments of the disclosure.

[0056] FIG. 8A shows an exemplary schematic diagram of the evaluation of rare disease diagnoses, according to some embodiments of the disclosure.

[0057] FIG. 8B shows an exemplary schematic diagram of the evaluation of emergency condition diagnoses, according to some embodiments of the disclosure.

[0058] FIG. 8C shows an exemplary schematic diagram of the evaluation of complex disease diagnoses, according to some embodiments of the disclosure.

[0059] FIG. 8D shows an exemplary schematic diagram of the evaluation of generated medical reports for multimodal images, according to some embodiments of the disclosure.

[0060] FIG. 9A shows an exemplary average diagnostic performance of EyeGPT, GPT-4, and ophthalmologists with varying experience levels for rare ophthalmic diseases, according to some embodiments of the disclosure.

[0061] FIG. 9B shows an exemplary distribution of scores (from -2 to 2) of individuals for rare ophthalmic disease diagnoses, according to some embodiments of the disclosure.

[0062] FIG. 9C shows an exemplary average accuracy and Fl scores of EyeGPT, GPT-4, and BERT in diagnosing rare diseases, according to some embodiments of the disclosure.

[0063] FIG. 9D shows exemplary statistical analyses and per-category results of EyeGPT, GPT-4, and BERT in diagnosing rare diseases, according to some embodiments of the disclosure.

[0064] FIG. 10A shows an exemplary average diagnostic performance of EyeGPT, GPT-4, and ophthalmologists with varying experience levels for ophthalmic emergencies, according to some embodiments of the disclosure.

[0065] FIG. 10B shows an exemplary distribution of scores (from -2 to 2) of individuals for ophthalmic emergency diagnoses, according to some embodiments of the disclosure.

[0066] FIG. 10C shows an exemplary average accuracy and Fl scores of EyeGPT, GPT-4, and BERT in diagnosing emergency diseases, according to some embodiments of the disclosure.

[0067] FIG. 10D shows exemplary statistical analyses and per-category results of EyeGPT, GPT-4, and BERT in diagnosing emergency diseases, according to some embodiments of the disclosure.

[0068] FIG. 11 shows exemplary accuracy of EyeGPT, Google LLM for DDx, and GPT-4 on case studies presented by the New England Journal of Medicine clinicopathological conferences, according to some embodiments of the disclosure.

[0069] FIG. 12A shows an exemplary multimodal medical report generation of EyeGPT, Med-Flamingo, and LLaVA-Med on the OCT imaging modality, according to some embodiments of the disclosure.

[0070] FIG. 12B shows an exemplary multimodal medical report generation of EyeGPT, Med-Flamingo, and LLaVA-Med on the Fundus imaging modality, according to some embodiments of the disclosure.

[0071] FIG. 12C shows an exemplary multimodal medical report generation of EyeGPT, Med-Flamingo, and LLaVA-Med on the FFA imaging modality, according to some embodiments of the disclosure.

[0072] FIG. 12D shows an exemplary multimodal medical report generation of EyeGPT, Med-Flamingo, and LLaVA-Med on the IGGA imaging modality, according to some embodiments of the disclosure.

[0073] FIG. 13A shows an exemplary comparative analysis of Al-generated and physician- composed medical image reports on the OCT imaging modality, according to some embodiments of the disclosure.

[0074] FIG. 13B shows an exemplary comparative analysis of Al-generated and physician- composed medical image reports on the Fundus imaging modality, according to some embodiments of the disclosure.

[0075] FIG. 13C shows an exemplary comparative analysis of Al-generated and physician- composed medical image reports on the FFA imaging modality, according to some embodiments of the disclosure.

[0076] FIG. 13D shows an exemplary comparative analysis of Al-generated and physician- composed medical image reports on the IGGA imaging modality, according to some embodiments of the disclosure.

[0077] FIG. 14A shows an exemplary diagnostic performance comparison among different models in rare ophthalmic diseases, according to some embodiments of the disclosure.

[0078] FIG. 14B shows an exemplary diagnostic performance comparison among different models in emergency ophthalmic diseases, according to some embodiments of the disclosure.

[0079] FIG. 15 shows an exemplary flow chart of a method for generating a medical diagnosis for a patient, according to some embodiments of the disclosure.

[0080] FIG. 16 shows an exemplary schematic diagram illustrating a system, which includes a computer system with a processing unit, a memory and one or more programs stored in the memory, and an optional electronic device having a display which communicates with the computer system, where the one or more programs stored in the memory may include instructions for performing the method shown in FIG. 15, according to some embodiments of the disclosure.

[0081] FIG. 17 shows the overall training procedure of an exemplary generative foundational model, which comprises four stages: pretraining, multimodal instruction finetuning, Reinforcement Learning from Human Feedback (RLHF) training, and application. Pretraining: the model was trained on large-scale datasets, including EHRs, academic articles, and medical books, using advanced processes like data aggregation, tokenization, pretrained language model (Qwen) initialization, and continual pretraining to build a strong foundation in medical knowledge. Multimodal instruction fine-tuning: this phase incorporated domain-specific datasets, such as medical QA, rare disease, and emergency EHRs and multimodal imaging data, to optimize task-specific outputs through text and image encoders. RLHF training: using evaluations and feedback from medical practitioners (e.g., ophthalmology specialists) on model outputs to optimize the model with Proximal Policy Optimization (PPO). Application: thetrained model was applied to diagnostic tasks, clinical decision support, and multimodal report generation. The evaluation included real-world cases from various clinical scenarios.

[0082] FIGS. 18A-18D show the comparison between the performance of Base Model and RL-enhanced Model in Diagnostic tasks and Treatment Recommendation tasks. RL: reinforcement learning. FIG. 18A shows score distributions of the two models in diagnostic tasks based on private cases. FIG. 18B shows score distributions of the two models in diagnostic tasks based on published cases. FIG. 18C shows score distributions of the two models in treatment recommendation tasks based on private cases. FIG. 18D shows score distributions of the two models in treatment recommendation tasks based on published cases.DETAILED DESCRIPTION OF THE DISCLOSURE

[0083] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.I. Generative Foundation Model for Medical Use

[0084] The development of LLMs in medicine has seen remarkable progress with the introduction of domain-specific models like PMC-LLaMA, BiomedGPT, and GatorTronGPT. These models, trained on specialized datasets such as PubMed articles, EHRs, medical textbooks, and imaging data, have advanced capabilities in tasks like clinical QA and medical report generation. However, they often fall short in handling complex and high-stakes scenarios, such as diagnosing rare diseases, managing critical emergency conditions, and integrating multimodal data for holistic patient care.

[0085] Disclosed herein are generative foundation models integrating electronic health records and multimodal imaging for addressing unmet clinical needs. To address the limitations of artificial intelligence in rare disease diagnosis and emergency condition identification, provided herein are generative foundation models trained on extensive datasets, including electronic health records, biomedical literature, and medical textbooks. One example is Meta General Practitioner (MetaGP), a 32-billion-parameter generative foundation model which demonstrates robust diagnostic capabilities, achieving accuracy comparable to experienced clinicians. In rare disease cases, it achieves an average diagnostic score of 1.57, surpassing GPT- 4’s 0.93. For emergency conditions, it improves diagnostic accuracy for junior and mid-level clinicians by 53% and 46%, respectively. MetaGP also excels in generating medical imaging reports, producing high-quality outputs for chest X-rays and computed tomography, often rated comparable to or superior to physician-authored reports. These findings highlight MetaGP’ s potential to transform clinical decision-making across diverse medical contexts.

[0086] In some embodiments, a trained generative foundation model disclosed herein addresses these limitations by excelling in the diagnosis of rare diseases and the identification of life-threatening emergency conditions — areas that demand both nuanced understanding and rapid decision-making. In some embodiments, through its multimodal capabilities, a trained generative foundation model disclosed herein integrates textual, imaging, and structured medical data, providing a comprehensive approach to complex diagnostics. Unlike prior models, which are often tailored to specific tasks or data types, the targeted fine-tuning and robust foundational design of the trained generative foundation model make it uniquely suited to tackle challenging medical cases across diverse scenarios, delivering precise, context-aware diagnostic support and setting a new standard for Al-driven clinical applications.

[0087] In some embodiments, a trained generative foundation model disclosed herein minimizes potentially harmful outputs compared to other generative foundation models such as GPT-4. In some embodiments, the trained generative foundation model produces fewer responses classified as “bad” or “dangerous,” which, if followed, could lead to adverse patient outcomes, for instance, on diagnosing rare diseases and emergency conditions. This suppression of harmful outputs is especially valuable in the medical field, where incorrect or misleading information can carry serious implications for patient health and safety. The reduction in harmful responses can be in some aspects attributed to the specialized training across comprehensive medical datasets, including EHRs, scholarly articles, and textbooks. By concentrating on domain-specific data, the trained generative foundation model has developed a more refined understanding of medical concepts and their interrelations, enabling it to generate highly accurate and contextually relevant diagnostic responses. This focused training allows the trained generative foundation model to more effectively distinguish beneficial from potentially harmful information, thereby reducing the likelihood of outputs that could negatively affect patient care. Moreover, the integration of insights with healthcare professionals across various levels of experience led to a further reduction in harmful responses. In some embodiments, the trained generative foundation model combines human expertise with Al-driven analysis provides an additional safeguard against errors and misinterpretations. The synergy of human and Al enhances the overall reliability and safety of diagnostic processes, ensuring a higher standard of patient care and reducing the risks associated with automated diagnostics.

[0088] In some embodiments, there is marked improvement in physicians’ diagnostic accuracy when supported by a trained generative foundation model disclosed herein. Across experience levels — from junior to senior practitioners — collaborating with the trained generative foundation model consistently elevated diagnostic accuracy for rare diseases and emergency conditions, demonstrating the value of Al-assisted decision support in clinical practice. In someembodiments, the trained generative foundation model’s extensive training on a diverse range of medical datasets enables it to draw from a comprehensive knowledge base, often including insights or patterns that individual physicians may not have encountered. By offering physicians additional perspectives and relevant considerations, the trained generative foundation model facilitates more thorough and informed diagnostic decisions. Moreover, the trained generative foundation model’s ability to quickly process and analyze large volumes of medical data allows it to recognize patterns and connections that might be missed by human observers, especially in complex or atypical cases where critical details may not be immediately evident. In some embodiments, the trained generative foundation model’s assistance had a more pronounced impact on junior and mid-level physicians than on their senior counterparts. This finding suggests that Al-assisted decision support is particularly beneficial for less experienced practitioners, who may lack the breadth of clinical exposure seen in more seasoned professionals. In some embodiments, by harnessing the trained generative foundation model’s expansive knowledge and analytical power, junior and mid-level physicians can bridge experience gaps and achieve greater diagnostic accuracy, ultimately elevating patient care.

[0089] In some embodiments, the trained generative foundation model demonstrates a significant advantage in generating clinically useful reports across multiple imaging modalities, underscoring its potential to enhance diagnostic workflows. Its consistent outperformance of established models in both internal and external validations highlights the model’s robustness and adaptability across various types of medical imaging. Human evaluation results further emphasize in some embodiments, the trained generative foundation model’s competitive edge, as it surpassed physician-generated reports in a notable proportion of cases across different imaging modalities. Interestingly, junior physicians exhibited a higher preference for AI- generated reports, suggesting a potential shift in reliance on Al tools in fields like radiology. In some embodiments, the trained generative foundation model can be used as a tool complementary to human expertise. By supporting medical professionals with precise, clinically relevant reports, the trained generative foundation model enhances the diagnostic process across a range of imaging modalities, ultimately contributing to improved clinical decision-making.

[0090] In some embodiments, the trained generative foundation model addresses potential biases prevalent in Al models trained on public datasets. By incorporating diverse training data from both public sources and private EHRs representing varied populations, the trained generative foundation model reduces the risk of demographic biases. Additionally, the model underwent fine-tuning using rare disease and emergency datasets curated by physicians, ensuring its robustness in high-stakes clinical scenarios. Comprehensive bias audits during the evaluation phase further ensured consistent performance across different population groups.These safeguards, combined with ongoing monitoring and refinement, position the trained generative foundation model as a reliable and equitable tool for clinical decision-making, addressing key ethical concerns in Al-driven healthcare.

[0091] In some embodiments, the trained generative foundation model exemplifies the integration of large, generative LLMs into medical workflows, the challenge remains in efficiently accommodating diverse user demands and clinical tasks while ensuring seamless integration into existing medical informatics systems. With continued monitoring of humanmachine interactions, the trained generative foundation model can positively impact physician behavior and patient care outcomes, with safeguards in place to uphold ethics and patient safety. As an example, MetaGP, which was built on the open-source Qwen-1.521 framework with a modest parameter count (32B), offers both versatility and agility, requiring significantly fewer computational resources. For pretraining, 120 NVIDIA A100 GPUs (graphics processing units) with 80 GB of VRAM (video random access memory) each were utilized over four weeks, followed by fine-tuning sessions using 48 Al 00 GPUs over five days per iteration. Randomized controlled clinical trials and extensive user feedback from integrating the trained generative foundation model into healthcare delivery are essential to fully validate its impact. Such evaluations would support the inclusion of task-specific, intervention-based recommendations aligned with predicted patient risk levels.

[0092] In some embodiments, increasing the size of a generative foundation model such as an LLM, specifically the number of parameters, leads to improved performance. Larger models generally demonstrate enhanced capabilities in learning complex patterns and representations, resulting in better accuracy and generalization. In some embodiments, however, to ensure computational efficiency, the model size is intentionally small. In some embodiments, the number of parameters of a generative foundation model disclosed herein is about 7 billion, about 10 billion, about 13 billion, about 16 billion, about 19 billion, about 22 billion, about 25 billion, about 28 billion, about 31 billion, about 34 billion, about 37 billion, or about 40 billion. In some embodiments, the number of parameters of a generative foundation model disclosed herein is no more than about 32 billion. In some embodiments, the number of parameters of an LLM disclosed herein is about 45 billion, about 50 billion, about 55 billion, about 50 billion, or about 65 billion. In some embodiments, the number of parameters of a generative foundation model disclosed herein is more than about 65 billion, such as 175 billion or about 176 billion.II. Generative Foundation Model for Ophthalmic Diagnostics

[0093] In some embodiments, the present disclosure relates to methods and systems employing a generative foundation model for diagnosing diseases or conditions inophthalmology. The present disclosure relates particularly but not exclusively to diagnosing rare, emergency and complex ophthalmic diseases or conditions.

[0094] The accurate diagnosis of ocular diseases or conditions poses significant challenges in ophthalmology. Accurate interpretation of imaging modalities such as optical coherence tomography (OCT), fundus photography, fluorescein angiography (FFA), and indocyanine green angiography (ICGA) are critical for effective disease management, yet they generate extensive and complex datasets that require timely and precise analysis to prevent vision loss and other serious complications. The diagnostic process is further complicated by the wide spectrum of ocular diseases, from prevalent conditions like diabetic retinopathy to rare genetic disorders such as corneal dystrophies, an inherited, progressive disorder of the cornea. Diagnosing rare diseases is particularly challenging due to their low incidence and often ambiguous symptoms, demanding specialized knowledge and nuanced understanding of medical literature. Complex ocular diseases, while not necessarily rare, involve multifactorial etiologies and systemic manifestations, presenting intricate diagnostic challenges. Traditional diagnostic approaches depend heavily on the expertise of ophthalmologists, leading to variability in diagnosis, and the growing volume of patient data increases the risk of errors and treatment delays. To address these challenges, innovative solutions are imperative to enhance diagnostic accuracy and efficiency in ophthalmic practice.

[0095] Artificial Intelligence (Al) has emerged as a transformative period in healthcare and medicine. Recent advancements have enabled Al tools to effectively analyze diverse medical datasets, including dermoscopic images, retinal images, electronic health records (EHRs), electrocardiograms, and data from oncology trials. The advent of generative large language models (LLMs) offers a promising solution. Trained on vast free-text, web-scale data, these models possess broad, multidisciplinary knowledge. Foundation models like ChatGPT have demonstrated notable performance, though they present risks of delivering inaccurate and harmful information, since many LLMs rely on general internet-based knowledge, lacking specialized medical data (EHRs, etc.). This poses a significant challenge when applying the models to nuanced medical applications. The dangers are two-fold: (i) the potential for misdiagnosis due to the lack of knowledge of medical practices and (ii) the susceptibility to over-reliance on unfiltered public data which might be misleading or factually incorrect. For Al to maximize its impact in medical diagnostics, it requires a foundation deeply rooted in vast and authentic medical data and expertise. Additionally, current Al models often struggle with crossdomain diagnostic tasks. For instance, an Al model trained on systemic disease indicators may overlook ocular manifestations like glaucoma or macular degeneration, leading to misseddiagnoses and incomplete patient care. Additionally, these Al models depend heavily on large volumes of structured data, requiring labor-intensive management that limits scalability.

[0096] Provided herein are methods and systems employing a generative foundation model for diagnosing ophthalmic diseases or conditions which ameliorate and / or overcome one or more disadvantages of existing arrangements, or at least provide a useful alternative.

[0097] In one aspect, the present disclosure provides a method for generating a medical diagnosis relating to one or more ophthalmic diseases or conditions for a patient, the method comprising: receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient; and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage, wherein the LLM is further trained using ophthalmic imaging data in the first stage and / or the second stage.

[0098] In some embodiments, the ophthalmic imaging data comprises image data from different imaging modalities.

[0099] In some embodiments, the LLM comprises at least two vision encoders for processing of the image data from different imaging modalities. The at least two vision encoders may comprise a Swin Transformer and a Vision Transformer.

[0100] In some embodiments, the LLM is trained using a reduced number of visual tokens obtained by dimension reduction of adjacent visual tokens of the ophthalmic imaging data.

[0101] In some embodiments, the LLM comprises one or more linear projectors to map visual representations of the reduced number of visual tokens to a language space for training of the LLM.

[0102] In some embodiments, the ophthalmic imaging data comprises one or more images selected from the group comprising: an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, and an indocyanine green angiography (ICGA) image.

[0103] In some embodiments, the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof.

[0104] In some embodiments, the LLM is trained using an extended context window in the first stage and / or the second stage. The extended context window may be provided by linear rotary position embedding (RoPE) scaling.

[0105] In some embodiments, the LLM is trained to minimize an auto-regressive loss of text corpora in the first stage and / or the second stage.

[0106] In some embodiments, the LLM is trained with a reduced number of parameters in the second stage by using Low-Rank Adaptation of Large Language Models (LoRA).

[0107] In some embodiments, the set of data related to the patient is a set of image data comprising one or more ophthalmic images, one or more radiological images, or any combination thereof. The one or more ophthalmic images may comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or any combination thereof. The one or more radiological images may comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

[0108] In some embodiments, the method further comprises generating a medical report for the patient comprising the medical diagnosis generated using the LLM.

[0109] In some embodiments, the method further comprises: comparing the medical diagnosis generated using the LLM with a medical diagnosis from a clinician; and determining a medical diagnosis for the patient based on the comparison.

[0110] In some embodiments, the one or more ophthalmic diseases or conditions comprise a retinal disorder, a visual pathway disorder, keratitis, a corneal scar and / or opacity condition, iridocyclitis, age-related cataract, cataract, a choroid disorder, a retinal detachment condition, a retinal vascular occlusion, a retinal disorder, glaucoma, a vitreous body disorder, and a globe disorder. The one or more ophthalmic diseases or conditions may comprise one or more ICD-10 codes.[OHl] In another aspect, the present disclosure provides a system, comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of the above aspect.

[0112] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform the method of the above aspect.

[0113] In another aspect, the present disclosure provides a method for generating a medical diagnosis for a patient, comprising: receiving a natural -language prompt for obtaining the medical diagnosis and a set of data related to the patient; and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage.

[0114] In some embodiments, the medical diagnosis comprises one or more ICD-10 codes.

[0115] In some embodiments, the medical diagnosis relates to one or more diseases. The one or more diseases may comprise an ophthalmic disease or a systemic disease.

[0116] In some embodiments, the medical diagnosis relates to one or more emergency conditions. The one or more emergency conditions may be related to an ophthalmic emergency.

[0117] In some embodiments, the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof.

[0118] In some embodiments, the LLM is trained to minimize an auto-regressive loss in the first stage. The LLM may be trained to minimize the autoregressive loss in the second stage.

[0119] In another aspect, the present disclosure provides a method for generating a medical report for a patient, comprising: receiving a natural -language prompt for obtaining the medical report and a set of image data related to the patient; and generating the medical report by inputting the prompt and the set of image data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage.

[0120] In some embodiments, the set of image data comprises one or more ophthalmic images, one or more radiological images, or any combination thereof.

[0121] In some embodiments, the one or more ophthalmic images comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or any combination thereof.

[0122] In some embodiments, the one or more radiological images comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

[0123] In some embodiments, the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof.

[0124] In some embodiments, the LLM is trained to minimize an auto-regressive loss in the first stage.

[0125] In some embodiments, the LLM is trained to minimize the autoregressive loss in the second stage.

[0126] In some embodiments, the LLM comprises one or more vision encoders. The one or more vision encoders may comprise a Swin transformer and / or a vision transformer.

[0127] In another aspect, the present disclosure provides a system, comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are storedin the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any one of the methods of the above aspect.

[0128] In another aspect, the present disclosure provides a non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform any one of the methods of the above aspect.

[0129] In another aspect, the present disclosure provides a generative foundation model trained over millions of health system-scale electronic health records along with web-scale medical text corpora to acquire knowledge of both medical practices and theories, and use of the generative model for rare disease diagnosis (including rare ophthalmic diseases and rare systemic diseases), emergency condition identification (including ophthalmic emergencies and systemic emergencies), complex disease solving (“diagnostic puzzles”), or generating multimodal medical imaging reports (including ophthalmic images and radiology images such as X-rays and CT scans), wherein the generative model involves the use of language data for pretraining, language data for supervised finetuning using an instruction tuning approach (e.g., QA pairs), and a human-machine hybrid evaluation strategy, wherein both the pre-training and supervised finetuning phases involve scaling to extend the context window.

[0130] In some embodiments, the human-machine hybrid evaluation strategy involves language data for automated evaluations, as well as evaluations by generalists and by different specialists (e.g., ophthalmologists and radiologists) of varying levels of experience.

[0131] Many existing LLMs rely on general internet-based knowledge instead of specialized medical data and / or expertise for medical diagnostics, especially for ophthalmology.Furthermore, existing LLMs are often trained on large volumes of structured data, which requires human management and limits scalability. To address these technical limitations, a foundational Al model may be trained on both specialized knowledge and a comprehensive understanding of health, with the ability to process both unstructured and structured data, to enable broader and more accurate diagnostic support for use in methods and systems of some embodiments of the disclosure.

[0132] In training existing LLMs with medical text corpora, including medical articles, text content from medical books and clinical data, it is necessary to tokenize the corpora for training of the model. This is the process of breaking down a piece of text (such as a sentence or paragraph) into smaller units called tokens. However, tokenizing medical text corpora results in large numbers of lengthy tokens which impairs the training and inference efficiency of existing LLMs. To address this technical limitation, the medical corpora can be tokenized, concatenatedinto a sequence of tokens, and optimized by minimizing auto-regressive loss for training of a foundational Al model for use in methods and systems of some embodiments of the disclosure.

[0133] Another technical problem with existing LLMs for medical scenarios is that the default size of the context window can cause technical limitations. For example, the default context window size may be exceeded for predictive tasks involving long texts, such as differential diagnosis based on extensive patient information. Where the total length of the patient information exceeds the default context window size of the LLM, significant performance degradation may occur. To address this technical limitation, an extended context window may be used for training of a foundational Al model to improve accuracy and performance, particularly for differential diagnosis tasks, for use in methods and systems of some embodiments of the disclosure.

[0134] Another technical problem in the training of existing LLMs for medical diagnostics is in the processing of image data from different modalities which have varying dimensions. To address this technical limitation, multiple vision encoders with different architectures can be employed for processing of medical images from different modalities for training of a foundational Al model for use in methods and systems of some embodiments of the disclosure.

[0135] A further technical problem is in the processing of image-based data which results in the generation of lengthy visual token counts or sequences. Long visual token counts impair the training and inference efficiency of existing LLMs. To address this technical limitation, the visual token account can be reduced through dimension reduction in order to reduce memory costs and speed up model training of a foundational Al model for use in methods and systems of some embodiments of the disclosure.

[0136] Yet another technical problem in the training of existing LLMs is in parameterefficient training to reduce the number of parameters which need to be updated during the training stage. To address this technical limitation, a Low-Rank Adaptation of Large Language Models (LoRA) strategy may be employed to reduce the number of trainable parameters requiring update during the training stage of a foundational Al model for use in methods and systems of some embodiments of the disclosure.

[0137] In some embodiments, disclosed herein is a generative foundation model, for example Eye Generative Pretrained Transformer (EyeGPT) for ophthalmology, to support a wide spectrum of complex clinical decision-making scenarios. The generative foundation model may be used in methods and systems for diagnosis of ophthalmic diseases or conditions, according to some embodiments of the disclosure.

[0138] Unlike previous endeavors and existing LLMs, the generative foundation model is trained on health system-scale electronic health records along with web-scale medical textcorpora to acquire knowledge of both medical practices and theories. In some embodiments, the model is evaluated across four critical clinical functions: rare disease diagnosis, emergency condition identification, complex disease diagnosis, and generation of medical reports. Through a rigorous and extensive evaluation by medical professionals on real-world cases, the generative foundation model not only surpasses GPT-4 in providing more accurate diagnostics but also significantly reduces potentially harmful outputs. Evaluated on real-world case studies from PubMed, the generative foundation model shows a markedly improved ability to diagnose rare ophthalmic diseases, achieving a diagnostic score of 1.23 out of a maximum of 2, outperforming GPT-4 by a large margin. When identifying emergency ophthalmic conditions, the generative foundation model achieves a score of 1.40, compared to GPT-4 scoring 1.02. When tackling complex diseases presented by the New England Journal of Medicine, the generative foundation model achieves a top-3 accuracy rate of 46%, surpassing the performance of both the Google LLM for DDx and GPT-4. In some embodiments, the generative foundation model is used for processing multimodal data by building a versatile medical report generator for multimodal imaging. Trained on over one million image-report pairs, the generative foundation model excels beyond top multimodal medical models and junior physicians in six imaging modalities, matching the performance of senior physicians. In some embodiments, the targeted training of the generative foundation model specifically on medical datasets uniquely equips it as a specialized, open-source solution for healthcare uses, setting it apart from the broad functionalities of GPT-4.

[0139] In some embodiments, the generative foundation model is trained using academic articles, medical images and associated reports, electronic health records, natural language data from the Internet, and medical textbooks, for instance, as shown in FIG. 7A. The resulting model can provide decision support for multimodal imaging report generation, rare disease diagnosis, emergency condition identification, and complex disease diagnostics in methods and systems according to some embodiments of the disclosure. FIG. 7B shows the composition of the text data for pretraining. FIG. 7C shows the composition of image sets from different modalities with the total number approximating 1 million. FIG. 7D shows the diagnostic performance scores achieved by the combination of senior physicians and EyeGPT, EyeGPT alone, and GPT-4 when evaluated on real -world clinical cases involving rare diseases and emergency conditions.

[0140] In some embodiments, disclosed herein is an evaluation schema that integrates human and automated evaluation methods, combining the expertise of medical professionals from relevant fields to conduct pertinent assessments. FIG. 8A shows evaluation of rare disease diagnoses. FIG. 8B shows evaluation of emergency condition diagnoses. FIG. 8C showsevaluation of complex disease diagnoses. FIG. 8D shows evaluation of generated medical reports for multimodal images.

[0141] In some embodiments, disclosed herein is a method comprising comparing the generative foundation model with human evaluation for rare disease diagnoses on case reports. FIG. 9A shows average diagnostic performance of EyeGPT, GPT-4, and ophthalmologists with varying experience levels for rare ophthalmic diseases. FIG. 9B shows distributions of scores (from -2 to 2) of individuals for rare ophthalmic disease diagnoses. FIG. 9C shows average accuracy and Fl scores of EyeGPT, GPT-4, and BERT in diagnosing rare diseases. FIG. 9D shows statistical analyses and per-category results of EyeGPT, GPT-4, and BERT in diagnosing rare diseases.

[0142] In some embodiments, disclosed herein is a method comprising comparing the generative foundation model with human evaluation of emergency conditions diagnoses on case reports. FIG. 10A shows average diagnostic scores of EyeGPT, GPT-4, and ophthalmologists with varying experience levels for ophthalmic emergencies. FIG. 10B shows distributions of scores (from -2 to 2) of individuals for ophthalmic emergency diagnoses. FIG. 10C shows average accuracy and Fl scores of EyeGPT, GPT-4, and BERT in diagnosing emergency diseases. FIG. 10D shows statistical analyses and per-category results of EyeGPT, GPT-4, and BERT in diagnosing emergency diseases.

[0143] In some embodiments, disclosed herein is a method comprising using the generative foundation model for automated evaluation for complex disease diagnoses. FIG. 11 shows the top-n accuracy of EyeGPT, Google LLM for DDx, and GPT-4 on case studies presented by the New England Journal of Medicine clinicopathological conferences.

[0144] In some embodiments, disclosed herein is a method comprising using the generative foundation model for multimodal medical report generation. In some embodiments, the report generation results of EyeGPT, Med-Flamingo, and LLaVA-Med on 4 imaging modalities are evaluated. FIGS. 12A-12D show external validation on OCT and retinal fundus photographs, with prospective validation conducted on retinal fundus photographs. Evaluation metrics include ROUGE-L and BertScore.

[0145] In some embodiments, disclosed herein is a method comprising comparing the generative foundation model with physician evaluation of medical reports. In some embodiments, disclosed herein is a comparative analysis of Al-generated and physician- composed medical image reports on multiple imaging modalities. In some embodiments, the imaging modalities comprise OCT (e.g., as shown in FIG. 13A), Fundus (e.g., as shown in FIG. 13B), FFA (e.g., as shown in FIG. 13C), and ICGA (e.g., as shown in FIG. 13D). For ophthalmic images, six raters, including one senior (rater 1) and five junior physicians can beemployed to do the pairwise human evaluation. In a masked assessment, where reports are deidentified, raters are tasked with judging the quality of the reports, determining which is superior or if there is a tie.

[0146] As will be described in the Example below, FIG. 14A and FIG. 14B illustrate exemplary diagnostic performance comparisons between different models in rare ophthalmic diseases and emergency ophthalmic diseases, respectively, according to some embodiments of the disclosure.

[0147] FIG. 15 shows an exemplary flow chart of a method 100 for generating a medical diagnosis for a patient, according to some embodiments of the disclosure. The medical diagnosis may relate to one or more ophthalmic diseases or conditions. The method 100 includes a step 102 of receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient. The method 100 further includes a step 104 of generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model (LLM).

[0148] In some embodiments and as will be described in the Example below, the LLM may be trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage. Where the medical diagnosis relates to one or more ophthalmic diseases or conditions, the LLM may be further trained using ophthalmic imaging data in the first stage and / or the second stage.

[0149] The method 100 may optionally include steps 106, 108 and 110 as illustrated in FIG. 15. In step 106, the method 100 may include a step of generating a medical report for the patient comprising the medical diagnosis generated using the LLM. The medical report may be generated automatically or by an operator or user (such as a clinician). In steps 108 and 110, the method 100 may include the steps of comparing the medical diagnosis generated using the LLM with a medical diagnosis from a clinician, and determining a medical diagnosis for the patient based on the comparison. The steps 108 and 110 may be performed automatically or by an operator or user (such as a clinician).

[0150] The method 100 according to some embodiments of the disclosure may be a computer-implemented method in which one or more computer processors perform the steps. For example, the natural -language prompt and the set of data may be received by a computer processor at step 102. The natural -language prompt and the set of data may be received at the computer processor via a network 230 or from an electronic device 224 having a display 226 (see FIG. 16). Referring to FIG. 15, once the natural-language prompt and the set of data are received, the one or more computer processors may then perform the method 100 including the steps 104, 106, 108 and 110 of the method 100 as shown and described in relation to FIG. 15.

[0151] In other embodiments, the method 100 may be a computer-assisted method in which one or more computer processors perform only some of the steps of the method 100. The computer-assisted method may perform at least the steps 102 and 104 of the method 100 of FIG. 15. The medical report generated at step 106 may be generated either by the one or more computer processors or by an operator or user (e.g., a clinician). Similarly, the steps 108 and 110 of comparing the medical diagnosis generated using the LLM with a medical diagnosis from a clinician, and determining a medical diagnosis for the patient based on the comparison may be performed either by the one or more computer processors or by an operator or user (e.g., a clinician). In this way, some embodiments of the disclosure may provide computer-assisted methods in which the LLM model generates a medical diagnosis for the patient and an operator or user (e.g., a clinician) uses the generated medical diagnosis to assist with their diagnosis of the patient. As discussed in the Example, these computer-assisted methods provided by the EyeGPT models have been shown to increase the accuracy of medical diagnosis of ophthalmic diseases or conditions provided by clinicians.

[0152] FIG. 16 provides a schematic diagram illustrating a system 200, which includes a computer system 220 with a processing unit 206 that receives the natural -language prompt and the set of data from the patient, and performs one or more steps of embodiments of the method 100 as shown and described above in relation to FIG. 15. The computer system 220 includes a memory 212 and one or more programs 210 stored in the memory 212 that may include instructions for performing one or more of the steps of the method 100 of embodiments of the disclosure. The system 200 may also optionally include an electronic device 224 having a display 226. The electronic device 224 can communicate with the computer system 220. In some embodiments (not shown), the computer system 220 is included in the electronic device 224.

[0153] Referring to FIG. 16, the system 200 includes a computer system 220 in the form a general -purpose computing device. The components of computer system / server 220 may include, but are not limited to, one or more processors or processing units 206, a system memory 212, and a bus 232 that couples various system components including system memory 212 to processor 206.

[0154] Bus 232 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnects (PCI) bus.

[0155] Computer system 220 typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system / server 220, and it includes both volatile and non-volatile media, as well as removable and nonremovable media.

[0156] System memory 212 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 214 and / or cache memory 216. Computer system / server 220 may further include other removable / non-removable, volatile / non- volatile computer system storage media. By way of example only, storage system 218 can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a “hard drive”), and other non-removable, non-volatile media (e.g., a “solid-state drive”). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a “floppy disk”), and an optical disk drive for reading from and / or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD- ROM or other optical media can be provided. In such instances, each can be connected to a bus 232 by one or more data media interfaces. As will be further described below, memory 212 may include a computer program product storing a set (e.g., at least one) of program modules 211 comprising computer readable instructions configured to carry out one or more steps of methods of the present disclosure.

[0157] Program 210, having a set (at least one) of program modules 211, may be stored in memory 212 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating systems, one or more application programs, other program modules, and program data or some combination thereof, may include an implementation of a networking environment. In some embodiments, program modules 211 are adapted to generally carry out the one or more functions and / or methodologies of one or more embodiments.

[0158] Computer system 220 may also communicate with one or more external devices such as a keyboard, a pointing device, a display, etc.; one or more devices that enable an operator to interact with computer system 220 or system 200; and / or any device (e.g., network card, modem, etc.) that enable computer system 220 to communicate with one or more other computing devices. Such communication can occur via Input / Output (VO) interfaces 222. FIG. 16 illustrates that the computer system 220 may communicate with an electronic device 224 including a display 226, and optionally other devices such as a keyboard, a pointing device etc. In some embodiments (not shown), the computer system 220 may be located on-board the electronic device 224.

[0159] The computer system 220 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via network adapter 228. As depicted, the network adapter 228 communicates with other components of computer system 220 via bus 232. It should be understood that although not shown, other hardware and / or software components could be used in conjunction with computer system 220. Examples include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0160] In another aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium 212 storing one or more programs 210 which comprise instructions, which when executed by one or more processers 206 of an electronic device 224 having a display 226, cause the electronic device 224 to perform the method 100 according to any one of the embodiments of the present disclosure. Furthermore, all steps and features described in relation to the method 100 of FIG. 15 may be included in embodiments of steps stored in program code of the one or more programs 210.

[0161] The computer-readable storage medium 212 can be a tangible device that can retain and store instructions for use by an instruction execution device, such as a memory device 212 shown in FIG. 16. The computer-readable storage medium 212 may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium 212 includes the following: a portable computer diskette, a hard disk, a random access memory (RAM) 214, a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer- readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0162] Computer-readable program instructions described herein can be downloaded to respective computing / processing devices 206 from a computer-readable storage medium 212 or to an external computer or external storage device via a network 230, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network 230 maycomprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface 228 in each computing / processing device 206 receives computer readable program instructions from the network 230 and forwards the computer readable program instructions for storage in a computer-readable storage medium 212 within the respective computing / processing device 206.

[0163] Computer-readable program instructions for carrying out operations may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language, Python or similar programming languages. The computer-readable program instructions may execute entirely on the operator’s computer, partly on the operator’s computer, as a stand-alone software package, partly on the operator’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the operator’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In one or more embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0164] Aspects are described herein with reference to flowchart illustrations and / or block diagrams of methods, systems, computer-readable storage media and computer program products according to one or more embodiments. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0165] These computer-readable program instructions may be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. Thesecomputer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0166] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0167] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, computer- readable storage media and computer program products according to various embodiments. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0168] In some embodiments, the application further comprises a software module making a medical treatment recommendation based on a medical diagnosis provided by a method disclosed herein. In some embodiments, the medical diagnosis is for a disease or disorder selected from the group consisting of age-related macular degeneration, diabetic retinopathy, glaucoma, cataract, myopia, retinal vein occlusions, kidney disease, hypertension, and stroke. In some embodiments, the subject has not manifested visible abnormalities or symptoms of the ophthalmic or systemic disease, disorder, or condition at the time the medical diagnosis is provided.

[0169] As used herein, a "disease" can refer to disease and additionally disorders and / or conditions. Examples of ophthalmic diseases include macular degeneration, diabetic retinopathy,glaucoma, cataract, myopia, and retinal vein occlusions. Examples of systemic diseases include diabetes and its kidney disease, hypertension and stroke. In some embodiments, the systems and methods disclosed herein diagnose a range of blinding diseases such as AMD, DR, HM, RVO, glaucoma, and cataracts. In some embodiments, disclosed herein is a model that shows high accuracy amongst all disease phenotypes, including in different cohorts such as in China, US, and Brazil. In some embodiments, the model is able to diagnose or differentiate diabetic patients without retinopathy from non-diabetic patients with normal retinas. In some embodiments, the model is able to detect underlying systemic complications by distinguishing diabetic patients with diabetic kidney disease from diabetic patients without any known complications. In some embodiments, the model demonstrates high accuracy in detecting unique patient characteristics such as age, gender, blood pressure and stroke. In some embodiments, these techniques serve as non-invasive screening tools as part of a health check-up or medical screening.

[0170] In certain aspects, the trained LLM disclosed herein is used for analyzing medical imaging data. In some embodiments, the medical imaging data comprises ophthalmic images, which can include images of the internal structure of the eye such as the retina and / or retinal vasculature, macula, and optic nerve. The framework described herein is applicable to various types of medical imaging including ophthalmic imaging. Ophthalmic imaging is a type of medical imaging that scans or captures one or more structures of the eye. In some embodiments, the trained LLM is used to analyze ophthalmic images generated using at least one ophthalmic medical imaging technique selected from optical coherence tomography (OCT), color fundus photography of the retina (CFP), corneal topography, slit-lamp photography, fluorescein angiography, indocyanine green angiography, fundus auto-fluorescence, optic nerve head analysis, endothelial cell-layer imaging, and external imaging. In some embodiments, ophthalmic images are generated using specialized imaging equipment. However, the requirement of specialized imaging equipment can increase the burden of obtaining a diagnosis, especially in low-income and / or underdeveloped areas. Accordingly, in some instances, nonspecialized imaging can be utilized to provide accurate detection or one or more ophthalmic or systemic diseases or disorders. For example, fundus imaging of the retina using an ophthalmoscope and a standard digital camera such as in a smart phone can be implemented to perform the analysis described herein to obtain detection or diagnosis or a disease or disorder. In some embodiments, a digital retinal camera is used for color fundus photography and / or fluorescein angiography. In some embodiments, an optical coherence tomography enables cross- sectional imaging of the retina such as for macular and optic nerve head imaging. In some embodiments, a scanning laser ophthalmoscope is used for fundus autofluorescence, fluorescein angiography and indocyanine green angiography. In some embodiments, photo slit-lampmicrography is used to photograph anterior eye structures (e.g. cornea, iris, conjunctiva, and lens). In some embodiments, corneal topography is used to measure the thickness, refractive power, and shape of the cornea. In some embodiments, an optic nerve head analyzer is used for optic nerve head imaging. In some embodiments, external photography is used to image the exterior of the eye, eyelid, or other structures in proximity to the eye. In some embodiments, a Rostock corneal module (RCM) is used to generate high-magnification images of the corneal layers, which allows counting of endothelial cells.EXAMPLESExample 1 - Generative Foundation Model for Medical Uses

[0171] Al tools are used to interpret various types of medical data, such as dermoscopic images, retina images, EHRs, electrocardiograms, and oncology trials. While these models are adept at their specialized tasks, they often struggle with diagnostic duties that span multiple disciplines. For instance, a cardiology-focused Al model might overlook neurological symptoms in neurology. Such “tunnel vision” can potentially lead to missed diagnoses or an incomplete understanding of a patient’s holistic health needs. Without a broad perspective or knowledge base, these tools may compromise comprehensive patient care. Moreover, the development of Al models requires the integration of a large amount of structured data, while the process of structuring medical data usually relies on extensive expertise and customized data processing procedures. For example, before applying Al to EHRs, one would normally need to convert the heterogeneous raw data into well-structured inputs, which is not only labor-intensive but also prone to information loss. Moreover, as more data are needed, this scheme could restrict the scalability of building more advanced Al systems. Addressing these challenges requires a foundation Al model that combines specialized insights with a holistic overview and requires minimal artificially structured data for training.

[0172] According to embodiments of the present disclosure, large language models (LLMs) such as GPT-4 (Generative Pre-trained Transformer-4) and BERT (Bidirectional Encoder Representations from Transformer) can be used in tasks like medical question answering (QA), report generation, and clinical decision support. When these models are trained predominantly on general internet-based knowledge, which often lacks the specialized context required for high-stakes medical applications, challenges remain in areas like rare disease diagnosis, emergency condition identification, and multimodal data integration.

[0173] Generative foundation models in the examples were developed through training on a vast array of EHRs across various health systems and an extensive collection of medical texts from both local and online sources. This ensures that the generative foundation models possess abroad and thorough understanding of medical theories and practices. In addition, the generative foundation models were trained to enhance their domain knowledge which leverage medical datasets like PubMed articles, EHRs, and textbooks. The generative foundation models have the potential to offer accurate decision support across a wide spectrum of diagnostic scenarios, tackling diverse challenges within the medical field, including rare disease diagnosis, emergency condition identification, and multimodal report generation. An example of the generative foundation models is Meta General Practitioner (MetaGP). FIG. 1A illustrates the three stages of MetaGP development: pretraining, multimodal instruction fine-tuning, and application.Pretraining: MetaGP was trained on large-scale datasets, including EHRs, academic articles, and medical books, using advanced processes like data aggregation, tokenization, pretrained language model (Qwen) initialization, and continual pretraining to build a strong foundation in medical knowledge. Multimodal instruction fine-tuning: this phase incorporated domain-specific datasets, such as medical QA, rare disease, and emergency EHRs and multimodal imaging data, to optimize task-specific outputs through text and image encoders. Application: MetaGP was applied to diagnostic tasks, clinical decision support, and multimodal report generation. The evaluation included real-world cases from various clinical scenarios.Developmental framework of MetaGP

[0174] The development of MetaGP began with Qwen-1.5 32B as its base architecture (FIG. IB), leveraging its advanced LLM capabilities. During the rigorous pretraining process, MetaGP utilized a diverse and comprehensive dataset comprising 8 million EHRs (as detailed in Table El below) (14.8 billion tokens), 5.4 million academic articles (48 billion tokens), and 15,731 medical books (8.6 billion tokens). This pretraining phase encompassed data aggregation, tokenization, model pretraining, and continual pretraining, enabling MetaGP to efficiently process long-context data and encode intricate medical knowledge. These processes established a solid foundation, empowering MetaGP to meet the complex demands of clinical diagnostics with precision and adaptability.

[0175] The fine-tuning phase further tailored the model for real-world clinical applications. MetaGP was refined with a focused dataset that included approximately 630,000 EHR records for rare diseases and emergency conditions; multimodal imaging data comprising over 600,000 chest X-rays (CXRs) and nearly 24,000 computed tomography (CT) scans; medical QA questions (Table E2) sourced from established databases such as PubMedQA, MedQA, and MedMCQA; and natural language data from datasets such as the OpenAssistant Conversations dataset (TOAC) and Chain-of-Thought Collection dataset (CoTC). This phase optimized MetaGP for specific tasks, such as recognizing rare disease patterns, identifying life-threateningemergencies, and processing complex clinical queries, ensuring its adaptability to diverse diagnostic challenges.Table El: EHR data characteristicsTable E2: Templates for applying multimodal instruction fine-tuning and evaluatingPubMed case reports

[0176] The generative foundation model was trained using a two-phased approach. In the example of MetaGP, in the initial phase an LLM was pretrained on an expansive corpus that included over 8 million EHRs, 5.4 million academic papers from PubMed, and 15 thousand medical textbooks. This phase ensured the model had an excellent grounding in a wealth of knowledge of medical theories. In the second phase, extensive question-answering (QA) datasetswere gathered and employed in conjunction with EHR data for cases involving rare diseases and emergency conditions in the real world. The health records include detailed case descriptions and decision-making processes. By finetuning the model with these data, more applied clinical insights were embedded, enhancing its ability to make differential diagnoses. For the multimodal medical report generation system, a comprehensive dataset of approximately 630,000 diagnostic reports were curated and paired with medical images from two imaging modalities: chest X-rays (CXR) and computed tomography (CT).

[0177] Pretraining data:

[0178] The pretraining dataset is derived from three main sources of medical data: EHRs, medical articles, and medical books. The EHRs provide real-world clinical insights, while the medical articles and books contribute scientific knowledge and certified medical information. By combining these diverse sources, the dataset aims to capture a comprehensive representation of the medical domain. The resulting dataset, after preprocessing and integration, contains a substantial total of 71.4 billion tokens, which forms a robust foundation for the pretraining process.

[0179] EHR data (public and private). The data utilized in this study were sourced from comprehensive hospitals. Consequently, the EHRs encompassed data from all departments. Emergency and rare diseases present considerable challenges in clinical practice. To address this, the ICD-10 codes were collected based on emergency department records and physician experience to mark emergency disease records. Keyword matching was employed against the Orphanet and OMIM databases for disease diagnosis, followed by manual curation by physicians to mark rare disease records. The EHRs contained comprehensive inpatient records, including outpatient clinic notes documenting patients’ symptoms, diagnoses, and treatment plans, as well as hospital admission and discharge records detailing treatments and discharge summaries. These records also incorporated crucial laboratory test results, such as blood tests and imaging studies, for diagnosing conditions and monitoring treatment effectiveness, along with prescription information and details of any surgical interventions or therapies.

[0180] Medical articles (public). A dataset of medical articles was constructed from PubMed Central (PMC), a subset of the PubMed online repository managed by the United States National Center for Biotechnology Information (NCBI). PMC provides open, full-text access to biomedical and life sciences literature. 5.4 million articles published up to June 20, 2023, were extracted, resulting in a total of 48 billion tokens after tokenization. The construction method was based on the PubMed Central subset from the Pile dataset, with updates to include more recent articles. Note that for evaluation purposes, case reports of rare diseases, emergencies, andcomplex diseases were removed from the pretraining data based on their unique identifiers (i.e., PMID).

[0181] Medical textbooks (public). To incorporate certified medical knowledge into the training process, a total of 15,731 publicly available books related to biomedicine were collected. These books were sourced from various channels, including university-subscribed platforms like ScienceDirect and the Wiley Online Library, as well as websites that provide open access to downloadable medical e-books, such as the Directory of Open Access Books. Keywords derived from the Medical Subject Headings (MeSH) were used to ensure the relevance and comprehensiveness of the search. For books that could not be directly extracted, the Nougat method was employed to extract the relevant content. The text was preprocessed to remove non-medical content, resulting in approximately 8.6 billion tokens.

[0182] Finetuning data:

[0183] In this phase, the aim is to elicit knowledge from the pretrained model by applying supervised finetuning with a collection of high-quality question-answering (QA) pairs. The finetuning data comprise medical QA data, natural language data, and EHR data.

[0184] Medical QA data (public). PubMedQA is a dataset designed for the development and evaluation of machine learning models in the QA domain, specifically targeting biomedical literature. It leverages abstracts from the vast PubMed database. MedQA (USMLE) focuses on medical examination questions derived from the United States Medical Licensing Examination (USMLE). It simulates the breadth and depth of medical knowledge required for licensure, making it an ideal dataset for training and evaluating Al models on medical reasoning and knowledge. MedMCQA, on the other hand, is a comprehensive multiple-choice QA dataset that covers a wide range of medical subjects.

[0185] Natural language data (public). The OpenAssistant Conversations dataset is a corpus of dialogues with 161,443 messages, used for basic model -based QA and multi -turn dialogue abilities, which are crucial for interacting with the model during diagnosis. The CoT- Collection dataset includes 1.88 million Chain-of-Thought (CoT) formatted QA data for 1,060 tasks, which helps improve the model’s performance on complex analytical tasks such as differential diagnosis. 500,000 samples from the CoT-Collection dataset were randomly selected and added to the training set for finetuning.

[0186] EHR data (private). In an effort to bridge the gap between QA and EHR data, a distinct model for EHRs of rare disease cases and emergency conditions was finetuned. The comprehensive descriptions, diagnostic basis, and treatment plans provided by physicians significantly aid the model in understanding the logic of diagnosis and treatment. Thus, the data were filtered based on its completeness, selecting records that included comprehensiveinformation on gender, age, chief complaint, medical history (present illness history, history, allergen history, personal history, family history), examination information (physical examination, specialist examination, auxiliary examination), main diagnosis, diagnosis basis, differentiated diagnosis description, and treatment plan for model finetuning. Following the filtration process, the finetuning rare disease data contained 264,445 records, and the finetuning emergency conditions data contained 368,369 records.

[0187] Evaluation data:

[0188] To avoid data contamination, the studies of rare diseases and emergencies were removed for evaluation purposes from the pretraining / finetuning data, based on their unique identifiers (i.e., PMID). A combination of Pubmed cases, as well as private EHR cases, was used as the validation cohorts.

[0189] Case reports (public). To enhance the real-world applicability of evaluation, case reports of rare diseases and emergency conditions from leading medical journals derived from PubMed were incorporated to calculate the subjective evaluation score (-2 ~ +2). 97 rare disease cases and 109 emergency cases were compiled from the American Journal of Emergency Medicine, Journal of American College of Emergency Physicians, North American Journal of Medical Sciences, and American Journal of Medical Genetics (2012-2022). These cases were evaluated by MetaGP, GPT-4 and physicians specializing in emergency medicine and medical genetics. The experts provided insights into the final diagnosis, differential diagnoses, and reasoning behind their conclusions.

[0190] EHR data (private). For objective evaluation, a subset of private EHR data, which was randomly divided into training and test sets, was utilized. The test set comprised 958 rare disease cases and 2,769 emergency cases. This data was used to evaluate the model's structured output, such as ICD-10 code predictions, by calculating metrics like accuracy, Fl score, precision, and recall.

[0191] Image-report data:

[0192] To train the multimodal report generator, CXR and CT from the private and public datasets were collected and integrated, resulting in a collection of over 63,000 image-report pairs spanning 2 imaging modalities: CXR, and CT.

[0193] CXR (public and private). A portion of CXR images and report data were from the China Consortium of Chest X-ray Image Investigation (CC-CXRI). The official test split of MIMIC-CXR was used for interval validation, while the lU-Xray dataset served as the external validation set.

[0194] CT (private). CT scans and report data were from cohorts of the China Consortium of Chest CT Image Investigation (CC-CCII).

[0195] Framework:

[0196] Although general-purpose LLMs like Qwen-1.521 and GPT-4 have showcased strong performance across various tasks in benchmarks such as BIG-bench, their utilization in the medical field necessitates adaptation and alignment with domain-specific data due to the inadequacy of domain knowledge. Hence, a two-stage training strategy comprising pretraining and supervised finetuning was implemented to enrich the language model with more medical knowledge and medical abilities.

[0197] Pretraining. MetaGP was initialized with the parameters of Qwen-1.5 32B and subsequently further pre-trained on medical text corpora, which encompassed medical articles from PubMed, text content from publicly available medical books (including guidebooks and medical genetic textbooks), and clinical text data (mainly EHRs). The corpora were tokenized and concatenated into a sequence of tokens, which were then chunked into fixed-length pieces for training. Given a token sequence} , the optimization objective is to minimize auto-regressive loss formulated asmodel, x<;denotes tokens before x, and N is the fixed sequence length.

[0198] Supervised finetuning. In the second training phase, the instruction tuning approach, defined as finetuning a pretrained large language model using a dataset comprising instructional inputs and their corresponding responses, was incorporated. The instructions used for medical QA and EHR data are presented in Table El. The training loss was the same auto-regressive loss in the pretraining stage and calculated based on the output tokens. For the remaining data, each sample consists of one or multiple rounds of dialogue content, where each round of dialogue includes instructions describing the task, optional instance inputs providing supplemental information required for task resolution, and the expected output corresponding to that round of dialogue. The training loss was computed based on the full dialog tokens.

[0199] Long context window. The default context window size of Qwen-1.5 is 32,678. However, in downstream medical scenarios, there are many predictive tasks involving long texts, such as differential diagnosis based on extensive patient information. In these situations, the total length of patient information exceeds the default context window size of Qwen-1.5. In both training phases, linear rope scaling was employed to extend the context window, preventing performance degradation when handling inputs that surpass the pretrained context window size during text completion. Specifically, position indices in Rotary Position Embedding (RoPE) were directly downscaled to match the previous context window limit and calculated the intermediate values of the position encodings between adjacent integer positions.

[0200] Medical report generator. To deal with the varying dimensions across different imaging modalities, two vision encoders with different architectures - Swin Transformer andViT - were appended on top of MetaGP. Swin Transformer is an architecture in the field of computer vision that utilizes shifted windows to provide a hierarchical and efficient way of processing images. This design enables the model to focus on both local and global features of an image, making it particularly effective for vision tasks. In practice, the developed medical report generator used a Swin Transformer to process medical images. However, the Swin Transformer was designed to handle 2D image data and is not capable of tackling 3D volumetric arrays. For 3D imaging data (e.g., CT scans), a ViT-3D encoder was used due to its efficiency in processing volumetric data, which is faster than treating CT scans as image sequences with Swin. The core idea is to first divide large 3D volumes into smaller ones, on top of which visual representation learning was further applied. In the pipeline, ViT-3D was employed to deal with CT to capture relations and dependencies across images or slices. After transforming raw pixels or voxels into latent embeddings using vision backbones, the obtained visual representations were forwarded to MetaGP. However, images or volumes are high-dimensional data, and converting them into dense embeddings creates a challenge due to the generation of lengthy token sequences (e.g., 768 tokens were generated by ViT-3D), which impairs the training and inference efficiency of the following MetaGP. To address this issue, dimension reduction was applied to adjacent visual tokens and effectively decreased the visual token count by a factor of 6 at least. In practice, this strategy can dramatically reduce memory costs and speed up model training. After obtaining the reduced visual tokens, linear projectors were used to map visual representations to the language space. In addition to visual tokens, the language input also comprises demographic information and textual tags that indicate the input modalities. All of these were combined and fed to the MetaGP in each forward pass. Since parameter finetuning was applied to MetaGP, this operation resulted in 20M trainable parameters, along with a 2D vision encoder that has 88M parameters and a 3D vision encoder that has 100M parameters. So the total number trainable parameters of the report generator is about 200M.

[0201] Implementations:

[0202] MetaGP. The scaling factor used in linear rope scaling was 2, resulting in a context window of 65,535 for MetaGP. During the training phase, a global batch size of 384 was used. Optimization was performed using AdamW with an initial learning rate of le-5. The loss function employed was autoregressive loss. DeepSpeed Zero-3 was utilized to accelerate the training process, based on the bfloatl6 format, while also implementing the gradient checkpointing strategy to increase the batch size. In the foundation phase, the training process was conducted for one epoch using 120 NVIDIA A100 GPUs (80G), while in the instructiontuning stage, the model was trained for 3 epochs using 48 A100 units. For GPT-4, the version of 2024-01-25-preview was used for model testing and evaluation.

[0203] MetaGP for report generation. The base version of Swin Transformer, which has 4 stages, a window size of 7, a patch size of 4, and an initial feature dimension of 128, was used. For data of CXR, a random portion of each image was first cropped, where the area of the cropped region with respect to that of the original image is between 0.5 and 1.0. Then, each crop was resized to 224 (Height) *224 (Width) pixels with 3 channels. Pretrained weights on ImageNet were used. The ViT-3D model has 12 layers and an embedding dimension of 768. Each input volume was first resized to 224 (Height) x 224 (Width) x 64 (Depth) and further divided into subvolumes with no overlap, where the size of each subvolume is 16 (Height) x 16 (Width) x 8 (Depth). To reduce the count of visual tokens, the adaptive average pooling strategy was used to fix the output length to 9. Two linear projectors for 2D and 3D data, respectively, were implemented, where each projector includes one fully connected layer, mapping each pooled visual token to a ID vector that has 4096 elements. Before the projector, layer normalization was added to avoid exploding gradients. End-to-end training was applied, where all parameters of vision encoders and linear projectors are trainable. For the LLM part, the LoRA, i.e., low-rank adaptation strategy, was used for parameter-efficient training. LoRA relies on the mathematical concept of low-rank matrix decomposition. In this approach, a large weight matrix of a neural network layer (which is typically high-rank) is approximated using two smaller matrices. This decomposition reduces the number of trainable parameters. By using low- rank matrices, LoRA modifies only a small fraction of the model’s parameters during the training stage. To achieve this, the rank and alpha value of LoRA was set to 16, respectively. AdamA was used as the default optimizer along with the cosine learning rate scheduler. The initial learning rate was set to le-4, and the minimum learning rate was 3e-6. The total number of training iterations is 80,000. Linear warmup was applied in the first 2,000 training iterations, and the warmup learning rate is le-7. The training stage was finished within 48 hours using 16 NVIDIA Al 00 GPUs (80G) with a global batch size of 64.Comprehensive evaluation framework for MetaGP

[0204] MetaGP demonstrated its utility in various clinical settings through practical applications. The model excelled in generating clinically relevant insights across diagnostic applications, including rare disease identification and emergency condition recognition. It also provided clinical decision support, aiding physicians by enhancing diagnostic accuracy and assessing risk severity. Additionally, MetaGP facilitated multimodal integration, such as CXR report generation and CT report generation, underscoring its versatility and potential to streamline diagnostic workflows across diverse clinical environments. These findings highlight MetaGP’ s role as a transformative tool in advancing medical Al.

[0205] To evaluate MetaGP’s clinical performance, a multi-faceted assessment approach was implemented, incorporating diverse datasets and rigorous methodologies (FIGS. 2A-2C). This framework thoroughly examined the model’s ability to diagnose rare diseases, identify emergency conditions, and generate multimodal medical reports. Evaluations combined manual grading by physicians with automated metrics, offering a balanced perspective on MetaGP’s capabilities in real-world applications.

[0206] MetaGP’s evaluation began with clinical datasets, including private EHR data and public case reports from PubMed (FIG. 2A). These datasets were categorized into rare diseases and emergency conditions, representing a wide range of medical challenges. Manual assessments compared Al performance with human diagnostic capabilities, while automated metrics such as Fl score, accuracy, recall, and precision provided quantitative insights. The scoring system, as summarized in FIG. 2B, ranged from -2 to +2, with +2 indicating a correct primary diagnosis. For example, in a case of arrhythmogenic right ventricular cardiomyopathy (FIG. 2B), MetaGP accurately identified the condition and received a high score (+2) from physicians, underscoring its effectiveness in handling complex diagnoses.

[0207] For radiologic imaging, including CXRs and CT scans, evaluations incorporated both manual grading and automated metrics (FIG. 2A). Manual assessments focused on diagnostic accuracy, clinical relevance, and quality, while metrics such as ROUGE-L (Recall-Oriented Understudy for Gisting Evaluation-Longest Common Subsequence) and BertScore validated performance across internal, external, and prospective studies. For instance, in CT imaging comparisons (FIG. 2C), MetaGP demonstrated superior accuracy in diagnosing bilateral lung infections with potential viral pneumonia. The Al-generated report aligned closely with clinical observations and outperformed the physician-generated reference report. This example highlights MetaGP’s ability to enhance diagnostic precision, particularly in radiologic assessments.

[0208] Evaluation criteria for clinical cases:

[0209] The quality scoring system evaluates diagnoses made by MetaGP, GPT-4, and participating physicians. Three senior physicians, each with over 20 years of experience, independently rated the final diagnoses produced by each model and physician in a blinded comparison to the actual diagnosis. For each case, the final evaluation score was determined by majority consensus among the three physicians. The scoring system is structured as follows: a score of +2 is assigned if the model’s or physician’s diagnosis precisely matches the actual primary diagnosis or an equivalent. A score of +1 is awarded if the actual diagnosis appears within the predicted differential diagnosis list, the primary diagnosis is contextually relevant, and includes partially accurate information. A score of 0 is given when the actual diagnosis isabsent from the differential list, and the primary diagnosis is irrelevant, but without any inaccurate or potentially harmful guidance. A score of -1 indicates a misdiagnosis that deviates from the actual diagnosis and could pose some risk to the patient’s health if the suggested treatment were implemented. Finally, a score of -2 reflects a significant misdiagnosis where the prescribed treatment, if followed, could result in severe harm or danger to the patient. Notably, only the primary diagnosis for each case study or health record is evaluated within this framework.

[0210] Evaluation criteria for imaging reports:

[0211] To rigorously assess the quality of medical reports generated by MetaGP, a structured physician grading system specifically for chest X-ray (CXR) and computed tomography (CT) scan reports was implemented. This task was assigned to eight radiologists, comprising (i) senior physicians with over 15 years of clinical experience in radiology and (ii) junior physicians (seven radiologists), each possessing at least 5 years of experience. These evaluators were carefully chosen to ensure depth of expertise and reliability in the assessment process. The evaluation was conducted as a blinded review, in which each radiologist was provided with both the original medical image and the accompanying report. To eliminate potential bias, all identifying information was removed from the reports. Radiologists assessed the reports, judging the Al-generated and human-written versions based on quality, diagnostic accuracy, and clinical relevance. They were also given the option to rate both reports as being of equal quality, enabling a balanced and thorough comparison. This method ensured a high- standard, unbiased evaluation of the model’s performance against traditional human reporting.

[0212] Automated evaluation criteria:

[0213] To assess diagnostic outcomes using electronic health records (EHRs), MetaGP and GPT-4 generated ICD-10 codes alongside final diagnoses. Predicted codes were matched to actual codes to calculate accuracy and Fl score, defined as: accuracy (ACC) is the proportion of correct predictions among all predictions, and Fl score is the harmonic mean of precision (true positives over predicted positives) and recall (true positives over actual positives). For medical reports, ROUGE-L and BertScore measured semantic similarity with reference texts: ROUGE-L assesses text overlap using the longest common subsequence, and BertScore uses BERT embeddings for precision, recall, and Fl, capturing semantic similarity. This framework rigorously evaluates model performance, ensuring transparency and reproducibility.

[0214] MetaGP vs. medical professionals:

[0215] MetaGP was compared with general practitioners and emergency physicians in assessing clinical cases of rare diseases and emergency conditions. Three groups of physicians were employed to participate in the study: 4 in the junior group with at least 5 years of clinicalexperience, 4 in the mid-level group with at least 10 years of experience, and 4 in the senior group with at least 15 years of clinical experience. The final evaluation score was established based on a consensus from an independent group of 3 senior physicians with 20 or more years of clinical experience.

[0216] MetaGP and physician collaboration:

[0217] To investigate whether MetaGP could help physicians improve their diagnostic performance, particularly in avoiding a misdiagnosis which may lead to harmful or dangerous outcomes, each participating physician was given the diagnostic output from MetaGP on each case and was asked to make a diagnosis with the assistance of the Al-generated result. To avoid a potential memorization bias, the follow-up Al-assisted diagnostic test was performed four weeks after the initial test.

[0218] Quantification and statistical analysis:

[0219] The study employed various statistical methods to evaluate MetaGP's diagnostic performance and multimodal capabilities. For rare diseases and emergency conditions, diagnostic accuracy was assessed using metrics such as Fl score, precision, recall and accuracy. These automated metrics were calculated based on the predictions of MetaGP compared to ground truth labels derived from expert annotations or clinical outcomes. Performance on imaging modalities, including chest X-rays and CT scans, was further validated using ROUGE- L and BertScore, quantifying the alignment between Al-generated reports and reference reports composed by experienced radiologists. Manual evaluations were conducted using a structured scoring framework ranging from -2 to +2, reflecting the clinical relevance, accuracy and potential risks of MetaGP’s diagnostic outputs. Comparative analyses were performed to benchmark MetaGP against GPT-4, BERT and human practitioners of varying experience levels.Example 2 - Diagnostic Performance of Generative Foundation Model in Rare Diseases

[0220] In the realm of medical diagnostics, identifying rare diseases poses a unique challenge due to their low incidence and often ambiguous symptom profiles. Diagnosing such diseases with Al demands highly specialized expertise and the capacity to detect subtle, nuanced patterns. Within this context, MetaGP leverages extensive training across diverse datasets to identify subtle, pattern-based correlations that might be missed through traditional diagnostic approaches.

[0221] In the manual evaluation section, a subjective score was obtained using 97 rare systemic disease cases from PubMed, comparing the diagnostic performance of MetaGP, GPT- 4, and general practitioners with varying levels of practicing experiences. MetaGP demonstratedsuperior diagnostic accuracy, achieving an average score of 1.57, compared to GPT-4’s score of 0.93 (FIG. 3A). General practitioners’ performance varied with experience. Junior practitioners scored 0.94; mid-level practitioners scored 1.02, surpassing GPT-4; and senior practitioners reached a score of 1.5. This outcome underscores the advantage of MetaGP’s specialized medical training, enabling it to surpass the diagnostic performance of GPT-4 and align closely with that of senior general practitioners. MetaGP’s efficacy is further emphasized by the fact that 76.3% of its diagnoses (74 cases) received the highest score of +2, indicating high accuracy and clinical relevance. Additionally, 8 cases (8.25%) achieved a score of +1 for providing useful diagnostic insights. While 12 cases (12.4%) received a score of 0, indicating neutral but nonharmful information, two cases received a score of -1 (“suboptimal”) and one case a score of -2 (“potentially harmful”), highlighting areas for further system refinement (FIG. 3B). These findings reinforce the need for ongoing development and validation of Al systems like MetaGP to ensure safe, effective integration into clinical practice.

[0222] To investigate whether the Al system could help general practitioners improve their diagnostic performance, particularly in avoiding misdiagnosis, each general practitioner was given the diagnostic output from MetaGP on each case and was asked to make a diagnosis with the assistance of the Al-generated result. As shown in FIG. 3A, the diagnostic performance of physicians with varying years of experience was mostly improved with the assistance of MetaGP. For rare systemic diseases, the average scores for the junior, mid-level, and senior groups rose from 0.94, 1.02, and 1.5 to 1.52, 1.62, and 1.88, respectively. These results demonstrate that MetaGP can assist general practitioners in improving their diagnostic accuracy when faced with rare diseases, regardless of their level of experience. A notable outcome of the human-AI collaboration is the significant reduction in potentially harmful responses. As shown in FIG. 3B, while physician-provided responses may still occasionally contain biases that could lead to adverse outcomes, integrating MetaGP’s diagnostic output markedly reduced the frequency of responses receiving negative scores.

[0223] During the automated metrics evaluation phase, a subset of local EHR data comprising 958 rare disease cases was utilized, as illustrated in FIG. 3C. MetaGP demonstrated a clear advantage across all metrics, achieving a mean accuracy of 0.698, significantly outperforming GPT-4 (0.329) and BERT (0.149). For Fl score, MetaGP achieved 0.754, markedly higher than GPT-4 (0.419) and BERT (0.232). Additionally, MetaGP excelled in recall and precision, achieving 0.714 and 0.867, respectively, compared to GPT-4’s 0.368 and 0.651 and BERT’s 0.160 and 0.509. Per-class results and additional comparative data are detailed in FIGS. 5A-5C. These results highlight the impact of MetaGP’s specialized training in rare disease diagnosis, allowing it to effectively capture complex relationships betweensymptoms, patient history, and diagnostic outcomes. The consistent performance advantage across all metrics underscores its potential as a reliable tool for tackling the challenges of rare disease identification.Example 3 - Diagnostic Performance of Generative Foundation Model in Emergency Conditions

[0224] In managing emergency conditions, particularly those that pose immediate threats to life, rapid and precise assessment is paramount. Emergency physicians often operate under severe time constraints, requiring swift decision-making to ensure patient safety. MetaGP addresses this need by efficiently analyzing patient data, offering valuable support to physicians for timely and accurate diagnoses.

[0225] In the manual evaluation of 109 cases on systemic emergency conditions, MetaGP showcased strong diagnostic performance, achieving an average score of 1.59, significantly surpassing GPT-4’s score of 1.19. (FIG. 4A). Emergency physician performance varied by experience: junior physicians scored an average of 0.97, and mid-level physicians aligned with GPT-4 at 1.15, while senior physicians achieved a significantly higher score of 1.61, underscoring MetaGP’ s comparability with seasoned medical professionals. MetaGP’ s diagnostic precision (FIG. 4B) is further emphasized by its highest score of +2 in 74.3% of cases (81 cases), reflecting exceptional accuracy and clinical utility. An additional 15 cases (13.76%) were rated +1 for providing valuable diagnostic insights. In 9 cases (8.26%), MetaGP’ s output was rated neutrally at 0, presenting neither particular benefit nor risk. Only 4 cases received a score of -1, indicating a “suboptimal” outcome, with no cases rated as “dangerous” (-2), demonstrating MetaGP’ s high reliability and minimal risk of adverse outcomes in emergency contexts. These findings, based on case reports from PubMed, highlight MetaGP’ s potential to support systemic emergency diagnosis with a substantial proportion of highly accurate responses, enhancing clinical decision-making in urgent scenarios.

[0226] How MetaGP’ s diagnostic capabilities impact the diagnostic accuracy of emergency medicine physicians was further examined. Adopting a methodology similar to that used for rare disease evaluation, physicians were enabled to review MetaGP’ s diagnostic suggestions prior to finalizing their own assessments. For systemic emergencies, diagnostic improvements were notable: junior physicians saw an increase of 53%, mid-level physicians improved by 46%, and senior physicians by 19% (FIG. 4A). The greater gains among junior and mid-level physicians suggest that MetaGP may be especially valuable in supporting less experienced clinicians in accurately diagnosing complex emergency cases. Moreover, MetaGP contributed to a reduction in potentially harmful diagnoses made by emergency physicians — a critical advancement inemergency care (FIG. 4B). This decrease in harmful diagnostics underscores MetaGP’s potential to enhance patient safety and optimize outcomes in high-stakes emergency settings, illustrating its capacity to serve as a valuable adjunct in acute medical decision-making.

[0227] During the automated metrics evaluation phase, a subset of local EHR data consisting of 2,769 emergency condition cases was utilized, as shown in FIG. 4C. MetaGP demonstrates a clear advantage across all metrics, achieving a mean accuracy of 0.702, compared to GPT-4’s 0.463 and BERT’s 0.219. Similarly, MetaGP’s Fl score reached 0.783, significantly outperforming GPT-4 (0.541) and BERT (0.322). For recall and precision, MetaGP achieved 0.715 and 0.889, respectively, further highlighting its superior diagnostic performance compared to GPT-4 (0.325 and 0.521) and BERT (0.155 and 0.458).

[0228] These results demonstrate MetaGP’s strong diagnostic accuracy, recall, and precision across a broad range of emergency scenarios, underscoring its promise for practical clinical applications. The consistent performance advantage across metrics and subcategories suggests MetaGP as a more reliable tool for high-stakes environments requiring rapid and accurate diagnostics. These findings highlight the impact of its specialized training in emergency condition tasks. Per-class results and additional comparative data are detailed in FIGS. 5D-5H.Example 4 - Performance of Generative Foundation Model in Multimodal Medical Reports

[0229] To evaluate the capabilities of MetaGP in handling multimodal data, a system was developed to generate medical reports for multimodal medical images, focusing on two widely used radiological modalities: CXR and CT.

[0230] The approach was benchmarked against two well-established multimodal medical foundation models, LLaVA-Med and Med-Flamingo, specifically assessing MetaGP’s performance in CXR and CT report generation. In CXR interpretation (FIG. 6A), MetaGP showed substantial advantages over both LLaVA-Med and Med-Flamingo. For CT data (FIG. 6B), MetaGP outperformed LLaVA-Med significantly in both internal and prospective validation.

[0231] Additionally, MetaGP -generated reports were evaluated by eight radiologists through a blinded review process, comparing Al-generated and physician-authored reports for CXRs and CT scans. Radiologists assessed report quality, diagnostic accuracy, and clinical relevance, with an option to rate both reports as equivalent. For CXR (FIG. 6C), MetaGP’s reports were either preferred or rated as equivalent in 53.8% of cases. For CT (FIG. 6D), ties were even more frequent, reflecting the comparable quality of Al and human reports. Interestingly, junior radiologists showed a higher preference for Al-generated reports compared to their seniorcounterparts (29.8% vs. 15.9%), highlighting MetaGP’s potential to assist radiologists across varying levels of expertise in interpreting complex imaging data.Example 5 - Generative Foundation Model EyeGPT for Diagnosing Ophthalmic Diseases

[0232] This example describes the development and deployment, in methods and systems of the present disclosure, of a generative foundation model (EyeGPT) for diagnosing ophthalmic diseases or conditions, particularly rare, emergency and complex ophthalmic diseases.

[0233] Eye Generative Pretrained Transformer (EyeGPT), a generative foundation model with 32 billion parameters for medicine, was developed through training on a vast array of electronic health records across various health systems and an extensive collection of medical texts from the internet. This ensures that EyeGPT possesses a broad and thorough understanding of medical theories and practices, thereby solving a technical problem in the lack of precise and reliable diagnostic support provided by existing LLMs within medicine, and specifically, ophthalmology. EyeGPT has the potential to offer accurate decision support across a wide spectrum of diagnostic scenarios, tackling diverse challenges within the medical field. As a proof of concept, the capabilities of EyeGPT were studied in addressing three complex, yet unmet clinical challenges: (i) rare disease diagnosis, (ii) emergency condition identification, and (iii) complex disease diagnostics, and specifically for ophthalmic diseases or conditions. Additionally, its performance in generating medical reports based on medical images was evaluated. To tackle the challenge of assessing the accuracy of generative medical Al predictions, stringent evaluation protocols were implemented and comprehensive testing with the help of healthcare experts was conducted.

[0234] EyeGPT’ s ability to enhance diagnostic performance is particularly important in the context of rare diseases and prevalent ophthalmological conditions. Glaucoma, a group of diseases characterized by damage to the optic nerve, is one of the leading causes of blindness worldwide. Similarly, diabetic retinopathy and age-related macular degeneration (AMD) are leading causes of vision loss among working-class adults and elderlies, respectively, and early identification of these diseases is crucial for managing progression and preventing severe central vision loss. Despite the scarcity of rare ophthalmic diseases, it is also a leading cause of visual loss in adolescents and in Europe, and timely diagnosis and intervention is essential to prevent vision loss. By synthesizing vast amounts of medical literature, imaging data, and clinical reports, EyeGPT assists in diagnosing these diseases by identifying subtle patterns in clinical presentations and imaging that may be overlooked. In doing so, EyeGPT can aid clinicians in making faster, more accurate diagnoses for both rare diseases and more common conditions, ultimately improving patient outcomes.

[0235] Materials and methods

[0236] Sex as a biological variable. The study examined male and female persons, and similar findings are reported for both sexes.

[0237] Model development overview. The training regime of EyeGPT employed a two- phased approach. In the initial phase, the EyeGPT was pretrained on an expansive corpus that included over 10 million EHRs, 5 million academic papers from PubMed, and 15 thousand medical textbooks. This phase ensured the model had an excellent grounding in a wealth of knowledge of medical theories. In the second phase, extensive question-answering (QA) datasets in conjunction with EHR data for cases involving rare diseases and emergency conditions in the real world were gathered and employed. The health records include detailed case descriptions and decision-making processes. By finetuning the model with these data, more applied clinical insights, enhancing its ability to make differential diagnoses were embedded. For the multimodal medical report generation system, a vast dataset comprising over 1 million diagnostic reports and associated medical images from 6 imaging modalities was curated, spanning disciplines such as ophthalmology and radiology.

[0238] Pretraining data:

[0239] The pretraining dataset is derived from three main sources of medical data: EHRs, medical articles, and medical books. The EHRs provide real-world clinical insights, while the medical articles and books contribute scientific knowledge and certified medical information. By combining these diverse sources, the dataset aims to capture a comprehensive representation of the medical domain. The resulting dataset, after preprocessing and integration, contains a substantial total of 71.4 billion tokens, which forms a robust foundation for the pretraining process.

[0240] EHR data (public and private). To encode knowledge of medical practices, 10,514,122 EHRs derived from the China Medical Al Investigation Consortium were collected, in which some of the clinical cohort data have been published. The data utilized in this study were sourced from comprehensive hospitals and specialized ophthalmology hospitals (Table E3). Consequently, the EHRs encompassed data from all departments, with a particular emphasis on ophthalmology. The inpatient records containing ophthalmic data comprise 597,039 records from 260,460 patients. Emergency and rare diseases present considerable challenges in clinical practice. To address this, the ICD-10 codes based on emergency department records and physician experience to mark emergency disease records were manually collected. Keyword matching against the Orphanet and OMIM databases for disease diagnosis was employed, followed by manual curation by physicians to mark rare disease records. Following this process, the rare disease included 447,925 records, while the emergency diseasecontained 1,114,554 records. The EHRs used in this study consist of comprehensive inpatient records, including outpatient clinic notes documenting patients’ symptoms, diagnoses, and treatment plans, as well as hospital admission and discharge records detailing treatments and discharge summaries. These records also incorporate crucial laboratory test results, such as blood tests and imaging studies, for diagnosing conditions and monitoring treatment effectiveness, along with prescription information and details of any surgical interventions or therapies. For the model to exhibit good generalizability across diverse populations including different countries and ethnicities, they were supplemented with 431,231 discharge summaries from MIMIC-IV. After removing duplications, about 14.8 billion tokens were obtained.Table E3: EHR Data Characteristics

[0241] Medical Articles (public). A dataset of medical articles from PubMed Central (PMC), a subset of the PubMed online repository managed by the United States National Center for Biotechnology Information (NCBI) was constructed. PMC provides open, full-text access to biomedical and life sciences literature. 5.4 million articles published up to June 20, 2023, were extracted, resulting in a total of 48 billion tokens after tokenization. The construction method was based on the PubMed Central subset from the Pile dataset, with updates to include more recent articles. Note that case reports of rare diseases, emergencies, and complex diseases were removed for evaluation purposes from the pretraining data, based on their unique identifiers (i.e., PMID). A combination of PubMed cases, as well as private EHR cases from China as the validation cohorts were used. The Fourier-transform infrared (FTIR) spectra of the samples were determined using an FTIR spectrometer (IRAFFINITY-IS, China) equipped with an attenuated total reflectance (ATR) accessory. The samples were loaded in the ATR device and measured over a scanning range from 4000 to 400 cm’1at a resolution of 1 cm’1.

[0242] Medical textbooks (public). To incorporate certified medical knowledge into the training process, a total of 15,731 publicly available books related to biomedicine were collected. These books were sourced from various channels, including university-subscribed platforms like ScienceDirect and the Wiley Online Library, as well as websites that provide open access to downloadable medical e-books, such as the Directory of Open Access Books.Keywords derived from the Medical Subject Headings (MeSH) were used to ensure the relevance and comprehensiveness of the search. For books that could not be directly extracted, the Nougat method was employed to extract the relevant content. The text was preprocessed to remove non-medical content, resulting in approximately 8.6 billion tokens.

[0243] Finetuning Data:

[0244] In this phase, the aim was to elicit knowledge from the pretrained model by applying supervised finetuning with a collection of high-quality question-answering (QA) pairs. Specifically, the finetuning data comprise: medical QA data, natural language data, and EHR data, essentially as described in Example 1.

[0245] Evaluation Data:

[0246] To avoid data contamination, studies of rare diseases, emergencies, and complex diseases were removed for evaluation purposes from the pretraining / finetuning data, based on their unique identifiers (i.e., PMID).

[0247] Case reports (public). To enhance the real -world applicability of the evaluation, case reports of rare diseases and emergency conditions from leading medical journals derived from PubMed were incorporated. 86 rare disease cases and 86 emergency cases from the American Journal of Ophthalmology (2010-2023) were curated, which were assessed by EyeGPT, GPT-4, and evaluating ophthalmologists. Additionally, 97 rare disease cases and 109 emergency cases from the American Journal of Emergency Medicine, Journal of American College of Emergency Physicians, North American Journal of Medical Sciences, and American Journal of Medical Genetics (2012-2022) were compiled. These cases were evaluated by EyeGPT, GPT-4, and physicians specializing in emergency medicine and medical genetics. The experts provided insights into the final diagnosis, differential diagnoses, and reasoning behind their conclusions.

[0248] Complex disease cases (public). To incorporate real-world complex disease cases, 80 common diseases and unusual cases published after (and including) 2020 in the NEJM CPCs were downloaded, each with the correct diagnosis, along with the full description of the case and the procedures performed.

[0249] EHR data (private). The EHR data was randomly divided into training and test sets, with the test set comprising 958 cases of rare diseases and 2,769 emergency cases, respectively.

[0250] Image-report data:

[0251] To train the multimodal report generator, four types of retinal images from the private and public datasets were collected, resulting in a collection of over 1 million imagereport pairs spanning 6 imaging modalities: OCT, retinal fundus photographs, FFA, ICGA.

[0252] Ophthalmic images (private). Image-report pairs from the China Consortium of Eye Image Investigation (CC-EII) were incorporated. The demographics and clinical information of the cohort participants are described in previous studies. OCT data was randomly split based on patient identifications to build the training and internal validation sets (the ratio of the training to the validation set is 9: 1) (71,558 data pairs). The external validation set consists of 1,775 data pairs. Different from OCT and fundus images, cases of FFA or ICGA usually comprised more than one image with different locations and time-sequence information. The FFA dataset consists of 4,895 reports associated with 67,402 images, and the ICGA data involves 1,131 reports alongside 20,132 images.

[0253] Framework:

[0254] Although general-purpose LLMs like Qwen-1.5 and GPT-4 have showcased strong performance across various tasks in benchmarks, their utilization in the medical field necessitates adaptation and alignment with domain-specific data due to the inadequacy of domain knowledge. Hence, a two-stage training strategy was implemented comprising pretraining and supervised finetuning to enrich the language model with more medical knowledge and medical abilities to solve the technical problem in the lack of precise and reliable diagnostic support within medicine, and specifically, ophthalmology.

[0255] Pretraining. The intention for building EyeGPT on an existing open-sourceQwenl.5 framework, is for its versatility and nimbleness, as well as its much smaller computing resource requirement. EyeGPT was initialized with the parameters of Qwenl.5-32B and subsequently further pre-trained on medical text corpora, which encompassed medical articles from PubMed, text content from publicly available medical books (including guidebooks and medical genetic textbooks), and clinical text data (mainly EHRs). The corpora were tokenized and concatenated into a sequence of tokens, which were then chunked into fixed-length pieces for training. Given a token sequenceA™ 3j v, the optimization objective is to minimize auto-regressive loss formulated asmodel, x<;denotes tokens before x, and N is the fixed sequence length.

[0256] As discussed above, tokenizing medical text corpora results in large numbers of lengthy tokens which impair the training and inference efficiency of existing LLMs. The embodiment of this Example addresses this technical problem through the steps in which the medical corpora is tokenized and concatenated into a sequence of tokens and minimizing autoregressive loss.

[0257] Supervised finetuning. In the second training phase, the instruction tuning approach was incorporated, defined as finetuning a pretrained large language model using a dataset comprising instructional inputs and their corresponding responses. The instructions used formedical QA and EHR data were essentially as described in Example 1, Table E2. The training loss was the same auto-regressive loss in the pretraining stage and calculated based on the output tokens. For the remaining data, each sample consists of one or multiple rounds of dialogue content, where each round of dialogue includes instructions describing the task, optional instance inputs providing supplementary information required for task resolution, and the expected output corresponding to that round of dialogue. The training loss was computed based on the full dialog tokens.

[0258] Long context window. The default context window size of Qwenl.5 32B is 32,768. However, in downstream medical scenarios, there are many predictive tasks involving long texts, such as differential diagnosis based on extensive patient information. In these situations, the total length of patient information exceeds the default context window size of Qwenl.5. In both training phases, the context window was extended to avoid significant performance degradation when handling inputs that surpass the pretrained context window size during text completion. In some instances, linear rope scaling was employed to extend the context window. Specifically, position indices were directly downscaled in the Rotary Position Embedding (RoPE) to match the previous context window limit and calculated the intermediate values of the position encodings between adjacent integer positions. These techniques to extend the context window solve a technical problem for building and training of EyeGPT, resulting in a model with improved accuracy and performance, particularly for differential diagnosis tasks.

[0259] Medical report generator. To deal with the varying dimensions across different imaging modalities, two vision encoders are used with different architectures - Swin Transformer and ViT - on top of EyeGPT. This technical problem is solved through the use of the two vision encoders of the Example to allow for image data to be used from different modalities for training of EyeGPT. Swin Transformer is an architecture in the field of computer vision that utilizes shifted windows to provide a hierarchical and efficient way of processing images. This design enables the model to focus on both local and global features of an image, making it particularly effective for vision tasks. In practice, the developed medical report generator used a Swin Transformer to process medical images, including OCT and retinal Fundus photographs. However, the Swin Transformer was designed to handle 2D image data and is not capable of tackling 3D volumetric arrays. To address this, ViT-3D, a transformerbased volumetric processing model, was used. The core idea is to first divide large 3D volumes into smaller ones, on top of which visual representation learning was further applied. In the pipeline, ViT-3D was employed to deal with FFA, ICGA, and CT to capture relations and dependencies across images or slices. After transforming raw pixels or voxels into latent embeddings using vision backbones, the obtained visual representations obtained wereforwarded to EyeGPT. However, images or volumes are high-dimensional data and converting them into dense embeddings creates a challenge due to the generation of lengthy token sequences (e.g., 768 tokens were generated by ViT-3D), which impairs the training and inference efficiency of the following EyeGPT. To address this technical problem, dimension reduction was applied to adjacent visual tokens and effectively decreased the visual token count by a factor of 6 at least. In practice, this strategy can dramatically reduce memory costs and speed up model training. After obtaining the reduced visual tokens, linear projectors were used to map visual representations to the language space. In addition to visual tokens, the language input also comprises demographic information and textual tags that indicate the input modalities. All of these were combined and fed to the EyeGPT in each forward pass. Since the parameter finetuning was applied to EyeGPT, this operation resulted in 20M trainable parameters, along with a 2D vision encoder that has 88M parameters and a 3D vision encoder that has 100M parameters. The total number of trainable parameters of the report generator is about 200M. Thus, the embodiment of this Example further addresses a technical problem in the processing of image-based data by reducing the visual token count, reducing memory costs and speeding up model training of EyeGPT.

[0260] Implementations:

[0261] EyeGPT. The scaling factor used in linear rope scaling was 4, resulting in a context window of 16,384 for EyeGPT. During the training phase, a global batch size of 384 was used. Optimization was performed using AdamW with an initial learning rate of le-5. The loss function employed was autoregressive loss. DeepSpeed Zero-3 was utilized to accelerate the training process, based on the bfloatl6 format, while also implementing the gradient checkpointing strategy to increase the batch size. In the foundation phase, the training process was conducted for one epoch using 120 NVIDIA A100 GPUs (80G), while in the instructiontuning stage, the model was trained for 3 epochs using 48 A100 units. For GPT-4, version 2024- 01-25-preview was used for model testing and evaluation.

[0262] EyeGPT for report generation. The base version of Swin Transformer, which has 4 stages, a window size of 7, a patch size of 4, and an initial feature dimension of 128 was used. For data of OCT, retinal fundus, a random portion of each image was first cropped, where the area of the cropped region with respect to that of the original image is between 0.5 and 1.0. Then, each crop was resized to 224 (Height) *224 (Width) pixels with 3 channels. Pretrained weights on ImageNet were used. The ViT-3D model has 12 layers and an embedding dimension of 768. Each input volume was first resized to 224 (Height) x 224 (Width) x 64 (Depth) and further divided into subvolumes with no overlap, where the size of each subvolume is 16 (Height) x 16 (Width) x 8 (Depth). To reduce the count of visual tokens, the adaptive averagepooling strategy was leveraged and the output length was fixed to 9. Two linear projectors for 2D and 3D data, respectively, where each projector includes one fully connected layer, mapping each pooled visual token to a ID vector that has 4096 elements. Before the projector, layer normalization was added to avoid exploding gradients. End-to-end training was applied, where all parameters of vision encoders and linear projectors are trainable. For the LLM part, the LoRA, i.e., low-rank adaptation strategy, was used for parameter-efficient training. The embodiment of this Example solves a technical problem in the number of parameters which need to be updated during training by improving parameter-efficient training of the EyeGPT model. LoRA relies on the mathematical concept of low-rank matrix decomposition. In this approach, a large weight matrix of a neural network layer (which is typically high-rank) is approximated using two smaller matrices. This decomposition reduces the number of trainable parameters. By using low-rank matrices, LoRA modifies only a small fraction of the model’s parameters during the training stage. To achieve this, the rank and alpha value of LoRA was set to 16, respectively. AdamW was used as the default optimizer along with the cosine learning rate scheduler. The initial learning rate was set to le-4, and the minimum learning rate was 3e-6. The total number of training iterations is 80,000. Linear warmup in the first 2,000 training iterations was applied, and the warmup learning rate was le-7. The training stage can be finished within 48 hours using 48 NVIDIA Al 00 GPUs (80G) with a global batch size of 64.

[0263] Human evaluation criteria for clinical text. A scoring system was implemented to evaluate the quality of diagnoses made by EyeGPT, GPT-4, and participating physicians. Three senior-level physicians, each with over 20 years of experience, independently graded the generated diagnoses in a masked manner, where evaluators did not know the source of each diagnosis and compared it to the actual diagnoses from the medical literature. A scoring system (detailed in the Methods) was used to assess the accuracy and relevancy of the generated diagnoses. For each case, the final score was determined based on the consensus between at least two of the graders.

[0264] Human evaluation criteria for imaging reports. To evaluate the quality of the medical reports generated by EyeGPT, a physician grading system was implemented involving six ophthalmologists, tasked with evaluating ophthalmic modalities such as OCT, Fundus, FFA, and ICGA. The grader panel consisted of (i) Senior physicians (one ophthalmologist and one radiologist) each with over 15 years of clinical experience in their respective field, and (ii) Junior physicians (five ophthalmologists and seven radiologists), each with at least 5 years of clinical experience in their respective fields. These physicians were carefully selected to ensure a comprehensive and reliable assessment of the generated reports. The evaluation process was conducted as a masked review, with all identifying details removed from the reports to preventbias. Physicians were then tasked to compare the quality, accuracy, and clinical relevance of reports, without knowing whether they are generated by Al or humans.

[0265] Automated evaluation criteria. To evaluate the diagnostic results on EHRs, the EyeGPT and GPT-4 were asked to generate the ICD-10 codes along with their final diagnoses. Then, ICD-10 codes were matched and predicted with the codes of the actual diagnoses and calculated the accuracy and Fl scores. For complex disease diagnostics, the top- / / accuracy was computed, where a differential diagnosis list was deemed correct if any of the first / / predicted diagnoses matched the actual diagnosis identified by the language model. The proportion of accurate differential diagnosis lists across all cases were determined to compute the top-n accuracy (where n ranged from 1 to 3). For the evaluation of generated medical report, the ROUGE-L and BertScore metrics were employed, which are widely adopted evaluation metrics for measuring the similarity between natural language sentences.

[0266] EyeGPT vs. medical professionals. To assess EyeGPT’s diagnostic performance relative to practicing ophthalmologists and emergency physicians, a comparative study was conducted involving three groups of physicians: 4 junior physicians (>5 years of clinical experience), 4 mid-level physicians with at least (>10 years of experience), and 4 senior physicians (>15 years of clinical experience). The diagnostic evaluation was reviewed by an independent panel of 3 senior physicians with over 20 or more years of clinical experience, who established the final diagnostic score based on consensus. The score used to evaluate diagnostic performance followed a structured approach based on the accuracy of the diagnosis and relevance of differential diagnosis list. A score of +2 is awarded when the final diagnosis made by the EyeGPT, GPT-4, or participating physician matched the actual primary diagnosis or its equivalent. A score of +1 is given if the actual diagnosis appeared in the differential diagnosis list and final diagnosis is relevant, with some correct information present. A score of 0 indicates that the actual diagnosis is not on the predicted differential diagnosis list and the final diagnosis is irrelevant to the case, though no harmful recommendations were made. A score of -1 was applied for a “bad” misdiagnosis that deviated from the actual diagnosis and had the potential to cause some harm if the suggested therapy was followed. Lastly, a score of -2 was reserved for severe misdiagnoses where the recommended therapy is “dangerous” could lead to significant harm to the patient.

[0267] EyeGPT and physician collaboration. To investigate whether EyeGPT could help physicians improve their diagnostic performance, particularly in avoiding a misdiagnosis which may lead to harmful or dangerous outcomes, each participating physician was given the diagnostic output from EyeGPT on each case and was asked to make a diagnosis with theassistance of the Al-generated result. To avoid a potential memorization bias, the follow-up AI- assisted diagnostic test was performed four weeks after the initial test.

[0268] Development of EyeGPT in ophthalmic diseases:

[0269] EyeGPT is a generative foundation model engineered to address four pivotal areas in medicine: multimodal report generation, rare disease diagnosis, emergency condition identification, and complex disease diagnostics (FIG. 7A). This model was pre-trained on an extensive corpus comprising 48 billion articles, 8.6 billion books, and 14.8 billion electronic health records (EHRs) (FIG. 7B). Furthermore, EyeGPT was trained on a comprehensive set of multimodal imaging data, including 315,731 fundus images, 72,668 OCT images, 67,402 FFA images, and 20,132 ICGA images (FIG. 7C). Each of these domains is represented as a distinct capability linked to EyeGPT, underscoring the model's broad applicability in enhancing diagnostic accuracy and efficiency across diverse medical contexts.

[0270] The findings indicate that EyeGPT outperforms GPT-4 in delivering more precise diagnoses and largely minimizes the likelihood of generating harmful advice (FIG. 7D). When assessed using real-world case studies from PubMed, EyeGPT demonstrated enhanced diagnostic capabilities for rare diseases, attaining a score of 1.40 out of a possible 2, significantly surpassing GPT-4 (0.76) by a large margin. These scores can be interpreted as that the diagnosis made by EyeGPT often matches the true diagnosis or includes it in the differential diagnosis list, while diagnosis made by GPT-4 is less aligned with the true diagnosis. In many cases, the true diagnosis appears in the differential diagnosis list, but it's not as consistently accurate as EyeGPT. In identifying emergency conditions, EyeGPT achieved a diagnostic score of 1.50, while the score of GPT-4 was lower at 1.11. Additionally, findings suggest that when EyeGPT partners with physicians across different career levels, it generally boosts their diagnostic abilities, sometimes leading to noticeable enhancements.

[0271] Evaluation framework and criteria for EyeGPT:

[0272] To comprehensively assess the performance of the Al model, an evaluation framework was designed centered on its four core functions. For rare diseases and emergency condition diagnosis, publicly available case studies curated from PubMed and electronic health records (EHRs) collected from hospitals with confirmed diagnosis were used (FIG. 8A and FIG. 8B). EyeGPT was presented with these case studies, and its ability to correctly identify the rare conditions was assessed. For the evaluation EyeGPT’ s performance in diagnosing complex disease, case studies from the New England Journal of Medicine (NEJM) clinicopathological conferences (CPCs) were used, recognized for their detailed exploration of diagnostically challenging conditions (FIG. 8C). The model’s ability to generate medical reports based on medical images was evaluated by inputting ophthalmic images, such as OCT, FFA, and ICGA.The model was tasked with producing comprehensive medical reports for each image, with the reports subsequently reviewed by ophthalmologists (FIG. 8D). To ensure robust assessment, experienced ophthalmologists were invited to evaluate the model’s diagnostic accuracy and report generation capabilities. Their input was complemented by an automated evaluation system designed to reduce clinician workload and provide objective performance metrics (detailed in Methods section). This dual evaluation approach ensured that the model was rigorously tested across both routine and complex ophthalmic tasks, facilitating a comprehensive analysis of its clinical applicability.

[0273] Rare ophthalmic disease diagnosis:

[0274] EyeGPT was evaluated on diverse datasets to help identify subtle, pattern-based correlations that might elude traditional diagnostic approaches.

[0275] EyeGPT diagnostic performance in rare ophthalmic disease. An evaluation ofPubMed case studies was performed. FIG. 9A and FIG. 9B present the average scores and score distributions of EyeGPT, GPT-4, and three groups of ophthalmologists with varying levels of experience (i.e., junior, mid-level, and senior). In an analysis of 86 case studies related to rare ophthalmic diseases, EyeGPT achieved an average score of 1.23, while the mean performance of GPT-4 was 0.58. In contrast, the average scores of junior, mid-level, and senior ophthalmologists are 0.57, 0.63, and 0.69, respectively. In a detailed evaluation (FIG. 9B), over half of the diagnoses made by EyeGPT, amounting to 45 instances, were awarded a score of +2. This rating signifies that these diagnoses not only aligned precisely with the evidence as per the Preferred Practice Pattern guidelines of the American Academy of Ophthalmology but also offered additional benefits to patients. Around 25.6% (22 cases) received a +1 rating, acknowledging the provision of valuable insights that could assist ophthalmologists in diagnosing patients. Approximately 16.3% (14 cases) were given a neutral score of 0, indicating that while they did not offer beneficial advice, they also did not contain any harmful information or recommendations that could adversely affect patient health or well-being. A smaller fraction, 4 cases (4.7%), were rated at -1, labeled as “bad” due to containing potentially harmful information. Lastly, 1 diagnostic outcome was deemed as “dangerous” with a -2 rating, suggesting it posed a risk to patient health or well-being.

[0276] Collaboration between EyeGPT and ophthalmologists. To investigate whether the Al system could help ophthalmologists improve their diagnostic performance, particularly in avoiding misdiagnosis, each ophthalmologist was given the diagnostic output from EyeGPT on each case and was asked to make a diagnosis with the assistance of the Al-generated result. As shown in FIG. 9A, the diagnostic performance of physicians with varying years of experience was mostly improved with the assistance of EyeGPT. For rare ophthalmic diseases, the averagescores of the junior, mid-level, and senior groups increased from 0.57, 0.63, and 0.69, respectively, to 1.02, 1.09, and 1.24. These results demonstrate that EyeGPT can assist ophthalmologists in improving their diagnostic accuracy when faced with rare diseases, regardless of their level of experience.

[0277] Avoiding harmful response. A particularly noteworthy aspect is the noticeable decrease in harmful responses achieved through the collaboration between humans and Al. As depicted in FIG. 9B, the responses provided by physicians may still contain certain biased information which may potentially lead to harmful or adverse outcomes. However, by incorporating the diagnostic output from EyeGPT, the number of responses that garnered negative scores was substantially reduced. Compared to GPT-4, EyeGPT presented a smaller proportion of diagnosis rated lower than 0 and considered to detrimentally impact patient health.

[0278] Diagnostic results on EHRs. The average accuracy and Fl scores of EyeGPT and GPT-4 in rare disease diagnoses (17 major categories, 66 subcategories) were compared (FIG. 9C). EyeGPT achieved a mean accuracy score of 0.698, surpassing GPT-4’ s 0.329 by a large margin. Meanwhile, the average Fl score of EyeGPT is 0.754, which is markedly higher than the mean Fl score of GPT-4 (0.419). The substantial difference in performance between EyeGPT and GPT-4 can be attributed to EyeGPT’ s specialized training in the domain of rare diseases. By focusing on a specific area of medicine, EyeGPT has been able to develop a deeper understanding of the complex relationships between symptoms, patient history, and rare disease diagnoses. A statistical comparison between EyeGPT and GPT-4 is presented in FIG. 9D, where the statistical significance using Mann Whitney Wilcoxon Test was computed. It was found that EyeGPT performs significantly better than GPT-4. Disease classes that are categorized by ICD 10 code can be found in Table E4.Table E4. EHR Disease Classes

[0279] Emergency ophthalmic conditions diagnosis:

[0280] EyeGPT performance was evaluated on analyzing patient data and aiding physicians in making quick and accurate diagnoses.

[0281] EyeGPT diagnostic performance in emergency ophthalmic disease. In a study examining the effectiveness of diagnostic tools in ophthalmic emergencies across 86 case studies, EyeGPT was found to be the most effective, with an average score of 1.40. This was higher than the average score of 1.02 achieved by GPT-4 (as shown in FIG. 10A and FIG. 10B) In contrast, junior, mid-level, and senior ophthalmologists received average scores of 1.02, 1.09, and 1.29, respectively. These results suggest that Al-powered diagnostic tools like EyeGPT can potentially outperform human experts in accurately diagnosing ophthalmic emergencies. Further analysis detailed in FIG. 10B shows that 60.5% of the diagnoses (52 cases) were given the highest score of +2, indicating they were considered “accurate.” About 23% of the cases (20 cases) were scored at +1, indicating they provided valuable insights towards a diagnosis. Around 12.8% (11 cases) were rated as neutral, neither aiding nor hindering the diagnostic process. A small number of diagnoses were rated negatively, with 2.3% (2 cases) receiving a -1 score for being slightly harmful, and 1 case was deemed highly dangerous, receiving a -2 score. The high proportion of accurate and insightful diagnoses highlights the potential of EyeGPT in assisting with ophthalmic emergencies, although the presence of a few harmful diagnoses underscores the need for further refinement and oversight. Note that the assessment was carried out on case studies from PubMed.

[0282] Collaboration between EyeGPT and emergency ophthalmologists. The impact of the diagnostic capabilities of EyeGPT on enhancing the diagnostic precision of ophthalmologistswas also explored. The evaluation was based on the approach used for rare diseases, allowing these healthcare providers to consider the diagnostic suggestions from EyeGPT before finalizing their diagnoses. For cases involving ophthalmic emergencies, it was observed that the diagnostic scores for physicians at various levels of experience - junior, mid-level, and senior - increased by 39%, 77%, and 30%, respectively (FIG. 10A). These substantial improvements indicate that EyeGPT can serve as a valuable tool to augment the diagnostic capabilities of ophthalmologists at all experience levels. The larger improvements seen in junior and mid-level physicians suggest that EyeGPT could be particularly beneficial in assisting less experienced ophthalmologists. EyeGPT also helped decrease the number of potentially harmful diagnostics of ophthalmologists, which is particularly vital when dealing with emergency conditions. This reduction in harmful diagnoses highlights the potential of EyeGPT to improve patient safety and outcomes in emergency settings.

[0283] Diagnostic results on EHRs. The comparison of average accuracy and Fl scores between EyeGPT and GPT-4 for the diagnosis of emergency conditions was explored (FIG. 10C), covering 18 major categories and 125 subcategories. EyeGPT largely outperforms GPT-4, achieving a mean accuracy of 0.702 compared to GPT-4’s 0.463. Additionally, the average Fl score of EyeGPT stands at 0.783, much higher than GPT-4’s average of 0.541. These results demonstrate EyeGPT’ s superior performance in accurately diagnosing emergency conditions across a wide range of categories, highlighting its potential for real-world clinical applications. Further details are expanded upon in FIG. 10D, which presents a statistical analysis comparing EyeGPT to GPT-4. The results indicate that EyeGPT consistently surpasses GPT-4 across different subcategories, evidencing its notable effectiveness in diagnosing emergency conditions. This consistent outperformance suggests that EyeGPT could be a more reliable tool for emergency diagnosis compared to GPT-4. Disease classes that are categorized by ICD 10 code can be found in Table E4.

[0284] Complex disease diagnostics:

[0285] Diagnosing complex diseases presents a formidable challenge to medical Al, which needs to interpret various patterns and distributions of clinical findings and integrate them with specific patient information to arrive at a diagnosis. The complexity of this task lies in the multifaceted nature of the data and the need for contextual understanding, making it a critical test for the capabilities of Al systems in healthcare. EyeGPT was used to formulate differential diagnoses for complex diseases presented by the NEJM CPCs. FIG. 11 presents a comparison of the top-n accuracy of three LLMs: EyeGPT, Google LLM for DDx, and GPT-4. Specifically, the model was asked to make the n most probable diagnostic predictions. If the correct answer appears in the prediction list, it is considered accurate. This approach allows for a morecomprehensive evaluation of the models’ diagnostic capabilities, as it accounts for the inherent uncertainty in complex disease diagnosis.

[0286] EyeGPT consistently outperformed the counterpart methods, maintaining the highest accuracy across all top-n values. Google LLM for DDx is the second most accurate, followed by GPT-4. EyeGPT scored a top-3 accuracy of 46%, surpassing the other two models by large margins. This superior performance highlights the potential of EyeGPT as a powerful tool for assisting physicians in diagnosing complex diseases, which could lead to improved patient outcomes and more efficient healthcare delivery. For all models, there is a clear trend that shows an increase in accuracy as the top-n value rises from 1 to 3, indicating that the models are more accurate when they are allowed to consider a broader set of possible outcomes. This suggests that the precision of these models’ predictions improves when they generate and consider more potential responses. This finding underscores the importance of designing Al systems that can generate and evaluate multiple hypotheses in complex diagnostic scenarios, mimicking the analytical process of experienced clinicians.

[0287] Multimodal medical report generation:

[0288] To explore the potential of handling multimodal data, a system based on EyeGPT was developed to compose medical reports for multimodal medical images. The system specifically processes ophthalmic images, including OCT, retinal fundus photography, FFA, and ICGA.

[0289] The approach was compared with two popular multimodal medical foundation models: LLaVA-Med 21 and Med-Flamingo 22. For ophthalmic report generation, EyeGPT scored the best on all four imaging modalities. In the OCT data (FIG. 12A), EyeGPT achieved ROUGE-L and BertScore scores of 0.739 (95% confidence interval (CI): 0.728, 0.747) and 0.828 (95% CI: 0.808, 0.849) in internal validation. LLaVA-Med scored the second best but was significantly worse than EyeGPT. Similar phenomena can be observed on the external sets of OCT (FIG. 12A) and fundus images (FIG. 12B). For generating reports for FFA images (FIG. 12C), EyeGPT achieved ROUGE-L and BertScore scores of 0.338 (95% CI: 0.331, 0.346) and 0.458 (95% CI: 0.445, 0.468) in the internal set, which are significantly better than Med-Flamingo (ROUGE-L: 0.105, 95% CI: 0.098, 0.120, BertScore: 0.172, 95% CI: 0.161, 0.182). Similarly, EyeGPT also surpassed Med-Flamingo by large margins in the ICGA data (FIG. 12D)

[0290] Next, the results of EyeGPT versus physician-composed medical image reports were compared and evaluated across 4 different imaging modalities. For ophthalmic images (FIGS. 13A-13D), EyeGPT performed better or on par with ophthalmologists in 56.1% of OCT data, 76.1% of fundus data, 56.7% of FFA data, and 62.3% of ICGA data, respectively. Evenignoring the tie results, the average rate of “Al is superior” in OCT is 32.7%, while the rate of “Ophthalmologist’s report is better” is 43.9%. Although ophthalmologists still perform better than Al, the performance gap is not large. Similar situations also existed in other types of ophthalmic images. Compared to senior radiologists (FIGS. 13A-13D), the junior ones were more likely to prefer Al-generated ophthalmic reports (Junior 32.6% vs. Senior 29.4%).

[0291] EyeGPT addresses the unmet clinical need for diagnosis of rare and emergency diseases and shows good potential for solving complex diagnostic challenges. Rare diseases often require comprehensive knowledge from diverse medical sources, a task where foundation models like EyeGPT can assist by synthesizing critical diagnostic insights. In this example, EyeGPT showcased notable superiority over state-of-the-art foundation Al in discerning nuanced patterns and symptoms related to rare diseases, indicating its ability to manage complexities in diagnosis. Moreover, its performance in emergency scenarios, especially in ophthalmology and immediate life-threatening conditions, suggests potential utility in providing swift and accurate assessments.

[0292] EyeGPT demonstrated good performance in broader emergency disease diagnoses and complex cases from NEJM CPCs, enhancing physician diagnostic accuracy across experience levels. While its ability to enhance diagnostic accuracy across physician experience levels is encouraging, these results should be interpreted with caution. Further studies are necessary to validate its robustness in diverse clinical settings and datasets.

[0293] EyeGPT can suppress harmful outputs. EyeGPT's ability to minimize harmful outputs is a key strength of its performance. Compared to GPT-4, EyeGPT produces significantly fewer responses categorized as "bad" or "dangerous," which could otherwise lead to adverse patient outcomes. This reduction is particularly crucial in the medical domain, where incorrect or misleading information can have severe consequences for patient care and wellbeing. The model's success in suppressing harmful outputs stems from its specialized training on extensive medical datasets, including electronic health records, academic articles, and medical textbooks. By focusing on domain-specific knowledge, EyeGPT has developed a nuanced understanding of medical concepts and their relationships, allowing it to generate more accurate and contextually relevant responses, thereby solving a technical problem in applying existing LLMs to medical applications. This targeted training helps EyeGPT effectively differentiate between beneficial and potentially harmful information, thereby minimizing risks to patient care. Moreover, collaboration between EyeGPT and healthcare professionals across different experience levels has led to a further decrease in harmful outputs. This synergy between human expertise and Al-generated insights acts as a powerful safeguard against errors andmisinterpretations. By leveraging the strengths of both human and artificial intelligence, the likelihood of harmful outputs is reduced, ensuring a higher standard of patient care and safety.

[0294] EyeGPT helps improve the diagnostic performance of ophthalmologists. Another promising finding of this study is the observed improvement in ophthalmologists’ diagnostic performance when assisted by EyeGPT. Across various levels of experience, from junior to senior ophthalmologists, collaboration with EyeGPT consistently led to higher diagnostic accuracy for both rare diseases and emergency conditions. This enhancement in diagnostic performance underscores the potential of Al-assisted decision support systems to augment and complement human expertise in clinical settings. The improvement in diagnostic accuracy can be attributed to several factors. First, extensive training on diverse medical datasets allows EyeGPT to draw upon a vast knowledge base, which may include information that individual physicians have not encountered or retained. By providing physicians with additional relevant insights and considerations, EyeGPT can help them make more informed and comprehensive diagnostic decisions. Second, the ability of EyeGPT to rapidly process and analyze large volumes of data enables it to identify patterns and connections that might be overlooked by human observers. This capability can be particularly valuable in complex or atypical cases, where the most relevant information may not be immediately apparent. Interestingly, the study found that the collaboration between EyeGPT and physicians led to a more pronounced improvement in diagnostic performance among junior and mid-level ophthalmologists compared to their senior counterparts. This finding suggests that Al-assisted decision support may be especially beneficial for less experienced physicians, who may not have encountered as many diverse cases in their clinical practice. By leveraging the knowledge and analytical capabilities of EyeGPT, junior and mid-level physicians can potentially bridge the experience gap and make more accurate diagnostic decisions. The findings suggest that integrating EyeGPT into clinical practice could serve as a powerful adjunct to physicians’ expertise, enhancing their ability to accurately diagnose rare diseases and ultimately improving patient outcomes.

[0295] EyeGPT can compose clinically useful reports for multiple imaging modalities. The model's success in outperforming existing counterparts in both internal and external validations demonstrates its robustness and versatility, suggesting that the visual interpretation system designed primarily for ophthalmic images can be extended seamlessly to other modalities. Human evaluation results further support these findings, showing EyeGPT's competitive edge over physicians in a significant number of cases across various imaging modalities. Notably, junior physicians exhibited a greater preference for Al-generated reports, indicating a potential shift in perception and reliance on Al assistance, particularly in fields like radiology. While ongoing performance improvements are necessary, this study emphasizes that EyeGPT is notintended to replace human expertise. Instead, it serves as a valuable augmentation tool, supporting medical professionals in generating accurate and clinically relevant reports across a wide range of imaging modalities.

[0296] This example shows that EyeGPT outperforms state-of-the-art Al models, including GPT-4 and Google LLM for DDx, in diagnosing rare and emergency ophthalmic conditions. This demonstrates the value of domain-specific fine-tuning for specialized medical applications. General models like GPT-4, trained on broad data, lack the nuanced understanding needed for fields like ophthalmology. PALM2, despite strong general medical capabilities, also struggles with specialized imaging and rare disease cases. By focusing on ophthalmic data, including EHRs, imaging, and literature, EyeGPT achieves higher diagnostic accuracy, especially for rare diseases. Unlike other models relying primarily on unstructured data, EyeGPT integrates structured medical records and multimodal imaging, providing a significant advantage. This ability to incorporate diverse data types sets EyeGPT apart from models like BioGPT and MedCAT, which, though successful in broader fields, do not address ophthalmology's complexities. This highlights the importance of tailored training and specialized datasets, positioning EyeGPT as a leading tool for ophthalmic diagnostics.

[0297] EyeGPT could be fine-tuned with domain-specific data to provide diagnostic support in broader clinical contexts. This adaptability is a hallmark of foundation models, as they can be retrained with minimal adjustments for cross-disciplinary applications, underscoring the broader relevance of EyeGPT in the wider healthcare landscape. The scalability of EyeGPT thus positions it not only as an asset within ophthalmology but also as a versatile tool in the medical Al ecosystem. This approach’s ultimate validation requires randomized controlled clinical trials and extensive user feedback from integrating EyeGPT into healthcare delivery, which can make medical recommendations including task-specific interventions based on predicted patient risk levels. This work showcases the potential of LLMs to enhance healthcare quality and affordability, paving the way for future advancements.Example 6 - Generative Foundation Model for Medical Uses

[0298] A Reinforcement Learning-enhanced multi-modal language model for disease diagnostic and treatment recommendation tasks was developed and evaluated. This example was built upon pretrained large language model architecture, optimized through multimodal integration and Reinforcement Learning from Human Feedback (RLHF) to improve performance in specialized medical domains, such as ophthalmology diagnosis.

[0299] In some instances, the training method employs a three-stage optimization strategy: continual pretraining with medical corpus, fine-tuning with specialized medical questionsanswering corpus, and applying Reinforcement Learning from Human Feedback (RLHF) to further enhance model performance.

[0300] Base Architecture and Extended Context Window: The foundation model utilized an extended context window design, applying linear RoPE scaling to expand the default 4,096- token context window to accommodate complex medical texts and extensive patient information required for comprehensive diagnosis.

[0301] Multimodal Medical Data Integration: The model incorporated diverse medical data sources, including academic literature, medical images with associated reports, electronic health records, and medical textbooks. Swin Transformer for processing 2D medical images and ViT-3D were employed for handling volumetric CT data, with dimension reduction techniques and linear projectors mapping visual representations to the language space.

[0302] Reinforcement Learning from Human Feedback (RLHF): Building upon multimodal pretraining, RLHF techniques were implemented to further optimize the model. By collecting evaluations and feedback from medical practitioners (e.g., ophthalmology specialists) on model outputs, a reward model was constructed and Proximal Policy Optimization (PPO) was employed to train the model to generate responses more aligned with clinical practice, with particular attention to diagnostic accuracy and the safety and appropriateness of treatment recommendations.

[0303] Results: RL-enhanced Model outperformed Base Model in both Diagnostic tasks and Treatment Recommendation tasks.

[0304] Two Al language models (the Base Model and the Reinforcement Learning-enhanced Model) were compared on two tasks (diagnosis and treatment recommendation). Two ophthalmic case sets (private cases and published cases) were used. Both models performed diagnostic tasks and treatment recommendation tasks on these cases, and experienced ophthalmologists scored their outputs. The results showed that the RL-enhanced Model significantly outperformed the Base Model in both tasks across both case sets.

[0305] Performance in Diagnostic tasks:

[0306] The RL-enhanced model demonstrated statistically significant improvements over the base model across both private and published ophthalmic case datasets. In private cases, the RL-enhanced model achieved a mean score of 1.67, outperforming the base model (mean = 1.31). Notably, 96% of the RL-enhanced model’s diagnoses scored >=1 (meeting clinically acceptable thresholds), compared to only 78% for the base model. Score distribution analysis revealed that the RL-enhanced model produced completely correct diagnoses (score = 2) in 61% of cases versus 37% for the base model, with incorrect but harmless diagnoses (score = 0 or 0.5) occurring in only 2% versus 18% of cases, respectively. (Table E5 and FIG. 18A). In publishedcases, the RL-enhanced model maintained superior performance (mean = 1.67 vs. 1.56 for base model), with 92% of scores >=1 (vs. 88%). The proportion of fully correct diagnoses (score = 2) was markedly higher for the RL-enhanced model (73% vs. 65% for base model), while misdiagnoses with potential harm (scores < 0) were rarer for the RL-enhanced model (2% vs. 5% for base model). (Table E6 and FIG. 18B).

[0307] Performance in Treatment Recommendation tasks:

[0308] The RL-enhanced model exhibited even greater advantages in generating treatment plans. In private cases, the mean scores for the RL-enhanced model reached 2.68 (vs. 1.74 for base model), with 95.3% of recommendations scoring >=1.5 (clinically adequate) compared to 77% for the base model. Strikingly, 64.4% of the RL-enhanced model’s plans achieved the maximum score (3, indicating fully appropriate measures), versus only 12.6% in the base model. Notably, the RL-enhanced model generated no treatment plans that were non-beneficial or potentially harmful (score <= -1), whereas the base model produced such recommendations in 6.9% of cases. (Table E7 and FIG. 18C). In published cases, The RL-enhanced model’s mean score was 2.60 (vs. 2.14 for base model), with 93.9% of scores >=1.5 (vs. 81.5% for base model). The RL-enhanced model achieved optimal treatment plans (score = 3) in 60.3% of published cases, while the base model achieved 36.2%. Both of the models did not produce non- beneficial or potentially harmful (score <= -1) treatment plans. (Table E8 and FIG. 18D).Table E5. Performance in Diagnostic Accuracy in Private CasesTable E6. Performance in Diagnostic Accuracy in Published CasesBase Model RL-enhanced ModelScore Case Proportio Accumulated Proportio Accumulated number n Proportion n Proportion2 133 65% 65% 73% 73%1.5 21 10% 75% 5% 78%1 26 13% 88% 14% 92%0.5 8 4% 92% 4% 96%0 7 3% 95% 3% 98%-0.5 3 2% 97% 1% 99%-1 3 2% 98%1% 99.5%1% 99% 1 0.50% 100%1% 100% 0% 100%Table E7. Performance in Treatment Recommendation in Private CasesTable E8. Performance in Treatment Recommendation in Published Cases

[0309] Where any or all of the terms “comprise”, “comprises”, “comprised” or “comprising” are used in this specification (including the claims) they are to be interpreted as specifying the presence of the stated features, integers, steps or components, but not precluding the presence of one or more other features, integers, steps or components or group thereof.

[0310] The present disclosure is not intended to be limited in scope to the particular disclosed embodiments, which are provided, for example, to illustrate various aspects of the present disclosure. Various modifications to the methods and systems described will become apparent from the description and teachings herein. Such variations may be practiced without departing from the true scope and spirit of the disclosure and are intended to fall within the scope of the present disclosure.

Claims

CLAIMS1. A method of generating a trained large language model for medical diagnosis, comprising:(a) pretraining a large language model (LLM) having a context window size using textual corpora to generate a pretrained LLM, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, medical books, or any combination thereof, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize auto-regressive loss; and(b) finetuning the pretrained LLM using a domain-specific dataset comprising: a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, a multimodal imaging dataset, or any combination thereof, thereby generating a trained LLM for medical diagnosis, wherein the pretraining in (a) comprises extending the context window size of the LLM to generate a pretrained LLM having an extended context window size, and / or wherein the finetuning in (b) comprises extending the context window size of the pretrained LLM.

2. The method of claim 1, wherein extending the context window size of the LLM and / or extending the context window size of the pretrained LLM comprises linear rope scaling.

3. The method of claim 2, wherein the linear rope scaling comprises downscaling position indices in Rotary Position Embedding (RoPE).

4. A method of generating a trained large language model for medical diagnosis, comprising:(a) pretraining a large language model (LLM) having a context window size using textual corpora to generate a pretrained LLM, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, and medical books, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize autoregressive loss; and(b) finetuning the pretrained LLM using a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, and a multimodal imaging dataset, thereby generating a trained LLM for medical diagnosis.

5. The method of any one of claims 1 to 4, wherein the pretraining in (a) and / or the finetuning in (b) comprises optimization to minimize auto-regressive loss.

6. The method of any one of claims 1 to 5, wherein the pretrained LLM is finetuned using the medical QA dataset, the rare disease EHR dataset, the emergency EHR dataset, and the multimodal imaging dataset.

7. The method of any one of claims 1 to 6, wherein the finetuning in (b) comprises using different image encoders each for processing image data from a different imaging modality.

8. The method of any one of claims 1 to 7, wherein the finetuning in (b) comprises using a first image encoder for processing 2D images and a second image encoder for processing 3D images.

9. The method of any one of claims 1 to 8, wherein the finetuning in (b) comprises using a Swin Transformer and a Vision Transformer.

10. The method of any one of claims 1 to 9, wherein the finetuning in (b) comprises using Low-Rank Adaptation of Large Language Models (LoRA) to reduce the number of trainable parameters.

11. The method of any one of claims 1 to 10, wherein the finetuning in (b) comprises using a reduced number of visual tokens obtained by dimension reduction of adjacent visual tokens of the multimodal imaging dataset.

12. The method of claim 11, comprising using one or more linear projectors to map visual representations of the reduced number of visual tokens to a language space.

13. The method of any one of claims 1 to 12, wherein the multimodal imaging dataset comprises one or more ophthalmic images, one or more radiological images, or any combination thereof.

14. The method of any one of claims 1 to 13, wherein the multimodal imaging dataset comprises one or more images selected from the group comprising: an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, and an indocyanine green angiography (ICGA) image.

15. The method of any one of claims 1 to 14, wherein the multimodal imaging dataset comprises a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

16. A method of generating a medical diagnosis for a patient, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient, and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model generated by:(a) pretraining a large language model (LLM) having a context window size using textual corpora to generate a pretrained LLM, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, and medical books, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize autoregressive loss; and(b) finetuning the pretrained LLM using a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, and a multimodal imaging dataset.

17. A method of generating a medical diagnosis for a patient, the method comprising receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient, and generating the medical diagnosis by inputting the prompt and the set of data in a trained large language model generated by:(a) pretraining a large language model (LLM) having a context window size, comprising: extending the context window size of the LLM, pretraining the LLM using textual corpora to generate a pretrained LLM having an extended context window size, wherein the textual corpora comprise electronic health records (EHRs) of a plurality of patients, medical articles, medical books, or any combination thereof, wherein the textual corpora are tokenized and concatenated into a token sequence having a fixed sequence length, and wherein the LMM is pretrained using the token sequence to minimize auto-regressive loss; and(b) finetuning the pretrained LLM using a domain-specific dataset comprising: a medical question-answering (QA) dataset comprising instructional inputs and their corresponding responses, a rare disease EHR dataset, an emergency EHR dataset, a natural language dataset, a multimodal imaging dataset, or any combination thereof.

18. The method of claim 16 or claim 17, wherein the set of data related to the patient is a set of image data comprising one or more ophthalmic images, one or more radiological images, or any combination thereof.

19. The method of claim 18, wherein the one or more ophthalmic images comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or any combination thereof.

20. The method of claim 18 or claim 19, wherein the one or more radiological images comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.

21. The method of any one of claims 16 to 20, further comprising: generating a medical report for the patient comprising the medical diagnosis generated using the trained LLM.

22. The method of any one of claims 16 to 21, further comprising: comparing the medical diagnosis generated using the trained LLM with a medical diagnosis from a clinician; and determining a medical diagnosis for the patient based on the comparison.

23. The method of any one of claims 16 to 22, wherein the medical diagnosis is for rare disease diagnosis, urgent care diagnosis, and / or complex disease diagnosis.

24. The method of any one of claims 16 to 23, wherein the medical diagnosis is for diagnosis of an ophthalmic disease or condition.

25. The method of claim 24, wherein the ophthalmic disease or condition comprises a retinal disorder, a visual pathway disorder, keratitis, a corneal scar and / or opacity condition, iridocyclitis, age-related cataract, cataract, a choroid disorder, a retinal detachment condition, a retinal vascular occlusion, a retinal disorder, glaucoma, a vitreous body disorder, and / or a globe disorder.

26. The method of any one of claims 1 to 25, further comprising:(c) collecting evaluations and feedback from one or more medical practitioners on outputs from the pretrained and finetuned LLM.

27. The method of claim 26, further comprising:(d) using a reward model and Proximal Policy Optimization (PPO) to train the pretrained and finetuned LLM to generate responses aligned with clinical practice.

28. A system comprising: at least one hardware processor; and one or more software modules configured to, when executed by the at least one hardware processor, perform the method of any one of claims 1 to 27.

29. A non-transitory computer-readable medium having instructions stored thereon, wherein the instructions, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 27.

30. A system comprising: at least one hardware processor; non-transitory computer-readable medium coupled to at least one hardware processor, optionally wherein the coupling is over a network; and instructions stored in the non-transitory computer-readable medium, wherein the instructions when implemented by the processor, configure the system to perform the method of any one of claims 1 to 27.

Citation Information

Patent Citations

  • Using deep learning techniques to determine the contextual reading order in a form document

    US20190188463A1

  • Hierarchical CNN-transformer based machine learning

    US20210183484A1

  • Transforming a lexicon that describes an information asset

    US20220300709A1

  • Method and system for automated generation of text captions from medical images

    US20230274420A1

Cited By

  • Rare disease risk screening model training system and method, screening system and medium

    CN121075621A

  • Multi-modal large model training method and system

    CN121170273A

  • A multi-modal large model training method and system

    CN121170273B

  • Skin disease auxiliary diagnosis method and system based on multi-mode optical image fusion

    CN121687461A