Generative foundation model for medical use
MetaGP, trained on extensive medical data, addresses the limitations of existing AI models by providing accurate and comprehensive medical diagnostics, enhancing physician performance and generating reliable medical reports across diverse imaging modalities.
Patent Information
- Application Number
- PCT/US2024/026697
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-27
- Publication Date
- 2025-10-30
Smart Images

Figure US2024026697_30102025_PF_FP_ABST
Abstract
Description
[0001] Generative foundation model for medical use Provided in some embodiments herein is a generative foundation model, Meta General Practitioner (MetaGP), which supports a wide spectrum of complex clinical decision-making scenarios. Unlike previous endeavors, MetaGP was trained on over ten million health system-scale electronic health records along with web-scale medical text corpora to acquire knowledge of both medical practices and theories. We validated the model on three complex, yet unmet clinical challenges that demand strong reasoning ability: rare disease diagnosis, emergency condition identification, and complex disease solving. Through a rigorous and extensive evaluation by medical professionals on real-world cases, MetaGP not only surpasses GPT-4 in providing more accurate diagnostics but also significantly reduces potentially harmful outputs. Evaluated on real-world case studies from PubMed, MetaGP showed a markedly improved ability to diagnose rare diseases, achieving a diagnostic score of 1.40 out of a maximum of 2, outperforming GPT-4 (0.76) by a large margin. When identifying emergency conditions, MetaGP achieved a score of 1.50, whereas GPT-4 scored 1.11. Furthermore, our research indicates that collaboration between physicians at various stages of their careers and MetaGP typically enhances the diagnostic skills of physicians, sometimes resulting in large improvements. When tackling complex diseases presented by the New England Journal of Medicine, MetaGP achieved a top- 3 accuracy rate of 46%, surpassing the performance of the Google LLM for DDx and GPT-4. We also explored the potential of MetaGP for processing multimodal data by building a versatile medical report generator for multimodal imaging. Trained on over one million image-report pairs, MetaGP excelled beyond top multimodal medical models and junior physicians in six imaging modalities, matching the performance of senior physicians. We believe the targeted training of MetaGP specifically on medical datasets uniquely equips it as a specialized, open-source solution for healthcare uses, setting it apart from the broad functionalities of GPT-4. sf-5927165 Introduction The emergence of Artificial Intelligence (AI) has ushered in a new era in healthcare and medicine. Recent breakthroughs have seen AI tools successfully interpret various types of medical data, such as dermoscopic images1,2, retina images3,4, electronic health records (EHRs)5-7, electrocardiograms8, and oncology trials9. While these models are adept at their specialized tasks, they often struggle with diagnostic duties that span multiple disciplines. For instance, a cardiology-focused AI model might overlook neurological symptoms in neurology. Such “tunnel vision” can potentially lead to missed diagnoses or an incomplete understanding of a patient's holistic health needs. Without a broad perspective or knowledge base, these tools may compromise comprehensive patient care. Moreover, the development of these AI models requires the integration of a large amount of structured data, while the process of structuring medical data usually relies on extensive expertise and customized data processing procedures. For example, before applying AI to EHRs, one would normally need to convert the heterogeneous raw data into well-structured inputs5,7, which is not only labor-intensive but also prone to information loss. Moreover, as more data are needed, this schema could restrict the scalability of building more advanced AI systems. Addressing these challenges requires a foundation AI model10that combines specialized insights with a holistic overview and requires minimal artificially structured data for training. The recent advent of generative large language models (LLMs)11-14provides a great opportunity to establish foundation AI, as they are trained on vast arrays of free-text web-scale data, possessing a breadth of knowledge spanning across diverse topics. Nevertheless, the training input of these AI models predominantly emanates from general internet-based knowledge, bereft of specialized medical data (EHRs, etc). This poses a significant challenge when applying the models to nuanced medical applications. The dangers are two-fold: (i) the potential for misdiagnosis due to the lack of knowledge of medical practices and (ii) the susceptibility to over-reliance on unfiltered public data which might be misleading or factually incorrect. For AI to maximize its impact in medical diagnostics, it requires a foundation deeply rooted in vast and authentic medical data and expertise. We introduce Meta General Practitioner (MetaGP), a generative foundation model for medicine, developed through training on a vast array of electronic health records across various health systems and an extensive collection of medical texts from the internet (Fig.1a-1c). This ensures that MetaGP possesses a broad and thorough understanding of medical theories and practices. MetaGP has the potential to offer accurate decision support across a wide spectrum of diagnostic scenarios, tackling diverse challenges within the medical field. As a proof of concept, we studied the capabilities of MetaGP in tackling three complex, yet unmet clinical challenges: (i) rare disease diagnosis, (ii) emergency condition identification, and (iii) complex disease diagnostics. To tackle the challenge of assessing the accuracy of generative medical AI predictions, we implemented stringent evaluation protocols and conducted comprehensive testing with the of healthcare experts. Our findings indicate that sf-5927165 MetaGP outperforms GPT-4 in delivering more precise diagnoses and largely minimizes the likelihood of generating harmful advice. When assessed using real-world case studies from PubMed (Fig.1d), MetaGP demonstrated enhanced diagnostic capabilities for rare diseases, attaining a score of 1.40 out of a possible 2, surpassing GPT-4 (0.76) by a large margin. In identifying emergency conditions, MetaGP achieved a diagnostic score of 1.50, while the score of GPT-4 was lower at 1.11. Additionally, our findings suggest that when MetaGP partners with physicians across different career levels, it generally boosts their diagnostic abilities, sometimes leading to noticeable enhancements. When it came to crafting differential diagnoses for challenging cases in the New England Journal of Medicine (NEJM) clinicopathological conferences (CPCs), MetaGP scored a top-3 accuracy of 46%. This performance is on par with Google LLM for DDx15and largely surpassed the capabilities of GPT-4. We also investigated the capabilities of MetaGP in handling multimodal data, demonstrating its potential through the creation of a versatile medical report generator designed for multimodal images. MetaGP showcased proficiency in generating consistent and accurate diagnostic reports, surpassing other multimodal report generators by large margins. Compared to humans, MetaGP outperforms junior physicians in all 6 imaging modalities, while performing comparably with the senior ones. Evaluation framework Fig.2 presents the evaluation framework for MetaGP. We have adopted a human- machine hybrid evaluation strategy with the aim of ensuring quality while controlling costs. For the diagnosis of rare diseases and emergency conditions, we validated the diagnostic output of MetaGP on publicly available case studies from PubMed and our collected EHRs from hospitals, respectively. To facilitate specialized understanding and promote targeted investigation, we further divided the case studies into subgroups to evaluate the MetaGP’s ability to be adapted as a specialist and act as a generalist – ophthalmic and systemic diseases – and sought the expertise of corresponding professionals to complete relevant assessments (Fig.2a and 2b). Specifically, senior ophthalmologists were asked to evaluate studies on rare diseases and emergency conditions in ophthalmology. Senior general practitioners and emergency medicine physicians reviewed the clinical cases featuring systemic characteristics of rare diseases and emergency conditions, respectively. In addition, automated evaluation was implemented to reduce the workload associated with human evaluation efforts across all four diagnostic scenarios. (Fig.2a-2d). Imaging studies from hospitals were categorized into two subgroups: ophthalmology and radiology. In this arrangement, ophthalmologists reviewed the AI-generated imaging reports associated with ophthalmic images, whereas radiologists were responsible for evaluating the generated reports related to chest X-rays (CXRs) and computed tomography (CT) scans. (Fig.2d). Human evaluation criteria for clinical text studies. We introduce a quality scoring scheme for diagnoses made by Three senior-level physicians (> 20 years) sf-5927165 independently graded the performance of MetaGP, GPT-4, and participating physicians blindly against the ground truth. For each sample, we took the major vote over the three physicians as the final evaluation score. Specifically, a score of +2 is awarded when the final diagnosis made by MetaGP matches the actual primary diagnosis or its equivalent. A score of +1 is given if the actual diagnosis is included in the model’s predicted differential diagnosis list, the final diagnosis is relevant, and there is at least some correct information present. A score of 0 indicates that the actual diagnosis is not on the differential diagnosis list and the final diagnosis is irrelevant to the case, but there is no incorrect information or recommendations that could harm the patient's health or well-being. A score of -1 signifies a misdiagnosis that deviates from the actual diagnosis and could cause some harm to the patient's health or well- being if the suggested therapy is pursued. A score of -2 is given when the final diagnosis of MetaGP is a severe misdiagnosis, and the therapy based on this diagnosis could cause severe harm or be significantly dangerous. Note that we only evaluate the primary diagnosis of each case study or a health record. Human evaluation criteria for imaging reports. For each imaging modality, we assigned three specialists: either ophthalmologists or radiologists, depending on the specific requirements of interpreting the modality. The specialists include 1 senior and 2 junior physicians. The evaluators performed a blinded review, examining both the image and the accompanying report, from which all identifying details had been removed, to evaluate the quality of the reports. They were tasked with determining which report was superior or if both reports were of equal quality. Automated evaluation criteria. To evaluate the diagnostic results on EHRs, we asked the MetaGP and GPT-4 to generate the ICD-10 codes along with their final diagnoses. Then, we matched the predicted ICD-10 codes with the codes of the actual diagnoses and calculated the accuracy and F1 scores. For complex disease diagnostics, we computed the top-n accuracy, where a differential diagnosis list was deemed correct if any of the first n predicted diagnoses matched the actual diagnosis identified by the language model. We determined the proportion of accurate differential diagnosis lists across all cases to compute the top-n accuracy (where n ranged from 1 to 3). For medical report evaluation, we employed the ROUGE-L and BertScore metrics, which are widely adopted evaluation metrics for measuring the similarity between natural language sentences. Results Rare disease diagnosis In the landscape of medical diagnostics, the identification and understanding of rare diseases present a significant challenge due to their low incidence and often ambiguous symptomatology. Diagnosing rare diseases with AI requires highly specialized knowledge and an ability to discern nuanced patterns. Within this context, MetaGP can leverage its extensive training on diverse datasets to help identify subtle, pattern-based correlations that elude traditional diagnostic approaches. sf-5927165 Rare ophthalmic diseases. We performed an evaluation of PubMed case studies. Fig. 3a and 3b present the average scores and score distributions of MetaGP, GPT-4, and three groups of ophthalmologists with varying levels of experience (i.e., junior, mid- level, and senior). In an analysis of 86 case studies related to rare ophthalmic diseases, MetaGP achieved an average score of 1.23, while the mean performance of GPT-4 was 0.58. In contrast, the average scores of junior, mid-level, and senior ophthalmologists are 0.57, 0.63, and 0.69, respectively. In a detailed evaluation (Fig. 3b), over half of the diagnoses made by MetaGP, amounting to 45 instances, were awarded a score of +2. This rating signifies that these diagnoses not only aligned precisely with the evidence as per the Preferred Practice Pattern guidelines of the American Academy of Ophthalmology but also offered additional benefits to patients. Around 25.6% (22 cases) received a +1 rating, acknowledging the provision of valuable insights that could assist ophthalmologists in diagnosing patients. Approximately 16.3% (14 cases) were given a neutral score of 0, indicating that while they did not offer beneficial advice, they also did not contain any harmful information or recommendations that could adversely affect patient health or well-being. A smaller fraction, 4 cases (4.7%), were rated at -1, labeled as “bad” due to containing potentially harmful information. Lastly, 1 diagnostic outcome was deemed as “dangerous” with a -2 rating, suggesting it posed a risk to patient health or well-being. Qualitative examples can be found in Supplementary Table 1. Rare systemic diseases. An evaluation was conducted on PubMed case studies. In a comparative analysis involving 97 case studies focused on rare systemic diseases, MetaGP, GPT-4, and general practitioners at various stages of their careers were also evaluated (Fig.3c and 3d). MetaGP led with an average score of 1.57, while GPT-4 followed with a mean score of 0.93. The performance of general practitioners showed variation across experience levels; junior members recorded a score of 0.94, mid-level members achieved a score of 1.02, outperforming GPT-4, and senior members had a score of 1.5. MetaGP exceeded the performance of GPT-4, junior, and mid-level general practitioners. It had 76.3% of its diagnoses (74 cases) receiving the highest score of +2. In addition, 8 cases (8.25%) were rated at +1 for providing helpful diagnostic insights.12 cases (12.4%) received a score of 0, indicating the information was neutral and did not negatively impact patient care. Regarding the samples with negative scores, 2 cases were marked as "bad" (-1 score), and 1 case was categorized as dangerous (-2 score). Qualitative examples can be found in Supplementary Table 2. Collaboration between MetaGP and ophthalmologists / general practitioners. To investigate whether our AI system could help ophthalmologists improve their diagnostic performance, particularly in avoiding a misdiagnosis which may lead to harmful or dangerous outcomes, each ophthalmologist / general practitioner was given the diagnostic output from MetaGP on each case and was asked to make a diagnosis with the assistance of the AI-generated result. To avoid a potential memorization bias, the follow-up AI-assisted diagnostic test was performed four weeks after the initial test. As shown in Fig.3, the diagnostic performance of physicians with varying years of experience was mostly improved with the assistance of MetaGP. For the diagnosis of rare ophthalmic diseases, the scores of the junior, mid-level, and senior sf-5927165 groups increase from 0.57, 0.63, and 0.69, respectively, to 1.02, 1.09, and 1.24. As for the rare systemic diseases, the average scores for the junior, mid-level, and senior groups rose from 0.94, 1.02, and 1.5 to 1.52, 1.62, and 1.88, respectively. A particularly noteworthy aspect is the noticeable decrease in harmful responses achieved through the collaboration between humans and AI. As depicted in Fig.3b and Fig.3d, the responses provided by physicians may still contain certain biased information which may potentially lead to harmful or adverse outcomes. However, by incorporating the diagnostic output from MetaGP, the number of responses that garnered negative scores was substantially reduced. Diagnostic results on EHRs. Fig.4 compares the average accuracy and F1 scores of MetaGP and GPT-4 in rare disease diagnoses (17 major categories, 66 subcategories). MetaGP achieved a mean accuracy score of 0.698, surpassing GPT-4’s 0.329 by a large margin. Meanwhile, the average F1 score of MetaGP is 0.754, which is markedly higher than the mean F1 score of GPT-4 (0.419). We give more details in Fig.6. A statistical comparison between MetaGP and GPT-4 is presented in Fig.6a, where we computed the statistical significance using Mann Whitney Wilcoxon Test. We found that MetaGP performs significantly better than GPT-4. Per-category results and comparisons between MetaGP and GPT-4 are provided in Fig.6. Our observation is that MetaGP led GPT-4 in most subcategories. Qualitative examples can be found in Supplementary Table 3. Emergency conditions diagnoses Emergency conditions, especially those concerning blinding or immediate life- threatening conditions, demand rapid and precise assessments. However, emergency physicians typically need to make decisions within quite limited time. MetaGP can swiftly analyze patient data, aiding physicians in making quick and accurate diagnostics. Emergency conditions related to ophthalmology. In a study examining the effectiveness of diagnostic tools in ophthalmic emergencies across 86 case studies, MetaGP was found to be the most effective, with an average score of 1.40. This was higher than the average score of 1.02 achieved by GPT-4 (as shown in Fig.5a and 5b). In contrast, junior, mid-level, and senior ophthalmologists received average scores of 1.02, 1.09, and 1.29, respectively. Further analysis detailed in Fig.5b shows that 60.5% of the diagnoses (52 cases) were given the highest score of +2, indicating they were considered “accurate.” About 23% of the cases (20 cases) were scored at +1, indicating they provided valuable insights towards a diagnosis. Around 12.8% (11 cases) were rated as neutral, neither aiding nor hindering the diagnostic process. A small number of diagnoses were rated negatively, with 2.3% (2 cases) receiving a -1 score for being slightly harmful, and 1 case was deemed highly dangerous, receiving a -2 score. Note that the assessment was carried out on case studies from PubMed. Qualitative examples can be found in Supplementary Table 1. Systemic emergency conditions. In a detailed analysis of 109 case studies on systemic emergency conditions, MetaGP again stood out by achieving an average score of 1.59, surpassing GPT-4, scored 1.19 (Fig.5c and 5d). The sf-5927165 performance among emergency physicians varied by their level of experience. Junior physicians scored 0.97, while the mid-levels matched the performance of GPT-4 at 1.15. The senior physicians outperformed GPT-4 by a large margin, achieving a mean score of 1.61. A large portion of the diagnoses, 74.3% or 81 cases, received the highest evaluation of +2. Furthermore, 15 cases, making up 13.76%, were acknowledged for providing valuable diagnostic insights with a score of +1.9 cases, or 8.26%, were assessed neutrally with a score of 0, indicating neither significant benefit nor detriment. A small number, 4 cases, were judged negatively with a score of -1, categorized as “bad,” yet notably, no diagnoses were deemed “dangerous” with a -2 score, highlighting the overall positive impact and reliability of the evaluations conducted. The above evaluation was based on case reports from PubMed. Qualitative examples can be found in Supplementary Table 2. Collaboration between MetaGP and general practitioners / emergency physicians. We also explored the impact of the diagnostic capabilities of MetaGP on enhancing the diagnostic precision of ophthalmologists and emergency medicine physicians. We structured our evaluation based on the approach used for rare diseases, allowing these healthcare providers to consider the diagnostic suggestions from MetaGP before finalizing their diagnoses. For cases involving ophthalmic emergencies, we observed that the diagnostic scores for physicians at various levels of experience – junior, mid- level, and senior – increased by 39%, 77%, and 30%, respectively (Fig.5a). When assessing systemic emergencies, the improvement in diagnostic scores for junior, mid- level, and senior emergency medicine physicians were recorded at 53%, 46%, and 19%, respectively (Fig.5c). MetaGP also helped decrease the number of potentially harmful diagnostics of general practitioners and emergency physicians, which is particularly vital when dealing with emergency conditions. Diagnostic results on EHRs. Fig.4 showcases the comparison of average accuracy and F1 scores between MetaGP and GPT-4 for the diagnosis of emergency conditions, covering 18 major categories and 125 subcategories. MetaGP largely outperforms GPT-4, achieving a mean accuracy of 0.702 compared to GPT-4's 0.463. Additionally, the average F1 score of MetaGP stands at 0.783, much higher than GPT-4's average of 0.541. Further details are expanded upon in Fig.6. In Fig.6b, we present a statistical analysis comparing MetaGP to GPT-4. The results indicate that MetaGP consistently surpasses GPT-4 across different subcategories, evidencing its notable effectiveness in diagnosing emergency conditions. Per-class results and comparisons are displayed in Fig.6d. Qualitative examples can be found in Supplementary Table 3. Complex disease diagnostics Diagnosing complex diseases presents a formidable challenge to medical AI, which needs to interpret various patterns and distributions of clinical findings and integrate them with specific patient information to arrive at a diagnosis. We used MetaGP to formulate differential diagnoses for complex diseases presented by the NEJM CPCs. Fig.7 presents a comparison of the top-n accuracy of three LLMs: MetaGP, Google LLM for DDx, and GPT-4. the model was asked to make the n most sf-5927165 probable diagnostic predictions. If the correct answer appears in the prediction list, it is considered accurate. MetaGP consistently outperformed the counterpart methods, maintaining the highest accuracy across all top-n values. Google LLM for DDx is the second most accurate, followed by GPT-4. MetaGP scored a top-3 accuracy of 46%, surpassing the other two models by large margins. For all models, there is a clear trend that shows an increase in accuracy as the top-n value rises from 1 to 3, indicating that the models are more accurate when they are allowed to consider a broader set of possible outcomes. This suggests that the precision of these models' predictions improves when they generate and consider more potential responses. Multimodal medical report generation To explore the potential of handling multimodal data, we developed a system based on MetaGP to compose medical reports for multimodal medical images. The system specifically processes ophthalmic images, including optical coherence tomography (OCT), retinal fundus photographs (Fundus), fundus fluorescein angiography (FFA), and indocyanine green angiography (ICGA). Additionally, it also handles radiological images from two commonly used imaging tests: chest X-ray (CXR) and computed tomography (CT). Examples of generated reports are displayed in Figure 10. We compared our approach with two popular multimodal medical foundation models: LLaVA-Med16and Med-Flamingo17. For ophthalmic report generation, MetaGP scored the best on all four imaging modalities. In the OCT data (Fig.8a), MetaGP achieved ROUGE-L and BertScore scores of 0.739 (95% confidence interval (CI): 0.728, 0.747) and 0.828 (95% CI: 0.808, 0.849) in internal validation. LLaVA- Med scored the second best but was significantly worse than MetaGP. Similar phenomena can be observed on the external sets of OCT (Fig.8a) and fundus images (Fig.8b). For generating reports for FFA images (Fig.8c), MetaGP achieved ROUGE-L and BertScore scores of 0.338 (95% CI: 0.331, 0.346) and 0.458 (95% CI: 0.445, 0.468) in the internal set, which are significantly better than Med-Flamingo (ROUGE-L: 0.105, 95% CI: 0.098, 0.120, BertScore: 0.172, 95% CI: 0.161, 0.182). Similarly, MetaGP also surpassed Med-Flamingo by large margins in the ICGA data (Fig.8d). We further tested MetaGP on two general-purpose medical image report generation (CXR and CT). For CXR interpretation (Fig.8e), the advantages of MetaGP became smaller but still significant. In the CT data (Fig.8f), MetaGP performed significantly better than the second-best LLaVA-Med in both internal and prospective validation. We next compared and evaluated the results of MetaGP versus physician-composed medical image reports across 6 different imaging modalities. For ophthalmic images (Fig.9a-9d), MetaGP performed better or on par with ophthalmologists in 64.5% of OCT data, 73.3% of fundus data, 56.7% of FFA data, and 61.3% of ICGA data, respectively. Even ignoring the tie results, the average rate of “AI is superior” in OCT is 34.5%, while the rate of “Ophthalmologist’s report is better” is 35.6%. Although ophthalmologists still perform better than AI, the performance gap is not large. Similar situations also existed in types of ophthalmic images. For CXR (Fig. sf-5927165 9e), MetaGP won or tied at 57.8% of cases on average, while the proportion of ties became higher on CT data (Fig.9f). Compared to senior radiologists (Fig.9a-9d), the junior ones were more likely to prefer AI-generated ophthalmic reports (Junior 34.7% vs. Senior 29.4%). This contrast was more amplified in CXR and CT data (Junior 29.2% vs. Senior 15.9%) (Fig.9e and 9f). Discussion MetaGP meets unmet clinical needs. MetaGP addresses the unmet clinical need for diagnosis of rare and emergency diseases while showing good potential for solving complex diagnostic challenges. The rarity and complexity of rare diseases necessitate a comprehensive understanding of disparate medical literature and case studies, a task that foundation models can navigate by synthesizing information to assist in uncovering critical diagnostic insights. In this study, MetaGP showcased notable superiority over state-of-the-art foundation AI in discerning nuanced patterns and correlating symptoms with corresponding rare diseases. Its ability to navigate these complexities not only highlights its efficacy but also signifies a promising step toward addressing the diagnostic hurdles posed by such conditions. Moreover, in emergency scenarios, especially those involving ocular health or immediate life-threatening conditions, MetaGP's precision in identifying critical conditions underscores its value in swift and accurate assessments, offering potential aid in critical healthcare scenarios. Beyond specific conditions, its superior performance in broader emergency disease diagnoses signifies its capacity to assist healthcare professionals in diverse medical domains. Furthermore, MetaGP showed a good performance in diagnosing complex disease cases presented in the NEJM CPCs. It is also encouraging to see a cross-board physician performance enhancement by MetaGP. However, further studies are necessary to validate its performance across varied clinical settings and datasets to ensure its robustness and reliability in real-world applications. Nevertheless, these findings offer a promising glimpse into the potential of LLM- driven diagnostics, suggesting that MetaGP could significantly impact the results of diagnosing rare and emergency conditions, thereby improving patient outcomes and healthcare efficacy. MetaGP can compose clinically useful reports for multiple imaging modalities. These capabilities underscore the impact of this approach, demonstrating MetaGP's superiority over established models across diverse medical imaging types. The model's ability to outperform existing counterparts in both internal and external validations indicates its robustness and versatility, suggesting that the visual interpretation system designed primarily for ophthalmic images can seamlessly extend to other modalities. Human evaluation results further support these findings, revealing MetaGP's competitive edge over physicians in a considerable proportion of cases across various imaging modalities. The nuanced preference patterns observed among junior physicians, favoring AI-generated reports to a greater extent, indicate a potential shift in perception and reliance on AI assistance, particularly in fields like radiology. While acknowledging need for consistent performance improvements, sf-5927165 it is important to note that our study does not propose a replacement for human expertise; rather, MetaGP presents itself as a valuable augmentation tool, providing support and aiding medical professionals in generating accurate and clinically relevant reports across a spectrum of medical imaging modalities. Limitations. The scaling law for neural language models18suggests that increasing the size of a model, specifically the number of parameters, often leads to improved performance19. Larger models tend to exhibit enhanced capabilities in learning complex patterns and representations, allowing them to achieve better accuracy and generalization. However, due to considerations of computational efficiency, the model size of MetaGP is much smaller compared to industry-standard large language models, such as the well-known GPT series11,20. Though it is convenient for deployment, the restricted model size might limit MetaGP's ability to encode and retain extensive medical knowledge, which larger models with significantly more parameters excel at. MetaGP's decision-making process might lack transparency, hindering its adoption in critical healthcare settings where understanding the rationale behind diagnostic outputs is crucial for healthcare professionals. Acknowledging potential ethical considerations, we stress the importance of optimized human-AI interactions and standardized assessments to mitigate over-reliance and biases leading to unfavorable outcomes. MetaGP can be further trained to improve its accuracy, factuality, consistency, safety, and mitigate bias21-23, in order for implementation and practical applications in real-world clinical scenarios. Our team is taking various initiatives to address these issues, including ongoing work that assesses the differences in sensitivity to token-level perturbations to clinical notes between language models and physicians. Prospects. As MetaGP presents examples and opportunities for large, generative LLMs’ integration into medical practice workflows, how to efficiently input various user demands and clinical task solutions, and seamless integration into existing medical informatics workflows presents a major challenge. With continuous supervision of human-machine interactions, we hope that MetaGP can make its key impact on physician behaviors and patient care outcomes while having a mitigating plan for safeguarding ethics and patient well-being. Our intention for building MetaGP on an existing open-source Llama2 framework using a small number of parameters (13B), is for its versatility and nimbleness, as well as its much smaller computing resource requirement. The pretraining of the foundation phase utilized 24 NVIDIA A100 GPUs, each equipped with 80GB of VRAM for 4 weeks, followed by subsequent finetuning procedures that employed 16 A100 GPUs for 5 days per run. The computational resources involved in our study, although currently beyond the capabilities of many research groups, given the rapid growth in computing power24,25, we believe that within one to two years, most research groups will be able to afford such computational expenses. As we have planned to make our model open-source and hope it can be further refined in many downstream tasks within weeks, we hope the medical community can pursue finetuning with specific medical data to achieve more accurate predictions in their target domains, so to tap into the full potential of multimodal AI. Our approach's ultimate validation requires randomized controlled sf-5927165 clinical trials and extensive user feedback from integrating MetaGP into healthcare delivery, which can make medical recommendations include task-specific interventions based on predicted patient risk levels. Our work showcases the potential of LLMs to enhance healthcare quality and affordability, paving the way for future advancements. Method details Model development overview The training regime of MetaGP employed a two-phased approach. In the initial phase, we pretrained the MetaGP on an expansive corpus that included over 10 million EHRs, 5 million academic papers from PubMed, and 15 thousand medical textbooks. This phase ensured the model had an excellent grounding in a wealth of knowledge of medical theories. In the second phase, we gathered and employed extensive question- answering (QA) datasets in conjunction with EHR data for cases involving rare diseases and emergency conditions in the real world. The health records include detailed case descriptions and decision-making processes. By finetuning the model with these data, we embedded more applied clinical insights, enhancing its ability to make differential diagnoses. For the multimodal medical report generation system, we curated a vast dataset comprising over 1 million diagnostic reports and associated medical images from 6 imaging modalities, spanning disciplines such as ophthalmology and radiology. AI vs. medical professionals. We compared MetaGP with practicing ophthalmologists, general practitioners, and emergency physicians in assessing clinical cases of rare diseases and emergency conditions. We employed three groups of physicians to participate in the study: 4 in the junior group with at least 5 years of clinical experience, 4 in the mid-level group with at least 10 years of experience, and 4 in the senior group with at least 15 years of clinical experience. The final evaluation score was established based on a consensus from an independent group of 3 senior physicians with 20 or more years of clinical experience. MetaGP and physician collaboration. To investigate whether our MetaGP could help physicians improve their diagnostic performance, particularly in avoiding a misdiagnosis which may lead to harmful or dangerous outcomes, each participating physician was given the diagnostic output from MetaGP on each case and was asked to make a diagnosis with the assistance of the AI-generated result. To avoid a potential memorization bias, the follow-up AI-assisted diagnostic test was performed four weeks after the initial test. Datasets Language data for pretraining. The pretraining dataset mainly comprises publicly available medical data of three sources: medical articles from PubMed, text content from publicly available medical books (including guidebooks and medical genetic textbooks), and clinical text data (mainly EHRs). For medical articles, we referred to sf-5927165 the construction method developed by the PubMed Central subset from the Pile dataset26and updated the article list to include more recent PubMed articles. PubMed Central (PMC) is a subset of the PubMed online repository for biomedical articles managed by the United States of America’s National Center for Biotechnology Information (NCBI), providing open, full-text access to biomedical and life sciences literature. We used PMC articles up to June 20, 2023, to extract 5.4 million articles, with a total token count of 48 billion after tokenization. Regarding certified medical knowledge, we collected a total of 15,731 publicly available books related to biomedicine. For preprocessing, we first extracted all text content from PDFs and then filtered the text to remove non-medical content such as references, author lists, etc. After processing, this part contains approximately 8.6 billion tokens. Since the subsequent tasks include differential diagnosis, we collected 10,514,122 EHRs derived from the China Medical AI Investigation Consortium which is composed of five tertiary hospitals in China (West China Hospital, China Southern Medical University, Guangzhou Women and Children’s Medical Center, The First Affiliated Hospital of Guangzhou Medical University, Peking University People’s Hospital), in which some of the clinical cohort data have been published5,27-29. In order for our model to exhibit good generalizability across diverse populations including different countries and ethnicities, we supplemented them with 431,231 discharge summaries from MIMIC-IV30. After removing duplications, we obtained about 14.8 billion tokens. Combining the above three parts, our first-stage dataset consists of a total of 71.4 billion tokens for pretraining. Note that we have removed case reports of rare diseases, emergencies, and complex diseases for evaluation purposes from the pretraining data, based on their unique identifiers (i.e., PMID). Language data for supervised finetuning. This phase aims to elicit knowledge from the pretrained model by applying supervised finetuning with a collection of high- quality question-answering (QA) pairs. Specifically, we combined the PubMedQA31, MedQA (USMLE)32, MedMCQA33, OpenAssistant Conversations34, and the CoT- Collection35. PubMedQA is a dataset designed for the development and evaluation of machine learning models in the QA domain, specifically targeting biomedical literature. It leverages abstracts from the vast PubMed database. MedQA (USMLE) focuses on medical examination questions derived from the United States Medical Licensing Examination (USMLE). It simulates the breadth and depth of medical knowledge required for licensure, making it an ideal dataset for training and evaluating AI models on medical reasoning and knowledge. MedMCQA, on the other hand, is a comprehensive multiple-choice QA dataset that covers a wide range of medical subjects. The OpenAssistant Conversations dataset is a corpus of dialogues with 161,443 messages, used for basic model-based QA and multi-turn dialogue abilities, which are crucial for interacting with the model during diagnosis. The CoT- Collection dataset includes 1.88 million CoT-formatted QA data for 1,060 tasks, which helps improve the model's performance on complex analytical tasks such as differential diagnosis. We randomly selected 500,000 samples from the CoT- Collection dataset and added them to the second stage. Combining the above parts, the instruction dataset comprises a total of 1.7 million training data entries, with a sf-5927165 total token count of 2 billion. Due to the gap between QA and EHR data, we finetuneda separate model on EHRs using 264 445 rare disease cases and 368,369 cases withemergency conditions. Language data for evaluation. The EHR data was randomly divided into training and test sets, with the test set comprising 958 cases of rare diseases and 2,769 emergency cases, respectively. To incorporate evaluations from real-world rare diseases and emergency conditions, we downloaded case reports from the American Journal of Ophthalmology from 1980-2020 for training.86 rare disease cases and 86 emergency cases in and after 2015 were used for evaluation by MetaGP, GTP-4, and ophthalmologists to provide a final diagnosis, a differential diagnosis list, and reasoning. Similarly, we downloaded case reports from the American Journal of Emergency Medicine, Journal of American College of Emergency Physicians, North American Journal of Medical Sciences, and American Journal of Medical Genetics from 1980-2020.97 rare disease cases and 109 emergency cases in and after 2015 were used for evaluation by MetaGP, GTP-4, and physicians to provide a final diagnosis, a differential diagnosis list, and reasoning. To incorporate real-world complex disease cases, we downloaded 521 common diseases and unusual cases published after (and including) 2010 in the NEJM CPCs, each with the correct diagnosis, along with the full description of the case and the procedures performed. Studies before 2020 were used for training, while the remaining 80 case records published in or after 2020 were used for validation. Note that all evaluation data had been excluded from the pretraining set. Image-report pairs. A portion of CXR images and report data were from the China Consortium of Chest X-ray Image Investigation (CC-CXRI), which were collected from multiple hospitals, including Sun Yat-sen Memorial Hospital and the Third Affiliated Hospital, both affiliated with Sun Yat-sen University, West China Hospital, Guangzhou Medical University First Affiliated Hospital, Nanjing Renmin Hospital, the First affiliated hospital of Anhui Medical University, and Yichang Central People’s Hospital. The demographics and clinical information of the cohort participants are described in previous studies28,29. CT scans and report data were from cohorts of the China Consortium of Chest CT Image Investigation (CC-CCII), which consists of Sun Yat-sen Memorial Hospital and Third Affiliated Hospital of Sun Yat- sen University, The First Affiliated Hospital of Anhui Medical University, West China Hospital, Nanjing Renmin Hospital, Yichang Central People’s Hospital, Peking University Renmin Hospital. The demographics and clinical information of the cohort participants were described in the previous study27. A large retinal image dataset was from cohorts of the China Consortium of Eye Image Investigation (CC-EII). The demographics and clinical information of the cohort participants are described in previous studies3,4,36. To train the multimodal report generator, we collected and integrated CXR, CT, and four types of retinal images from the above datasets and public datasets, resulting in a collection of over 1 million image-report pairs spanning 6 imaging modalities: OCT, retinal fundus photographs, FFA, ICGA, CXR, and CT. For OCT data, we incorporated image-report pairs from China Eye Image Investigation and randomly sf-5927165 split the data based on patient identifications to build the training and internal validation sets (the ratio of the training to the validation set is 9:1) (71,558 data pairs)4. We used data from the Eye Hospital of Wenzhou for external validation (1,110 data pairs). For fundus images, the training and internal validation data include 313,956 fundus photograph-report pairs3,4. For prospective validation, we used image-report pairs collected in or after May 2018 from Wenzhou Eye Hospital.10% of the remaining data were split out as the internal validation set. The external validation set consists of 1,775 data pairs. Different from OCT and fundus images, cases of FFA or ICGA usually comprised more than one image with different locations and time- sequence information. The FFA dataset consists of 4,895 reports associated with 67,402 images, and the ICGA data involves 1,131 reports alongside 20,132 images. As for CXR data, we combined the public MIMIC-CXR dataset (377,110), IU-Xray dataset (7,470), and our internal data from the China Consortium of Chest X-ray Image Investigation (CC-CXRI) (222,308)28. The official test split of MIMIC-CXR was used for interval validation, while the IU-Xray dataset served as the external validation set. For CT data, we integrated 23,690 volume-report pairs from the China Consortium of Chest CT Image Investigation (CC-CCII)27where 10% of data were used for internal validation. To build the prospective validation set, we exploited the data collected in or after June 2019. Approvals for the research were secured from the Institutional Review Board / Ethics Committees of all participating hospitals or institutes. Consent forms were signed by all involved patients. The study was carried out in accordance with the United States Health Insurance Portability and Accountability Act. Furthermore, it was in compliance with the principles of the Declaration of Helsinki and conformed to the policies of the Chinese Center for Disease Control and Prevention, as well as the Chinese Health Laws. Framework Although general-purpose LLMs like LLaMA-2 and GPT-4 have showcased strong performance across various tasks in benchmarks such as BIG-bench37, their utilization in the medical field necessitates adaptation and alignment with domain- specific data due to the inadequacy of domain knowledge. Hence, we implemented a two-stage training strategy comprising pretraining and supervised finetuning to enrich our language model with more medical knowledge and medical abilities. Pretraining. MetaGP was initialized with the parameters of Llama2-13B and subsequently further pre-trained on our medical text corpora, which encompassed medical articles from PubMed, text content from publicly available medical books (including guidebooks and medical genetic textbooks), and clinical text data (mainly EHRs). The corpora were tokenized and concatenated into a sequence of tokens, which were then chunked into fixed-length pieces for training. Given a tokensequence , , … , , the optimization objective is to minimize auto-regressive loss formulated as | , where is our model,denotes tokens before and is the fixed sequence length. sf-5927165 Supervised finetuning. In the second training phase, we incorporated the instruction tuning approach38, defined as finetuning a pretrained large language model using a dataset comprising instructional inputs and their corresponding responses. The instructions that we used for medical QA and EHR data are presented in Table 1. The training loss was the same auto-regressive loss in the pretraining stage and calculated based on the output tokens. For the remaining data, each sample consists of one or multiple rounds of dialogue content, where each round of dialogue includes instructions describing the task, optional instance inputs providing supplementary information required for task resolution, and the expected output corresponding to that round of dialogue. The training loss was computed based on the full dialog tokens. Long context window. The default context window size of Llama2 is 4,096. However, in downstream medical scenarios, there are many predictive tasks involving long texts, such as differential diagnosis based on extensive patient information. In these situations, the total length of patient information exceeds the default context window size of Llama2. Additionally, previous studies have demonstrated a significant performance degradation when handling inputs that surpass the pretrained context window size during text completion38. Therefore, in both training phases, we employed linear rope scaling to extend the context window. Specifically, we directly downscaled position indices in the Rotary Position Embedding (RoPE) to match the previous context window limit and calculated the intermediate values of the position encodings between adjacent integer positions. Medical report generator. To deal with the varying dimensions across different imaging modalities, we appended two vision encoders with different architectures – Swin Transformer39and ViT40– on top of Me. Swin Transformer is an innovative architecture in the field of computer vision that utilizes shifted windows to provide a hierarchical and efficient way of processing images. This design enables the model to focus on both local and global features of an image, making it particularly effective for vision tasks39. In practice, the developed medical report generator used a Swin Transformer to process medical images, including OCT, retinal Fundus photographs, and CXR. However, the Swin Transformer was designed to handle 2D image data and is not capable of tackling 3D volumetric arrays. To address this, we used ViT-3D, a transformer-based volumetric processing model41. The core idea is to first divide large 3D volumes into smaller ones, on top of which visual representation learning was further applied. In our pipeline, ViT-3D was employed to deal with FFA, ICGA, and CT to capture relations and dependencies across images or slices. After transforming raw pixels or voxels into latent embeddings using vision backbones, the obtained visual representations were forwarded to an LLM to generate medical reports. However, images or volumes are high-dimensional data, and converting them into dense embeddings creates a challenge due to the generation of lengthy token sequences (e.g., 768 tokens were generated by ViT-3D), which impairs the training and inference efficiency of the following MetaGP. To address this issue, we applied dimension reduction to adjacent visual tokens and effectively decreased the visual token count by a factor of 6 at least. In practice, this strategy can dramatically reduce memory costs and speed up model After obtaining the reduced visual tokens, sf-5927165 we used linear projectors to map visual representations to the language space. In addition to visual tokens, the language input also comprises demographic information and textual tags that indicate the input modalities. All of these were combined and fed to the MetaGP in each forward pass. Implementations. The scaling factor used in linear rope scaling was 4, resulting in a context window of 16,384 for MetaGP. During the training phase, we used a global batch size of 384. Optimization was performed using AdamW42with an initial learning rate of 1e-5. The loss function employed was autoregressive loss. We utilized DeepSpeed Zero-343to accelerate the training process, based on the bfloat16 format, while also implementing the gradient checkpointing strategy to increase the batch size. In the foundation phase, the training process was conducted for one epoch using 24 NVIDIA A100 GPUs (80G), while in the instruction-tuning stage, we trained the model for 3 epochs using 16 A100 units. For GPT-4, we used the version of 2024-01- 25-preview for model testing and evaluation. To build the report generator, we used the base version of Swin Transformer, which has 4 stages, a window size of 7, a patch size of 4, and an initial feature dimension of 128. For data of OCT, retinal fundus, and CXR, we first cropped a random portion of each image, where the area of the cropped region with respect to that of the original image is between 0.5 and 1.0. Then, we resized each crop to 224 (Height) ×224 (Width) pixels with 3 channels. We used pretrained weights on ImageNet. The ViT- 3D model has 12 layers and an embedding dimension of 768. Each input volume was first resized to 224 (Height) × 224 (Width) × 64 (Depth) and further divided into subvolumes with no overlap, where the size of each subvolume is 16 (Height) × 16 (Width) × 8 (Depth). To reduce the count of visual tokens, we leveraged the adaptive average pooling strategy and fixed the output length to 9. We implemented two linear projectors for 2D and 3D data, respectively, where each projector includes one fully connected layer, mapping each pooled visual token to a 1D vector that has 4096 elements. Before the projector, layer normalization44was added to avoid exploding gradients. End-to-end training was applied, where all parameters of vision encoders and linear projectors are trainable. For the LLM part, we used the LoRA, i.e., low- rank adaptation strategy45, for parameter-efficient training. LoRA relies on the mathematical concept of low-rank matrix decomposition. In this approach, a large weight matrix of a neural network layer (which is typically high-rank) is approximated using two smaller matrices. This decomposition reduces the number of trainable parameters. By using low-rank matrices, LoRA modifies only a small fraction of the model's parameters during the training stage. To achieve this, we set the rank and alpha value of LoRA to 16, respectively. AdamW42was used as the default optimizer along with the cosine learning rate scheduler. The initial learning rate was set to 1e-4, and the minimum learning rate was 3e-6. The total number of training iterations is 80,000. We applied linear warmup in the first 2,000 training iterations, and the warmup learning rate is 1e-7. The training stage can be finished within 48 hours using 16 NVIDIA A100 GPUs (80G) with a global batch size of 64. Data Availability sf-5927165 All used academic articles can be downloaded from the PubMed Central portal (https: / / www.ncbi.nlm.nih.gov / pmc / ). We also incorporated text data from the following publicly available datasets: PubMedQA (https: / / github.com / pubmedqa / pubmedqa), MedQA (https: / / github.com / jind11 / MedQA), MedMCQA (https: / / github.com / MedMCQA / MedMCQA), MIMIC-IV (https: / / physionet.org / content / mimiciv / 2.2 / ), OpenAssistant Conversations (https: / / huggingface.co / datasets / OpenAssistant / oasst1), and CoT-Collection (https: / / github.com / kaistAI / CoT-Collection). For image-report pairs, we used MIMIC-CXR (https: / / physionet.org / content / mimic-cxr / 2.0.0 / ) and IUxray (https: / / openi.nlm.nih.gov / ). Restrictions apply to the availability of the EHR data, which were used with the permission of the participants for the current study. De- identified data is available for research purposes from the corresponding authors upon reasonable request. sf-5927165 Reference 1. Liu, Y., et al. A deep learning system for differential diagnosis of skin diseases. Nat Med 26, 900-908 (2020). 2. Daneshjou, R., et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci Adv 8, eabq6147 (2022). 3. Zhang, K., et al. Deep-learning models for the detection and incidence prediction of chronic kidney disease and type 2 diabetes from retinal fundus images. Nat Biomed Eng 5, 533-545 (2021). 4. Kermany, D.S., et al. Identifying Medical Diagnoses and Treatable Diseases by Image- Based Deep Learning. Cell 172, 1122-1131.e1129 (2018). 5. Liang, H., et al. Evaluation and accurate diagnoses of pediatric diseases using artificial intelligence. Nat Med 25, 433-438 (2019). 6. Jiang, L.Y., et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357-362 (2023). 7. Rajkomar, A., et al. Scalable and accurate deep learning with electronic health records. NPJ Digit Med 1, 18 (2018). 8. Hannun, A.Y., et al. Cardiologist-level arrhythmia detection and classification in ambulatory electrocardiograms using a deep neural network. Nat Med 25, 65-69 (2019). 9. Liu, R., et al. Evaluating eligibility criteria of oncology trials using real-world data and AI. Nature 592, 629-633 (2021). 10. Bommasani, R., et al. On the Opportunities and Risks of Foundation Models. (arXiv, 2022). 11. Brown, T.B., et al. Language Models are Few-Shot Learners. (arXiv, 2020). 12. Singhal, K., et al. Large language models encode clinical knowledge. Nature (2023). 13. Touvron, H., et al. Llama 2: Open Foundation and Fine-Tuned Chat Models. (arXiv, 2023). 14. Touvron, H., et al. LLaMA: Open and Efficient Foundation Language Models. (arXiv, 2023). 15. McDuff, D., et al. Towards Accurate Differential Diagnosis with Large Language Models. (arXiv, 2023). 16. Li, C., et al. LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day. (arXiv, 2023). 17. Moor, M., et al. Med-Flamingo: a Multimodal Medical Few-shot Learner. (arXiv, 2023). 18. Kaplan, J., et al. Scaling Laws for Neural Language Models. (arXiv, 2020). 19. Tu, T., et al. Towards generalist biomedical ai. NEJM AI 1, AIoa2300138 (2024). 20. Ouyang, L., et al. Training language models to follow instructions with human feedback. Vol.35 (eds. Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K. & Oh, A.) 27730-27744 (Curran Associates, Inc., 2022). 21. Abràmoff, M.D., et al. Considerations for addressing bias in artificial intelligence for health equity. NPJ Digit Med 6, 170 (2023). 22. Zou, J., Gichoya, J.W., Ho, D.E. & Obermeyer, Z. Implications of predicting race variables from medical Science 381, 149-150 (2023). sf-5927165 23. Seyyed-Kalantari, L., Zhang, H., McDermott, M.B.A., Chen, I.Y. & Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat Med 27, 2176-2182 (2021). 24. Perry, T.S. Move over, moore's law. Make way for huang's law [Spectral lines]. IEEE Spectrum 55, 7-7 (2018). 25. Hobbhahn, M. & Besiroglu, T. Predicting GPU performance. (2022). 26. Gao, L., et al. The Pile: An 800GB Dataset of Diverse Text for Language Modeling. (arXiv, 2020). 27. Zhang, K., et al. Clinically Applicable AI System for Accurate Diagnosis, Quantitative Measurements, and Prognosis of COVID-19 Pneumonia Using Computed Tomography. Cell 182, 1360 (2020). 28. Wang, G., et al. A deep-learning pipeline for the diagnosis and discrimination of viral, non-viral and COVID-19 pneumonia from chest X-ray images. Nat Biomed Eng (2021). 29. Zhou, H.-Y., et al. A transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics. Nat Biomed Eng, 1-13 (2023). 30. Johnson, A.E.W., et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data 10, 1 (2023). 31. Jin, Q., Dhingra, B., Liu, Z., Cohen, W.W. & Lu, X. Pubmedqa: A dataset for biomedical research question answering. arXiv . 32. Jin, D., Pan, E., Oufattole, N., Weng, W.-H., Fang, does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences 11, 6421 (2021). 33. Pal, A., Umapathi, L.K. & Sankarasubbu, M. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. in Conference on health, inference, and learning 248-260 (PMLR, 2022). 34. Köpf, A., et al. OpenAssistant Conversations -- Democratizing Large Language Model Alignment. (arXiv, 2023). 35. Kim, S., et al. The CoT Collection: Improving Zero-shot and Few-shot Learning of Language Models via Chain-of-Thought Fine-Tuning. (arXiv, 2023). 36. Cai, W., et al. EyeHealer: A large-scale anterior eye segment dataset with eye structure and lesion annotations. Precis. Clin. Med.4, 85-92 (2021). 37. Srivastava, A., et al. Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models. Transactions on Machine Learning Research (2023). 38. Wei, J., et al. Finetuned Language Models Are Zero-Shot Learners. (arXiv, 2022). 39. Liu, Z., et al. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. in 2021 IEEE / CVF International Conference on Computer Vision (ICCV) 9992-10002 (2021). 40. Dosovitskiy, A., et al. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. (arXiv, 2021). 41. Chen, Z., Agarwal, D., Aggarwal, K., Safta, W., Balan, M.M. & Brown, K. Masked image modeling advances 3d medical image analysis. in Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision 1970-1980 (2023). 42. Loshchilov, I. & Hutter, F. Decoupled Weight Decay Regularization. (arXiv, 2019). sf-5927165 43. Rasley, J., Rajbhandari, S., Ruwase, O. & He, Y. DeepSpeed: System optimizations enable training deep learning models with over 100 billion parameters. 3505-3506 (Association for Computing Machinery, 2020). 44. Ba, J.L., Kiros, J.R. & Hinton, G.E. Layer Normalization. (arXiv, 2016). 45. Hu, E.J., et al. LoRA: Low-Rank Adaptation of Large Language Models. (arXiv, 2021). sf-5927165 Figure legend Figure 1 | Overview of development of MetaGP. a, MetaGP was trained using academic articles, medical images and associated reports, electronic health records, natural language data from the Internet, and medical textbooks. The resulting model can provide decision support for multimodal imaging report generation, rare disease diagnosis, emergency condition identification, and complex disease diagnostics. b, Composition of the text data for pretraining. c, Composition of image sets from different modalities with the total number approximating 1 million. d, Diagnostic performance scores achieved by the combination of senior physicians and MetaGP, MetaGP alone, and GPT-4 when evaluated on real-world clinical cases involving rare diseases and emergency conditions. In each category, we average the diagnostic scores of two subgroups: ophthalmic and systemic diseases. Figure 2 | Evaluation schema. The schema integrates human and automated evaluation methods, combining the expertise of medical professionals from relevant fields to conduct pertinent assessments. a, Evaluation of rare disease diagnoses. b, Evaluation of emergency condition diagnoses. c, Evaluation of complex disease diagnoses. d, Evaluation of generated medical reports for multimodal images. Figure 3 | Human evaluation for rare disease diagnoses on case reports. a, Average diagnostic performance of MetaGP, GPT-4, and ophthalmologists with varying experience levels for rare ophthalmic diseases. b, Distributions of scores (from -2 to 2) of individuals for rare ophthalmic disease diagnoses. c, Average diagnostic performance of MetaGP, GPT-4, and general practitioners with varying experience levels for rare systemic diseases. d, Distributions of scores (from -2 to 2) of individuals for rare systemic disease diagnoses. Figure 4 | Average accuracy and F1 scores of MetaGP and GPT-4 in diagnosing rare diseases and emergency conditions. Figure 5 | Human evaluation of emergency conditions diagnoses on case reports. a, Average diagnostic scores of MetaGP, GPT-4, and ophthalmologists with varying experience levels for ophthalmic emergencies. b, Distributions of scores (from -2 to 2) of individuals for ophthalmic emergency diagnoses. c, Average diagnostic scores of MetaGP, GPT-4, and emergency physicians with varying experience levels for systemic emergencies. d, Distributions of scores (from -2 to 2) of individuals for systemic emergency diagnoses. Figure 6 | Statistical analyses and per-category results of MetaGP and GPT-4 in diagnosing rare diseases and emergency conditions. The statistic significances were calculated by Mann Whitney Wilcoxon Test. Figure 7 | Automated evaluation for complex disease diagnoses. We report the top-n accuracy of MetaGP, Google LLM for DDx, and GPT-4 on case studies presented by the New England Journal of Medicine clinicopathological conferences. Figure 8 | Multimodal medical report generation. We evaluated the report generation results of MetaGP, Med-Flamingo, and LLaVA-Med on 6 imaging modalities. Specifically, we did external validation on OCT, retinal fundus photographs, and sf-5927165 CXR, while prospective validation was conducted on retinal fundus photographs and CT. Evaluation metrics include the ROUGE-L and BertScore. Figure 9 | Physician evaluation of medical reports. This image shows a comparative analysis of AI-generated and physician-composed medical image reports on 6 imaging modalities (a, OCT, b, Fundus, c, FFA, d, ICGA, e, CXR, f, CT). For each modality, we employed three raters, including one senior (rater 1) and two junior (rater 2 and 3) physicians, to do the pairwise human evaluation. In a blinded assessment, where reports were deidentified, raters were tasked with judging the quality of the reports, determining which was superior or if there was a tie. Figure 10 | Examples of MetaGP-generated medical reports. sf-5927165 Table Table 1 | Templates for applying supervised finetuning and performing evaluation on PubMed case reports and samples from NEJM CPCs. Medical question-answering ### Instruction: Choose the correct answer from the given options for the question. Input: [Question] [Options] Output: [Correct option] EHR ### Instruction: You are an empathetic and experienced doctor. Based on the given patient's record, complete the following three tasks: 1. Determine the most probable diagnosis. 2. Offer a detailed explanation supporting the diagnosis. 3. Develop an appropriate treatment plan according to the patient's condition and your diagnosis. Input: [Demographic characteristics] [Chief complain] [Present medical history] [Past medical history] [Family medical history] [Medical examination results] Output: [Diagnosis] [Basis of diagnosis] [Treatment plan] OpenAssistant Conversations For i in 1, …, m ### Instruction: [Human question i] ### Response: [Model answer i ] CoT-Collection ### Instruction: [Source] ### Response: The answer is [Target]. Let’s think step by step. [Rationale] PubMed case report ### Instruction: Based on the given case report, complete the following four tasks: 1. Provide the ten most likely differential diagnoses, with the more probable diagnoses listed first. 2. Based on the above differential diagnoses, provide your best final diagnosis. 3. Provide a detailed analysis of the above two steps, totaling no fewer than 500 words. 4. Based on the given case report and your final diagnosis, provide the corresponding treatment, totaling no fewer than 200 words. The format should be: Top 10 differential diagnosis: 1. ddx1 2. ddx2 3.... Final diagnosis: ddx1, Detailed analysis: XXX, Treatment: XXX NEJM CPC sf-5927165 ### Instruction: What are the top 10 most likely diagnoses? Be precise, listing one diagnosis per line, and try to cover many unique possibilities (at least 10). The top 10 diagnoses are: Supplementary Table 1 | Qualitative examples of ophthalmic diseases. Supplementary Table 2 | Qualitative examples of systemic diseases. Supplementary Table 3 | Qualitative examples of EHRs. sf-5927165 Ophthalmic Disease Representative CasesID 1A 72-year-old woman developed nonpainful blurred vision in her left eye. The patient had a history of angina pectoris and well_x0002_controlled diabetes mellitus. She also had a history of liver enzyme elevation secondary to systemic corticosteroids; the patient did not remember the reason for which she was taking the corticosteroids and had not been on systemic steroids since then. She had no history of ophthalmologic disease. At the first ophthalmological consultation, she had optic disc swelling and choroidal folds in both eyes and subretinal fluid in the left eye. Her best-corrected visual acuity (BCVA) was 20 / 16 in the right eye and 20 / 29 in the left eye. The critical flicker-fusion frequency (CFF), which is impaired in patients with optic nerve damage or retinal damage, was low at 24 Hz in the right eye and 22 Hz in the left eye (reference range, 52.5 ± 4.4 Hz). She was followed up by the same doctor for 1 month with steroid eye drops, but she was not treated with systemic therapy. Her symptom did not improve, and she was referred to Hiroshima University Hospital 1 month after symptom onset. At the first visit to our clinic, she still complained of blurred vision. Her BCVA was 20 / 20 in the right eye and 20 / 40 in the left eye. The intraocular pressure was 17 mmHg in both eyes, and the CFF had improved to 33 Hz in both eyes. The pupillary direct light reflex was normal in both eyes, and no relative Case afferent pupillary defect was seen in either eye. No conjunctival injection and no cells in the anterior segment were observed in either eye. However, 1 + vitreous cells were seen in the left eye. She was found to have optic disc swelling in both eyes, but it was improved compared with the findings of the previous doctor. The bilateral choroidal folds and the subretinal fluid in the left eye had resolved without any treatment. Fluorescence angiography (FA) showed slight hyperfluorescence in both optic nerves but no leakage or pooling. Indocyanine green (ICG) angiography showed no leakage or dark spots. Laboratory tests for syphilis, tuberculosis, anti- aquaporin 4 antibody, and angiotensin-converting enzyme were normal or negative. Bilateral hilar lymphadenopathy, which is seen in sarcoidosis, was not detected on a chest radiograph. A visual field test showed decreased sensitivity in the left eye. A B-scan ultrasound image showed slight fluid in sub-Tenon's space and mild increase in scleral thickness in the left eye; additionally, the optic nerve sheath was displayed as a highly reflective circle with shadowing around the optic disc in both eyes.e. MRI with interpretation by a radiologist was then planned. MRI with intravenous contrast revealed increased signal intensity around the optic nerve (i.e., the “tram track sign”) in both eyes.Final diagnosis Bilateral Optic perineuritis with posterior scleritis
[0002] Top 10 differential diagnosis: 1. Optic neuritis 2. Acute macular neuroretinopathy 3. Posterior scleritis with bilateral Optic perineuritis 4. Acute zonal occult outer retinopathy (AZOOR) 5. Multiple evanescent white dot syndrome (MEWDS) 6. Toxic amblyopia 7. Acute idiopathic blind spot enlargement (AIBSE) 8. Retinitis pigmentosa 9. Leber hereditary optic neuropathy (LHON) 10. Sarcoidosis Final diagnosis: Posterior scleritis with bilateral Optic perineuritis Detailed analysis: In this case, the patient presented with nonpainful blurred vision in the left eye, accompanied by optic disc swelling and choroid folds. The critical flicker-fusion frequency was low in both eyes, suggesting optic nerve involvement. The patient was initially treated with steroid eye drops, but there was no improvement MetaGP diagnosis in her symptoms. One month later, she was referred to our hospital, where we performed a comprehensive evaluation, including fluorescein angiography, indocyanine green angiography, laboratory tests, and imaging studies. Fluorescein angiography showed slight hyperfluorescence in both optic nerve heads, but no leakage or pooling. Indocyane green angiography showed no leakage or dark spot characteristic for VKH. Laboratory tests for syphilis, tissue bacilli, anti-aquaporin 4 antigen, and angiotensin-converting en-zyme were all within normal limits. Chest X-ray showed no evidence of hilar lymphadenopathy. MRI with intravenous contrast revealed increased T2 signal intensity around the optic nerve (the "tram track sign") in both eyes, consistent with posterior scleritis. This finding was confirmed by B-scan ultrasound, which showed fluid in the sub-Tenon's space and increased scleral thickness. Posterior scleritis is an inflammation of the sclera that can cause pain, redness, and blurred vision. It is often associated with systemic autoimmune diseases such as rheumatoid arthritis, systemic lupus erythematosus, and Sjogren's syndrome. In this case, the patient had a history of angina pectoralis and well-controlled diabetes mellitus, but no other signs or symptoms suggestive of an underlying autoimmune disorder.MetaGP score 2
[0003] Top 10 differential diagnosis: 1. Optic neuritis 2. Papilledema 3. Idiopathic intracranial hypertension 4. Central retinal vein occlusion 5. Ischemic optic neuropathy 6. Choroidal neovascularization 7. Retinal detachment 8. Multiple sclerosis 9. Vasculitis 10. Optic nerve sheath meningioma Final diagnosis: Optic neuritis Detailed analysis: The patient presented with nonpainful blurred vision in her left eye, with a history of angina pectoris, diabetes mellitus, and previous use of systemic corticosteroids. The initial ophthalmological GPT4 diagnosis examination showed optic disc swelling and choroidal folds in both eyes, along with subretinal fluid in the left eye. BCVA and CFF were impaired, indicating optic nerve damage or retinal damage. The patient was initially treated with steroid eye drops but did not improve. At the visit to Hiroshima University Hospital, her symptoms persisted, and she still had optic disc swelling, although it had improved compared to the previous examination. The choroidal folds and subretinal fluid had resolved. FA showed slight hyperfluorescence in both optic nerves, while ICG angiography showed no leakage or dark spots characteristic of VKH disease. Laboratory tests for other possible causes were normal or negative. The visual field test showed decreased sensitivity in the left eye. B-scan ultrasound revealed slight fluid in sub-Tenon's space and scleral thickness, indicative of posterior scleritis, and the optic nerve sheath appeared abnormal in both eyes. MRI with intravenous contrast confirmed the diagnosis of optic neuritis, showing increased signal intensity around the optic nerve in both eyes. Treatment: Optic neuritis is often treated with systemic corticosteroids to reduce inflammation and preserve vision. The patient should be started on high-dose intravenous corticosteroids, followed by a tapering regimen of oral corticosteroids. Close monitoring of visual acuity, color vision, and visual fields is recommended during and after treatment. Additionally, the patient should be educated about the potential for recurrence and the need for regular follow-up with an ophthalmologist. Other possible treatments, such as disease-modifying therapies for multiple sclerosis, may be considered based on further evaluation and consultation with a neurologist.GPT4 score 1Ophthalmologist Retinal vasculitis diagnosis Ophthalmologist 0 score MetaGP assisted ophthalmologist's Scleritis diagnosis MetaGP assisted Ophthalmologist 2 scoreID 2A 3-day-old Caucasian boy who was born at term was referred for bilateral leukocoria. In utero, congenital heart defects had been diagnosed, including a large ventricular septal defect, an atrial septal defect, and a hypoplastic aortic arch. At birth, he was also found to have microcephaly, ventriculomegaly, bowel dilatation, and hearing loss. External examination showed elongated palpebral fissures, wide and arched eyebrows, and low-set, posteriorly rotated ears. On ophthalmic examination, the baby did not wince to light in both eyes. The anterior chambers were formed but shallow, and there were bilateral vascular pupillary membranes. Dilation was poor, but there appeared to be cataracts and retrolental fibrovascular plaques with no view to the retina. B-scan Case echography demonstrated hyperechoic lesions extending from the optic nerve to the posterior surface of the lens, concerning for bilateral funnel RD. None of the family had undergone genetic testing. After the baby was born, focused chromosomal microarray (performed by a CLIA-certified laboratory) demonstrated a 1.9 Mb hemizygous deletion of chromosome Xp11.3 that included the NDP and KDM6A genes. This microdeletion was inherited from his heterozygous mother who also underwent genetic testing that confirmed an identical microdeletion. MRI imaging of the brain revealed microcephaly and hydrocephalus. Electroencephalogram confirmed the presence of seizure episodes.Final diagnosis Norrie disease together with X- Kabuki syndrome Top 10 differential diagnosis: 1. Kabuki syndrome with Norrie Disease 2. Trisomy 18 3. Down syndrome 4. Fetal alcohol spectrum disorder 5. Edward syndrome 6. Patau syndrome 7. Turner syndrome Ophthalmologist -1 score MetaGP assisted ophthalmologist's Kabuki syndrome with Norrie Disease dia nosis resection was performed along with partial lamellar sclerouvectomy, pars plana vitrectomy, and silicon oil Case tamponade. The tumour was completely removed during the operation. Pathological examination was performed vesicular nuclei, vacuolar chromatin and prominent nucleoli, and rich in interstitial mucus-like with neutrophil infiltration. Spindle-shaped myofibroblasts are loosely arranged, surrounded by edema-like mucus-like background, with a large number of blood vessels and inflammatory cells,plasma cells, lymphocytes, and eosinophils. No clear zonal maturation was seen under the microscope. The cells are spindle-shaped or oval. Nuclear atypia was not obvious. Nucleoli are not prominent. Lymphocyte infiltration was scattered around, and more lymphocytes aggregated around small blood vessels. The proliferation index of tumor cells is 2%(ki-67 2%). Immunohistochemical analysis revealed positive staining results for anaplastic lymphoma kinase (ALK), desmin, and vimentin, as well as negative staining results for smooth muscle actin (SMA), S-100, Ki-67, CK, CD68, and calponin. Additional immunohistochemical experiments were performed for the patient, and the patient’s p53 was wild-type. The patient was followed up regularly after discharge, and the condition of body and eyes was stable. The silicone oil removal operation was performed in our hospital, and the operation went well. Final diagnosis Intraocular inflammatory Myofibroblastoma Tumors (IMTs)
[0004] Top 10 differential diagnosis: 4. Ciliary body lymphoma 5. Ciliary body schwannoma / neurofibroma 6. Ciliary body metastasis 7. Ciliary body osteoma 8. Retinoblastoma 9. Ciliary body hemangioma 10. Ciliary body hemangiopericytoma Final diagnosis: Ciliary body melanoma GPT4 diagnosis Detailed analysis: The patient presented with progressive, painless vision loss and a well-defined, ovoid amelanotic mass in the inferonasal quadrant of the left eye. The mass was found to be in the ciliary body and showed medium to high echodensity on ultrasound and enhancement on MRI. Pathological examination revealed a collagenized fibrous tumor with spindle cell proliferation and chronic inflammatory cell infiltration. Immunohistochemical analysis showed positive staining for ALK, desmin, and vimentin, and negative staining for SMA, S-100, CK, CD68, and calponin. These findings are consistent with a diagnosis of ciliary body melanoma. Treatment: The patient underwent trans-scleral resection, lamellar sclerouvectomy, pars plana vitrectomy, and silicon oil tamponade. The tumor was completely removed during the operation. Regular follow-up and silicone oil removal were performed without complications. Further treatment may be required based on the patient's condition and the risk of recurrence or metastasis.GPT4 score -2Ophthalmologist Intraocular melanoma diagnosis Ophthalmologist -1 score , glucose at 100 mL / h) during dialysis to prevent IOP rise. Differences between IOP before and after i.v. glucose and panretinal photocoagulation treatments. In the last ophthalmological evaluation on 17 February 2015, BCVA was 2 / 20 in RE, anterior segment examination showed reduction of iris rubeosis in RE, IOP was 8 mmHg RE and 9 mmHg LE and fundoscopic examination revealed successful argon laser treatment of ischaemic peripheral retinal areas in both eyes. The optic disk appear unchanged and the visual field stable after the two treatments. Alpha 2 -agonist (Alphagan®) and pilocarpine hydrochloride 2 % eye drops wereFinal diagnosis Neovascular glaucoma
[0005]
[0006] assisted 2 doctor scoreID 2
[0007] Final Left hemopneumothorax diagnosis 8. Lung abscess with secondary infection and bleeding 9. Esophageal rupture with mediastinitis and hemothorax 10. Lung cancer with associated hemorrhage Final diagnosis: Traumatic pneumothorax with iatrogenic hemothorax GPT4 Detailed analysis: The patient presented with acute onset chest pain and shortness of breath, which is consistent diagnosis with a pneumothorax. The chest x-ray showed a hydropneumothorax with a large pneumothorax component and left-sided effusion, suggesting the presence of both air and blood in the thoracic cavity. The patient's history of smoking may have contributed to the development of a spontaneous pneumothorax, but the presence of recent trauma (tube thoracostomy) suggests an iatrogenic cause for the pneumothorax. The large amount of blood drained from the chest after tube thoracostomy supports the diagnosis of iatrogenic hemothorax. The patient's hypotension and need for massive transfusion further suggest significant blood loss. Treatment: The patient should continue to receive supportive care, including monitoring of vital signs, oxygen supplementation, and pain management. The pneumothorax and hemothorax should be further evaluated and managed. A repeat chest x-ray or CT scan may be needed to assess the adequacy of tube thoracostomy and identify any ongoing bleeding or complications such as lung injury or infection. If significant bleeding continues or if there are signs of ongoing instability, surgical intervention such as thoracotomy may be required to control bleeding and repair any underlying injury. The patient should also receive appropriate antibiotics and prophylaxis for infection, considering the recent positive COVID-19 PCR result. Close monitoring of the patient's respiratory status, hemodynamics, and laboratory values should continue throughout their hospital stay.GPT4 score 2 , , in severe cases. Genetic counseling and screening of family members is also recommended. In this case, the patient's family was informed about the diagnosis and referred for further genetic testing. It is important to educate patients and their families about the risks and symptoms of ARVC so they can seek medical attention promptly if any signs or symptoms arise. Regular follow-up appointments with a cardiologist are also essential to monitor the progression of the disease and adjust treatment plans accordingly. scoreID 5
[0008] MetaGP 2 score tsesses ses ses ses sessessessessessessessessessessessessessessesses ses sessessessessessessessesses sessessessessesses ses ses ses sessesses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses ses E14 Unspecified diabetes mellitus Endocrine, Nutritional, and Metabolic DiseasesEmergency conditions E16 Disorder of pancreatic internal secretion, Endocrine, Nutritional, and Metabolic Diseases Emergency unspecified conditionsE26 Hyperaldosteronism, unspecified Endocrine, Nutritional, and Metabolic DiseasesEmergency conditionsE71 Other peroxisomal disorders Endocrine, Nutritional, and Metabolic DiseasesEmergency conditions Disorder of amino-acid metabolism, Emergency E72 unspecified Endocrine, Nutritional, and Metabolic Diseases conditionsG00 Bacterial meningitis, unspecified Diseases of the Nervous SystemEmergency conditionsG11 Hereditary ataxia, unspecified Diseases of the Nervous SystemEmergency conditions Diseases of the Musculoskeletal System and Myositis, unspecifiedEmergency Connective Tissue conditions Diseases of the Musculoskeletal System and Fibroblastic disorder, unspecifiedConnective Emergency Tissue conditions Diseases of the Musculoskeletal System and Disorder of continuity of bone, unspecifiedConnective Emergency Tissue conditions Renal tubulo-interstitial disease, unspecified Diseases of Genitourinary SystemEmergency conditions Abnormal uterine and vaginal bleeding, Diseases of Genitouri Emergency unspecified nary System conditions u osaca sp e a pe s, sequea o e e a causes co os
Claims
Claims 1. A method for generating a medical diagnosis for a patient, comprising: receiving a natural-language prompt for obtaining the medical diagnosis and a set of data related to the patient; and generating the medical diagnosis by Inputting the prompt and the set of data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage.
2. The method of claim 1, wherein the medical diagnosis comprises one or more ICD-10 codes.
3. The method of any one of the preceding claims, wherein the medical diagnosis relates to one or more diseases.
4. The method of claim 3, wherein the one or more diseases comprise an ophthalmic disease or a systemic disease.
5. The method of any one of the preceding claims, wherein the medical diagnosis relates to one or more emergency conditions.
6. The method of claim 5, wherein the one or more emergency conditions are related to an ophthalmic emergency.
7. The method of any one of the preceding claims, wherein the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof.
8. The method of any one of the preceding claims, wherein the LLM is trained to minimize an auto-regressive loss in the first stage.
9. The method of claim 8, wherein the LLM is trained to minimize the auto- regressive loss in the second stage.
10. A method for generating a medical report for a patient, comprising: receiving a natural-language prompt for obtaining the medical report and a set of image data related to the patient; and generating the medical report by Inputting the prompt and the set of image data in a trained large language model (LLM), wherein the LLM is trained using a textual corpus in a first stage and trained using a question-answering (QA) dataset in a second stage.
11. The method of claim 10, wherein the set of image data comprises one or more ophthalmic images, one or more radiological images, or any combination thereof. sf-592716512. The method of claim 11, wherein the one or more ophthalmic images comprise an optical coherence tomography (OCT) image, a retinal fundus photograph, a fundus fluorescein angiography (FFA) image, an indocyanine green angiography (ICGA) image, or any combination thereof.
13. The method of claim 11, wherein the one or more radiological images comprise a chest X-ray (CXR) image, a computed tomography (CT) image, or any combination thereof.
14. The method of any one of claims 10-13, wherein the textual corpus comprises one or more electronic health records, one or more academic papers, one or more medical textbooks, or any combination thereof.
15. The method of any one of claims 10-14, wherein the LLM is trained to minimize an auto-regressive loss in the first stage.
16. The method of claim 15, wherein the LLM is trained to minimize the auto- regressive loss in the second stage.
17. The method of any one of claims 10-16, wherein the LLM comprises one or more vision encoders.
18. The method of claim 17, wherein the one or more vision encoders comprise a Swin transformer and / or a vision transformer.
19. A system, comprising one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of methods 1-18.
20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device having a display, cause the electronic device to perform any of methods 1-18.
21. A generative foundation model trained over millions of health system-scale electronic health records along with web-scale medical text corpora to acquire knowledge of both medical practices and theories, and use of the generative model for rare disease diagnosis (including rare ophthalmic diseases and rare systemic diseases), emergency condition identification (including ophthalmic emergencies and systemic emergencies), complex disease solving (“diagnostic puzzles”), or generating multimodal medical imaging reports (including ophthalmic images and radiology images such as X-rays and CT scans), wherein the generative model involves the use of language data for pre-training, language data for supervised finetuning using a instruction tuning approach (e.g., QA pairs), and a human-machine hybrid evaluation sf-5927165strategy, wherein both the pre-training and supervised finetuning phases involve scaling to extend the context window.
22. The generative foundation model of claim 21, wherein the human-machine hybrid evaluation strategy involves language data for automated evaluations, as well as evaluations by generalists and by different specialists (e.g., ophthalmologists and radiologists) of varying levels of experience. sf-5927165
Citation Information
Patent Citations
Information processing apparatus, information processing method, and storage medium
US20220005584A1
System with report analysis and methods for use therewith
US20220253592A1
Cue-based medical reporting assistance
US20230076903A1
Systems and methods for generating a text report and simulating health care journey
US20240029848A1
Cited By
Clinical decision-making method, system and equipment for respiratory system diseases and medium
CN122224478A