Method and apparatus for processing condition image

By constructing a diagnostic bootstrapping and medical image description dataset, and combining hybrid expert diagnosis and retrieval to enhance the diagnostic mechanism, the collaboration between GFM and specialist models is realized, which solves the problem of insufficient accuracy of generalist models in medical image diagnosis and improves the model's analytical capabilities under various treatment modes.

CN121964074APending Publication Date: 2026-05-01THE HONG KONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE HONG KONG UNIV OF SCI & TECH
Filing Date
2025-07-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing generalist basic models lack accuracy in medical image diagnosis tasks and cannot effectively combine visual and textual information for comprehensive analysis. While specialist models are accurate in specific tasks, their generalization ability is insufficient.

Method used

By constructing diagnostic-guided bootstrapping and medical image description datasets, a generalist base model is trained, and a hybrid expert diagnosis and retrieval-enhanced diagnostic mechanism is combined to achieve collaboration between GFM and the specialist model, utilizing visual and textual information for comprehensive analysis.

Benefits of technology

It improves the accuracy and generalization ability of medical image diagnosis, and can efficiently process medical images in multiple treatment modes, surpassing the performance of traditional GFM and specialist models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121964074A_ABST
    Figure CN121964074A_ABST
Patent Text Reader

Abstract

A computer-implemented method (300) for processing a sample image of a condition is disclosed, the method comprising: processing (305) a sample image by a set of computer vision models for treatment patterns associated with a condition to obtain a set of decisions associated with a plurality of predictive diagnoses for the condition; processing (310) the sample image to determine a relevant visual embedding to be used as a query for querying a database; querying (315) a database to retrieve first k most similar entries from the database based on vector similarities between the determined visual embedding and corresponding keys in the plurality of entries; and processing (320) the sample image by a medically trained talent base model (GFM) using the obtained set of decisions as a reference context and images corresponding to the retrieved first k most similar entries to make a predictive decision on a condition displayed by the sample image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to Provisional Application No. 63 / 714,144, filed with the United States Patent and Trademark Office on October 31, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application generally relates to image processing, and more specifically, to a computer-implemented method for processing sample images of conditions or diseases. Background Technology

[0004] With the rapid development of large-scale language models (LLM) [References 1-3] and visual language pre-training [References 4-6], large-scale visual language models (LVLM) [References 7-9] have demonstrated outstanding performance in a wide variety of tasks (e.g., visual question answering [Reference 10] and image captioning generation [Reference 11]), thus establishing their status as general foundational models (GFM).

[0005] In the medical / healthcare field, GFM [References 12-18] has also demonstrated impressive generalization capabilities across various tasks, such as visual question answering [References 19-20] and radiology report generation [References 21-24]. GFM's superior generalization ability can be attributed to two key aspects. First, the extensive and diverse training corpus endows these models with comprehensive medical knowledge. Furthermore, GFM's strong instruction compliance and contextual learning capabilities enhance its versatility and flexibility, facilitating its application across numerous tasks.

[0006] While GFM appears to exhibit excellent generalization ability, specialist models excel in accuracy-based tasks. These specialist models are tailored for specific downstream tasks, possessing deep domain-specific knowledge, which allows them to focus on a narrower scope and deliver more accurate image analysis results. For example, in medical image diagnostic tasks [References 25-27], specialist models outperform GFM and demonstrate superior performance [References 15, 28-29]. Therefore, GFM is characterized by generalization and flexibility, while specialist models possess expertise and accuracy.

[0007] Therefore, a solution is needed that can address at least one of the aforementioned problems in the prior art and / or provide an option that is useful in the art. Summary of the Invention

[0008] The techniques described herein may relate to a method and apparatus for processing sample images of conditions or diseases.

[0009] According to a first aspect, this application discloses a computer-implemented method for training a generalist foundational model (GFM), the method comprising: generating multiple medical reports related to different conditions using the GFM; providing the medical reports as a first dataset, wherein each medical report is generated based on a relevant image displaying the condition, a classification label containing a textual description of the condition, and a treatment pattern associated with the condition; generating textual descriptions using the GFM for corresponding images displaying different conditions, wherein the textual descriptions describe the condition based on diagnostically relevant visual features displayed in the corresponding images, and wherein the generated textual descriptions and the corresponding images are provided as a second dataset; configuring a training dataset to include the first dataset, the second dataset, a third dataset associated with medical image diagnostics (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA); and training the GFM based on the training dataset to obtain a medically trained GFM.

[0010] Additionally or alternatively, GFM may include one of the following: RadFM, LLaVA-Med, Med-Flamingo, MedDr, and InternVL.

[0011] Additionally or alternatively, each medical report may be configured to include information on medical findings and preliminary diagnoses related to the condition as shown by the relevant images, based on the processing of the GFM.

[0012] Alternatively or additionally, the corresponding images showing different conditions can be obtained from OpenI.

[0013] Additionally or alternatively, treatment modalities may include one of the following: radiology, pathology, dermatology, ophthalmology, gastroenterology, fundus examination, chest X-ray, and endoscopy.

[0014] According to a second aspect, this application discloses a computer-implemented method for processing sample images of a condition, the method comprising: processing the sample images using a set of computer vision models for treatment patterns associated with the condition to obtain a set of determinations associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, wherein each set of computer vision models is pre-trained using downstream datasets specific to various treatment patterns for different conditions; processing the sample images using an embedding model to determine associated visual embeddings, the associated visual embeddings being used as queries for querying a database, wherein the database includes entries of multiple value-key pairs associated with corresponding images of different conditions; wherein the key The method involves: representing a visual embedding of an image, and providing associated values ​​that provide a textual description of the condition displayed in the image; querying a database based on vector similarity between the determined visual embedding and corresponding keys in the plurality of entries using at least one computer vision model from a selected set of computer vision models to retrieve the top k most similar entries from the database, where k is a positive integer; and processing the sample image using a medically trained generalist foundation model (GFM), using an obtained set of judgments as a reference context and images corresponding to the retrieved top k most similar entries, to make a predictive judgment on the condition displayed by the sample image, wherein the medically trained generalist foundation model is trained according to the method described in the first aspect.

[0015] Alternatively, vector similarity calculations can be performed based on cosine similarity.

[0016] According to a third aspect, this application discloses a computing device for training a generalist foundational model (GFM), comprising: one or more memories having executable code; and one or more processors coupled to the one or more memories and configured to execute the code to cause the computing device to perform the following operations: generating multiple medical reports related to different conditions via the GFM, the medical reports being provided as a first dataset, wherein each medical report is generated based on a relevant image displaying the condition, a classification label containing a text description of the condition, and a treatment pattern related to the condition; generating text descriptions via the GFM for corresponding images displaying different conditions, wherein the text descriptions describe the condition based on diagnostically relevant visual features displayed in the corresponding images, and wherein the generated text descriptions and the corresponding images are provided as a second dataset; configuring a training dataset to include the first dataset, the second dataset, a third dataset associated with medical image diagnosis (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA); and training the GFM based on the training dataset to obtain a medically trained GFM.

[0017] Additionally or alternatively, GFM may include one of the following: RadFM, LLaVA-Med, Med-Flamingo, MedDr, and InternVL.

[0018] Additionally or alternatively, the corresponding images can show different conditions obtained from OpenI.

[0019] Additionally or alternatively, the treatment modality may include one of the following: radiology, pathology, dermatology, ophthalmology, gastroenterology, fundus examination, chest X-ray, and endoscopy.

[0020] According to a fourth aspect, this application discloses a computing device for processing sample images of a condition, comprising: one or more memories having executable code; and one or more processors coupled to the one or more memories, configured to execute the code to cause the computing device to perform the following operations: processing the sample images by means of a set of computer vision models for treatment patterns associated with the condition to obtain a set of decisions associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, wherein each set of computer vision models is pre-trained using downstream datasets of various treatment patterns specific to different conditions; and processing the sample images by means of an embedding model to determine associated visual embeddings, which will be used as queries for querying a database, wherein the database includes data related to… The entries are multiple value-key pairs associated with corresponding images of different conditions; wherein the key represents the visual embedding of the image, and the associated value provides a textual description of the condition displayed in the image; using at least one computer vision model from a selected set of computer vision models, a database is queried based on the vector similarity between the determined visual embedding and the corresponding key in the multiple entries to retrieve the top k most similar entries from the database, where k is a positive integer; and the sample images are processed using a medically trained generalist foundation model (GFM), using an obtained set of judgments as a reference context, and images corresponding to the retrieved top k most similar entries to make a predictive judgment on the condition displayed by the sample images, wherein the medically trained generalist foundation model is trained according to the method described in the first aspect.

[0021] Alternatively, vector similarity can be performed based on cosine similarity.

[0022] According to a fifth aspect, this application discloses a non-transitory computer-readable medium containing executable code that, when executed by a processor of a computing device, causes the computing device to perform the method according to the first aspect.

[0023] According to a sixth aspect, this application discloses a non-transitory computer-readable medium containing executable code that, when executed by a processor of a computing device, causes the computing device to perform the method according to any one of the second aspects.

[0024] Other benefits and advantages of the various aspects of this disclosure will be readily apparent from the specification and drawings. These benefits and / or advantages may be obtained individually from the various aspects and features of the specification and drawings, and one or more such benefits and / or advantages need not be provided in their entirety. Attached Figure Description

[0025] In the accompanying drawings, the same reference numerals refer to the same or similarly functional elements in the various views, and the drawings, together with the following detailed description, are incorporated in and form part of the specification to illustrate various aspects and explain the various principles and advantages according to this disclosure.

[0026] Figure 1a This is a schematic diagram of an exemplary framework for training a generalist foundational model (GFM) to obtain a medically trained GFM, based on various aspects of this disclosure.

[0027] Figure 1b This is a schematic diagram illustrating the distribution of the training dataset used to train the General Foundation Model (GFM) according to various aspects of this disclosure.

[0028] Figure 1c This is a schematic diagram of an exemplary framework for a method for processing images of medical condition samples, based on various aspects of this disclosure.

[0029] Figure 1d This is a graph illustrating the performance of a trained GFM according to various aspects of this disclosure.

[0030] Figure 1e This is a graph illustrating the overall performance of a trained GFM on a medical image diagnostic task according to various aspects of this disclosure.

[0031] Figure 2 This is a flowchart illustrating a method for training a generalist foundational model (GFM) based on various aspects of this disclosure.

[0032] Figure 3 This is a flowchart illustrating a method for processing images of disease samples according to various aspects of this disclosure.

[0033] Figures 4a-4t The conventional GFM and its counterparts for processing medical image diagnostic datasets are illustrated according to various aspects of this disclosure. Figure 3 Experimental results between the methods shown.

[0034] Figure 5a This illustrates aspects of conventional GFM and those according to this disclosure. Figure 3 Experimental results of the methods shown on the VQA-RAD, Slake-VQA, Path-VQA, PMC-VQA, and VQA-Med datasets for closed-ended questions in terms of accuracy and labeled F1 score.

[0035] Figure 5b This illustrates aspects of conventional GFM and those according to this disclosure. Figure 3 Experimental results of the generalist-based model on the Omnimed VQA dataset regarding treatment patterns and outcome distribution.

[0036] Figure 5c This illustrates aspects of conventional GFM and those according to this disclosure. Figure 3 Experimental results between the methods on the medical report generation datasets MIMIC-CXR and IU-Xray for ROUGE-L, METEOR, and F1 radiographs (F1-RadGraph).

[0037] Figure 5d An example of generating a medical report on a chest X-ray image is shown in accordance with various aspects of this disclosure.

[0038] Figure 6a This illustrates aspects of conventional GFM and those according to this disclosure. Figure 3 The experimental results show the overall ranking and sum of rankings of different models or mechanisms on medical image diagnostic datasets across eight domains.

[0039] Figure 6b This illustrates aspects of conventional GFM and those according to this disclosure. Figure 3 The experimental results show the overall ranking and sum of rankings of different models or mechanisms on 12 out-of-domain datasets for the methods shown on the medical image diagnostic dataset.

[0040] Figure 6c This demonstrates the best expert (computer vision) model, DenseNet-121, and its various aspects according to this disclosure. Figure 3 Performance comparison between the methods shown.

[0041] Figure 7 (a)-(f) in the figures respectively illustrate examples of hybrid expert diagnosis and retrieval-enhanced diagnosis based on various aspects of this disclosure on a downstream medical image diagnostic dataset.

[0042] Figure 8 Hints for a bootstrapping method for diagnostic guidance are shown in accordance with various aspects of this disclosure, where {Modality} and {Disease} represent placeholders for the corresponding information.

[0043] Figure 9 (a) illustrates the following situation under various aspects of this disclosure: a generalist basic model in a general domain possesses knowledge about the disease, but struggles to correlate this knowledge with the queried image, leading to a misdiagnosis.

[0044] Figure 9 (b) shows the aspects of this disclosure. Figure 2 The method generates detailed medical reports containing findings and preliminary diagnoses, guided by a correct diagnosis.

[0045] Figure 9 (c) shows the various aspects of this disclosure. Figure 2 Examples of medical reports generated using methods across different treatment modalities.

[0046] Figure 10 (a)-(b) show the various aspects of the disclosure. Figure 3 The method employs hybrid expert diagnosis to process disease sample images.

[0047] Figure 11 (a) illustrates the use of an embedding model to extract visual embeddings of images in medical image-text pairs according to various aspects of this disclosure.

[0048] Figure 11 (b) shows that each entry in the database contains metadata and its index embedding, according to various aspects of this disclosure.

[0049] Figure 11 (c) illustrates the use of embedded test images to query a database and retrieve similar samples during the reasoning process, according to various aspects of this disclosure.

[0050] Figure 12 (a)-(e) illustrate examples of hybrid expert diagnosis of five downstream medical image diagnostic datasets in accordance with various aspects of this disclosure.

[0051] Figure 13 (a)-(e) illustrate examples of enhanced diagnostic retrieval on five downstream medical image diagnostic datasets in accordance with various aspects of this disclosure.

[0052] Figure 14 Detailed results are shown on a binary classification dataset for medical image diagnostic tasks.

[0053] Figure 15 Detailed results are shown on a multi-class classification dataset for medical image diagnostic tasks.

[0054] Figure 16 Detailed results are shown on a multi-label classification dataset for medical image diagnostic tasks.

[0055] Figure 17 The performance of the generalist basic model on a domain-specific visual question answering dataset is shown.

[0056] Figure 18 The performance of the generalist basic model on the VQA-Med dataset is shown.

[0057] Figure 19The performance of the generalist basic model on the OmniMedVQA dataset is shown.

[0058] Figure 20 This disclosure illustrates various aspects of the disclosure. Figure 3 Ablation study using the method shown.

[0059] Figure 21 The performance of the generalist basic model on the medical report generation task is demonstrated.

[0060] Figure 22 The performance of the retrieved enhanced diagnostics is demonstrated.

[0061] Figure 23 The document displays detailed information about medical image diagnostic datasets in the field, including dataset name, classification type, dataset size, and label set.

[0062] Figure 24 The details of the out-of-domain medical image diagnostic dataset are shown, including the dataset name, classification type, dataset size, and label set.

[0063] Figure 25 This disclosure illustrates various aspects of the disclosure. Figure 2 The method involves adjusting the prompt template of the dataset for different instructions.

[0064] Figure 26 This disclosure illustrates various aspects of the present disclosure. Figure 3 The method provides instructions. Specifically, {Modality} is a placeholder for the names of different treatment modalities, {Label Set} represents the candidate label set for the classification task, and {RAD} and {MoED} are the results of retrieving enhanced diagnosis and mixed expert diagnosis.

[0065] Figure 27 Specifications for 10 selected computer vision models are shown according to various aspects of this disclosure, covering example parameters such as the number of parameters when using PneumoniaMNIST, gigabytes per second multiplication-addition operations (GMAC), and training time per epoch on the dataset.

[0066] Figure 28 Hyperparameters for training each computer vision model are shown according to various aspects of this disclosure.

[0067] Figure 29 The availability of training datasets according to various aspects of this disclosure is shown.

[0068] Figure 30 The availability of data for the benchmark datasets in accordance with various aspects of this disclosure is shown.

[0069] Figure 31 A list of networks using public code in accordance with various aspects of this disclosure is shown.

[0070] Figure 32 This is a block diagram of a training manager for training a generalist foundational model (GFM) based on various aspects of this disclosure.

[0071] Figure 33 This is a block diagram of a coordination manager for processing disease sample images, based on various aspects of this disclosure.

[0072] Figure 34 It is applicable to the execution of various aspects of this disclosure. Figure 2 Method or Figure 3 A schematic diagram of an exemplary computing device for the method.

[0073] Figure 35 It is applicable to the execution of various aspects of this disclosure. Figure 2 Method or Figure 3 A schematic diagram of an exemplary computing device for the method. Detailed Implementation

[0074] Some parts described below will be presented, explicitly or implicitly, as algorithms and functional or symbolic representations of operations on data in computer memory. These algorithmic descriptions and functional or symbolic representations are means by which those skilled in the art of data processing most effectively communicate their work to others skilled in the art. An algorithm is generally considered as a series of self-consistent steps to achieve a desired result. These steps require physical operations on physical quantities, such as electrical, magnetic, or optical signals that can be stored, transmitted, combined, compared, and otherwise manipulated.

[0075] Unless otherwise expressly stated and as will be apparent from the following, it should be understood that throughout this specification, discussions using terms such as “scan,” “calculate,” “determine,” “replace,” “generate,” “initialize,” “output,” etc., refer to the operations and processes of a computer system or similar electronic device that manipulate and convert data expressed in physical quantities within the computer system into other data expressed in similar physical quantities within the computer system or other information storage, transmission, or display device.

[0076] This specification also discloses apparatus for performing the various operations of the method. Such apparatus may be specifically constructed for the desired purpose, or may include a computer or other device selectively activated or reconfigured by a computer program stored in a computer. The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various machines may be used with the programs taught herein. Alternatively, it is also appropriate to construct more specialized apparatus to perform the required method steps. The architecture of a conventional computer will be shown in the following description.

[0077] Furthermore, this specification implicitly discloses a computer program, and it will be apparent to those skilled in the art that the various steps of the methods described herein can be implemented using computer code. This computer program is not limited to any particular programming language or its implementation. It should be understood that various programming languages ​​and their encodings can be used to implement the teachings disclosed herein. Moreover, this computer program is not limited to any particular control flow. Many other variations of this computer program exist, which may use different control flows without departing from the spirit or scope of the invention.

[0078] Furthermore, one or more steps of the computer program can be executed in parallel rather than sequentially. Such a computer program can be stored on any computer-readable medium. This computer-readable medium may include storage devices such as disks or optical discs, memory chips, or other storage devices suitable for interfacing with a computer. The computer-readable medium may also include hardwired media (e.g., exemplified in Internet systems) or wireless media (e.g., exemplified in GSM, GPRS, 3G, 4G, 5G, NR mobile communication systems and other wireless communication systems / standards (e.g., Bluetooth, ZigBee, or Wi-Fi)). When the computer program is loaded and executed on such a computer, it effectively creates a device for implementing various aspects of this disclosure.

[0079] Various aspects of this disclosure can be implemented as hardware modules. More specifically, in a hardware sense, a module is a functional hardware unit designed for use with other components or modules. For example, a module can be implemented using discrete electronic components, or it can be part of an entire electronic circuit, such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). Many other possibilities are known in the art. Those skilled in the art will also understand that the system can be implemented as a combination of hardware and software modules.

[0080] This disclosure provides methods and corresponding apparatus for training a generalist foundational model (GFM), as well as methods and corresponding apparatus for processing sample images of (different) condition data, the latter being performed at least in part based on a trained GFM. For the avoidance of doubt, a condition can generally refer to a medical illness, any specific health problem or symptom, which can be diagnosed by a healthcare provider based on symptoms, medication use, or diagnostic tests. The disclosed subject matter will be set forth below in accordance with various aspects of this disclosure.

[0081] For the sake of completeness, it should be clarified that any reference to the format definition “[Reference X]” in any paragraph of this manual should be interpreted as referring to the corresponding reference “X” in the “References” section of this manual. For example, [Reference 10] refers to the reference listed in the “References” section of this manual

[10] , while [References 21-24] accordingly refer to references

[21] -

[24] with the necessary modifications.

[0082] Specifically, according to various aspects of this disclosure, the method for processing sample images is based on an exemplary collaborative framework called Generalist-Specialist Collaboration (hereinafter referred to as "GSCo") to explore the synergy between GFM and Specialist (computer vision) models. Figure 1a This is a schematic diagram of a framework for training a GFM to obtain a medically trained GFM, based on various aspects of this disclosure. The GFM may be based on GFMs known in the art (e.g., see Reference 9), but should not be construed as limiting. The GSCo framework may include two phases: the construction (i.e., training) of the GFM and the specialist model, and subsequent collaborative inference for downstream tasks.

[0083] During the construction phase, a medically trained GFM (hereinafter referred to as "MedDr") can be developed based on a large-scale medical image-text pair training corpus spanning multiple modalities. Simultaneously, a series of lightweight specialist models can be selected and customized for specific downstream tasks to significantly reduce computational costs. It is understood that downstream tasks may include medical image analysis tasks, such as pneumonia detection, lesion classification, or similar tasks.

[0084] In the collaborative reasoning phase, in accordance with various aspects of this disclosure, this paper discloses two mechanisms, namely Hybrid Expert Diagnosis (hereinafter referred to as "MoED") and Retrieval Enhancement Diagnosis (hereinafter referred to as "RAD"), to achieve collaborative cooperation between MedDr and the expert model.

[0085] Especially during the construction phase, the focus is on developing advanced medical GFM. To manage a large-scale multimodal training corpus, two datasets are introduced: the first is called the Diagnostic Guided Bootstrapping (DGB) dataset, and the second is called the Medical Image Description (DES) dataset. The DGB dataset can be constructed based on rich medical image diagnostic data [References 30–35], aiming to enhance the intrinsic disease diagnostic capabilities of GFM. Unlike traditional methods that rely solely on textual information from image-text pairs to generate instruction-adjusted data [References 12, 14], the DGB dataset is generated by integrating both visual and textual information. GFM can be used in general domains to generate detailed medical reports, including findings and conclusions, based on diagnostic information such as classification labels. Guided by human-verified annotations, the generated data is not only more reliable but also significantly enriches the depth and diversity of visual information conveyed in the text.

[0086] Furthermore, the DES dataset can be broadened by integrating image-based case studies from OpenI [Reference 36] corresponding to various conditions. GFM [Reference 9] can also be used to rewrite the text descriptions of case studies and remove any information that may originate from the corresponding images. In addition to the DGB and DES datasets, GFM can also be trained by integrating Medical Image Diagnosis (CLS) [Reference 31], Medical Report Generation (MRG) [Reference 21], and Visual Question Answering (VQA) [Reference 37] datasets. In total, the training corpus may contain more than 2 million samples across five different types of training datasets, covering a wide range of treatment patterns. Figure 1b The distribution of the training corpus according to this disclosure and two samples from the DGB and DES datasets, respectively, is shown.

[0087] MedDr can be developed based on the training corpus. It is considered the largest open-source general medical model, containing 40 billion parameters. Compared with other traditional general medical models [References 12-14], MedDr demonstrates superior capabilities in medical image analysis, covering a more diverse range of treatment modalities, including radiology, pathology, dermatology, ophthalmology, and gastroenterology.

[0088] Furthermore, MedDr demonstrates advanced contextual learning and instruction-following capabilities and facilitates collaboration between trained generalist foundational models (GFM) and specialist models.

[0089] In the collaborative reasoning phase, the synergistic relationship between MedDr and the specialist model is described below. To fully leverage the contextual learning capabilities of generalists and the domain expertise of specialists, Hybrid Expert Diagnosis (MoED) and Retrieval-Enhanced Diagnosis (RAD) are proposed as collaborative mechanisms between the trained GFM and the specialist model. MoED diagnosis aims to supplement and enhance GFM by utilizing the prediction results of the specialist model. The MoED paradigm [Reference 38] integrates the insights of multiple experts, which has been previously explored in the field to enhance the robustness of predictions [References 39–41]. In MoED, the output of the specialist model serves as the reference context and is then provided to MedDr along with the test image. MedDr then needs to provide a predictive diagnosis by considering both the content of the test image and the results of the specialist model. Unlike MoED, which utilizes the inherent expertise of the specialist model, Retrieval-Enhanced Diagnosis is also proposed to fully leverage the extensive medical knowledge contained in the existing data.

[0090] Unlike previous generalist medical models that relied solely on internal model knowledge for diagnosis [References 12-15], the proposed GSCO framework introduces Retrieval Enhanced Generation (RAG) [Reference 42] to leverage external knowledge, thereby improving the model's accuracy and reliability. In the proposed RAD mechanism, each specialist model can act as a retrieval unit, using the visual embedding of the test image as a query to retrieve the "K" most similar samples from the database, where K is a positive integer. In various examples, the best specialist model (among all those selected for RAD) can act as the retrieval unit. Information extracted from the retrieved samples is then provided to MedDr as a contextual reference to assist in medical image analysis. Thus, MoED and RAD can jointly provide beneficial guidance to MedDr based on the specialists' expertise. Simultaneously, as a decision-maker, MedDr combines its internal and external knowledge to arrive at a final diagnosis.

[0091] To comprehensively evaluate MedDr and the proposed GSCo framework, this application governs a large-scale benchmark set. Compared to traditional literature [References 12-15, 43], the governed benchmark set demonstrates superior diversity and scale. Figure 1cAs shown, this benchmark set covers various medical datasets, including medical image diagnosis, visual question answering, and medical report generation tasks. Specifically, the benchmark set includes 14 in-domain datasets and 14 out-of-domain datasets, totaling 250,000 samples. To evaluate the improvements of GSCO over GFM on medical image diagnosis datasets, where specialists may achieve better results due to their specific domain knowledge, this application integrates 20 different medical image diagnosis datasets covering 11 treatment modalities. Experiments were conducted on the standard GFM dataset for comparison to demonstrate the superiority of MedDr.

[0092] Figure 1d The comparison results of GFM on 10 benchmark datasets (covering different medical tasks and modalities) are shown. It can be seen that MedDr consistently outperforms other traditional GFMs by a significant margin in both the medical and general domains. Furthermore, experiments were conducted to validate the effectiveness of GSCO. Figure 1e The performance of MedDr on medical image diagnostic tasks is demonstrated. While specialist models are competitive due to their domain-specific knowledge, the proposed GSCo framework further enhances MedDr's performance, outperforming all specialist models. These experimental results highlight the importance of the proposed GSCo framework, representing a paradigm shift in the clinical application of GFM. This shift moves away from using individual models independently to complete medical tasks and instead promotes collaboration between GFM and specialist models.

[0093] The GSCo framework offers two key advantages: First, GSCo is efficient. Compared to using GFM or specialist models independently, GSCo demonstrates superior performance, particularly on out-of-domain datasets, showcasing its advanced generalization capabilities. Second, GSCo is highly efficient. When faced with out-of-domain tasks or data, it eliminates the need for extensive resource allocation to fine-tuning GFM, instead allowing for efficient tuning of lightweight specialist models with minimal power consumption, demonstrating its scalability and sustainability.

[0094] The contributions of the framework proposed in this application are described below in accordance with various aspects of this disclosure.

[0095] To the best of the inventors' knowledge, the Generalist-Specialist Collaboration (GSCo) proposed in this application is the first collaborative framework for exploring the synergy between GFM and specialist models. GSCo combines the contextual learning capabilities of generalists with the domain expertise of specialists, enabling precise medical image analysis for a wide range of medical tasks. This collaborative paradigm not only expands the functionality of GFM and achieves efficient resource utilization but also ensures scalability and sustainability, thereby advancing the development of general AI in the medical field.

[0096] This article also introduces MedDr, currently the largest open-source medical generalist foundational model. During MedDr's development, Diagnostic Guided Bootstrapping (DGB) and Medical Image Description (DES) were introduced to enhance the diversity of the training corpus. Therefore, MedDr can handle various treatment modalities and tasks, achieving state-of-the-art performance in downstream tasks and surpassing other traditional GFMs. Notably, downstream tasks can include medical image analysis tasks such as pneumonia detection, lesion classification, or similar tasks. Furthermore, MedDr excels in instruction following and contextual learning, laying a better foundation for collaboration with specialist models.

[0097] To facilitate collaboration between MedDr and the expert model, this paper proposes two collaborative mechanisms: Hybrid Expert Diagnosis (MoED) and Retrieval Enhanced Diagnosis (RAD). MoED incorporates expert diagnoses as guidance, while RAD utilizes expert information to retrieve the most similar cases for reference. The results of MoED and RAD are combined as contextual information to guide MedDr.

[0098] This paper also establishes a large-scale benchmark set containing 28 datasets with 250,000 test samples, covering more than ten treatment modalities across various medical tasks. Extensive experiments were conducted on this benchmark set, demonstrating MedDr's superior capabilities and validating the effectiveness of the proposed GSCo framework.

[0099] Figure 2 This is a flowchart illustrating a method 200 for training a generalist foundational model (GFM) according to various aspects of this disclosure. As an example, the GFM used herein may be based on a model known in the art [e.g., Reference 9]. Method 200 corresponds to the construction (i.e., training) phase of the GFM and specialist model within the GSCO framework. Method 200 can be implemented as a computer-implemented method. Operation of method 200 can be performed on computer devices 3400, 3500 (e.g., ...). Figures 34-35 The computer devices 3400 and 3500 (as shown) or their components are implemented and executed thereon. In some examples, the computer devices 3400 and 3500 may execute a set of instructions to control the functional elements of the computer devices 3400 and 3500 to perform the functions described below. Additionally or alternatively, the computer devices 3400 and 3500 may use dedicated hardware to perform the functions described below.

[0100] At 205, method 200 may include: generating multiple medical reports related to different conditions via GFM, and providing the medical reports as a first dataset, wherein each medical report is generated based on a relevant image showing the condition, a classification label containing a textual description of the condition, and a treatment pattern related to the condition.

[0101] At 210, method 200 may include: generating text descriptions for corresponding images displaying different conditions via GFM, wherein the text descriptions describe the conditions based on diagnostic-related visual features displayed in the corresponding images, and wherein the generated text descriptions and corresponding images are provided as a second dataset.

[0102] At 215, method 200 may include: configuring the training dataset to include a first dataset, a second dataset, a third dataset associated with medical image diagnosis (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA).

[0103] At 220, method 200 may include: training a GFM based on a training dataset to obtain a medically trained GFM.

[0104] Figure 3 This is a flowchart illustrating a method 300 for processing patient sample images according to various aspects of this disclosure. Method 300 corresponds to the inference phase within the GSCo framework. Method 300 can be implemented as a computer-implemented method. The operation of method 300 can be performed as follows: Figures 34-35 The computer devices 3400, 3500 or their components are implemented and performed thereon. In some examples, the computer devices 3400, 3500 may execute a set of instructions to control the functional elements of the computer devices 3400, 3500 to perform the functions described below. Additionally or alternatively, the computer devices 3400, 3500 may use dedicated hardware to perform aspects of the functions described below.

[0105] At 305, method 300 may include processing sample images using a set of computer vision models for condition-related treatment modalities to obtain a set of decisions related to multiple predicted diagnoses of the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, each set of computer vision models being pre-trained using downstream datasets specific to different treatment modalities for different conditions. It should be clarified that the multiple different sets of computer vision models refer to different specialized models within the GSCo framework.

[0106] In some examples, method 300 may alternatively select a computer vision model set from multiple different computer vision model sets before performing step 305.

[0107] At 310, method 300 may include processing the sample image by means of an embedding model to determine an associated visual embedding to be used as a query for querying a database, wherein the database includes entries of multiple value-key pairs associated with corresponding images of different conditions; wherein the key represents the visual embedding of the image, and the associated value provides a textual description of the condition displayed in the image.

[0108] At 315, method 300 may include: querying a database based on vector similarity between a determined visual embedding and a corresponding key in a plurality of entries, using at least one computer vision model in a selected set of computer vision models, to retrieve the top k most similar entries from the database, where k is a positive integer.

[0109] At 320, method 300 may include: processing sample images using a medically trained generalist foundational model (GFM), combined with a decision set associated with multiple predicted diagnoses as a reference context, and images corresponding to the top k most similar entries retrieved, to make a predictive judgment on the condition shown by the sample images. The medically trained GFM is based on a reference context... Figure 2 The method discussed was used for training 200. It should be noted that medically trained GFM refers to the MedDr proposed in this paper.

[0110] ·result

[0111] Medical Comprehensive Assessment Criteria

[0112] In this study, to comprehensively evaluate the model's performance in medical tasks, a large-scale benchmark was established, encompassing 28 public datasets containing approximately 250,000 samples. Compared to previous studies [References 12-15, 43], this benchmark excels in both diversity and scale. The datasets in the benchmark were carefully selected to include clinically relevant tasks such as medical image diagnosis, visual question answering, and medical report generation. Furthermore, the benchmark was configured to cover a variety of medical image modalities, including radiology, pathology, dermatology, ophthalmology, gastroenterology, etc., to ensure a comprehensive evaluation of the model's capabilities across diverse conditions.

[0113] Medical image diagnosis

[0114] Medical image diagnosis is one of the most fundamental tasks in the medical field, requiring models to diagnose query images within a predefined set of labels. This benchmark integrates 20 different medical image diagnosis datasets, covering approximately 100,000 test samples and 11 treatment modalities. This benchmark has two prominent features. First, it covers a wide range of medical conditions. For example, VinDr-SpineXR [Reference 25] and FMC-Chest [Reference 27] are challenging radiology datasets, focusing on spine and chest X-ray images, respectively. HAM10000 [Reference 31] and DermNet [Reference 44] are both dermatology-related datasets, the former focusing more on dermoscopy images and the latter more on clinical images. RetOCT [Reference 45] and BRSET [Reference 46] focus on ocular diseases, although involving different treatment modalities. Second, the benchmark includes different classification types. For example, PneumoniaMNIST [Reference 26] and BreastMNIST [Reference 26] are binary classification datasets. While providing a usable answer (i.e., positive or negative) for binary classification tasks is straightforward for GFM, achieving accurate predictions is indeed challenging. FMC-Endo [Reference 27] and OCTMNIST [Reference 26] are multi-class classification datasets with label sets consisting of 5 and 4 labels, respectively.

[0115] For each given image, the model only needs to predict one label from a predefined label set. The most challenging task is multi-label classification, where each test image may have multiple labels, or even no labels. For example, ChestMNIST [Reference 33] is a multi-label classification dataset containing 15 classes, where each sample may contain one or more positive labels. Compared to binary and multi-class classification tasks, such tasks place higher demands on the model's disease diagnosis and task understanding capabilities. For ease of comparative analysis, these medical image diagnosis datasets are divided into two groups based on whether their training sets are included in MedDr's training corpus: in-domain datasets and out-of-domain datasets. Figure 23-24 The data tables within provide further details about these datasets. Accuracy and macro-F1 scores were used as evaluation metrics for these medical image diagnostic tasks. For binary classification datasets, results are presented as accuracy. For multi-class and multi-label classification datasets, due to the imbalanced distribution of samples across different classes, results are reported as macro-F1 scores. Further details regarding the metrics used will be discussed in the following sections.

[0116] Visual Q&A

[0117] Visual question answering (VQA) tasks require models to answer questions based on given images, which necessitates a thorough understanding of both visual and textual information. Experiments were conducted on six different VQA datasets, including four in-domain datasets: VQA-RAD [Reference 37], Slake-VQA [Reference 47], Path-VQA [Reference 19], and PMC-VQA [Reference 48], and two out-of-domain datasets: VQA-Med [Reference 20] and OmniMedVQA [Reference 28]. The inherent diversity of these benchmark datasets enabled this application to extract valuable insights from various domains and facilitated a comprehensive evaluation of model performance. For example, the VQA-RAD, Slake-VQA, and VQA-Med datasets primarily focus on radiological data, such as CT, MRI, and X-ray images, while Path-VQA primarily focuses on pathological data. In contrast, the PMC-VQA and OmniMedVQA datasets are larger in scale and cover a wider range of treatment modalities. PMC-VQA is built on PubMed, while OmniMedVQA is derived from a wide range of medical datasets. Notably, except for OmniMedVQA, which includes multiple-choice questions, the other datasets contain free-form questions. In evaluating OmniMedVQA, this application uses accuracy as the metric. For other datasets, this application follows MultiMedEval [Reference 49], using both natural language generation (NLG) and classification metrics to evaluate the results.

[0118] Medical report generation

[0119] Medical report generation (MRG) involves the model's ability to enumerate all observations and make diagnoses based on the analysis of medical images. This task presents a significant challenge because the model needs to accurately capture the inherent complexity in the medical images being evaluated. To evaluate the performance of MRG, this application conducted experiments on two benchmark datasets, MIMIC-CXR [Reference 21] and IU-Xray [Reference 22]. Both datasets focus on chest X-rays and provide detailed medical reports summarizing the patient's condition. To evaluate the model, this application used both the NLG metric and the model-based metric from MultiMedEval [Reference 49]. Further details regarding the metrics used in the VQA and MRG tasks will be outlined in the following sections.

[0120] MedDr demonstrates superior performance in medical image diagnostics.

[0121] First, the performance of GFM in medical image diagnosis tasks was discussed. It should be noted that some GFMs failed to generate appropriate responses for evaluation on specific datasets and were therefore not included in the comparison. If MedDr's performance is best compared to other GFMs, the p-value is shown. Accordingly, the experimental results are shown in Figure 4. Specifically, Figures 4a-4h The results within the domain were described, while Figure 4i-4t Out-of-domain results are depicted. For binary classification datasets (e.g., PCam200), results are reported as accuracy. For multi-class datasets (e.g., DermNet) and multi-label datasets (e.g., ChestMNIST), results are presented as macro F1 scores.

[0122] Overall, MedDr significantly outperforms other GFMs, demonstrating its advantages in three aspects. First, MedDr excels at following instructions. Notably, RadFM, LLaVA-Med, and Med-Flamingo sometimes struggle to generate appropriate outputs for tasks involving multiple labels. For example, in multi-class classification tasks with only one label, RadFM, LLaVA-Med, and Med-Flamingo might output the entire label set or simply restate the instructions, indicating poor instruction following ability. In contrast, MedDr effectively follows instructions across various tasks, consistently generating coherent and context-appropriate output.

[0123] Secondly, MedDr can handle a wider range of treatment modalities. RadFM [Reference 14] is a GFM focused on radiology data, and therefore achieves commendable performance on radiology datasets such as VinDr-SpineXR [Reference 25] and VinDr-PCXR [Reference 30]. However, RadFM's performance degrades when applied to other treatment modalities, indicating its limited generalization ability. In contrast, MedDr can handle multiple treatment modalities, including radiology, pathology, dermatology, ophthalmology, gastroenterology, and more.

[0124] Third, MedDr performs well in medical image diagnosis. LLaVA-Med [Reference 12] and Med-Flamingo [Reference 13] are primarily trained on visual question-answering datasets. The overall performance of LLaVA-Med and Med-Flamingo in medical image diagnosis tasks is still far from satisfactory. Meanwhile, InternVL [Reference 9] may follow instructions well, but its performance still lags far behind MedDr due to its limited inherent medical knowledge. For example, on the FMC-Endo [Reference 27] dataset, MedDr achieves a macro F1 score of 32.0%, significantly outperforming InternVL (11.7% macro F1 score, P<0.001, see [reference 27]). Figure 4r ).

[0125] It is noteworthy that in some binary classification datasets, such as CBIS-DDSM(CALC) [Reference 50] and CBIS-DDSM(MASS) [Reference 50], LLaVA-Med, Med-Flamingo, and InternVL struggle to distinguish between the two classes, consistently assigning the same label to all samples. While these models may achieve higher accuracy, their F1 scores are 0. This lack of discriminative power suggests that these models may fail to capture the relevant features required for effective classification, thus compromising the reliability of their results. In contrast, MedDr performs exceptionally well across various treatment modalities and classification tasks, highlighting its state-of-the-art capabilities in medical image analysis.

[0126] MedDr excels in visual question answering and medical report generation.

[0127] This section will discuss the evaluation of GFM in visual question answering and medical report generation tasks. Results of the VQA task are shown... Figures 5a-5b In the middle, and the data table is as follows Figure 17-18 As shown.

[0128] Overall, MedDr outperformed other GFM datasets across all datasets. Specifically, on the VQA-RAD dataset, MedDr achieved a BLEU-1 score of 59.62% (P = 0.062) and an F1 score of 61.10% (P = 0.054), significantly outperforming the finely tuned LLaVA-Med. On the Slake-VQA dataset, MedDr achieved 83.38% accuracy and 77.26% recall on closed-ended questions, substantially surpassing RadFM (P < 0.001). Figure 5a On the most challenging dataset PMC-VQA, MedDr achieved a recall of 27.30% and an accuracy of 14.94% on open-ended questions, significantly outperforming other models (P<0.001, see [link]). Figure 17 (Data table). On the OmniMedVQA dataset, such as Figure 5b and Figure 18 As shown in the data table, MedDr also performs well across all treatment modalities, achieving an overall accuracy of 63.0%. MedDr's overall performance in the visual question answering task demonstrates that the model proposed in this application is not only adept at understanding both visual and textual information, but also capable of handling a wider variety of treatment modalities.

[0129] The results of the MRG task show that... Figure 5c and Figure 21 The data is presented in the table. It should be noted that, apart from RadFM [Reference 14] and MedDr, other models (such as LLaVA-Med [Reference 12], Med-Flamingo [Reference 13], and InternVL [Reference 9]) have not yet been trained on MRG task datasets. MedDr outperforms RadFM [Reference 14] (overall, P < 0.001), and this model, primarily focused on radiology tasks, performs well on almost all evaluation metrics across both benchmark datasets. For example, MedDr achieves a ROUGE-L score of 22.59 on the MIMIC-CXR dataset and 28.35 on the IU-Xray dataset, highlighting its ability to comprehensively interpret medical images.

[0130] Figure 5d This is an example of a medical report generation task based on chest X-ray images. MedDr was compared to RadFM [Reference 14] on the benchmark dataset MIMIC-CXR [Reference 21], a GFM specifically designed for radiology. Most findings generated by RadFM only describe the patient's normal (healthy) condition, while abnormal information is more critical for the medical report. In contrast, MedDr lists both normal findings (e.g., "no evidence of pulmonary edema" and "no pneumothorax or large pleural effusion") and abnormal findings (e.g., "increased shadows at the right lung base, which may indicate atelectasis or early pneumonia").

[0131] GSCo facilitates the generalization of precision disease diagnosis and medical AI.

[0132] This section presents experiments on medical image diagnostic datasets to demonstrate the effectiveness of the proposed GSCO framework. Based on previous research [Reference 51], ten representative computer vision models (e.g., ResNet [Reference 52] and ViT [Reference 53]) were selected as base (specialist) models, which have significantly fewer parameters compared to GFM. Details regarding the selected models are provided below. These base models were fine-tuned on each of 20 medical image diagnostic datasets, resulting in a total of 200 specialist models. Furthermore, for fairness, a benchmarking method called "Voting" was introduced, which aggregates and votes on the specialist predictions, effectively leveraging the naive collaborative approach to utilize the results of the specialist models.

[0133] Quantitative analysis

[0134] Figures 6a-6c The ranking of various methods on in-domain and out-of-domain datasets is shown separately. A summary of their rankings on different tasks is also presented. First, this application observes that MedDr outperforms other GFM models on in-domain datasets, even surpassing some specialist models. Compared to specialist models, GFM's overall performance is lower, with most GFM models ranking relatively poorly. This observation suggests that while GFM demonstrates superior generalization ability by using a unified model to perform various tasks, specialist models achieve higher accuracy on specific datasets due to their domain-specific fine-tuning. Notably, MedDr outperforms several specialist models on in-domain datasets. For example, on the ChestMNIST and PCam200 datasets, MedDr surpasses most specialist models, ranking fourth and fifth respectively, highlighting its superior intrinsic diagnostic capabilities.

[0135] Secondly, this application finds that GSCo achieves the highest overall performance, significantly outperforming other methods. As a direct collaborative approach, "voting" outperforms specialist models on most datasets, highlighting its effectiveness. However, while "voting" performs well on binary and multi-class classification datasets, it performs poorly on multi-label datasets. This discrepancy suggests that it may struggle to effectively utilize the predictions of specialist models in multi-label classification tasks. This is because, as a simple approach, "voting" relies solely on the outputs of specialist models and treats their suggestions equally. Consequently, when faced with highly challenging multi-label classification tasks, "voting" may arrive at incorrect diagnoses if the majority of specialist model predictions are incorrect. In contrast, the GSCo framework proposed in this application not only considers the predictions of specialist models but also leverages the inherent knowledge of MedDr, thus achieving superior performance and robustness.

[0136] Figure 6c The overall best-performing specialist model, DenseNet-121, was compared with the GSCo framework proposed in this application. For clarity, the results are divided into three groups based on their classification task type. Compared with the best specialist model, GSCo demonstrates a significant performance advantage even on out-of-domain datasets, highlighting its superiority and generalization ability.

[0137] Therefore, the GSCo framework proposed in this application demonstrates a synergistic relationship between GFM and specialist models. Specialist models, leveraging their domain-specific knowledge, guide GFM, significantly improving its performance, especially on out-of-domain datasets. Furthermore, fine-tuning specialist models on specific downstream datasets is computationally efficient and can be implemented with minimal additional resources. On the other hand, GFM acts as a decision-maker possessing rich intrinsic medical knowledge. Unlike "voting" methods that rely solely on the output of specialist models, MedDr integrates guidance from these specialists while retaining its diagnostic capabilities. This collaborative strategy effectively improves performance across a wide range of medical tasks.

[0138] Qualitative analysis

[0139] Experimental results for Hybrid Expert Diagnosis (MoED) and Search Enhanced Diagnosis (RAD) are presented in a visual format. Figure 7 (a)-(c) and Figure 12 (a)-(e) show example cases of MoED on downstream datasets of various treatment modalities. First, this application finds that aggregating predictions from multiple specialist models can provide accurate and robust guidance. In medical image diagnosis, the key to accurate disease identification often lies in the subtle differences present in the images. Due to this inherent difficulty, among all specialist models, only EfficientNet-B4 [Reference 54] consistently produces correct diagnostic results in all example cases. Meanwhile, aggregating predictions from multiple specialist models (e.g., “voting”) yields even more accurate results. For example, on the RetOCT dataset ( Figure 7 (a) and CBIS-DDSM(CALC) dataset ( Figure 7 In (b)), most specialist models made correct predictions, thus providing effective guidance for MedDr to arrive at accurate diagnoses. Secondly, it is noteworthy that MedDr can still arrive at correct diagnoses even when the aggregation results fail to provide the correct reference. For example, in the FMC-Colon dataset ( Figure 7In (c)), half of the specialist models predicted a "positive" result, while the other half gave the opposite result. In this predicament, a simple "voting" strategy failed, and MedDr correctly arrived at a "negative" result. These observations validate the superiority of MedDr and the effectiveness of the MoED proposed in this application. MedDr not only utilizes the reference diagnosis provided by the specialist models but also leverages its inherent knowledge to make the final decision. MoED represents an effective collaboration that integrates the strengths of both generalist and specialist models, thereby improving diagnostic accuracy.

[0140] Figure 7 (d)-(f) and Figure 13 Images (a)-(e) illustrate examples of RAD on downstream datasets covering a wide range of modalities. First, this application discusses the effectiveness of this retrieval strategy. As... Figure 7 (d) and Figure 13 As shown in (e), the retrieved images share the same labels as the query images and can therefore serve as reliable references. It is noteworthy that although the expert may make incorrect diagnoses as predictors, the majority of the retrieved images still provide accurate diagnostic information, demonstrating the expert's robustness as a retriever. For example, in Figure 7 In (e), the specialist model predicted "malignant," while most retrieved samples were "normal or benign." These observations demonstrate that, in most cases, the retrieved samples provide accurate guidance for MedDr, thus validating the accuracy and robustness of the retrieval strategy. Secondly, this application discusses the effectiveness of RAD. Figure 7 (d)-(f) and Figure 13 In most cases (a)-(e), the retrieved samples provided correct guidance. Furthermore, this application observes that even when the expert model's predictions and retrieved items contain interfering information, MedDr is still able to make accurate diagnoses based on its inherent disease diagnostic capabilities. For example, as... Figure 7 As shown in (f), although most of the retrieved images were “pneumonia” and thus provided incorrect guidance, MedDr successfully arrived at a “normal” diagnosis. These findings highlight that MedDr not only cleverly utilizes the external knowledge provided by the retrieved samples but also leverages its inherent diagnostic capabilities, thereby demonstrating the effectiveness of the RAD method proposed in this application.

[0141] Ablation Research

[0142] This section discusses ablation studies using the MoED and RAD mechanisms proposed in this application. Experiments were conducted on nine medical image diagnostic datasets covering various treatment modalities and tasks. Figure 20The experimental results are shown. For binary classification, the accuracy is reported. Otherwise, the macro F1 score is reported. Both MoED and RAD showed consistent performance improvements over MedDr, with average improvements of 0.2030 and 0.2065, respectively. RAD showed superior performance compared to MoED, which can be attributed to its use of not only domain-specific knowledge from specialists but also information from the training database, thus providing a more reliable reference for MedDr.

[0143] To further evaluate the effectiveness and generalization ability of RAD, experiments were conducted on a wider range of downstream tasks. Developing a specialist model that consistently outperforms GFM in visual question answering and medical report generation tasks is particularly challenging. If specialist models perform poorly, they may fail to provide reliable guidance for GFM, potentially leading to a decline in overall performance. Therefore, in the following experiments, MedDr's visual encoder was used as the retrieval engine. For the medical image diagnosis task, the labels of the five most similar cases were retrieved, as is the usual practice. In the visual question answering and medical report generation tasks, the most similar images were retrieved, and their corresponding annotations were then incorporated into the input. In addition, Med-Flamingo [Reference 13] was also introduced as a baseline model, which demonstrated impressive few-show learning capabilities.

[0144] Figure 22 The data fields depict the results of medical image diagnosis, visual question answering, and medical report generation tasks. The "Vote" column reflects the results obtained through a voting mechanism based on the retrieved sample labels. For medical report generation, the top-ranked retrieved report is used as the "Vote" result. Notably, both Med-Flamingo and MedDr show significant and consistent improvements across most datasets and tasks, highlighting the versatility of RAD. However, Med-Flamingo's performance is limited by the "Vote" upper bound, indicating that it heavily relies on the retrieved results in the final diagnosis without adequately considering the image content. For example, in the medical report generation task, Med-Flamingo tends to directly rewrite or even copy the retrieved samples with only minor modifications. In contrast, MedDr consistently outperforms other methods in the "Vote" results on most downstream datasets, suggesting that MedDr not only considers the retrieved results but also leverages its inherent knowledge to diagnose test images. These experiments validate the effectiveness of RAD, demonstrating that it can enhance GFM's capabilities even without specialized expertise, thereby expanding its application scenarios.

[0145] ·discuss

[0146] The GSCo framework proposed in this application is the first research result to investigate the synergy between GFM and specialist models. The GSCo framework aims to fully leverage the contextual learning capabilities of generalists and the domain-specific knowledge of specialists. GSCo comprises two phases: the construction of GFM and specialists, and collaborative reasoning for downstream tasks. In the construction phase, MedDr, the largest open-source GFM tailored for medicine, was developed, enabling it to handle various medical tasks and patterns. MedDr also demonstrates superior capabilities in both instruction compliance and contextual learning, providing a solid foundation for synergy with specialist models. Simultaneously, a series of lightweight specialists (models) are customized for specific downstream tasks, resulting in lower computational overhead. In the collaborative reasoning phase, Hybrid Expert Diagnosis (MoED) and Retrieval Enhanced Diagnosis (RAD) are proposed as the core mechanisms for collaboration. MoED integrates predictions from specialists as reference diagnoses, while RAD utilizes these specialists to retrieve similar cases, jointly providing contextual information to MedDr to facilitate medical image analysis. To evaluate MedDr and GSCo, this application manages the largest benchmark in medical GFM, comprising 28 datasets and approximately 250,000 test samples, covering a wide range of treatment modalities and tasks. Extensive qualitative and quantitative experiments highlight several aspects of the research in this application.

[0147] This application finds that MedDr excels in understanding and analysis. Compared to traditional models [References 9, 12-15], MedDr demonstrates significant advantages in two key areas. First, as a generalist foundational model, MedDr exhibits superior generalization ability in medical image analysis, enabling it to handle a wider range of treatment modalities, including radiology, pathology, dermatology, ophthalmology, and gastroenterology, and achieving state-of-the-art performance in multiple tasks such as visual question answering, medical report generation, and medical image diagnosis. Second, MedDr demonstrates superior instruction following and contextual learning capabilities. This enhances the model's flexibility in handling different tasks and leveraging external knowledge, laying a solid foundation for effective collaboration with specialist models. These advantages are attributed to a well-managed training corpus and a larger MedDr scale. Specifically, the training corpus covers over 2 million samples across five different task types, thus expanding MedDr's application scope. Furthermore, MedDr boasts 40 billion parameters, surpassing previous models and endowing it with superior intrinsic capabilities. In the future, this application plans to integrate a wider variety of training corpora and a larger base model.

[0148] This application also finds that instruction compliance and contextual learning enable collaboration between generalist and specialist models. In most traditional models [References 9, 12-15, 26, 43], both GFM and specialist models independently utilize their inherent capabilities to handle downstream tasks. GFM is known for its generalization ability and flexibility, which allows them to perform a variety of medical tasks across different modalities using a single model. In contrast, specialist models are known for their accuracy and efficiency, achieving satisfactory performance by tailoring them to specific downstream tasks with lower computational costs. The GSCo framework proposed in this application aims to explore collaboration between generalist and specialist models, where instruction compliance and contextual learning (which have been neglected by previous methods) are crucial for facilitating their integration. Experimental results show that MedDr can effectively collaborate with specialist models and achieve state-of-the-art (SOTA) performance in downstream tasks thanks to its superior instruction compliance and contextual learning capabilities.

[0149] The GSCo framework provides a new paradigm for the clinical application of GFM and specialist models. Collaboration among healthcare professionals is not only common but also crucial in clinical practice. The GSCo framework proposed in this application demonstrates that superior performance on downstream tasks can be achieved through synergistic cooperation between GFM and specialist models. GSCo provides a new paradigm for the clinical application of GFM and specialist models. Specifically, when faced with tasks or data outside the domain, it is more efficient to tune lightweight specialist models with minimal resource consumption than to invest significant resources in fine-tuning GFM. Furthermore, since most medical data (including Protected Health Information (PHI)) is subject to strict privacy regulations, directly using data from multiple sources to fine-tune GFM is often impractical. Instead, under the GSCo framework, knowledge from these private datasets can be integrated by independently training each specialist model within the institution where the model is located, thereby protecting the confidentiality of medical data.

[0150] This application discovers that hybrid expert diagnosis and retrieval-enhanced diagnosis provide an effective collaborative strategy between generalist and specialist models. For the research, this application discloses two mechanisms—hybrid expert diagnosis and retrieval-enhanced diagnosis—to provide specialist guidance to MedDr. Hybrid expert diagnosis utilizes the predictions of the specialist model as a reference, while retrieval-enhanced diagnosis treats the specialist model as a retrieval agent to access relevant information in the training dataset. In this collaboration, the roles of MedDr and the expert are distinctly different. The specialist model plays a crucial role in providing reference guidance to GFM to enhance its generalization ability, especially for out-of-domain tasks. Simultaneously, MedDr acts as a decision-maker with extensive medical knowledge. Through extensive experiments, this application demonstrates that these mechanisms enable GFM to benefit from the expertise of the specialist model, thereby creating a powerful and generalizable AI (artificial intelligence) in the medical / healthcare field.

[0151] This paper discusses the limitations and future directions of the framework proposed in this application. Despite some progress, the experimental results also reveal some limitations, pointing the way for future improvements and research. First, while GSCo has shown satisfactory performance on various public datasets, further exploration of its application in clinical practice is needed to confirm the superiority of the method. Second, multimodal retrieval strategies should be further explored. GSCo demonstrates the great potential of retrieval enhancement generation in medical image analysis. However, current vision-based retrieval methods still struggle to ensure the accuracy of retrieved samples. Finally, diverse collaborative paradigms need to be further explored. While GSCo has validated the effectiveness of collaboration, more interaction patterns between generalists and specialists are expected to be observed. For example, combining it with Chain of Thought (CoT) [Reference 55] may be a feasible approach to promote collaboration between generalists and specialists.

[0152] ·method

[0153] GFM and the building of specialists

[0154] Diagnostic-guided bootstrapping and medical image description datasets

[0155] Based on various aspects of this disclosure, this paper discusses the Diagnostic Guided Bootstrapping (DGB) and Medical Image Description (DES) datasets proposed in this application. To leverage the rich medical image diagnostic datasets, the DGB dataset is first introduced. Traditional solutions [References 12, 14] have constructed instruction-adjusted datasets primarily derived from PubMed research articles to alleviate the scarcity of visual-language data in the medical field. While these strategies have successfully constructed large-scale training corpora, they also have two significant limitations. First, they rely solely on textual information, ignoring visual elements, which can lead to inconsistencies in description. Second, the content of research articles may lack reliability and accuracy, thus introducing noise into the training corpus.

[0156] Conversely, based on various aspects of this disclosure, a method for generating instruction-adjusted datasets based on medical image diagnostic datasets is proposed to leverage multimodal information and manually verified annotations. For example... Figure 9 As shown in (a), due to extensive training on multiple corpora, this application observes that GFM in the general domain [References 8,9] demonstrates a comprehensive understanding of disease-related information and associated symptoms. However, these models encounter difficulties in associating this knowledge with specific medical images, leading to inaccurate diagnostic predictions. Conversely, when provided with specific disease and modality information along with a given image, GFM in the general domain is able to generate high-quality medical reports with accurate diagnoses. For example, Figure 9 (b) A generated report on “ulcerative colitis” was described, in which the findings listed the observations in the images, and the preliminary diagnosis summarized the above conclusions.

[0157] Based on the above observations, this application proposes a diagnostic-guided bootstrapping strategy that utilizes both visual and textual information to construct an instruction adjustment dataset. Specifically, the instruction can be as follows: Figure 8 The formatting is as shown, but not limited to. The model contains information about the modalities and diseases of the medical images and requires the generation of detailed reports.

[0158] Compared to traditional solutions that generate data solely from textual information [References 12, 14], the proposed method offers two significant advantages. First, it facilitates the use of large, manually verified, label-level annotated datasets in the medical field. Second, the combination of visual and textual information ensures that the generated information remains relevant to the accompanying images. Following this approach, a DGB dataset covering multiple medical image modalities was constructed.

[0159] To enhance the diversity of training data (for MedDr), this paper provides a Medical Image Description (DES) dataset. Specifically, image-based case studies were collected from OpenI [Reference 36], and image descriptions were written from a medical perspective using GFM [Reference 9] to integrate both image and related textual information. Furthermore, to focus on diagnostically relevant visual features, this application removed details that could not be directly inferred from the images, such as the patient's name or age.

[0160] Medical Order Adjustment

[0161] To enhance the various functionalities of GFM, this application follows a conventional solution [References 12-14], incorporating visual question answering datasets [References 14, 19, 37, 47, 56] and medical report generation datasets [References 21, 22] into the training corpus. Furthermore, to enhance data diversity, this application also collected image-based case studies from OpenI [Reference 36]. Overall, as... Figure 1b As shown, the training corpus covers five different types of projects: Medical Image Diagnosis (CLS), Medical Report Generation (MRG), Visual Question Answering (VQA), Diagnostic Guided Bootstrapping (DGB), and Medical Image Description (DES). It has been proposed to craft specific prompt templates for each type, such as... Figure 25 As shown. More details about the training dataset will be provided below. Language modeling loss is used as the loss function for training the model.

[0162] Collaborative reasoning for downstream tasks

[0163] Hybrid expert diagnosis

[0164] When tackling specific downstream tasks, training lightweight specialist models can often be more practical than using GFM. This advantage lies primarily in its ability to achieve satisfactory performance while significantly reducing training overhead. This study explores the collaborative potential between generalist and specialist models to enhance outcomes for downstream tasks.

[0165] like Figure 10 As shown in (a)-(b), several lightweight models are selected and trained on downstream datasets, and these models are designated as specialist models. Unlike GFM, these specialist models are tailored for specific downstream tasks, leveraging expert knowledge to achieve superior performance on these tasks. During the inference phase, the test image is first input into the specialist model, and the prediction of the specialist model is used as a guiding reference, which is then incorporated into the instruction. MedDr integrates both the test image and the reference prediction to arrive at the final diagnosis.

[0166] Enhanced diagnostic search

[0167] GFM has demonstrated powerful capabilities, but it still faces challenges when processing out-of-domain data [References 9, 57]. To fully utilize the training corpus and enhance its capabilities on out-of-domain datasets, a collaborative mechanism, Retrieval-Augmented Diagnosis, is proposed to assist MedDr.

[0168] Figure 11 Images (a)-(c) illustrate the collaborative mechanism proposed in this application, in which a database is constructed based on training data spanning multiple medical tasks and modalities. To construct the database, for a given image-text pair, the image is encoded by an embedding model, with the visual embedding set as the key and the text as the value. Thus, each image is associated with its visual embedding. During inference, the test image is encoded using the embedding model. The database is then queried using the resulting visual embedding of the test image by calculating the vector (e.g., cosine) similarity between the query and the key in the database. The top k most similar items (k being a positive integer) are retrieved from the database, and their meta-information is incorporated into the instruction as additional textual cues to aid the model in making medical decisions.

[0169] This study first explores the use of a specialist model as a retrieval engine. After fine-tuning for a specific downstream dataset, this specialist model is able to generate more discriminative embeddings, thereby improving retrieval accuracy. It is understood that the best specialist model (among all selected specialist models for RAG) can be used as the retrieval engine in various examples. Furthermore, the use of MedDr's visual encoder as a retrieval engine is explored, which offers two advantages. First, it effectively eliminates the need for an additional embedding module, thus reducing the associated computational cost. By utilizing MedDr's visual encoder as the embedding module, intermediate results can be directly used as embeddings during inference for querying (in the database) without any additional overhead. Second, this approach is applicable to a wider range of scenarios, especially where obtaining qualified specialist models is challenging.

[0170] Implementation Plan Details

[0171] While this should not be considered a limitation, InternVL [Reference 9], the most state-of-the-art GFM in the general domain, was chosen as the base model for MedDr. It contains approximately 40 billion parameters, consisting of 6 billion visual encoders and 34 billion large language models. The input image size was resized to 448×448 pixels. The model was fine-tuned based on both collected and generated data. The number of training samples was approximately 2 million. See the preceding discussion for details on the training dataset and instruction prompts. The instruction tuning recipe followed the recommendations provided by InternVL. All parameters were fixed except for the LoRA component, which consists of approximately 100 million parameters, representing 0.4% of the total parameters. DeepSpeed ​​ZeRO Stage 3 [Reference 58] was also used to optimize the training process. This model was trained for two epochs on 16 NVIDIA H800 GPUs over 72 hours. The specialist model was trained on a single NVIDIA 4090 GPU, with the corresponding training recipe as follows: Figure 28 The data fields are shown in the image.

[0172] The most advanced generalist and specialist models

[0173] This study selected four open-source models from both the general and medical domains as the baseline GFM.

[0174] RadFM [Reference 14] focuses primarily on the radiographic modality. It consists of PMC-LLaMA-13B [Reference 59] as the visual backbone and LLM as the PMC. The model is first trained on 16 million noisy pre-training data points and then fine-tuned on 3 million intra-domain data points.

[0175] LLaVA-Med [Reference 12] is built upon pre-trained LLaVA [Reference 8]. It was fine-tuned using 8 A100 GPUs over a day on approximately 600K concept-aligned samples and 60K instruction-tuning samples.

[0176] Med-Flamingo [Reference 13] is developed based on OpenFlamingo-9B [Reference 60] and can process multiple images interleaved with text. It is trained on a large-scale interleaved dataset based on medical textbooks and the PMC-OA dataset [Reference 61].

[0177] InternVL [Reference 9] is one of the most powerful open-source large-scale visual language models in the general domain. It achieves state-of-the-art performance in general-domain multimodal tasks.

[0178] For these GFMs, the above models are reproduced based on their open-source checkpoints, and the models are evaluated using the same test data. Test hints are set according to their official implementation (see [link to official implementation]). Figure 31 (Data table in the middle).

[0179] For the selection of specialist models, following [Reference 51], ten representative computer vision models were selected and trained on specific downstream datasets:

[0180] VGG16 [Reference 62] is a convolution-based model. It consists of 16 layers and is known for its simplicity and effectiveness.

[0181] AlexNet [Reference 63] is a convolution-based model. It won the ImageNet Large Scale Visual Recognition Challenge (ILSVRC) in 2012 and popularized the use of convolutional neural networks (CNNs).

[0182] ResNet-18 [Reference 52] is a convolution-based model. It introduces residual connections, which allow gradients to flow directly through the network, enabling the training of very deep networks.

[0183] DenseNet-121 [Reference 64] is a convolutional-based model. It encourages feature reuse by connecting each layer to every other layer in a feedforward manner.

[0184] EfficientNet-B4 [Reference 54] is a convolution-based model. It balances depth, width, and resolution to achieve high performance with fewer parameters. It is efficient in both accuracy and computation.

[0185] ViT-B / 16 [Reference 53] is a Transformer-based model. It applies a transformer architecture to image data and achieves competitive results on various vision tasks.

[0186] CLIP ViT-B / 16 [Reference 4] is a transformer-based model. It learns to associate images and text in a joint embedding space, enabling zero-shot classification and other multimodal tasks.

[0187] EVA-02ViT-B / 16 [Reference 65] is a transformer-based model. It improves upon CLIP's training techniques at scale and achieves superior performance with significantly reduced training costs.

[0188] DINO ViT-B / 16 [Reference 66] is a transformer-based model. It utilizes clustering and momentum encoders and achieves powerful performance without using labeled data.

[0189] SAM ViT-B / 16 [Reference 67] is a transformer-based model designed to be a general-purpose model for image segmentation, capable of segmenting any object in an image with minimal user input.

[0190] Figure 27 Detailed information about these specialist models is provided. Specifically, Figure 27 The specifications for ten selected computer vision models are detailed, including the number of parameters, gigabytes per second multiplication and addition (GMAC), and training time per epoch on the dataset, using PneumoniaMNIST as an example. Parameters are quantized in millions (M), and "IN1K" refers to "ImageNet-1K".

[0191] Figure 28 The fine-tuning recipes for the downstream dataset are shown, listing the hyperparameters used to train each vision model. Specific parameters include input dimension, hidden dimension, and dropout rate. Common parameters include batch size, number of epochs, optimizer, learning rate, scheduler, and weight decay. Compared to GFM, these specialist models have significantly fewer parameters and can be trained on consumer-grade hardware such as the NVIDIA 4090 GPU.

[0192] Training dataset

[0193] This article details the training dataset, and Figure 25 A prompt template is described for each dataset.

[0194] Visual question answering dataset

[0195] SLAKE [Reference 47] is a bilingual radiology VQA dataset containing 642 images and 14,000 questions. Only the English portion of the training set was used, which contains 4,919 question-answer pairs.

[0196] VQA-RAD [Reference 37] is an artificially constructed dataset in which clinicians ask questions about naturally occurring radiological images and provide reference answers. It uses 3064 question-answer pairs from the training set, as officially defined.

[0197] Path-VQA [Reference 19] contains 32,799 open-ended questions from 4,998 pathological images, each of which has been manually reviewed to ensure accuracy. Following the official classification, 19,755 question-answer pairs from the training set were used.

[0198] PMC-VQA [Reference 56] is a large-scale medical visual question-answering dataset constructed from image-text pairs from PubMedCentral, covering a wider range of medical image modalities. Following the official partitioning, 152,603 ​​question-answer pairs from the training set were used.

[0199] PMC-CaseReport [Reference 14] is a visual question-answering dataset automatically generated from case reports in the PMC-Inline dataset. According to the official classification, 254,105 question-answer pairs from the training set were used.

[0200] Medical report generation dataset

[0201] MIMIC-CXR [Reference 21] provided 371,920 chest X-rays from 227,943 imaging examinations of 65,079 patients. Training was performed using 337,292 cases based on the RadFM [Reference 14] and R2Gen [Reference 68] datasets.

[0202] IU-Xray [Reference 22] is a set of chest X-ray images paired with their corresponding diagnostic reports. The dataset contains 7470 image-report pairs. 4730 cases from the training set were used in R2Gen [Reference 68].

[0203] Medical image diagnostic dataset

[0204] VinDr-SpineXR [Reference 25] is a large, annotated medical image dataset used to detect and classify spinal lesions from radiographs. It was trained using 6129 samples, based on RADFM [Reference 68].

[0205] VinDr-PCXR [Reference 30] is an open-source, large-scale pediatric chest X-ray dataset used to interpret common chest diseases. It was trained using 4585 samples according to RADFM [Reference 68].

[0206] VinDr-Mammo [Reference 69] is a large benchmark dataset for computer-aided detection and diagnosis in full-view digital mammography. It was trained using 6047 samples according to RADFM [Reference 68].

[0207] VinDr-CXR [Reference 70] is an open, large-scale chest X-ray dataset with annotations from radiologists. The training set contains 15,000 scans, each independently annotated by three radiologists. 45,000 samples were used for training, according to the official partitioning.

[0208] CheXpert [Reference 71] is a large public dataset for interpreting chest X-rays, containing 224,316 chest X-ray images from 65,240 patients. 223,414 samples were used for training, according to the official dataset.

[0209] ChestX-ray14 [Reference 33] is a medical imaging dataset containing 112,120 frontal X-ray images of 30,805 patients, labeled with 14 common diseases obtained through text mining. It was trained using 86,524 samples according to the official dataset.

[0210] PCam200 [Reference 72] is a public pathological H&E image dataset, created in the same way from the Camelyon2016 challenge dataset [Reference 73]. According to the official partitioning, 28,539 samples were used for training.

[0211] PAD-UFES-20 [Reference 34] is a dermatology classification dataset containing 2298 images for six different diagnoses. All 2298 samples were used for training.

[0212] DermNet [Reference 44] contains dermatological images of 23 skin diseases extracted from DermNet. It was trained using 15,557 samples, according to the official classification.

[0213] HAM10000 [Reference 31] is a large collection of multi-source dermoscopic images of pigmented lesions. According to the official classification, 10015 samples were used for training.

[0214] ISIC2020 [Reference 74] is the dataset for the SIIM-ISIC Melanoma Classification Challenge 2020. This dataset contains 33,126 dermoscopic training images from over 2,000 patients, including unique benign and malignant skin lesions. 33,126 samples were used for training according to the official classification.

[0215] Kvasir [Reference 75] is a multi-class image dataset for computer-aided detection of gastrointestinal diseases. It was trained using 8000 samples according to the official classification.

[0216] The Kvasir Capsule [Reference 35] is an endoscopic dataset containing 47,238 images with annotations of anatomical landmarks, pathological findings, and normal findings. It was trained using 47,238 samples according to the official partitioning.

[0217] WCE [Reference 76] is a curative colorectal disease dataset based on Kvasir [Reference 75] and ETIS-Larib-Polyp DB [Reference 77]. It was trained using 3200 samples according to the official partitioning.

[0218] GastroVision [Reference 78] is a multicenter, open-access gastrointestinal (GI) endoscopy dataset containing various anatomical landmarks, pathological abnormalities, polyp removal cases, and normal findings of the gastrointestinal tract. All 8000 samples were used for training.

[0219] ODIR [Reference 79] is a structured ophthalmological database of 5,000 patients, including age, color fundus photographs of both eyes, and diagnostic keywords from doctors. It was trained using 6,392 samples according to the official classification.

[0220] Fundus1000 [Reference 80] contains 1000 fundus images, divided into 39 categories. All 1000 samples were used for training.

[0221] RFMiD2.0 [Reference 32] is a multi-label dataset containing approximately 860 retinal fundus images annotated by three ophthalmologists. 455 samples were used for training, according to the official classification.

[0222] The Retinal OCT-C8 [Reference 45] is a large-scale dataset used for ophthalmic research, containing 24,000 optical coherence tomography (OCT) images divided into 8 categories. 18,000 samples were used for training, based on the official classification.

[0223] UltraBreast is a proprietary breast ultrasound dataset containing 45,896 cases labeled as benign or malignant.

[0224] Synthetic dataset

[0225] The Diagnosis-Guided Bootstrapping Dataset is part of a synthetic dataset. As previously mentioned, this application constructs a large-scale medical report dataset covering multiple treatment modalities. Specifically, a total of 196,760 samples were generated based on the VinDr-SpineXR [Reference 25], VinDr-PCXR [Reference 30], VinDr-Mammo [Reference 69], VinDr-CXR [Reference 70], ChestX-ray14 [Reference 33], PADUFES-20 [Reference 34], Dermnet [Reference 44], Kvasir [Reference 75], WCE [Reference 76], Kvasir Capsule [Reference 35], ODIR [Reference 79], Fundus1000 [Reference 80], and RFMiD2.0 [Reference 32] datasets.

[0226] The medical image description dataset is considered part of the synthetic dataset. To enhance the diversity of the training data, this application collected 245,371 image-based case studies from OpenI [Reference 36]. Case names and image titles were summarized based on these images using InternVL [Reference 9], resulting in high-quality images and corresponding text descriptions.

[0227] External benchmark datasets

[0228] Visual question answering dataset

[0229] VQA-Med [Reference 20] focuses on radiographic images and includes four main categories of questions: modality, planar, organ systems, and anomalies. It uses 500 questions for testing, according to the official classification.

[0230] OmniMedVQA [Reference 28] is a large-scale comprehensive evaluation benchmark dataset for medical GFM. Due to overlap with the training data, some data was excluded to prevent data leakage, and only a subset of disease diagnoses was used. The total number of test samples is 51,977.

[0231] Medical image diagnostic dataset

[0232] PneumoniaMNIST [Reference 26] is a binary classification dataset for chest X-rays. According to the official breakdown, 624 items were used for testing.

[0233] BreastMNIST [Reference 26] is a binary classification dataset for breast ultrasound examinations. According to the official classification, 156 samples were used for testing.

[0234] OrganAMNIST [Reference 26] is a multi-class classification dataset for abdominal CT scans. According to the official classification, 17,778 items were used in the test.

[0235] PathMNIST [Reference 26] is a multi-class classification dataset for colon pathology. According to the official classification, 7180 items were used for testing.

[0236] OCTMNIST [Reference 26] is a multi-class classification dataset for retinal OCT. According to the official classification, 1000 items were used for testing.

[0237] ChestMNIST [Reference 26] is a multi-label dataset for chest X-rays. According to the official classification, 22,433 samples were used for testing.

[0238] CBIS-DDSM [Reference 50] contains images used for screening mammograms. The original dataset contains images of three types of breast cancer cases: benign, benign without callback, and malignant. Since the information in the text is insufficient to distinguish between benign and benign without callback, it is defined as a binary classification task. Following the official split, it was tested using 326 samples from a subset of CALC and 378 samples from MASS.

[0239] FMC-Colon [Reference 27] is a pathological tumor tissue classification dataset that requires the model to determine whether a sample is positive or negative. According to the official classification, 4355 samples were used for testing.

[0240] FMC-Endo [Reference 27] is a colonoscopy lesion classification dataset that includes four different lesion types. Based on the official classification, 2055 samples were used for testing.

[0241] FMC-Chest [Reference 27] is a chest disease screening dataset that covers 19 common chest abnormalities. According to the official classification, 2708 samples were used for testing.

[0242] Derm7pt [Reference 81] is a dataset used to evaluate computer-based image predictions based on a 7-point checklist for the malignancy of skin lesions. According to the official classification, 395 samples were used for testing.

[0243] BRSET [Reference 46] is a multi-label ophthalmology dataset designed to foster the development of the science and technology community and to validate machine learning models. The dataset is randomly divided into training and test sets in an 8:2 ratio. The test set contains 3254 samples.

[0244] Evaluation indicators

[0245] Medical image diagnosis

[0246] For medical image diagnostic datasets, this application uses accuracy and F1 score for evaluation. Accuracy is calculated based on formula (1):

[0247]

[0248] Where y is the target value tensor. Let be the predicted value tensor. The F1 score is defined based on recall and precision according to the following formulas (2)-(4):

[0249]

[0250] Here, TP and FP represent the number of true positives and false positives, respectively. Specifically, for multi-class and multi-label classification datasets, both the macro F1 score and the micro F1 score are calculated.

[0251] Visual Q&A

[0252] For datasets consisting of multiple-choice questions [Reference 28], accuracy is calculated. For other visual question-answering datasets, according to MultiMedEval [Reference 49], both predictions and answers are first tokenized, and then precision and recall are calculated. For closed-ended questions, a prediction is considered correct if the recall is at least 0.5. For open-ended questions, a prediction is considered correct if the recall is at least 0.5. For open-ended questions, a prediction is considered correct if the recall is at least 0.75. This application reports the accuracy for both closed-ended and open-ended questions. Furthermore, according to [Reference 14], the BLEU-1 score is calculated using formula (5):

[0253]

[0254] Where BP stands for brevity penalty, p n For the accuracy of n-grams, w nBP is the weight for the accuracy of n-grams. If the length c of the predicted result is greater than the reference length r, then BP = 1. If c ≤ r, then BP = exp(1 – r / c). This ensures that shorter predictions are penalized to prevent the system from biasing towards overly concise outputs. Since there is only one type of n-gram, w1 = 1.

[0255] Medical report generation

[0256] For medical report generation tasks, common n-gram-based metrics can be used, such as BLEU-1, BLEU-4, ROUGE-1, ROUGE-L, and METEOR [Reference 82].

[0257] Here, ROUGE-1 is defined by formula (6):

[0258]

[0259] Where |Recall∩Reference| represents the number of overlapping unigrams between the generated report and the reference report, while |Reference| represents the total number of unigrams in the reference text. ROUGE-L is defined according to formula (7):

[0260]

[0261] Among them, F LCS This represents the F1 score of the longest common subsequence. The METEOR score is calculated according to formula (8):

[0262]

[0263] Where m represents the number of gold standard (reference) sentences, and Precision(g,h) represents the precision score between a specific gold standard (g) and a hypothetical sentence (h), forming a set of all gold standard sentences (gold) and a set of all hypothetical sentences (hyp).

[0264] In addition, the F1 radiograph (F1-RadGraph) was evaluated, which is based on formula (9) and measures the F1 score between entities extracted from the reference and the generated report using RADGraph [Reference 83]:

[0265]

[0266] CheXbert vector similarity [reference 84] is also calculated using cosine similarity between the embedded reference and the generated report, according to formula (10):

[0267]

[0268] Where A and B are vectors for the reference report and the generated report, respectively.

[0269] Data availability

[0270] The dataset used to build the training dataset is in Figure 29 The evaluation benchmark dataset is listed in [the document / reference]. Figure 30 Listed in. Regarding in Figures 29-30 The labels in the "Access" column are important to note. "Open Access" datasets are publicly accessible and free of charge, while "Restricted Access" datasets require contacting the relevant dataset provider to obtain access permissions. "Credential Access" datasets require specific permissions, and "Private" datasets are not publicly accessible.

[0271] Figure 32 This is a block diagram of a training manager 3205 for training a generalist foundation model (GFM) according to various aspects of this disclosure. The training manager 3205 can be implemented in software as described herein. Figure 2 The method described in the text. The Training Manager 3205 can be... Figures 34-35 The computer devices 3400, 3500, or components thereof shown perform this operation. The training manager 3205 may include a generation component 3210, a configuration component 3215, and a training component 3220. Each of these components may communicate directly or indirectly with each other 3225 (e.g., via one or more buses).

[0272] The generation component 3210 can generate multiple medical reports related to different conditions through GFM and provide them as a first dataset, wherein each medical report is generated based on relevant images displaying the condition, classification labels containing textual descriptions of the condition, and treatment patterns related to the condition.

[0273] The generation component 3210 can further generate text descriptions for corresponding images displaying different conditions via GFM, wherein the text descriptions describe the conditions based on diagnostic-related visual features displayed in the corresponding images, and wherein the generated text descriptions and corresponding images are provided as a second dataset.

[0274] Configuration component 3215 can configure training datasets to include a first dataset, a second dataset, a third dataset associated with medical image diagnosis (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA).

[0275] The training component 3220 can train the GFM based on the training dataset to obtain a medically trained GFM.

[0276] Alternatively, each of the aforementioned components 3210, 3215, 3220 in the training manager 3205 (including the training manager itself) can be implemented as a specific hardware module (e.g., an application-specific integrated circuit (ASIC)) performing the same operation. Nevertheless, the implementation of components 3210, 3215, 3220 can also be achieved as needed through a combination of hardware and software modules.

[0277] Figure 33 This is a block diagram of a coordination manager 3305 for processing patient sample images according to various aspects of this disclosure. The coordination manager 3305 may be used as described herein. Figure 3 The software implementation of the method. The coordination manager 3305 can be implemented by computer devices 3400, 3500 (e.g., ...). Figures 34-35 (as shown) or its components. The coordination manager 3305 may include a processing component 3310 and a query component 3315. Each of these components may communicate directly or indirectly with each other 3320 (e.g., via one or more buses).

[0278] Processing component 3310 can process sample images through a set of computer vision models for condition-related treatment patterns to obtain a set of judgments related to multiple predictive diagnoses of the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, wherein each set of computer vision models is pre-trained using downstream datasets specific to the corresponding treatment patterns for different conditions.

[0279] Processing component 3310 can process sample images using an embedding model to determine relevant visual embeddings to be used as a query for querying a database, wherein the database includes entries of multiple value-key pairs associated with individual images of different conditions; and wherein the key represents the visual embedding of the image, while the associated value provides a textual description of the condition displayed in the image.

[0280] The query component 3315 can query the database based on the vector similarity between the determined visual embedding and the corresponding key in a plurality of entries to retrieve the top k most similar entries from the database, where k is a positive integer.

[0281] Processing component 3310 can process sample images by using a medically trained generalist foundational model (GFM) in conjunction with a decision set related to multiple predictive diagnoses as a reference context, and images corresponding to the top k most similar entries retrieved, to make a predictive judgment on the condition displayed by the sample images. The medically trained GFM is based on... Figure 2 The method discussed is 200 trained. For clarity, medically trained GFM refers to MedDr proposed in this application.

[0282] Alternatively, each of the aforementioned components 3310, 3315 in the coordination manager 3305 (including the coordination manager itself) can be implemented as a specific hardware module (e.g., an ASIC) performing the same operation. Furthermore, the implementation of components 3310, 3315 can optionally be achieved through a combination of hardware and software modules as needed.

[0283] Figure 34 It is applicable to the execution of various aspects of this disclosure. Figure 2 Method 200 or Figure 3 A schematic diagram of an exemplary (first) computing device 3400 of method 300.

[0284] The computing device 3400 includes a keyboard 3402, a touchscreen 3404, a microphone 3406, a speaker 3408, and an antenna 3410. Users can operate the computing device 3400 to perform various functions / tasks, such as answering phone calls, sending SMS messages, browsing the internet, sending emails, and / or providing satellite navigation.

[0285] The computing device 3400 includes hardware for performing communication functions (e.g., telephone or data communication), an application processor, and corresponding supporting hardware to enable the computing device 3400 to perform other functions, such as messaging, internet browsing, email functionality, or similar functions. The communication hardware includes a radio frequency (RF) processor 3412 that provides RF signals to an antenna 3410 for transmitting and receiving data signals. Additionally, a baseband processor 3414 is provided that provides signals to and receives signals from the RF processor 3412. The baseband processor 3414 can also interact with a Subscriber Identity Module (SIM) 3416, as is known in the art. The communication subsystem enables the computing device 3400 to communicate via various communication protocols, including 3G, 4G, 5G, New Radio (NR), GSM, WiFi, Bluetooth, and / or CDMA. The communication subsystem of the computing device 3400 is not within the scope of this invention.

[0286] The keyboard 3402 and touchscreen 3404 are controlled by the application processor 3418. A power and audio controller 3420 supplies power from the battery 3422 to the communication subsystem, application processor 3418, and other hardware. The power and audio controller 3420 also controls input from the microphone 3406 and audio output via the speaker 3408. A Global Positioning System (GPS) antenna and associated receiver elements 3424 are also provided, controlled by the application processor 3418, and capable of receiving GPS signals for satellite navigation functions of the computing device 3400.

[0287] Various types of memory can be provided in the computing device 3400 to complement the operation of the application processor 3418. The computing device 3400 may include random access memory (RAM) 3426 coupled to the application processor 3418, into which data and program code can be written and read. Code stored in RAM 3426 can be executed by the application processor 3418 from RAM 3426. RAM 3426 represents a type of volatile memory in the computing device 3400.

[0288] The computing device 3400 may also be equipped with a non-volatile (long-term) storage device 3428 coupled to the application processor 3418. The storage device 3428 may be logically divided into three partitions: an operating system (OS) partition 3430, a system partition 3432, and a user partition 3434. The storage device 3428 may represent the non-volatile memory of the computing device 3400.

[0289] In this example, OS partition 3430 may contain the firmware of computing device 3400, which includes the operating system. Other computer programs, such as applications (also known as apps), may also be stored in storage device 3428. Specifically, applications essential to the operation of computing device 3400, such as communication applications and similar programs in smartphones, are typically stored in system partition 3432. Applications stored in system partition 3432 are typically programmed into computing device 3400 during factory settings.

[0290] Applications that users subsequently add and install on computing device 3400 can typically be stored in user partition 3434.

[0291] Alternatively, Figure 34 The various functional components shown can be combined into a single component. For example, storage device 3428 may include NAND flash memory, NOR flash memory, hard disk drive, or a combination thereof.

[0292] Figure 35 It is applicable to the execution of various aspects of this disclosure. Figure 2 Method 200 or Figure 3 A schematic diagram of an exemplary (second) computing device 3500 for method 300. The following description of the computing device 3500 is provided by way of example only and is not intended to be limiting.

[0293] like Figure 35As shown, the example computing device 3500 includes a processor 3504 for executing software routines. Although a single processor is shown for clarity, the computing device 3500 may also include a multiprocessor system. The processor 3504 is connected to a communication infrastructure 3506 to communicate with other components of the computing device 3500. The communication infrastructure 3506 may include, for example, a communication bus, a crossbar network, or a network.

[0294] The computing device 3500 also includes a main memory 3508 (e.g., random access memory (RAM)) and a secondary memory 3510. The secondary memory 3510 may include, for example, a hard disk drive 3512 and / or a removable storage drive 3514 (e.g., a floppy disk drive, magnetic tape drive, optical disk drive, or similar drive). As is known in the art, the removable storage drive 3514 reads from and / or writes to the removable storage unit 3518. The removable storage unit 3518 may include a floppy disk, magnetic tape, optical disk, or similar storage unit that can be read from and written to by the removable storage drive 3514. It will be understood in the art that the removable storage unit 3518 includes a computer-readable storage medium in which computer-executable program code instructions and / or data are stored.

[0295] In other respects, the auxiliary storage 3510 may additionally or alternatively include other similar means for allowing computer programs or other instructions to be loaded into the computing device 3500. Such means may include, for example, a removable storage unit 3522 and an associated interface 3520. Examples of removable storage units 3522 and interfaces 3520 may include program cartridges and cartridge interfaces (e.g., program cartridge interfaces found in video game console devices), removable memory chips (e.g., EPROM or PROM) and associated slots, and other exemplary removable storage units 3522 and interfaces 3520 that enable the transfer of software programs and / or data between the removable storage unit 3522 and the computer system 3500.

[0296] The computing device 3500 also includes at least one communication interface 3524. The communication interface 3524 allows software programs and data to be transferred between the computing device 3500 and an external device 3526 via a communication path. In various aspects, the communication interface 3524 allows data to be transferred between the computing device 3500 and a data communication network (e.g., a public or private data communication network). The communication interface 3524 can be used to exchange data between different computing devices 3500 that may collectively form part of an interconnected computer network. Examples of the communication interface 3524 may include a modem, a network interface (e.g., an Ethernet card), a communication port, an antenna with associated circuitry, or a similar device. The communication interface 3524 can be wired or wireless. Software and data transmitted via the communication interface 3524 exist in the form of signals, which may be electronic signals, electromagnetic signals, optical signals, or other signals that can be received by the communication interface 3524. These signals are provided to the communication interface via the communication path 3526.

[0297] The computing device 3500 also includes a display interface 3502 and an audio interface 3532. The display interface is configured to perform the operation of rendering an image to the associated display 3530, while the audio interface 3532 is used to perform the operation of playing audio content via the associated speaker 3534.

[0298] As used herein, the term "computer program product" may refer in part to removable storage unit 3518, removable storage unit 3522, hard disk installed in hard disk drive 3512, or carrier wave carrying software to communication interface 3524 via communication path 3526 (wireless link or cable). Computer-readable storage medium refers to any non-transitory tangible storage medium that provides recorded instructions and / or data to computing device 3500 for execution and / or processing. Examples of such storage media include floppy disks, magnetic tapes, CD-ROMs, DVDs, and Blu-ray discs. TM Optical discs, hard disk drives, ROMs or integrated circuits, USB storage devices, magneto-optical discs or computer-readable cards (such as PCMCIA cards) or similar storage devices, whether such devices are located inside or outside the computing device 3500. Examples of temporary or non-tangible computer-readable transmission media that may also be involved in providing software, applications, instructions and / or data to the computing device 3500 include radio or infrared transmission channels and network connections to another computer or networked device, as well as the Internet or intranets (including email transmissions and information recorded on websites) and similar media.

[0299] Computer programs (also referred to as computer program code / instructions) are stored in main memory 3508 and / or auxiliary memory 3510. Computer programs may also be received via communication interface 3524. When executed, such computer programs enable computing device 3500 to perform one or more aspects of the present disclosure. In various aspects of the present disclosure, when executed, computer programs enable processor 3504 to perform one or more aspects of the present disclosure. Therefore, such computer programs can represent the controller of computer system 3500.

[0300] The software can be stored in a computer program product and loaded into a computing device 3500 using a removable storage drive 3514, a hard disk drive 3512, or an interface 3520. Alternatively, the computer program product can be downloaded into the computer system 3500 via a communication path 3526. When executed by a processor 3504, the software causes the computing device 3500 to perform various aspects of this disclosure.

[0301] It should be understood that Figure 35 The computing device 3500 shown in the text is by way of example only. Therefore, in some aspects, one or more features of the computing device 3500 may be omitted. Furthermore, in other aspects, one or more features of the computing device 3500 may be combined or juxtaposed. Moreover, in some aspects, one or more features of the computing device 3500 may be divided into one or more components.

[0302] It should be understood that Figure 35 The components shown can be used to provide execution Figure 2 Method 200 or Figure 3 The various functional devices of method 300 are as described in various aspects of this disclosure. Furthermore, the term "computing device" 3400, 3500 may include or be referred to as a mobile device, wireless device, remote device, handheld device, tablet computer, laptop computer, computer server, computer terminal, blade server, and other instances. The computer devices 3400, 3500 described herein may be capable of communicating with various types of devices, such as other computer devices 3400, 3500 that may sometimes act as repeaters, or may be configured to work together as a computer cluster performing high-performance computing.

[0303] All methods described herein depict possible implementations, while operations and steps may be rearranged or otherwise modified, and other implementations are also possible. Furthermore, aspects of two or more methods may be combined, if applicable.

[0304] The information and signals described herein can be represented using any of a variety of different techniques and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout this specification can be represented by voltage, current, electromagnetic waves, magnetic fields or particles, light fields or particles, or any combination thereof.

[0305] Various exemplary blocks and components related to the disclosure herein may be implemented or performed by a general-purpose processor, DSP, ASIC, CPU, FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof, which are intended to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may also be any processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration).

[0306] The functions described herein can be implemented by hardware, software executed by a processor, firmware, or any combination thereof. If implemented by software executed by a processor, these functions can be stored on or transmitted via a computer-readable medium as one or more instructions or code. Other examples and implementations are within the scope of this disclosure and the appended claims. For example, due to the nature of software, the functions described herein can be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination thereof. Features implementing these functions can also be physically located in different locations, including a distributed arrangement, such that some portions of the functions are implemented in different physical locations.

[0307] Computer-readable media includes both non-transitory computer storage media and communication media that encompasses any medium that facilitates the transfer of a computer program from one place to another. Non-transitory storage media can be any available medium accessible by a general-purpose or special-purpose computer. For example, but not limited to, non-transitory computer-readable media can include RAM, ROM, electrically erasable programmable EPROM, flash memory, optical disc (CD) ROM or other optical disc storage devices, magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to carry or store required program code in the form of instructions or data structures and is accessible by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Furthermore, any connection can be appropriately referred to as computer-readable media. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are all included in the definition of computer-readable media. As used in this article, "disk" and "optical disc" include CDs, laser discs, optical discs, digital multifunction discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs use lasers to copy data optically. Combinations of both are also included within the scope of computer-readable media.

[0308] The word "or" used in the list of items herein (including the claims) (e.g., a list of items beginning with phrases such as "at least one" or "one or more") indicates an inclusive list, such as a list of at least one of A, B, or C indicating A or B or C or AB or AC or BC or ABC (e.g., A and B and C). Furthermore, the phrase "based on" as used herein should not be construed as referring to a closed set of conditions. For example, an example step described as "based on condition A" may be based on both condition A and condition B without departing from the scope of this disclosure. In other words, the phrase "based on" as used herein should be interpreted in the same way as the phrase "at least partially based on".

[0309] In the accompanying drawings, similar components or features may have the same reference numerals. Furthermore, various components of the same type can be distinguished by adding a dash after the reference numeral and a second numeral to differentiate similar components. If only the first reference numeral is used in the specification, the description applies to any similar component having the same first reference numeral, regardless of what the second or other subsequent reference numerals are.

[0310] The specification described herein, in conjunction with the accompanying drawings, describes exemplary configurations and does not represent all implementable examples or all examples falling within the scope of the claims. The term "example" as used herein means "serving as an example, instance, or illustration," and not "preferred" or "superior to other examples." The detailed description includes specific details intended to aid in understanding the techniques described. However, these techniques can be implemented without these specific details. In some cases, known structures and apparatuses are shown in block diagram form to avoid obscuring the concepts of the examples.

[0311] The description herein is intended to enable those skilled in the art to make or use the present disclosure. Various modifications to the present disclosure will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other variations without departing from the scope of this disclosure. Therefore, this disclosure is not limited to the examples and designs described herein, but should be given the widest scope consistent with the principles and novel features described herein.

[0312] Example

[0313] The following examples are disclosed in accordance with various aspects of this disclosure.

[0314] Example 1: A computer-implemented method for training a generalist foundational model (GFM), the method comprising: generating multiple medical reports related to different conditions via the GFM, the medical reports being provided as a first dataset, wherein each medical report is generated based on a relevant image displaying the condition, a classification label containing a textual description of the condition, and a treatment pattern associated with the condition; generating textual descriptions via the GFM for corresponding images displaying the different conditions, wherein the textual descriptions describe the condition based on diagnostically relevant visual features displayed in the corresponding images, and wherein the generated textual descriptions and the corresponding images are provided as a second dataset; configuring a training dataset to include the first dataset, the second dataset, a third dataset associated with medical image diagnostics (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA); and training the GFM based on the training dataset to obtain a medically trained GFM.

[0315] Example 2: According to the method of Example 1, wherein the GFM includes one selected from the following: RadFM, LLaVA-Med, Med-Flamingo, and InternVL.

[0316] Example 3: According to the method of Example 1, each medical report is configured to include information on medical findings and preliminary diagnoses related to the condition as shown by the relevant images, based on the processing of the GFM.

[0317] Example 4: According to the method of Example 1, wherein the corresponding images showing the different conditions are obtained from OpenI.

[0318] Example 5: According to the method of Example 1, wherein the treatment modality includes one of the following: radiology, pathology, dermatology, ophthalmology, gastroenterology, fundus examination, chest X-ray, and endoscopy.

[0319] Example 6: A computer-implemented method for processing sample images of a condition, the method comprising: processing the sample images by means of a set of computer vision models for treatment patterns associated with the condition to obtain a set of determinations associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, wherein each set of computer vision models is pre-trained using downstream datasets of treatment patterns specific to different conditions; processing the sample images by means of an embedding model to determine associated visual embeddings, the visual embeddings being used as queries for querying a database, wherein the database comprises entries of multiple value-key pairs associated with corresponding images displaying different conditions; wherein the key represents the visual value of the image. The visual embedding, and the associated value, provides a textual description of the condition displayed in the image; using at least one computer vision model from a selected set of computer vision models, a database is queried based on the vector similarity between the determined visual embedding and the corresponding key in the plurality of entries to retrieve the top k most similar entries from the database, where k is a positive integer; and the sample image is processed using a medically trained generalist foundation model (GFM) by using a set of obtained judgments as a reference context and images corresponding to the retrieved top k most similar entries to make a predictive judgment on the condition displayed in the sample image, wherein the medically trained generalist foundation model is trained according to the method described in Example 1.

[0320] Example 7: The method of Example 6, wherein the vector similarity is performed based on cosine similarity.

[0321] Example 8: A computing device for training a generalist foundational model (GFM), comprising: one or more memories having executable code; and one or more processors coupled to the one or more memories, the processors being configured to execute the code to cause the computing device to perform the following operations: generating multiple medical reports related to different conditions via the GFM, providing the medical reports as a first dataset, wherein each medical report is generated based on a relevant image displaying the condition, a classification label containing a textual description of the condition, and a treatment pattern related to the condition; generating textual descriptions via the GFM for corresponding images displaying different conditions, wherein the textual descriptions describe the condition based on diagnostically relevant visual features displayed in the corresponding images, and wherein the generated textual descriptions and the corresponding images are provided as a second dataset; configuring a training dataset to include the first dataset, the second dataset, a third dataset associated with medical image diagnostics (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA); and training the GFM based on the training dataset to obtain a medically trained GFM.

[0322] Example 9: The computing device according to Example 8, wherein the GFM is selected from one of the following: RadFM, LLaVA-Med, Med-Flamingo, and InternVL.

[0323] Example 10: The computing device according to Example 8, wherein the corresponding images displaying the different conditions are acquired from OpenI.

[0324] Example 11: According to the computing device of Example 8, the treatment mode includes one of the following: radiology, pathology, dermatology, ophthalmology, gastroenterology, fundus examination, chest X-ray examination, and endoscopy.

[0325] Example 12: A computing device for processing sample images of a condition, comprising: one or more memories having executable code; and one or more processors coupled to the one or more memories, the processors being configured to execute the code to cause the computing device to perform the following operations: processing the sample images with a set of computer vision models for treatment patterns associated with the condition to obtain a set of decisions associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, wherein each set of computer vision models is pre-trained using downstream datasets of various treatment patterns specific to different conditions; processing the sample images with an embedding model to determine associated visual embeddings, which will be used as queries for querying a database, wherein the database includes data related to display... The method involves generating multiple value-key pairs associated with corresponding images of different conditions; wherein the key represents a visual embedding of the image, and the associated value provides a textual description of the condition displayed in the image; using at least one computer vision model from a selected set of computer vision models, the method queries the database based on the vector similarity between the determined visual embedding and the corresponding key in the multiple entries to retrieve the top k most similar entries from the database, where k is a positive integer; and using a medically trained generalist foundational model (GFM), the method processes the sample images using a set of judgments obtained as a reference context and images corresponding to the retrieved top k most similar entries to make a predictive judgment on the condition displayed by the sample images, wherein the medically trained generalist foundational model is trained according to the method described in Example 1.

[0326] Example 13: The computing apparatus according to Example 12, wherein the vector similarity calculation is performed based on cosine similarity.

[0327] Example 14: A non-transitory computer-readable medium comprising executable code that, when executed by a processor of a computing device, causes the computing device to perform any one of Examples 1-5.

[0328] Example 15: A non-transitory computer-readable medium comprising executable code that, when executed by a processor of a computing device, causes the computing device to perform the method described in any one of Examples 6-7.

[0329] Example 16: A computational apparatus for training a generalist foundational model (GFM), comprising: means for generating multiple medical reports related to different conditions via the GFM, the multiple medical reports being provided as a first dataset, wherein each medical report is generated based on a relevant image displaying the condition, a classification label containing a textual description of the condition, and a treatment pattern associated with the condition; means for generating textual descriptions via the GFM for corresponding images displaying different conditions, wherein the textual descriptions describe the condition based on diagnostically relevant visual features displayed in the corresponding images, and wherein the generated textual descriptions and the corresponding images are provided as a second dataset; means for configuring a training dataset to include the first dataset, the second dataset, a third dataset associated with medical image diagnostics (CLS), a fourth dataset associated with medical report generation (MRG), and a fifth dataset associated with visual question answering (VQA); and means for training the GFM based on the training dataset to obtain a medically trained GFM.

[0330] Example 17: A computational apparatus for processing sample images of a condition, comprising: means for processing the sample images by means of a set of computer vision models for treatment patterns associated with the condition to obtain a set of determinations associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, wherein each set of computer vision models is pre-trained using downstream datasets of treatment patterns specific to different conditions; means for processing the sample images by means of an embedding model to determine associated visual embeddings, the visual embeddings being used as queries for querying a database, wherein the database includes multiple value-key pairs of entries associated with corresponding images displaying different conditions; wherein the keys represent visual embeddings of the images. The associated value provides a textual description of the condition displayed in the image; means for querying a database based on vector similarity between a determined visual embedding and a corresponding key in the plurality of entries, using at least one computer vision model from a selected set of computer vision models, to retrieve the top k most similar entries from the database, where k is a positive integer; and means for processing the sample image using a medically trained generalist foundation model (GFM), using an obtained set of judgments as a reference context, and images corresponding to the retrieved top k most similar entries, to make a predictive judgment on the condition displayed by the sample image, wherein the medically trained generalist foundation model is trained according to the method described in Example 1.

[0331] Reference List

[0332] [1] Achiam, J. et al., “GPT-4 technical report”, arXiv preprint, arXiv:2303.08774 (2023).

[0333] [2]Touvron, H. et al., “Llama2: Open foundation and fine-tuned chat models”, arXiv preprint, arXiv:2307.09288 (2023).

[0334] [3] Anil, R. et al., “Palm2 technical report”, arXiv preprint, arXiv:2305.10403 (2023).

[0335] [4] Radford, A. et al., Learning transferable visual models from natural language supervision, 8748–8763 (PMLR, 2021).

[0336] [5]Tung, C., Lin, Y., Yin, J., Ye, Q. & Chen, H., “Exploring vision language pretraining with knowledge enhancement via large language model”, 81–91 (Springer, 2024).

[0337] [6] Xu, Y. et al., “A multimodal knowledge-enhanced whole-slide pathologyfoundation model”, arXiv preprint, arXiv:2407.15362 (2024).

[0338] [7] Zhu, D., Chen, J., Shen, X., Li, X. & Elhoseiny, M., “Minigpt-4: Enhancing vision-language understanding with advanced large language models”, arXiv preprint, arXiv:2304.10592 (2023).

[0339] [8] Liu, H., Li, C., Wu, Q. & Lee, YJ, “Visual instruction tuning”, Advances in neural information processing systems

[0340] 36 (Advances in Neural Information Processing Systems 36) (2024).

[0341] [9] Chen, Z. et al., “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks”, arXiv preprint.

[0342] arXiv:2312.14238(2023).

[0343]

[10] Antol, S. et al., “VQA: Visual Question Answering”

[0344] 2425–2433 (2015).

[0345]

[11] Lin, T.-Y. et al., “Microsoft coco: Common objects in context

[0346] (Microsoft Cocoa: Common Objects in Context), 740–755 (Springer, 2014).

[0347]

[12] Li, C. et al., “Llava-med: Training a large language-and-vision assistant for biomedicine in one day”, Advances in Neural Information Processing Systems 36 (2024).

[0348]

[13] Moor, M. et al., “Med-flamingo: a multimodal medical few-shot learner”, 353–367.

[0349] (PMLR, 2023).

[0350]

[14] Wu, C., Zhang, X., Zhang, Y., Wang, Y. & Xie, W., “Towards generalistfoundation model for RADiology”, arXiv preprint, arXiv:2308.02463 (2023).

[0351]

[15] Tu, T. et al., “Towards generalist biomedical AI”, NEJM AI1, AIoa2300138 (2024).

[0352]

[16] Moor, M. et al., “Foundation models for generalist medical artificial intelligence”, Nature 616,

[0353] 259–265 (2023).

[0354]

[17] Saab, K. et al., “Capabilities of gemini models in medicine”, arXiv preprint, arXiv:2404.18416 (2024).

[0355]

[18] Yang, L. et al., “Advancing multimodal medical capabilities of gemini”, arXiv preprint, arXiv:2405.03162 (2024).

[0356]

[19] He, X., Zhang, Y., Mou, L., Xing, E. & Xie, P., “PathVQA: 30000+ questions formedical visual question answering”, arXiv preprint, arXiv:2003.10286 (2020).

[0357]

[20] Ben Abacha, A. et al., “Vqa-med: Overview of the medical visual question answering task at imageclef2019”, Vol. 2380 of CEUR Workshop Proceedings (CEUR-WS.org, Lugano, Switzerland, 2019), URL: https: / / ceur-ws.org / Vol-2380 / paper272.pdf.

[0358]

[21] Johnson, AE et al., “Mimic-cxr, a de-identified publicly available database of chest RADiographs with free-text reports”, Scientific Data 6, 317 (2019).

[0359]

[22] Demner-Fushman, D. et al., “Preparing a collection of RADiology examinations for distribution and retrieval”, Journal of the American Medical Informatics Association 23, 304–310 (2016).

[0360]

[23] Jin, H., Che, H., Lin, Y. & Chen, H., “Promptmrg: Diagnosis-driven prompts for medical report generation”, Vol. 38, 2607–2615 (2024).

[0361]

[24] Chen, Z., Luo, L., Bie, Y. & Chen, H., “Dia-llama: Towards large languagemodel-driven ct report generation”, arXiv preprint, arXiv:2403.16386 (2024).

[0362]

[25] Nguyen, HT et al., “Vindr-spinexr: A deep learning framework for spinal lesions detection and classification from RADiographs”, 291–301 (Springer, 2021).

[0363]

[26] Yang, J. et al., “Medmnist v2-a large-scale lightweight benchmark for 2d and 3d biomedical image classification”, Scientific Data 10, 41 (2023).

[0364]

[27] Wang, D. et al., “A real-world dataset and benchmark for foundationmodel adaptation in medical image classification”, Scientific Data 10, 574 (2023).

[0365]

[28] Hu, Y. et al., “OmnimedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM”, arXiv preprint, arXiv:2402.09181 (2024).

[0366]

[29] Chen, P. et al., Gmai-mmbench: A comprehensive multimodal evaluation benchmark towards general medical AI, arXiv preprint, arXiv:2408.03361 (2024).

[0367]

[30] Pham, HH, Tran, TT & Nguyen, HQVindr-pcxr: An open, large-scale pediatric chest x-ray dataset for interpretation of common thoracic diseases, PhysioNet (version 1.0.0) 10 (2022).

[0368]

[31] Tschandl, P., Rosendahl, C. & Kittler, H., “The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skinlesions”, Scientific data 5, 1–9 (2018).

[0369]

[32] Panchar, S. et al., “Retinal fundus multi-disease image dataset (rfmid) 2.0: A dataset of frequently and rarely identified diseases”, Data8 (2023), URL https: / / www.mdpi.com / 2306-5729 / 8 / 2 / 29.

[0370]

[33] Wang, X. et al., “Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases”, 2097–2106 (2017).

[0371]

[34] Pacheco, AG et al., “Pad-ufes-20: A skin lesion dataset composed of patient data and clinical images collected from smartphones”, Data in brief 32, 106221 (2020).

[0372]

[35] Smedsrud, PH et al., “Kvasir-Capsule, a video capsule endoscopy dataset”, Scientific Data 8, 142 (2021).

[0373]

[36] Demner-Fushman, D., Antani, S., Simpson, M. & Thoma, GR, “Design and development of a multimodal biomedical information retrieval system”, Journal of Computing Science and Engineering 6, 168–177 (2012).

[0374]

[37] Lau, JJ, Gayen, S., Demner, D. & Abacha, AB, “Visual question answering in RADiology (vqa-rad)”, Open Science Framework (2018).

[0375]

[38] Jacobs, RA, Jordan, MI, Nowlan, SJ & Hinton, GE, “Adaptive mixtures of local experts”, Neural computation 3, 79–87 (1991).

[0376]

[39] Fedus, W., Zoph, B. & Shazeer, N., “Switch transformers: Scaling totrillion parameter models with simple and efficient sparsity”, Journal of Machine Learning Research 23, 1–39 (2022).

[0377]

[40] Luo, L. et al., “Towards non-invasive and personalized management of breast cancer patients from multiparametric MRI via a large mixture-of-modality-experts model”, arXiv preprint, arXiv:2408.12606 (2024).

[0378]

[41] Xiong, C. et al., “Mome: Mixture of multimodal experts for cancer survival prediction”, arXiv preprint, arXiv:2406.09696 (2024).

[0379]

[42] Lewis, P. et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks”, Advances in Neural Information Processing Systems 33, 9459–9474 (2020).

[0380]

[43] Zhang, K. et al., “A generalist vision–language foundation model for diverse biomedical tasks”, Nature Medicine 1–13 (2024).

[0381]

[44] Goel, S. Dermnet., https: / / www.kaggle.com / datasets / shubhamgoel27 / dermnet (2020).

[0382]

[45] Subramanian, M., Shanmugavadivel, K., Naren, OS, Premkumar, K. & Ranlish, K., “Classification of retinal oct images using deep learning”, 1–7 (2022).

[0383]

[46] Nakayama, LF et al., “A Brazilian multilabel ophthalmological dataset (brset)”, PhysioNet https: / / doi.org / 1013026 (2023).

[0384]

[47] Liu, B. et al., “Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering”, 1650–1654 (IEEE, 2021).

[0385]

[48] ​​Zhang, X. et al., “Pmc-vqa: Visual instruction tuning for medical visual question answering”, arXiv preprint, arXiv:2305.10415 (2023).

[0386]

[49] Royer, C.,Menze, B. & Sekuboyina, A., “Multimedeval: A benchmark and a toolkit for evaluating medical vision-language models” (2024), 2402.09262.

[0387]

[50] Sawyer-Lee, R., Gimenez, F., Hoogi, A. & Rubin, D., “Curated breastimaging subset of digital database for screening mammography (cbis-ddsm) [skuppodataka]”, The cancer imaging archive (2016).

[0388]

[51] Doerrich, S., Di Salvo, F., Brockmann, J. & Ledig, C., “Rethinking model prototyping through the medmnist+dataset collection”, arXiv preprint, arXiv:2404.15786 (2024).

[0389]

[52] He, K., Zhang, X., Ren, S. & Sun, J., “Deep residual learning for image recognition”, 770–778 (2016).

[0390]

[53] Dosovitskiy, A. et al., “An image is worth 16x16 words: Transformers for image recognition at scale”, arXiv preprint, arXiv:2010.11929 (2020).

[0391]

[54] Tan, M. & Le, Q., “Efficientnet: Rethinking model scaling for convolutional neural networks”, 6105–6114 (PMLR, 2019).

[0392]

[55] Wei, J. et al., “Chain-of-thought prompting elicits reasoning in large language models”, Advances in neural information processing systems 35, 24824–24837 (2022).

[0393]

[56] Zhang, X. et al., “Pmc-vqa: Visual instruction tuning for medical visual question answering”, arXiv preprint, arXiv:2305.10415 (2023).

[0394]

[57] Van Veen, D. et al., “RADadapt: ​​RADiology report summarization via lightweight domain adaptation of large language models”, arXiv preprint, arXiv:2305.01146 (2023).

[0395]

[58] Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S. & He, Y., “Zero-infinity: Breaking the GPU memory wall for extreme scale deep learning”, 1–14 (2021).

[0396]

[59] Wu, C., Zhang, X., Zhang, Y., Wang, Y. & Xie, W., “Pmc-llama: Further fine-tuning llama on medical papers”, arXiv preprint, arXiv:2304.14454 (2023).

[0397]

[60] Awadalla, A. et al., “Openflamingo: An open-source framework for training large autoregressive vision-language models”, arXiv preprint, arXiv:2308.01390 (2023).

[0398]

[61] Lin, W. et al., “Pmc-clip: Contrastive language-image pre-training using biomedical documents”, 525–536 (Springer, 2023).

[0399]

[62] Simonyan, K. & Zisserman, A., “Very deep convolutional networks for large-scale image recognition”, arXiv preprint, arXiv:1409.1556 (2014).

[0400]

[63] Krizhevsky, A., Sutskever, I. & Hinton, GE, “ImageNet classification with deep convolutional neural networks”, Advances in neural information processing systems 25 (2012).

[0401]

[64] Huang, G., Liu, Z., Van Der Maaten, L. & Weinberger, KQ, “Denselyconnected convolutional networks”, 4700–4708 (2017).

[0402]

[65] Fang, Y. et al., “Eva-02: A visual representation for neon genesis”, arXiv preprint, arXiv:2303.11331 (2023).

[0403]

[66] Caron, M. et al., “Emerging properties in self-supervised vision transformers”, 9650–9660 (2021).

[0404]

[67] Kirillov, A. et al., “Segment anything”, 4015–4026 (2023).

[0405]

[68] Chen, Z., Song, Y., Chang, T.-H. & Wan, X., “Generating RADiology reports via memory-driven transformer” (2020).

[0406]

[69] Nguyen, HT et al., “Vindr-mammo: A large-scale benchmark dataset for computer-aided diagnosis in full-field digital mammography”, Scientific Data 10, 277 (2023).

[0407]

[70] Nguyen, HQ et al., “Vindr-cxr: An open dataset of chest x-rays with RADiologist's annotations” (2020), 2012.15029.

[0408]

[71] Irvin, J. et al., “Chexpert: A large chest RADiograph dataset with uncertainty labels and expert comparison”, Vol. 33, 590–597 (2019).

[0409]

[72] Kawai, M., Ota, N. & Yamaoka, S., “Large-scale pretraining on pathological images for fine-tuning of small pathological benchmarks”, 257–267 (Springer, 2023).

[0410]

[73] Bejnordi, BE et al., “Diagnostic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer”, Jama 318, 2199–2210 (2017).

[0411]

[74] Rotemberg, V. et al., “A patient-centric dataset of images and metadata for identifying melanomas using clinical context”, Scientific Data 8, 34 (2021).

[0412]

[75] Pogorelov, K. et al., “Kvasir: A multi-class image dataset for computer-aided gastrointestinal disease detection”, 164–169 (2017).

[0413]

[76] Montalbo, FJ, “WCE curated colon disease dataset deep learning” https: / / www.ka ggle.com / datasets / francismon / curated-colon-data set-for-deep-learning(2022).

[0414]

[77] Silva, J., Histace, A., Romain, O., Dray, X. & Granado, B., “Toward embedded detection of polyps in wce images for early diagnosis of colorectal cancer”, International journal of computer assisted RADiology and surgery 9, 283–293 (2014).

[0415]

[78] Jha, D. et al., “Gastrovision: A multi-class endoscopy image dataset for computer aided gastrointestinal disease detection”, 125–140 (Springer, 2023).

[0416]

[79] Li, N., Li, T., Hu, C., Wang, K. & Kang, H., “A benchmark of ocular disease intelligent recognition: One shot for multi-disease detection”, 177–193 (Springer, 2021).

[0417]

[80] Cen, L.-P. et al., “Automatic detection of 39 fundus diseases and conditions in retinal photographs using deep neural networks”, Nature Communications 12, 4828 (2021).

[0418]

[81] Kawahara, J., Daneshvar, S., Argenziano, G. & Hamarneh, G., “Seven-point check-list and skin lesion classification using multitask multimodal neuralnets”, IEEE journal of biomedical and health informatics 23, 538–546 (2018).

[0419]

[82] Banerjee, S. & Lavie, A., “Meteor: An automatic metric for machine translation evaluation with improved correlation with human judgments”, 65–72 (2005).

[0420]

[83] Jain, S. et al., “RADgraph: Extracting clinical entities and relations from RADiology reports”.

[0421]

[84] Yu, F. et al., “Evaluating progress in automatic chest x-ray RADiologyreport generation”, Patterns 4 (2023).

Claims

1. A computer-implemented method for training a generalist foundational model (GFM), the method comprising: The GFM generates multiple medical reports related to different conditions, which are provided as a first dataset. Each medical report is generated based on relevant images showing the condition, classification labels containing textual descriptions of the condition, and treatment patterns associated with the condition. The GFM generates text descriptions for corresponding images displaying different disease conditions, wherein the text descriptions describe the disease conditions based on diagnostic-related visual features displayed in the corresponding images, and wherein the generated text descriptions and the corresponding images are provided as a second dataset; Configure the training dataset to include the first dataset, the second dataset, the third dataset associated with medical image diagnosis (CLS), the fourth dataset associated with medical report generation (MRG), and the fifth dataset associated with visual question answering (VQA); as well as The GFM is trained based on the training dataset to obtain a medically trained GFM.

2. The method according to claim 1, wherein, The GFM includes one of the following: RadFM, LLaVA-Med, Med-Flamingo, and InternVL.

3. The method according to claim 1, wherein, Each medical report is configured to include information on medical findings and preliminary diagnosis of the condition as shown by the relevant images, based on the processing of the GFM.

4. The method according to claim 1, wherein, The corresponding images used to display different disease conditions were acquired from OpenI.

5. The method according to claim 1, wherein, The treatment modalities include one of the following: radiology, pathology, dermatology, ophthalmology, gastroenterology, fundus examination, chest X-ray, and endoscopy.

6. A computer-implemented method for processing sample images of a medical condition, the method comprising: The sample images are processed by a set of computer vision models for treatment patterns associated with the condition to obtain a set of judgments associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, each set of computer vision models being pre-trained using downstream datasets specific to the corresponding treatment patterns for different conditions. The sample images are processed using an embedding model to determine associated visual embeddings, which will be used as queries for a database, wherein the database includes entries of multiple value-key pairs associated with corresponding images of different conditions; wherein the key represents the visual embedding of the image, and the associated value provides a textual description of the condition displayed in the image; Using at least one computer vision model from a selected set of computer vision models, the database is queried based on vector similarity between the determined visual embedding and the corresponding key in the plurality of entries to retrieve the k most similar entries from the database, where k is a positive integer; and Using a medically trained generalist foundational model (GFM), a set of obtained judgments as a reference context, and images corresponding to the top k most similar entries, the sample images are processed to make a predictive judgment on the disease condition displayed in the sample images. The medically trained generalist basic model is trained according to the method described in claim 1.

7. The method according to claim 6, wherein, The vector similarity is performed based on cosine similarity.

8. A computing device for training a generalist foundation model (GFM), comprising: One or more memories containing executable code; as well as One or more processors, coupled to the one or more memories and configured to execute the code to cause the computing device to perform the following operations: The GFM generates multiple medical reports related to different conditions, which are provided as a first dataset. Each medical report is generated based on relevant images displaying the condition, classification labels containing textual descriptions of the condition, and treatment patterns related to the condition. The GFM generates text descriptions for corresponding images displaying different disease conditions, wherein the text descriptions describe the disease conditions based on diagnostic-related visual features displayed in the corresponding images, and wherein the generated text descriptions and the corresponding images are provided as a second dataset; The training dataset is configured to include the first dataset, the second dataset, the third dataset associated with medical image diagnosis (CLS), the fourth dataset associated with medical report generation (MRG), and the fifth dataset associated with visual question answering (VQA); as well as The GFM is trained based on the training dataset to obtain a medically trained GFM.

9. The computing device according to claim 8, wherein, The GFM includes one of the following: RadFM, LLaVA-Med, Med-Flamingo, and InternVL.

10. The computing device according to claim 8, wherein, The corresponding images showing different disease conditions were acquired from OpenI.

11. The computing device according to claim 8, wherein, The treatment modalities include one of the following: radiology, pathology, dermatology, ophthalmology, gastroenterology, fundus examination, chest X-ray, and endoscopy.

12. A computing device for processing sample images of a disease condition, comprising: One or more memories containing executable code; as well as One or more processors, coupled to the one or more memories and configured to execute the code to cause the computing device to perform the following operations: The sample images are processed by a set of computer vision models for treatment patterns associated with the condition to obtain a set of judgments associated with multiple predicted diagnoses for the condition, wherein the set of computer vision models is selected from multiple different sets of computer vision models, each set of computer vision models being pre-trained using downstream datasets of corresponding treatment patterns specific to different conditions. The sample images are processed by an embedding model to determine associated visual embeddings, which will be used as queries for a database, wherein the database includes entries of multiple value-key pairs associated with corresponding images of different conditions; wherein the key represents the visual embedding of the image and the associated value provides a textual description of the condition displayed in the image; Using at least one computer vision model from a selected set of computer vision models, the database is queried based on vector similarity between the determined visual embedding and the corresponding key in the plurality of entries to retrieve the k most similar entries from the database, where k is a positive integer; and Using a medically trained generalist foundational model (GFM), a set of obtained judgments as a reference context, and images corresponding to the top k most similar entries retrieved, the sample images are processed to make a predictive judgment on the disease condition displayed in the sample images. The medically trained generalist basic model is trained according to the method described in claim 1.

13. The computing device according to claim 12, wherein, The vector similarity calculation is performed based on cosine similarity.

14. A non-transitory computer-readable medium comprising executable code that, when executed by a processor of a computing device, causes the device to perform the method of any one of claims 1-5.

15. A non-transitory computer-readable medium comprising executable code that, when executed by a processor of a computing device, causes the device to perform the method of any one of claims 6-7.