Systems and methods for disease classification and treatment recommendation using large language models

The system addresses the challenges of inconsistent disease classification and treatment recommendations by using large language models to analyze medical images and integrate patient data, achieving accurate and patient-preference-informed outcomes.

WO2025101973A1PCT designated stage expired Publication Date: 2025-05-15CEDARS SINAI MEDICAL CENT +1

Patent Information

Application Number
PCT/US2024/055229
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2024-11-08
Publication Date
2025-05-15

AI Technical Summary

Technical Problem

Current methods for disease classification from medical images are time-consuming, require specialized expertise, and often lead to inconsistent diagnoses and treatment recommendations, failing to adequately incorporate patient preferences.

Method used

A system utilizing large language models (LLMs) to analyze medical images, integrate patient data, and generate disease classifications and treatment recommendations, while also considering patient preferences through retrieval-augmented generation techniques.

Benefits of technology

The system provides accurate, consistent, and objective disease classification and treatment recommendations, improving patient outcomes by integrating patient preferences and leveraging advanced retrieval-augmented generation techniques.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024055229_15052025_PF_FP_ABST
    Figure US2024055229_15052025_PF_FP_ABST
Patent Text Reader

Abstract

In some implementations, the system may include one or more processors. In addition, the system may include a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: store a plurality of data embeddings characterizing and corresponding to an external knowledge base; receive first input data, the first input data characterizing a first patient; receive one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics present within the one or more medical images; identify first portions of the knowledge base based on the first input data; and generate, based on at least the first portions of the knowledge base, the one or more medical images, and the corresponding prompt, the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR DISEASE CLASSIFICATION AND TREATMENT RECOMMENDATION USING LARGE LANGUAGE MODELSCROSS-REFERENCE TO RELATED APPLICATION

[0001] This Patent Application claims priority to U.S. Provisional Patent Application No. 63 / 597,485, filed on November 9, 2023, and entitled “Systems and Methods for Identifying and Classifying Disease Activity in Images Using Large Language Models.” The disclosure of the prior Application is considered part of and is incorporated by reference into this Patent Application.FIELD

[0002] The present invention relates generally to optimizing large language models for disease classification and treatment recommendation.BACKGROUND

[0003] The process of classifying disease characteristics from medical images requires specialized expertise and can be time-consuming. Even expert physicians with extensive training have difficulty consistently identifying and classifying medical conditions based on medical images. In many instances, even if expert physicians agree on the diagnosis of a medical condition based on an analysis of a medical image, physicians often disagree on the degree or severity of the medical condition. Further, failure to identify a medical condition or improperly characterizing the severity of an identified medical condition can lead to inappropriate treatment recommendations, adversely affecting patient outcomes.Furthermore, physicians providing care to patients often exhibit a lack of consistency in the degree to which patient preferences are incorporated into a recommended treatment modality, even if the medical condition is properly characterized in type and severity.

[0004] In summary, there is a critical need for systems equipped to assist physicians in the identification and classification of disease, as well as systems configured to generate suggested treatment modalities based not only on current health policy, but additionally influenced by patient preference in an objective, consistent, and repeatable manner.

[0005] Traditional approaches, such as artificial neural networks (ANNs) and large language models trained on additional domain-specific data, are limited by their dependance on supervised learning, which entails providing algorithms with sample labeled input to achieve a desired output. A key limitation to supervised learning is its dependence on large,labeled datasets, which can be costly and time-consuming. Furthermore, these methods do not allow for the integration of patient clinical data like symptoms, laboratory values and past medical history in addition to endoscopic images in order to allow for a more holistic clinical evaluation.

[0006] Embodiments consistent with the present disclosure are directed to these and other considerations.SUMMARY

[0007] A system of one or more computers can be configured to perform particular operations or actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by virtue of including instructions that, when executed by data processing apparatus, cause the apparatus to perform the actions.

[0008] In one general aspect, a system may include one or more processors. The system may also include a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: store a plurality of data embeddings characterizing and corresponding to an external knowledge base; receive first input data, the first input data characterizing a first patient; receive one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics present within the one or more medical images; identify first portions of the knowledge base based on the first input data; and generate, based on at least the first portions of the knowledge base, the one or more medical images, and the corresponding prompt, the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0009] In one general aspect, a system may include one or more processors. The system may also include a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: receive first input data having a plurality of medical images and a corresponding treatment modality and medical outcome for each of the plurality of medical images; receive a first prompt for generating a classification scheme for the plurality of medical images; andgenerate the classification scheme and group the plurality of images according to the classification scheme based at least on the first prompt, the first input data, the corresponding treatment modality, and the corresponding medical outcomes. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The disclosure, and its advantages and drawings, will be better understood from the following description of representative embodiments together with reference to the accompanying drawings. These drawings depict only representative embodiments, and are therefore not to be considered as limitations on the scope of the various embodiments or claims.

[0011] FIG. l is a block diagram of a system for classifying disease activity, according to certain aspects of the present disclosure.

[0012] FIG. 2 is a process for classifying disease activity, according to certain aspects of the present disclosure.

[0013] FIG. 3 is a process for generating disease classifications, according to certain aspects of the present disclosure.

[0014] FIG. 4A illustrates a word cloud including a variety of descriptors used by an Al model for describing positive findings suggesting colitis.

[0015] FIG. 4B illustrates a word cloud including a variety of descriptors used by an Al model for describing negative findings suggesting normal mucosa.DETAILED DESCRIPTION

[0016] The present disclosure relates to a system 100 for classifying diseases and generating treatment recommendations using large language models (LLMs). In some embodiments, implementations of the present disclosure capture images, photos, or videos of associated with gastrointestinal endoscopy. The captured images are analyzed using an artificial intelligence (Al) algorithm, e.g., a large language model (LLM). Prior to providing the captured images for analysis, an instruction is provided to the LLM. The instruction provides guidelines on how the LLM should provide disease activity information associated with the captured images. The LLM. For example, if one of the captured images is identified as being associated with a specific disease, the LLM can provide a disease severity scoreusing one of many accepted standardized scores (e.g., MES, etc.). In addition to (or in lieu of) characterizing the captured images by standardized scores, the LLM can describe the captured images to highlight the specific abnormalities in the captured images. In some cases, the LLM further provides a list of potential diagnoses to explain any abnormal findings in the captured images. In some embodiments, the output generated by the LLM is further enriched with retrieval augmentation techniques, as described herein.

[0017] Various embodiments are described with reference to the attached figures, where like reference numerals are used throughout the figures to designate similar or equivalent elements. The figures are not necessarily drawn to scale and are provided merely to illustrate aspects and features of the present disclosure. Numerous specific details, relationships, and methods are set forth to provide a full understanding of certain aspects and features of the present disclosure, although one having ordinary skill in the relevant art will recognize that these aspects and features can be practiced without one or more of the specific details, with other relationships, or with other methods. In some instances, well-known structures or operations are not shown in detail for illustrative purposes. The various embodiments disclosed herein are not necessarily limited by the illustrated ordering of acts or events, as some acts may occur in different orders and / or concurrently with other acts or events. Furthermore, not all illustrated acts or events are necessarily required to implement certain aspects and features of the present disclosure.

[0018] For purposes of the present detailed description, unless specifically disclaimed, and where appropriate, the singular includes the plural and vice versa. The word “including” means “including without limitation.” Moreover, words of approximation, such as “about,” “almost,” “substantially,” “approximately,” and the like, can be used herein to mean “at,” “near,” “nearly at,” “within 3-5% of,” “within acceptable manufacturing tolerances of,” or any logical combination thereof. Similarly, terms “vertical” or “horizontal” are intended to additionally include “within 3-5% of’ a vertical or horizontal orientation, respectively. Additionally, words of direction, such as “top,” “bottom,” “left,” “right,” “above,” and “below” are intended to relate to the equivalent direction as depicted in a reference illustration; as understood contextually from the object(s) or element(s) being referenced, such as from a commonly used position for the object(s) or element(s); or as otherwise described herein.

[0019] As shown in FIG. 1, The system 100 comprises several components, including a device 102, a server 110, a data store 130, an LLM repository 120, and various subcomponents that facilitate the processing and analysis of medical data.

[0020] The device 102 is used by a user 104 to input data into the system. This device can be any suitable computing device, such as a smartphone, tablet, or computer, capable of transmitting data to the server 110 and receiving inputs from user 104. Notably, device 102 may be configured to receive user input information in a number of different forms including written form and spoken communication. Further, in some embodiments, device 102 comprises logic and programming that enables it to transmit messages to user 104 in various forms, including written form and spoken form. In this regard, the device 102 can include various input / output components, such as a mouse, keyboard, graphical display, touchscreen, and the like. In some implementations, the device 102 includes an endoscope for capturing images to be provided to the server 110 and further includes a computer or a smartphone for providing instructions to the server 110. The identify of user 104 may vary across embodiments. For example, user 104 may be a patient seeking diagnosis and / or treatment. In some embodiments, user 104 may be a technician associated with the entity operating system 100. In some embodiments, user 104 may be a healthcare provider (e.g., a nurse, physician, hospital administrator) that is providing Al-assisted diagnosis and / or treatment recommendations to a patient.

[0021] The server 110 is a central component of the system 100 and is responsible for processing the input data and generating disease classifications and treatment recommendations. The server 110 includes an API 140, language logic 142, embedding logic 144, and a data store 150.

[0022] The API 140 serves as an interface to facilitate communication between the device 102 and the server 110, allowing for the transmission of data from the device 102 to the server 110. API 140 additionally allows for the exchange of data between server 110 and data store 150, as well as for the exchange of data between server 110 and LLM repository 120. The API 140 packages data packets to (and from) the server 110 and various other components of system 100, to facilitate a bidirectional information flow between the server 110 and other components of system 100, such as device 102. The API 140 can package information received from the device 104 so that these provided information can be processed by the server 110. In some implementations, the API 140 is a web service compatible with hypertext transfer protocol (HTTP) and machine-readable file formats such as extensible markup language (XML) and JavaScript object notation (JSON). The language logic 142 is responsible for processing natural language input from the user 104 and converting it into a format that can be analyzed by the system. Similarly, server 110 mayleverage language logic 142 to convert written communications into verbal communications, which may be provided to the user 104 via device 102.

[0023] The data store 130 is another component of the system 100, which may house a broad range of data types. Data store 130 includes EHR data 132, clinical data 134, policy data 136, and image data 138, but it should be understood that additional, or fewer data types can be included in data store 130. In operation, server 110 has access to the data contained within data store 130, which performs the role of an external knowledge base which improves the capability of the system to perform disease classification and treatment recommendation. EHR data 132 can include medical records associated with a patient or a plurality of patients, such as those generated by a hospital healthcare records system, that have consented to share medical information with the entity operating server 110. Clinical data 134 may include clinical treatment records of physicians practicing in a particular field of medicine. In this regard, clinical data 134 may include physician case notes, diagnoses, treatment recommendations, outcomes, etc. for a given patient or patient(s). Policy data 136 can include publications specifying current medical treatment policy for a variety of conditions and diseases. For example, in the field of gastroenterology, the American College of Gastroenterology publishes a set of guidelines with evidence-based recommendations and based practices for healthcare practitioners operating in the field of gastroenterology. Policy data 136 may maintain such guidelines for various fields of medicine. Image data 138 may include medical images annotated by an expert practitioner to identify disease characteristics. Image data may be of any suitable type, depending on the type of disease being classified by the system 100. For example, image data 138 may include endoscopic images, MRI images, CT scans, X-rays, and the like.

[0024] Notably, although data store 130 is shown as a single data store, it should be understood that system 100 can include a plurality of data stores, with disparate types of information being housed in distinct, separate data stores. The data store 130 can be housed at a separate location from the server 110 and / or owned by a different entity than the server 110.

[0025] Returning to server 110, it can include multiple computing devices, networked across different physical locations, for example, by using the Internet. In some implementations, computing device(s) can host a chat interface or can receive requests via application programming interfaces (e.g., API 140) for interacting with the large language model repository 120.

[0026] According to embodiments consistent with the present disclosure, the server 110 leverages embedding logic 144 to create and maintain embeddings (e.g., a machine-readable format) of the data store in data store 130. In some embodiments, embedding logic 144 converts the data stored in data store 130 into vectors which are stored as EHR embedding 152, clinical data embedding 154, policy data embedding 156, and image data embedding 158 within server 110. These embeddings can be understood as vector representations of the various types of data stored in data store 130, which allows server 110 to enhance, augment, or enrich the disease classification and treatment modality recommendations generated by the system, as described in more detail below with respect to processes 200 and 300.

[0027] Additionally, according to at least some embodiments, server 110 may be configured to intermittently update the stored embeddings 152, 154, 156, 158, etc. by performing routine API calls (via API 140) to data store 130. In such a way, embeddings stored in data store 150 accurately reflect the most up-to-date information within data store 130. The server 110 can make such API calls (e.g., using API 140) on a routine basis, for example, every hour, every day, every week, every month, etc.

[0028] The LLM repository 120 may contain multiple LLM models, denoted as LLM model 120-1 through LLM model 120-N. These models may be generally pre-trained large language models but without domain specific knowledge within the field of disease classification or treatment modality recommendation. It should be understood that server 110 can be configured to use various LLM models (e.g., LLM model 120-1) for a particular stage of processes 200 or 300 depending on the model suitable for a particular task. Each LLM model in the repository 120 can be specialized for different types of medical data or diseases, allowing the system 100 to provide precise and relevant recommendations. The LLM repository 120 can be housed at a separate location from the server 110 and / or owned by a different entity than the server 110. Further, although shown as a single repository, it should be understood that LLM repository 120 may comprise a single, or multiple distinct repositories each housing a respective LLM model 120. Example of large language models include any version of generative pretrained transformer (GPT), large language model meta Al (LLaMA), Google Gemini, Google pathways language model (PaLM), Microsoft Orca, etc.

[0029] At a high level, according to at least some embodiments consistent with the present disclosure, the user 104 inputs data into the device 102, which transmits the data to the server 110 via the API 140. The language logic 142 processes the natural language input, and the embedding logic 144 further identifies embeddings stored within data store 150 thatcan provide additional context, or otherwise augment the input data provided by user 104. The embedding logic 144 may further allow server 110 to identify a portion, or a number of portions, of data stored within the data store 130 that are relevant to the input received from the user 104 and may be used to enrich a response that the server 110 can generate using one of the available LLM models 120-N from LLM repository 120. The server 110 may generate a response using one of the given LLM models 120-N and enrich the response using the relevant data from the data store 150. Based on this analysis, the system 100 classifies the disease and generates treatment recommendations, which may be transmitted back to the device 102 for the user 104 to review.

[0030] In some embodiments, after receiving the disease classification and treatment recommendation (interchangeably referred to herein as a treatment modality), the user may provide further input to system 100 via interactions with device 102. In some embodiments, the further input can correspond to patient preferences with respect to treatment of disease. For example, server 110 may provide a disease diagnosis with a certain confidence score, and a plurality of alternative treatment recommendations, each with a corresponding confidence score. These LLM generated responses may guide a healthcare practitioner, such as treating physician, in diagnosing and providing treatment recommendations to a patient. The further input from user 104 can be provided to server 110, and embedding logic 144 may identify embeddings (152, 154, 156, 158, etc.) stored within data store 150 that can provide additional context, or otherwise augment the additional input data provided by user 104. The embedding logic 144 may further allow server 110 to identify a portion, or a number of portions, of data stored within the data store 130 that are relevant to the further input received from the user 104 and may be used to enrich a response that the server 110 can generate using one of the available LLM models 120-N from LLM repository 120 (e.g., selected portions of EHR data 132, clinical data 134, policy data 136, image data 138, etc.). In response, server 110 may provide a new treatment modality recommendation, or select one of the previously identified treatment modality recommendations based on disease identification, severity classification, and integrating patient preferences and further augmented by selected portions of data from data store 130. Notably, additional input may be received from user 104 in a variety of formats. For example, in natural language prompt response format via interaction with device 102, or input as text, etc.

[0031] In another embodiment, system 100 may be configured to generate disease classifications based on annotated medical images and known corresponding disease outcomes. For example, system 100 may receive a plurality of medical images at server 110.The annotated medical images may be accompanied by a corresponding treatment modality that is implemented to address the disease outcomes suggested by the medical image. Additionally, the input may include, for each annotated medical image, a medical outcome resulting from the corresponding treatment modality. The information may be received by server 110 from device 102 and / or from data store 130 via one or more API calls facilitated by API 140. Additionally, a prompt may be provided to server 110 that includes instructions to generate a classification scheme for the received medical images. In response, server 110 may perform operations consistent with the disclosed embodiments to generate a classification scheme and group the medical images according to the classification. In this way, server 110 may be able to generate classification of disease based on visual cues that a human may not be capable of identifying due to the biological limitations of the human eye.

[0032] LLMs (e.g., LLM models 120-1, . . . , 120-N) are used here as an example because LLMs have potential to offer Al support throughout healthcare, with a growing body of literature demonstrating versatility in the ability of Al systems to address clinical questions across multiple medical specialties. Although LLMs were originally designed for text-based natural language processing tasks, more recent iterations of LLM models, such as Generative Pre-Trained Transformer (GPT)-4, has recently expanded its functionalities to include image recognition, which is now termed GPT-4V(ision). Such image recognition tools may have the ability to address limitations discussed earlier in endoscopic evaluation of IBD. Although GPT-4V is generally trained, and its training is not domain specific for medical evaluation and / or endoscopic evaluation, GPT-4V can still be used to successfully classify disease activity from images. GPT-4V is provided as part of ChatGPT having a chat function that facilitates real-time, text and image-based interactions between users and the model. Each chat session operates independently, meaning that two separate chats are not related to each other in context or content. Although GPT-4V is used here as an example, other LLMs (e.g., Meta's Llama 2, Google's Med-PaLM 2 and Bard, Anthropic's Claude, etc.) can be augmented for classifying disease activity. Additionally, the ability of LLM to classify disease and provide treatment recommendations may be further enhanced with retrieval augmentation techniques, as described herein.

[0033] Thus, according to some embodiments, system 100 leverages advanced retrieval- augmented generation techniques to provide accurate and timely medical insights, improving patient outcomes and aiding healthcare professionals in their decision-making processes.

[0034] FIG. 2 is a flowchart of an example process 200. In some implementations, one or more process blocks of FIG. 2 may be performed by various components of system 100,including device 102, server 110, LLM repository 120, and data store 130. As shown in FIG. 2, process 200 may include storing a plurality of data embeddings characterizing and corresponding to an external knowledge base (block 202). In this regard, server 110 may, via API 140, make one or more calls to data store 130 to ingest one or more of EHR data 132, clinical data 134, policy data 136, and / or image data 138. Embedding logic 144 of server 110 may transform the information contained within the data store 130 into a machine-readable format. This may be accomplished by the embedding logic 144 generating data embeddings 152, 154, 156, and / or 158 corresponding to EHR data 132, clinical data 134, policy data 136, and image data 138. In at least some embodiments, embedding logic 144 vectorizes data from data store 130 and stores these vectors as the data embeddings 152, 154, 156, and 158. Notably, as the data in data store is modified, updated, etc., embedding logic 144 is configured to intermittently, on a predetermined regular interval, update, delete, modify, or add to the embeddings stored in data store 150. The regular intervals for updating the embeddings in data store 150 can be of any desired frequency, for example, every minute, every hour, every day, every week, every month, etc.

[0035] As also shown in FIG. 2, process 200 may include receiving first input data, the first input data characterizing a first patient (block 204). In this regard, first input data may include a unique identifier associated with a patient (e.g., a social security number, a patient ID, etc.) In some embodiments the first input data may include various other identifying characteristics of the patient, such as the patient’s weight, height, BMI, medication record, and the like. According to some embodiments first input data may be collected by the system 100 via a series of prompt-response sessions in which the server 110 uses LLM models from LLM repository 120 to cause device 102 to generate spoken prompts for user 104 to respond to. User 104 may provide responses in spoken form, and server 110 (e.g., via language logic 142 and embedding logic) may translate the spoken responses into written responses, and ultimately, into a machine-readable format. In other embodiments, first input data may be provided to system 100 in written form, and the written input may be translated into machine- readable format by embedding logic 144.

[0036] As further shown in FIG. 2, process 200 may include receiving one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics present within the one or more medical images (block 206). In this regard, the images may be directly sent by user 104 via device 102 to server 110 for processing. In another embodiment, the user 104 may provide the server 110 a pointer or link to a data store where the one or more medical images are available for processing.

[0037] As also shown in FIG. 2, process 200 may include identifying first portions of the knowledge base based on the first input data (block 208). In this regard, embedding logic 144 may be configured to convert the first input data into a machine-readable format. By using an appropriate method, embedding logic 144 can determine one or more portions of the knowledge base that is relevant to the first input data received in block 204. For example, embedding logic 144 can convert first input data into first input embedding data. By using an appropriate similarity metric, the first input embedding data can be compared to the embeddings 152, 154, 156, 158, etc. in order to identify portions of relevant domain knowledge present within data store 130. In some examples, the similarity metric can be based on cosine similarity, a Euclidean distance, a minimized loss function, and the like.

[0038] As further shown in FIG. 2, process 200 may include generating, based on at least the first portions of the knowledge base, the one or more medical images, and the corresponding prompt, the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics (block 210). In this regard, server 110 generates a response that is attributable to the implementation of one of LLM models contained in LLM repository 120. In addition, server 110 augments the response generated by one of LLM models contained in LLM repository 120 with the identified first portions of the knowledgebase (e.g., identified portions of data store 130). Thus, the response is enhanced with retrieval augmented techniques consistent with the disclosed embodiments.

[0039] Additionally, process 200 may optionally include providing the classification and one or more suggested treatment modalities to the first patient (block 212). As further shown in FIG. 2, process 200 may optionally include receiving from the first patient second input data characterizing patient treatment preferences (block 214). In response to receiving patient treatment preferences, process 200 may optionally include identifying a first suggested treatment modality of the one or more suggested treatment modalities based on at least the classification, the one or more suggested treatment modalities, and the second input data (block 216). In this regard, system 100 may undergo proceed with a similar process as described above with respect to block 210 to generate the first suggested treatment modality. As discussed above with respect to FIG. 1, server 110 may provide a new treatment modality recommendation, or select one of the previously identified treatment modality recommendations based on disease identification, severity classification, and integrating patient preferences and further augmented by selected portions of data from data store 130. Accordingly, aspects of process 200 effectively, objectively, and consistently is able tointegrate patient treatment preferences into a recommendation of a treatment modality in comparison to trained physicians, who struggle to consistently integrate patient preferences when providing treatment modality recommendations to patients. It should be noted that, according to some embodiments, server 110 may optionally perform another retrieval augmentation process by comparing portions of second input data to the stored embeddings 152, 154, 156, 158 etc. before generating the first suggested treatment modality in block 216.

[0040] Although FIG. 2 shows example blocks of process 200, in some implementations, process 200 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in FIG. 2. Additionally, or alternatively, two or more of the blocks of process 200 may be performed in parallel.

[0041] FIG. 3 is a flowchart of an example process 300. In some implementations, one or more process blocks of FIG. 3 may be performed by may be performed by various components of system 100, including device 102, server 110, LLM repository 120, and data store 130.

[0042] As shown in FIG. 3, process 300 may include receiving first input data having a plurality of medical images and a corresponding treatment modality and medical outcome for each of the plurality of medical images (block 302). First input data maybe received directly from a user 104 via device 102, or alternatively, server 110 may be provided a link or a pointer to an external knowledge base, such as data store 130, that contains the relevant medical images, treatment modalities (e.g., treatment plans) and medical outcomes of those treatment plans.

[0043] As also shown in FIG. 3, process 300 may include receiving a first prompt for generating a classification scheme for the plurality of medical images (block 304). In this context, a classification scheme may be understood as a method to categorize a given disease’s progression based on identifiable factors found within medical images.

[0044] As further shown in FIG. 3, process 300 may include generating the classification scheme and grouping the plurality of images according to the classification scheme based at least on the first prompt, the first input data, the corresponding treatment modality, and the corresponding medical outcomes (block 306). In this regard, the system may group together medical images that the server 110 determines, based on an application of a selected LLM model 120-N from LLM repository 120, that a given group of images are correlated with a similar severity level of disease. Notably, server 110 may accomplish this goal using an LLM model 120-N that is not specifically trained to accomplish this task.

[0045] As also shown in FIG. 3, process 300 may include the optional step of receiving second input data, the second input data characterizing a first patient (block 308). For example, after server 110 has generated a classification scheme using a given LLM model from LLM repository 120, the server 110 can apply the generated model to a novel input. In this regard, second input data can include at least a patient identifier, such that embedding logic 144 may convert the second input data into a machine readable format, and identify a relevant portion of data store 130 based on comparing the second input data to embeddings 152, 154, 156, and 158, etc., as described above with respect to process 200. Process 300 may optionally include receiving one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics may include with the classification scheme (block 310). Process 300 may also optionally include generating the classification of disease characteristics may include with the classification scheme (block 312).

[0046] Although FIG. 3 shows example blocks of process 300, in some implementations, process 300 may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those depicted in FIG. 3. Additionally, or alternatively, two or more of the blocks of process 300 may be performed in parallel.EXAMPLE

[0047] Endoscopic images were obtained from the comprehensive and multi-class image and video dataset for gastrointestinal endoscopy, HyperKvasir to test LLM model accuracy consistent with the present disclosure. HyperKvasir dataset is the largest collection of annotated endoscopic images and videos currently available, encompassing a wide variety of upper and lower gastrointestinal pathologies. The selected images were annotated by at least one experienced gastroenterologist and one or more qualified individuals in the medical field such as a junior doctor or Ph.D. student. The selected images were annotated for the presence and severity of ulcerative colitis, ranging in severity from Mayo Endoscopic Score (MES) 1 to MES 3. From these images, 100 images were randomly selected associated with each MES score, totaling 300 images. An additional 100 images that were classified as Boston Bowel Preparation Score (BBPS) 2-3 were selected to serve as normal colonic mucosa controls. A gastroenterology fellow and a practicing academic gastroenterologist evaluated the BBPS 2-3 images to confirm the absence of any endoscopic abnormalities.Data Collection

[0048] A total of 15 randomly selected images displaying UC MES 1-3 and 5 randomly selected images of normal colonic mucosa were selected for testing efficacy of LLM models.Images were provided to the LLM one by one with a corresponding prompt instructing the system to examine the image for evidence of normal and abnormal endoscopic findings and to produce a written assessment of its findings and identify if the image demonstrates any evidence of colitis. If the image contains evidence of colitis, the model was instructed to grade the colitis based on the MES system. Each image in the test set, along with the final optimized prompt, was independently provided to GPT-4 and GPT-4o via an API (e.g., API 140). Images were not altered prior to upload to API. One response was generated per image. Default model settings were used such as temperature and maximum token size.Statistical Analysis

[0049] Accuracy, sensitivity, and specificity were calculated for the binary classification of colitis (present versus absent), binary classification of colitis severity (normal / mild versus moderate / severe), and multiclass classification of colitis severity (MES score). The binary classification of colitis severity was selected based on previous studies. The chi-square test and Fisher exact test were performed to determine differences in performance. A 0.05 significance level was used throughout.ResultsColitis Detection

[0050] GPT-4o outperformed GPT-4 in detecting colitis and classifying colitis disease severity in endoscopic images. GPT-4o and GPT-4 correctly identified normal mucosa in 94% and 91% of images displaying normal mucosa controls (p<0.001), respectively. Overall, GPT-4o correctly assigned the MES in 273 / 300 (91%) endoscopic images of colitis, compared to 255 / 300 (85.0%) for GPT-4 (p<0.001). GPT-4o (MES 1 : 82%, MES 2: 91%, MES 3: 100%) and GPT-4 (MES 1 :76%, MES 2: 79%, MES 3: 100%) both had progressively better detection (presence / absence) of colitis with increasing disease severity (Table 1). The overall accuracy, specificity, and sensitivity for GPT-4o and GPT-4 in detecting were 91.8 %, 94%, 91.8%, and 85%, 91%, 86.5%, respectively (Table 2).Table 1: Comparison of the Performance of GPT-4o and GPT-4 in Detecting and Classifying the Severity of Colitis in Endoscopic Images Displaying Mayo Endoscopic Scores (MES) 1-3 and Normal Colonic Mucosa Controls.GPT-4o: Generative pre-trained transformer-4 (Omni)GTPT-4: Generative pre-trained transformer-4Disease Severity Classification

[0051] GPT-4o outperformed GPT-4 when classifying disease severity based on the MES. GPT-4o and GPT-4 correctly classified inactive / mild disease 72.5% and 70% of the time (p<0.001), respectively (Table 1). Moderate / severe disease was correctly classified by GPT-4o and GPT-4 82% and 70% of the time (p<0.001), respectively. The accuracy, sensitivity and specificity for disease severity classification for GPT-4o and GPT-4 were 76%, 82%, 70%, and 80.3%, 88%, 72.5%, respectively.Table 2: Accuracy, Sensitivity, and Specificity of GPT-4 and GPT-4o in Detecting and Classifying the Severity Colitis in Endoscopic Images Displaying Mayo Endoscopic Scores (MES) 1-3 and Normal Colonic Mucosa Controls.GPT-4o: Generative pre-trained transformer-4 (Omni)GTPT-4: Generative pre-trained transformer-4Qualitative Analysis of GPT-4O Output

[0052] The model described an array of findings associated with colitis from the endoscopic images, with examples of complete model outputs shown in Table 3. For example, the model used descriptors such as “The mucosa appears erythematous (red and inflamed) extensively...”, “.. .areas of friability as seen by the presence of small mucosal bleeding points ...”, “.. .vascular pattern is obliterated due to the inflammation...” and “...yellowish, mucopurulent exudate coating several area..”. The model also described pertinent negatives when examining normal mucosa, using descriptors such as “.. .visible vascular pattern is intact, and there are no abnormal growths, strictures, or other abnormalities.” and “. . .no noticeable swelling or edema in the mucosal lining .”. In each output, the model appeared to systematically evaluate and describe the colonic mucosa, much like a gastroenterologist. FIG. 4A shows an example word cloud associated with a positive finding of colitis, and FIG. 4B shows an example word cloud associated with an absence of colitis.Table 3: Examples of GPT-4o Outputs when Assessing Endoscopic Images.

[0053] Although the example provided above deals with IBD, the application of Al-aided image recognition using some embodiments of the present disclosure extends beyond IBD. For example, embodiments of the present disclosure can be used in lesion detection during colonoscopy. Embodiments of the present disclosure can be used to detect polyps in captured images. Embodiments of the present disclosure can be used in identifying normal and abnormal findings on capsule endoscopy. Capsule endoscopy has a critical shortage of physicians that can read images well, therefore, LLMs can be used with images, tailored prompts, and augmented generation retrieval to mitigate effects associated with expertise shortage in capsule endoscopy.

[0054] The disclosed embodiments may be implemented at least according to the following clauses:

[0055] 1 : A system, may include: one or more processors; and a non-transitory computer- readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: store a plurality of data embeddings characterizing and corresponding to an external knowledge base; receive first input data, the first input data characterizing a first patient; receive one or more medical images associatedwith the first patient and a corresponding prompt requesting a classification of disease characteristics present within the one or more medical images; identify first portions of the knowledge base based on the first input data; and generate, based on at least the first portions of the knowledge base, the one or more medical images, and the corresponding prompt, the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics.

[0056] 2 : The system as paragraph 1 describes, where the non-transitory computer- readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: provide the classification and one or more suggested treatment modalities to the first patient; receive, from the first patient, second input data characterizing patient treatment preferences; and based on at least the classification, the one or more suggested treatment modalities, the second input data, identify a first suggested treatment modality of the one or more suggested treatment modalities.

[0057] 3 : The system as either of paragraphs 1 or 2 describe, where the non-transitory computer-readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: identify second portions of the knowledge base based on the second input data; and where identifying the first suggested treatment modality is further based on the second portions of the knowledgebase.

[0058] 4 : The system as any of paragraphs 1-3 describe, where receiving input data may include: generating one or more natural language prompts; and receiving, in response to the one or more natural language prompts, a natural language response from the first patient.

[0059] 5 : The system as any of paragraphs 1-4 describe, where the first portions of the knowledge base are selected from patient electronic health record data, clinical data, policy data, image data, and combinations thereof.

[0060] 6 : The system as any of paragraphs 1-5 describe, where the policy data may include standard medical policy guidelines for treatment of disease.

[0061] 7 : The system as any of paragraphs 1-6 describe, where the clinical data may include one or more clinical records of a physician that is trained to treat a disease associated with the classification of disease characteristics.

[0062] 8 : The system as any of paragraphs 1-7 describe, where the image data may include medical images associated with the classification of disease characteristics.

[0063] 9 : The system as any of paragraphs 1-8 describe, where each of the classification and the one or more suggested treatment modalities are associated with a probabilistic confidence score.

[0064] 10: The system as any of paragraphs 1-9 describe, where the one or more processors are configured to implement a trained large language model to generate the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics.

[0065] 11 : The system as paragraph 10 describes, where the large language model is selected from GPT-4, GPT-4o, and combinations thereof.

[0066] 12: The system as any of paragraphs 1-11 describe, where the classification of disease characteristics may include a classification of colitis based on a metric.

[0067] 13: The system as any of paragraphs 1-12 describe, where the metric may include a Mayo Endoscopic Score (MES).

[0068] 14: The system as any of paragraphs 1-13 describe, where the classification of disease characteristics may include (i) the MES, (ii) qualitative disease severity, (iii) potential diagnosis, (iv) descriptors for positive findings, (v) descriptors for negative findings, (vi) further recommendations for checking clinical history, (vii) further recommendations for histopathological examination of biopsied tissue, (viii) further recommendations for a followup with healthcare provider, and (ix) any combination thereof.

[0069] 15: A system may include: one or more processors; and a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: receive first input data may include a plurality of medical images and a corresponding treatment modality and medical outcome for each of the plurality of medical images; receive a first prompt for generating a classification scheme for the plurality of medical images; and generate the classification scheme and group the plurality of medical images according to the classification scheme based at least on the first prompt, the first input data, the corresponding treatment modalities, and the corresponding medical outcomes.

[0070] 16: The system as paragraph 15 describes, where the non-transitory computer- readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: store a plurality of data embeddings characterizing and corresponding to an external knowledge base; identify first portions of the knowledge base based on the first input data or the first prompt; and where generating the classification scheme is further based on the identified first portions of the knowledge base.

[0071] 17: The system as either of paragraphs 15 or 16 describe, where the plurality of medical images may include endoscopic medical images.

[0072] 18: The system as any of paragraphs 15-17 describe, where the generated classification scheme is predictive of a medical outcome, an optimal treatment modality, or combinations thereof.

[0073] 19: The system as any of paragraphs 15-18 describe, where the non-transitory computer-readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: receive second input data, the second input data characterizing a first patient; receive one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics may include with the classification scheme; and generate the classification of disease characteristics may include with the classification scheme.

[0074] 20: The system as any of paragraphs 15-19 describe, where the classification scheme may include a model for evaluating severity of colitis.

[0075] The foregoing disclosure provides illustration and description but is not intended to be exhaustive or to limit the implementations to the precise form disclosed. Modifications may be made in light of the above disclosure or may be acquired from practice of the implementations. As used herein, the term “component” is intended to be broadly construed as hardware, firmware, or a combination of hardware and software. It will be apparent that systems and / or methods described herein may be implemented in different forms of hardware, firmware, and / or a combination of hardware and software. The actual specialized control hardware or software code used to implement these systems and / or methods is not limiting of the implementations. Thus, the operation and behavior of the systems and / or methods are described herein without reference to specific software code - it being understood that software and hardware can be used to implement the systems and / or methods based on the description herein. As used herein, satisfying a threshold may, depending on the context, refer to a value being greater than the threshold, greater than or equal to the threshold, less than the threshold, less than or equal to the threshold, equal to the threshold, and / or the like, depending on the context. Although particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of various implementations. In fact, many of these features may be combined in ways not specifically recited in the claims and / or disclosed in the specification.

[0076] Although each dependent claim listed below may directly depend on only one claim, the disclosure of various implementations includes each dependent claim in combination with every other claim in the claim set. No element, act, or instruction used herein should be construed as critical or essential unless explicitly described as such. Also,as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Further, as used herein, the article “the” is intended to include one or more items referenced in connection with “the article “the” and may be used interchangeably with “the one or more.” Furthermore, as used herein, the term “set” is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, and / or the like), and may be used interchangeably with “one or more.” Where only one item is intended, the phrase “only one” or similar language is used. Also, as used herein, the terms “has,” “have,” “having,” or the like are intended to be open-ended terms. Further, the phrase “based on” is intended to mean “based, at least in part, on” unless explicitly stated otherwise. Also, as used herein, the term “or” is intended to be inclusive when used in a series and may be used interchangeably with “and / or,” unless explicitly stated otherwise (e.g., if used in combination with “either” or “only one of’).

Claims

CLAIMSWHAT IS CLAIMED IS:

1. A system, comprising: one or more processors; and a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: store a plurality of data embeddings characterizing and corresponding to an external knowledge base; receive first input data, the first input data characterizing a first patient; receive one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics present within the one or more medical images; identify first portions of the knowledge base based on the first input data; and generate, based on at least the first portions of the knowledge base, the one or more medical images, and the corresponding prompt, the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics.

2. The system of claim 1, wherein the non-transitory computer-readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: provide the classification and one or more suggested treatment modalities to the first patient; receive, from the first patient, second input data characterizing patient treatment preferences; and based on at least the classification, the one or more suggested treatment modalities, the second input data, identify a first suggested treatment modality of the one or more suggested treatment modalities.

3. The system of claim 1, wherein receiving input data comprises: generating one or more natural language prompts; and receiving, in response to the one or more natural language prompts, a natural language response from the first patient.

4. The system of claim 1, wherein the first portions of the knowledge base are selected from patient electronic health record data, clinical data, policy data, image data, and combinations thereof.

5. The system of claim 4, wherein the policy data comprises standard medical policy guidelines for treatment of disease.

6. The system of claim 4, wherein the clinical data comprises one or more clinical records of a physician that is trained to treat a disease associated with the classification of disease characteristics.

7. The system of claim 4, wherein the image data comprises medical images associated with the classification of disease characteristics.

8. The system of claim 1, wherein each of the classification and the one or more suggested treatment modalities are associated with a probabilistic confidence score.

9. The system of claim 2, wherein the non-transitory computer-readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: identify second portions of the knowledge base based on the second input data; and wherein identifying the first suggested treatment modality is further based on the second portions of the knowledge base.

10. The system of claim 1, wherein the one or more processors are configured to implement a trained large language model to generate the classification of disease characteristics and one or more suggested treatment modalities based on the classification of disease characteristics.

11. The system of claim 10, wherein the large language model is selected from GPT-4, GPT-4o, and combinations thereof.

12. The system of claim 1, wherein the classification of disease characteristics comprises a classification of colitis based on a metric.

13. The system of claim 12, wherein the metric comprises a Mayo Endoscopic Score (MES).

14. The system of claim 13, wherein the classification of disease characteristics comprises (i) the MES, (ii) qualitative disease severity, (iii) potential diagnosis, (iv) descriptors for positive findings, (v) descriptors for negative findings, (vi) further recommendations for checking clinical history, (vii) further recommendations for histopathological examination of biopsi ed tissue, (viii) further recommendations for a follow-up with healthcare provider, and (ix) any combination thereof.

15. A system comprising: one or more processors; and a non-transitory computer-readable storage medium containing instructions that, when executed by the one or more processors, are configured to cause the system to: receive first input data comprising a plurality of medical images and a corresponding treatment modality and medical outcome for each of the plurality of medical images; receive a first prompt for generating a classification scheme for the plurality of medical images; and generate the classification scheme and group the plurality of medical images according to the classification scheme based at least on the first prompt, the first input data, the corresponding treatment modalities, and the corresponding medical outcomes.

16. The system of claim 15, wherein the non-transitory computer-readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: store a plurality of data embeddings characterizing and corresponding to an external knowledge base; identify first portions of the knowledge base based on the first input data or the first prompt; and wherein generating the classification scheme is further based on the identified first portions of the knowledge base.

17. The system of claim 15, wherein the plurality of medical images comprise endoscopic medical images.

18. The system of claim 15, wherein the generated classification scheme is predictive of a medical outcome, an optimal treatment modality, or combinations thereof.

19. The system of claim 15, wherein the non-transitory computer-readable storage medium contains instructions that, when executed by the one or more processors, are configured to cause the system to: receive second input data, the second input data characterizing a first patient; receive one or more medical images associated with the first patient and a corresponding prompt requesting a classification of disease characteristics consistent with the classification scheme; and generate the classification of disease characteristics consistent with the classification scheme.

20. The system of claim 15, wherein the classification scheme comprises a model for evaluating severity of colitis.

Citation Information

Patent Citations

  • Systems and methods for point of care guidance

    US20150242580A1

  • PREDICTING OUTCOME OF TREATMENT WITH AN ANTI-alpha4beta7 INTEGRIN ANTIBODY

    US20200155673A1

  • Aligning image data of a patient with actual views of the patient using an optical code affixed to the patient

    US20210057080A1

  • Method and system for machine learning classification based on structure or material segmentation in an image

    US20220147757A1

  • System for transcribing and performing analysis on patient data

    US20230210610A1

Cited By

  • Method and system for generating diagnosis and treatment pathways based on patient case twins

    US20260120869A1