Techniques for optimizing summary generation using generative artificial intelligence models

A computing system using trained generative AI models optimizes HPI summaries and outcome predictions by integrating subspecialty data and multimodal analysis, addressing inefficiencies and enhancing clinical accuracy and efficiency in healthcare.

WO2025145006A1PCT designated stage expired Publication Date: 2025-07-03MAYO FOUNDATION FOR MEDICAL EDUCATION & RESEARCH

Patent Information

Application Number
PCT/US2024/062059
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-27
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing healthcare systems face challenges in efficiently generating accurate and comprehensive History of Present Illness (HPI) summaries and outcome predictions, relying heavily on manual input and review of medical records, which is labor-intensive and prone to errors, leading to inefficiencies and clinician burnout.

Method used

A computing system utilizing trained generative AI models to aggregate and analyze patient data, including subspecialty information, prior inputs, and test values to generate optimized HPI summaries and outcome predictions, such as differential diagnoses and potential complications, by integrating foundation and multimodal prediction models for enhanced accuracy and efficiency.

Benefits of technology

The system significantly reduces manual tasks, enhances predictive capabilities, and improves clinical decision-making by providing timely and accurate HPI summaries and outcome predictions, thereby reducing clinician workload and improving patient care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024062059_03072025_PF_FP_ABST
    Figure US2024062059_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Techniques for optimizing summary generation using generative artificial intelligence (Al) models are disclosed herein. An example computing system comprises: one or more processors; and one or more memories. The one or more memories have stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user; generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, and determining the output based on the input, the subspeciality of the provider, the prior input, and the test value; and cause the HPI summary to be displayed at a device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNIQUES FOR OPTIMIZING SUMMARY GENERATION USING GENERATIVE ARTIFICIAL INTELLIGENCE MODELSCROSS-REFERENCE TO RELATED APPLICATION[1] This application claims priority to U.S. Provisional Patent Application No. 63 / 615,439, entitled “Techniques for Optimizing Summary Generation Using Generative Artificial Intelligence Models,” filed on December, 28 2023, the disclosure of which is hereby incorporated herein by reference.TECHNICAL FIELD[2] The present disclosure is generally directed to methods and systems for methods and systems for optimizing summary generation, and more particularly, to techniques for training and operating one or more generative artificial intelligence (Al) models to optimize summary and outcome prediction generation.BACKGROUND[3] In modern healthcare systems, optimizing and streamlining patient care using innovative tools has become paramount. Such tools will ideally reduce clerical burden, thereby enabling physicians and other caregivers to focus on patients and enhance human contact. Moreover, these tools should reduce clerical burden and burnout and learn from interactions with healthcare professionals to become increasingly useful.[4] A key piece of providing high-quality patient care is the History of Present Illness (HPI) . Broadly, the HPI details a patient's primary symptoms and provides a description of the development of the patient’s present illness. In particular, the HPI is usually a chronological description of the progression of the patient's present illness from the first sign and symptom to the present. However, HPI creation is often labor-intensive, relying heavily on human input and review of medical records for specific, specialty-specific facts and details.[5] Therefore, there is an opportunity for improved techniques / platforms for generating summaries (e.g., an HPI), outcome predictions, and other data using generative AI / ML models, capable of delivering more timely, accurate, and comprehensive outcomes for healthcare providers and patients.BRIEF SUMMARY[6] In some aspects, the techniques described herein relate to a computing system for optimizing summary generation using generative artificial intelligence (Al) models, including: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, and determining the output based on the input, the subspeciality of the provider, the prior input, and the test value, and cause the HPI summary to be displayed at a device.[7] In some aspects, the techniques described herein relate to a computing system, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to determine the output by: selectively including a subset of data from one or more of the input, the prior input, or the test value in the HPI summary based on the subspeciality of the provider.[8] In some aspects, the techniques described herein relate to a computing system, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: generate, by executing the trained ML model, an outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.[9] In some aspects, the techniques described herein relate to a computing system, wherein the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value.

[0010] In some aspects, the techniques described herein relate to a computing system, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: receive a subsequent input related to the output from a subsequent provider; generate, by executing the trained ML model, a subsequent output including a subsequent HPI summary by: identifying a subsequent subspeciality of the subsequent provider, wherein the subsequent subspeciality is different from the subspeciality of the provider, aggregating (i) the input and (ii) the output, and determining the subsequent HPI summary based on the subsequent input, the subsequent subspeciality of the subsequent provider, the input, and the output; and cause the subsequent HPI summary to be displayed at a subsequent device.

[0011] In some aspects, the techniques described herein relate to a computing system, wherein the input includes a subspeciality request indicating a second subspeciality that is different fromthe subspeciality of the provider, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to generate the output by: determining, by executing the trained ML model, a subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the subset of data.

[0012] In some aspects, the techniques described herein relate to a computing system, wherein the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an HPI source and (b) a multi-modal ML model.

[0013] In some aspects, the techniques described herein relate to a computing system, wherein the first data further includes an X-ray image associated with the user, and the trained ML model is a foundation model configured to: receive X-ray images of users; and determine image features of the X-ray images.

[0014] In some aspects, the techniques described herein relate to a computing system, further including a prediction model that is configured to: receive the image features from the foundation model; and determine a predicted metric value based on the image features.

[0015] In some aspects, the techniques described herein relate to a computing system, wherein the prediction model is further configured to generate a predicted multimedia output based on the image features that correspond to the predicted value.

[0016] In some aspects, the techniques described herein relate to a computing system, wherein the prediction model is a multimedia model that is trained using historical multimedia data of a plurality of patients and historical metric data of the plurality of patients.

[0017] In some aspects, the techniques described herein relate to a computing system, wherein: the foundation model is configured to receive one or more of: (i) X-ray images, (ii) computed tomography (CT) scans, (iii) magnetic resonance imaging (MRI) images, (iv) ultrasound images, (v) positron emission tomography (PET) scans, (vi) single photon emission computed tomography (SPECT) scans, (vii) optical coherence tomography (OCT) scans, or (viii) digital pathology slide; and the image features are associated with (i) an ejection fraction of the user, (ii) a cardiac output of the user, (iii) a ventricular volume of the user, (iv) a wall thickness of the user, (v) a motion abnormality of the user, (vi) a valvular function of the user, (vi) a cardiac morphology of the user, (vii) a blood flow pattern of the user, and (viii) a blood flow velocity of the user.

[0018] In some aspects, the techniques described herein relate to a computing system for optimizing summary generation using generative artificial intelligence (Al) models, including: one or more processors; and one or more memories having stored thereon computerexecutable instructions that, when executed by the one or more processors, cause the computing system to: receive an input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including an outcome prediction for the user by: aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, determining the output based on the input, the prior input, and the test value, and wherein the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value, and cause the outcome prediction to be displayed at a device.

[0019] In some aspects, the techniques described herein relate to a computing system, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: receive a subsequent input related to the output; generate, by executing the trained ML model, a subsequent output including a subsequent outcome prediction by: aggregating (i) the input and (ii) the output, determining the subsequent output based on the subsequent input, the input, and the output, and wherein the subsequent outcome prediction includes one or more of: (i) a subsequent differential diagnosis, (ii) a subsequent potential complication, (iii) a subsequent operative value, (iv) a subsequent length of stay value, or (v) a subsequent rehospitalization value; and cause the subsequent outcome prediction to be displayed at a subsequent device.

[0020] In some aspects, the techniques described herein relate to a computing system, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: aggregate a set of outcome predictions for a plurality of users at a first location; generate, by executing the trained ML model, a capacity prediction for the first location based on the set of outcome predictions; and cause the capacity prediction to be displayed at the device.

[0021] In some aspects, the techniques described herein relate to a computing system, wherein the input is received from a provider with a subspeciality, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to determine the output by: selectively including a subset of data from one or more of the input, the prior input, or the test value in theoutput based on the subspeciality of the provider, and wherein the output includes a history of present illness (HP I) summary based on the subset of data.

[0022] In some aspects, the techniques described herein relate to a computing system, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: generate, by executing the trained ML model, the outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.

[0023] In some aspects, the techniques described herein relate to a computing system, wherein the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to generate the output by: determining, by executing the trained ML model, a second subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the second subset of data.

[0024] In some aspects, the techniques described herein relate to a computing system, wherein the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an outcome source and (b) a multi-modal ML model.

[0025] In some aspects, the techniques described herein relate to a computing system for optimizing summary generation using generative artificial intelligence (Al) models, including: one or more processors; and one or more memories having stored thereon computerexecutable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary and an outcome prediction by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, determining the HPI summary based on the input, the subspeciality of the provider, the prior input, and the test value, and determining the outcome prediction based on the HPI summary, and cause the HPI summary and the outcome prediction to be displayed at a device.BRIEF DESCRIPTION OF THE FIGURES

[0026] The figures described below depict various aspects of the system and methods disclosed therein. It should be understood that each figure depicts one aspect of a particular aspect of the disclosed system and methods, and that each of the figures is intended to accord with a possible aspect thereof. Further, wherever possible, the following description refers to the reference numerals included in the following figures, in which features depicted in multiple figures are designated with consistent reference numerals.

[0027] FIG. 1 depicts an example computing environment in which the techniques disclosed herein may be implemented, in accordance with various embodiments herein.

[0028] FIG. 2 depicts an example model-controller configuration for generating summaries, outcome predictions, and other data for providers with various subspecialities, in accordance with various embodiments herein.

[0029] FIGS. 3A-3C depict example computer-implemented methods for optimizing summary generation using trained AI / ML models, in accordance with various embodiments herein.

[0030] FIG. 4 illustrates an example user interface configured to display and provide interactive access to the generated summaries and other data output by the trained AI / ML models, in accordance with various embodiments herein.

[0031] FIG. 5 depicts an example foundation model and multimodal prediction model workflow for generating predicted metric values and / or predicted multimedia outputs, in accordance with various embodiments herein.

[0032] The figures depict preferred aspects for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative aspects of the systems and methods illustrated herein may be employed without departing from the principles of the invention described herein.DETAILED DESCRIPTIONOverview

[0033] The present techniques provide methods and systems for, inter alia, optimizing summary and outcome prediction generation, and more particularly, to techniques for training and operating one or more AI / ML models to optimize such summary / outcome prediction generation using generative AI / ML models (e.g., in a clinical setting). Specifically, the present techniques may include technologies for training one or more ML models e.g., one or more language models, one or more artificial neural networks, etc.) to approximate human decision makers, forexample, using a corpus of history of present illness ( H P I) data, written history and physical (H&P) data, electronic health record (EHR) data, correspondence, and / or other metadata. Once trained, these models may generate outputs in response to user prompts related to healthcare data contained in, for example, an HPI corresponding to a particular patient. In some embodiments, the output may include one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value in a natural language, textual format. Further, these outputs may be customized for different specialties (e.g., structural cardiologist, electrophysiologist, neurologist, etc.) in some embodiments.

[0034] The present techniques may further include one or more client computing devices, or applications that may be installed on client computing devices, which enable users to access one or more trained models by providing input prompts that are processed by one or more trained models. These prompts may be propagated to the trained models for processing. The client computing device may receive responses from these models, and convey those responses to the user, optionally after performing pre-processing and / or post-processing steps.

[0035] As discussed herein, the systems and methods of the present disclosure may seamlessly integrate with an electronic medical record system, and thereby reduce / eliminate manual tasks by automatically generating HPI summaries, H&P documents, output predictions, and / or other data. More specifically, the natural language interface of the summary and prediction module described herein, powered by LLMs and / or other ML models, offers a unique capability to, for example, automatically generate HPIs specifically configured for subspecialty practices, generate differential diagnoses and / or otherwise create germane predictive capabilities that exceed expert human capabilities, and / or other similar functions. As a result, the summary and prediction module described herein may obviate certain clinical and administrative tasks, which are often burdensome resulting in operational inefficiencies, user burnout and increased fatigue, and potential human errors through manual, time-consuming tasks

[0036] Broadly, the summary and prediction module described herein may utilize an intelligent orchestration layer which may execute and / or otherwise leverage LLMs trained to answer and / or otherwise respond to prompts in a manner that most benefits the specific provider interacting with the module. To illustrate, the summary and prediction module may leverage patient notes and history (e.g., prior HPI, H&P), test values (e.g., from laboratory tests), and / or output from other Al models (e.g., Al configured to analyze ECGs) and offer users an option to input new symptoms and details from a current encounter with a patient. The ML models of the summary and prediction module may then process this data to generate a holistic HPI, H&P,predictive information (e.g., outcome prediction(s)), and / or any other suitable data or combinations thereof. Further, these ML models may be trained to generate outputs tailored to provide perspective from ( / .e., emulate) and / or for different specialties (e.g., structural cardiologist, electrophysiologist, neurologist, etc.).

[0037] As mentioned, the summary and prediction module described herein may further assist busy clinicians by utilize medical data, such as summaries (e.g., HPIs), lab results, imaging studies, and / or other data to predict various medical parameters. Namely, the ML models described herein may analyze input data to generate outcome predictions indicating a predicted length of stay, mortality, a differential diagnosis, operative risks, potential complications, rehospitalization probability, and / or other vital patient outcome parameters that are relevant to a clinician / patient. These outputs may include natural language text, tables, annotated images, annotated video, etc. Moreover, the summary and prediction module may efficiently generate a unified user interface configured to seamlessly integrate the predictive insights of the ML models and consequently offer clinicians an intuitive interactive interface.

[0038] Further, existing techniques often suffer from several drawbacks that limit their effectiveness in diagnosing and predicting patient outcomes, such as a lack of accuracy, efficiency, and predictive power. These drawbacks may be attributed to several factors. For example, existing methods may rely on manual features or machine learning models that cannot capture the complex patterns present in medical images. This can lead to inaccuracies in identifying and diagnosing conditions, as these methods frequently miss subtle but clinically significant features. Moreover, processing and analyzing medical imaging data using existing techniques is often time-consuming, as such techniques typically require significant manual intervention or are not optimized / capable of handling the large volumes of data generated in modern medical practices, thereby resulting in inefficiencies that can delay diagnoses and treatment planning, impacting patient care. Additionally, existing techniques are frequently constrained to analyzing images in isolation, without the ability to integrate data from multiple modalities or generate predictive simulations. This limits the depth of analysis and reduces the predictive power of the system, making it difficult to forecast patient outcomes or simulate the progression of diseases. Additionally, the same or similar information can often be expressed in equivalent formats e.g., using different units.

[0039] By contrast, the present techniques overcome these (among other) challenges by utilizing a foundation model and a multimodal prediction model. Namely, these models address these drawbacks associated with existing techniques and significantly improve the functioning of a computing device / system in several ways. The foundation model, trained on a vast array ofmedical imaging data, can extract nuanced features with high precision. When this detailed feature data is fed into the multimodal prediction model, the system can leverage deep learning algorithms to make highly accurate predictions. Such a dual-model approach captures complex patterns and relationships in the data that existing methods often miss, leading to more accurate diagnoses and assessments.

[0040] Further, by automating the feature extraction and prediction processes, this configuration significantly reduces the time required to analyze medical images. The foundation model efficiently processes diverse imaging data to extract relevant features, while the multimodal prediction model quickly generates predictions and / or simulations (e.g., predicted multimedia outputs) based on these features. Such a streamlined process minimizes manual intervention and accelerates the analysis even relative to existing techniques that leverage machine learning techniques, enabling faster decision-making in clinical settings. Additionally, the multimodal prediction model's ability to generate metric values (e.g., values associated with clinically relevant conditions and / or diagnostic operations, such as ejection fraction, cardiac output, ventricular volume, wall thickness, motion abnormality, valvular function, cardiac morphology, blood flow pattern, and / or blood flow velocity) and complex outputs like videos or images based on the extracted features greatly enhances the system's predictive capabilities relative to existing techniques. This allows for dynamic simulations of patient conditions, offering insights into disease progression and potential outcomes that were not possible with existing techniques. Such predictive simulations can aid in personalized treatment planning and patient monitoring, providing a more comprehensive understanding of patient health.

[0041] Thus, the integration of a foundation model and a multimodal prediction model overcomes the limitations of existing techniques by offering improved accuracy, efficiency, and predictive power. This advanced configuration enables a more nuanced and comprehensive analysis of medical imaging data, leading to better-informed clinical decisions and ultimately improving patient care outcomes.

[0042] In accordance with the above, and with the disclosure herein, the present disclosure includes improvements in computer functionality or in improvements to other technologies at least because the disclosure describes that, e.g., a hosting server (e.g., server computing device), or otherwise computing device (e.g., a client computing device), is improved where the intelligence or predictive ability of the hosting server or computing device is enhanced by a trained machine learning model. This model, executing on the hosting server or user computing device, is able to accurately and efficiently generate highly curated summaries of patient medical data in response to user queries / prompts (e.g., from a medical / healthcare provider).That is, the present disclosure describes improvements in the functioning of the computer itself or “any other technology or technical field” because a hosting server or user computing device, is enhanced with a trained machine learning model to accurately detect, evaluate, predict, and generate HPI summaries, outcome predictions, and other data configured to improve a provider’s diagnostic / treatment efforts. This improves over the prior art at least because existing systems lack such evaluative and / or predictive functionality and are generally unable to accurately analyze such medical data on a real-time basis to output predictive and / or otherwise recommended HPI summaries and / or outcome predictions designed to improve a provider's overall diagnostic / treatment efforts.

[0043] As mentioned, the model(s) may be trained using machine learning and may utilize machine learning during operation. Therefore, in these instances, the techniques of the present disclosure may further include improvements in computer functionality or in improvements to other technologies at least because the disclosure describes such models being trained with a plurality of training data {e.g., 10,000s of training data corresponding to HPIs, outcome predictions, user queries / prompts, output / response prompts, etc.) to generate and output the relevant prompt responses configured to improve the provider’s diagnostic / treatment efforts.

[0044] Moreover, the present disclosure includes effecting a transformation or reduction of a particular article to a different state or thing, e.g., transforming or reducing the interaction processing demand of an medical records system (and associated subsystems / components / devices) and clinician occupation times evaluating such medical records {e.g., HPIs) from a non-optimal or error state to an optimal state by automatically retrieving, analyzing, and summarizing such medical records based on provider-specific subspecialties and generating outcome predictions from HPIs and other data while eliminating untimely / erroneous clinician searching through such medical records systems.

[0045] Still further, the present disclosure includes specific features other than what is well- understood, routine, conventional activity in the field, or adding unconventional operations that demonstrate, in various embodiments, particular useful applications, e.g., receiving an input from a provider (e.g., via user query / prompt or other input, directly from the provider or from a provider computer terminal communicating the input), the input including first data associated with a user, generating, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, and determining the output based on the input, thesubspeciality of the provider, the prior input, and the test value, and / or causing the HPI summary to be displayed at a device, among others.Example Computing Environment

[0046] FIG. 1 depicts an example computing environment 100 in which the techniques disclosed herein may be implemented, according to some aspects. The environment 100 may include computing resources for training and / or operating machine learning models to perform optimized summary generation, outcome prediction generation, and / or other data evaluation, prediction, and / or generation, in some aspects.

[0047] The computing environment 100 may include a client computing device 102, a server computing device 104, an electronic network 106, and a model database 112. The computing environment may further include one or more cloud application programming interfaces (APIs) 1 14. The components of the computing environment 100 may be communicatively connected to one another via the electronic network 106, in some aspects.

[0048] The client computing device 102 may be implemented as one or more computing devices (e.g., one or more servers, one or more laptops, one or more mobile computing devices, one or more tablets, one or more wearable devices, one or more cloud-computing virtual instances, etc.). In some aspects, a plurality of client computing devices may be part of the environment100 - for example, a first user may access a client computing device 102 that is a laptop, while a second user accesses the client computing device 102 that is a smart phone, while yet a third user accesses a client computing device 102 that is a wearable device. Each of these respective users may access the AI / ML-based services via their use of their respective client computing device 102.

[0049] The client computing device 102 may include one or more processors 120, one or more network interface controllers 122, one or more memories 124, an input device 126, an output device 128 and a client API 130. The one or more memories 124 may have stored thereon one or more modules 140 (e.g., one or more sets of instructions).

[0050] In some aspects, the one or more processors 120 may include one or more central processing units, one or more graphics processing units, one or more field-programmable gate arrays, one or more application-specific integrated circuits, one or more tensor processing units, one or more digital signal processors, one or more neural processing units, one or more RISC-V processors, one or more coprocessors, one or more specialized processors / accelerators for artificial intelligence or machine learning-specific applications, one or more microcontrollers, etc.

[0051] The client computing device 102 may include one or more network interface controllers 122, such as Ethernet network interface controllers, wireless network interface controllers, etc. The network interface controllers 122 may include advanced features, in some aspects, such as hardware acceleration, specialized networking protocols, etc.

[0052] The memories 124 of the client computing device 102 may include volatile and / or nonvolatile storage media. For example, the memories 124 may include one or more random access memories, one or more read-only memories, one or more cache memories, one or more hard disk drives, one or more solid-state drives, one or more non-volatile memory express, one or more optical drives, one or more universal serial bus flash drives, one or more external hard drives, one or more network-attached storage devices, one or more cloud storage instances, one or more tape drives, etc.

[0053] As noted, the memories 124 may have stored thereon one or more modules 140, for example, as one or more sets of computer-executable instructions. In some aspects, the modules 140 may include additional storage, such as one or more operating systems {e.g., Microsoft Windows, GNU / Linux, Mac OSX, etc.). The operating systems may be configured to run the modules 140 during operation of the client computing device 102 - for example, the modules 140 may include additional modules and / or services for receiving and processing data from one or more other components of the environment 100 such as the one or more cloud APIs 114 or the server computing device 104. The modules 140 may be implemented using any suitable computer programming language(s) {e.g., Python, JavaScript, C, C++, Rust, C#, Swift, Java, Go, LISP, Ruby, Fortran, etc.).

[0054] The modules 140 may include an API module 144, an input processing module 146, an authentication / security module 148, a context module 150 and a controller module 152, in some aspects. In some aspects, more or fewer modules 140 may be included. The modules 140 may be configured to communicate with one another e.g., via inter-process communication, via a bus, via sockets, pipes, message queues, etc.).

[0055] The API module 144 may include one or more sets of computer executable instructions for accessing one or more remote APIs, and / or for enabling one or more other components within the environment 100 to access functionality of the client computing device 102. For example, the API module 144 may enable a user to query the electronic objects corresponding to generative AI / ML models that are configured and stored in the model database 112. In some aspects, the API module 144 may enable other client applications {i.e., not applications facilitated by the modules 140) to connect to the client computing device 102, for example, to send queries or prompts, and to receive responses from the client computing device 102.

[0056] As noted, the client computing device 102 may enable one or more users to access one or more trained models by providing input prompts that are processed by one or more trained models. The input processing module 146 may perform pre-processing of user prompts prior to being input into one or more models, and / or post-processing of outputs output by one or more models. For example, the input processing module 146 may process data input into one or more input fields, voice inputs or other input methods (e.g., file attachments) depending upon the application. The input processing module 146 may receive inputs directly via the input device 126, in some aspects.

[0057] In some aspects, the input processing module 146 may perform post-processing of input received from one or more trained models. In some aspects, post-processing (and / or preprocessing) may include formatting outputs from the trained ML models in a manner consistent with the intended display that may be interpreted from the user intent. The input processing module 146 may include instructions for handling the display of such post-processed content to users (e.g., via the output device 128). The input processing module 146 may cause one or more graphical user interfaces to be displayed, for example to enable the user to enter information directly via a text field.

[0058] The authentication / security module 148 may include one or more sets of computerexecutable instructions for implementing access control mechanisms for one or more trained models, ensuring that the model and / or the underlying EHR data can only be accessed by those who are authorized to do so, and that the access of those users is private and secure. It should be appreciated that the security module 148 may permission users / agents based upon their respective membership and / or otherwise authorization to access information / records stored in a patient HPI, H&P, EHR, and / or other suitable record(s). For example, a user that has not been added and / or otherwise authorized to access certain sensitive records / information may not be able to access the contents of that information contained in a particular HPI via the client computing device 102.

[0059] Typically, trained models, especially trained models, may utilize state information in order to meaningfully carry on a dialogue with a user or with another trained model. For example, if a user prompts a trained model with a question such as “What is the weather in Chicago today?” followed by a second prompt “And how about tomorrow?” the model should understand that, in context, the second query relates to the first query, insofar as the user is asking about the weather tomorrow in the same location (Chicago).

[0060] However, language models (e.g., large language models (LLMs)) are generally stateless, meaning that after they process a prompt, they have no internal record or memory of theinformation that was input, or the information that was generated as part of the language model’s processing. Thus, some embodiments of this disclosure may add statefulness to models using context information. This may be implemented using sliding context windows, wherein a predetermined number of tokens (e.g., 4096 maximum tokens in the case of GPT 3.5, equivalent to about 3000 words) may be “remembered” by the LLM and can be used to enrich multiple sequential prompts input into the LLM (for example, when the LLM is used in a chat mode).

[0061] The context module 150 may include one or more sets of computer-executable instructions for maintaining state of the type found in this example, and other types of state information. The context module 150 may implement sliding window context, in some aspects. In other aspects, the context module 150 may perform other types of state maintaining strategies. For example, the context module 150 may implement a strategy in which information from the immediately preceding prompt is part of the window, regardless of the size of that prior prompt.

[0062] In some aspects, the context module 150 may implement a strategy in which one or more prior prompts are included in each current prompt. This prompt stuffing technique, or prompt concatenation, may be limited by prompt size constraints — once the total size of the prompt exceeds the prompt limit, the model immediately loses state information related to parts of the prompt truncated from the prompt.

[0063] The controller module 152 may include one or more sets of computer-executable instructions for accessing one or more trained models (e.g., via the model database 112 and / or the HPI database 108). The controller module 152 may include instructions for creating / adding electronic model objects, for adding controllers to those respective model objects and / or for configuring one or more links between the electronic model objects and / or the one or more controllers. The controller module 152 may access the one or more trained models via the network 106 and the server computing device 104, in some aspects.

[0064] The server computing device 104 may include one or more processors 160, one or more network interface controllers 162, one or more memories 164, an input device (not depicted), an output device (not depicted) and a server API 166. The one or more memories 164 may have stored thereon one or more modules 170 (e.g., one or more sets of instructions).

[0065] In some aspects, the one or more processors 160 may include one or more central processing units, one or more graphics processing units, one or more field-programmable gate arrays, one or more application-specific integrated circuits, one or more tensor processing units, one or more digital signal processors, one or more neural processing units, one or more RISC-Vprocessors, one or more coprocessors, one or more specialized processors / accelerators for artificial intelligence or machine learning-specific applications, one or more microcontrollers, etc.

[0066] The server computing device 104 may include one or more network interface controllers 162, such as Ethernet network interface controllers, wireless network interface controllers, etc. The network interface controllers 162 may include advanced features, in some aspects, such as hardware acceleration, specialized networking protocols, etc.

[0067] The memories 164 of the server computing device 104 may include volatile and / or nonvolatile storage media. For example, the memories 164 may include one or more random access memories, one or more read-only memories, one or more cache memories, one or more hard disk drives, one or more solid-state drives, one or more non-volatile memory express, one or more optical drives, one or more universal serial bus flash drives, one or more external hard drives, one or more network-attached storage devices, one or more cloud storage instances, one or more tape drives, etc.

[0068] As noted, the memories 164 may have stored thereon one or more modules 170, for example, as one or more sets of computer-executable instructions. In some aspects, the modules 170 may include additional storage, such as one or more operating systems (e.g., Microsoft Windows, GNU / Linux, Mac OSX, etc.). The operating systems may be configured to run the modules 170 during operation of the server computing device 104 - for example, the modules 170 may include additional modules and / or services for receiving and processing data from one or more other components of the environment 100 such as the one or more cloud APIs 114 or the client computing device 102. The modules 170 may be implemented using any suitable computer programming language(s) (e.g., Python, JavaScript, C, C++, Rust, C#, Swift, Java, Go, LISP, Ruby, Fortran, etc.).

[0069] In some aspects, the modules 170 may include a data collection module 172, a data preprocessing module 174, a model pretraining module 176, a fine-tuning module 178, a model training module 180, a checkpointing module 182, a hyperparameter tuning module 184, a validation and testing module 186, an auto-prompting module 188, and a summary and prediction module 190. In some aspects, more or fewer modules 170 may be included. The modules 170 may be configured to communicate with one another (e.g., via inter-process communication, via a bus, via sockets, pipes, message queues, etc.). The modules 170 may respond to network requests (e.g., via the API 166) or other requests received via the network106 (e.g., via the client computing device 102 or other components of the environment 100).

[0070] The data collection module 172 may be configured to collect information used to train one or more modules. In general, the information collected may be any suitable informationused for training a language model. The data collection module 172 may collect data via web scraping, via API calls / access, via database extract-transform-load (ETL) processes, etc. Sources accessed by the data collection module 172 include social media websites, books, websites, academic publications, web forums / interest sites (e.g., Reddit, Facebook, bulletin boards, etc.), etc. The data collection module 172 may access data sources by active means (e.g., scraping or other retrieval) or may access existing corpuses. The data collection module 172 may include sets of instructions for performing data collection in parallel, in some aspects. The data collection module 172 may store collected data in one or more electronic databases, such as a database accessible via the cloud APIs 1 14 or via a local electronic database (not depicted). The data may be stored in a structured and / or unstructured format. In some aspects, the data collection module 172 may store large data volumes used for training one or more models (i.e., training data). For example, the data collection module 172 may store terabytes, petabytes, exabytes or more of initial training data. This vast collection of initial training data may be used, e.g., to train an initial, foundation model to have general understanding of language and / or knowledge to be applied across a wide range of use cases.

[0071] In some aspects, the data collection module 172 may further retrieve data from the HPI database 108 to further train and / or fine-tune the foundational model to create a customized ML model. For example, the data collection module 172 may process the retrieved / received data and sort the data into multiple subsets based on information included within the HPI database 108. For example, the data collection module 172 may receive one or more sets of unstructured text (e.g., chronological descriptions of a patient’s illness development, contextual background, medications provided / taken, past medical history, etc.). The data collection module 172 may also chunk the data according to time (e.g., hourly, daily, quarterly, etc.).

[0072] The data preprocessing module 174 may include instructions for pre-processing data collected by the data collection module 172. In particular, the data preprocessing module 174 may perform text extraction and / or cleaning operations on data collected by the data collection module 172. The data pre-processing module 174 may perform preprocessing operations, such as lexical parsing, tokenizing, case conversions and other string splitting / munging. In some aspects, the data collection module 172 may perform data deduplication, filtering, annotation, compliance, version control, validation, quality control, etc. In some aspects, one or more human reviewers may be looped into the process of pre-processing data collected by the data pre-processing module 174. For example, a distributed work queue may be used to transmit batch jobs and receive human-computed responses from one or more human workers. Oncepre-processed, the data pre-processing module 174 may store copied and / or modified copies of the training data in an electronic database.

[0073] In some aspects the data pre-processing module 174 may include instructions for parsing the unstructured text received by the data collection module 172 to structure the text. For example, when the text relates to a plurality of clinician notes from different clinicians (e.g., nurses, cardiologist, radiologist, etc.) from an HPI or H&P document, the data pre-processing module 174 may generate a time series data structure in which each encounter with a patient is represented by one or more timestamps, and at each timestamp, text from one or more clinicians is labeled. The data pre-processing module 174 may also label the data according to the identity of one or more clinician, one or more topic, and / or one or more specialty / subspeciality of the clinician associated with the data.

[0074] For example, the time series data may be labeled according to one or more clinicians associated with textual and / or verbal notes, in a language transcript form. The time series data may include one or more keywords associated with the transcript. In some aspects, the present techniques may use a separate trained text summarization module to generate keywords used for this purpose. In this way, the data pre-processing module 174 may generate structured data corresponding to unstructured clinician notes, HPIs, etc., such that the structured data is enriched with information about the individual patient encounters that is suitable for training. This structured data may be processed by downstream processes / modules.

[0075] Generally, the present techniques may train one or more models to perform language generation tasks that include token generation. Both training inputs and model outputs may be tokenized. Herein, tokenization refers to the process by which text used for training is divided into units such as words, subwords or characters. Tokenization may break a single word into multiple subwords (e.g., “LLM” may be tokenized as “L” and “LM”). The present techniques may train one or more models using a set of tokens (e.g., a vocabulary) that includes many (e.g., thousands or more) of tokens. These tokens may be embedded into a vector. This vector of token or “embeddings” may include numerical representations of the individual tokens in the vocabulary in high-dimensional vector space. The modules 170 may access and modify the embeddings during training to learn relationships between tokens. These relationships effectively represent semantic language meaning.

[0076] In some aspects, a specialized database (e.g., a vector store, a graph database, etc.) may be used to store and query the embeddings. Embedding databases may include specialized features, such as efficient retrieval, similarity search and scalability. For example, the server computing device 104 may include a local electronic embedding database (notdepicted). In some aspects, a remote embedding database service may be used (e.g., via the cloud APIs 1 14). Such a remote embedding database service may be based on an open source or proprietary model (e.g., Milvus, Pinecone, Redis, Postgres, MongoDB, Facebook Al Similarity Search (FAISS), etc.). The server computing device 104 may include instructions (e.g., in the data collection module 172) for adding training data to one or more specialized databases, and for accessing it to train models.

[0077] The present techniques may include language modeling, wherein one or more deep learning models are trained by processing token sequences using an LLM architecture. For example, in some aspects, a transformer architecture may be used to process a sequence of tokens. Such a transformer model may include a plurality of layers including self-attention and feedforward neural networks. This architecture may enable the model to learn contextual relationships between the tokens, and to predict the next token in a sequence, based upon the preceding tokens. During training, the model is provided with the sequence of tokens and it learns to predict a probability distribution over the next token in the sequence. This training process may include updating one or more model parameters (e.g., weights or biases) using an objective function that minimizes the difference between the predicted distribution and a true next token in the training data.

[0078] Alternatives to the transformer architecture may include recurrent neural networks, long short-term memory networks, gated recurrent networks, convolutional neural networks, recursive neural networks, and / or other modeling architectures.

[0079] More generally speaking, programmable chatbots, such as an example embodiment of the summary and prediction module 190, may provide tailored, conversational-like abilities when interacting with a user. The chatbot may be capable of understanding user requests / responses, providing relevant information, etc. Additionally, the chatbot may generate data from user interactions which the enterprise may use to personalize future support and / or improve the chatbot's functionality, e.g., when retraining and / or fine-tuning the chatbot.

[0080] A ML chatbot may provide advance features as compared to a non-ML chatbot, which may include, and / or derive functionality from, a large language model (LLM). The ML chatbot may be trained on a server, such as the server computing device 104, using large training datasets of text which may provide sophisticated capability for natural-language tasks, such as answering questions and / or holding conversations. The ML chatbot may include a general- purpose pretrained LLM which, when provided with a starting set of words (prompt) as an input, may attempt to provide an output (response) of the most likely set of words that follow from the input. In one aspect, the prompt may be provided to, and / or the response received from, the MLchatbot and / or any other ML model, via a user interface of the server. This may include a user interface device operably connected to the server via an I / O module, such as the input device 126 and / or the output device 128. Example user interface devices may include a touchscreen, a keyboard, a mouse, a microphone, a speaker, a display, and / or any other suitable user interface devices.

[0081] Multi-turn ( / .e., back-and-forth) conversations may require LLMs to maintain context and coherence across multiple user utterances and / or prompts, which may require the ML chatbot to keep track of an entire conversation history as well as the current state of the conversation. The ML chatbot may rely on various techniques to engage in conversations with users, which may include the use of short-term and long-term memory. Short-term memory may temporarily store information (e.g., in the memory 164 of the server computing device 104) that may be required for immediate use and may keep track of the current state of the conversation and / or to understand the user’s latest input in order to generate an appropriate response. Long-term memory may include persistent storage of information (e.g., on the model database 112 and / or the HPI database 108 connected to the server computing device 104) which may be accessed over an extended period of time. The long-term memory may be used by the ML chatbot to store information about the user (e.g., preferences, chat history, etc.) and may be useful for improving an overall user experience by enabling the ML chatbot to personalize and / or provide more informed responses.

[0082] The system and methods to generate and / or train an ML chatbot model (e.g., via the training module 180 of the server computing device 104) which may be used by the ML chatbot, may consist of three steps: (1 ) a supervised fine-tuning (SET) step where a pretrained language model (e.g., an LLM) may be fine-tuned on a relatively small amount of demonstration data curated by human labelers to learn a supervised policy (SET ML model) which may generate responses / outputs from a selected list of prompts / inputs. The SET ML model may represent a cursory model for what may be later developed and / or configured as the ML chatbot model; (2) a reward model step where human labelers may rank numerous SFT ML model responses to evaluate the responses which best mimic preferred human responses, thereby generating comparison data. The reward model may be trained on the comparison data; and / or (3) a policy optimization step in which the reward model may further fine-tune and improve the SFT ML model. The outcome of this step may be the ML chatbot model using an optimized policy. In one aspect, step one may take place only once, while steps two and three may be iterated continuously, e.g., more comparison data is collected on the current ML chatbot model, which may be used to optimize / update the reward model and / or further optimize / update the policy.

[0083] More specifically, the modules 170 may include instructions for performing pretraining of a language model (e.g., an LLM), for example, in a pretraining module 176. The pretraining module 176 may include one or more sets of instructions for performing pretraining, which as used herein, generally refers to a process that may span pre-processing of training data via the data pre-processing module 174 and initialization of an as-yet untrained language model. In general, a pre-trained model is one that has no prior training of specific tasks. For example, the model pretraining module 176 may include instructions that initialize one more model weights.

[0084] In some aspects, model pretraining module 176 may initialize the weights to have random values. The model pretraining module 176 may train one or more models using unsupervised learning, wherein the one or more models process one or more tokens (e.g., preprocessed data output by the data pre-processing module 174) to learn to predict one or more elements (e.g., tokens). The model pretraining module 176 may include one or more optimizing objective functions that the model pretraining module 176 applies to the one or more models, to cause the one or more models to predict one or more most-likely next tokens, based on the likelihood of tokens in the training data. In general, the model pretraining module 176 causes the one or more models to learn linguistic features such as grammar and syntax. The pretraining module 176 may include additional steps, including training, data batching, hyperparameter tuning and / or model checkpointing.

[0085] The model pretraining module 176 may include instructions for generating a model that is pretrained for a general purpose, such as general text processing / understanding. This model may be known as a “base model” or a “foundation model” in some aspects. The base model may be further trained by downstream training process(es), for example, those training processes described with respect to the fine-tuning module 178. The model pretraining module 176 generally trains foundational models that have general understanding of language and / or knowledge. Pretraining may be a distinct stage of model training in which training data of a general and diverse nature (i.e., not specific to any particular task or subset of knowledge) is used to train the one or more models. In some aspects, a single model may be trained and copied. Copies of this model may serve as respective base models for a plurality of fine-tuned models.

[0086] The modules 170 may include a fine-tuning module 178. The fine-tuning module 178 may include instructions that train the one or models further to perform specific tasks. For example, the fine-tuning module 178 may train each of a plurality of models to generate one or more outputs that are based on respective personalities (e.g., specialties / subspecialities of clinicians or other healthcare providers), characteristics, and / or sets of tasks and / or tools withwhich the model is expected to interact / perform. Specifically, the fine-tuning module 178 may include instructions that train one or more models to generate respective language outputs (e.g., text generation), summarization, question answering or translation activities based on the characteristics of each respective model.

[0087] Continuing the example, the fine-tuning module 178 may include sets of instructions for retrieving one or more structured data sets, such as time series generated by the data preprocessing module 174. These structured data sets may be sorted by time and / or speaker to train one or more machine learning models (e.g., one or more language models) that may be used within the environment 100 to provide accurate, timely healthcare data management. For example, the fine-tuning module 178 may include instructions for configuring an objective function for performing a specific task, such as generating text that is similar to text found within the corpus of training data associated with a particular individual by role.

[0088] In some aspects, the fine-tuning module 178 may include user-selectable parameters that affect the fine-tuning of the one or more models. For example, a “caution” bias parameter may be included that represents medical conservativeness. This bias parameter may be adjusted to affect the cautiousness with which the resulting trained model approaches medical decision-making (e.g., diagnosis, prognosis), record generation / updating (e.g., HPI summary generation / updating), and / or predictions / estimations (e.g., outcome predictions, capacity predictions). Additional models may be trained, for additional personas / tasks, as discussed below.

[0089] In some aspects, to manage complexity of fine-tuning and other machine learning operations of the server computing device 104, one or more open-source frameworks may be used. Example frameworks include TensorFlow, Keras, MXNet, Caffe, SciKit learn, PyTorch. Specifically for training and operating language models, frameworks such as OpenLLM and LangChain may be used, in some aspects. The fine-tuning module 178 may use an algorithm such as stochastic gradient descent or another optimization technique to adjust weights of the pretrained model.

[0090] Fine-tuning may be an optional operation, in some aspects. In some aspects, training may be performed by the training module 180 after pretraining by the model pretraining module 176. In some aspects, the model training module 180 may perform task-specific training like the fine-tuning module 178, on a smaller scale or with a more tailored objective.

[0091] The training module 180 may include one or more submodules, including the checkpointing module 182, the hyperparameter tuning module 184, the validation and testing module 186 and the auto-prompting module 188. The checkpointing module 182 may performcheckpointing, which is saving of a model’s parameters. The checkpointing module 182 may store checkpoints during training and at the conclusion of training, for example, in the model database 112. In this way, the model may be run (e.g., for testing and validation) at multiple stages and its training parameters loaded, and also retrained from a checkpoint. In this way, the model can be run and trained forward without being re-trained from the beginning, which may save significant time (e.g., days of computation).

[0092] The hyperparameter tuning module 184 may include hyperparameters such as batch size, model size, learning rate, etc. These hyperparameters may be adjusted to influence model training. The hyperparameter tuning module 184 may include instructions for tuning hyperparameters by successive evaluation. The validation and testing module 186 may include sets of instructions for validating and testing one or more machine learning models, including those generated by the model pretraining module 176, the fine-tuning module 178 and the model training module 180. The auto-prompting module 188 may include sets of instructions for performing auto-prompting of one or more models. Specifically, the auto-prompting module 188 may enrich a prompt with additional information. The auto-prompting module 188 may include additional information in a prompt, so that the model receiving the prompt has additional context or directions that it can use. This may allow the auto-prompting module 188 to fine-tune a base model using one-shot or few-shot learning, in some aspects. The auto-prompting module 188 may also be used to focus the output of the one or more models.

[0093] In some aspects, the training module 180 may include instructions for training one or more additional machine learning models, such as supervised or unsupervised machine learning models. For example, as discussed below, in some aspects, the present techniques may include processing imaging data (e.g., X-rays, CT scans) retrieved from the HPI database 108. In that case, the patient’s imaging data may be processed by a model (e.g., a convolutional neural network) and the results processed further (e.g., by a language model) and / or provided to the client computing device 102. The training module 180 may train such a supervised model separately from training one or more language models. Further, the server computing device 104 may select one or more trained models at runtime based on data about a specific patient, based upon data contained in a prompt or based on other conditions that may be preprogrammed into the server computing device 104.

[0094] In some aspects, the training module 180 may train multi-modal models. For example, the training module 180 may train a plurality of models each capable of drawing from multimodal data types such as written text, imaging data, laboratory data, real-time monitoring data, pathology images, etc. In some cases, the training module 180 may train a single modelcapable of processing the multimodal data types. In some aspects, a trained multi-modal model may be used in conjunction with another model (e.g., a large language model) to provide nontext data interactions with users. Non-text data may be analyzed and integrated into the chat functions discussed herein.

[0095] The summary and prediction module 190 may operate one or more trained models. Specifically, the summary and prediction module 190 may initialize one or more trained models, load parameters into the model(s), and provide the model(s) with inference data (e.g., prompt inputs). In some aspects, the summary and prediction module 190 may deploy one or more trained model (e.g., a pretrained model, a fine-tuned model and / or a trained model) onto a cloud computing device (e.g., via the API 166). The summary and prediction module 190 may receive one or more inputs, for example from the client computing device 102, and provide those inputs (e.g., one or more prompts) to the trained model.

[0096] In some aspects, the API 166 may include elements for receiving requests to the model, and for generating outputs based on model outputs. For example, the API 166 may include a RESTful API that receives a GET or POST request including a prompt parameter. The summary and prediction module 190 may receive the request from the API 166 and pass the prompt parameter into the trained model and receive a corresponding input. For example, the prompt parameter may be an update to a patient’s HPI indicating “Patient X is experiencing symptoms A, B, and C”. The prompt output may be and / or include an HPI summary indicating “Patient X has experienced symptoms A and B for the past two months. Further, patient X experiences the most pain while standing and suffered an acute injury two months ago leading to the initial symptoms A and B.” In certain aspects, outputs may also include numeric confidence levels.

[0097] Broadly, the summary and prediction module 190 may include one or more ML models trained to emulate and / or otherwise optimally serve users with particular specialties / subspecialties and / or may serve as a conversational agent. In some embodiments, a user (e.g., a clinician) may provide an input prompt to the summary and prediction module 190, and the summary and prediction module 190 may determine a user intent based on the input prompt. For example, a user intent may indicate that the user intends to retrieve the data from patient X’s HPI, and more specifically, that the user intends to retrieve a summary of patient X’s HPI to quickly determine subsequent diagnostic steps or actions to be taken that may benefit the patient.

[0098] In some embodiments, in response to receiving such an input prompt and identifying such a user intent, the summary and prediction module 190 may determine and / or otherwisedesignate a tool configured to retrieve pertinent information (e.g., patient X’s HF I) related to the input prompt. When the tool retrieves the relevant data, the summary and prediction module 190 may format the data into a response / output prompt, and the user may engage in an LLM- based Q&A session with the summary and prediction module 190 and / or may simply receive the HPI information without further engagement with the summary and prediction module 190. For example, such a Q&A session may include the summary and prediction module 190 visualizing the data graphically for the user or collating the data for comprehensive insights. In this manner, the summary and prediction module 190 may succinctly summarize intricate notes (e.g., a patient’s entire HPI), thereby making user healthcare data consumption more efficient than conventional techniques requiring users to filter through extensively complex HPIs, H&Ps, EHRs, etc. For instance, complex tasks such as outside record summarization or retrieval augmented generation, may be streamlined through the distilled LLMs of the summary and prediction module 190.

[0099] In particular, in these embodiments, the summary and prediction module 190 integration layer may be equipped with LLM models configured to decipher user prompts and uses a series of tools / plugins to fulfil the users’ needs, as determined from the user’s input prompt. These models may determine and / or otherwise analyze the intent behind an input prompt and decide what type of database (e.g., HPI database 108) to search or what kind of application to call to fulfill the user’s intent. This intent analysis ensures accurate data retrieval and suggests further inquiries, thereby enhancing the depth of data interaction. Each piece of information generated by the summary and prediction module 190 may also be backed by links and / or references that enable users to delve deeper or verify the data's accuracy.

[0100] As part of this data retrieval / analysis, the summary and prediction module 190 is configured to connect with both external and internally stored digital tools via APIs (e.g., APIs 166, 114). For example, a specific specialty tool may be configured to predict hospital occupancy / bed availability based on patient HPIs, and the summary and prediction module 190 may seamlessly link the user's input prompt to this tool to receive relevant feedback. This seamless integration with Al tools through suitable APIs 166, 1 14 enables the summary and prediction module 190 to be a continuously evolving system, such that insights and data results retrieved and / or otherwise generated by the Al tools may be used to re-train, update, and / or otherwise influence the outputs of the summary and prediction module 190 models.

[0101] Further, the summary and prediction module 190 may provide an outcome prediction feature, which may predict / estimate various outcomes pertinent to a particular patient. For example, the outcome prediction feature may receive inputs related to a patient’s illness (e.g.,via HPI database 108), and may output predictions related to a differential diagnosis, potential complications associated with the patient’s illness, operative risks, a projected length of hospital stay, a rehospitalization rate, and / or other parameters associated with the patient’s illness.

[0102] Additionally, the summary and prediction module 190 may automatically update patient HPIs, H&P documents, and / or other medical records by collating retrieved items into a formatted document {e.g., PDF), thereby facilitating the creation of updated HPIs, H&Ps, document submissions, and / or otherwise enhancing the patient's medical record with relevant interval data {e.g., new lab results, new scan images, new symptoms, etc.) from a current encounter with a clinician. This functionality may also enable the summary and prediction module 190 to integrate radiology image links, electrocardiogram (ECG) results, and / or other relevant image / audio data into output / response prompts, which enables users to view exam results with a single click.

[0103] As part of these functionalities, the summary and prediction module 190 may include any suitable configuration ML models. For example, the summary and prediction module 190 may include an existing LLM that is fine-tuned for two tasks: pre-authorization and outside record summarization. These pre-existing models may be distilled down for these specific tasks and incorporated into the summary and prediction module 190. However, it should be understood that the summary and prediction module 190 may include any suitable number of LLMs, and / or other ML models, and that such LLMs and / or other ML models may be configured to perform {e.g., fine-tuned) any suitable number and / or types of tasks.

[0104] In any event, in certain aspects, the summary and prediction module 190 and / or other modules described herein may utilize multi-modal modeling. The data preprocessing module may, for example, process and understand image data, audio data, video data, etc. The server computing device 104 may interpret and respond to queries that involve understanding content from these different modalities. For example, the server computing device 104 may include an image processing module (not depicted) including instructions for performing image analysis on images provided by users, or images retrieved from patient HPI data. In some aspects, the server computing device 104 may generate outputs in modalities other than text. For example, the server computing device 104 may generate an audio response, an image, etc. Still further, the model objects may process both text data and other data e.g., an image uploaded by a user, an image from patient HPI, etc.) during a data stream with a user. Combining multi-modal data may enable the present models to perform more comprehensive analysis of patient conditions, based on information processed in multiple different modes simultaneously.

[0105] The summary and prediction module 190 may include a set of computer-executable instructions that when executed by one or more processors (e.g., the processors 160) cause a computer e.g., the server computing device 104) to perform retrieval-augmented generation. Specifically, the summary and prediction module 190 may perform retrieval-augmented generation based upon inputs or queries received from the user. This allows the summary and prediction module 190 to tailor responses of a model based on the specific input and context, such as the medical issue under discussion and / or the provider specialty being considered / evaluated. For example, one or more models may be pre-trained, fine-tuned and / or trained as discussed above. During that training, the model may learn to generate tokens based on general language understanding as well as application-specific training. Such a model at that point may be static, insofar as it cannot access further information when presented with an input query.

[0106] When the model is used at runtime, however, such as when deployed in the environment 100, the summary and prediction module 190 may perform retrieval operations, such as searching or selecting information from a document, a database, or another source. The summary and prediction module 190 may include instructions for processing user input and for performing a keyword search, a regular expression search, a similarity search, etc. based upon that user input. The summary and prediction module 190 may input the results of that search, along with the user input, into the trained model. Thus, the trained model may process this additional retrieved information to augment, or contextualize, the generation of tokens that represent responses to the user’s query. In sum, retrieval augmented generation applied in this manner allows the model to dynamically generate outputs that are more relevant to the user’s input query at runtime. Information that may be retrieved may include data corresponding to a patient e.g., patient demographic information, medical history, clinical notes, diagnoses, medications, allergies, immunizations, laboratory results, oncology information, radiation and imaging information, vitals, etc.) and additional training information, such as medical journals, notes or speech transcripts from symposia or other meetings / conferences, etc.

[0107] The summary and prediction module 190 may also include generating an HPI summary or abstract for a patient. Essential information may be generated at this stage. For example, a patient’s KRAS mutation status may be essential for determining how aggressive to be in treating a patient. Thus, the summary and prediction module 190 may generate summary statistics for a patient based on the information included in the patient’s HPI and / or other relevant documentation related to the patient's condition. This HPI summary may be provided to the clinicians providing input prompts by the summary and prediction module 190, to provideadditional context. Additional example summary information may include, for example, chemotherapy types and duration that the patient has had, if any. In this case, the summary and prediction module 190 may be asked “Which radiation therapy has the patient received?” and in response, may generate “The patient has received radiation therapy X, so a suggested dosage at this stage is 50 gy in 6 daily fractions.” The output of the summary and prediction module 190 may include links that guide a user (e.g., clinician) from the patient’s HPI summary to the particular medical literature document(s) and / or to various standard operating procedure document(s) supporting this generated response.

[0108] The present techniques may trigger retrieval augmented generation by processing a prompt, in some aspects. For example, a prompt may be processed by the input processing module 146 of the client computing device 102, prior to processing the prompt by the one or more generative models. The input processing module 146 may trigger retrieval augmented generation based on the presence of certain inputs, such as patient information, or a request for specific information, in the form of keywords. The input processing module 146 may perform entity recognition or other natural language processing functions to determine whether the prompt should be processed using retrieval augmented generation prior to being provided to the trained model.

[0109] Further, the input prompts may request and / or otherwise indicate that a response for or from a particular clinician with a particular specialty / subspeciality is required. For example, the input processing module 146 may analyze an input prompt providing interval data for a patient, and may determine that the input prompt is from a cardiologist. The input processing module 146 and / or the summary and prediction module 190 may determine / identify the specialty / subspeciality of the user based on data input and / or requested by the user, log-on credentials of the user accessing the computing environment 100, semantic / syntactic interpretation of the prompt indicating a particular interest of the user, and / or by any other suitable means or combinations thereof.

[0110] Continuing the prior example, the input processing module 146 and / or the summary and prediction module 190 may determine that the input prompt is from a cardiologist because the semantic and / or syntactic content / context of the input prompt suggests that the user providing / requesting ECG results and interpretations is a cardiologist. As another example, the input processing module 146 and / or the summary and prediction module 190 may determine that an input prompt is from a radiologist because the semantic and / or syntactic content / context of the input prompt suggests that the user is providing / requesting CT scan results. In any event, the summary and prediction module 190 may then leverage a ML model (e.g., multimodalprediction model 504) trained to generate outputs formatted and / or otherwise configured to mimic a cardiologist (or radiologist, etc.) and / or provide information from the patient’s HPI or other data that is particularly relevant to a cardiologist. The summary and prediction module 190 may thereby provide an output highlighting the patient’s cardiac history, particularly as it relates to the patient’s current illness and / or is otherwise relevant / recorded in the patient’s HPI.

[0111] As discussed above, prompts may be received via the input processing module 146 of the client computing device 102 and transmitted to the server computing device 104 via the electronic network 106. In some aspects, the output of the model may be modulated prior to being transmitted, output or otherwise displayed to a user.

[0112] The client computing device 102 and the server computing device 104 may communicate with one another via the network 106. In some aspects, the client computing device 102 and / or the server computing device 104 may offload some or all of their respective functionality to the one or more cloud APIs 1 14. In aspects, the one or more cloud APIs 114 may include one or more public clouds, one or more private clouds and / or one or more hybrid clouds. The one or more cloud APIs 114 may include one or resources provided under one or more service models, such as Infrastructure as a Service (laaS), Platform as a Service (PaaS), Software as a Service (SaaS), and Function as a Service (FaaS). For example, the one or more cloud APIs 114 may include one or more cloud computing resources, such as computing instances, electronic databases, operating systems, email resources, etc. The one or more cloud APIs 1 14 may include distributed computing resources that enable, for example, the model pretraining module 176 and / or other of the modules 170 to distribute parallel model training jobs across many processors.

[0113] In some aspects, the one or more cloud APIs 1 14 may include one or more language operation APIs, such as OpenAI, Bing, Claude.ai, etc. In other aspects, the one or more cloud APIs 114 may include an API configured to operate one or more open-source models, such as Llama 2.

[0114] The electronic network 106 may be a collection of interconnected devices, and may include one or more local area networks, wide area networks, subnets, and / or the Internet. The network 106 may include one or more networking devices such as routers, switches, etc. Each device within the network 106 may be assigned a unique identifier, such as an IP address, to facilitate communication. The network 106 may include wired (e.g., Ethernet cables) and wireless (e.g., Wi-Fi) connections. The network 106 may include a topology such as a star topology (devices connected to a central hub), a bus topology (devices connected along a single cable), a ring topology (devices connected in a circular fashion), and / or a mesh topology(devices connected to multiple other devices). The electronic network 106 may facilitate communication via one or more networking protocols, such as packet protocols (e.g., Internet Protocol (IP)) and / or application-layer protocols (e.g., HTTP, SMTP, SSH, etc.). The network 106 may perform routing and / or switching operations using routers and switches. The network 106 may include one or more firewalls, file servers and / or storage devices. The network 106 may include one or more subnetworks such as a virtual LAN (VLAN).

[0115] The environment 100 may include one or more electronic databases, such as a relational database that uses structured query language (SQL) and / or a NoSQL database or other schema-less database suited for the storage of unstructured or semi-structured data.

[0116] For example, the HP I database 108 may be an electronic database that stores medical records of a patient, particularly those related to a present illness of a patient. This data may describe past treatment of patients, past measurements associated with the patient (e.g., blood pressure, heart rate, lab results, radiology images, etc.) and / or other information related to patient care. The HPI database 108 may also include correspondence between physicians (e.g., a plurality of sub-specialists). The present techniques may include populating the HPI database 108 with data (e.g., HPI summaries, outcome predictions, etc.) relating to a specific patient during an interaction (e.g., discussion) between a healthcare provider and the summary and prediction module 190 regarding the specific patient. This patient-specific information may include imaging data related to the patient, treatment data, biometric data, and / or other suitable data.

[0117] The present techniques may store training data, training parameters and / or trained models in an electronic database such as the model database 112. Specifically, one or more trained machine learning models may be serialized and stored in a database (e.g., as a binary, a JSON object, etc.). Such a model can later be retrieved, deserialized and loaded into memory and then used for predictive purposes. The one or more trained models and their respective training parameters (e.g., weights) may also be stored as blob objects. Cloud computing APIs may also be used to stored trained models, via the cloud APIs 114. Examples of these services include AWS SageMaker, Google Al Platform and Azure Machine Learning.

[0118] In operation, a user may access a unified graphical user interface via the client computing device 102. The unified graphical user interface may be configured, generated, and / or displayed by the input processing module 146 and / or the summary and prediction module 190 via the output device 128. Specifically, the input processing module 146 may configure the unified graphical user interface to accept prompts and display corresponding prompt outputs generated by one or more models by processing the outputs. The inputprocessing module 146 may also be configured to transmit the prompts input by the user via the network electronic network 106 to the API 166 of the server computing device 104. The API 166 may process the user inputs via one or more trained models.

[0119] At the time the user accesses the unified graphical user interface, one or more models may already be trained, including pretraining and fine-tuning. These trained models may be selectively loaded into the one or more model objects based on configuration parameters, and / or based upon the content of the user’s input prompts. Further, in some aspects, the user may engage in a question-answer session with the client computing device 102.Example Model-Controller Configuration

[0120] FIG. 2 depicts an example model-controller configuration 200, according to one or more aspects. The example model-controller configuration 200 may include a controller 202 that receives one or more prompts 204 and provides the prompts 204 to one or more model objects 206A-N. The controller 202 may correspond to the controller module 152 of FIG. 1 , in some aspects. The example model-controller configuration 200 may comprise the one or more model objects 206A-N and additional information, as discussed above.

[0121] Generally, the one or more model objects 206A-N correspond, respectively, to one or more trained ML models configured to emulate and / or otherwise interact with a specific subspeciality of provider (e.g., medical / healthcare provider). For example, each of the one or more model objects 206A-N may include a trained model and additional metadata, such as a profile name, training date, checkpoint information, etc. The controller 202 may receive the respective responses of one, some, or all of the one or more model objects 206A-N and generate one or more controller responses 208 based on the object responses. In some aspects, the one or more controller responses 208 may include each model object’s 206A-N response (e.g., a batched response). In some aspects, the one or more controller responses 208 may include one or more individual model object’s 206A-N responses.

[0122] The model-controller configuration 200 may include any number of model objects. For example, the model-controller configuration 200 depicts a model object 206-A, which is the first model object, a model object 206-B, which is the second model object, a model object 206-C, which is the third model object, a model object 206-D, which is the fourth model object and a model object 206-E, which is the nth model object, wherein n is a positive natural number.

[0123] For example, the first model object 206-A may include a ML model that is trained / configured to summarize data and / or otherwise generate an HPI and / or HPI summary that is most beneficial to a neurologist. Thus, the first model object 206-A, when receiving theprompt 204, may filter data from the patient’s HPI that is less relevant to the types of symptoms that may be of interest to a neurologist and / or may be less important / critical to any resulting diagnosis, prognosis, and / or treatment determined by the neurologist. The first model object 206-A may also aggregate any relevant test values from the patient’s HPI (e.g., brain angiography result, electroencephalogram (EEG) result) that may be of interest to the neurologist, and may determine the output prompt (e.g., HPI summary) by highlighting and / or otherwise emphasizing these specific elements of the patient’s HPI. In this manner, the neurologist may quickly analyze the relevant data presented in the HPI summary and make a well-informed decision regarding subsequent actions or treatment options for a patient based on the most relevant data from the patient’s HPI.

[0124] Further, when the first model object 206-A receives the input prompt from the neurologist (e.g., indicating interval history that updates patient neurological symptoms, etc.), the first model object 206-A may generate an HPI update to automatically format and apply an update to the patient’s HPI in the HPI database (e.g., HPI database 108). Moreover, if the input prompt from the neurologist indicates an error and / or other discontinuity between the interval history and the recorded HPI, the first model object 206-A may update the HPI retroactively / retrospectively to create an accurate record of the patient’s illness history. Additionally, as described in more detail above, this input may also be used to retrain / re-fine tune the ML model to produce updated versions of the ML model.

[0125] Similarly, as another example, the second model object 206-B may include a ML model that is trained / configured to summarize data and / or otherwise generate an HPI and / or HPI summary that is most beneficial to a cardiologist caring for an arrhythmia patient. In this example, the cardiologist may provide an input prompt indicating interval history that may include data describing and / or otherwise relating to heart palpitations, syncope, etc. and the second model object 206-B may leverage the input prompt, and the patient’s HPI to provide specialist germane background to the cardiologist related to any cardioversions, anti-arrhythmic drugs tried, previous ablation experienced by the patient, etc.

[0126] More generally, any of the model objects 206A-N may correspond to and / or be trained to interact with any suitable specialty / subspeciality. Moreover, in certain embodiments and as mentioned, the model objects 206A-N may be configured to interact with and / or otherwise utilize additional tools (e.g., a digital stethoscope (not shown)) to automatically populate physical exam results within a patient’s HPI i.e. , in some embodiments, the HPI may comprise sections generated by the ML model, sections directly populated using output from the additional tools, and sections directly populated using output from specialized prediction models (not shown).The model objects 206A-N may also automatically push / update the patient’s more comprehensive medical records (e.g., EHR) with such updated HPIs. Further, as previously mentioned, the model objects 206A-N may include and / or otherwise communicate with additional AI / ML tools / models (not shown) configured to provide analysis of other data or aspects of a patient’s medical history. For example, the model objects 206A-N may be configured to interact with a ML model configured to automatic analyze and generate annotations and / or interpreted results of a patient’s CT scan and / or other tests, and the output from these ML models may feed into the model objects 206A-N as inputs for inclusion and / or consideration when generating HPI summaries, outcome predictions, H&P documents, and / or any other suitable outputs or combinations thereof.

[0127] Of course, it should be appreciated that the example embodiments of the model objects 206A-N described in reference to FIG. 2 are for the purposes of discussion only. Accordingly, it should be appreciated that the model objects 206-A-N may each be configured to perform any suitable number and / or combination of the functions described herein, such as generating HPIs, HPI summaries, outcome predictions, and / or any other data evaluation / generation described herein.Example Computer-Implemented Methods

[0128] FIG. 3A depicts an example computer-implemented method 300 for optimizing summary generation using trained AI / ML models, in accordance with various embodiments herein.

[0129] At block 302, the method 300 includes receiving an input from a provider, the input including first data associated with a user. The method 300 further includes generating, by executing a trained ML model, an output including a history of present illness (HPI) summary by: identifying a subspeciality of the provider (block 304), aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input (block 306), and determining the output based on the input, the subspeciality of the provider, the prior input, and the test value (block 308). The method 300 further includes causing the HPI summary to be displayed at a device (block 310).

[0130] In some aspects, the method 300 further includes selectively including a subset of data from one or more of the input, the prior input, or the test value in the HPI summary based on the subspeciality of the provider. Such selective inclusion may be achieved by the trained ML model, by a report generator process receiving an initial HPI summary from the trained ML model, or some combination thereof.

[0131] In some aspects, the method 300 further includes generating, by executing the trained ML model, an outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and causing the outcome prediction and the HPI summary to be displayed at the device.

[0132] In some aspects, the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value.

[0133] In some aspects, the method 300 further includes: receiving a subsequent input related to the output from a subsequent provider; generating, by executing the trained ML model, a subsequent output including a subsequent HPI summary by: identifying a subsequent subspeciality of the subsequent provider, wherein the subsequent subspeciality is different from the subspeciality of the provider, aggregating (i) the input and (ii) the output, and determining the subsequent HPI summary based on the subsequent input, the subsequent subspeciality of the subsequent provider, the input, and the output; and causing the subsequent HPI summary to be displayed at a subsequent device.

[0134] In some aspects, the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider. In these aspects, the method 300 may further include generating the output by: determining, by executing the trained ML model, a subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the subset of data.

[0135] In some aspects, the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an HPI source and (b) a multi-modal ML model.

[0136] FIG. 3B depicts an example computer-implemented method 320 for optimizing summary generation using trained AI / ML models, in accordance with various embodiments herein.

[0137] At block 302, the method 320 includes receiving an input from a provider, the input including first data associated with a user. The method 320 further includes generating, by executing a trained ML model, an output including an outcome prediction for the user by: aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input (block 324), and determining the output based on the input, the prior input, and the test value (block 326). The method 320 further includes causing the outcome prediction to be displayed at a device (block 328).

[0138] In some aspects, the method 320 further includes: receiving a subsequent input related to the output; generating, by executing the trained ML model, a subsequent output including a subsequent outcome prediction by: aggregating (i) the input and (ii) the output, determining the subsequent output based on the subsequent input, the input, and the output, and wherein the subsequent outcome prediction includes one or more of: (i) a subsequent differential diagnosis, (ii) a subsequent potential complication, (iii) a subsequent operative value, (iv) a subsequent length of stay value, or (v) a subsequent rehospitalization value; and causing the subsequent outcome prediction to be displayed at a subsequent device.

[0139] In some aspects, the method 320 further includes: aggregating a set of outcome predictions for a plurality of users at a first location; generating, by executing the trained ML model, a capacity prediction for the first location based on the set of outcome predictions; and causing the capacity prediction to be displayed at the device.

[0140] In some aspects, the input is received from a provider with a subspeciality, and the method 320 further includes: selectively including a subset of data from one or more of the input, the prior input, or the test value in the output based on the subspeciality of the provider, and wherein the output includes a history of present illness (HPI) summary based on the subset of data.

[0141] In some aspects, the method 320 further includes: generating, by executing the trained ML model, the outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.

[0142] In some aspects, the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the method 320 further includes generating the output by: determining, by executing the trained ML model, a second subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the second subset of data.

[0143] In some aspects, the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an outcome source and (b) a multi-modal ML model.

[0144] FIG. 3C depicts an example computer-implemented method 340 for optimizing summary generation using trained AI / ML models, in accordance with various embodiments herein.

[0145] At block 342, the method 340 includes receiving an input from a provider, the input including first data associated with a user. The method 340 further includes generating, by executing a trained ML model, an output including a history of present illness (HPI) summary and an outcome prediction by: identifying a subspeciality of the provider (block 344), aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input (block 346), determining the HPI summary based on the input, the subspeciality of the provider, the prior input, and the test value (block 348), and determining the outcome prediction based on the HPI summary (block 350). The method 340 further includes causing the HPI summary and the outcome prediction to be displayed at a device (block 352).

[0146] In some aspects, the method 340 further includes: receiving a subsequent input related to the output; generating, by executing the trained ML model, a subsequent output including a subsequent HPI summary and a subsequent outcome prediction by: aggregating (i) the input and (ii) the output, determining the subsequent HPI summary based on the subsequent input, the input, and the output, and determining the subsequent outcome prediction based on the subsequent HPI summary; and causing the subsequent outcome prediction and the subsequent HPI summary to be displayed at a subsequent device.

[0147] In some aspects, the method 340 further includes: aggregating a set of outcome predictions for a plurality of users at a first location; generating, by executing the trained ML model, a capacity prediction for the first location based on the set of outcome predictions; and causing the capacity prediction to be displayed at the device.

[0148] In some aspects, the method 340 further includes determining the HPI summary by: selectively including a subset of data from one or more of the input, the prior input, or the test value in the HPI summary based on the subspeciality of the provider.

[0149] In some aspects, the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the method 340 further includes generating the output by: determining, by executing the trained ML model, a second subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the second subset of data.

[0150] In some aspects, the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value.

[0151] In some aspects, the first data may further comprise an X-ray image associated with the user, and the trained ML model may be a foundation model or a fine-tuned ML model configured to: receive X-ray images of users; and determine image features of the X-ray images.

[0152] In some aspects, the method 340 further includes a prediction model being configured to: receive the image features from the foundation model; and determine a predicted metric value based on the image features.

[0153] In some aspects, the method 340 further includes the prediction model being further configured to generate a predicted multimedia output based on the image features that correspond to the predicted value.

[0154] In some aspects, the prediction model is a multimedia model that is trained using historical multimedia data of a plurality of patients and historical metric data of the plurality of patients.

[0155] In some aspects, the foundation model is configured to receive one or more of: (i) X-ray images, (ii) computed tomography (CT) scans, (iii) magnetic resonance imaging (MRI) images, (iv) ultrasound images, (v) positron emission tomography (PET) scans, (vi) single photon emission computed tomography (SPECT) scans, (vii) optical coherence tomography (OCT) scans, or (viii) digital pathology slide; and the image features are associated with (i) an ejection fraction of the user, (ii) a cardiac output of the user, (iii) a ventricular volume of the user, (iv) a wall thickness of the user, (v) a motion abnormality of the user, (vi) a valvular function of the user, (vi) a cardiac morphology of the user, (vii) a blood flow pattern of the user, and (viii) a blood flow velocity of the user.Example User Interfaces

[0156] FIG. 4 illustrates an example user interface 400 configured to display and provide interactive access to the generated summaries and other data output by the trained AI / ML models, in accordance with various embodiments herein.

[0157] The example user interface 400 includes an HPI summary section 402, an outcome prediction section 404, an H&P and HPI interaction button 406, and a location occupancy predictor section 408. Each of these sections 402, 404, 408 and / or buttons 406 may be interactable, such that a user viewing the example user interface 400 may interact with (e.g., click, tap, swipe, voice command, gesture, etc.) each individual section 402, 404, 408 and / or button 406 to receive additional information and / or to open additional windows / browsers.

[0158] The HPI summary section 402 may include an HPI summary generated by a trained ML model, as described herein. For example, the HPI summary featured in the HPI summarysection 420 states “Patient X has experienced symptoms 1 , 2, 3 for the past 2 months. Patient X experiences the most pain while standing and suffered an acute injury leading to initial symptoms by performing action Z. Patient X previously received medication K and undergone tests A, B, and C. Further, based on the interval history provided, patient X should likely receive medication J.” Thus, the HPI summary section 420 provides a brief, succinct summary of the relevant data from a patient’s HPI to quickly inform a clinician (e.g., doctor, nurse, etc.) of the patient’s recent medical history surrounding a particular illness / illnesses. In this manner, the clinician may quickly and efficiently acquaint themselves with the patient’s history and thereby make accurate medical decisions without requiring the clinician to shuffle through a patient’s cumbersome, voluminous HPI.

[0159] More specifically, the HPI summary section 402 may include HPI summaries that are specifically tailored to clinicians with particular specialties / subspecialities. For example, and as described previously, the trained ML models of the present disclosure may evaluate input prompts from a clinician, identify a specialty of the clinician, and curate the resulting HPI summary based on the clinician’ specialty. As a result, the HPI summary displayed in the HPI summary section 402 may feature specific medical data for a patient that is of particular relevance to the clinician interacting with the trained ML model. Of course, if the clinician desires to review other / additional data related to the patient’s current illness, the clinician may subsequently interact with the HPI summary section 402 to receive additional information / data from the patient’s HPI. This additional information / data may also be curated for the clinician, such that the clinician may always receive data that is predicted to be more relevant to the clinician than the remaining corpus of data in the patient’s HPI.

[0160] The outcome prediction section 404 includes an outcome prediction related to the patient’s present illness. For example, the outcome prediction in the outcome prediction section 404 states, “Based on the CT results, Patient X is likely suffering from condition Y. This condition can lead to complications D and E, and operations to resolve condition Y have risks F and G. Patients experiencing condition Y typically stay hospitalized for 3-4 days and have a relatively low risk of rehospitalization.” Thus, the outcome prediction provided in the outcome prediction section 404 provides estimates for a differential diagnosis (e.g., condition Y), a potential complication (e.g., complications E and G), an operative value (e.g., risks F and G), a length of stay value (e.g., hospitalized for 3-4 days), and a rehospitalization value (e.g., relatively low risk of rehospitalization). Of course, as mentioned herein, it should be appreciated that the outcome prediction section 404 may include any suitable number, type, and / or combination of predicted outcomes.

[0161] The H&P and HPI interaction button 406 may be a selectable button that enables the viewing clinician / user to view H&P documents and / or HPIs for a particular patient or group of patients. As the trained ML models described herein update and / or summarize a patient's H&P documents and HPIs, the treating clinicians may desire to view these prior updates and / or summaries. Accordingly, the clinician may interact with the H&P and HPI interaction button 406 to view any / all prior updates and / or summaries for a patient’s (or group of patients) illness.

[0162] The location occupancy predictor section 408 generally provides clinicians with a predicted occupancy and / or bed availability for a hospital or other specific location (e.g., emergency room) based on the HPIs of current patients. For example, the occupancy prediction of the location occupancy predictor section 408 states, “Based on the current patient occupancy rate and prognoses of hospitalized patients, hospital S likely has 15 beds available for new patients.” Thus, the trained ML models described herein may evaluate the HPIs, outcome predictions, and / or other data corresponding to patients at a particular location (e.g., hospital) and may generate a predicted availability of beds, medical equipment, clinician time / appointments, and / or any other suitable resources within the particular location.

[0163] It should be appreciated that, while many of the embodiments herein are described in the context of generating HPIs, HPI summaries, outcome predictions, etc., the techniques of the present disclosure may be implemented to accomplish a variety of purposes. For example, the trained ML models may be configured to generate dismissal summaries, letters to physicians, and / or any other suitable purpose. Furthermore, while many of the embodiments herein are described in the context of a hospital, the techniques of the present disclosure may apply to a wide variety of contexts outside of a hospital setting.

[0164] FIG. 5 depicts an example foundation model and multimodal prediction model workflow 500 for generating predicted metric values and / or predicted multimedia outputs, in accordance with various embodiments herein. Generally, the workflow 500 includes a first stage 506 that comprises a training sequence for a foundation model 502 and a second stage 508 that comprises an implementation of the foundation model 502 with a multimodal prediction model 504.

[0165] The first stage 506 more specifically comprises training the foundation model 502 using training data, as described herein (e.g., as part of any of the pretraining module 176, fine-tuning module 178, and / or the training module 180). The foundation model 502 may incorporate various types of ML models, including but not limited to convolutional neural networks (CNNs), recurrent neural networks (RNNs), and transformers, and each may be suited to different aspects of image analysis and feature extraction using the training data. For example, thefoundation model 502 may include one or more CNNs, which may be particularly well-suited for image data analysis due to their ability to automatically and adaptively learn spatial hierarchies of features from images. For medical imaging (e.g., X-rays, CT scans, MRI, ultrasound), the CNNs included as part of the foundation model 502 may be trained to recognize patterns, shapes, and / or anomalies in the images, making them ideal for extracting features relevant to diagnosing conditions that may be inferred, predicted, and / or otherwise interpreted from such image data.

[0166] As another example, the foundation model 502 may include one or more RNNs, which may be beneficial when the data (e.g., medical image data) involves sequences or time-series, such as in echocardiograms or dynamic contrast-enhanced scans. RNNs, e.g., those with Long Short-Term Memory (LSTM) units, may analyze temporal changes in the medical images, which may enable the foundation model 502 to track the progression of a condition or the movement of organs. As still another example, the foundation model 502 may include one or more transformers. The ability of transformers to handle sequences makes them suitable for analyzing series of images or volumetric data (e.g., such as 3D scans) by treating the image or scan slices as sequences. This may thereby enable the foundation model 502 to more accurately understand complex spatial relationships and dependencies in medical images.

[0167] During the first stage 506, the foundation model 502 may be trained using any suitable ML technique(s), such as supervised learning, where the model 502 may learn from a labeled dataset containing medical images paired with annotations or labels indicating the features or conditions present in the images. For example, to train the foundation model 502 to output features relevant to determining the ejection fraction of a user’s heart, the training dataset may include images (e.g., X-rays) labeled with the corresponding ejection fraction values or annotations indicating key features (e.g., ventricular volumes or wall thickness) that are directly related to calculating the ejection fraction.

[0168] Additionally, or alternatively, the first stage 506 may include transfer learning to leverage pre-trained models on large datasets of medical images to improve the foundation model's 502 performance. Such transfer learning may generally involve fine-tuning a pretrained model on a specific dataset relevant to the task at hand, such as images specifically related to cardiac health, to create a fine-tuned foundation model 502. This approach may be particularly advantageous to achieve high accuracy results, particularly when the available training data is limited.

[0169] More specifically, the training process illustrated by the first stage 506 may involve iteratively adjusting the foundation model's 502 parameters to minimize the difference betweenthe predicted outputs and the actual labels in the training dataset. In certain instances, this parametric optimization is achieved through backpropagation and optimization algorithms, such as stochastic gradient descent and / or the like. As the training iterations increase, the foundation model 502 may learn to recognize patterns and features in the images that are predictive of the labels, allowing it to output specific features of new, unseen patient image data that are relevant to determining predicted values like an ejection fraction.

[0170] For example, the foundation model 502 may receive training data (e.g., X-ray data of a plurality of users) and output training features (e.g., features of the X-ray data that may be relevant to computing an ejection fraction). Based on the differences between the output training features and the known features of interest and / or the labels associated with those known features, the foundation model 502 may receive feedback that adjusts one or more parameters of the model 502 for subsequent iterations. This process may continue for as many iterations as necessary to achieve a desired alignment between the training features and the known feature values, until a threshold number of training iterations has taken place, and / or for any suitable number of iterations. Regardless, by combining these ML techniques and training approaches, the first stage 506 may produce a trained foundation model 502 that is configured to accurately analyze various types of input data (e.g., medical images, patient EHR data) and extract meaningful features that aid in the diagnosis and monitoring of health conditions.

[0171] In certain embodiments, the foundation model is configured to receive one or more of: (i) X-ray images, (ii) computed tomography (CT) scans, (iii) magnetic resonance imaging (MRI) images, (iv) ultrasound images, (v) positron emission tomography (PET) scans, (vi) single photon emission computed tomography (SPECT) scans, (vii) optical coherence tomography (OCT) scans, and / or (viii) digital pathology slides. Thus, the foundation model 502 may be generally configured to receive any suitable medical imaging data and / or other medical data (e.g., patient EHR data, provider data) to generate features associated with such data that is specifically configured to be utilized by the multimodal prediction model 504 to determine predicted metric values and / or predicted multimedia outputs associated with diagnosing / predicting the presence of any medical condition or value (e.g., ejection fraction, cardiac output, ventricular volume, wall thickness, motion abnormality, valvular function, cardiac morphology, blood flow pattern, and / or blood flow velocity).

[0172] While illustrated as the foundation model 502 being trained as part of the first stage 506, it should be appreciated that the first stage 506 may additionally, or alternatively, include training the multimodal prediction model 504. For example, the training of the multimodal prediction model 504 may involve a comprehensive approach utilizing advanced ML techniques.Generally speaking, the multimodal prediction model 504 may make predictions based on the extracted features from the foundation model 502, which may range from quantifiable metrics like the ejection fraction of a patient's heart to generating multimodal outputs such as videos simulating the patient's heart movement. Such model 504 may include several types of ML models and techniques, including deep learning models like CNNs, RNNs, generative adversarial networks (GANs), diffusion models, and / or any other suitable models or combinations thereof.

[0173] As an example, the multimodal prediction model 504 may include one or more GANs which may be utilized for generating realistic multimodal outputs, such as videos, based on the learned features. Additionally, or alternatively, the model 504 may include one or more diffusion models, which is a generative model configured to generate high-quality, realistic images and videos. Such diffusion models may generally work by gradually learning to reverse a diffusion process, which transforms data into a Gaussian distribution, to generate data from noise. Thus, for the multimodal prediction model 504, a diffusion model may be trained to generate detailed and realistic simulations of a patient's heart movement based on the feature data extracted from static medical images (e.g., X-rays).

[0174] Similar to the foundation model 502, the multimodal prediction model 504 may be trained using any suitable training approaches / techniques, such as supervised learning, where the model 504 may learn from a training dataset containing feature data and the corresponding outcomes and / or annotations. Specifically, the multimedia prediction model 504 may be trained using historical multimedia data (e.g., features of X-rays, CT scans, etc.) of a plurality of patients and historical metric data (e.g., ejection fraction, videos of the patient’s heart beating) of the plurality of patients. In certain instances, training the multimodal prediction model 504 may further include transfer learning to enhance the model's 504 ability to generate complex outputs or improve prediction accuracy by leveraging pre-trained models. Specifically, training the multimodal prediction model 504 when the model 504 includes a diffusion model may include teaching the model 504 to gradually denoise data, starting from a random noise distribution to the structured output (e.g., a video of a heart beating). This process may include a careful balancing of the fidelity and diversity of the generated outputs to ensure that the model's 504 parameters minimize the difference between its predictions or generated outputs and the actual data.

[0175] Such optimization may be achieved through algorithms that adjust the model 504 parameters / hyperparameters / etc. based on a calculated loss function. Moreover, the performance of the multimodal prediction model 504 may be evaluated using any suitablemetrics, such as accuracy for predictive tasks and specialized metrics like structural similarity index measure (SSIM) or peak signal-to-noise ratio (PSNR) for assessing the quality of generated videos. By integrating these ML models and training methodologies, the multimodal prediction model 504 may effectively utilize the feature data provided by the foundation model 502 to make accurate health predictions and generate realistic simulations of patient conditions. The multimodal prediction model 504 may thereby create highly realistic and detailed simulations based on medical imaging data, enhancing the model's 504 utility in diagnostic and treatment planning scenarios.

[0176] The second stage 508 more specifically comprises executing / applying the foundation model 502 in combination with a multimodal prediction model 504 to receive user image data as inputs and ultimately output predicted metric values and / or predicted multimedia outputs, which may be included as part of the HPI summary (e.g., HPI summary section 402), outcome predictions (e.g., outcome prediction section 404), and / or otherwise displayed to a user.

[0177] For example, executing the multimodal prediction model 504 may include receives feature data from a foundation model 502 as inputs and processing this input data to output predicted metric values (e.g., ejection fraction) and / or predicted multimedia outputs (e.g., generated videos or images). This process may generally leverage the intricate patterns and features extracted from the patient's imaging data to make accurate predictions or generate realistic simulations.

[0178] As an example, the multimodal prediction model 504 may output predicted metric values that include (i) an ejection fraction of the user, (ii) a cardiac output of the user, (iii) a ventricular volume of the user, (iv) a wall thickness of the user, (v) a motion abnormality of the user, (vi) a valvular function of the user, (vi) a cardiac morphology of the user, (vii) a blood flow pattern of the user, (viii) a blood flow velocity of the user, and / or any other suitable value associated with a user.

[0179] The ejection fraction may be used to assess the user’s cardiac function by indicating the percentage of blood ejected from the ventricles with each heartbeat. The multimodal prediction model 504 may predict the user’s ejection fraction based on features related to heart size, shape, and motion extracted from the user’s imaging data (e.g., X-rays). The cardiac output may be the volume of blood the heart pumps in a minute and may be a direct function of both the heart rate and the stroke volume (the amount of blood ejected by each heartbeat). The ventricular volumes, including both end-diastolic volume (EDV) (e.g., the total volume of blood in the ventricles at the end of diastole (heart relaxation)) and end-systolic volume (ESV) (e.g., the volume of blood remaining in the ventricles after contraction), may be included in the outputpredicted metric values for use, e.g., in calculating the ejection fraction. The wall thickness and motion abnormalities may indicate conditions such as hypertrophy (e.g., thickening of the heart muscle) or areas of the heart that may be underperforming due to ischemic injuries or other pathologies. The valvular function may include the detection of valvular stenosis (e.g., narrowing) or regurgitation (e.g., leakage), which may significantly affect cardiac efficiency and overall cardiovascular health. The cardiac morphology (e.g., including the shape, size, and structure of the heart and its chambers) may indicate one or more abnormalities that may be indicative of congenital heart defects, cardiomyopathies, and / or other conditions that may impact heart function. The blood flow patterns and velocities may indicate issues like turbulent flow due to stenosis or abnormal flow patterns that might suggest shunts or other circulatory issues. Thus, the model 504 may output any of these values and / or may utilize these values to generate any of the other values and / or to generate predicted multimedia outputs.

[0180] In some circumstances, the multimodal prediction model 504 may be configured to receive feature data from the foundation model 502 and output predicted metric values associated with any suitable values, such as a tumor size and classification. For example, the model 504 may predict the size of a tumor and its classification (e.g., benign or malignant) based on features extracted from imaging data, such as texture, shape, and the presence of irregularities. As another example, the model 504 may be configured to output predicted metric values associated with a user’s bone density, such as by predicting bone density metrics from features extracted from CT scans or X-rays, aiding in the assessment of fracture risk.

[0181] Additionally, or alternatively, the multimodal prediction model 504 may output predicted multimedia outputs, such as generated videos, images, and / or any other suitable media or combinations thereof. For example, the multimodal prediction model 504 may utilize a diffusion model to generate videos simulating the patient's heart beating and / or other dynamic processes within the body. Further, the model 504 may generate images, such as predicting the progression of a disease by simulating how a tumor might grow or how a condition like macular degeneration could advance, providing a visual forecast of the patient's condition. In another example, the model 504 may generate 3D reconstructions of anatomical structures (e.g., for surgical planning or diagnosis) from 2D image data, offering a comprehensive view of the area of interest. The multimodal prediction model 504 may generate such multimedia outputs by learning to reverse a process that adds noise to data, starting from noise and gradually refining it into a coherent video sequence based on the feature data.

[0182] More generally, the multimodal prediction model 504 may receive the feature data from the foundation model 502, which may broadly include data about the patient's anatomy, asidentified in the medical images. The multimodal prediction model 504 may then apply learned patterns and relationships between the features and the metrics of interest (e.g., ejection fraction) to output precise predictions. In certain embodiments, when the multimodal prediction model 504 includes a diffusion model, the prediction generation process may involve iteratively refining noise into a structured output that matches the feature data. The model 504 may leverage the extracted features to guide the generation process, ensuring that the final output (e.g., video, image, 3D reconstruction) accurately reflects the patient's condition, as depicted in the original imaging data. In this manner, the model 504 accurately interprets the feature data and applies the learned patterns to generate predictions or outputs that are clinically relevant and useful for patient care.Aspects of the Disclosure

[0183] Aspects of the techniques described in the present disclosure may include any of the following aspects, either alone or in combination:

[0184] Example 1 . A computing system for optimizing summary generation using generative artificial intelligence (Al) models, comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, and determining the output based on the input, the subspeciality of the provider, the prior input, and the test value, and cause the HPI summary to be displayed at a device.

[0185] Example 2. The computing system of example 1 , the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to determine the output by: selectively including a subset of data from one or more of the input, the prior input, or the test value in the HPI summary based on the subspeciality of the provider.

[0186] Example 3. The computing system of example 1 or 2, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: generate, by executing the trained ML model, an outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.

[0187] Example 4. The computing system of example 3, wherein the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value.

[0188] Example 5. The computing system of any of examples 1 through 4, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: receive a subsequent input related to the output from a subsequent provider; generate, by executing the trained ML model, a subsequent output including a subsequent HPI summary by: identifying a subsequent subspeciality of the subsequent provider, wherein the subsequent subspeciality is different from the subspeciality of the provider, aggregating (i) the input and (ii) the output, and determining the subsequent HPI summary based on the subsequent input, the subsequent subspeciality of the subsequent provider, the input, and the output; and cause the subsequent HPI summary to be displayed at a subsequent device.

[0189] Example 6. The computing system of any of examples 1 through 5, wherein the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the one or more memories having stored thereon computerexecutable instructions that, when executed by the one or more processors, further cause the computing system to generate the output by: determining, by executing the trained ML model, a subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the subset of data.

[0190] Example 7. The computing system of any of examples 1 through 6, wherein the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an HPI source and (b) a multi-modal ML model.

[0191] Example 8. The computing system of any of examples 1 through 7, wherein the first data further comprises an X-ray image associated with the user, and the trained ML model is a foundation model or a fine-tuned ML model configured to: receive X-ray images of users; and determine image features of the X-ray images.

[0192] Example 9. The computing system of example 8, further comprising a prediction model that is configured to: receive the image features from the foundation model; and determine a predicted metric value based on the image features.

[0193] Example 10. The computing system of example 9, wherein the prediction model is further configured to generate a predicted multimedia output based on the image features that correspond to the predicted value.

[0194] Example 11 . The computing system of example 10, wherein the prediction model is a multimedia model that is trained using historical multimedia data of a plurality of patients and historical metric data of the plurality of patients.

[0195] Example 12. The computing system of any of examples 8 through 11 , wherein: the foundation model is configured to receive one or more of: (i) X-ray images, (ii) computed tomography (CT) scans, (iii) magnetic resonance imaging (MRI) images, (iv) ultrasound images, (v) positron emission tomography (PET) scans, (vi) single photon emission computed tomography (SPECT) scans, (vii) optical coherence tomography (OCT) scans, or (viii) digital pathology slide; and the image features are associated with (i) an ejection fraction of the user, (ii) a cardiac output of the user, (iii) a ventricular volume of the user, (iv) a wall thickness of the user, (v) a motion abnormality of the user, (vi) a valvular function of the user, (vi) a cardiac morphology of the user, (vii) a blood flow pattern of the user, and (viii) a blood flow velocity of the user.

[0196] Example 13. A computing system for optimizing summary generation using generative artificial intelligence (Al) models, comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including an outcome prediction for the user by: aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, determining the output based on the input, the prior input, and the test value, and wherein the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value, and cause the outcome prediction to be displayed at a device.

[0197] Example 14. The computing system of example 13, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: receive a subsequent input related to the output; generate, by executing the trained ML model, a subsequent output including a subsequent outcome prediction by: aggregating (i) the input and (ii) the output, determining the subsequent output based on the subsequent input, the input, and the output, and wherein the subsequent outcome prediction includes one or more of: (i) a subsequent differential diagnosis,(ii) a subsequent potential complication, (iii) a subsequent operative value, (iv) a subsequent length of stay value, or (v) a subsequent rehospitalization value; and cause the subsequent outcome prediction to be displayed at a subsequent device.

[0198] Example 15. The computing system of example 13 or 14, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: aggregate a set of outcome predictions for a plurality of users at a first location; generate, by executing the trained ML model, a capacity prediction for the first location based on the set of outcome predictions; and cause the capacity prediction to be displayed at the device.

[0199] Example 16. The computing system of any of examples 13 through 15, wherein the input is received from a provider with a subspeciality, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to determine the output by: selectively including a subset of data from one or more of the input, the prior input, or the test value in the output based on the subspeciality of the provider, and wherein the output includes a history of present illness (HPI) summary based on the subset of data.

[0200] Example 17. The computing system of example 16, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: generate, by executing the trained ML model, the outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.

[0201] Example 18. The computing system of example 16 or 17, wherein the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to generate the output by: determining, by executing the trained ML model, a second subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the second subset of data.

[0202] Example 19. The computing system of any of examples 13 through 18, wherein the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an outcome source and (b) a multi-modal ML model.

[0203] Example 20. A computing system for optimizing summary generation using generative artificial intelligence (Al) models, comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary and an outcome prediction by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, determining the HPI summary based on the input, the subspeciality of the provider, the prior input, and the test value, and determining the outcome prediction based on the HPI summary, and cause the HPI summary and the outcome prediction to be displayed at a device.Additional Considerations

[0204] The following considerations also apply to the foregoing discussion. Throughout this specification, plural instances may implement operations or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.

[0205] It should also be understood that, unless a term is expressly defined in this patent using the sentence "As used herein, the term " " is hereby defined to mean . . . " or a similar sentence, there is no intent to limit the meaning of that term, either expressly or by implication, beyond its plain or ordinary meaning, and such term should not be interpreted to be limited in scope based on any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of this patent is referred to in this patent in a manner consistent with a single meaning, that is done for sake of clarity only so as to not confuse the reader, and it is not intended that such claim term be limited, by implication or otherwise, to that single meaning. Finally, unless a claim element is defined by reciting the word "means" and a function without the recital of any structure, it is not intended that the scope of any claim element be interpreted based on the application of 35 ll.S.C. § 112(f).

[0206] Unless specifically stated otherwise, discussions herein using words such as "processing," "computing," "calculating," "determining," "presenting," "displaying," or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities withinone or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.

[0207] As used herein any reference to "one aspect" or "an aspect" means that a particular element, feature, structure, or characteristic described in connection with the aspect is included in at least one aspect. The appearances of the phrase "in one aspect" in various places in the specification are not necessarily all referring to the same aspect.

[0208] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having" or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).

[0209] In addition, use of "a" or "an" is employed to describe elements and components of the aspects herein. This is done merely for convenience and to give a general sense of the invention. This description should be read to include one or at least one and the singular also includes the plural unless it is obvious that it is meant otherwise.

[0210] The various embodiments described above can be combined to provide further embodiments. All U.S. patents, U.S. patent application publications, U.S. patent application, foreign patents, foreign patent application and non-patent publications referred to in this specification and / or listed in the Application Data Sheet are incorporated herein by reference, in their respective entireties, for all purposes. Aspects of the embodiments can be modified if necessary to employ concepts of the various patents, applications, and publications to provide yet further embodiments.

[0211] These and other changes can be made to the embodiments in light of the abovedetailed description. In general, in the following claims, the terms used should not be construed to limit the claims to the specific embodiments disclosed in the specification and the claims but should be construed to include all possible embodiments along with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the disclosure.

[0212] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs for implementing the concepts disclosed herein, through the principles disclosed herein. Thus, while particular aspects and applications havebeen illustrated and described, it is to be understood that the disclosed aspects are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in the arrangement, operation and details of the method and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.

Claims

What is claimed is:1 . A computing system for optimizing summary generation using generative artificial intelligence (Al) models, comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary by: identifying a subspeciality of the provider, aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, and determining the output based on the input, the subspeciality of the provider, the prior input, and the test value, and cause the HPI summary to be displayed at a device.

2. The computing system of claim 1 , the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to determine the output by: selectively including a subset of data from one or more of the input, the prior input, or the test value in the HPI summary based on the subspeciality of the provider.

3. The computing system of claim 1 , the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: generate, by executing the trained ML model, an outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.

4. The computing system of claim 3, wherein the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value.

5. The computing system of claim 1 , the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: receive a subsequent input related to the output from a subsequent provider; generate, by executing the trained ML model, a subsequent output including a subsequent HPI summary by: identifying a subsequent subspeciality of the subsequent provider, wherein the subsequent subspeciality is different from the subspeciality of the provider, aggregating (i) the input and (ii) the output, and determining the subsequent HPI summary based on the subsequent input, the subsequent subspeciality of the subsequent provider, the input, and the output; and cause the subsequent HPI summary to be displayed at a subsequent device.

6. The computing system of claim 1 , wherein the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to generate the output by: determining, by executing the trained ML model, a subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the subset of data.

7. The computing system of claim 1 , wherein the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an HPI source and (b) a multi-modal ML model.

8. The computing system of claim 1 , wherein the first data further comprises an X- ray image associated with the user, and the trained ML model is a foundation model or a finetuned ML model configured to: receive X-ray images of users; and determine image features of the X-ray images.

9. The computing system of claim 8, further comprising a prediction model that is configured to:receive the image features from the foundation model; and determine a predicted metric value based on the image features.

10. The computing system of claim 9, wherein the prediction model is further configured to generate a predicted multimedia output based on the image features that correspond to the predicted value.11 . The computing system of claim 10, wherein the prediction model is a multimedia model that is trained using historical multimedia data of a plurality of patients and historical metric data of the plurality of patients.

12. The computing system of claim 8, wherein: the foundation model is configured to receive one or more of: (i) X-ray images, (ii) computed tomography (CT) scans, (iii) magnetic resonance imaging (MRI) images, (iv) ultrasound images, (v) positron emission tomography (PET) scans, (vi) single photon emission computed tomography (SPECT) scans, (vii) optical coherence tomography (OCT) scans, or (viii) digital pathology slide; and the image features are associated with (i) an ejection fraction of the user, (ii) a cardiac output of the user, (iii) a ventricular volume of the user, (iv) a wall thickness of the user, (v) a motion abnormality of the user, (vi) a valvular function of the user, (vi) a cardiac morphology of the user, (vii) a blood flow pattern of the user, and (viii) a blood flow velocity of the user.

13. A computing system for optimizing summary generation using generative artificial intelligence (Al) models, comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including an outcome prediction for the user by: aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, determining the output based on the input, the prior input, and the test value, andwherein the outcome prediction includes one or more of: (i) a differential diagnosis, (ii) a potential complication, (iii) an operative value, (iv) a length of stay value, or (v) a rehospitalization value, and cause the outcome prediction to be displayed at a device.

14. The computing system of claim 13, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: receive a subsequent input related to the output; generate, by executing the trained ML model, a subsequent output including a subsequent outcome prediction by: aggregating (i) the input and (ii) the output, determining the subsequent output based on the subsequent input, the input, and the output, and wherein the subsequent outcome prediction includes one or more of: (i) a subsequent differential diagnosis, (ii) a subsequent potential complication, (iii) a subsequent operative value, (iv) a subsequent length of stay value, or (v) a subsequent rehospitalization value; and cause the subsequent outcome prediction to be displayed at a subsequent device.

15. The computing system of claim 13, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: aggregate a set of outcome predictions for a plurality of users at a first location; generate, by executing the trained ML model, a capacity prediction for the first location based on the set of outcome predictions; and cause the capacity prediction to be displayed at the device.

16. The computing system of claim 13, wherein the input is received from a provider with a subspeciality, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to determine the output by: selectively including a subset of data from one or more of the input, the prior input, or the test value in the output based on the subspeciality of the provider, andwherein the output includes a history of present illness (HPI) summary based on the subset of data.

17. The computing system of claim 16, the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to: generate, by executing the trained ML model, the outcome prediction for the user based on one or more of the input, the subspeciality of the provider, the prior input, or the test value; and cause the outcome prediction and the HPI summary to be displayed at the device.

18. The computing system of claim 16, wherein the input includes a subspeciality request indicating a second subspeciality that is different from the subspeciality of the provider, and the one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, further cause the computing system to generate the output by: determining, by executing the trained ML model, a second subset of data from one or more of the input, the prior input, or the test value based on the second subspeciality; and generating, by executing the trained ML model, the output based on the second subset of data.

19. The computing system of claim 13, wherein the trained ML model is (a) a large language model (LLM) pre-trained, fine-tuned and / or trained using information from an outcome source and (b) a multi-modal ML model.

20. A computing system for optimizing summary generation using generative artificial intelligence (Al) models, comprising: one or more processors; and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: receive an input from a provider, the input including first data associated with a user, generate, by executing a trained machine learning (ML) model, an output including a history of present illness (HPI) summary and an outcome prediction by: identifying a subspeciality of the provider,aggregating (i) a prior input including second data associated with the user and (ii) a test value corresponding to the prior input, determining the HPI summary based on the input, the subspeciality of the provider, the prior input, and the test value, and determining the outcome prediction based on the HPI summary, and cause the HPI summary and the outcome prediction to be displayed at a device.

Citation Information

Patent Citations

  • Records access and management

    CN110462654A

  • Multi-label disease auxiliary diagnosis system based on physician decision-making pattern recognition

    CN116386856B

  • Cosmetic Composition Containing Mixture Extracts of Phaseolus Radiatus Seed, Betula Alba Juice and Rumex Crispus Root

    KR1020210025757A

  • System and method for medical image interpretation

    US10445462B2

  • Platforms for conducting virtual trials

    WO2019144116A1

Cited By

  • Cross-modal learning preference optimization enhanced three-dimensional face generation method and system

    CN120510260A