Healthcare documents generated with natural language processing

Advanced NLP systems generate personalized healthcare documents using patient data to address clarity and completeness issues in existing healthcare documents, improving patient understanding and adherence.

WO2025240180A1PCT designated stage Publication Date: 2025-11-2023ANDME INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/028124
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2025-05-07
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Healthcare documents, such as after visit summaries, lack clarity, precision, completeness, and personalization due to freeform writing by healthcare providers, leading to patient confusion and suboptimal health outcomes.

Method used

Utilizing advanced natural language processing systems, like large language models, to generate personalized healthcare documents by combining patient medical and genetic data, ensuring accurate and empathetic communication of health recommendations.

Benefits of technology

Improves document clarity and completeness, reduces wastage of resources, and enhances patient understanding and adherence to medical recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028124_20112025_PF_FP_ABST
    Figure US2025028124_20112025_PF_FP_ABST
Patent Text Reader

Abstract

An example embodiment may involve receiving, by way of a user interface of a computing system, an indication of a healthcare patient profile and further input; based on the indication of the healthcare patient profile, obtaining, by the computing system, health records and genetic data of a patient associated with the healthcare patient profile; based on the health records and the genetic data of the patient and the further input, generating, by the computing system, a natural language prompt, wherein the natural language prompt contains instructions to generate a healthcare document for the patient; providing, by the computing system, the natural language prompt to a generative natural language model; receiving, by the computing system, the healthcare document from the generative natural language model; and providing, by way of the user interface, a version of the healthcare document.
Need to check novelty before this filing date? Find Prior Art

Description

Healthcare Documents Generated with Natural Language ProcessingCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. provisional patent application no. 63 / 648,376, filed May 16, 2024, which is hereby incorporated by reference in its entirety.BACKGROUND

[0002] Healthcare documents (such as after visit summaries (AVSs)) and instructions that are provided to patients regarding important health-related factors (such as diet, exercise, lifestyle, medications, and disease risks) need to be clear, precise, complete, and correct. In today’s healthcare systems, however, these healthcare documents are largely written in a freeform fashion by health care providers (e.g., physicians, residents, nurse practitioners, clinicians, and / or administrative personnel). As a result, healthcare documents can vary dramatically in quality and content. This, in turn, can result in patient confusion or - even worse - patients not following best medical practices because these practices are not clearly set forth. Further, modem healthcare documents lack analysis of genetic data that may inform health care providers and patients of their risks for various diseases and conditions.SUMMARY

[0003] The embodiments herein involve using advanced natural language processing systems, such as large language models (LLMs), to generate healthcare documents. The generated documents may include AVSs, personalized health plans (PHPs), medication dosage instructions, and so on. A natural language processing system may be prompted with information relating to a patient’s medical chart (e.g., identifying information, medical history, medication list, allergy information, immunization record, family history, physician notes, patient comments, treatment plans, discharge summaries, and so on) as well as the patient’s genetic data (e.g., single nucleotide polymorphisms (SNPs) indicative of disease risk and / or polygenic risk scores (PRSs)).

[0004] The combination of this information can be used to provide detailed and personalized healthcare documents that are more accurate and easier to understand than those traditionally supplied. An LLM, for example, can provide output that is personal and empathic, suggesting how a patient can maintain or improve their baseline health as well as justify any healthcare recommendation through consideration of a combination of risk factors.

[0005] In particular, the embodiments herein may develop sophisticated andcomprehensive prompts that, when submitted to an LLM, cause the LLM to generate such documents. As a consequence, not only are these documents improved, but patient health outcomes may also be improved correspondingly. Further, the clarity and completeness of the documents generated through these natural language processing systems results in less wastage and improved utilization of computational resources. For example, patients are less likely to refresh pages and navigate extensively through a healthcare provider web site seeking information to bridge the gap between a physician’s terse notes and actionable healthcare recommendations .

[0006] A first example embodiment may involve receiving, by way of a user interface of a computing system, an indication of a healthcare patient profile and further input; based on the indication of the healthcare patient profile, obtaining, by the computing system, health records and genetic data of a patient associated with the healthcare patient profile; based on the health records and the genetic data of the patient and the further input, generating, by the computing system, a natural language prompt, wherein the natural language prompt contains instructions to generate a healthcare document for the patient; providing, by the computing system, the natural language prompt to a generative natural language model; receiving, by the computing system, the healthcare document from the generative natural language model; and providing, by way of the user interface, a version of the healthcare document.

[0007] In a second example embodiment, an article of manufacture may include a non- transitory computer-readable medium, having stored thereon program instructions that, upon execution by a computing system, cause the computing system to perform operations in accordance with any previous embodiment.

[0008] In a third example embodiment, a computing system may include at least one processor, as well as memory and program instructions. The program instructions may be stored in the memory, and upon execution by the at least one processor, cause the computing system to perform operations in accordance with any previous embodiment.

[0009] In a fourth example embodiment, a system may include various means for carrying out each of the operations of any previous embodiment.

[0010] These, as well as other embodiments, aspects, advantages, and alternatives, will become apparent to those of ordinary skill in the art by reading the following detailed description, with reference where appropriate to the accompanying drawings. Further, this summary and other descriptions and figures provided herein are intended to illustrate embodiments by way of example only and, as such, that numerous variations are possible. For instance, structural elements and process steps can be rearranged, combined, distributed,eliminated, or otherwise changed, while remaining within the scope of the embodiments as claimed.BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 illustrates a schematic drawing of a computing device, in accordance with example embodiments.

[0012] Figure 2 illustrates a schematic drawing of a server device cluster, in accordance with example embodiments.

[0013] Figure 3 depicts an LLM interface, in accordance with example embodiments.

[0014] Figure 4 depicts a workflow for generating healthcare documents, in accordance with example embodiments.

[0015] Figure 5 depicts an iterative evaluation process, in accordance with example embodiments.

[0016] Figure 6 depicts evaluation of synthetic patient notes and synthetic patient AVSs, in accordance with example embodiments.

[0017] Figure 7A depicts a document generation prompt template, in accordance with example embodiments.

[0018] Figure 7B depicts a schema generation prompt template, in accordance with example embodiments.

[0019] Figures 8A and 8B provide example LLM-generated AVSs, in accordance with example embodiments.

[0020] Figure 9 is a flow chart, in accordance with example embodiments.DETAILED DESCRIPTION

[0021] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any embodiment or feature described herein as being an “example” or “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or features unless stated as such. Thus, other embodiments can be utilized and other changes can be made without departing from the scope of the subject matter presented herein.

[0022] Accordingly, the example embodiments described herein are not meant to be limiting. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined,separated, and designed in a wide variety of different configurations. For example, the separation of features into “client” and “server” components may occur in a number of ways.

[0023] Further, unless context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be generally viewed as component aspects of one or more overall embodiments, with the understanding that not all illustrated features are necessary for each embodiment.

[0024] Additionally, any enumeration of elements, blocks, or steps in this specification or the claims is for purposes of clarity. Thus, such enumeration should not be interpreted to require or imply that these elements, blocks, or steps adhere to a particular arrangement or are carried out in a particular order.I. Example Computing Devices and Cloud-Based Computing Environments

[0025] Figure 1 is a simplified block diagram exemplifying a computing device 100, illustrating some of the components that could be included in a computing device arranged to operate in accordance with the embodiments herein. Computing device 100 could be a client device (e.g., a device actively operated by a user), a server device (e.g., a device that provides computational services to client devices), or some other type of computational platform. Some server devices may operate as client devices from time to time in order to perform particular operations, and some client devices may incorporate server features.

[0026] In this example, computing device 100 includes processor 102, memory 104, network interface 106, and input / output unit 108, all of which may be coupled by system bus 110 or a similar mechanism. In some embodiments, computing device 100 may include other components and / or peripheral devices (e.g., detachable storage, printers, and so on).

[0027] Processor 102 may be one or more of any type of computer processing element, such as a central processing unit (CPU), a co-processor (e.g., a mathematics, graphics, or encryption co-processor), a digital signal processor (DSP), a network processor, and / or a form of integrated circuit or controller that performs processor operations. In some cases, processor 102 may be one or more single-core processors. In other cases, processor 102 may be one or more multi-core processors with multiple independent processing units. Processor 102 may also include register memory for temporarily storing instructions being executed and related data, as well as cache memory for temporarily storing recently-used instructions and data.

[0028] Memory 104 may be any form of computer-usable memory, including but not limited to random access memory (RAM), read-only memory (ROM), and non-volatile memory (e.g., flash memory, hard disk drives, solid state drives, compact discs (CDs), digital video discs (DVDs), and / or tape storage). Thus, memory 104 represents both main memoryunits, as well as long-term storage. Other types of memory may include biological memory.

[0029] Memory 104 may store program instructions and / or data on which program instructions may operate. By way of example, memory 104 may store these program instructions on a non-transitory, computer-readable medium, such that the instructions are executable by processor 102 to carry out any of the methods, processes, or operations disclosed in this specification or the accompanying drawings.

[0030] As shown in Figure 1, memory 104 may include firmware 104A, kernel 104B, and / or applications 104C. Firmware 104A may be program code used to boot or otherwise initiate some or all of computing device 100. Kernel 104B may be an operating system, including modules for memory management, scheduling and management of processes, input / output, and communication. Kernel 104B may also include device drivers that allow the operating system to communicate with the hardware modules (e.g., memory units, networking interfaces, ports, and buses) of computing device 100. Applications 104C may be one or more user-space software programs, such as web browsers or email clients, as well as any software libraries used by these programs. Memory 104 may also store data used by these and other programs and applications.

[0031] Network interface 106 may take the form of one or more wireline interfaces, such as Ethernet (e.g., Fast Ethernet, Gigabit Ethernet, and so on). Network interface 106 may also support communication over one or more non-Ethemet media, such as coaxial cables or power lines, or over wide-area media, such as Synchronous Optical Networking (SONET) or digital subscriber line (DSL) technologies. Network interface 106 may additionally take the form of one or more wireless interfaces, such as IEEE 802.11 (WIFI), BLUETOOTH, global positioning system (GPS), or a wide-area wireless interface. However, other forms of physical layer interfaces and other types of standard or proprietary communication protocols may be used over network interface 106. Furthermore, network interface 106 may comprise multiple physical interfaces. For instance, some embodiments of computing device 100 may include Ethernet, BLUETOOTH, and WIFI interfaces.

[0032] Input / output unit 108 may facilitate user and peripheral device interaction with computing device 100. Input / output unit 108 may include one or more types of input devices, such as a keyboard, a mouse, a touch screen, and so on. Similarly, input / output unit 108 may include one or more types of output devices, such as a screen, monitor, printer, and / or one or more light emitting diodes (LEDs). Additionally or alternatively, computing device 100 may communicate with other devices using a universal serial bus (USB) interface or high-definition multimedia interface (HDMI) port, for example.

[0033] One or more computing devices like computing device 100 may be deployed to support the embodiments herein. The exact physical location, connectivity, and configuration of these computing devices may be unknown and / or unimportant to client devices. Accordingly, the computing devices may be referred to as “cloud-based” devices that may be housed at various remote data center locations.

[0034] Figure 2 depicts a cloud-based server cluster 200 in accordance with example embodiments. In Figure 2, operations of a computing device (e.g., computing device 100) may be distributed between server devices 202, data storage 204, and routers 206, all of which may be connected by local cluster network 208. The number of server devices 202, data storages 204, and routers 206 in server cluster 200 may depend on the computing task(s) and / or applications assigned to server cluster 200.

[0035] For example, server devices 202 can be configured to perform various computing tasks of computing device 100. Thus, computing tasks can be distributed among one or more of server devices 202. To the extent that these computing tasks can be performed in parallel, such a distribution of tasks may reduce the total time to complete these tasks and return a result. For purposes of simplicity, both server cluster 200 and individual server devices 202 may be referred to as a “server device.” This nomenclature should be understood to imply that one or more distinct server devices, data storage devices, and cluster routers may be involved in server device operations.

[0036] Data storage 204 may be data storage arrays that include drive array controllers configured to manage read and write access to groups of hard disk drives and / or solid state drives. The drive array controllers, alone or in conjunction with server devices 202, may also be configured to manage backup or redundant copies of the data stored in data storage 204 to protect against drive failures or other types of failures that prevent one or more of server devices 202 from accessing units of data storage 204. Other types of memory aside from drives may be used.

[0037] Routers 206 may include networking equipment configured to provide internal and external communications for server cluster 200. For example, routers 206 may include one or more packet-switching and / or routing devices (including switches and / or gateways) configured to provide (i) network communications between server devices 202 and data storage 204 via local cluster network 208, and / or (ii) network communications between server cluster 200 and other devices via communication link 210 to network 212.

[0038] Additionally, the configuration of routers 206 can be based at least in part on the data communication requirements of server devices 202 and data storage 204, the latencyand throughput of the local cluster network 208, the latency, throughput, and cost of communication link 210, and / or other factors that may contribute to the cost, speed, faulttolerance, resiliency, efficiency, and / or other design goals of the system architecture.

[0039] As a possible example, data storage 204 may include any form of database, such as a structured query language (SQL) database. Various types of data structures may store the information in such a database, including but not limited to tables, arrays, lists, trees, and tuples. Furthermore, any databases in data storage 204 may be monolithic or distributed across multiple physical devices.

[0040] Server devices 202 may be configured to transmit data to and receive data from data storage 204. This transmission and retrieval may take the form of SQL queries or other types of database queries, and the output of such queries, respectively. Additional text, images, video, and / or audio may be included as well. Furthermore, server devices 202 may organize the received data into web page or web application representations. Such a representation may take the form of a markup language, such as HTML, the extensible Markup Language (XML), JavaScript Object Notation (JSON), or some other standardized or proprietary format. Moreover, server devices 202 may have the capability of executing various types of computerized scripting languages, such as but not limited to Perl, Python, PHP Hypertext Preprocessor (PHP), Active Server Pages (ASP), JAVASCRIPT, and so on. Computer program code written in these languages may facilitate the providing of web pages to client devices, as well as client device interaction with the web pages. Alternatively or additionally, JAVA may be used to facilitate generation of web pages and / or to provide web application functionality.II. Example Large Language Models

[0041] Various embodiments herein may employ LLMs to perform certain tasks. Doing so is advantageous because these models have capabilities that surpass previous techniques in the fields of natural language understanding, natural language generation, knowledge aggregation, information retrieval, pattern recognition, and data analysis. Thus, before describing the main embodiments in detail, it is helpful to consider the operation and capabilities of LLMs.

[0042] An LLM is an advanced computational model, primarily functioning within the domain of natural language processing and machine learning. An LLM can be configured to understand, interpret, generate, and respond to human language in a manner that is both contextually relevant and syntactically coherent. The underlying structure of an LLM is typically based on a neural network architecture, more specifically, a variant of the transformer model. Transformers are notable for their ability to process sequential data, such as text, withhigh efficiency.

[0043] The operation of a neural-network-based LLM involves layers of interconnected processing units, known as neurons, which collectively form a deep neural network. This network can be trained on vast datasets comprising text from diverse sources, thereby enabling the LLM to learn a wide array of language patterns, structures, and colloquial nuances for prose, poetry, and program code. The training process involves adjusting the weights of the connections between neurons using algorithms such as backpropagation, in conjunction with optimization techniques like stochastic gradient descent, to minimize the difference between the LLM’s output and expected output.

[0044] An aspect of an LLM’s functionality is its use of attention mechanisms, particularly self-attention, within the transformer architecture. These mechanisms allow the model to weigh the relative importance of different parts of the input text differently, enabling the model to focus on relevant aspects of the data when generating responses or analyzing language. The self-attention mechanism facilitates the model’s ability to generate contextually relevant and coherent text by understanding the relationships and dependencies between words or tokens in a sentence (or longer parts of texts), regardless of their position.

[0045] Upon receiving an input, such as a text query or a prompt, the LLM may process this input through multiple layers, generating a probabilistic model of the language therein. It predicts the likelihood of each word or token that might follow the given input, based on the patterns it has learned during its training. The model then generates an output, which could be a continuation of the input text, an answer to a query, or other relevant textual content, by selecting words or tokens that have the highest probability of being contextually appropriate.

[0046] Further, an LLM can be fine-tuned after initial training for specific applications or tasks. This fine-tuning process involves additional training (e.g., with reinforcement from humans), usually on a smaller, task-specific dataset, which allows the model to adapt its responses to suit particular use cases more accurately. This adaptability makes LLMs highly versatile and applicable in various domains, including but not limited to, chatbot development, content creation, language translation, and sentiment analysis.

[0047] Some LLMs are multimodal in that they can receive prompts in formats other than text and can produce outputs in formats other than text. Thus, while LLMs are predominantly designed for understanding and generating textual data, multimodal LLMs extend this functionality to include multiple data modalities, such as visual and auditory inputs, in addition to text.

[0048] A multimodal LLM can employ an advanced neural network architecture, oftena variant of the transformer model that is specifically adapted to process and fuse data from different sources. This architecture integrates specialized mechanisms, such as convolutional neural networks for visual data and recurrent neural networks for audio processing, allowing the model to effectively process each modality before synthesizing a unified output.

[0049] The training of a multimodal LLM involves multimodal datasets, enabling the model to learn not only language patterns but also the correlations and interactions between different types of data. This cross-modal training results in multimodal LLMs being adept at tasks that require an understanding of complex relationships across multiple data forms, a capability that text-only LLMs do not possess. This makes multimodal LLMs particularly suited for advanced applications that necessitate a holistic understanding of multimodal information, such as chatbots that can interpret and produce images and / or audio.

[0050] Some LLMs can employ a form of reasoning by transforming user-provided natural language input into a structured internal representation, which is then used to generate a response based on learned statistical associations and contextual dependencies. For complex queries requiring multi-step inference (e.g., mathematical problems, logical deductions, or temporal event analysis), the LLM can decompose the problem into intermediate steps by generating latent representations that reflect intermediate reasoning states. These states are influenced by the training data and are shaped during inference through techniques such as masked self-attention, positional encoding, and next-token prediction, which collectively allow the LLM to simulate deductive or abductive reasoning processes. In some embodiments, explicit chain-of-thought prompting is used to guide the LLM in generating a sequence of intermediate natural language steps, enhancing the LLM’s ability to produce structured reasoning outputs. This capability enables the LLM to generate responses that are not merely retrieved facts but the result of multi-step inferential synthesis over the input data and the LLM’s internal knowledge representation.

[0051] Given that a significant portion of patient medical information and genetic information is in textual or image form, an LLM may be able to assist with the generation of healthcare documents from this information. However, naively attempting to use an LLM in this manner (e.g., where the LLM has been trained as a foundational model using only general- purpose knowledge) might not produce the desired results.

[0052] Notably, a general-purpose LLM can be designed to perform a wide range of language understanding and generation tasks across various domains without specialization. Such an LLM may have been trained on a diverse and broad dataset covering multiple fields, topics, and types of language use. This makes general -purpose LLMs versatile and capable ofhandling a wide array of tasks, but they may not produce domain-specific results with the desired proficiency and nuance. Therefore, it may be desirable to employ a “wrapper” around such an LLM so that intent determination from LLM prompt input may be performed more accurately and robustly.

[0053] Figure 3 depicts such an architecture. LLM interface 300 may include LLM prompt pre-processor 302, LLM response post-processor 304, patient health records 306A, patient genetic data 306B, and / or other data 306C. Additional features of LLM interface 300 may also be present.

[0054] Patient health records 306A, patient genetic data 306B, and / or other data 306C may influence the operation of LLM prompt pre-processor 302 and LLM response postprocessor 304. Examples of content that may be stored for each of these items is discussed below.

[0055] Patient health records 306A may include a wide array of information relating to comprehensive medical care. These records may include personal identification details, such as name, age, and / or contact information; a thorough medical history documenting past and current illnesses, surgeries, and / or any chronic conditions; a detailed medication list with dosages and / or prescribing details; known allergies; immunization records; and / or family medical history. They may also contain records of vital signs over time; diagnostic test results such as blood tests, and / or imaging studies; progress notes from healthcare providers; reports from consultations with specialists; and / or any treatment plans or interventions. Additionally, health records may include past after visit summaries, discharge summaries from hospital stays, consent forms for treatments and procedures, and / or instructions for ongoing care or rehabilitation.

[0056] Patient genetic data 306B may include representations of the patient’s SNPs, which are the most common type of genetic variation among people. SNPs can influence how patients respond to drugs, affect a patient’s risk of developing certain diseases, and guide personalized treatment strategies. Another type of genetic data is a PRS, which aggregates the effects of many SNPs to estimate an individual’s predisposition to various diseases, such as cardiovascular diseases, diabetes, and / or cancer. PRSs can be instrumental in preventive healthcare by identifying individuals at high risk of certain diseases or conditions who may benefit from early interventions. Other genetic information may include whole genome sequencing data, gene expression profiles, and / or epigenetic markers, all of which contribute to a deeper understanding of a patient’s health status and potential health risks.

[0057] Other data 306C may include any additional information that can be useful inevaluating a patient’s health and providing an effective after visit summary. Lifestyle information, such as dietary habits, physical activity levels, alcohol consumption, and / or smoking status, offers insights into risk factors and preventative health measures. Environmental data, including exposure to pollutants or occupational hazards, can also significantly influence health assessments. Psychological and social factors, such as stress levels, mental health history, social support systems, and / or economic status, are useful for a holistic health evaluation, impacting both physical health and the effectiveness of medical interventions. Additionally, patient-reported outcomes and self-monitored health data from devices like fitness trackers, blood glucose monitors, or blood pressure monitors can provide real-time insights into a patient’s daily health status and / or compliance with prescribed treatments.

[0058] As noted, LLM interface 300 serves as a “wrapper” around LLM 310. In other words, representations of provider input (e.g., from a healthcare provider) may be received by LLM prompt pre-processor 302. LLM prompt pre-processor 302 may generate an LLM prompt from healthcare provider input and information from one or more of patient health records 306A, patient genetic data 306B, and / or other data 306C. LLM prompt pre-processor 302 may transmit the LLM prompt to LLM 310. In turn, LLM 310 may perform natural language processing tasks to determine an LLM response to the LLM prompt. LLM 310 may provide this LLM response to LLM interface 300, which routes it to LLM response post-processor 304. LLM response post-processor 304 may modify, edit, and / or select words, tokens, and / or other items from the LLM response and provide it as a healthcare document (e.g., an after visit summary).

[0059] In some embodiments, LLM 310 may be a general-purpose LLM that is not specifically trained to interpret health care records or to generate healthcare documents. In other embodiments, LLM 310 may have been trained and / or fine-tuned to perform these tasks.

[0060] Notably, LLM interface 300 may operate on a client device (e.g., a desktop computer, laptop computer, smartphone, or tablet operated by the provider. Alternatively, LLM interface 300 may operate on a server device. LLM 310 may be disposed on such a server device or within a remote cloud-based network accessible to the server device.

[0061] Some embodiments may involve LLM response post-processor 304 providing at least part of the LLM response back to LLM prompt pre-processor 302. In this manner, a sequence of LLM prompts may be generated, each building upon the previous LLM response until the healthcare document is sufficiently completed. In such embodiments, one or more routines may be performed by LLM response post-processor 304 to determine whether an LLMresponse is sufficiently completed (e.g., and ready to be output by the LLM interface 300 as a healthcare document) or whether one or more portions of the LLM response should be fed back into the LLM prompt pre-processor 302 in order to generate another LLM prompt.III. Healthcare Document Generation Workflow

[0062] Figure 4 depicts an example healthcare document generation workflow 400. Workflow 400 may involve the components of Figure 3 (e.g., LLM interface 300 and LLM 310) in various arrangements. In other words, the functions of the components of Figure 3 may be distributed in various ways across the steps and components of Figure 4. For example, LLM prompt pre-processor 302 and LLM response post-processor 304 may be incorporated into or called by backend 406.

[0063] In Figure 4, healthcare practitioner 402 (e.g., a clinician or other individual) creates a patient profile for a patient and sends it to backend 406 via practitioner user interface 404 (e.g., a dedicated application or a web-based interface). Here, backend 406 represents any medical software that maintains or has access to point of care documents 408. Point of care documents 408 may include patient health records 306A, patient genetic data 306B, and / or other data 306C.

[0064] Practitioner user interface 404 and / or backend 406 may authenticate the healthcare practitioner (e.g., using one or more login credentials) and establish the patient profile (e.g., create database entries for the content of the patient profile, such as point of care documents 408). Healthcare practitioner 402 may also provide content and / or instructions for this system to generate one or more specific healthcare documents.

[0065] Backend 406 may arrange these inputs into a manner that is consistent with the database schema and retrieve one or more documents from point of care documents 408. Point of care documents 408 a database that, as noted, stores patient health records 306A, patient genetic data 306B, and / or other data 306C. Alternatively, in some embodiments, point of care documents 408 may be separate from a further repository that stores patient health records 306A, patient genetic data 306, and / or other data 306C.

[0066] Backend 406 may also construct an LLM prompt based on this aggregated information (e.g., by way of LLM prompt pre-processor 302) and provide the LLM prompt to LLM 310. LLM 310 may generate a representation of the requested healthcare documents and produce an LLM response containing this representation. Backend 406 may receive the representation, conduct any further processing and / or editing of the representation (e.g., by way of LLM response post-processor 304).

[0067] In some cases, this representation includes after visit summary (AVS) data 410,possibly including a personalized health plan (PHP). Note that the term “PHP” in this context is different from the term “PHP Hypertext Preprocessor” discussed above. Based on content, use of the term “PHP” herein can be easily disambiguated.

[0068] Backend 406 may provide the healthcare documents of the representation to patient 414 by way of the patient user interface 412. In some cases, edits by the healthcare provider or another LLM (see below) may be made before the healthcare documents are provided to the patient.

[0069] Notably, practitioner user interface 404 and patient user interface 412 may have different functionality. For instance, practitioner user interface 404 may allow practitioner 404 to view and generate AVSs for multiple patients, while patient user interface may only patient 414 to view their own AVSs but not those of other patients.

[0070] Figure 5 depicts an example of how the LLM 310 can be trained (and / or finetuned) and then evaluated. A number of synthetic patient records 502 can be generated (e.g., a synthetic dataset of patients, test results, and visits), each with the type of information that would be found in actual patient health records 306A, patient genetic data 306B, and / or other data 306C. For example, an LLM (possibly other than LLM 310) may generate these synthetic patient records, perhaps with human editing and assistance. Then backend 406 can be used to generate AVS prompt 504 (a form of LLM prompt) requesting healthcare documents (e.g., an AVS) from LLM 310. LLM 310, in turn, may generate an LLM response with such documents for evaluation 506 (some forms of manual, automated, or semi-automated evaluation could be used). The result of evaluation 506 may involve updates to AVS prompt 504 and / or other procedures. This prompt / response / evaluation / updating cycle may iterate a number of times until backend 406 is capable of generating high quality healthcare documents. The determination of whether the generated documents are “high quality” may be based on specific metrics, rubrics, and / or the presence or absence of certain types of information in the documents (e.g., the presence of recommendations that match those of best practice reference documents and / or the absence of hallucinations).

[0071] In some cases, an automated evaluation framework may be employed in which the generated documents may be automatically checked for certain desirable or undesirable characteristics. Doing so is helpful because updates to LLM 310 and backend 406 may be made over time, and these updates may introduce defects or unpredictable LLM behavior. Further, it is known that LLMs sometimes generate incorrect, irrelevant, biased, or nonsensical information. In some cases, the automated evaluation framework may be performed by a further LLM that is capable of or specifically trained to evaluate generated documents basedon pre-established metrics.

[0072] Examples of metrics used to evaluate generated documents include: hallucinations (e.g., is LLM 310 providing factually correct information?), personalization (e.g., is LLM 310 able to combine multiple pieces of information and personalize the response?), utility (e.g., is the system as a whole improving productivity such as time taken for review, number of edits), bias (e.g., is LLM 310 providing biased outputs towards any particular groups of people?), and / or edited-version similarity (e.g., how similar are the LLM-generated documents and those documents after editing?). Of these metrics, six (five relating to hallucination and one relating to personalization) were found to be sufficient for general usage. In testing against other techniques (e.g., based on the Ragas or MLflow frameworks), the automated evaluation framework presented herein provides better results.

[0073] This automated evaluation framework may break down the metrics calculated into several distinct categories. They can be as follows. More detail on aspects and alternative embodiments of the automated evaluation framework are discussed in the section below.

[0074] Category 1 : ground-truth based similarity - how similar is the LLM-generated version to the post-editing version? This metric compares a healthcare-provider-approved AVS to the LLM-generated AVS, directly captures the quality of the LLM-generated AVS. It may directly interrogate some of the edits made to the LLM-generated content, acting as a proxy for the categories below, as well as relevance, formatting, and tone. This metric might not capture any inaccuracies, biases, or otherwise inappropriate information that is approved by the clinician, for example, if the LLM-generated content contains a factual inaccuracy, and the healthcare provider does not catch the inaccuracy and approves the AVS to be sent to the patient, this category of metrics might not measure that.

[0075] This type of metric may be evaluated based on using other models to evaluate the similarity of the LLM-generated content and the clinician-edited and approved content. This may initially be prototyped using the AVS data as well as the stock phrases to develop a topology. Also, traditional statistical similarity metrics such as word count, semantic similarity, etc. may be used.

[0076] Category 2: medical reasoning - is the model using effective medical reasoning given the provided inputs? Despite best attempts to provide LLM 310 with the necessary medical information to craft a well-reasoned set of recommendations, it is possible that LLM 310 may experience lapses in reasoning. While reasoning, itself, is something regularly measured in LLM-generated content, these metrics directly interrogate reasoning specifically in the medical domain by using a set of patient vignettes that require LLM 310 to reason aboutsome of the medical knowledge. An example of a lapse in medical reasoning could be when a male patient presents with a BRCA 1 variant, and the LLM-generated AVS recommends regular transvaginal ultrasounds to screen for ovarian cancer.

[0077] This type of metric may be evaluated using other models to test the reasoning of an AVS output given patient vignette(s) and a scoring topology. Also, a limited benchmark dataset with patient vignettes that have been crafted to specifically test certain medical reasoning tasks may be used.

[0078] Category 3: Hallucination evaluation - is the model providing relevant and correct information? The model can pull in external information and said information is generally correct. These metrics directly interrogate the LLM-generated AVS without any additional information provided. These metrics will use an evaluation LLM (e.g., different from LLM 310) to analyze the AVS. These metrics may not capture inaccuracies that the evaluation LLM cannot recognize, such as some detailed medical information. Alternatively, the model properly uses in-context information (context adherence). These metrics compare the LLM- generated AVS to the context provided in an attempt to detect any inaccuracies. These metrics may use another evaluation LLM as well as basic statistical / comparison metrics. These metrics may not capture inaccuracies that are not explicitly outlined in the context.

[0079] This type of metric may be evaluated based on the correctness and context adherence of the output AVS. Various methods may be used, employing both statistical comparison-based methods as well as using other models to evaluate the output AVSs.

[0080] Category 4: Bias metrics - does the model contain any biases against any particular groups of people, conditions, etc.? It is known that LLMs can exhibit biases, and that the healthcare system, as a whole, also contains biases. While these metrics might not be fully comprehensive in this space, such metrics can capture some of the known and common biases in the medical domain in the LLM-generated recommendations. These metrics may use a comparison-based methodology where patient vignettes are perturbed to analyze what difference those perturbations have on the output recommendations. An example of potential bias in recommendations: A female, middle-aged patient presents as somewhat overweight and her genetic results show she is at higher risk for coronary artery disease. Instead of prescribing a preventative medication coupled with lifestyle recommendations, the LLM heavily focuses on weight loss recommendations and only mentions the possibility of helpful medications as a side-note. This may depict bias due to the patient’s gender and also due to the patient being obese.

[0081] This type of metric may be evaluated based on a benchmark dataset with patientvignetes that have been perturbed in such a way to test potential biases in the LLM-generated output (e.g. the same person where the only difference is ethnicity). These metrics will compare the outputs of the baseline patient vignete to the outputs of the perturbed patient vignetes to determine differences.

[0082] Category 5 : Utility metrics - is the system as a whole improving productivity? Initially, utility metrics may be determined based on anecdotal reports from clinicians, such as time saved and feedback on the LLM-generated AVSs themselves. In the longer term, utility metrics may be determined by monitoring and / or may be automatically collected.

[0083] These metrics may be evaluated based on initial time estimates of a productivity increase when clinicians use the LLM-generated AVSs compared to when clinicians manually write the AVSs. Also, basic metrics on the outputs of LLM 310, such as response length, correct formating, etc. may be used.

[0084] In practice, this approach has resulted in the embodiments herein producing healthcare documents that are personalized, empathic, and empowering for patients. Further, these healthcare documents are typically generated and reviewed by a healthcare practitioner in less than 5 minutes, whereas healthcare practitioners would have taken 15-30 minutes to write such documents (e.g., an AVS) previously. The embodiments herein have proven to be particularly well suited for suggesting improvements from a patient’s baseline and providing justification for a recommendation by combining multiple risk factors. Thus, beter results are produced in a more efficient fashion.

[0085] Figure 6 provides examples of synthetic patient notes (generated for purposes of testing the embodiments herein), synthetic patient AVSs (generated by the embodiments herein), and an automated evaluation of the AVSs (performed by the automated evaluation framework). In some cases, edits were made or errors were introduced to the AVSs in order to test the automated evaluation framework. The results show that the automated evaluation framework is able to produce accurate assessments of the AVSs.

[0086] Figure 7A depicts example content 700 for an LLM prompt in accordance with these embodiments. Content 700 is defined programmatically, with a “_type:” section defining the type of data (a prompt) and its input (documents and patient information). The prompt content is provided in markup language (e.g., XML, though other structured formats such as JSON could be used). The {documents} token would be replaced with data from a patient’s healthcare records, and the {patient} token would be replaced with data from the patient’s profile (e.g., demographic information). Any information from patient health records 306A, patient genetic data 306B, and / or other data 306C may be used to populate such data.

[0087] Figure 7B depicts example content 750 for another LLM prompt in accordance with these embodiments. Content 750 instructs an LLM to generate a JSON schema from patient data.

[0088] Figures 8A and 8B provide example LLM-generated AVSs, AVS 800 and AVS 850, respectively. These examples employ specific AVS formats, and other AVSs may be formatted differently and / or contain different content. AVSs 800 and 850 demonstrate that LLMs can be trained and / or prompted to provide more or less detail and / or with different conversational tones.IV. Automatic Evaluation Framework

[0089] In some embodiments, the automated evaluation framework includes four categories of metrics to capture common failures in generation of AVSes. These are: hallucination, personalization, similarity, and bias. NLP metrics as well as LLM-as-a-judge metrics can be used as or with the evaluator LLM. Testing has shown that these metrics were able to accurately identify hallucinations, incorrect personalizations, and biases. Further, explanations of output differences generated by an LLM acting as a judge were able to identify differences between AVSs that otherwise were deemed similar.A. Hallucination

[0090] The hallucination metrics may use LLM-as-a-judge methods to determine whether there is incorrect content in the AVS. These are divided into open domain and closed domain methodologies. Hallucinations are defined here as information that is nonsensical or inaccurate. Closed domain hallucinations are incorrect with respect to the given documents or context, whereas open domain hallucinations are generally incorrect regardless of context. Given the stochastic nature of LLM-as-a-judge metrics, having multiple metrics mitigates excessive variability in the overall scores. Five distinct hallucination metrics can be defined as follows.

[0091] Two metrics, one for closed domain and one for open domain, were derived using a modified version of Chainpoll (an LLM-based evaluation tool that can detect hallucinations in model output by combining chain-of-thought reasoning and multiple polls of the LLM-based evaluator). First, the AVS is split into sentences using an LLM, and then a fewshot chain-of-thought prompting strategy is used to fetch multiple responses (e.g., n=5 samples). Each time, all the sentences are provided as a numbered list and the LLM returns a comma separated list of sentence numbers it deems incorrect. This provides per-sentence scores of either 0 or 1 where 0 indicates a hallucination, and 1 indicates that the sentence is correct.The per-sentence scores are then averaged over all samples. With closed domain evaluation, the prompt uses the reference text for comparison.

[0092] A metric similar to G-eval (not unlike Chainpoll but includes a form-filling technique to provide specific assessments of the AVS) is used for both closed and open domains, with and without the reference text respectively. This metric queries the LLM once and requests a rating from 1-5 for the entire AVS along with a qualitative rationale using chain-of-thought prompting. The rating is then calculated as a weighted sum of token probabilities. The rationale provides additional insight behind the LLM’s reasoning.

[0093] A modified, efficient version of SelfCheckGPT (tests an LLM for hallucinations by issuing it the same prompt multiple times and comparing answers) is used for closed domain evaluation. Multiple AVS samples (e.g., n=3) are generated using the exact same inputs and parameters. The nondeterministic nature of LLMs is such that each sample and the original are slightly different. The original AVS is then parsed into numbered sentences using an LLM. Then, for each additional sample provided, the numbered list of sentences is provided along with the additional sample and context to the LLM in another request using a few-shot chain- of-thought prompting strategy. The LLM returns a comma separated list of incorrect sentence numbers, and the same scoring scheme is applied.B. Personalization

[0094] The point of care documents used for generating AVSs contain generic guidelines and are not personalized to the patient’s background. It is desirable for the AVSs to be personalized to the patient; for example, it is not helpful to suggest a patient stop eating red meat if they are already a vegetarian. The personalization metric seeks to find congruence between the assumptions embedded in the AVS and the information in the context and is implemented as follows.

[0095] A single-shot chain-of-thought prompt can used to request that an LLM break down the AVS into a numbered list of assumptions made in the content with a focus on the patient’s health and lifestyle. Next, another single-shot chain-of-thought prompt can be used to assess congruence between the assumptions and the context (patient information and point of care documents), but the AVS itself is not included. The LLM then returns a numbered list of scores for each assumption, where 0 indicates that the assumption is not congruent, and 1 indicates otherwise. A rationale for each ‘0’ scored assumption can also be requested.

[0096] The numeric scores may then be averaged to produce an overall personalization score for the entire AVS. The rationale is used for manual validation and additional insights into the scores provided.C. Similarity

[0097] Since clinicians are able to edit the LLM-generated AVS before sending it to patients, this provides another opportunity to evaluate the LLM-generated AVSs. The LLM- generated AVSs can be compared with their post-edited versions to quantify other incidental issues that may arise such as tone, wording, redundancy, etc., based on a clinician’s individual preferences. A variety of classical similarity metrics may be used, as well as an LLM-as-a- judge metric to include a rationale. If a detrimental system change were to be made in the future causing a significant amount of edits in the AVSs, this metric would provide a concise summary of what sort of edits are being made and therefore expedite troubleshooting significantly. The similarity metrics consist of: word count difference, cosine similarity using text embeddings, Jaccard similarity, TF-IDF similarity, Rouge3 precision, recall, f-score, RougeL precision, recall, f-score, and, an LLM-as-a-judge similarity score. In evaluating similarity, one or more of these metrics may be calculated and (except for the word count difference) averaged together to obtain a score between 0.0 and 1.0, where 1.0 is identical.D. Bias

[0098] LLMs can encode societal biases. Thus, it is desirable to include some assessment of health equity in the automated evaluation framework. To accomplish this, synthetic patient profiles can be altered and an LLM-as-a-judge metric can be used to evaluate bias in the resulting AVSs. Three bias axes (age, ethnicity, and gender) were identified for which the expected recommendations should be the same for both the original and altered profile. The AVSs for the original patients were manually reviewed, and they contained little to no bias. The evaluation may proceed as follows: the LLM is prompted with a task to identify and score two AVSs, where one is unbiased (original profile) and one may or may not be biased (altered profile). It is also provided a brief overview of some common medical biases. The LLM provides a score and a rationale which was used in manual validation.V. Usage of the Automated Evaluation Framework

[0099] Genetic results presented to patients have associated point of care documents that serve as references for respective genetically-related conditions. Giving the LLM access to the specialized knowledge in point of care documents results in higher-quality AVSs as illustrated below.

[0100] Providing clear and specific instructions in LLM prompts helps with interpretation and improves results. The AVS generation prompts can be lengthy and include task instructions, test result specific point of care documents and patient information. Delimiters are used to demarcate the various sections of text that need to be processed differently and to schematize patient information with XML tags and JSON formatting. Schematization aids with information extraction and task understanding, leading to more personalized responses. The following examples based on synthetic patient outputs illustrate the difference in response quality with and without schematization.VI. Example Patient Profile and Recommendations

[0101] A patient profile can be based on data from patient health records 306A, patient genetic data 306B, and / or other data 306C. An example is shown below.• Age: Required• Sex: Required• Ethnicity: Required• Reason for the visit: Optional• Genetic results: Optional o Gene: APOE, Variant: c.388T>C (e4) (1 copy)- Alzheimer’s disease• Pharmacogenetics results: Optional o CYP2C19 - Predicted rapid metabolizer o SLCO IB 1 - Predicted normal function o DPYP - Predicted normal metabolizer• Polygenic risk score results: Optional o Age-Related Macular Degeneration- variant detected, not likely at increased risk o ADHD- increased likelihood o Gout- increased likelihood o Insomnia- increased likelihood o Skin cancer- increased likelihood o Triglycerides- increased likelihood• Past medical history: Optional o Borderline high cholesterol (diagnosed age 32, last labs 2 years ago were normal per pt) o Borderline high blood pressure (diagnosed age 32, last BP 1 year ago and in normal range per pt) o Depression / Anxiety (diagnosed age 35) o Hair loss o ED o Deviated septum.• Family History: Required o Mother (age 73)- High blood pressure (diagnosed in 50s), High cholesterol (diagnosed in 50s), Macular degeneration (diagnosed in 60s), Panic attacks (diagnosed in 50s), possible ADHD, cataracts, hip replacements. o Father (age 77)- High blood pressure (diagnosed in 50s), High cholesterol (diagnosed in 50s), Melanoma (diagnosed in 70s), Gout (diagnosed in 60s), Dementia / Parkinsonism with Progressive Supranuclear Palsy (diagnosed age 75), Hearing / Vision loss in 60s. o Brother (age 35)- Type 1 diabetes (diagnosed age 19). o Son- age 6. GERD, Trouble with sleep. o Daughter- age 9. Healthy. o Maternal uncle (deceased in 50s from colon cancer)- Colon cancer (diagnosed in 50s), CAD (diagnosed in 50s). o Maternal uncle (age late 60s)- CAD (diagnosed late 50s / early 60s). o Maternal uncle (age 70s)- Stroke (diagnosed in 70s)o Maternal uncle (age 65)- Healthy. o Maternal aunt (age 75)- High blood pressure, Obesity. o Maternal aunt (deceased in 40s from suicide)- Bipolar, Schizophrenia. o Maternal grandfather (deceased in late 50s / early 60s from car accident)- CAD (diagnosed in late 30s / early 40s), Alcoholism. o Maternal grandmother (deceased in 80s)- Alzheimer’s dementia (diagnosed in 60s). o Paternal uncle (age early 70s)- CAD (diagnosed in 60s), High blood pressure (diagnosed in 50s), Stroke (diagnosed in 50s), Diverticulitis (50s), Alcoholism. o Paternal aunt (age 80)- Obesity, High cholesterol, Vision loss. o Paternal grandfather (deceased in 70s)- CAD (diagnosed in 70s). o Paternal grandmother (deceased in 90s)- Small stroke (diagnosed in 90s), Macular degeneration.• Lifestyle: Required o Smoking: no o Alcohol: 5 drinks / week o Exercise: 4 days a week, vigorous exercise for 45 min o Diet: pretty good, fats and sweets 1-3 days a week, daily vegetables and fruit• Medications: Optional o Zoloft, Propecia, Regain, Flonase, Viagra.• Laboratory findings: Required o No labs available.• Clinical Assessment: Optional o BP: N / A o HR: N / A o Height: 5’ 11” o Weight: 175 pounds• Assessment and Plan: Required• Confirmatory testing recommended — No• Although there is a family history of CAD (and presumed hyperlipidemia), the majority family members with CAD did not have premature CAD. That, inaddition to the patient having borderline high cholesterol that improved with dietary changes, makes familial hyperlipidemia less likely.

[0102] An example of healthcare provider notes relating to this patient is shown below.• Recommendations:• For Late-Onset Alzheimer’s Disease Risk Reduction: o Annual physical with labs yearly to include CBC, CMP, Lipid panel, HgbAlC, Lp(a), ApoB, Vitamin D. Monitor for blood pressure control. o Optimize existing metabolic risk factors (HTN, DM, cholesterol control) o Vitamin D replacement o Diet: Evidence for value of mediterranean diet specifically, omega-3.Needs to decrease sweets and decrease fat in diet. o Exercise: men 50 / 50 cardio and resistance training o Avoid smoking o Regular sleep o Regular participation in mentally challenging leisure activities such as reading, playing board games, doing puzzles, or playing an instrument• Discussed genetic reports with increased likelihood including Age-Related Macular Degeneration, ADHD, Gout, Insomnia, Skin caner, Triglycerides. Discussed recommendations for each.• Discussed CYP2C19 Drug Metabolism- Predicted rapid metabolizer and how this can affect medication he is prescribed, particular his anxiety medications.• Discussed Age-Related Macular Degeneration- Variant detected, not likely at increased risk, but would continue to monitor and recommend annual eye exams given his family history of ARMD.• Consider evaluation for sleep apnea

[0103] This information may be provided in a prompt to an LLM, such as LLM 310, with instructions to generate an AVS.VII. Example Technical Improvements

[0104] As noted above, this invention is a technical solution to a technical problem. Generation of accurate AVSs can lead to reduction of computing resource (processor, memory, network, and / or power) usage in a web-based or app-based system, as patients are less likely to refresh web / app pages and navigate extensively seeking information to bridge the gap between a physician’s terse notes and actionable healthcare recommendations.

[0105] Moreover, by introducing an integrated “second-look” LLM-as-a-judge that vets the LLM-generated AVSs, the system can intelligently gate which records actually need full reprocessing. For example, if the initial AVS-generating LLM produces an AVS whose metadata (e.g., confidence scores, token counts, accuracy metrics) all fall within pre-defined safe bounds, the vetting LLM-as-a-judge may be able to skip in-depth bias and hallucination checks altogether. This conditional invocation avoids wasteful full-model executions on straightforward cases, cutting down processor and GPU cycles. Further, by caching intermediate embeddings and prompt templates for patients with similar visit profiles, the system can reuse vector representations instead of regenerating them, reducing memory allocations and network calls to external embedding services. Together, these factors shrink average processor and GPU utilization per AVS by avoiding redundant LLM passes and redundant data transfers between subsystems.

[0106] At the hardware and orchestration layer, the heavier bias and hallucination analyses can be scheduled to execute in bulk during off-peak hours, leveraging batch processing to amortize I / O and model-load overheads across many AVSs. Additionally, using shared memory buffers between the AVS-generating LLM and the LLM-as-a-judge allows them to exchange embeddings and tokenized text without serialization through disk or RPC layers. Doing so saves storage I / O and power.

[0107] Other technical improvements may be possible. Thus, this statement of technical improvements is not intended to be comprehensive or limiting.VIII. Example Operations

[0108] Figure 9 is a flow chart 900 illustrating an example embodiment. The process illustrated by Figure 9 may be carried out by a computing device, such as computing device 100, and / or a cluster of computing devices, such as server cluster 200. However, the process can be carried out by other types of devices or device subsystems. For example, the process could be carried out by a computational instance of a remote network management platform or a portable computer, such as a laptop or a tablet device.

[0109] The embodiments of Figure 9 may be simplified by the removal of any one or more of the features shown therein. Further, these embodiments may be combined with features, aspects, and / or implementations of any of the previous figures or otherwise described herein.

[0110] Block 902 may involve receiving, by way of a user interface of a computing system, an indication of a healthcare patient profile and further input.[oni] Block 904 may involve, based on the indication of the healthcare patient profde, obtaining, by the computing system, health records and genetic data of a patient associated with the healthcare patient profile.

[0112] Block 906 may involve, based on the health records and the genetic data of the patient and the further input, generating, by the computing system, a natural language prompt, wherein the natural language prompt contains instructions to generate a healthcare document for the patient.

[0113] Block 908 may involve providing, by the computing system, the natural language prompt to a generative natural language model.

[0114] Block 910 may involve receiving, by the computing system, the healthcare document from the generative natural language model.

[0115] Block 912 may involve providing, by way of the user interface, a version of the healthcare document.

[0116] Some embodiments may further involve, before providing the version of the healthcare document: (i) providing the healthcare document as produced by the generative natural language model to a further natural language model with instructions to evaluate the healthcare document and produce suggested edits to the version of the healthcare document, and (ii) receiving the suggested edits to the version of the healthcare document from the further natural language model, wherein the version of the healthcare document is based on the suggested edits to the version of the healthcare document.

[0117] In some embodiments, the instructions specify pre-determined metrics comprising one or more of: accuracy, bias, and personalization metrics and thresholds thereof, and wherein the version of the healthcare document was edited by the further natural language model to determine or meet the thresholds of the accuracy, bias, and personalization metrics.

[0118] Some embodiments may further involve, before providing the version of the healthcare document: (i) providing the healthcare document as produced by the generative natural language model to a further natural language model with instructions to evaluate the healthcare document and produce a scored version of the healthcare document, and (ii) receiving the scored version of the healthcare document from the further natural language model, wherein the version of the healthcare document is based on the scored version.

[0119] In some embodiments, the instructions specify pre-determined metrics including one or more of: accuracy, bias, and personalization metrics and thresholds thereof, and wherein the scored version of the healthcare document was scored by the further naturallanguage model to determine whether it met the thresholds of the accuracy, bias, and personalization metrics.

[0120] In some embodiments, the instructions to evaluate the healthcare document comprise separating the healthcare document into a plurality of sections, scoring each of the plurality of sections of the healthcare document, and statistically combining the scores for each of the plurality of sections of the healthcare document to determine an overall score, and determining whether the overall score meets the thresholds.

[0121] In some embodiments, the accuracy metric includes determining presence or absence of one or more hallucinations.

[0122] In some embodiments, the accuracy metric includes open domain hallucinations, closed domain hallucinations, and a statistical combination thereof.

[0123] Some embodiments may further involve receiving, by way of the user interface, edits to the version of the healthcare document, wherein the version of the healthcare document is provided to the patient as a final version thereof.

[0124] In some embodiments, the indication of the healthcare patient profile is received from a healthcare provider, and wherein the computing system authenticates the healthcare provider for access to the healthcare patient profile.

[0125] In some embodiments, the healthcare document comprises a draft of an after visit summary for clinician editing and review.

[0126] In some embodiments, the generative natural language model is a transformerbased large language model.

[0127] Some embodiments may further involve, based on the indication of the healthcare patient profile, obtaining, by the computing system, one or more point of care documents with expert health related information relevant to the healthcare patient profile; and including the point of care documents in the natural language prompt.

[0128] In some embodiments, the healthcare document is a draft personalized after visit health summary for clinician editing and review.

[0129] Some embodiments may further involve, comparing the version of the healthcare document to an edited version of the healthcare document that was edited by a clinician to generate a similarity metric.

[0130] Some embodiments may further involve, statistically analyzing similarity metrics for a plurality of healthcare documents to determine an aggregate similarity metric. IX. Closing

[0131] The present disclosure is not to be limited in terms of the particularembodiments described in this application, which are intended as illustrations of various aspects. Many modifications and variations can be made without departing from its scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those described herein, will be apparent to those skilled in the art from the foregoing descriptions. Such modifications and variations are intended to fall within the scope of the appended claims.

[0132] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying figures. The example embodiments described herein and in the figures are not meant to be limiting. Other embodiments can be utilized, and other changes can be made, without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0133] With respect to any or all of the message flow diagrams, scenarios, and flow charts in the figures and as discussed herein, each step, block, and / or communication can represent a processing of information and / or a transmission of information in accordance with example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, operations described as steps, blocks, transmissions, communications, requests, responses, and / or messages can be executed out of order from that shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved. Further, more or fewer blocks and / or operations can be used with any of the message flow diagrams, scenarios, and flow charts discussed herein, and these message flow diagrams, scenarios, and flow charts can be combined with one another, in part or in whole.

[0134] A step or block that represents a processing of information can correspond to circuitry that can be configured to perform the specific logical functions of a herein-described method or technique. Alternatively or additionally, a step or block that represents a processing of information can correspond to a module, a segment, or a portion of program code (including related data). The program code can include one or more instructions executable by a processor for implementing specific logical operations or actions in the method or technique. The program code and / or related data can be stored on any type of computer readable medium such as a storage device including RAM, a disk drive, a solid-state drive, or another storage medium.

[0135] The computer readable medium can also include non-transitory computer readable media such as non-transitory computer readable media that store data for short periodsof time like register memory and processor cache. The non-transitory computer readable media can further include non-transitory computer readable media that store program code and / or data for longer periods of time. Thus, the non-transitory computer readable media may include secondary or persistent long-term storage, like ROM, optical or magnetic disks, solid-state drives, or compact disc read only memory (CD-ROM), for example. The non-transitory computer readable media can also be any other volatile or non-volatile storage systems. A non- transitory computer readable medium can be considered a computer readable storage medium, for example, or a tangible storage device.

[0136] Moreover, a step or block that represents one or more information transmissions can correspond to information transmissions between software and / or hardware modules in the same physical device. However, other information transmissions can be between software modules and / or hardware modules in different physical devices.

[0137] The particular arrangements shown in the figures should not be viewed as limiting. It should be understood that other embodiments could include more or less of each element shown in a given figure. Further, some of the illustrated elements can be combined or omitted. Yet further, an example embodiment can include elements that are not illustrated in the figures.

[0138] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purpose of illustration and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method comprising: receiving, by way of a user interface of a computing system, an indication of a healthcare patient profde and further input; based on the indication of the healthcare patient profde, obtaining, by the computing system, health records and genetic data of a patient associated with the healthcare patient profde; based on the health records and the genetic data of the patient and the further input, generating, by the computing system, a natural language prompt, wherein the natural language prompt contains instructions to generate a healthcare document for the patient; providing, by the computing system, the natural language prompt to a generative natural language model; receiving, by the computing system, the healthcare document from the generative natural language model; and providing, by way of the user interface, a version of the healthcare document.

2. The computer-implemented method of claim 1, further comprising: before providing the version of the healthcare document: (i) providing the healthcare document as produced by the generative natural language model to a further natural language model with instructions to evaluate the healthcare document and produce suggested edits to the version of the healthcare document, and (ii) receiving the suggested edits to the version of the healthcare document from the further natural language model, wherein the version of the healthcare document is based on the suggested edits to the version of the healthcare document.

3. The computer-implemented method of claim 2, wherein the instructions specify pre-determined metrics comprising one or more of: accuracy, bias, and personalization metrics and thresholds thereof, and wherein the version of the healthcare document was edited by the further natural language model to determine or meet the thresholds of the accuracy, bias, and personalization metrics.

4. The computer-implemented method of claim 1, further comprising: before providing the version of the healthcare document: (i) providing the healthcaredocument as produced by the generative natural language model to a further natural language model with instructions to evaluate the healthcare document and produce a scored version of the healthcare document, and (ii) receiving the scored version of the healthcare document from the further natural language model, wherein the version of the healthcare document is based on the scored version.

5. The computer-implemented method of claim 4, wherein the instructions specify pre-determined metrics including one or more of: accuracy, bias, and personalization metrics and thresholds thereof, and wherein the scored version of the healthcare document was scored by the further natural language model to determine whether it met the thresholds of the accuracy, bias, and personalization metrics.

6. The computer-implemented method of claim 5, wherein the instructions to evaluate the healthcare document comprise separating the healthcare document into a plurality of sections, scoring each of the plurality of sections of the healthcare document, and statistically combining the scores for each of the plurality of sections of the healthcare document to determine an overall score, and determining whether the overall score meets the thresholds.

7. The computer-implemented method of claim 5, wherein the accuracy metric includes determining presence or absence of one or more hallucinations.

8. The computer-implemented method of claim 7, wherein the accuracy metric includes open domain hallucinations, closed domain hallucinations, and a statistical combination thereof.

9. The computer-implemented method of claim 1, further comprising: receiving, by way of the user interface, edits to the version of the healthcare document, wherein the version of the healthcare document is provided to the patient as a final version thereof.

10. The computer-implemented method of claim 1, wherein the indication of the healthcare patient profile is received from a healthcare provider, and wherein the computing system authenticates the healthcare provider for access to the healthcare patient profile.

11. The computer-implemented method of claim 1, wherein the healthcare document comprises a draft of an after visit summary for clinician editing and review.

12. The computer-implemented method of claim 1, wherein the generative natural language model is a transformer-based large language model.

13. The computer-implemented method of claim 1, further comprising: based on the indication of the healthcare patient profile, obtaining, by the computing system, one or more point of care documents with expert health related information relevant to the healthcare patient profile; and including the point of care documents in the natural language prompt.

14. The computer-implemented method of claim 1, wherein the healthcare document is a draft personalized after visit health summary for clinician editing and review.

15. The computer-implemented method of claim 14, further comprising: comparing the version of the healthcare document to an edited version of the healthcare document that was edited by a clinician to generate a similarity metric.

16. The computer-implemented method of claim 15, further comprising: statistically analyzing similarity metrics for a plurality of healthcare documents to determine an aggregate similarity metric.

17. A non-transitory computer-readable medium storing program instructions that, when executed by one or more processors of a computing system, cause the computing system to perform the operations of any of claims 1-16.

18. A computing system comprising: one or more processors; memory; and program instructions, stored in the memory, that upon execution by the one or more processors cause the computing system to perform the operations of any of claims 1-16.

Citation Information

Patent Citations

  • Systems and methods of clinical trial evaluation

    US20200381087A1

  • Systems and methods for machine-learning-assisted cognitive evaluation and treatment

    US20230255564A1

  • Informatics platform for integrated clinical care

    WO2017042396A1

  • Systems for predictive data analytics, and related methods and apparatus

    WO2018075995A1

  • KR20220135387A