Privacy-protecting large-language model output
The method trains and fine-tunes LLMs to identify and deidentify Pll in clinical analyses, addressing privacy concerns and maintaining functionality, thus ensuring compliance with privacy laws.
Patent Information
- Application Number
- PCT/US2024/062279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-30
- Publication Date
- 2025-07-03
AI Technical Summary
Existing large-language models (LLMs) used in generating clinical notes and recommendations risk exposing personally identifiable information (Pll) during training, refinement, or usage, violating privacy laws like HIPAA.
A method involving training an LLM on clinical records, fine-tuning it with labeled Pll data, and using it to identify and deidentify Pll instances within clinical analyses, followed by secondary deidentification processes to ensure privacy.
The method effectively protects patient privacy by ensuring Pll is not traceable in clinical analyses, while maintaining the LLM's functionality for authorized clinicians, and is more efficient than traditional dictionary-based methods.
Smart Images

Figure US2024062279_03072025_PF_FP_ABST
Abstract
Description
PRIVACY-PROTECTING LARGE-LANGUAGE MODEL OUTPUTCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to and the filing date of provisional U.S. Patent Application No. 63 / 616,160 entitled “PRIVACY-PROTECTING LARGE-LANGUAGE MODEL OUTPUT,” filed on December 29, 2023, the entire contents of which is hereby expressly incorporated herein by reference.TECHNICAL FIELD
[0002] The present disclosure is generally directed to systems and methods for deidentifying personally identifiable information (PH), and more particularly to deidentifying protected health information (PHI) from clinical analyses.BACKGROUND
[0003] Privacy laws and regulations, such as the Health Insurance Portability and Accountability Act (HIPAA), restrict the use of personally identifiable information (Pll), such as protected health information (PHI), by covered entities. Large-language models (LLMs) can be used to generate clinical notes, differential diagnoses, and other valuable summaries, as well as to provide clinical recommendations in the medical field. However, there are concerns that during the training, refinement, or usage of these LLMs, identifiable clinical information could be leaked, resulting in the public exposure of Pll for individuals whose data was used in the training process.BRIEF SUMMARY
[0004] In an embodiment, a computer-implemented method for deidentifying personally identifiable information (Pll) is provided. The method includes (1 ) training an LLM on a set of clinical records to generate a clinical analysis, wherein the first set of clinical records comprises unstructured text including Pll; (2) fine tuning the LLM on a set of Pll training data to identify Pll, wherein the set of Pll training data comprises Pll examples labeled with one or more Pll categories; (3) receiving from a client device a clinical analysis request and a patient identifier; (4) providing to the LLM the clinical analysis request and one or more patient records associated with the patient identifier, causing the LLM to: (a) generate a clinical analysis of the one or more patient records and (b) identify one or more Pll instances in the clinical analysis, (5) receiving from the LLM the via the one or more processors, the clinical analysis and the one or more Pll instances; (6) deidentifying, by the one or more processors, at least one of the one or more Pll instances in the clinical analysis to generate a deidentified clinical analysis; and / or (7) transmitting the deidentified clinical analysis to the client device.
[0005] In another embodiment, a computer-implemented method for deidentifying PH is provided. The method includes (1 ) training an LLM on a set of clinical records to generate clinical analyses wherein the set of clinical records comprises unstructured text including PH; (2) fine tuning the LLM on a set of PH training data to identify PH, wherein the set of PH training data comprises Pll examples labeled with one or more PH categories; (3) receiving from a client device a clinical analysis request and a patient identifier; (4) providing to the LLM the clinical analysis request and one or more patient records associated with the patient identifier, causing the LLM to: (a) generate a clinical analysis of the one or more patient records, (b) identify one or more Pll instances in the clinical analysis, and (c) compile the one or more PH instances into a Pll list; (5) receiving from the LLM via the one or more processors the clinical analysis and the Pll list; (6) deidentifying, using the PH list, the one or more PH instances from the clinical analysis to generate a deidentified clinical analysis; and / or (7) transmitting the deidentified clinical analysis to the client device.
[0006] In another embodiment, a computing system for deidentifying PH is provided. The computing system includes: (i) one or more processors, and (ii) one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: (1 ) train an LLM on a set of clinical records to generate a clinical analysis, wherein the set of clinical records comprises unstructured text including PH; (2) fine tune the LLM on a set of PH training data to identify PH, wherein the set of Pll training data comprises Pll examples labeled with one or more Pll categories; (3) receive from a client device a clinical analysis request and a patient identifier; (4) provide to the LLM the clinical analysis request and one or more patient records associated with the patient identifier, causing the LLM to: (a) generate a clinical analysis of the one or more patient records, and (b) identify one or more PH instances in the clinical analysis, (5) receive from the LLM the clinical analysis and the one or more PH instances; (6) deidentify at least one of the one or more PH instances in the clinical analysis to generate a deidentified clinical analysis; and / or (7) transmit the deidentified clinical analysis to the client device.
[0007] In another embodiment, a computing system for deidentifying PH is provided. The computing system includes: (i) one or more processors, and (ii) one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to: (1 ) train an LLM on a set of clinical records to generate clinical analyses wherein the first set of clinical records comprises unstructured text including PH; (2) fine tune the LLM on a second set of Pll training data to identify Pll, wherein the second set of Pll training data comprises Pll examples labeled with one or more Pllcategories; (3) receive from a client device a clinical analysis request and a patient identifier; (4) provide to the LLM the clinical analysis request and one or more patient records associated with the patient identifier, causing the LLM to: (a) generate a clinical analysis of the one or more patient records, (b) identify one or more PH instances in the clinical analysis, and (c) compile the one or more Pll instances into a Pll list; (5) receive from the LLM the clinical analysis and the Pll list; (6) deidentify, using the Pll list, the one or more Pll instances from the clinical analysis to generate a deidentified clinical analysis; and / or (7) transmit the deidentified clinical analysis to the client device.BRIEF DESCRIPTION OF THE FIGURES
[0008] The figures described below depict various aspects of the system and methods disclosed therein. It should be understood that each figure depicts one aspect of a particular aspect of the disclosed system and methods, and that each of the figures is intended to accord with a possible aspect thereof. Further, wherever possible, the following description refers to the reference numerals included in the following figures, in which features depicted in multiple figures are designated with consistent reference numerals.
[0009] FIG. 1 depicts an example computing environment in which the techniques disclosed herein may be implemented, according to some aspects, according to some aspects.
[0010] FIG. 2 depicts a combined block and logic diagram in which example computer- implemented methods and systems for training an LLM are implemented, according to one embodiment.
[0011] FIG. 3 depicts a combined block and logic diagram in which example computer- implemented methods and systems for fine tuning and testing an LLM are implemented, according to another embodiment.
[0012] FIG. 4A depicts a flow diagram of an example computer-implemented method for deidentifying Pll using an LLM, according to one embodiment.
[0013] FIG. 4B depicts a flow diagram of an example computer-implemented method for deidentifying Pll using an LLM, according to another embodiment.
[0014] FIG. 5A depicts an example clinical analysis including Pll, according to some aspects.
[0015] FIG. 5B depicts the example clinical analysis with Pll redacted, according to some aspects.
[0016] The figures depict preferred aspects for purposes of illustration only. One skilled in the art will readily recognize from the following discussion that alternative aspects of the systemsand methods illustrated herein may be employed without departing from the principles of the invention described herein.DETAILED DESCRIPTIONOverview
[0017] The present techniques provide methods and systems for, inter alia, deidentifying PH from LLM output. Examples of PH include patient name, birth date, social security number or other government identifier, home address, e-mail address, telephone number, etc. A computing system may receive from a client device a patient identifier and a request for a clinical analysis. The computing system may transmit patient records associated with the patient identifier and the clinical analysis request to a trained LLM model. The trained LLM model may generate a clinical analysis from the patient records. In various examples, the trained LLM model may identify one or more instances of Pll in the clinical analysis, deidentify the Pll instances, and return a deidentified clinical analysis (containing the deidentified PH instances) to the computing system. In various other examples, the trained LLM model may identify one or more instances of Pll in the clinical analysis, generate a list of the Pll instances, and return the clinical analysis and the list of the Pll instances. From that output of the LLM, a subsequent post-processing may be done at the computing system to perform deidentification on the clinical analysis and PH instances, thereby generating a deidentified clinical analysis. Such post-processing deidentification may be performed by a secondary, independent, privacypreserving trained LLM model, for example, that could be applied to any computing system for preserving privacy. That is, the post-processing deidentification may be performed on the same computing system having the trained LLM model or an external computing system communicatively coupled thereto. The computing system may transmit the deidentified clinical analysis to the client device. Examples of clinical analyses may include clinical notes, differential diagnoses, clinical recommendations, clinical summaries, etc.
[0018] Thus, in various examples, computing systems are provided in which trained LLMs generate deidentified clinical analyses where at least some Pll data has been excluded from the clinical analyses and / or modified by the LLM to ensure non-traceability. And computing systems are provided in which trained LLMs collate Pll data into a dedicated list, for example, to generate a collated clinical analyses that is subsequently deidentified, by replacing or redacting at least some Pll data with generic placeholders or modified values or by encrypting access to at least some of the PH data, through a process after the LLM, even a process performed by a subsequent LLM.
[0019] In various examples, the proposed frameworks herein are able to harness the capabilities of LLMs in the medical field and others while protecting the privacy and confidentiality of patients and patient data through deidentification of some or all PH data. The proposed frameworks provide technical benefits, among them the ability to allow authorized clinicians to have access to full set PH data, while protecting Pll data from traceability, as that Pll data is collected into clinical analyses data that is to be accessed to others outside of the authorized clinicians, thereby ultimately preventing disclosure of private patient information.
[0020] Advantageously, the disclosed methods and systems provide improvements in computer functionality. Relative to existing techniques that use dictionary-based or rule-based named entity recognition (NER) to detect Pll instances, the present techniques use a finetuned LLM, which is faster and more computationally efficient. The advantages of using a finetuned LLM to detect PH instances are the ability to detect PH instances not previously encountered and lower memory usage due to not storing a dictionary in memory. The advantages of a finetuned LLM over rule-based NER are the ability to detect Pll instances that do not match a predefined pattern and lower processor usage due to not performing pattern matching on each token or word.
[0021] As is clear from the above and the description that follows, the present disclosure provides specific features that are not well-understood, routine, or conventional activity in the field (e.g., using a finetuned LLM to detect Pll). Other advantages will also be apparent to those of ordinary skill in the relevant art upon reviewing this disclosure.
[0022] While the various examples herein are described in different examples, it will be appreciated that the proposed processes and methods herein (including those above) can be sequentially implemented to serve as a dual-layered safeguard. For example, initially, an LLM may produce deidentified data. Following that initial layer safeguard, in a subsequent layer safeguard, a PH detection mechanism may be activated to identify and rectify any inadvertently generated identifiable information from the initial layer, ensuring comprehensive protection against unintentional data exposure.Example Computing Environment
[0023] FIG. 1 depicts an example computing environment 100 in which the techniques disclosed herein may be implemented, according to some aspects. The environment 100 may include computing resources for training and / or operating machine learning models to manage discharge of a patient.
[0024] The computing environment 100 may include a client computing device 102, a server computing device 104, an electronic network 106, an electronic health record (EHR) database 108, and a model database 112. The computing environment may further include one or more cloud application programming interfaces (APIs) 1 14. The components of the computing environment 100 may be communicatively connected to one another via the electronic network 106, in some aspects.
[0025] The client computing device 102 may implement, inter alia, operation of one or more applications for transmitting clinical analysis requests and receiving clinical analyses. In some aspects, the client computing device 102 may be implemented as one or more computing devices (e.g., one or more servers, one or more laptops, one or more mobile computing devices, one or more tablets, one or more wearable devices, one or more cloud-computing virtual instances, etc.). In some aspects, a plurality of client computing devices may be part of the environment 100 - for example, a first user may access a client computing device 102 that is a laptop, while a second user accesses the client computing device 102 that is a smart phone, while yet a third user accesses a client computing device 102 that is a wearable device.
[0026] The client computing device 102 may include one or more processors 120, one or more network interface controllers 122, one or more memories 124, an input device 126, an output device 128 and a client API 130. The one or more memories 124 may have stored thereon one or more modules 140 (e.g., one or more sets of instructions).
[0027] In some aspects, the one or more processors 120 may include one or more central processing units, one or more graphics processing units, one or more field-programmable gate arrays, one or more application-specific integrated circuits, one or more tensor processing units, one or more digital signal processors, one or more neural processing units, one or more RISC-V processors, one or more coprocessors, one or more specialized processors / accelerators for artificial intelligence (Al) or machine learning (ML) -specific applications, one or more microcontrollers, etc.
[0028] The client computing device 102 may include one or more network interface controllers 122, such as Ethernet network interface controllers, wireless network interface controllers, etc. In some aspects, the network interface controllers 122 may include advanced features, such as hardware acceleration, specialized networking protocols, etc.
[0029] The memories 124 of the client computing device 102 may include volatile and / or nonvolatile storage media. For example, the memories 124 may include one or more random access memories, one or more read-only memories, one or more cache memories, one or more hard disk drives, one or more solid-state drives, one or more non-volatile memory express, oneor more optical drives, one or more universal serial bus flash drives, one or more external hard drives, one or more network-attached storage devices, one or more cloud storage instances, one or more tape drives, etc.
[0030] As noted, the memories 124 may have stored thereon one or more modules 140, for example, as one or more sets of computer-executable instructions. In some aspects, the modules 140 may include additional storage, such as one or more operating systems (e.g., Microsoft Windows, GNU / Linux, Mac OSX, etc.). The operating systems may be configured to run the modules 140 during operation of the client computing device 102 - for example, the modules 140 may include additional modules and / or services for receiving and processing data from one or more other components of the environment 100 such as the one or more cloud APIs 114 or the server computing device 104. The modules 140 may be implemented using any suitable computer programming language(s) (e.g., Python, JavaScript, C, C++, Rust, C#, Swift, Java, Go, LISP, Ruby, Fortran, etc.).
[0031] The modules 140 may include an API module 144, an input processing module 146, an authentication / security module 148, and a context module 150 in some aspects. In some aspects, more or fewer modules 140 may be included. The modules 140 may be configured to communicate with one another (e.g., via inter-process communication, via a bus, via sockets, pipes, message queues, etc.).
[0032] The API module 144 may include one or more sets of computer executable instructions for accessing one or more remote APIs, and / or for enabling one or more other components within the environment 100 to access functionality of the client computing device 102. In some aspects, the API module 144 may enable other client applications (i.e., not applications facilitated by the modules 140) to connect to the client computing device 102, for example, to send queries or prompts, and to receive responses from the client computing device 102. The API module 144 may include instructions for authentication, rate limiting, and error handling.
[0033] As noted, the client computing device 102 may enable one or more users to access one or more trained LLMs by providing input prompts that are processed by one or more trained LLMs. The input processing module 146 may perform pre-processing of user prompts prior to being input into one or more LLMs, and / or post-processing of outputs generated by one or more LLMs. For example, the input processing module 146 may process data input into one or more input fields, voice inputs, or other input methods (e.g., file attachments) depending upon the application. The input processing module 146 may receive inputs directly via the input device 126, in some aspects.
[0034] In some aspects, the input processing module 146 may perform post-processing of input received from one or more trained LLMs. In some aspects, post-processing (and / or preprocessing) may include implementing content moderation mechanisms, to prevent misuses of trained LLMs or inappropriate content generation. The input processing module 146 may include instructions for handling errors and for displaying errors to users (e.g., via the output device 128). The input processing module 146 may cause one or more graphical user interfaces to be displayed, for example to enable the user to enter information directly via a text field.
[0035] The authentication / security module 148 may include one or more sets of computerexecutable instructions for implementing access control mechanisms for one or more trained LLMs, ensuring that the LLM can only be accessed by those who are authorized to do so, and that the access of those users is private and secure. It should be appreciated that the security module 148 may permission users / agents based upon their respective membership in panels or other groupings.
[0036] Generally, trained models, especially trained LLMs, require state information in order to meaningfully carry on a dialogue with a user or with another trained LLM. For example, if a user prompts a trained LLM with a question such as “What is the weather in Chicago today?” followed by a second prompt “And how about tomorrow?,” the LLM should understand that, in context, the second query relates to the first query, insofar as the user is asking about the weather tomorrow in the same location (Chicago).
[0037] However, LLMs are generally stateless, meaning that after they process a prompt, they have no internal record or memory of the information that was input, or the information that was generated as part of the LLM’s processing. Thus, many systems add statefulness to models using context information. This may be implemented using sliding context windows, wherein a predetermined number of tokens (e.g., 4096 maximum tokens in the case of GPT 3.5, equivalent to about 3000 words) may be “remembered” by the LLM and can be used to enrich multiple sequential prompts input into the LLM (for example, when the LLM is used in a chat mode). GPT is a trademark of OpenAI Corp, of San Francisco, CA, and is currently available at openai.com / .
[0038] The context module 150 may include one or more sets of computer-executable instructions for maintaining state of the type found in this example, and other types of state information. The context module 150 may implement sliding window context, in some aspects. In other aspects, the context module 150 may perform other types of state maintaining strategies. For example, the context module 150 may implement a strategy in which informationfrom the immediately preceding prompt is part of the window, regardless of the size of that prior prompt.
[0039] In some aspects, the context module 150 may implement a strategy in which one or more prior prompts are included in each current prompt. This prompt stuffing technique, or prompt concatenation, may be limited by prompt size constraints — once the total size of the prompt exceeds the prompt limit, the model immediately loses state information related to parts of the prompt truncated from the prompt.
[0040] The server computing device 104 may include one or more EHR databases 108, model databases 112, processors 160, one or more network interface controllers 162, one or more memories 164, an input device (not depicted), an output device (not depicted), and / or a server API 166. The one or more memories 164 may have stored thereon one or more modules 170 (e.g., one or more sets of instructions).
[0041] The EHR database 108 may comprise health records for a plurality of patients. The EHR database 108 may include a SQL and / or NoSQL database. The health records may include clinical notes, laboratory test results, medical history, etc. The health records may comprise text, audio recordings, video recordings, and / or images. The health records may include instances of PH. The model database 112 may store one or more LLMs, including LLMs in various stages of training and / or tuning.
[0042] In some aspects, the one or more processors 160 may include one or more central processing units, one or more graphics processing units, one or more field-programmable gate arrays, one or more application-specific integrated circuits, one or more tensor processing units, one or more digital signal processors, one or more neural processing units, one or more RISC-V processors, one or more coprocessors, one or more specialized processors / accelerators for artificial intelligence or machine learning-specific applications, one or more microcontrollers, etc.
[0043] The server computing device 104 may include one or more network interface controllers 162, such as Ethernet network interface controllers, wireless network interface controllers, etc. In some aspects, the network interface controllers 162 may include advanced features, such as hardware acceleration, specialized networking protocols, etc.
[0044] The memories 164 of the server computing device 104 may include volatile and / or nonvolatile storage media. For example, the memories 164 may include one or more random access memories, one or more read-only memories, one or more cache memories, one or more hard disk drives, one or more solid-state drives, one or more non-volatile memory express, one or more optical drives, one or more universal serial bus flash drives, one or more external harddrives, one or more network-attached storage devices, one or more cloud storage instances, one or more tape drives, etc.
[0045] As noted, the memories 164 may have stored thereon one or more modules 170, for example, as one or more sets of computer-executable instructions. In some aspects, the modules 170 may include additional storage, such as one or more operating systems (e.g., Microsoft Windows, GNU / Linux, Mac OSX, etc.). The operating systems may be configured to run the modules 170 during operation of the server computing device 104 - for example, the modules 170 may include additional modules and / or services for receiving and processing data from one or more other components of the environment 100 such as the one or more cloud APIs 114 or the client computing device 102. The modules 170 may be implemented using any suitable computer programming language(s) (e.g., Python, JavaScript, C, C++, Rust, C#, Swift, Java, Go, LISP, Ruby, Fortran, etc.).
[0046] In some aspects, the modules 170 may include a data collection module 172, a data preprocessing module 174, an LLM pretraining module 176, a fine-tuning module 178, an LLM training module 180, a checkpointing module 182, a hyperparameter tuning module 184, a validation and testing module 186, an auto-prompting module 188, an operating module 190, a chatbot module 192, and a deidentification module 194. In some aspects, more or fewer modules 170 may be included. The modules 170 may be configured to communicate with one another (e.g., via inter-process communication, via a bus, via sockets, pipes, message queues, etc.). The modules 170 may respond to network requests (e.g., via the API 166) or other requests received via the network 106 (e.g., via the client computing device 102 or other components of the environment 100).
[0047] The data collection module 172 may be configured to collect information used to train one or more LLMs. The data collection module 172 may collect data via web scraping, via API calls / access, via database extract-transform-load (ETL) processes, etc. Sources accessed by the data collection module 172 include patient health records, social media websites, books, websites, academic publications, web forums / interest sites (e.g., Reddit, Facebook, bulletin boards, etc.), etc. The data collection module 172 may access data sources by active techniques (e.g., scraping or other retrieval) or may access existing corpuses. The data collection module 172 may include sets of instructions for performing data collection in parallel, in some aspects. The data collection module 172 may store collected data in one or more electronic databases, such as a database accessible via the cloud APIs 114 or via a local electronic database (not depicted). The data may be stored in a structured and / or unstructured format. In some aspects, the data collection module 172 may store large data volumes used fortraining one or more models (i.e., training data). For example, the data collection module 172 may store terabytes, petabytes, exabytes or more of training data.
[0048] In some aspects, the data collection module 172 may retrieve data from the electronic health record database 108. For example, the data collection module 172 may process the retrieved / received data and sort the data into multiple subsets based on information included within the electronic health record database 108. For example, the data collection module 172 may receive one or more sets of unstructured text (e.g., unstructured clinical notes). The data collection module 172 may chunk the data according to time (e.g., hourly, daily, quarterly, etc.).
[0049] The model preprocessing module 174 may include instructions for pre-processing data collected by the data collection module 172. In particular, the model preprocessing module 174 may perform text extraction, optical character recognition (OCR), and / or cleaning operations on data collected by the data collection module 172. The data pre-processing module 174 may perform preprocessing operations, such as lexical parsing, tokenizing, case conversions and other string splitting / munging. In some aspects, the data collection module 172 may perform data deduplication, filtering, annotation, compliance, version control, validation, quality control, etc. In some aspects, one or more human reviewers may be looped into the process of preprocessing data collected by the data pre-processing module 174. For example, a distributed work queue may be used to transmit batch jobs and receive human-computed responses from one or more human workers. Once pre-processed, the data pre-processing module 174 may store copied and / or modified copies of the training data in an electronic database.
[0050] In some aspects, the data pre-processing module 174 may include instructions for parsing the unstructured text received by the data collection module 172 to structure the text. For example, when the text relates to meeting minutes (e.g., text transcripts) of meetings of a clinical review board, the data pre-processing module 174 may generate a time series data structure in which each meeting is represented by one or more timestamps, and at each timestamp, text of one or more speakers is labeled. The data pre-processing module 174 may also label the data according to the identity of one or more speaker and / or one or more topic. For example, the time series data may be labeled according to one or more speakers associated with textual speech, in a language transcript form. The time series may include one or more keywords associated with the transcript. In some aspects, the present techniques may use a separate trained text summarization module to generate keywords used for this purpose. In this way, the data pre-processing module 174 may generate structured data corresponding to unstructured meeting minutes, such that the structured data is enriched with information aboutthe meeting that is suitable for training. This structured data may be processed by downstream processes / modules.
[0051] Generally, the present techniques may train one or more LLMs to perform language generation tasks that include token generation. Both training inputs and model outputs may be tokenized. Herein, tokenization refers to the process by which text used for training is divided into units such as words, subwords, or characters. Tokenization may break a single word into multiple subwords (e.g., “LLM” may be tokenized as “L” and “LM”). The present techniques may train one or more models using a set of tokens (e.g., a vocabulary) that includes many (e.g., thousands or more) of tokens. These tokens may be embedded into a vector. This vector of token or “embeddings” may include numerical representations of the individual tokens in the vocabulary in high-dimensional vector space. The modules 170 may access and modify the embeddings during training to learn relationships between tokens. These relationships effectively represent semantic language meaning.
[0052] In some aspects, a specialized database (e.g., a vector store, a graph database, etc.) may be used to store and query the embeddings. Embedding databases may include specialized features, such as efficient retrieval, similarity search and scalability. For example, the server computing device 104 may include a local electronic embedding database (not depicted). In some aspects, a remote embedding database service may be used (e.g., via the cloud APIs 1 14). Such a remote embedding database service may be based on an open source or proprietary model (e.g., Milvus, Pinecone, Redis, Postgres, MongoDB, Facebook Al Similarity Search (FAISS), etc.). The server computing device 104 may include instructions (e.g., in the data collection module 172) for adding training data to one or more specialized databases, and for accessing it to train models.
[0053] The present techniques may include language modeling, wherein one or more deep learning models are trained by processing token sequences using an LLM architecture. For example, in some aspects, a transformer architecture may be used to process a sequence of tokens. Such a transformer model may include a plurality of layers including self-attention and feedforward neural networks. This architecture may enable the LLM to learn contextual relationships between the tokens, and to predict the next token in a sequence, based upon the preceding tokens. During training, the LLM is provided with the sequence of tokens and it learns to predict a probability distribution over the next token in the sequence. This training process may include updating one or more model parameters (e.g., weights or biases) using an objective function that minimizes the difference between the predicted distribution and a truenext token in the training data. Particular techniques for training are discussed in further detail below in Fig. 2.
[0054] Alternatives to the transformer architecture may include recurrent neural networks, long short-term memory networks, gated recurrent networks, convolutional neural networks, recursive neural networks, and other modeling architectures.
[0055] In some aspects, the modules 170 may include instructions for performing pretraining of an LLM in a pretraining module 176. The pretraining module 176 may include one or more sets of instructions for performing pretraining, which as used herein, generally refers to a process that may span pre-processing of training data via the data pre-processing module 174 and initialization of an as-yet untrained LLM. In general, a pre-trained model is one that has no prior training of specific tasks. For example, the pretraining module 176 may include instructions that initialize one more model weights. In some aspects, pretraining module 176 may initialize the weights to have random values. The pretraining module 176 may train one or more LLMs using unsupervised learning, wherein the one or more LLMs process one or more tokens (e.g., preprocessed data output by the data pre-processing module 174) to learn to predict one or more elements (e.g., tokens). The pretraining module 176 may include one or more optimizing objective functions that the pretraining module 176 applies to the one or more LLMs, to cause the one or more LLMs to predict one or more most-likely next tokens, based on the likelihood of tokens in the training data. In general, the pretraining module 176 causes the one or more LLMs to learn linguistic features such as grammar and syntax. The pretraining module 176 may include additional steps, including training, data batching, hyperparameter tuning, and / or model checkpointing.
[0056] The pretraining module 176 may include instructions for generating an LLM that is pretrained for a general purpose, such as general text processing / understanding. This LLM may be known as a “base model” in some aspects. The base model may be further trained by downstream training process(es), for example, those training processes described with respect to the fine-tuning module 178. The pretraining module 176 generally trains foundational LLMs that have general understanding of language and / or knowledge. Pretraining may be a distinct stage of LLM training in which training data of a general and diverse nature (i.e. , not specific to any particular task or subset of knowledge) is used to train the one or more LLMs. In some aspects, a single LLM may be trained and copied. Copies of this LLM may serve as respective base models for a plurality of fine-tuned LLMs.
[0057] In some aspects, base models may be trained to have specific levels of knowledge common to more advanced agents. For example, the pretraining module 176 may train amedical student base model that may be subsequently used to fine tune an internist LLM, a surgeon LLM, a resident LLM, etc. In this way, the base model can start from a relatively advanced stage, without requiring pretraining of each more advanced LLM individually. This strategy represents an advantageous improvement, because pretraining can take a long time (many days), and pretraining the common base model only requires that pretraining process to be performed once.
[0058] The modules 170 may include a finetuning module 178. The finetuning module 178 may include instructions that train the one or LLMs further to perform specific tasks. Specifically, the fine-tuning module 178 may include instructions that train one or more LLMs to generate respective language outputs (e.g., text generation), summarization, question answering, NER, and / or translation activities based on healthcare data.
[0059] Continuing the example, the fine-tuning module 178 may include sets of instructions for retrieving one or more structured data sets, such as time series generated by the data preprocessing module 174. These structured data sets may be sorted by time and / or speaker to train one or more LLMs that may be used within the environment 100 to simulate debate. For example, the fine-tuning module 178 may include instructions for configuring an objective function for performing a specific task, such as generating text that is similar to text found within the corpus of training data associated with a particular individual by role. For example, the fine- tuning module 178 may include instructions for fine-tuning a pathologist LLM, based on a base language model. These fine-tuning instructions may select statements of pathologists from one or more databases including electronic health record database 108 (or another data source). A medical resident LLM may be fine-tuned by the fine-tuning module 178, wherein the base model is the same used to fine-tune the pathologist LLM. The fine-tuning module 178 may train many (e.g., hundreds or more) additional LLMs.
[0060] In some aspects, the fine-tuning module 178 may include sets of instructions for retrieving one or more labeled data sets. The labeled data sets may comprise text including instances of Pll annotated by entity labels, such as patient name, birth date, home address, telephone number, e-mail address, etc. The fine-tuning module 178 may use the labeled data sets to train one or more LLMs to identify instances of Pll in a corpus of text using NER.
[0061] In some aspects, the fine-tuning module 178 may include user-selectable parameters that affect the fine-tuning of the one or more LLMs. For example, a “caution” bias parameter may be included that represents medical conservativeness. This bias parameter may be adjusted to affect the cautiousness with which the resulting trained LLM (i.e., agent) approachesmedical decision-making. Additional LLMs may be trained, for additional personas / tasks, as discussed below.
[0062] In some aspects, to manage complexity of fine-tuning and other machine learning operations of the server computing device 104, one or more open source frameworks may be used. Example frameworks include TensorFlow, Keras, MXNet, Caffe, SciKit learn, PyTorch. Specifically for training and operating LLMs, frameworks such as OpenLLM and LangChain may be used, in some aspects. The fine-tuning module 178 may use an algorithm such as stochastic gradient descent or another optimization technique to adjust weights of the pretrained LLM.
[0063] Fine-tuning may be an optional operation, in some aspects. In some aspects, training may be performed by the training module 180 after pretraining by the pretraining module 176. In some aspects, the training module 180 may perform task-specific training like the fine-tuning module 178, on a smaller scale or with a more tailored objective. For example, whereas the fine-tuning module 178 may fine tune a model to learn knowledge corresponding to a surgeon, the model training module 180 may further train the model to learn knowledge of a plastic surgeon, an orthopedic surgeon, etc.
[0064] The training module 180 may include one or more submodules, including the checkpointing module 182, the hyperparameter tuning module 184, the validation and testing module 186, and the auto-prompting module 188. The checkpointing module 182 may perform checkpointing, which is saving of an LLM’s parameters. The checkpointing module 182 may store checkpoints during training and at the conclusion of training, for example, in the model electronic database 1 12. In this way, the LLM may be run (e.g., for testing and validation) at multiple stages and its training parameters loaded, and also retrained from a checkpoint. In this way, the LLM can be run and trained forward without being re-trained from the beginning, which may save significant time (e.g., days of computation). The hyperparameter tuning module 184 may include hyperparameters such as batch size, model size, learning rate, etc. These hyperparameters may be adjusted to influence LLM training. The hyperparameter tuning module 184 may include instructions for tuning hyperparameters by successive evaluation. The validation and testing module 186 may include sets of instructions for validating and testing one or more LLMs, including those generated by the pretraining module 176, the fine-tuning module 178 and the training module 180. The auto-prompting module 188 may include sets of instructions for performing auto-prompting of one or more LLMs. Specifically, the autoprompting module 188 may enrich a prompt with additional information. The auto-prompting module 188 may include additional information in a prompt, so that the LLM receiving theprompt has additional context or directions that it can use. This may allow the auto-prompting module 188 to fine-tune a base model using one-shot or few-shot learning, in some aspects. The auto-prompting module 188 may also be used to focus the output of the one or more LLMs.
[0065] In some aspects, the training module 180 may train multi-modal LLMs. For example, the training module 180 may train a plurality of LLMs each capable of drawing from multimodal data types such as written text, imaging data, laboratory data, real-time monitoring data, pathology images, etc. In some cases, the training module 180 may train a single LLM capable of processing the multimodal data types.
[0066] The operating module 190 may operate one or more trained LLMs. Specifically, the operation module 190 may initialize one or more trained LLMs, load parameters into the LLM(s), and provide the LLM(s) with inference data (e.g., prompt inputs). In some aspects, the operation module 190 may deploy one or more trained LLMs (e.g., a pretrained LLM, a fine tuned LLM, and / or a trained LLM) onto a cloud computing device (e.g., via the API 166). The operation module 190 may receive one or more inputs, for example from the client computing device 102, and provide those inputs (e.g., one or more prompts) to the trained LLM. In some aspects, the API 166 may include elements for receiving requests to the LLM, and for generating outputs based on LLM outputs. For example, the API 166 may include a RESTful API that receives a GET or POST request including a prompt parameter. The operation module 190 may receive the request from the API 166, pass the prompt parameter into the trained LLM, and receive a corresponding output. For example, the prompt parameter may be “What is the smallest bone in the human body?” The prompt output may be “The stapes bone of the inner ear is the smallest bone in the human body.”
[0067] The operation module 190 may operate LLMs in different modes. For example, in a first mode, the operation module 190 may receive a prompt input via the client computing device 102 and provide that input to one or more of a plurality of LLMs for processing. The output of the or more LLMs may be collected and transmitted back to the client computing device 102 for display. The outputs may be labeled according to an identifier of each LLM (e.g., “pathologist,” “surgeon,” “medical student,” etc.). In the first mode, the operation module 190 may receive additional inputs from the client computing device 102 that enable the user to interact with the one or more trained LLMs in a question-answer format.
[0068] A physician may want to use the present techniques to request and receive a clinical analysis of a patient. Clinical analysis data herein may include Pll instances, such as name, identification number, birthdate, birthplace, age, race, and other demographic data or other data for identifying a patient. Clinical analysis data herein may include clinical notes, differentialdiagnoses, clinical recommendations, clinical summaries, etc., and the like. The physician may access the client computing device 102. The input processing module 146 may identify the patient via the name or other identifying information provided and transmit the patient identifier and the clinical analysis request to the server computer device 104 via the API 166 or chatbot 192. In some embodiments, the physician may specify what categories of PH, e.g., name, identification number, and / or birthdate, should be removed from the clinical analysis output and what categories of PH, e.g., age and / or sex, should remain in the clinical analysis output. In some embodiments, a provider may specify that certain categories of PH be removed from the clinical analysis output by default.
[0069] The server computing device 104, using the data collection module 172, may retrieve patient information (e.g., electronic health records including test data, diagnostic data, medical condition data, pathology data, medical treatment data, prescription data, time stamp data for various other data, etc.) from the EHR database 108 using the patient identifier. The server computing device 104 may send the patient information and clinical analysis request to one or more LLMs via the operating module 190 or chatbot 192. In response to the patient information and clinical analysis request being sent, the one or more LLMs may generate a clinical analysis.
[0070] The clinical analysis may include one or more instances of PH, such as Pll obtained from training data. The LLM may identify the PH instances in the clinical analysis. The LLM may split up the clinical analysis text into tokens, where each token represents a word or a sub-word. The LLM may generate embeddings of the tokens using a technique, such as Word2Vec or ELMO, to capture the semantic meaning and the contexts of words or tokens. The LLM may analyze the embeddings to determine the patterns and relationships within the clinical analysis text to accurately identify the PH instances.
[0071] In one aspect, the LLM may deidentify the PH instances in the clinical analysis. For example, the LLM may replace “Mark Thompson” and “05 / 17 / 1980” with “<PATIENT NAME>” and “<PATIENT BIRTH DATE>” or with “John Doe” and “MM / DD / 1980.” In another aspect, the LLM may generate and output a list of Pll instances. The server computing device 104, via the operating module 190 or chatbot 192, may receive the clinical analysis, deidentified clinical analysis, and / or Pll list from the LLM. The server computing device 104 may deidentify the clinical analysis using the operating module 190. The server computing device 104 may transmit the deidentified clinical analysis to the client computing device 102.
[0072] As discussed, in some aspects, multi-modal modeling may be used. The data preprocessing module 174 may, for example, process and understand image data, audio data, video data, etc. The server computing device 104 may interpret and respond to queries thatinvolve understanding content from these different modalities. For example, the server computing device 104 may include an image processing module (not depicted) including instructions for performing image analysis on images provided by users, or images retrieved from patient EHR data. In some aspects, the server computing device 104 may generate outputs in modalities other than text. For example, the server computing device 104 may generate an audio response, an image, etc. Combining multi-modal data may enable the present LLMs to perform more comprehensive analysis of patient conditions, based on information processed in multiple different modes simultaneously.
[0073] The operating module 190 may include a set of computer-executable instructions that when executed by one or more processors (e.g., the processors 160) cause a computer (e.g., the server computing device 104) to perform retrieval-augmented generation. Specifically, the operating module 190 may perform retrieval-augmented generation based upon inputs or queries received from the user. This allows the operating module 190 to tailor responses of an LLM based on the specific input and context, such as the medical issue under discussion. For example, one or more LLMs may be pre-trained, fine-tuned and / or trained as discussed above. During that training, the LLM may learn to generate tokens based on general language understanding as well as application-specific training. Such an LLM at that point may be static, insofar as it cannot access further information when presented with an input query.
[0074] When the LLM is used at runtime, however, such as when deployed in the environment 100, the operating module 190 may perform retrieval operations, such as searching or selecting information from a document, a database, or another source. The operating module 190 may include instructions for processing user input and for performing a keyword search, a regular expression search, a similarity search, etc. based upon that user input. The operating module 190 may input the results of that search, along with the user input, into the trained LLM. Thus, the trained LLM may process this additional retrieved information to augment, or contextualize, the generation of tokens that represent responses to the user’s query. In sum, retrieval augmented generation applied in this manner allows the LLM to dynamically generate outputs that are more relevant to the user’s input query at runtime. Information that may be retrieved may include data corresponding to a patient (e.g., patient demographic information, medical history, clinical notes, diagnoses, medications, allergies, immunizations, laboratory results, oncology information, radiation and imaging information, vitals, etc.) and additional training information, such as medical journals, notes or speech transcripts from symposia or other meetings / conferences, etc.
[0075] The present techniques may trigger retrieval augmented generation by processing a prompt, in some aspects. For example, a prompt may be processed by the input processing module 146 of the client computing device 102, prior to processing the prompt by the one or more LLMs. The input processing module 146 may trigger retrieval augmented generation based on the presence of certain inputs, such as patient information, or a request for specific information, in the form of keywords. The input processing module 146 may perform natural language processing functions to determine whether the prompt should be processed using retrieval augmented generation prior to being provided to the trained model.
[0076] As discussed above, prompts may be received via the input processing module 146 of the client computing device 102 and transmitted to the server computing device 104 via the electronic network 106. In some aspects, the output of the LLM may be modulated prior to being transmitted, output, or otherwise displayed to a user.
[0077] The modules 170 may include a chatbot 192 which may be programmed to simulate human conversation, interact with users, understand their needs, and recommend an appropriate line of action with minimal and / or no human intervention, among other things. This may include providing the best response of any query that it receives and / or asking follow-up questions. For example, a patient may input a question to the chatbot such as “What medication was prescribed in my treatment plan?" and the chatbot may respond “Tylenol.”
[0078] In some embodiments, the chatbot 192 may be configured to utilize Al and / or ML techniques. For instance, the chatbot 192 may be a CHATGPT chatbot, an INSTRUCTGPT bot, a CODEX bot, or a GOOGLE BARD bot. The chatbot 192 may employ supervised or unsupervised ML techniques, which may be followed by and / or used in conjunction with reinforced or reinforcement learning techniques. The chatbot 192 may employ the techniques utilized for CHATGPT, INSTRUCTGPT bot, CODEX bot, or GOOGLE BARD bot.
[0079] The modules 170 may include a deidentification module 194. In some aspects, the deidentification module 194 may be configured to receive a list of PH instances found by the LLM in the clinical analysis. The deidentification module 194 may search for the PI I instances in the clinical analysis and replace with placeholder or altered text to produce a deidentified clinical analysis.
[0080] The client computing device 102 and the server computing device 104 may communicate with one another via the network 106. In some aspects, the client computing device 102 and / or the server computing device 104 may offload some or all of their respective functionality to the one or more cloud APIs 1 14. In aspects, the one or more cloud APIs 114 may include one or more public clouds, one or more private clouds and / or one or more hybridclouds. The one or more cloud APIs 114 may include one or resources provided under one or more service models, such as Infrastructure as a Service (laaS), Platform as a Service (PaaS), Software as a Service (SaaS), and Function as a Service (PaaS). For example, the one or more cloud APIs 114 may include one or more cloud computing resources, such as computing instances, electronic databases, operating systems, email resources, etc. The one or more cloud APIs 1 14 may include distributed computing resources that enable, for example, the pretraining module 176 and / or other modules 170 to distribute parallel LLM training jobs across many processors.
[0081] In some aspects, the one or more cloud APIs 114 may include one or more language operation APIs, such as OPENAI, BING, CLAUDE.AI, etc. In other aspects, the one or more cloud APIs 1 14 may include an API configured to operate one or more open source LLMs, such as LLAMA 2.
[0082] The electronic network 106 may be a collection of interconnected devices, and may include one or more local area networks, wide area networks, subnets, and / or the Internet. The network 106 may include one or more networking devices such as routers, switches, etc. Each device within the network 106 may be assigned a unique identifier, such as an IP address, to facilitate communication. The network 106 may include wired (e.g., Ethernet cables) and wireless (e.g., Wi-Fi) connections. The network 106 may include a topology such as a star topology (devices connected to a central hub), a bus topology (devices connected along a single cable), a ring topology (devices connected in a circular fashion), and / or a mesh topology (devices connected to multiple other devices). The electronic network 106 may facilitate communication via one or more networking protocols, such as packet protocols (e.g., Internet Protocol (IP)) and / or application-layer protocols (e.g., HTTP, SMTP, SSH, etc.). The network 106 may perform routing and / or switching operations using routers and switches. The network 106 may include one or more firewalls, file servers and / or storage devices. The network 106 may include one or more subnetworks such as a virtual LAN (VLAN).
[0083] The environment 100 may include one or more electronic databases, such as a relational database that uses structured query language (SQL) and / or a NoSQL database or other schema-less database suited for the storage of unstructured or semi-structured data.
[0084] The present techniques may store training data, training parameters, and / or trained LLMs in an electronic database such as the model database 1 12. Specifically, one or more trained LLMs may be serialized and stored in a database (e.g., as a binary, a JSON object, etc.). Such an LLM can later be retrieved, deserialized, and loaded into memory and then used for predictive purposes. The one or more trained LLMs and their respective training parameters(e.g., weights) may also be stored as blob objects. Cloud computing APIs may also be used to stored trained LLMs via the cloud APIs 114. Examples of these cloud computing APIs include AWS SAGEMAKER, GOOGLE Al PLATFORM and AZURE MACHINE LEARNING.Example Training and / or Fine Tuning of an LLM
[0085] Figure 2 depicts an example combined block and logic diagram 200 for training and / or fine tuning an LLM, in which the techniques described herein may be implemented, according to some embodiments, including for example methods herein such as shown in Figure 4A and / or Figure 4B. Some of the blocks in Figure 2 may represent hardware and / or software components, other blocks may represent data structures or memory storing these data structures, registers, or state variables (e.g., 212), and other blocks may represent output data (e.g., 225). Input and / or output signals may be represented by arrows labeled with corresponding signal names and / or other identifiers. The methods and systems may include one or more servers 202, 204, 206, such as the server computing device 104, cloud APIs 114, or an external computing device.
[0086] In one aspect, the server 202 may fine-tune a pretrained LLM 210. The pretrained LLM 210 may be obtained by the server 202 and be stored in a memory, such as the model database 1 12 or memory 164. The pretrained LLM 210 may be loaded into a training module, such as the fine-tuning module 178 or training module 180, by the server 202 for retraining / fine-tuning. A supervised training dataset 212 may be used to fine-tune the pretrained LLM 210 wherein each data input prompt to the pretrained LLM 210 may have a known output response for the pretrained LLM 210 to learn from. The supervised training dataset 212 may be stored in a memory of the server 202, e.g., the memory 164 or the EHR database 108. In one aspect, data labelers may create the supervised training dataset 212 prompts and appropriate responses.The pretrained LLM 210 may be fine-tuned using the supervised training dataset 212 resulting in the shrink and fine tune (SFT) LLM 215 which may provide appropriate responses to user prompts once trained. The trained SFT LLM 215 may be stored in a memory of the server 202, e.g., model database 1 12 or memory 164.
[0087] In some embodiments, the server 202 may fine-tune the pretrained LLM 210 using a set of vectors associated with a set of training data. In some instances, the set of training data may include prompts associated with questions and documents, and responses associated with the prompts. Creating the set of vectors may include (1 ) splitting the text of the prompts, associated questions, and / or associated documents into semantic clusters, and (2) encoding the semantic clusters as the set of vectors. The semantic clusters may be one or more words, a portion of a word, or a character. A distance between the vectors (e.g., a cosine distance, a Euclideandistance) may depend on a relevance between the semantic clusters corresponding to the vectors.
[0088] In one aspect, training the LLM may include the server 204 training a reward model 220 to provide as an output a scaler value / reward 225. The reward model 220 may be required to leverage Reinforcement Learning with Human Feedback (RLHF) in which an LLM (e.g., LLM 250) learns to produce outputs which maximize its reward 225, and in doing so may provide responses which are better aligned to user prompts.
[0089] Training the reward model 220 may include the server 204 providing a single prompt 222 to the SFT ML model 215 as an input. The input prompt 222 may be provided via an input device (e.g., a keyboard) of the server 204. The prompt 222 may be previously unknown to the SFT ML model 215, e.g., the labelers may generate new prompt data, the prompt 222 may include health records stored in EHR database 108, and / or any other suitable prompt data. The SFT ML model 215 may generate multiple, different output responses 224A, 224B, 224C, 224D to the single prompt 222. The server 204 may output the responses 224A, 224B, 224C, 224D via an I / O module to a user interface device, such as a display (e.g., as text responses), a speaker (e.g., as audio / voice responses), and / or any other suitable manner of outputting of the responses 224A, 224B, 224C, 224D for review by the data labelers.
[0090] The data labelers may provide feedback via the server 204 on the responses 224A, 224B, 224C, 224D when ranking 226 them from best to worst based upon the prompt-response pairs. The data labelers may rank 226 the responses 224A, 224B, 224C, 224D by labeling the associated data. The ranked prompt-response pairs 228 may be used to train the reward model 220. In one aspect, the server 204 may load the reward model 220 via an operating module (e.g., the operating module 190) and train the reward model 220 using the ranked response pairs 228 as input. The reward model 220 may provide as an output the scalar reward 225.
[0091] In one aspect, the scalar reward 225 may include a value numerically representing a human preference for the best and / or most expected response to a prompt, i.e., a higher scaler reward value may indicate the user is more likely to prefer that response, and a lower scalar reward may indicate that the user is less likely to prefer that response. For example, inputting the “winning” prompt-response (i.e., input-output) pair data to the reward model 220 may generate a winning reward. Inputting a “losing” prompt-response pair data to the same reward model 220 may generate a losing reward. The reward model 220 and / or scalar reward 225 may be updated based upon labelers ranking 226 additional prompt-response pairs generated in response to additional prompts 222.
[0092] In one example, a data labeler may provide to the SFT ML model 215 as an input prompt 222, “Describe the sky.” The input may be provided by the labeler via the client computing device 102 over network 106 to the server 204 running a chatbot application utilizing the SFT ML model 215. The SFT ML model 215 may provide as output responses to the labeler via the user device 102: (i) “the sky is above” 224A; (ii) “the sky includes the atmosphere and may be considered a place between the ground and outer space” 224B; and (iii) “the sky is heavenly” 224C. The data labeler may rank 226, via labeling the prompt-response pairs, prompt-response pair 222 / 224B as the most preferred answer; prompt-response pair 222 / 224A as a less preferred answer; and prompt-response 222 / 224C as the least preferred answer. The labeler may rank 226 the prompt-response pair data in any suitable manner. The ranked prompt-response pairs 228 may be provided to the reward model 220 to generate the scalar reward 225.
[0093] While the reward model 220 may provide the scalar reward 225 as an output, the reward model 220 may not generate a response (e.g., text). Rather, the scalar reward 225 may be used by a version of the SFT LLM 215 to generate more accurate responses to prompts, i.e., the SFT LLM 215 may generate the response such as text to the prompt, and the reward model 220 may receive the response to generate a scalar reward 225 of how well humans perceive it. Reinforcement learning may optimize the SFT LLM 215 with respect to the reward model 220 which may realize the configured LLM 250.
[0094] In one aspect, the server 206 may train the LLM 250 (e.g., via the fine-tuning module 178 or training module 180) to generate a response 234 to a random, new and / or previously unknown user prompt 232. To generate the response 234, the LLM 250 may use a policy 235 (e.g., algorithm) which it learns during training of the reward model 220, and in doing so may advance from the SFT LLM 215 to the LLM 250. The policy 235 may represent a strategy that the LLM 250 learns to maximize the reward 225. As discussed herein, based upon promptresponse pairs, a human labeler may continuously provide feedback to assist in determining how well the LLM’s 250 responses match expected responses to determine the rewards 225. The rewards 225 may feed back into the LLM 250 to evolve the policy 235. Thus, the policy 235 may adjust the parameters of the LLM 250 based upon the rewards 225 it receives for generating good responses. The policy 235 may update as the LLM 250 provides responses 234 to additional prompts 232.
[0095] In one aspect, the response 234 of the LLM 250 using the policy 235 based upon the reward 225 may be compared using a cost function 238 to the SFT LLM 215 (which may not use a policy) response 236 of the same prompt 232. The cost function 238 may be trained in asimilar manner and / or contemporaneous with the reward model 220. The server 206 may compute a cost 240 based upon the cost function 238 of the responses 234, 236. The cost 240 may reduce the distance between the responses 234, 236, i.e., a statistical distance measuring how one probability distribution is different from a second, in one aspect the response 234 of the LLM 250 versus the response 236 of the SFT LLM 215. Using the cost 240 to reduce the distance between the responses 234, 236 may avoid a server over-optimizing the reward model 220 and deviating too drastically from the human-intended / preferred response. Without the cost 240, the LLM 250 optimizations may result in generating responses 234 which are unreasonable but may still result in the reward model 220 outputting a high reward 225.
[0096] In one aspect, the responses 234 of the LLM 250 using the current policy 235 may be passed by the server 206 to the rewards model 220, which may return the scalar reward 225. The response 234 from the LLM 250 may be compared via the cost function 238 to the SFT ML LLM’s 215 response 236 by the server 206 to compute the cost 240. The server 206 may generate a final reward 242 which may include the scalar reward 225 offset and / or restricted by the cost 240. The final reward 242 may be provided by the server 206 to the LLM 250 and may update the policy 235, which in turn may improve the functionality of the LLM 250.
[0097] To optimize the LLM 250 over time, RLHF via the human labeler feedback may continue ranking 226 responses of the LLM 250 versus outputs of earlier / other versions of the SFT LLM 215, i.e., providing positive or negative rewards 225. The RLHF may allow the servers (e.g., servers 204, 206) to continue iteratively updating the reward model 220 and / or the policy 235. As a result, the LLM 250 may be retrained and / or fine-tuned based upon the human feedback via the RLHF process, and throughout continuing conversations may become increasingly efficient.
[0098] Although multiple servers 202, 204, 206 are depicted in the example block and logic diagram 200, each providing one of the three steps of the overall LLM 250 training, fewer and / or additional servers may be utilized and / or may provide the one or more steps of the LLM 250 training. In one aspect, one server, e.g., server computing device 104, may provide the entire LLM 250 training.
[0099] Figure 3 depicts a combined block and logic diagram 300 for training and / or fine tuning an LLM, in which the techniques described herein may be implemented, according to some embodiments, including for example methods herein such as shown in Figure 4A and / or Figure 4B. Some of the blocks in Figure 3 may represent hardware and / or software components, other blocks may represent data structures or memory storing these data structures, registers, or state variables, and other blocks may represent output data. Input and / or output signals may berepresented by arrows labeled with corresponding signal names and / or other identifiers. The methods and systems may include one or more servers, such as the server computing device 104, cloud APIs 1 14, or an external computing device.
[0100] In one aspect, the fine-tuning module 178 may fine tune a pretrained LLM 31 OA. The pretrained LLM 31 OA may be obtained by server computing device 104 and be stored in a memory, such as the model database 112 or memory 164. The pretrained LLM 31 OA may be loaded into a training module, such as the fine-tuning module 178 or training module 180, by the server 104 for retraining / fine-tuning.
[0101] The pretrained LLM 31 OA may be configured with a set of initial hyperparameters 320. For an LLM, for example, the set of initial hyperparameters 320 may include specified values for the number of epochs, learning rate, etc.
[0102] In one aspect, the data pre-processing module 174 may retrieve a labeled dataset from the EHR database 108 and split the labeled dataset into a training dataset 330A and a test dataset 330B. The training dataset 330A and test dataset 330B may include unstructured text comprising one or more instances of PH. The one or more instances of PH may be labeled with a ground truth Pll classification, such as patient name, patient birthdate, patient identifier, etc.
[0103] In one aspect, the fine-tuning module 178 may provide the training dataset 330A to the pretrained LLM 31 OA in a fine-tuning step. The training dataset 330A propagates forward through the pretrained LLM 31 OA, causing the pretrained LLM 31 OA to generate one or more predictions. A prediction may be a Pll identification, i.e., PH or not Pll, and / or a Pll classification for a plurality of text chunks in the training dataset 330A.
[0104] In some embodiments, the fine-tuning module 178 computes a loss metric 340 by comparing the predictions generated by the pretrained LLM 31 OA to the ground truth labels in the training dataset 330A. For example, the fine-tuning module 178 may apply cross entropy as the loss function. The loss may be backpropagated through the pretrained LLM 31 OA to adjust the weights and biases of the model. The fine-tuning, which may occur over a plurality, e.g., 25, of epochs, may cause the pretrained LLM 31 OA to adjust model weights to minimize the loss function.
[0105] In some embodiments, after fine-tuning is complete, the pretrained LLM 31 OA becomes a finetuned LLM 31 OB. The training module 180 may provide the test dataset 330B as input to the finetuned LLM 31 OB in a testing step. The testing step may determine if thefinetuned LLM 31 OB has been overfitted on the training dataset 330A. The finetuned LLM 31 OB generates one or more Pll predictions from the test dataset 330B input.
[0106] In some embodiments, the training module 180 may calculate a prediction error 340 by comparing the predicted Pll instances and / or Pll classifications to the manually labeled ground truth Pll labels in the test dataset 330B. The prediction error 480 may include the F- score, precision, and recall metrics. Optionally, in some embodiments, the training and / or fine- tuning operations described herein may be periodically repeated (e.g., daily, weekly, monthly) to incorporate new input data and / or feedback from users / providers. This retraining may produce updated versions of the respective models. In this way, some embodiments may improve over time to produce outputs that more closely match the particularized uses and preferences of its users / providers.Example Computer-Implemented Methods
[0107] FIGs. 4A and 4B depict example computer-implemented methods 400A and 400B for deidentifying Pll using LLMs, in accordance with various embodiments herein, and as may be implemented by the example computing environment 100.
[0108] At block 410, the methods 400A and 400B include training an LLM, such as pretrained LLM 31 OA, to generate a clinical analysis. The LLM may be trained on a first set of training data, in particular clinical records. The clinical records may include health information about a plurality of patients. The clinical records may comprise unstructured text that includes Pll. The data pre-processing module 176 may clean and parse the unstructured text of the clinical records. In some examples, the LLM is trained on fully-identified, full-context clinical records. This may be particularly useful for troubleshooting and clinical evaluation within a host institution.
[0109] At block 420, the methods 400A and 400B include training the LLM, such as the pretrained LLM 31 OA, to identify Pll. The LLM may be trained on a second set of training data, in particular Pll training data. The Pll training data may include Pll examples labeled with one or more Pll categories. The one or more Pll categories may include patient name, birth date, social security name, address, e-mail address, telephone number, etc.
[0110] At block 430, the methods 400A and 400B include receiving a clinical analysis request and a patient identifier. The clinical analysis request may be received from the client computing device 102 to the server computing device 104. The patient identifier may include a patient name, patient number, or any other suitable identifier.
[0111] At block 440, the methods 400A and 400B include sending the clinical analysis request and one or more patient records to the LLM, such as finetuned LLM 31 OB. The one or more patient records may be associated with the patient identifier. The one or more patient records may be retrieved from the EHR database 108. In one aspect, to validate the PH identification ability of the LLM, the one or more patient records include one or more known PI I instances.
[0112] In one aspect, sending the clinical analysis request and one or more patient records to the LLM may cause the LLM to generate a clinical analysis of the one or more patient records. In one aspect, sending the clinical analysis request and one or more patient records to the LLM may cause the LLM to identify instances of Pll in the generated clinical analysis. In one aspect, in identifying the Pll instances, the LLM splits the generated clinical analysis into tokens. In one aspect, in identifying the Pll instances, the LLM generates embeddings of the tokens that capture the tokens’ semantic meaning and contextual usage. In one aspect, the LLM assigns a Pll category to each instance of the Pll instances.
[0113] In method 400A, the LLM deidentifies the Pll instances in the clinical analysis to generate a deidentifed clinical analysis. In one aspect, deidentifying includes replacing the Pll instances with placeholder text. In one aspect, deidentifying includes altering the Pll instances. At block 450A, the method 400A includes receiving the deidentifed clinical analysis from the LLM. In one aspect, deidentifying includes excluding specific data within the Pll instances.
[0114] In method 400B, the LLM compiles the identified Pll instances into a Pll list, which may be generated from collating all Pll instances into a dedicated list. At block 450B, the method 400B includes receiving the clinical analysis and the Pll list from the LLM. At block 452B, the method 400B includes deidentifying the clinical analysis. In one aspect, that post LLM deidentification may be performed by a secondary, independent LLM that could be applied to any system to make the resulting data safe for preserving privacy. In one aspect, the post LLM deidentification may be performed by the deidentification module 194.
[0115] In one aspect, to validate the Pll identification ability of the LLM, the methods 400A and 400B include searching the deidentified clinical analysis for the one or more known Pll instances.
[0116] At block 460, the methods 400A and 400B include transmitting the deidentified clinical analysis to the client device.
[0117] While the methods 400A and 400B are shown separately, in some examples, the two methods 400A and 400B may be performed sequentially, e.g., implemented as a dual-layered safeguard. Initially, the process 400A may be performed to produce deidentified data. Then, the method 400B may be performed to identify and rectify any inadvertently generated identifiable information still present in the deidentified data from the process 400A, ensuring comprehensive protection against unintentional data exposure. Such a dual-layered configuration may be implemented in scenarios where a first provider sets the trained deidentification operations by training an LLM at 400A, while different providers may set different trained deidentification operations by training a subsequent LLM at 400B according to different training data. For example, some providers may determine that some Pll data is need not be deidentified, while other providers may determine that the same Pll data should be deidentified. A dual-layered configuration allows the second provider to deploy a trained LLM that further affects the output of the trained LLM of a first provider.Example Clinical Analysis
[0118] FIG. 5A depicts an example clinical analysis 500A including Pll, according to some aspects. In some embodiments, an LLM, such as the finetuned LLM 31 OB, generated the example clinical analysis 500A, including a provider name 510a, a patient name 510b, the patient’s age 510c, and the patient’s gender 51 Od. The example clinical analysis 500A includes dates of medical treatment / testing (51 Od and 51 Oe) with a corresponding description for each date. Also included is a name 51 Of of a next of kin. FIG. 5B depicts an example clinical analysis 500A similar to that of analysis 500A but wherein Pll data has been redacted by the finetuned LLM 310B.Additional Considerations
[0119] The following considerations also apply to the foregoing discussion. Throughout this specification, plural instances may implement operations or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0120] It should also be understood that, unless a term is expressly defined in this patent using the sentence "As used herein, the term " " is hereby defined to mean . . . " or a similar sentence, there is no intent to limit the meaning of that term, either expressly or by implication, beyond its plain or ordinary meaning, and such term should not be interpreted to be limited in scope based on any statement made in any section of this patent (other than the language of the claims). To the extent that any term recited in the claims at the end of thispatent is referred to in this patent in a manner consistent with a single meaning, that is done for sake of clarity only so as to not confuse the reader, and it is not intended that such claim term be limited, by implication or otherwise, to that single meaning. Finally, unless a claim element is defined by reciting the word "means'1and a function without the recital of any structure, it is not intended that the scope of any claim element be interpreted based on the application of 35 U.S.C. § 1 12(f).
[0121] Unless specifically stated otherwise, discussions herein using words such as "processing," "computing," "calculating," "determining," "presenting," "displaying," or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or a combination thereof), registers, or other machine components that receive, store, transmit, or display information.
[0122] As used herein any reference to "one aspect" or "an aspect" means that a particular element, feature, structure, or characteristic described in connection with the aspect is included in at least one aspect. The appearances of the phrase "in one aspect" in various places in the specification are not necessarily all referring to the same aspect.
[0123] As used herein, the terms "comprises," "comprising," "includes," "including," "has," "having" or any other variation thereof, are intended to cover a non-exclusive inclusion. For example, a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, unless expressly stated to the contrary, "or" refers to an inclusive or and not to an exclusive or. For example, a condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or not present), A is false (or not present) and B is true (or present), and both A and B are true (or present).
[0124] In addition, use of "a" or "an" is employed to describe elements and components of the aspects herein. This is done merely for convenience and to give a general sense of the invention. This description should be read to include one or at least one and the singular also includes the plural unless it is obvious that it is meant otherwise.
[0125] Upon reading this disclosure, those of skill in the art will appreciate still additional alternative structural and functional designs for implementing the concepts disclosed herein, through the principles disclosed herein. Thus, while particular aspects and applications have been illustrated and described, it is to be understood that the disclosed aspects are not limited to the precise construction and components disclosed herein. Various modifications, changes and variations, which will be apparent to those skilled in the art, may be made in thearrangement, operation and details of the method and apparatus disclosed herein without departing from the spirit and scope defined in the appended claims.
Claims
What is claimed is:1 . A computer-implemented method for deidentifying personally identifiable information (Pll), the method comprising: training, via one or more processors, a large language model (LLM) on a set of clinical records to generate a clinical analysis, wherein the set of clinical records comprises unstructured text including Pll; fine tuning, via the one or more processors, the LLM on a set of Pll training data to identify Pll, wherein the set of Pll training data comprises Pll examples labeled with one or more Pll categories; receiving, from a client device via the one or more processors, a clinical analysis request and a patient identifier; providing, to the LLM via the one or more processors, the clinical analysis request and one or more patient records associated with the patient identifier, causing the LLM to: generate a clinical analysis of the one or more patient records, identify one or more Pll instances in the clinical analysis, and receiving, from the LLM via the one or more processors, the clinical analysis and the one or more Pll instances; deidentifying, by the one or more processors, at least one of the one or more Pll instances in the clinical analysis to generate a deidentified clinical analysis; and transmitting, via the one or more processors, the deidentified clinical analysis to the client device.
2. The computer-implemented method of claim 1 , wherein providing the clinical analysis request and the one or more patient records associated with the patient identifier to the LLM further causes the LLM to: assign a Pll category from one or more Pll categories to at least one Pll instance of the one or more Pll instances.
3. The computer-implemented method of claim 1 or 2, wherein the deidentifying the at least one of the one or more PH instances in the clinical analysis to generate a deidentified clinical analysis is performed in a subsequent LLM.
4. The computer-implemented method of any one of claims 1 -3, wherein deidentifying the PH instances in the clinical analysis to generate the deidentified clinical analysis comprises replacing the PH instances with placeholder text.
5. The computer-implemented method of any one of claims 1 -3, wherein deidentifying the PH instances in the clinical analysis to generate the deidentified clinical analysis comprises altering the PH instances.
6. The computer-implemented method of any one of claims 1 -5, wherein the one or more patient records include one or more known PH instances, and further comprising: searching, via the one or more processors, the deidentified clinical analysis for the one or more known PH instances.
7. The computer-implemented method of any one of claims 1 -6, wherein identifying the one or more PH instances in the clinical analysis comprises: splitting the clinical analysis into a plurality of tokens; and generating an embedding for each token of the plurality of tokens, wherein the embedding captures the semantic meaning and contextual usage of the token.
8. A computer-implemented method for deidentifying personally identifiable information (Pll) comprising: training, via one or more processors, a large language model (LLM) on a set of clinical records to generate clinical analyses wherein the set of clinical records comprises unstructured text including PH;fine tuning, via the one or more processors, the LLM on a set of PH training data to identify PH, wherein the set of Pll training data comprises Pll examples labeled with one or more Pll categories; receiving, from a client device via the one or more processors, a clinical analysis request and a patient identifier; providing, to the LLM via the one or more processors, the clinical analysis request and one or more patient records associated with the patient identifier, causing the LLM to: generate a clinical analysis of the one or more patient records, identify one or more Pll instances in the clinical analysis, and compile the one or more Pll instances into a Pll list; receiving, from the LLM via the one or more processors, the clinical analysis and the Pll list; deidentifying, via the one or more processors using the Pll list, the one or more Pll instances from the clinical analysis to generate a deidentified clinical analysis; and transmitting, via the one or more processors, the deidentified clinical analysis to the client device.
9. The computer-implemented method of claim 8, wherein providing to the LLM further causes the LLM to: assign a Pll category to at least one Pll instance of the one or more Pll instances.
10. The computer-implemented method of claim 8 or 9, wherein the one or more Pll categories comprise patient name and patient birth date.11 . The computer-implemented method of any one of claims 8-10, wherein deidentifying the Pll instances in the clinical analysis to generate the deidentified clinical analysis comprises replacing the Pll instances with placeholder text.
12. The computer-implemented method of any one of claims 8-10, wherein deidentifying the PH instances in the clinical analysis to generate the deidentified clinical analysis comprises altering the PH instances.
13. The computer-implemented method of any one of claims 8-12, wherein the one or more patient records include one or more known PH instances, and further comprising: comparing, via the one or more processors, the Pll list to the one or more known PH instances.
14. The computer-implemented method of any one of claims 8-13, wherein identifying the one or more PH instances in the clinical analysis comprises: splitting the clinical analysis into a plurality of tokens; and generating an embedding for each token of the plurality of tokens, wherein the embedding captures the semantic meaning and contextual usage of the token.
15. A computing system for deidentifying personally identifiable information (PH), the system comprising: one or more processors, and one or more memories having stored thereon computer-executable instructions that, when executed by the one or more processors, cause the computing system to perform the method of any one of claims 1 -14.
Citation Information
Cited By
Method for medical record data extraction and summarization using large language artificial intelligence models
US20260162783A1