Machine learning-enabled determination of diagnostic codes

WO2025049485A3PCT designated stage expired Publication Date: 2025-05-08RGT UNIV OF CALIFORNIA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/044049
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-28
Filing Date
2024-08-27
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

The increasing burden of documentation in healthcare, driven by the need to input diagnostic codes into electronic health records (EHRs), leads to physician burnout and undermines face-to-face patient care.

Method used

A machine learning model is trained to automatically determine diagnostic codes from clinical notes using natural language processing (NLP) techniques, such as those employed by BERT models, to reduce the documentation burden and improve efficiency.

Benefits of technology

The model effectively reduces the time spent on documentation by automating the process of assigning diagnostic codes, thereby alleviating physician burnout and enhancing patient care interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024044049_08052025_PF_FP_ABST
    Figure US2024044049_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are methods, systems, and computer-readable media for machine learning enabled determination of diagnostic codes. Various embodiments describe training a machine learning model, testing and / or optimizing a machine learning model, and using a machine learning model. Such embodiments may include inputting clinical notes into a trained machine learning model and obtaining a diagnostic code from the machine learning model. In various embodiments, the machine learning model is trained on descriptions associated with diagnostic codes. In various embodiments, the clinical notes are obtained from electronic health records (EHR). Certain embodiments may utilize a similarity matrix to determine differences between a clinical note and the diagnostic code description. Certain embodiments describe optimizing a trained model, which may include iteratively adjusting a parameter or hyperparameter associated with the model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 MACHINE LEARNING-ENABLED DETERMINATION OF DIAGNOSTIC CODES CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No.63 / 534,990, filed August 28, 2023, which application is incorporated herein by reference in its entirety. INTRODUCTION Electronic Health Records (EHR) revolutionized medicine by enabling the creation of legible healthcare records and streamlining their access for continuity of care. This came at a cost, however, as increased screen time during medical visits has undermined personalized face- to-face patient care and vital doctor-patient interactions. Research from the National Academy of Medicine has found that nurses and doctors spend 50% of their workday on documentation. This increased work burden is a significant contributor to physician burnout. SUMMARY Provided are methods, systems, and computer-readable media for machine learning enabled determination of diagnostic codes. Various embodiments describe training a machine learning model, testing and / or optimizing a machine learning model, and using a machine learning model. Such embodiments may include inputting clinical notes into a trained machine learning model and obtaining a diagnostic code from the machine learning model. In various embodiments, the machine learning model is trained on descriptions associated with diagnostic codes. In various embodiments, the clinical notes are obtained from electronic health records (EHR). Certain embodiments may utilize a similarity matrix to determine differences between a clinical note and the diagnostic code description. Certain embodiments describe optimizing a trained model, which may include iteratively adjusting a parameter or hyperparameter associated with the model. BRIEF DESCRIPTION OF THE FIGURES FIGs. 1A-1B provide an overview of an AI model used in accordance with various embodiments. FIG.2 illustrates the pipeline for testing an AI model used in accordance with various embodiments. FIG.3 provides a distribution of length of clinical notes and diagnosis code descriptions used in an exemplary embodiment. FIG.4 provides a mean squared error (MSE) between learned representations of similar and distinct clinical notes used in an exemplary embodiment. Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 FIG. 5 illustrates a conventional BERT-based classifier used in an exemplary embodiment. DETAILEDDESCRIPTIONBefore the methods, systems and computer-readable media of the present disclosure are described in greater detail, it is to be understood that the methods, systems and computer- readable media are not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the methods, systems and computer-readable media will be limited only by the appended claims. Where a range of values is provided, it is understood that each intervening value, to the tenth of the unit of the lower limit unless the context clearly dictates otherwise, between the upper and lower limit of that range and any other stated or intervening value in that stated range, is encompassed within the methods, systems and computer-readable media. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the methods, systems and computer-readable media, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the methods, systems and computer-readable media. Certain ranges are presented herein with numerical values being preceded by the term “about.” The term “about” is used herein to provide literal support for the exact number that it precedes, as well as a number that is near to or approximately the number that the term precedes. In determining whether a number is near to or approximately a specifically recited number, the near or approximating unrecited number may be a number which, in the context in which it is presented, provides the substantial equivalent of the specifically recited number. Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the methods, systems and computer-readable media belong. Although any methods, systems and computer- readable media similar or equivalent to those described herein can also be used in the practice or testing of the methods, systems and computer-readable media, representative illustrative methods, systems and computer-readable media are now described. All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the materials and / or methods in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present methods, systems and computer-readable media are not entitled to antedate Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 such publication, as the date of publication provided may be different from the actual publication date which may need to be independently confirmed. It is noted that, as used herein and in the appended claims, the singular forms “a”, “an”, and “the” include plural referents unless the context clearly dictates otherwise. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements, or use of a “negative” limitation. It is appreciated that certain features of the methods, systems and computer-readable media, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the methods, systems and computer-readable media, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. All combinations of the embodiments are specifically embraced by the present disclosure and are disclosed herein just as if each and every combination was individually and explicitly disclosed, to the extent that such combinations embrace operable processes and / or compositions. In addition, all sub-combinations listed in the embodiments describing such variables are also specifically embraced by the present methods, systems and computer-readable media and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein. As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present methods. Any recited method can be carried out in the order of events recited or in any other order that is logically possible. NATURAL LANGUAGE PROCESSING FOR DETERMINING DIAGNOSTIC CODES Natural Language Processing (NLP) is a branch of artificial intelligence (AI) that focuses on extracting text-based data. Recent advancements in transformers and large language models have proved their ability to interpret language for various applications. Traditionally, NLP studies for clinical notes are done in a task-specific manner. These models are trained to predict fixed, predetermined categories, and require additional labeled data and training when introducing new categories into the task. This poses a significant burden for medical applications as ground truth creation requires labeling of vast amounts of EHR data by specialized experts. Therefore, models that can transfer learned representations to new tasks are highly desirable. In fields involving imaging data, efforts have leveraged natural language supervision to learn visual concepts. The resultant representations can be applied to new tasks with little additional training. Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 The present disclosure provides an AI model trained to achieve semantic representation of clinical notes through pairs of diagnosis codes and their attendant text descriptions. The code- description pairings are available in clinical databases, for example ICD-9, ICD-10, and / or any other relevant clinical database. Given the vast number of diagnostic codes in some of these databases (e.g., > 70,000 codes in ICD-10), AI-based automation can greatly alleviate provider documentation burden, restore therapeutic patient-doctor relationships, and optimize allocation of scarce human resources towards providing better healthcare. Methods of the present disclosure include inputting a clinical note into a machine learning model. A clinical note can be input directly into the machine learning model, or such a note can be input into a EHR system, which automatically inputs the clinical note into a machine learning model. Such interfacing can be local (e.g., on the same device) or remote (e.g., connected via a network). Remote connections can be a local network, where a machine learning model is located on a computing device (e.g., server) within the same building, institution, or facility. Alternatively (or in addition to local networks), remote connections can be across an external network or cloud-based, such that the clinical notes are transmitted to a computing device not controlled by or located within the facility or institution from which the clinical notes originate. In various embodiments, the machine learning model is capable of performing NLP to identify a diagnostic code. In certain embodiments, the model is based on one or more text encoders, such as a Bidirectional Encoder Representations from Transformers (BERT) model. Additional encoders that can be used in combination with or instead of BERT are Index-Based Encoding, Bag of Words, Term Frequency-Inverse Document Frequency (TF-IDF) Encoding, Word2Vector Encoding, Multilayer Perceptron (MLP), and / or any other relevant methodology for performing NLP. Figure 1A provides an example of a model architecture 100 for various embodiments, where a first text encoder 102 creates latent representations (^^…^^) 104 for diagnostic code descriptions, while a second text encoder 106 creates latent representations (^^…^^) 108 for clinical notes. As illustrated in Figure 1B, encoders can comprise multiple submodules, such as a transformer (e.g., BERT) and an MLP, such as illustrated. It should be noted that the examples illustrated in Figures 1A-1B are merely for illustrative and demonstrative purposes and are not meant to be limiting on the scope of the all embodiments pondered herein. The specific settings, parameters, and / or hyperparameters for a model or submodule can be based on power and / or performance of a model and / or the computing device or media holding or executing the model. Such settings can include model width, parameters used, embedding dimensions, number of heads, number of layers, connection types, and type of activation (e.g., GeLu activation). For example, various models can use a transformer model (e.g., BERT) with 4 layers, 8 layers, 12 layers, 16 layers, 20 layers, 24 layers, or more. Additionally, models can use any number of attention heads such as 4 attention heads, 8 attention heads, 12 attention heads, Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 16 attention heads, 20 attention heads, 24 attention heads, or more. Model widths can be selected from 96, 192, 384, 768, 1536, or more. Once a clinical note is input into a machine learning model, the machine learning model can identify a diagnostic code and / or diagnostic code description best matching the clinical note. Such a diagnostic code can be obtained (or received) by an individual inputting the note or another individual affiliated with the same facility or institution as the inputter. Such individuals can include medical practitioners (e.g., physicians, doctors, nurses, surgeons, orderlies, technicians, and / or any other practitioner) or administrators (e.g., billing associates, medical coders, facility managers, business managers, accountants, and / or any other administrator). Once a diagnostic code is obtained, certain embodiments may provide a treatment (e.g., administration of a therapeutic agent), treatment plan, treatment regimen, or other method of alleviating the condition associated with the diagnosis code. In some instances, providing may be performing a treatment to a subject (e.g., surgery, respiratory treatment, physical therapy, providing a medicament). In some scenarios, providing may be prescribing a treatment or medication and directing the subject on the appropriate utilization (e.g., when, how, and how long to take a medication). In certain cases, providing may be directing a caretaker or assistant to provide a treatment. Such caretakers or assistants may be a nurse, an orderly, a technician, a candy striper, a therapist, and / or other person who can provide care. For example, a returned diagnostic code of “M06.9–Rheumatoid Arthritis” may lead to the administration to the subject of an effective amount of a therapeutic agent to treat the rheumatoid arthritis, such as a non- steroidal anti-inflammatory drug (e.g., aspirin, acetaminophen, ibuprofen, etc.), a steroid, an opioid, etc. In certain situations, a history of diagnostic codes may lead to a different treatment decision, such as “R781–Finding of opiate in blood” may contraindicate treatment with an opioid. As used herein, the terms "treatment," "treating," and the like, refer to obtaining a desired pharmacologic and / or physiologic effect. The effect may be therapeutic in terms of a partial or complete cure for a disease and / or one or more symptoms attributable to the disease. "Treatment," as used herein, covers any treatment of a disease in a mammal, including in a human, and includes: (a) inhibiting the disease, i.e., arresting its development; and (b) relieving the disease, i.e., causing regression of the disease. A “therapeutically effective amount” or “efficacious amount” refers to the amount of a therapeutic agent that, when administered to a mammal or other subject for treating a disease, is sufficient to affect such treatment for the disease. The “therapeutically effective amount” will vary depending on the therapeutic agent, the disease and its severity and the age, weight, etc., of the subject to be treated. An effective amount may be administered in one or more administrations. In other embodiments, obtained diagnostic codes may lead to an individual in a business office, billing department, etc. to generate a bill. Such a bill may be generated for 1) a subject to Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 whom the diagnostic code relates or 2) a health insurance company providing insurance to the subject. Certain bills may be altered for appropriate maximum approved maximum approved cost values as determined by agreements with the health insurance company. Such bills may further be transmitted to a recipient. Such transmission can be electronic and / or physical. Electronic means include uploading a document into a billing platform, email, text message (e.g., SMS / MMS), and / or any other electronic method for communication. Physical means can include a mail and / or courier service, such as U.S. Postal Service (USPS), Canada Post, DHL, FedEx, UPS, Purolator, any other national or international mail or courier service, and combinations thereof. As noted, such code determination can automatically be determined from clinical note entry into EHR. As such, certain embodiments can be integrated into an EHR system. EHR systems can be used by a medical practice, facility, and / or other institution to manage all aspects of patient care, including medical history, billing, insurance compliance, and communications with the parties (e.g., insurance, patient, doctor, accounting / billing departments, etc.). As can be expected from the description herein, AI models as described herein can be integrated into EHR systems to facilitate the methods and processes as described herein. MODEL TRAINING, TESTING, AND OPTIMIZING AI models used in various methods, systems, and methods, systems and computer- readable media can be trained based on diagnostic codes and diagnostic code descriptions. As noted previously, such codes are widely available in clinical code databases, including ICD-9, ICD-10, and / or any other database for medical coding. Data may be preprocessed for training, as some models may need training data in a particular format or pattern. For example, some models may utilize prompt-based training–for example, “sepsis” alone may be insufficient for training, but a prompt of “clinical note of sepsis” may be sufficient for training. This prompt-based training is illustrated in the example in Figure 1A, where the diagnostic code descriptions have been converted to prompts. Various models may further tokenize inputs. Such tokenization may create a sequence of any relevant number of tokens, such as 32 tokens, 64 tokens, 128 tokens, 256 tokens, 512 tokens, 1024 tokens, 2048 tokens, or more. In some situations, zero padding may be used to increase the number of tokens for one input to the number of tokens used in the model (e.g., if tokenizing an input produces only 400 tokens, zero padding may be used to increase the tokens to 512 tokens). Tokens can be bracketed, as necessary, for a particular model, such as using classification [CLS] and separators [SEP]. Once inputs are preprocessed, as described above, models, including respective text encoders, implemented in various embodiments may be trained with the preprocessed data. In certain embodiments, a first text encoder 102 may use the preprocessed diagnostic code descriptions (or prompt-converted versions thereof) to obtain a latent representation of the Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 diagnostic code descriptions 104, while a second text encoder 106 may use preprocessed clinical notes to obtain a latent representation of the clinical notes 108. Some embodiments produce a similarity matrix 110 based on latent representations 104, 108. Various methods may be used to determine similarities between pairs of representations. Similarity determinations can be made based on physical similarity, string similarity, kernel functions, and / or any other relevant methodology for determining such similarities. Various methods, systems, and computer-readable media utilize a cosine similarity to generate a similarity matrix 110. Cosine similarity can be calculated via various methodologies known in the art. Certain methods, systems, and computer-readable media utilize one or more loss functions and a target matrix representing ground truth to train the model. In some instances, loss functions are unidirectional or bidirectional regarding the training inputs (e.g., diagnostic code descriptions and clinical notes). For example, loss functions can determine 1) loss when mapping a clinical note to a diagnosis code and / or 2) loss when mapping a diagnosis code description to a clinical note. A model, as implemented into methods, systems, and / or computer- readable media described herein, may minimize the sum of the one or more loss functions during training. Additionally, loss functions may utilize hyperparameters, such as temperature, ^, within the functionality to affect final probabilities. Learning rate optimization may utilize any known and relevant means, such as Adam, AdamW, LARS, RMSprop. AdaGrad, SGD, and / or any other optimizer. Additionally, learning rates and weight decay may be selected for efficiency in training and / or power in the resultant model. Turning to Figure 2, model testing can utilize inputs not previously used in the model, such as held-back or reserved inputs. Such inputs can be used to determine the efficacy and / or accuracy of a model used in various methods, systems, and / or computer-readable media. Testing can be performed following training a model or on an obtained model (e.g., a model that was trained by another entity, including a person, a corporation, an institution, etc.). In testing, the reserved inputs can include one or both of diagnostic codes 202 (and / or diagnostic code descriptions or prompts of the diagnostic code descriptions) and clinical notes 204. Such testing can include transforming the diagnostic code inputs into a new set of representations 206, then sequentially assessing each clinical note representation 208 for the similarity to the diagnostic representations 206. Such assessment can use a similarity determination, such as used in model training. For example, a cosine similarity. Training optimization can further be performed to increase accuracy. Optimization can be performed following training a model, testing a model, and / or on an obtained model, such as described prior. Such optimization can iteratively alter one or more parameters or hyperparameters of training to improve the accuracy and / or efficacy of a model. Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 Such optimization may be based on different initialization methods, such as random weight initialization, general initialization, and medical domain initialization. However, such methods may result in different levels of accuracy. Furthermore, hyperparameters, such as temperature, ^, weight decay, and / or learning rates may be altered, and the model retrained. Such parameters and / or hyperparameters can be of any relevant rate. For example, temperature, ^, can be elected from 0.001, 0.01, 0.1, 1, 10, 100, 1000, etc. to improve accuracy, while learning rate and / or weight decay can be selected from 1e-7, 1e-6, 1e-5, 1e-4, 1e-3, 0.01, 0.1, 1, etc. In certain instances, parameters and / or hyperparameters can be altered incrementally (e.g., 0.1 to 1 to 10, etc.) or in a more stochastic fashion (e.g., 0.1 to 10 to 1, etc.). Furthermore, in some situations, only one parameter and / or hyperparameter is altered at a time, while other situations may alter multiple parameters and / or hyperparameters at a time (e.g., altering both temperature and weight decay). Altering a parameter and / or hyperparameter can be iterated until reaching a predefined threshold (e.g., a specified number of iterations, a minimum threshold of accuracy and / or efficacy, etc.) or until a maximum accuracy and / or efficacy is identified. COMPUTER-READABLEMEDIA ANDDEVICESAlso provided by the present disclosure are computer-readable media and devices. For example, one or more steps of any of the methods of the present disclosure (e.g., methods of determining diagnostic codes, methods of training and / or optimizing a machine learning model, training a text encoder, minimizing a loss function, determining similarities, and / or the like) may be computer-implemented. By “computer-implemented” generally means at least one step of the method is implemented using one or more processors and one or more non-transitory computer- readable media. The computer-implemented methods of the present disclosure may further comprise one or more steps that are not computer-implemented, e.g., obtaining a blood microsample from a subject, performing one or more “wet lab” steps on the sample (e.g., manual extraction of proteins, lipids and / or metabolites from the sample), and / or the like. By way of example, provided are computer-implemented methods of machine learning enabled determination of diagnostic codes, the methods being implemented using one or more processors and one or more non-transitory computer-readable media comprising instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to: receive clinical notes and / or diagnostic codes, train one or more text encoders based on these inputs, and determine a diagnostic code based on a subsequently input clinical note. Such determinations can use an artificial intelligence model and determine cosine similarities between diagnostic codes and clinical notes. As will be appreciated with the benefit of the present disclosure, any of the methods of the present disclosure amenable to computer- implementation may be implemented in a similar manner employing one or more processors and one or more non-transitory computer-readable media comprising instructions stored thereon, Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 which when executed by the one or more processors, cause the one or more processors to perform one or more steps of such methods. Also provided are devices for performing any of the methods of the present disclosure. By way of example, provided are devices for machine learning enabled determination a diagnostic code, such devices comprising one or more processors and one or more computer- readable media. The one or more computer-readable media comprise instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to: receive clinical notes and / or diagnostic codes, train one or more text encoders based on these inputs, and determine a diagnostic code based on a subsequently input clinical note. Such determinations can use an artificial intelligence model and determine cosine similarities between diagnostic codes and clinical notes. As will be appreciated with the benefit of the present disclosure, any of the methods of the present disclosure amenable to computer-implementation may be implemented in a similar manner by a device of the present disclosure comprising one or more processors and one or more non-transitory computer-readable media comprising instructions stored thereon, which when executed by the one or more processors, cause the one or more processors to perform one or more steps of such methods. A variety of processor-based devices may be employed to implement the embodiments of the present disclosure. Such devices may include device architecture wherein the components of the device are in electrical communication with each other using a bus. Device architecture can include a processing unit (CPU or processor), as well as a cache, that are variously coupled to the device bus. The bus couples various device components including device memory, (e.g., read only memory (ROM) and random access memory (RAM), to the processor. Device architecture can include a cache of high-speed memory connected directly with, in close proximity to, or integrated as part of the processor. Device architecture can copy data from the memory and / or the storage device to the cache for quick access by the processor. In this way, the cache can provide a performance boost that avoids processor delays while waiting for data. These and other modules can control or be configured to control the processor to perform various actions. Other device memory may be available for use as well. Memory can include multiple different types of memory with different performance characteristics. Processor can include any general purpose processor and a hardware module or software module, such as first, second and third modules stored in the storage device, configured to control the processor as well as a special-purpose processor where software instructions are incorporated into the actual processor design. The processor may essentially be a completely self-contained computing device, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric. To enable user interaction with the computing device architecture, an input device can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 screen for gesture or graphical input, keyboard, mouse, motion input, speech and so forth. An output device can also be one or more of a number of output mechanisms. In some instances, multimodal devices can enable a user to provide multiple types of input to communicate with the computing device architecture. A communications interface can generally govern and manage the user input and device output. There is no restriction on operating on any particular hardware arrangement and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed. The storage device is typically a non-volatile memory and can be a hard disk or other types of computer-readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs), read only memory (ROM), and hybrids thereof. The storage device can include software modules for controlling the processor. Other hardware or software modules are contemplated. The storage device can be connected to the device bus. In one aspect, a hardware module that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as the processor, bus, output device, and so forth, to carry out various functions of the disclosed technology. Embodiments within the scope of the present disclosure may also include tangible and / or non-transitory computer-readable storage media or devices for carrying or having computer- executable instructions or data structures stored thereon. Such tangible computer-readable storage devices can be any available device that can be accessed by a general purpose or special purpose computer, including the functional design of any special purpose processor as described above. By way of example, and not limitation, such tangible computer-readable devices can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other device which can be used to carry or store desired program code in the form of computer-executable instructions, data structures, or processor chip design. When information or instructions are provided via a network or another communications connection (either hardwired, wireless, or combination thereof) to a computer, the computer properly views the connection as a computer-readable medium. Thus, any such connection is properly termed a computer-readable medium. Combinations of the above should also be included within the scope of the computer-readable storage devices. Computer-executable instructions include, for example, instructions and data which cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Computer-executable instructions also include program modules that are executed by computers in stand-alone or network environments. Generally, program modules include routines, programs, components, data structures, objects, and the functions inherent in the design of special-purpose processors, etc. that perform tasks or Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 implement abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of the program code means for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps. Other embodiments of the disclosure may be practiced in network computing environments with many types of computer device configurations, including personal computers, hand-held devices, multi-processor devices, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. Embodiments may also be practiced in distributed computing environments where tasks are performed by local and remote processing devices that are linked (either by hardwired links, wireless links, or by a combination thereof) through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices. For purposes of completeness, various aspects of the present disclosure are set out in the following numbered clauses. Clause 1. A method for training a machine learning model for determining diagnostic code, comprising: training a first text encoder based on a plurality of diagnosis code descriptions to obtain a latent representation for the diagnosis code descriptions; training a second text encoder based on a plurality of clinical notes to obtain a latent representation for the clinical notes; producing a similarity matrix between the latent representation for the diagnosis code descriptions and the latent representation for the clinical notes; and minimizing a loss function representing entropy loss between the similarity matrix and a ground truth of the matrix. Clause 2. The method of Clause 1, wherein at least one of the first encoder and the second encoder comprises a transformer-based Bidirectional Encoder Representations from Transformers (BERT) and a multilayer perceptron (MLP). Clause 3. The method of Clause 1 or 2, wherein the first encoder and the second encoder each comprise a transformer-based Bidirectional Encoder Representations from Transformers (BERT) and a multilayer perceptron (MLP). Clause 4. The method of Clause 2 or 3, wherein the BERT comprises a 110M-parameter, 12-layer and 768-wide model with 12 attention heads. Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 Clause 5. The method of any one of clauses 2-4, wherein the MLP comprises two linear layers with GeLU activation and a residual connection. Clause 6. The method of any one of Clauses 1-5, wherein each diagnosis code description in the plurality of diagnosis code descriptions is tokenized. Clause 7. The method of Clause 6, wherein each diagnosis code is represented as a sequence of up to 512 tokens. Clause 8. The method of any one of Clauses 1-7, wherein each clinical note in the plurality of clinical notes is tokenized. Clause 9. The method of Clause 9, wherein each clinical note is represented as a sequence of up to 512 tokens. Clause 10. The method of any one of Clauses 1-9 wherein the similarity matrix is based on a cosine similarity between the latent representation for the diagnosis code descriptions and the latent representation for the clinical notes. Clause 11. The method of Clause 10, wherein the cosine similarity is determined by: ^ ^^· ^ <^^, ^^> = where ^^represents clinical note code description. Clause 12. The method of any one of Clauses 1-11, wherein the loss function comprises a first loss function (^^→^) representing a loss when mapping a clinical note to a diagnosis code and a second loss function (^^→^) representing a loss when mapping a diagnosis code to a clinical note. Clause 13. The method of Clause 12, wherein the first and second loss functions are represented as: ^^^ (< ^^, ^^> / ^) ^^→^= − ^^^ ^^and exp#< ^ , ^ > / ^% ^ − ^^^$ $^$ Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 where ^ denotes the temperature hyperparameter; ^^and ^$are ground truth values; and <> denotes the cosine similarity. Clause 14. A method for optimizing a machine learning model, comprising: a) obtaining a machine learning model comprised of a first text encoder trained to generate a representation of a diagnostic code description and the second text encoder is trained to generate a representation of a clinical note; b) transforming a plurality of diagnostic code descriptions using the first trained text encoder to generate a set of representations of the plurality of diagnostic code descriptions; c) transforming a clinical note using a second trained text encoder to generate a representation of the clinical note; d) determining a cosine similarity between the representation of the clinical note and each representation in the set of representations of the plurality of diagnostic code descriptions; and e) selecting a diagnostic code description wherein the diagnostic code description is associated with representation of the diagnostic code description having a greatest similarity to the representation of the clinical note. Clause 15. The method of Clause 14, further comprising: f) altering a hyperparameter associated with one or more of the first trained text encoder and the second trained text encoder; and g) iterating through steps b)-f) until an accuracy reaches a predefined threshold. Clause 16. The method of Clause 14, further comprising: f) altering a hyperparameter associated with one or more of the first trained text encoder and the second trained text encoder; and g) iterating through steps b)-f) until an accuracy reaches a maximum. Clause 17. The method of any one of Clauses 14-16, wherein the obtained machine learning model is trained according to any one of Clauses 1-13. Clause 18. A method of machine learning enabled diagnostic code determination, comprising: inputting a clinical note into a machine learning model, wherein the machine learning model is comprised of a first text encoder trained to generate a representation of a diagnostic code description and the second text encoder is trained to generate a representation of a clinical note; Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 obtaining a diagnostic code from the machine learning model. Clause 19. The method of Clause 18, wherein inputting the clinical comprises inputting the clinical note into an electronic health record, wherein the electronic health record is input into the machine learning model. Clause 20. The method of Clause 18 or 19, wherein the machine learning model was trained in accordance with a method described by at least one of the methods described in Clauses 1-13. Clause 21. The method of any one of Clauses 18-20, wherein the machine learning model was optimized in accordance with a method described by at least one of the methods described in Clauses 14-17. Clause 22. The method of any one of Clauses 18-21, further comprising providing a treatment based on the obtained diagnostic code. Clause 23. The method of Clause 22, wherein the treatment is not contraindicated by another diagnostic code. Clause 24. The method of any one of Clause 18-23, further comprising generating a bill based on the obtained diagnostic code. Clause 25. The method of Clause 24, further comprising transmitting the bill to a recipient. Clause 26. A non-transitory machine-readable medium comprising instructions that when executed by a processor direct the processor to execute the instructions, the instructions comprising the methods of any one of Clauses 1-25. Clause 27. A system for training a machine learning model for determining diagnostic code, comprising a processor and a memory, wherein the memory comprises instructions that when executed direct the processor to perform the method of any one of Clauses 1-25. The following examples are offered by way of illustration and not by way of limitation. Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 EXPERIMENTAL Example 1 – Clinical Dataset Electronic health record (EHR) data from the University of California San Francisco (UCSF), a tertiary medical center, were used to identify patients for this study. All patient clinical notes since UCSF’s transition to EHR in June 2012, until March 2023, were included for analyses. Data were accessed from UCSF’s de-identified clinical data warehouse, which is based on the Epic Caboodle Data Warehouse and updated monthly.18 types of protected health information (PHI) were removed according to the Safe Harbor Method. Examples of training data are shown in Tables 1A-1B. In total, 89,957 data pairs of clinical notes and their billed diagnosis codes & text descriptions, spanning 100 distinct clinical diagnoses, are used. The diagnosis codes are in ICD-10-CM format and mapped one-to-one with their corresponding textual descriptions. The dataset was randomly divided into training (90%), testing (5%), and evaluation (5%) sets, each containing 80959, 4499, and 4499 data pairs respectively. This study was approved by the Institutional Review Board of UCSF and adhered to the tenets of the Declaration of Helsinki. Pairs of clinical notes and their billed diagnosis codes with text descriptions were extracted from the de-identified electronic health records (EHR) of the University of California, San Francisco (UCSF). A random 100 of the top 300 most common diagnoses were included for analysis, totaling 89,957 data pairs. This dataset was collected from de-identified patient encounters at UCSF from May 2012 to March 2023.90% of this dataset (80959 data pairs) were applied toward model training, 5% (4499 data pairs) towards model testing and 5% (4499 data pairs) towards model evaluation. The distributions of the clinical note length and the diagnosis code description length are shown in Figure 3. The length of clinical notes has a mean of 1906 words and a standard deviation of 1345 words. The length of textual diagnosis descriptions has a mean of 14 words and a standard deviation of 7 words. Example 2 – Data Preprocessing The ClinicalCLIP model (an exemplary embodiment) expects inputs to be of uniform length. A BERT model operates with a max sequence length of 512 tokens, on a lower-cased byte pair encoding representation of the text with a 28,996 vocab size. Thus, all texts were converted to lower cases and randomly sampled 512 tokens from each clinical note and diagnosis description. Importantly, the order of the sampled tokens is reserved. For entries shorter than 512 tokens (which include all the diagnosis code descriptions and some clinical notes), zero paddings were added to expand their length to 512 tokens. Text sequences were bracketed with classification ([CLS]) and separator ([SEP]) tokens. All CLIP models require prompt-based training, whereby input into the encoder needs to be in the format of a prompt, such as “clinical note of sepsis”; the original diagnosis description Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 of “sepsis” by itself cannot serve as an input. Thus, the prefix “clinical note of” is added to the beginning of all diagnosis code descriptions. Representation of the entire text of clinical note or diagnosis description is wholly contained within the first token of the output of BERT, thus only the first token is used in our model. This token serves as input into MLP, which outputs a 256-dimensional representation of each of ^^…^^and ^^…^^for contrastive learning. Example 3 – Model Training Figures 1A-1B provide an overview of the model and its training. As noted above, clinical notes and diagnosis code descriptions were preprocessed into uniform-sized vectors (512 tokens each). These preprocessed vectors were used to train corresponding text encoders. The clinical note encoder learns a latent representation for the clinical notes (^^…^^), while the diagnosis description encoder learns a latent representation for the diagnosis code descriptions (^^…^^). Figure 1B illustrates the architecture used in each encoder, where each encoder includes two submodules: a 12-attention layer, a transformer-based BERT model, and a multilayer perceptron (MLP) consisting of two linear layers with GeLU activation and a residual connection. For the transformer-based BERT in the encoders, a 110M-parameter, 12-layer and 768-wide model with 12 attention heads was used. A cosine similarity between ^^and ^^was calculated to produce a similarity matrix. Cosine similarities were calculated using Equation 1: < ^ , ^^&·^&^ ^> = ||^&|| ||^&|| (1) A target matrix providing a 2: (2) Loss functions can be one or both of Equations 3a and 3b: ^^→^= − ^^^'() (*^&,^&+ / ^)- ' ^^(3a) Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 In the above Equations 3a and 3b, ^ denotes the temperature hyperparameter; ^^and ^$are target matrix ground truth values; and <> denotes the cosine similarity. The first loss function (Equation 3a; ^^→^) is used to calculate the loss when mapping clinical note to diagnosis code; the second function (Equation 3b; ^^→^) calculates the loss when mapping diagnosis code description to clinical note. The two loss functions represent cross entropy loss and the model training aims to minimize their sum. In ClinicalCLIP, the natural language-supervised learning scheme is agnostic to the implementation details. The initialization of the parameters of the BERT model can be based on the downstream task. In this exemplary embodiment, the weights from ClinicalBERT

[0022] achieved the best performance. This was not surprising given the similarities of our training data (both are de-identified healthcare care). During model training, AdamW was applied with a cosine learning rate of 1e-5 and a weight decay of 1e-6. Example 4 – Model Testing Figure 2 illustrates the pipeline for testing ClinicalCLIP. There is 1-to-1 mapping of diagnosis codes to their corresponding textual descriptors, accessible through a look-up table. In this dataset, 100 unique diagnosis codes were randomly selected out from the top 300 most common diagnosis codes in the database, resulting in 100 generated prompts (e.g. diagnosis code of sepsis), which were transformed by trained ClinicalCLIP diagnosis code encoders into 100 representations (^^, ^1… ^^22). During model testing, each input clinical note is transformed by trained ClinicalCLIP note encoders into a representation (^3). ClinicalCLIP then computes the cosine similarities between ^3and ^^, ^1… ^^22, and outputs the prompt with the highest similarity, as well as its corresponding diagnosis code. In order to evaluate the ability of ClinicalCLIP to reduce provider documentation burden, its performance in predicting diagnosis codes was tested from clinical notes using the pipeline shown in Figure 2. Results showed that ClinicalCLIP significantly outperformed benchmark language models, with a 92.71% accuracy. During ClinicalCLIP model optimization, it was found that proper weight initialization of the encoders before training is important for achieving high performance. Experiments were conducted using three distinct initialization methods: 1. random weight initialization; 2. general initialization (which used weights obtained from pretraining on large-scale corpora such as Wikipedia and BookCorpus); and 3. medical domain initialization, which used weight initialization from ClinicalBERT (a competitive large language model-based classifier trained for clinical note- related tasks). The results are shown in Table 2. ClinicalCLIP with medical-domain initialization achieved the highest level of accuracy at 92.05% in predicting diagnosis codes from clinical notes. By contrast, the accuracy is significantly lower with random initialization (30.17%). Additional improvements in prediction accuracy were achieved by optimizing model hyperparameters. Other CLIP models have demonstrated that temperature ^ significantly impacts Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 performance. Thus, experiments were conducted to explore the effect of different ^ values, ranging from 0.01 to 100, on model performance. A temperature of 10 achieved the highest accuracy of 92.71%, while temperatures of 0.01, 0.1, 1, and 100 achieved accuracies of 85.74%, 86.19%, 92.05%, and 85.83%, respectively (Table 3). To further characterize the quality of the learned representations, mean squared error (MSE) were calculated to assess the similarity between two textural representations. Higher MSEs indicate a greater difference in the representations. Results showed that clinical notes with the same diagnosis codes (Figure 4 top and bottom left) had low MSE while clinical notes with different diagnosis codes (Figure 4 top and bottom right) had high MSE. These results indicate that ClinicalCLIP effectively learns semantic representations of clinical notes by leveraging natural language supervision of the diagnosis code descriptions. Finally, ClinicalCLIP was benchmarked against a competitive model in clinical note NLP, ClinicalBERT (Figure 5). The ClinicalBERT model consists of a BERT model, a fully connected layer for output and was configured with medical domain weight initialization. After training on the same datasets, ClinicalBERT was able to predict diagnosis codes using clinical notes with a peak accuracy of 50.11%, which is 42.60% lower than the peak accuracy of ClinicalCLIP. References [1] National Academies of Sciences, Engineering, and Medicine. (2019). Taking action against clinician burnout: A systems approach to professional well-being. Washington, DC: The National Academies Press. https: / / doi.org / 10.17226 / 25521 [2] Kong HJ. Managing unstructured big data in healthcare system. Healthc Inform Res 2019; 25 (1): 1–2. [3] Chapman WW, Nadkarni PM, Hirschman L, D’Avolio LW, Savova GK, Uzuner O. Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutions. Journal of the American Medical Informatics Association.2011; 18(5):540– 543. [4] Huang K, Altosaar J, Ranganath R. ClinicalBERT: Modeling clinical notes and predicting hospital readmission. arXiv preprint arXiv:1904.05342 (2019). [5] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. In NeurIPS 2017. [6] Devlin J, Chang MW, Lee K, and Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. In NAACL 2019. [7] Raffel C, Shazeer N, Roberts A, et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res.21, 1, Article 140 (January 2020), 67 pages. [8] Taori R, Gulrajani I, Zhang T, et al., Stanford Alpaca: An Instruction-following LLaMA model. arXiv preprint arXiv:2302.13971 (2023). Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 [9] Fan L, Li L, Ma Z, Lee S, Yu H, Hemphill L. A Bibliometric Review of Large Language Models Research from 2017 to 2023. arXiv preprint arXiv:2304.02020 (2023).

[0010] Khin K, Burckhardt P, and Padman R. A Deep Learning Architecture for De-identification of Patient Notes: Implementation and Evaluation. arXiv preprint arXiv:1810.01570 (2018).

[0011] Zhu H, Paschalidis IC, and Tahmaseb Ai. Clinical Concept Extraction with Contextual Word Embedding. arXiv preprint arXiv:1810.10566 (2018).

[0012] Si Y, Wang J, Xu H, and Roberts K. Enhancing Clinical Concept Extraction with Contextual Embedding. Journal of the American Medical Informatics Association 26.11 (2019): 1297-1304.

[0013] Wu J, Ye X, Mou C, and Dai W, Fineehr: Refine clinical note representations to improve mortality prediction.202311th International Symposium on Digital Forensics and Security (ISDFS). IEEE, 2023.

[0014] Liang JJ, Tsou CH , and Devarakonda MV. Ground truth creation for complex clinical nlp tasks – an iterative vetting approach and lessons learned. AMIA Summits on Translational Science Proceedings 2017: 203.

[0015] Radford A, Kim JW, Hallacy C, et al. Learning transferable visual models from natural language supervision. International conference on machine learning. PMLR, 2021.

[0016] Li LH, Zhang P, Zhang H, et al. Grounded language-image pre-training. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2022.

[0017] Zareian A, Rosa KD, Hu DH, Chang SF. Open-vocabulary object detection using captions. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021.

[0018] Li B, Weinberger KQ, Belongie S, Koltun V, Ranftl R. Language-driven semantic segmentation. In ICLR 2022.

[0019] Xu J, Mello SD, Liu S, Byeon W, Breuel T, Kautz J, Wang X. Groupvit: Semantic segmentation emerges from text supervision. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition.2022.

[0020] Zhang Y, Jiang H, Miura Y, Manning CD, Langlotz CP. Contrastive learning of medical visual representations from paired images and text. Machine Learning for Healthcare Conference. PMLR, 2022.

[0021] Wang X, Xu Z, Tam L, Yang D, Xu D. Self-supervised image-text pre-training with mixed data in chest X-rays. arXiv preprint arXiv:2103.16022 (2021).

[0022] Alsentzer E, Murphy JR, Boag W, Weng WH , Jin D, Naumann T, McDermott MB. Publicly available clinical BERT embeddings. arXiv preprint arXiv:1904.03323 (2019).

[0023] Wachter RM. The digital doctor: hope, hype, and harm at the dawn of medicine's computer age. New York: McGraw-Hill Education, 2015.

[0024] Haug CJ, Drazen JM. Artificial intelligence and machine learning in clinical medicine. New England Journal of Medicine, 388(13):1201–1208, 2023. Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226

[0025] Hochreiter S and Schmidhuber J. Long short-term memory. Neural computation 9.8 (1997): 1735-1780.

[0026] Schuster M and Paliwal KK. “Bidirectional recurrent neural networks”. In IEEE transactions on Signal Processing 45.11 (1997): 2673-2681.

[0027] Peters ME, Neumann M, Iyyer M, Gardner M, Clark C, Lee K, and Zettlemoyer L. “Deep contextualized word representations”. In NAACL 2018.

[0028] Holmes C. The problem list beyond meaningful use Part 2: fixing the problem list. Journal of AHIMA 82.3 (2011): 32-35. Accordingly, the preceding merely illustrates the principles of the present disclosure. It will be appreciated that those skilled in the art will be able to devise various arrangements which, although not explicitly described or shown herein, embody the principles of the invention and are included within its spirit and scope. Furthermore, all examples and conditional language recited herein are principally intended to aid the reader in understanding the principles of the invention and the concepts contributed by the inventors to furthering the art, and are to be construed as being without limitation to such specifically recited examples and conditions. Moreover, all statements herein reciting principles, aspects, and embodiments of the invention as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, it is intended that such equivalents include both currently known equivalents and equivalents developed in the future, i.e., any elements developed that perform the same function, regardless of structure. The scope of the present invention, therefore, is not intended to be limited to the exemplary embodiments shown and described herein.

[0002] Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 TABLE 1A: EXAMPLE CLINICAL NOTE DATA Diagnosis Code: C79.31 Diagnosis Description: Secondary malignant neoplasm of brain and spinal cord e: 0 , n f s ic y al s s Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 either metastasis at this time per primary team, w / the pt continuing ICU care, w / wound I&D planned for tomorrow. PLAN: [ ] no role for emergent XRT today [ ] evaluate al , s t * s s t's e n Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 Lactate normal. Recent ICU stay for septic shock that improved with abx's though all cx's negative then. Remains afebrile though still mildly tachycardic, concerning for - t g nt b H Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 referral received from medical team to clarify potential family needs during this admission. This assessment based on phone conversation with ***** CPS ***** ***** l le *, nt

[0003] Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 TABLE 1A: EXAMPLE CLINICAL NOTE DATA ICD_code Descriptors H91.90 Hearing disorder e Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 S06.0X0A Concussion without loss of consciousness initial encounter H90.5 Sensorineural hearing loss d s Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 L70.0 Acne vulgaris K62.82 Dysplasia of anus

[0004] Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 TABLE 2: PREDICTION ACCURACY WITH DIFFERENT INITIALIZATION METHODS Initialization Method Prediction Accuracy TABLE 3: EFFECT OF 4 TEMPERATURES ON PERFORMANCE Temperature 4 Prediction Accuracy

Claims

Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 WHAT IS CLAIMED IS:

1. A method for training a machine learning model for determining diagnostic code, comprising: training a first text encoder based on a plurality of diagnosis code descriptions to obtain a latent representation for the diagnosis code descriptions; training a second text encoder based on a plurality of clinical notes to obtain a latent representation for the clinical notes; producing a similarity matrix between the latent representation for the diagnosis code descriptions and the latent representation for the clinical notes; and minimizing a loss function representing entropy loss between the similarity matrix and a ground truth of the matrix.

2. The method of claim 1, wherein at least one of the first encoder and the second encoder comprises a transformer-based Bidirectional Encoder Representations from Transformers (BERT) and a multilayer perceptron (MLP).

3. The method of claim 1, wherein the first encoder and the second encoder each comprise a transformer-based Bidirectional Encoder Representations from Transformers (BERT) and a multilayer perceptron (MLP).

4. The method of claim 2, wherein the BERT comprises a 110M-parameter, 12-layer and 768-wide model with 12 attention heads.

5. The method of claim 2, wherein the MLP comprises two linear layers with GeLU activation and a residual connection.

6. The method of claim 1, wherein each diagnosis code description in the plurality of diagnosis code descriptions is tokenized.

7. The method of claim 6, wherein each diagnosis code is represented as a sequence of up to 512 tokens.

8. The method of claim 1, wherein each clinical note in the plurality of clinical notes is tokenized.

9. The method of claim 8, wherein each clinical note is represented as a sequence of up to 512 tokens.Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 10. The method of claim 1, wherein the similarity matrix is based on a cosine similarity between the latent representation for the diagnosis code descriptions and the latent representation for the clinical notes.

11. The method of claim 10, wherein the cosine similarity is determined by: ^^· ^ < ^ , ^ > =^^ ^||^^|| ||^^|| where ^^represents clinical note code description.

12. The method of claim 1, wherein the loss function comprises a first loss function (^^→^) representing a loss when mapping a clinical note to a diagnosis code and a second loss function (^^→^) representing a loss when mapping a diagnosis code to a clinical note.

13. The method of claim 12, wherein the first and second loss functions are represented as: ^^^ (< ^^, ^ > / ^) ^ = −^^→^^^^ ^ ∑^^ ^^^^^^(< ^^, ^^> / ^) andexp#< ^ , ^ ^ = − ^$ $^→^^^ ^$where ^ denotes the truth values; and <>denotes the cosine similarity.

14. A method for optimizing a machine learning model, comprising: a) obtaining a machine learning model comprised of a first text encoder trained to generate a representation of a diagnostic code description and the second text encoder is trained to generate a representation of a clinical note; b) transforming a plurality of diagnostic code descriptions using the first trained text encoder to generate a set of representations of the plurality of diagnostic code descriptions; c) transforming a clinical note using a second trained text encoder to generate a representation of the clinical note; d) determining a cosine similarity between the representation of the clinical note and each representation in the set of representations of the plurality of diagnostic code descriptions; and e) selecting a diagnostic code description wherein the diagnostic code description is associated with representation of the diagnostic code description having a greatest similarity to the representation of the clinical note.Atty. Docket: UCSF-752WO UCSF Ref: SF2023-226 15. The method of claim 14, further comprising: f) altering a hyperparameter associated with one or more of the first trained text encoder and the second trained text encoder; and g) iterating through steps b)-f) until an accuracy reaches a predefined threshold.

16. The method of claim 14, further comprising: f) altering a hyperparameter associated with one or more of the first trained text encoder and the second trained text encoder; and g) iterating through steps b)-f) until an accuracy reaches a maximum.

17. A method of machine learning enabled diagnostic code determination, comprising: inputting a clinical note into a machine learning model, wherein the machine learning model is comprised of a first text encoder trained to generate a representation of a diagnostic code description and the second text encoder is trained to generate a representation of a clinical note; obtaining a diagnostic code from the machine learning model.

18. The method of claim 17, wherein inputting the clinical comprises inputting the clinical note into an electronic health record, wherein the electronic health record is input into the machine learning model.

19. A non-transitory machine-readable medium comprising instructions that when executed by a processor direct the processor to execute the instructions, the instructions comprising the method of claim 1.

20. A system for training a machine learning model for determining diagnostic code, comprising a processor and a memory, wherein the memory comprises instructions that when executed direct the processor to perform the method of claim 1.

Citation Information

Patent Citations

  • Machine learning techniques for cross-domain text classification

    US20230119402A1

  • Multi-task adapters and task similarity for efficient extraction of pathologies from medical reports

    US20230161978A1