Machine learning models for extracting diagnosis, treatment, and key dates.

JP7917534B2Active Publication Date: 2026-09-08FLATIRON HEALTH INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023553939
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-05
Filing Date
2022-03-04
Publication Date
2026-09-08
Estimated Expiration
2042-03-04

Smart Images

  • Figure 0007917534000001
    Figure 0007917534000001
  • Figure 0007917534000002
    Figure 0007917534000002
  • Figure 0007917534000003
    Figure 0007917534000003
Patent Text Reader

Abstract

A model-assisted system for determining dates of patient events may include a processor that may be programmed to access a database storing medical records related to patients, the medical records including unstructured data, analyze the unstructured data to identify a plurality of pieces of information in the medical records related to the patient events, determine a date associated with each of the plurality of pieces, identify a plurality of query periods related to the patient events, and generate, for each of the query periods, a probability of whether the patient event occurred during the query periods based on the plurality of pieces and the associated dates.
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] Background Cross-Reference to Related Applications

[0001] This application claims the benefit of priority from U.S. Provisional Patent Application No. 63 / 157,369, filed on March 5, 2021. The entire contents of the above application are incorporated herein by reference.

[0002] Technical Field

[0002] The present disclosure relates to analyzing medical records, and more specifically, to extracting key dates and other information from unstructured medical data. [[Background Art]]

[0003] Background Information

[0003] In today's healthcare systems, analyzing patient diagnosis, treatment, testing, and other healthcare data across large patient populations can provide useful insights for understanding diseases, developing new forms of therapies and treatments, and evaluating the effectiveness of existing therapies and treatments. Specifically, it can be useful to identify specific dates associated with key events or stages during a patient's diagnosis and / or treatment. For example, it can be beneficial for researchers to identify patients diagnosed with a particular disease and the date the disease was diagnosed or dates associated with specific stages of the disease. It can be further beneficial to extract other dates, such as the date of a later advanced diagnosis (e.g., due to recurrence or an advanced stage of the disease). This can allow researchers to make decisions across large patient populations, for example, for selecting patients to include in clinical trials.

[0004]

[0004] Patient information may be contained within electronic medical records (EMRs). However, in many cases, information regarding the date of diagnosis or other key events is represented in unstructured data (e.g., notes from a doctor's visit, a research assistant's report, or other text-based data), which can make it difficult for a computer to extract the relevant date information. For example, a doctor may include notes about a medical diagnosis in several documents within a patient's medical record without explicitly including the date of diagnosis. Therefore, determining the exact date of the diagnosis (or similar event) based on ambiguous notes may involve piecing together several fragments of information. Furthermore, the enormous amount of data that researchers must examine makes it impossible to manually extract dates or other information. For example, such manual extraction may involve searching through thousands, tens of thousands, hundreds of thousands, or even millions of patient medical records, each of which may contain hundreds of pages of unstructured text. Therefore, it can be extremely time-consuming and difficult, if not impossible, for a human reviewer to process that amount of data. Thus, extracting key dates from patient medical records using conventional techniques, especially for large patient populations, can quickly become an insurmountable task.

[0005]

[0005] In light of these and other shortcomings of current techniques, there is a need for technical solutions to more accurately extract key dates related to patient diagnosis and treatment. Specifically, the solution should favorably enable the extraction of specific dates (e.g., date of initial diagnosis, date of advanced diagnosis, start date of treatment, end date of treatment, etc.) from unstructured data within a large set of patient EMRs. [Overview of the initiative]

[0006] overview

[0006] Embodiments consistent with the present disclosure include systems and methods for determining the dates of patient events. In one embodiment, the model-reliever system may include at least one processor. The processor may be programmed to access a database storing medical records related to a patient, the medical records including unstructured data, to parse the unstructured data to identify multiple fragments of information in the medical records related to a patient event, to determine dates associated with each of the multiple fragments, to identify multiple query periods related to a patient event, and to generate probabilities for each query period that the patient event occurred during a query period based on the multiple fragments and associated dates.

[0007]

[0007] In one embodiment, a method for determining the date of a patient event is disclosed. This method may include accessing a database that stores a medical record related to a patient, the medical record including unstructured data, parsing the unstructured data to identify multiple fragments of information in the medical record related to a patient event, determining the date associated with each of the multiple fragments, identifying multiple query periods related to the patient event, and generating a probability for each query period that the patient event occurred during the query period based on the multiple fragments and the associated dates.

[0008]

[0008] In accordance with other embodiments disclosed, the non-temporary computer-readable storage medium is executed by at least one processor and may store program instructions that execute any of the methods described herein.

[0009] Brief explanation of the drawing

[0009] The accompanying drawings incorporated into this specification and constituting part of this specification serve to illustrate and illustrate the principles of various exemplary embodiments in conjunction with this description. [Brief explanation of the drawing]

[0010] [Figure 1]

[0010] This block diagram shows an exemplary system environment for implementing an embodiment consistent with the present disclosure. [Figure 2]

[0011] This block diagram shows an exemplary medical record for a patient, consistent with the disclosed embodiments. [Figure 3]

[0012] An example of a process for extracting text fragments from unstructured data in patient medical records, consistent with the disclosed embodiments, is provided. [Figure 4]

[0013] An example of a set of documents that can be analyzed to determine patient-related dates consistent with the disclosed embodiments is shown. [Figure 5]

[0014] This is a schematic diagram of an example of a trained model and inputs and outputs to the model, consistent with the disclosed embodiments. [Figure 6]

[0015] This is a schematic diagram of an example of a process for determining query output for a query date, consistent with the disclosed embodiments. [Figure 7]

[0016] This is a schematic diagram of an example of a process for generating probabilities based on a query output vector, consistent with the disclosed embodiments. [Figure 8]

[0017] This flowchart illustrates an example of a process for extracting patient information, consistent with the disclosed embodiments. [Modes for carrying out the invention]

[0011] Detailed explanation

[0018] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numerals are used in the drawings and the following description to refer to the same or similar parts. While several exemplary embodiments are described herein, modifications, adaptations, and other implementations are possible. For example, the illustrated components may be replaced, added, or modified, and the exemplary methods described herein may be modified by replacing, rearranging, removing, or adding steps to the disclosed method. Accordingly, the following detailed description is not limited to the embodiments and examples disclosed. Rather, the appropriate scope is determined by the appended claims.

[0012]

[0019] Embodiments disclosed herein include computer-implemented methods, tangible non-temporary computer-readable media, and systems. Computer-implemented methods may be executed by, for example, at least one processor (e.g., a processing unit) receiving instructions from a non-temporary computer-readable storage medium. Similarly, a system consistent with this disclosure may include at least one processor (e.g., a processing unit) and memory, where memory may be a non-temporary computer-readable storage medium. As used herein, non-temporary computer-readable storage medium refers to any type of physical memory in which information or data readable by at least one processor can be stored. Examples include random-access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage media. Singular terms such as “memory” and “computer-readable storage medium” may further refer to multiple structures, such as multiple memories and / or computer-readable storage media. As used herein, “memory” may include any type of computer-readable storage medium unless otherwise specified. A computer-readable storage medium can store instructions for execution by at least one processor, including instructions for causing a processor to perform steps or stages consistent with the embodiments herein. In addition, one or more computer-readable storage media can be used when implementing a method implemented by a computer. The term “computer-readable storage medium” should be understood to include tangible items that remove carrier and transient signals.

[0013]

[0020] The disclosed systems and methods can automate the analysis of patient or patient population medical records to identify dates related to key events during a patient's diagnosis and treatment. For example, researchers, physicians, clinicians, or other users may be interested in identifying patients diagnosed with a particular disease and their estimated date of diagnosis, or patients receiving a particular medication and dates related to such treatment. This could enable users to efficiently make various decisions about individual patients within a large population based on EMR analysis. For example, researchers can identify patients diagnosed with an advanced stage of disease and the date of the advanced diagnosis, and this information may indicate whether that patient could be included in a clinical trial or other form of cohort study.

[0014]

[0021] Figure 1 shows an example of a system environment 100 for implementing an embodiment consistent with the present disclosure, which will be described in detail below. As shown in Figure 1, the system environment 100 may include several components, including a client device 110, a data source 120, a system 130, and a network 140. It will be understood from this disclosure that the number and arrangement of these components are illustrative and shown for illustrative purposes only. Other arrangements and numbers of components may also be used without departing from the teachings and embodiments of this disclosure.

[0015]

[0022] As shown in FIG. 1, an exemplary system environment 100 may include a system 130. The system 130 is one or more server systems configured to receive information from entities via a network, process the information, store the information, and display / transmit the information to other entities via the network, It may include a database and / or a computing system. Therefore, in some embodiments, a network may facilitate cloud-based sharing, storage, and / or computation. In one embodiment, system 130 may include a processing engine 131 and one or more databases 132, which are illustrated within the region bounded by the dashed line representing system 130. The processing engine 131 may include one or more general-purpose processors, such as a central processing unit (CPU), a graphics processing unit (GPU), etc., and / or one or more special-purpose processors, such as an application specific integrated circuit (ASIC), It may include at least one processing device such as a rewritable gate array (FPGA).

[0016]

[0023] The various components of system environment 100 may include assemblies of hardware, software, and / or firmware, including memory, a central processing unit (CPU), and / or a user interface. Memory may include any type of RAM or ROM implemented by a physical storage medium such as magnetic storage including floppy disks, hard disks, or magnetic tapes, semiconductor storage such as solid-state disks (SSDs) or flash memory, optical storage, or magneto-optical disk storage. The CPU may include one or more processors for processing data according to a set of programmable instructions or software stored in memory. The functions of each processor may be provided by a single dedicated processor or by multiple processors. Furthermore, processors may, without limitation, include digital signal processor (DSP) hardware or any other hardware capable of running software. Optional user interfaces may include any type or combination of input / output devices such as a display monitor, keyboard, and / or mouse. Users of system environment 100 may include any individual who wishes to access and / or analyze patient data. Accordingly, throughout this disclosure, references to “users” in the disclosed embodiments may include any individual, such as physicians, researchers, quality assurance departments in healthcare institutions, and / or any other individual.

[0017]

[0024] Data transmitted and / or exchanged within the system environment 100 may occur over a data interface. As used herein, a data interface may include any boundary through which two or more components of the system environment 100 exchange data. For example, environment 100 can exchange data between software, hardware, databases, devices, people, or any combination thereof. Furthermore, it will be understood that any suitable configuration of software, processors, data storage devices, and networks can be selected to implement the features of the components of the system environment 100 and the related embodiments.

[0018]

[0025] Components of environment 100 (including system 130, client device 110, and data source 120) can communicate with each other or with other components via network 140. Network 140 may include various types of networks, such as the Internet, wired wide area networks (WAN), wired local area networks (LAN), wireless WAN (e.g., WiMAX), wireless LAN (e.g., IEEE 802.11, etc.), mesh networks, mobile / cellular networks, corporate or private data networks, storage area networks, virtual private networks using public networks, short-range wireless communication techniques (e.g., Bluetooth, infrared, etc.), or various other types of network communication. In some embodiments, communication may be performed across two or more of these forms of networks and protocols.

[0019]

[0026] System 130 may be configured to receive and store data transmitted via network 140 from various data sources including data source 120, process the received data, and transmit data and results based on the processing to client device 110. For example, system 130 may be configured to receive patient data from data source 120 or other sources within network 140. In some embodiments, patient data may include medical information stored in the form of one or more medical records. Each medical record may be associated with a specific patient. Data source 120 may be associated with a wide variety of sources of medical information about a patient. For example, data source 120 may include the patient's medical providers such as doctors, nurses, counselors, hospitals, clinics, etc. Data source 120 may also be associated with laboratories such as radiology or other imaging laboratories, hematology laboratories, pathology laboratories, etc. Data source 120 may also be associated with insurance companies or any other patient data source.

[0020]

[0027] System 130 can further communicate with one or more client devices 110 via the network 140. For example, System 130 can provide the client devices 110 with results based on the analysis of information from the data source 120. Client devices 110 may include any entity or device that can receive or transmit data via the network 140. For example, client devices 110 may include a server or computing device such as a desktop or laptop computer. Client devices 110 may also include other devices such as mobile devices, tablets, wearable devices (i.e., smartwatches, implantable devices, fitness trackers, etc.), virtual machines, IoT devices, or various other technologies. In some embodiments, client devices 110 can transmit queries to System 130 via the network 140 for information about one or more patients, such as patients having or associated with specific attributes, dates on which patients are associated with attributes, dates of events associated with patients, or various other information about patients.

[0021]

[0028] In some embodiments, System 130 may be configured to analyze patient medical records (or other forms of unstructured data) to determine dates related to major events during a patient's diagnosis and / or treatment. For example, System 130 may analyze a patient's medical records to determine the date a patient was diagnosed with a particular condition (e.g., a metastasis in a particular area of ​​the body), the date of a test related to that condition, the date a patient tested positive or negative for that condition, the start or end date of a particular treatment or type of treatment (e.g., taking a particular medication), the date of an operation (e.g., surgery), or various other dates. As will be further described below, System 130 may be configured to use one or more machine learning models to perform this analysis. While patient medical records are used as illustrative examples throughout this disclosure, it will be understood that in some embodiments, systems, methods, and / or techniques disclosed may be used similarly to identify patients exhibiting a condition from other forms of records.

[0022]

[0029] To efficiently extract a specific disease or treatment along with key dates and stages related to the disease or treatment, the system 130 may be configured to access a database that stores medical records related to one or more patients. Medical records may refer to any form of document containing data relating to a patient's diagnosis and / or treatment. In some embodiments, a patient may be associated with multiple medical records. For example, a doctor, nurse, physical therapist, pathologist, radiologist, etc., associated with the patient may each generate medical records relating to the patient.

[0023]

[0030] Figure 2 is a block diagram showing an exemplary medical record 200 for a patient consistent with the disclosed embodiment. The medical record 200 is received from data source 120 and may be processed by system 130 to identify dates relevant to the patient as described above. As shown in Figure 2, the record received from data source 120 (or elsewhere) may include either or both unstructured data 210 and structured data 220. Structured data 220 may include quantifiable or classifiable data relating to the patient, such as sex, age, race, weight, vital signs, laboratory results, date of diagnosis, type of diagnosis, stage classification (e.g., claim code), timing of treatment, procedures performed, date of visit, type of clinic, insurer and start date, medication instructions, drug administration, or any other measurable data relating to the patient.

[0024]

[0031] As described above, much of the information used to make decisions about a patient, such as the date of treatment or diagnosis, can be stored in the unstructured data of the patient's medical record. When used herein, unstructured data may include information about the patient that cannot be quantified or easily classified, such as physician's notes or patient's test reports. For example, unstructured data 210 may include information such as a physician's description of the treatment plan, notes describing what happened at the time of the visit, the patient's statements or conversations, subjective assessments or descriptions of the patient's health, radiology reports, pathology reports, test reports, or any other form of information that is not stored in a structured format.

[0025]

[0032] Within the data received from data source 120, each patient may be represented by one or more records generated by one or more healthcare professionals or by the patient. For example, a doctor associated with the patient, a nurse associated with the patient, a physical therapist associated with the patient, etc., may each generate medical records about the patient. In some embodiments, one or more records may be matched and / or stored in the same database. In other embodiments, one or more records may be distributed across multiple databases. In some embodiments, records may be stored and / or given multiple electronic data representations. For example, patient records may be represented as one or more electronic files such as text files, Portable Document Format (PDF) files, Extensible Markup Language (XML) files, etc. If the document is stored as a PDF file, an image, or another file without text, the electronic data representation may also include text related to the document derived from an optical character recognition process. In some embodiments, unstructured data may be captured by an extraction process, while structured data may be entered by healthcare professionals or computed using algorithms.

[0026]

[0033] In some embodiments, unstructured data may include data relating to a particular patient's condition. As used herein, a patient's “condition” may refer to any attribute or characteristic relating to the patient’s health or well-being. For example, a condition may refer to a patient’s diagnosed condition, such as whether the patient has been diagnosed with a particular disease or illness. In some embodiments, a condition may refer to a stage or situation of a particular diagnosed condition. For example, in a patient diagnosed with cancer, the condition may be a specific site of metastasis of the cancer (e.g., whether the patient has been diagnosed with metastases in the brain, liver, bones, lungs, adrenal glands, peritoneum, or various other sites).

[0027]

[0034] System 130 may be configured to extract dates related to various patient conditions. For example, such dates may include the date the condition developed, was observed, tested, diagnosed, or treated, or any other date related to the condition. For example, a date may be the start or end date of a particular treatment or therapy plan for a patient. As used herein, a therapy plan (or treatment plan) may refer to a therapy used to treat a particular disease or condition. For example, a therapy plan may include the administration of a particular drug to the patient (e.g., pharmacotherapy, chemotherapy, etc.), surgical procedures, gene therapy, immunotherapy, changes in the patient's diet, radiotherapy, physiotherapy, counseling or psychotherapy, meditation, sleep therapy, or various other forms of treatment that may be prescribed to the patient. In some embodiments, the dates of interest may vary depending on the specific application, such as the type of cohort selected, the type of study, the type of condition being analyzed, or various other factors.

[0028]

[0035] In many cases, diagnostic information or disease stage related to a specific disease, or treatment information and key dates related to a specific treatment, may be represented within unstructured data and may not be clearly linked to specific dates in the patient's medical records. For example, regarding the diagnosis of a patient with metastatic non-small cell lung cancer (NSCLC), a physician may refer to the diagnosis in multiple notes before and after the date of diagnosis. For instance, before an advanced diagnosis, a physician might include notes such as "Symptoms of NSCLC are present. No evidence of metastasis" and "Possible metastasis to the liver." Various documents after an advanced diagnosis may include phrases such as "Biopsy showed metastasis to the liver" and "Patient with metastatic NSCLC." Therefore, using conventional techniques, it may be difficult for an automated system to verify the exact date of the advanced diagnosis or disease stage based on these notes. Similar problems can arise with treatment information related to the disease. For example, before initiating a specific treatment for the disease, a physician may include notes indicating the diagnosis of the disease and / or notes indicating that various treatment options were discussed with the patient. After treatment is initiated, documents in the patient's records may show the response to the treatment but may not indicate the specific date on which treatment was initiated. Therefore, extracting the exact date of treatment can sometimes be difficult using conventional techniques.

[0029]

[0036] To overcome these and other difficulties, system 130 can extract text fragments related to events of interest. As discussed above, system 130 may be configured to determine dates related to a particular event or condition. Taking the example of the date of diagnosis for a patient with metastatic non-small cell lung cancer above, such fragments may include the fragments: “Symptoms of NSCLC are present. No evidence of metastasis,” “Possible metastasis to the liver,” “Tissue biopsy showed metastasis to the liver,” and “Patient with metastatic NSCLC.” In some embodiments, system 130 can perform a search on a set of documents (which may be a set of patient medical records) for keywords related to a particular event or condition. For example, system 130 can perform a search 510 on one or more unstructured medical record documents to extract fragments related to the date of diagnosis regarding a patient's condition. In the case of a diagnosis of NSCLC, relevant terms may include “NSCLC,” “lung cancer,” “metastatic,” “metastasis,” “metastatic,” “spread,” or various other words, symbols, acronyms, or phrases that may be related to the particular event. These sentences or fragments can be tokenized and represented as a series of tokenized vectors. These tokenized vectors can be processed to generate the corresponding vectorized sentences.

[0030]

[0037] Figure 3 shows an example of a process 300 for extracting text fragments from unstructured data of patient medical records, consistent with the embodiments disclosed. In step 310, the system 130 can perform a search on a set of documents for keywords related to a particular event or condition. In this example, the date of interest may be the date of diagnosis regarding brain metastases. Thus, the keywords may include “brain,” “temporal,” “occipital,” “frontal,” or other terms that may generally be related to the brain. Thus, the system 130 can identify the term 322 in the unstructured text, as shown in step 320. The system 130 can then extract text fragments around the relevant term in the unstructured data. For example, as shown in step 330, fragments 332 around term 322 can be extracted from the unstructured text. The length or structure of the fragments can be specified in various ways. In some embodiments, fragments 332 can be defined based on a predetermined window. For example, a fragment can be defined based on a predetermined number of characters before and after a target term 322 in the text (e.g., 20, 50, 60 characters, or any appropriate number of characters to capture the context of the term's usage). The window can also be defined to take word boundaries into account, for example, by widening or narrowing the window so that it ends at a word boundary or sentence boundary (e.g., based on punctuation, etc.), so that the ends of the fragment do not contain partial words. In some embodiments, the window can be defined based on a predetermined number of words or other variables.

[0031]

[0038] These sentences or fragments can be tokenized and represented as a series of tokenized vectors 340. For example, system 130 can replace term 322 with a tokenized term. This ensures that the patient's condition is expressed using the same terminology in each extracted fragment. For example, both a document containing "brain" and a document containing "cerebrum" may result in extracted fragments containing the term "_brain_" or similar tokens. Using tokens can also improve the performance of machine learning models by reducing feature sparsity, accelerating training time, and enabling model convergence with a more limited set of labeled data. The same tokenization process can be performed on other words within a fragment so that each word is represented by a token. Thus, each fragment can be represented as a vector of values, each value being a tokenized representation of a word contained within the fragment. In some embodiments, the fragment vector may have a predetermined size so that the extracted fragments have a uniform size. As a result, a vector of tokenized fragments 340 can be extracted from a document. In some embodiments, this may involve applying a gated regression unit (GRU) network followed by attention and feedforward layers. An example of a process for generating these fragment vectors is described in U.S. Patent Publication No. 2021 / 0027894 A1 and International Publication No. 2020 / 092316, both assigned to the same applicant as this application. The contents of these applications are incorporated herein by reference in their entirety. The resulting fragment vectors can be input as input vectors to a trained machine learning model.

[0032]

[0039] The system can further associate each input vector (i.e., vectorized sentence) with a date. In some embodiments, such a date may include a date related to the document from which the fragment was extracted. Alternatively, various other dates can be associated with each input vector. For example, if the text in or around a fragment contains a specific date, that date can be used instead of the document date. Figure 4 shows an example of a set of documents that can be analyzed to determine a date related to a patient, consistent with the embodiments disclosed. In this example, a set of documents 410, 420, 430, 440, and 450 can be analyzed to determine the date on which the patient was diagnosed with liver metastases or exhibited symptoms of liver metastases. For illustrative purposes, documents 410, 420, 430, 440, and 450 are shown along a timeline 400 corresponding to the document date of each document. Each of documents 410, 420, 430, 440, and 450 may be associated with a document date. The document date can refer to the date on which the document was created and can be extracted from metadata or other data related to the document. The document date can refer to various other dates related to the document, such as the date the document was updated, revised, received, published, or any other relevant date. For example, document 410 may be related to May 1, 2016, and therefore any fragment extracted from document 410 may be related to this date.

[0033]

[0040] In some embodiments, various other dates can be associated with the fragment. For example, if the text within or around the fragment contains a specific date, that date can be used instead of the document date. As shown in Figure 4, document 450 may contain a fragment that "transferred from February 15, 2017". The date of document 450 is March 13, 2018, but the February 2017 date may better represent the date of interest. Therefore, this fragment may be associated with the February 2017 date instead of the March 2018 date. In some embodiments, the document can be parsed against a specific censorship date 460 (in this case, February 14, 2018), and documents associated with dates after the censorship date can be disregarded. However, using the February 2017 date instead of the March 2018 date may allow the fragment of document 450 to be included as input to the model.

[0034]

[0041] The resulting input vector and associated input dates can be input to a trained machine learning model configured to determine specific diseases and dates associated with an event of interest. Figure 5 is a schematic diagram of a trained model 540 and an example of inputs and outputs to the model, consistent with the disclosed embodiment. The trained model 540 may be trained to receive a vector 510 of tokenized fragments and paired dates 520 as input and output probabilities 550 indicating the probability that a particular event occurred within a given date range. In some embodiments, the trained model 540 may include a feedforward network. A variety of machine learning algorithms may be used, including neural networks, logistic regression, linear regression, regression, random forests, K-nearest neighbors (KNN) models (e.g., as described above), K-means models, decision trees, Cox proportional hazards regression models, naive Bayes models, support vector machine (SVM) models, gradient boosting algorithms, or any other form of machine learning model or algorithm.

[0035]

[0042] The fragment vector 510 can be represented as a vectorized fragment extracted from unstructured data, as described above with respect to Figure 3. For example, the fragment vector 510 can be represented in the form of a vectorized fragment 340. The paired date 520 can represent the date associated with each of the vectorized fragments, as described above with respect to Figure 4. In this example, the fragment vector 512 may be associated with the paired date 522, and another fragment vector 514 may be associated with the paired date 524. The paired date 520 can be represented in vector form, similar to the fragment vector 510. These dates can be represented in various ways. In some embodiments, the representation may include standardized date formats such as YYYY / DD / MM (where "YYYY" represents the year, "MM" represents the month, and "DD" represents the day). In some embodiments, as shown in Figure 3, the date can be represented as a base date or as the number of days away from the base date. For example, the date can be expressed as the number of days before or after the censoring date (e.g., date 460 above), a date relevant to the cohort (e.g., a censoring date for a specific diagnosis), the current date, or any other date that can be used as a reference point.

[0036]

[0043] In some embodiments, one or more query dates 530 can be input to the model in addition to the fragment vector 510 and paired dates 520. Each query date is a point in time, which may represent a point in time for which the trained model 540 should make predictions in relation to it. In some embodiments, queries may be a series of dates spaced evenly apart. For example, queries may be a series of dates spaced seven days apart. In some embodiments, query dates 530 may be automatically generated to encompass the paired dates 520. Alternatively, query dates 530 may depend at least in part on user input. For example, query dates 530 may each be manually defined based on user input. As another example, the user may input an interval for query dates 530, and the system 130 may generate query dates 530 to encompass the paired dates 520 along with the interval defined by the user. In some embodiments, this may involve presenting one or more elements within a user interface (e.g., by a client device 110) and receiving user input through one or more elements of the user interface. While weekly queries are used as an example, various other time periods can be used. For example, the period can be daily, several days, every other week, monthly, yearly, or any other appropriate period.

[0037]

[0044] For each of the query days 530, the trained model 540 can generate a prediction of whether the date of interest occurred within a certain range or period relative to the query day. For example, such a range may include a range encompassing the query day (e.g., within 1 day, 5 days, 10 days, 30 days, or any other appropriate range), a range before the query day, and a range after the query day. Thus, based on the fragment vector 510 and the paired dates 520, the trained model 540 may output multiple probabilities associated with each of the query days 530. For example, query day 530 may include query day 532, which in this example could be a date 14 days before the base date. The trained model 540 may output a set of probabilities 552 indicating whether the date of interest occurred before query day 532, during query day 532 (or within the range of query day 532, such as within 30 days of query day 532), or after query day 532. In the illustrated example, probability 552 could include a 90% probability that the date of interest occurred before query day 532 (or within the range encompassing query day 532), a 9% probability that the date of interest occurred during query day 532 (or within the range encompassing query day 532), and a 1% probability that the date of interest occurred after query day 532 (or within the range encompassing query day 532). As a result, the model can generate a distribution of probabilities over a series of query days, with each query returning a probability of whether the date of interest occurred during that query.

[0038]

[0045] In some embodiments, the fragment vector 510 can be further processed before being input to the trained model 540. For example, to determine the probability of a given query (e.g., to determine the probability 552 for query day 532), the system 130 can analyze the fragment vector 510 within several time windows for the query. In some embodiments, such analysis may include applying one or more aggregate functions to the relevant fragments for each window. Figure 6 is a schematic diagram of an example of a process 600 for determining query outputs for query day 532, consistent with the embodiments disclosed. In the example shown in Figure 6, one or more query outputs 640 can be generated based on query day 532. Generating in this way may include evaluating the fragment vector 510 against a set of time windows 610, as shown. The time windows may be a range of time relative to a base date 612 (corresponding to query day 532). For example, as shown in Figure 6, the time windows could be 365 days or more before the query date, between 365 days and 30 days before the query date, between 30 days and 7 days before the query date, less than 7 days before the query date, within 7 days after the query date, between 7 days and 30 days after the query date, between 30 days and 365 days after the query date, and 365 days or more after the query date. These time windows are given as examples, and any other suitable time window may be used.

[0039]

[0046] For each time window, the vector of fragments related to the input dates contained within that time window can be analyzed according to one or more aggregate functions. For example, aggregate functions may include the sum function 620, the mean function 622, and the LogSumExp function 624. The sum function 620 can represent the vector sum of the input vectors related to the dates within the time window. For example, a matrix M can be generated as all the vectors of the input vectors relating to the model. A logical matrix D can be generated, with elements indicating whether the corresponding input vector is contained within the time window, and thus D * Multiplication by M results in the sum of related vectors.

[0040]

[0047] In the example shown in Figure 6, the sum function 620, mean function 622, and LogSumExp function 624 are executed on time window 614. In this example, time window 614 represents a time window from 7 days after query date 532 to 30 days after query date 532. Therefore, when applied to time window 614, sum function 620 becomes the sum of vectors of any fragments related to dates from 7 days after query date 532 to 30 days after query date 532. In this example, query date 532 is represented as a date 14 days from the base date. Therefore, time window 614 ranges from 21 days from the base date to 44 days from the base date. Referring to the example of fragment vector 510 and paired date 520 shown in Figure 5, the paired dates 522 and 524 are included in the range defined by time window 614, and thus this range includes the sum of fragment vectors 512 and 514. As shown in Figure 6, applying the sum function 620 within the time window 614 involves generating the sum of the fragment vectors 512 and 514, which may result in the query output 630.

[0041]

[0048] Similarly, the mean function 622 can represent the mean of the input vectors relative to the time window. For example, a logical matrix D can be divided by its sum along the second dimension, and therefore D * Multiplication by M yields the mean of the related vectors. Although not shown in Figure 6 for simplicity, the mean function 622 is applied to the fragment vectors 512 and 514 to generate the corresponding query output for the time window 614. The LogSumExp function 624 may be a smoothing approximation to a maximum function (i.e., the "RealSoftMax" or "TrueSoftMax" function) defined as the logarithm of the sum of the exponents of the arguments. For example, the torch.exp() function can be applied to the matrix M, and the resulting matrix can be (D * It can be multiplied by the logical matrix D (where M is the matrix). The torch.log() function can then be applied to the resulting output. The LogSumExp function 624 is also applied to the fragment vectors 512 and 514 to generate the corresponding query output for the time window 614.

[0042]

[0049] The sum function 620, mean function 622, and LogSumExp function 624 are applied to each time window 610, respectively. As a result, multiple aggregates can be generated for each time window for the input vector associated with the time window. This can be repeated for each time window of a given query to generate the output vector of the query. This output vector can be input into a feedforward network (or a similar form of regression neural network) to generate the probability that the date of interest occurred within the query date range (or date range). While the sum function 620, mean function 622, and LogSumExp function 624 are shown as examples, those skilled in the art will understand that they are different aggregate functions (or combinations of aggregate functions can be used).

[0043]

[0050] Figure 7 is a schematic diagram of an example process for generating probabilities based on a query output vector, consistent with the disclosed embodiment. As shown in Figure 7, the query output vector 710 can be input to a model (i.e., a trained model 540) to generate probabilities 552. The query output vector 710 can be generated by applying the sum function 620, the mean function 622, and the LogSumExp function 624 to a vector of fragments 510 for each of the time windows 610. In other words, for each of the time windows 610, the sum function 620, the mean function 622, and the LogSumExp function 624 can be applied to a vector of fragments related to paired dates contained in the time window. For example, the query output vector 710 includes, in particular, query output 532 among several query outputs as shown.

[0044]

[0051] As shown in Figure 7, the query output vector 710 can be input to the trained model 540. For example, the trained model may be a feedforward network (or a similar form of regressive neural network) for generating a probability 552 for query date 532. Thus, the trained model 540 can be trained using a set of training data, which may consist of a set of query output vectors generated based on the fragment vector and associated date, as described above, along with a known date of interest. This training vector can be input to a training algorithm to generate the trained model 540. Thus, the subsequent output vectors can be input to the trained model 540 to generate probabilities for various query dates. In some embodiments, additional layers can be applied. For example, the trained model 540 may output a score 720, which may be an equal / unequal score indicating whether a date matches a diagnostic date or another date of interest. These scores may be real numbers and can be converted to probabilities 552 using the Softmax function 730. The various layers shown in Figure 7 are examples, and other possible layers for manipulating the output of the trained model 540 will be understood by those skilled in the art.

[0045]

[0052] The probability 552 can be expressed in various forms. For example, the probability can be expressed as a value within a range (e.g., 0 to 1, 0 to 100, etc.), as a percentage, as a value within a series of stepwise values ​​(e.g., 0, 1, 2, 3, etc.), or as a text-based representation of the probability (e.g., low probability, high probability, etc.). In some embodiments, the model can also output the probability that a date of interest does not occur within the query dates, which may be the inverse of the other probability. For example, for a given query date, the model may output a 0.98 probability that the date of interest occurs within the query dates, and a 0.02 probability that the date of interest does not occur within the query dates. The above process can be repeated for each query to generate a probability distribution 550 showing the probability that the date of interest occurs across a series of query dates. Various other outputs can be generated, such as the overall probability that a patient is diagnosed with a particular disease, confidence levels associated with the distribution, etc.

[0046]

[0053] In some embodiments, the model may generate probabilities for multiple dates of interest. For example, continuing the NSCLC diagnosis date example above, the model can output, for each query date, the probability that the initial diagnosis of NSCLC occurred within the range of that query date, and that a more advanced diagnosis (e.g., stage 3b or higher, a lower stage with distant metastasis, etc.) occurred within the range of that query date. Thus, multiple feedforward layers can be applied to the output vector of each query, resulting in multiple probabilities being generated.

[0047]

[0054] The disclosure system and methods generally describe specific diseases and dates related to the diagnosis, as well as examples of various states of the disease, but it should be understood that the same or similar process can be applied to dates of other events. For example, such dates may include the start and end dates of a particular treatment or therapy, whether a particular drug was taken along with the dosage and date, and the dates of a particular diagnosis made and related information. Furthermore, various other inputs, such as the type of document, the format of the document, or other metadata of the document, can also be provided to the trained model.

[0048]

[0055] In some embodiments, a model can be trained by identifying a cohort of patients with a specific disease and identifying various stages of the disease and associated diagnosis dates. For example, input to the model may include a set of sentences containing keywords related to advanced NSCLC and extracted from the patient's EHR document. A date can be associated with each sentence using the document's timestamp or, if any, a date explicitly mentioned in the sentence. The sentences can be processed by a GRU network, which can be trained to predict the probability of each diagnosis across a series of points in time, and it can be used to extract whether a patient has been diagnosed with a specific disease and, if so, the diagnosis date.

[0049]

[0056] In some embodiments, the model can be trained for both the diagnosis and treatment of diseases. For example, the input to the model may include one or more sentences with keywords related to the diagnosis of a disease, as well as sentences with keywords related to the treatment of a disease. The input may also include a training dataset of dates related to the diagnosis and treatment of a disease. For example, the training dataset may include the dates when a disease (or different stages of a disease) was diagnosed. Furthermore, the training dataset may include the dates when a particular treatment plan was initiated, as well as other dates related to the treatment (e.g., increasing dosage, changing treatment). The model can be trained using stochastic gradient descent or a similar method for training a model using a set of labeled training data. As a result, the trained model may be configured to extract specific types of diseases and / or treatments, as well as dates related to the diagnosis and treatment of the disease.

[0050]

[0057] Figure 8 is a flowchart showing an example of a process 800 for extracting patient information, consistent with the embodiments disclosed. Process 800 may be executed by at least one processing unit, such as the processing engine 131 described above. Throughout this disclosure, the term “processor” should be understood as an abbreviation for “at least one processor.” In other words, a processor may include one or more structures that perform logical operations, whether such structures are located together, connected, or distributed. In some embodiments, instructions that cause the processor to perform process 800 when executed by the processor may include non-transient computer-readable media. Furthermore, process 800 is not necessarily limited to the steps shown in Figure 8, and any steps or processes of the various embodiments described throughout this disclosure, including those described above with respect to Figures 3, 4, 5, 6, and 7, may also be included within process 800.

[0051]

[0058] In step 810, process 800 includes accessing a database that stores medical records related to the patient. For example, system 130 can access patient medical records from a local database 132 or from an external data source such as data source 120. Medical records may include one or more electronic files such as text files, image files, PDF files, XLM files, and YAML files. One or more medical records may correspond to the medical records 200 discussed above. In some embodiments, medical records may include unstructured data such as unstructured data 210.

[0052]

[0059] In step 820, process 800 includes parsing unstructured data to identify multiple fragments of information within the medical record related to a patient event. For example, this step may include identifying fragment 332 as described above. As described above, a patient event can be any event that may be related to the patient's care. For example, a patient event may include at least one of the date of diagnosis, the date of treatment, or the date of surgery. In some embodiments, process 800 may further include generating a vector of multiple fragments based on the fragments, such as the fragment vector 510.

[0053]

[0060] In step 830, process 800 includes determining the dates associated with each of the multiple fragments. For example, this step may include determining the paired dates 520 as described above. In some embodiments, determining the dates associated with each of the multiple fragments may include identifying dates based on the metadata of the document containing the fragments. For example, as described above with respect to Figure 4, such dates may include the date the document was created, stored, published, or various other dates related to the document. Alternatively, determining the dates associated with each of the multiple fragments may also include identifying dates mentioned within the fragments. For example, the fragments themselves may include a date that better reflects the diagnosis date or other events discussed within the fragments.

[0054]

[0061] In step 840, process 800 includes identifying multiple query periods related to a patient event. For example, this step may include determining a query day 530 as described above. Thus, identifying multiple query periods includes identifying multiple query days, where each query period includes at least one period for each query day. For example, the at least one period for each query day may include a period encompassing the query day, a period before the query day, and a period after the query day. Query days can be identified in various ways. In some embodiments, the multiple query days may be dates spaced one week apart, but any other suitable interval may be used. In some embodiments, the query days, the range of query days, the interval between query days, or various other factors may be specified by user input as described above. Thus, identifying multiple query periods related to a patient event may include receiving at least one user input related to a query period via a user interface. For example, the user interface may be displayed on the client device 110 as described above.

[0055]

[0062] In step 850, process 800 includes generating a probability for each query period that a patient event occurred during the query period, based on multiple fragments and associated dates. For example, this step may include generating a probability of 550 as described above. In some embodiments, this step may include various steps described with respect to Figures 6 and 7. For example, step 850 may include evaluating multiple fragments across multiple time windows for each of multiple query dates, such as time window 610. In some embodiments, evaluating multiple fragments across multiple time windows may include processing fragments associated with dates contained in the time windows for each of the multiple time windows using one or more aggregate functions, as described in more detail above with respect to Figure 6. For example, one or more aggregate functions include at least one of a sum function, a mean function, or a LogSumExp function. Step 850 may further include inputting the results of the multiple functions (i.e., query output vector 710) into a feedforward network or other form of trained machine learning model.

[0056]

[0063] The above description is for illustrative purposes only. The above description is not exhaustive and is not limited to the exact form or embodiment disclosed. Modifications and adaptations will become apparent to those skilled in the art by examining this specification and practicing the disclosed embodiments. In addition, although aspects of the disclosed embodiments have been described as being stored in memory, those skilled in the art will understand that these aspects can also be stored on secondary storage devices, such as hard disks or CD-ROMs, or on other types of computer-readable media such as other forms of RAM or ROM, USB media, DVDs, Blu-rays, 4K Ultra HD Blu-rays, or other optical drive media.

[0057]

[0064] Computer programs based on the descriptions and methods disclosed are within the scope of the skills of an experienced developer. Various programs or program modules can be created using any technique known to those skilled in the art, or can be designed in relation to existing software. For example, a program section or program module may be designed in or by .Net Framework, .Net Compact Framework (and related languages ​​such as Visual Basic, C, etc.), Java, Python, R, C++, Objective-C, HTML, HTML / AJAX combinations, XML, or HTML with included Java applets.

[0058]

[0065] Furthermore, while exemplary embodiments have been described herein, the scope of any and all embodiments having equivalent elements, modifications, omissions, combinations (e.g., aspects across various embodiments), adaptations and / or changes, as understood by those skilled in the art based on this disclosure, is not limited to the examples described herein or examples in the course of this application. Such examples should be construed as non-exclusive. Furthermore, the steps of the disclosed methods can be modified in any way, including rearranging the steps and / or inserting or deleting steps. Accordingly, this specification and the examples are considered solely illustrative, and the true scope and spirit are intended to be shown by the entire scope of the appended claims and their equivalents.

Claims

1. A model-assisted system for determining the date of a patient's event, Accessing a database that stores medical records related to a patient, wherein the medical records include unstructured data. Identifying multiple fragments of information within the medical record related to the patient's events in the unstructured data, To determine the date associated with each of the aforementioned multiple fragments, Identifying multiple query periods related to the patient's events, and For each of the aforementioned query periods, the probability of whether the patient's event occurred during the query period is calculated based on the plurality of fragments and the associated dates. To generate query output for the query period by applying at least one aggregate function to one or more fragments related to one or more dates included within the query period, Based on the query output, the probability of whether the patient's event occurred during the query period is generated. To generate at least one processor programmed to do the following A model-assisted system, including one mentioned above.

2. The model-reliable system according to claim 1, wherein the patient's events include at least one of the diagnosis date, treatment date, or surgery date.

3. The model-reliable system according to claim 1, wherein determining the date associated with each of the plurality of fragments includes identifying the date based on metadata of the document containing the fragment.

4. The model-reliable system according to claim 1, wherein determining the date associated with each of the plurality of fragments includes identifying the date mentioned within the fragment.

5. The model-assisted system according to claim 1, wherein the at least one processor is further programmed to generate a plurality of fragment vectors based on the fragments.

6. The model-reliable system according to claim 1, wherein identifying the plurality of query periods includes identifying a plurality of query days, and the plurality of query periods includes at least one period for each of the query days.

7. The model-reliable system according to claim 6, wherein the at least one period for each of the query days includes a period encompassing the query day, a period prior to the query day, and a period after the query day.

8. The model-reliable system according to claim 6, wherein the plurality of query dates include dates spaced one week apart.

9. The model-assisted system according to claim 6, wherein generating the probabilities includes evaluating the plurality of segments across a plurality of time windows for each of the plurality of query days.

10. The model-assisted system according to claim 9, wherein evaluating the plurality of fragments across a plurality of time windows includes processing the date-related fragments contained in the time windows for each of the plurality of time windows using one or more aggregate functions.

11. The model-assisted system according to claim 10, wherein the one or more aggregation functions include at least one of a sum function, a mean function, or a LogSumExp function.

12. The model-assisted system according to claim 10, wherein generating the aforementioned probabilities includes inputting the results of a plurality of functions into a feedforward network.

13. A method for determining the date of a patient's event, wherein at least one processor, Accessing a database that stores medical records related to a patient, wherein the medical records include unstructured data. Identifying multiple fragments of information within the medical record related to the patient's events in the unstructured data, To determine the date associated with each of the aforementioned multiple fragments, Identifying multiple query periods related to the patient's events, and For each of the aforementioned query periods, the probability of whether the patient's event occurred during the query period is calculated based on the plurality of fragments and the associated dates. To generate query output for the query period by applying at least one aggregate function to one or more fragments related to one or more dates included within the query period, Based on the query output, the probability of whether the patient's event occurred during the query period is generated. To generate Methods that include...

14. The method according to claim 13, wherein the patient's events include at least one of the diagnosis date, treatment date, or surgery date.

15. The method according to claim 13, wherein identifying the plurality of query periods includes identifying a plurality of query days, and the plurality of query periods includes at least one period for each of the query days.

16. The method according to claim 15, wherein the at least one period for each of the query days includes at least one period encompassing the query day, at least one period prior to the query day, and at least one period after the query day.

17. The method according to claim 15, wherein generating the probabilities includes evaluating the plurality of segments across a plurality of time windows for each of the plurality of query days.

18. The method according to claim 17, wherein evaluating the plurality of fragments across a plurality of time windows includes processing the fragments related to dates contained in the time windows for each of the plurality of time windows using a plurality of functions.

19. The method according to claim 18, wherein the plurality of functions include at least one of a sum function, a mean function, or a LogSumExp function.

20. A non-temporary computer-readable medium, which includes instructions causing one or more processors to perform a method for determining the date of a patient's event when executed by one or more processors, wherein the method is: Accessing a database that stores medical records related to a patient, wherein the medical records include unstructured data. Identifying multiple fragments of information within the medical record related to the patient's events in the unstructured data, To determine the date associated with each of the aforementioned multiple fragments, Identifying multiple query periods related to the patient's events, and For each of the aforementioned query periods, the probability of whether the patient's event occurred during the query period is calculated based on the plurality of fragments and the associated dates. To generate query output for the query period by applying at least one aggregate function to one or more fragments related to one or more dates included within the query period, Based on the query output, the probability of whether the patient's event occurred during the query period is generated. To generate Non-temporary computer-readable media, including [specific examples of such media].

21. The model-assisted system according to claim 5, wherein generating the plurality of fragment vectors includes replacing at least one term in the plurality of fragments with a tokenized representation of the at least one term.

22. The aforementioned at least one processor further, For each of the aforementioned multiple query periods, one or more aggregate functions are applied to the aforementioned multiple fragment vectors and associated dates in order to generate multiple query outputs, For each of the aforementioned query periods, the probability of whether the patient's event occurred during the query period is generated based on at least one of the plurality of query outputs. The model-assisted system according to claim 21, which is programmed to perform the following.

23. The one or more aggregate functions mentioned above are: A first function representing the sum of the plurality of fragment vectors for at least one time window, A second function representing the average of the plurality of fragment vectors for at least one time window, and A third function for at least one time window, wherein the third function is different from the first function and the second function. The model-reliever system according to claim 22, comprising at least one of the following.

Citation Information

Patent Citations

  • Text analytics on relational medical data

    US20170329931A1

  • Systems and methods for model-assisted event prediction

    WO2020081495A1