System and method for extracting dates associated with patient status
The system uses machine learning to analyze unstructured patient data, addressing the challenge of identifying metastatic sites and dates, facilitating effective cohort generation and treatment insight.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FLATIRON HEALTH INC
- Filing Date
- 2021-06-11
- Publication Date
- 2026-05-25
AI Technical Summary
Existing healthcare systems face challenges in processing large volumes of unstructured patient data to identify relevant information such as metastatic sites and diagnosis dates, limiting the effectiveness of data analysis and missing rare conditions.
A system utilizing machine learning models to analyze unstructured patient data, identifying patient conditions and associated dates, enabling the extraction of large cohorts of patients with specific metastatic sites and therapy timelines.
Enables the generation of statistically significant patient cohorts for analysis, providing insights into treatment effectiveness and patient responses, overcoming limitations of manual data extraction.
Smart Images

Figure 0007864644000001 
Figure 0007864644000002 
Figure 0007864644000003
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications
[0001] This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 038,397, filed on Jun. 12, 2020. The content of the above application is hereby incorporated by reference in its entirety into this specification.
[0002] Background Technical Field
[0002] The present disclosure relates to identifying patient states within a large set of unstructured data and dates associated with those states, and more particularly, to an architecture of a deep learning model configured to determine attributes of patient states.
Background Art
[0003] Background Information
[0003] In today's healthcare systems, access to patient diagnoses, treatments, tests, and other healthcare data across a large number of patients can provide useful insights for understanding diseases and developing new forms of therapies and treatments. For example, researchers can use data to evaluate patient responses to specific treatments, understand differences in patient treatment outcomes among patients with similar conditions, and identify patients to include in a cohort for a treatment. As one example, researchers may be interested in data regarding the sites of metastasis for cancer patients. In particular, researchers may be interested in data regarding the timing at which these metastases are diagnosed at different sites, especially data regarding the timing of diagnosis and related events and the associated timing of treatment lines.
[0004]
[0004] This information may be contained in the patient's electronic medical record (EMR). Each medical record may contain a large amount of data associated with the patient. Therefore, processing these records to identify relevant information about a statistically significant set of patients in a cohort can quickly become a task that cannot be overcome by manual methods. For example, researchers who want to perform statistical analysis on patients' medical data often need relatively large datasets (e.g., thousands, tens of thousands, hundreds of thousands, or millions of patients, or more) to extract significant insights from the data. It is virtually impossible for human reviewers to process this amount of data. Therefore, patient data collected by manual data extraction by human reviewers typically results in a small set of patients, thereby limiting the effectiveness of the data. Furthermore, due to the size limitations of the dataset, rare conditions that occur in only a small percentage of patients may not be represented in the dataset.
[0005]
[0005] Therefore, computer-based extraction of data from EMRS may be required to obtain a statistically significant set of data. However, relevant information within EMR may be stored in unstructured notes or various other forms of unstructured information (e.g., physician's notes, clinical laboratory technician reports, or other text-based data). This can make computer-based extraction of relevant information (e.g., information indicating the site of metastasis, the date of onset of those metastases, etc.) difficult and impossible without system capabilities to recognize important and / or relevant information from sources containing unstructured information.
[0006]
[0006] Given these and other shortcomings of the current technology, there is a need for technical solutions to more accurately identify patients by specific diagnoses, test results, or other characteristics related to specific dates, such as the date on which treatment or therapy is initiated. In particular, the solution should enable the identification of metastatic sites and the date of diagnosis of the corresponding metastatic sites, based on automated analysis of very large sets of patient data. [Overview of the project]
[0007] overview
[0007] Embodiments consistent with the present disclosure include systems and methods for extracting patient information. In embodiments, a model-assisted system may include at least one processor. The processor may be programmed to access a database storing one or more medical records associated with a patient, to use a first machine learning model to determine, based on unstructured information contained in one or more medical records, whether the patient is associated with a condition, to identify dates associated with the patient, to use a second machine learning model to determine, based on unstructured information, whether the patient is associated with a date-related condition, and to generate output indicating whether the patient is associated with a condition and whether the patient is associated with a date-related condition.
[0008]
[0008] In another embodiment, a computer implementation method for extracting patient information is disclosed. The method may include: accessing a database that stores one or more medical records associated with a patient; using a first machine learning model to determine, based on unstructured information contained in one or more medical records, whether the patient is associated with a condition; identifying dates associated with the patient; using a second machine learning model to determine, based on unstructured information, whether the patient is associated with a date-related condition; and generating an output indicating whether the patient is associated with a condition and whether the patient is associated with a date-related condition.
[0009]
[0009] Consistent with other disclosed embodiments, a non-temporary computer-readable storage medium may store program instructions, which are executed by at least one processing device performing any of the methods described herein.
[0010] Brief explanation of the drawing
[0010] The accompanying drawings, which are incorporated into and constitute part of this specification, serve to illustrate and illustrate the principles of various exemplary embodiments, together with the description. [Brief explanation of the drawing]
[0011] [Figure 1]
[0011] This block diagram shows an exemplary system environment for carrying out embodiments consistent with the present disclosure. [Figure 2]
[0012] A block diagram illustrating an exemplary medical record of a patient, consistent with the disclosed embodiments. [Figure 3]
[0013] In line with the disclosed embodiments, a timeline is provided as an example illustrating multiple therapy lines associated with a patient. [Figure 4]
[0014] This block diagram, consistent with the disclosed embodiments, is an example illustrating a process for extracting patient information associated with one or more dates. [Figure 5]
[0015] In line with the disclosed embodiments, the process is shown as an example for extracting features based on text snippets within unstructured data of patient medical records. [Figure 6A]
[0016] In line with the disclosed embodiments, a set of example documents that may be analyzed to determine whether a date-related patient condition exists is provided. [Figure 6B]
[0017] In line with the disclosed embodiments, the technique is presented as an example for analyzing the set of documents in Figure 6A. [Figure 6C]
[0017] In accordance with the disclosed embodiments, the technique is shown as an example for analyzing the set of documents in Figure 6A. [Figure 7]
[0018] This flowchart shows an example process for extracting patient information, consistent with the disclosed embodiments. [Modes for carrying out the invention]
[0012] Detailed explanation
[0019] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numerals are used to refer to the same or similar parts in the drawings and the following description. While several exemplary embodiments are described herein, modifications, adaptations, and other embodiments are possible. For example, components shown in the drawings may be substituted, added, or modified, and the exemplary methods described herein may be modified by substituting, rearranging, removing, or adding steps to the disclosed methods. Therefore, the following detailed description is not limited to the disclosed embodiments and examples. Instead, the appropriate scope is defined by the appended claims.
[0013]
[0020] Embodiments in this specification include computer implementations, tangible non-temporary computer-readable media, and systems. A computer implementation may be performed, for example, by at least one processor (e.g., a processing device) that receives instructions from a non-temporary computer-readable storage medium. Similarly, a system consistent with this disclosure may include at least one processor (e.g., a processing device) and memory, where memory may be a non-temporary computer-readable storage medium. As used herein, non-temporary computer-readable storage medium refers to any type of physical memory in which information or data readable by at least one processor can be stored. Embodiments include random-access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD-ROMs, DVDs, flash drives, disks, and any other known physical storage media. Singular terms such as “memory” and “computer-readable storage medium” may additionally refer to multiple structures, such as multiple memories and / or computer-readable storage media. As used herein, “memory” may include any type of computer-readable storage medium unless otherwise specified. A computer-readable storage medium may store at least one processor execution instruction, including instructions for causing a processor to perform steps or stages consistent with the embodiments herein. Additionally, one or more computer-readable storage media may be used when carrying out a computer implementation. The term “computer-readable storage medium” should be understood to include tangible items, excluding carrier and transient signals.
[0014]
[0021] The disclosed embodiments may automate the analysis of patient or patient population medical records to identify patient characteristics associated with one or more clinically relevant dates. For example, researchers, physicians, or other users may be interested in identifying patients diagnosed with brain metastases (or metastases in various other sites) in relation to one or more lines of therapy provided to the patient. This may provide insights, in particular, into how effective a particular line of therapy is. For example, if a patient is managed with three different cancer treatments, researchers may be interested in determining when and where different metastases appeared in relation to the start and end dates of these treatments. This may indicate how well a patient or group of patients responded to a particular therapy. Thus, the disclosed embodiments may enable the generation of large patient cohorts associated with metastatic sites based on the analysis of medical records. In particular, the system may automatically analyze patient medical records (including unstructured patient data) to identify large cohorts (e.g., up to several million patients or more) of patients exhibiting a certain type of metastasis, or exhibiting a certain type of metastasis in conjunction with a particular line of therapy. Additional techniques for selecting cohorts are described in detail in U.S. Patent No. 10,304,000 and PCT International Publication No. 2020 / 092316 (and corresponding U.S. Patent Application No. 16 / 971,238), assigned to the same applicant as this application, and are incorporated herein by reference in their entirety.
[0015]
[0022] Once this initial cohort is identified, the system may further filter the cohort based on other patient characteristics. In particular, the system may automatically make determinations related to the metastatic site, the date on which these metastases were diagnosed, and the relationship between these dates and other dates or characteristics such as the patient's treatment date for the metastatic site. For example, the system may identify patients with brain metastases who have been given a particular therapy, patients who have been diagnosed with brain metastases within a particular time range associated with a particular therapy, and the like. Although the diagnosis of the brain and other metastatic sites is used as an example throughout the present disclosure, the disclosed techniques may be applied to other patient states such as other sites of metastasis, other diagnostic types, test results (e.g., PDL1 gene mutation tests, etc.), clinical test results, or various conditions or events that may be represented in a patient medical record.
[0016]
[0023] FIG. 1 shows a system environment 100 as an example for implementing an embodiment consistent with the present disclosure, which will be described in detail below. As shown in FIG. 1, the system environment 100 may include several components including a client device 110, a data source 120, a system 130, and / or a network 140. The number and arrangement of these components are exemplary and will be understood from the present disclosure to be provided for purposes of illustration. Other arrangements and numbers of components may be used without departing from the teachings and embodiments of the present disclosure.
[0017]
[0024] As shown in Figure 1, an exemplary system environment 100 may include a system 130. System 130 may include one or more server systems, databases, and / or computing systems configured to receive information from entities over a network, process the information, store the information, and display / transmit the information to other entities over the network. Thus, in some embodiments, the network may facilitate cloud sharing, cloud storage, and / or cloud computing. In one embodiment, system 130 may include a processing engine 131 and one or more databases 132, which are shown in the area enclosed by the dashed line representing system 130. The processing engine 131 may include one or more general-purpose processors, such as a central processing unit (CPU) or graphics processing unit (GPU), and / or at least one processing device, such as one or more dedicated processors, such as an application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA).
[0018]
[0025] The various components of the system environment 100 may include an assembly of hardware, software, and / or firmware, including memory, a central processing unit (CPU), and / or a user interface. The memory may include any type of RAM or ROM embodied in a physical storage medium, such as a magnetic storage device including a floppy disk, hard disk, or magnetic tape, a semiconductor storage device such as a solid state disk (SSD) or flash memory, an optical disk storage device, or a magneto-optical disk storage device. The CPU may include one or more processors for processing data according to a set of programmable instructions or software stored in the memory. The functionality of each processor may be provided by a single dedicated processor or multiple processors. Further, the processor may include, but is not limited to, digital signal processor (DSP) hardware, or any other hardware capable of executing software. Any user interface may include any type or combination of input / output devices, such as a display monitor, keyboard, and / or mouse. The user of the environment 100 may include any individual who may desire to access and / or analyze patient data. Thus, throughout this disclosure, references to "user" of the disclosed embodiments may include any individual, such as a physician, researcher, quality assurance department of a healthcare facility, and / or any other individual.
[0019]
[0026] Data transmitted and / or exchanged within the system environment 100 may occur via a data interface. As used herein, a data interface may include any boundary across which two or more components of the system environment 100 exchange data. For example, the environment 100 may exchange data between the aforementioned software, hardware, database, devices, people, or any combination. Further, it should be understood that any suitable configuration of software, processors, data storage devices, and networks may be selected to implement the components of the system environment 100 and the features of the related embodiments.
[0020]
[0027] The components of environment 100 (including system 130, client device 110, and data source 120) can communicate with each other or with other components through network 140. Network 140 may include various types of networks, such as the Internet, wired wide area network (WAN), wired local area network (LAN), wireless WAN (e.g., WiMAX), wireless LAN (e.g., IEEE 802.11), mesh network, mobile / cellular network, enterprise or private data network, storage area network, virtual private network using a public network, near-field communication technology (e.g., Bluetooth, infrared, etc.), or various other types of network communication. In some embodiments, communication may take place over two or more of these forms of networks and protocols.
[0021]
[0028] System 130 may be configured to receive and store data transmitted via the network 140 from various data sources, including data source 120, process the received data, and transmit the data and results based on the processing to client devices 110. For example, system 130 may be configured to receive patient data from data source 120 or other sources on the network 140. In some embodiments, patient data may include medical information stored in the form of one or more medical records. Each medical record may be associated with a particular patient. Data source 120 may be associated with a variety of sources of medical information about a patient. For example, data source 120 may include the patient's healthcare providers, such as physicians, nurses, specialists, consulting physicians, hospitals, and clinics. Data source 120 may also be associated with laboratories, such as radiology or other imaging laboratories, hematology laboratories, and pathology laboratories. Data source 120 may also be associated with insurance companies or any other sources of patient data.
[0022]
[0029] System 130 may further communicate with one or more client devices 110 via Network 140. For example, System 130 may provide results from Data Source 120 to the client devices 110 based on the analysis of information. Client devices 110 may include any entity or device capable of receiving or transmitting data via Network 140. For example, client devices 110 may include a server or computing device such as a desktop or laptop computer. Client devices 110 may also include other devices such as mobile devices, tablets, wearable devices (i.e., smartwatches, implantable devices, fitness trackers, etc.), virtual machines, IoT devices, or various other technologies. In some embodiments, client devices 110 may send queries to System 130 via Network 140 about information concerning one or more patients, such as queries about patients having or associated with specific attributes, patients associated with specific attributes related to a specified date, or various other information about patients.
[0023]
[0030] In some embodiments, System 130 may be configured to analyze patient medical records (or other forms of unstructured data) to determine whether a patient is associated with a particular condition. For example, System 130 may analyze a patient's medical records to determine whether a patient has been diagnosed with a particular condition (e.g., metastasis in a particular area of the body), whether they have been tested for a particular condition, whether they have tested positive or negative for a particular condition, or various other characteristics. System 130 may be further configured to determine whether a patient exhibits a particular condition related to a date. For example, System 130 may determine whether a patient exhibits a particular condition prior to the start or end date of a treatment line, or any other date that may be relevant. System 130 may be configured to perform this analysis using one or more machine learning models, as further described below. While patient medical records are used as exemplary embodiments throughout this disclosure, it should be understood that in some embodiments, the disclosed systems, methods, and / or techniques may be similarly used to identify patients exhibiting a condition from other forms of records.
[0024]
[0031] Figure 2 is a block diagram showing an exemplary medical record 200 for a patient, consistent with the disclosed embodiments. The medical record 200 is received from data source 120 as described above and may be processed by system 130 to identify whether the patient is associated with a particular attribute. Records received from data source 120 (or elsewhere) may include one or both of unstructured data 210 and structured data 220, as shown in Figure 2. Structured data 220 may include quantifiable or classifiable data about the patient, such as sex, age, race, weight, life response, clinical test results, date of diagnosis, type of diagnosis, stage of illness (e.g., billing code), timing of therapy, treatments performed, date of consultation, type of treatment, insurance company and start date, medication instructions, medication management, or any other measurable data about the patient.
[0025]
[0032] As described above, much of the information relevant to making decisions about a patient, such as whether the patient was diagnosed with a particular condition, the date of diagnosis, and other similar information, can be stored in unstructured data of the patient's medical record. As used herein, unstructured data may include information about a patient that is not quantifiable or easily classifiable, such as physician's notes or the patient's clinical laboratory reports. For example, unstructured data 210 may include information such as a physician's explanation of a treatment plan, notes describing what happened during a consultation, statements or explanations from the patient, subjective assessments or explanations of the patient's health status, radiology reports, pathology reports, or any other form of information that is not stored in a structured format.
[0026]
[0033] In the data received from data source 120, each patient may be represented by one or more records generated by one or more healthcare professionals or by the patient. For example, a physician associated with the patient, a nurse associated with the patient, a physical therapist associated with the patient, etc., may each generate a medical record about the patient. In some embodiments, one or more records may be collated and / or stored within the same database. In other embodiments, one or more records may be distributed across multiple databases. In some embodiments, a record may be stored and / or provided in multiple electronic data representations. For example, a patient record may be represented as one or more electronic files, such as a text file, a Portable Document Format (PDF) file, or an Extended Markup Language (XML) file. If the document is stored as a PDF file, an image, or another file without text, the electronic data representation may also include text associated with the document derived from an optical character recognition process. In some embodiments, unstructured data may be captured by an extraction process, while structured data may be entered by healthcare professionals or computed using algorithms.
[0027]
[0034] In some embodiments, unstructured data may include data associated with a particular patient condition. As used herein, patient “condition” may refer to any attribute or characteristic associated with a patient’s health or health status. For example, condition may refer to a diagnosed condition of a patient, such as whether the patient has been diagnosed with a particular disease or illness. In some embodiments, it may refer to a stage or state of a particular diagnosed condition. For example, in the case of a patient diagnosed with cancer, condition may be a specific site of metastasis of the cancer (e.g., whether the patient has been diagnosed with metastasis to the brain, liver, bones, lungs, adrenal glands, peritoneum, or various other sites of metastasis). While metastasis sites are used as examples throughout this disclosure, it should be understood that the disclosed embodiments and techniques may be applied to other conditions as well, and that this disclosure is not limited to any particular condition.
[0028]
[0035] Whether a patient is associated with a condition of interest and whether the condition was observed up to a specific date can be extracted from the unstructured data 210. For example, the condition may include brain metastases, and the system 130 may analyze the patient's medical records to determine whether the patient has been diagnosed with brain metastases. In some embodiments, this extraction may occur through embodiments of one or more trained models. For example, one or more models (e.g., one or more machine learning models) may be trained to identify documents, such as patient medical records, that indicate whether the patient has been diagnosed with brain metastases and whether the brain metastases occurred in relation to one or more dates (e.g., before or after the date of interest).
[0029]
[0036] Various types of dates may be analyzed in relation to a patient's condition. These target dates may vary depending on the specific use, such as the type of cohort selected, the type of study, the type of condition being analyzed, or various other factors. Dates or a set of dates may be specified in various ways. In some embodiments, a user interface may be provided so that the user can input a reference date to identify patients diagnosed with a condition prior to that date. For example, the user interface may be presented on one or more client devices 110. Alternatively or additionally, dates may be identified in relation to another date associated with a diagnosis, treatment, or other form of patient care. For example, a date may be the start or end date of a particular treatment or therapy line for a patient. As used herein, a therapy line (or treatment line) may refer to a therapy employed to treat a particular disease or condition. For example, a therapy line may include the management of specific medications for a patient (e.g., pharmacotherapy, chemotherapy, etc.), surgical procedures, gene therapy, immunotherapy, changes in the patient's diet, radiation therapy, physiotherapy, counseling or psychotherapy, meditation, sleep therapy, or various other forms of treatment that may be prescribed for a patient. In some embodiments, the date may be an index date that defines inclusion criteria for a cohort. For example, a researcher or other user may define a date such that patients exhibiting a condition associated with the index date are eligible to be included in the study.
[0030]
[0037] In some embodiments, multiple dates may be specified. For example, a patient may receive multiple lines of therapy, and the onset of a condition may be analyzed against one or more lines of therapy. Thus, for each patient condition, a timeline may be developed showing a list of dates with indicators indicating whether the condition occurred before a given date. For example, these dates may include the date of advanced diagnosis, the date the patient started the first line of therapy, the date the patient finished the first line of therapy, the date the patient started the second line of therapy, the date the patient finished the second line of therapy, and so on. The model may output a table or similar data structure indicating whether the condition was diagnosed before these dates. This process may be performed across a pool of patient data to extract information that may be statistically robust to the researcher.
[0031]
[0038] Figure 3 shows a timeline 300 as an example illustrating multiple therapy lines associated with a patient, consistent with the disclosed embodiments. For example, one or more patients may receive a first therapy line 301 treated with “Medication X”, a second therapy line 302 treated with “Medication Y”, and a third therapy line 303 treated with “Medication Z”. As shown in timeline 300, the first patient ("Patient 1") may develop brain metastases during therapy line 301, while the second patient ("Patient 2") may develop brain metastases during therapy line 302.
[0032]
[0039] System 130 may be configured to analyze medical records associated with patients 1 and 2 to determine that patients 1 and 2 have been diagnosed with brain metastases. System 130 may further determine whether brain metastases occurred for patients 1 and 2 relative to the respective start dates of therapy lines 301, 302, and 303. For example, System 130 may generate output 310 indicating that patient 1 was diagnosed with brain metastases that were not present by the start date of therapy line 301 but had occurred by the start dates of therapy lines 302 and 303. This can provide insights to, for example, physicians, researchers, or other users regarding the effectiveness of therapy lines 301, 302, and / or 303. In some embodiments, output 310 may provide information about multiple patients. For example, as shown in Figure 3, output 310 may also indicate that patient 2 has brain metastases that were not present before the start of therapy lines 301 and 302 but had occurred before the start of therapy line 303. Output 310 may contain similar information about a larger set of patients. Therefore, patients may be selected for a cohort, or analyzed based on whether the patient was diagnosed with brain metastases or whether the condition occurred in relation to one or more dates.
[0033]
[0040] In some embodiments, this process may use two or more separately trained models (e.g., two or more distinct machine learning models). For example, a first machine learning model may be trained to identify patients diagnosed with the target metastatic site. A second machine learning model may then be trained to determine other information, such as the type of metastasis, the location of the metastasis, the date the metastasis appeared or was diagnosed, the type of treatment the patient received, the dates of these treatments, and other potentially relevant information. Based on this information, researchers may use the disclosed systems and methods to identify patients who have certain characteristics, such as whether the metastasis occurred before one or more dates. As mentioned above, these dates may correspond to dates of therapy lines associated with the patient, such as the dates on which a particular treatment was initiated or ended. Thus, researchers may identify patients with metastatic diagnosis dates relative to these therapy lines, thereby providing insights or similar insights into the patient's response to these therapies.
[0034]
[0041] Figure 4 is a block diagram as an example illustrating a process 400 for extracting patient information associated with one or more dates, consistent with the disclosed embodiments. Process 400 may be performed based on a set of medical records 410. In some embodiments, medical records 410 may be associated with a specific patient. Thus, process 400 may be used to determine whether a patient has been diagnosed with brain metastases and whether the metastases occurred before or after a specified date. Alternatively or additionally, medical records 410 may be associated with multiple patients. Thus, process 400 may be used to identify patients diagnosed with brain metastases, and further, which patients developed brain metastases before or after a specified date. As a result, process 400 may enable analysis of a large set of patient medical records (including unstructured patient data) to identify cohorts of patients exhibiting a certain type of metastasis, or a combination of several types of metastases, and whether the metastases occurred before initiating a certain line of therapy or before other target dates. Although brain metastases are used as an example, process 400 may be similarly applied to other metastatic sites or other conditions.
[0035]
[0042] Medical records 410 may be input to a first trained model 420. The trained model 420 may be configured to identify patients diagnosed with metastases in specific locations, such as brain metastases. For example, a training algorithm such as an artificial neural network may receive training data from medical records in the form of unstructured data. The training data may be labeled to indicate the specific condition to which the patient has been diagnosed, associated with the unstructured data. As a result, the model may be trained to determine, based on the unstructured data in the patient medical records, whether the patient has been diagnosed with brain metastases or not. In accordance with this disclosure, a variety of other machine learning algorithms may be used, including logistic regression, linear regression, regression, random forest, K-nearest neighbors (KNN) models (e.g., as described above), K-means models, decision trees, Cox proportional hazards regression models, naive Bayes models, support vector machine (SVM) models, gradient boosting algorithms, or any other form of machine learning model or algorithm. An example training process is described in more detail below with reference to Figure 5.
[0036]
[0043] As shown in Figure 4, the trained model 420 can determine on a patient-by-patient basis whether the patient has been diagnosed with brain metastases (or various other conditions). If not diagnosed, output 422 may indicate that brain metastases are not associated with the patient. On the other hand, if the trained model 420 determines that the patient has been diagnosed with brain metastases, the data associated with the patient may be input into a second trained model 430, as shown. In embodiments where the medical records 310 include a large set of patients, the trained model 420 may thereby filter out patients (and patient medical records) that do not include the specified condition. In some embodiments, documents identified by the trained model 420 as being associated with a condition may be selected for input to the trained model 430. Thus, including the trained model 420 may enable more efficient application of the trained model 430.
[0037]
[0044] The trained model 430 can extract information to further reduce the set of documents. For example, as mentioned above, the trained model 430 could enable researchers to identify patients who have been diagnosed with a particular condition by a certain date. Similar to the trained model 420, the trained model 430 can be trained on a training dataset of documents known to contain a diagnosis for a particular condition, along with the date of diagnosis for the condition. As a result, the second model can be trained to include the date on which the relevant diagnosis was made. A similar training process can be applied to determine treatment lines, the start date of treatment lines, and so on.
[0038]
[0045] As a result, the trained model 430 can determine, on a patient-by-patient basis, whether the patient has been diagnosed with brain metastases (or various other conditions) by a certain date. For example, that date may be the start date for a particular line of therapy, or various other dates that may be relevant. If not diagnosed, output 432 may indicate that metastases in the brain have been diagnosed but have not occurred by the date in question. If the patient has developed brain metastases by the specified date, this may be reflected in output 434. In some embodiments, the trained model 430 may determine whether the patient condition has occurred for multiple dates, as shown in Figure 3. Alternatively or additionally, the trained model 430 may include multiple trained models, each of which may determine whether the patient developed the condition by different dates. In embodiments where the medical record 310 includes a large set of patients, models 420 and 430 may be used to narrow the set of patients to identify only those patients who developed the condition related to a specified date. Thus, process 400 may be used to select patients for a cohort or to identify a specific subset of patients. The above description refers to a process performed by two separate models, but in some embodiments, this process may be performed by a single model or by two or more models.
[0039]
[0046] Figure 5 shows process 500 as an example for extracting features based on text snippets in unstructured data of patient medical records, consistent with the disclosed embodiments. As described above, the trained model 420 may be configured to determine whether a patient exhibits a particular condition. In some embodiments, system 130 may perform a search on a set of training documents (which may be a set of patient medical records) known to contain information about diagnosing a particular condition to extract features that can be input into a model, such as a logistic regression algorithm. For example, system 130 may perform a search 510 on one or more unstructured medical record documents to extract snippets associated with a patient condition. In the case of brain metastases, relevant terms may include “brain,” “temporal,” “occipital,” “frontal,” or other terms that may generally be associated with the brain. Accordingly, system 130 may identify terms 522, as shown in step 520. System 130 may then extract text snippets around the relevant terms in the unstructured data. For example, as shown in step 530, a snippet 532 surrounding a term 522 may be extracted from unstructured text. The length or structure of the snippet may be specified in various ways. In some embodiments, the snippet 532 may be defined based on a predefined window. For example, the snippet may be defined based on a predetermined number of characters before and after the target term 522 in the text (e.g., 20, 50, 60 characters, or any number of characters appropriate to capture the context for the use of the term). The window may also be defined to consider word boundaries so that subwords are not included at the end of the snippet, for example, by expanding or contracting the window to the end of the word boundary. In some embodiments, the window may be defined based on a predefined number of words or other variables.
[0040]
[0047] In some embodiments, system 130 may replace term 522 with tokenized term. This can ensure that patient conditions are represented using the same technical terminology in each extracted snippet. For example, documents containing "brain" and documents containing "cerebrum" may both result in the extraction of snippets containing the term "[_brain_]" or similar tokens. The use of tokens can also improve the performance of machine learning models by reducing feature sparsity, speeding up training time, and allowing the model to converge on a more limited set of labeled data.
[0041]
[0048] Next, the system may extract other words or phrases from within a snippet that may be relevant to identifying diagnostic information within a document. These words or phrases may be represented as features 540, as shown in Figure 5. For example, the snippet may include the terms “MRI,” “metastasis,” “tumor metastasis,” “lesion,” and “radiation.” For each of these terms (or phrases), a feature may be generated. The features may be input into a normalized logistic representation to determine the relative weights for each of the features. Thus, a model for identifying documents containing brain metastasis diagnoses may be generated and trained. Process 500 provides a general overview of the process for extracting feature vectors for the purpose of determining whether a patient is associated with a particular condition, but various other techniques may be used. Examples of machine learning techniques for determining whether a patient is associated with a particular trait are described in detail in U.S. Patent Application Publication 2021 / 0027894A1 and PCT International Publication 2020 / 092316, which are assigned to the same applicant as this application. The contents of these applications are incorporated herein by reference in their entirety.
[0042]
[0049] In some embodiments, irrationality such as the diagnosis being represented by unstructured data may make it difficult to determine the date of diagnosis in some cases. For example, a physician may include a note indicating that brain metastases were observed "one month ago" or "three months after the start of [treatment line]." Taking this into consideration, the documents input into the trained model 430 may be restricted and / or modified to improve the accuracy of determining patient status associated with one or more dates.
[0043]
[0050] Figure 6A shows a set of example documents that can be analyzed to determine whether a date-related patient condition exists, consistent with the disclosed embodiments. The set of documents 602, 604, 606, 608, and 610 can be analyzed to determine whether a patient had liver metastases before a particular date 620. Each of documents 602, 604, 606, 608, and 610 can be associated with a document date. The document date may refer to the date the document was generated, or it may be extracted from metadata or other data associated with the document. The document date may refer to various other dates associated with the document, such as the date the document was updated, revised, filed, published, or any other relevant date. For example, document 602 may be associated with the date May 1, 2016. For illustrative purposes, documents 602, 604, 606, 608, and 610 are shown along timeline 600, corresponding to the document date for each document. Researchers may be interested in determining whether the liver metastases occurred before date 620, in this case February 14, 2018.
[0044]
[0051] Documents 602, 604, 606, 608, and 610 may each contain text relevant to the occurrence of liver metastasis. Therefore, references to events both before and after its occurrence may exist. Furthermore, although explicit dates are not required, unstructured text data may indicate various stages of occurrence. For example, document 602 may indicate that the patient has breast cancer and that no metastasis has occurred in the liver (or other potential metastatic sites). Documents 606 and 608 may indicate that liver metastasis has been diagnosed, thereby indicating that the metastasis occurred before those dates. For each document, the trained model 430 may predict the probability that the condition occurred before, after, or on date 620.
[0045]
[0052] In some cases, as mentioned above, dates may be expressed in relative terms. For example, a document may indicate that liver metastases occurred relative to another event (e.g., initiating a line of treatment). In the example shown in Figure 6A, document 610 may include an explicit date on which liver metastases occurred. Other documents, on the other hand, may express dates using related terms such as "three months ago." Including these documents that express relative dates can lead to errors or inaccuracies in determining whether the condition occurred before a given date. For example, the trained model 430 may not properly interpret the relative terms used, which can distort the analysis.
[0046]
[0053] Therefore, the trained model 430 may only consider documents associated with document dates prior to date 620. Figure 6B shows a limited set of documents as an example for input to the model, consistent with the disclosed embodiments. In particular, a deadline 632 may be specified so that documents after deadline 632 are not analyzed for the trained model 430. In the embodiment shown in Figure 6B, document 610 may be excluded from the documents input to the trained model 430. In some embodiments, deadline 632 may be identical to date 620. This may mitigate the problem of ambiguous descriptions of diagnostic dates in unstructured data by focusing only on documents prior to the reference date. In some embodiments, a buffer 630 may also be applied so that deadline 632 is before or after date 620. For example, the buffer 630 may be a buffer of one day, seven days, two weeks, or any other suitable period after date 620. This may ensure that documents with dates immediately following the reference date, which may still contain relevant information, are not excluded from the analysis.
[0047]
[0054] In some embodiments, system 130 may be configured to generate one or more “pseudo-documents” for input to a trained model 430. As described above, by applying the deadline 632, the trained model 430 may miss explicit dates in document text that could provide an accurate indicator of when the condition occurred. For example, as shown in Figure 6B, the trained model 430 may not be used to analyze document 610, which could indicate that liver metastases occurred on or around February 17, 2017. Taking this into consideration, system 130 may generate a pseudo-document 634, as shown in Figure 6C. The pseudo-document may contain text from the original document, but may be associated with a document date contained in the text. For example, pseudo-document 634 may contain text from document 610 that refers to an explicit date, but may be associated with a document date that matches the explicit date, in this case February 15, 2017. Thus, the pseudo-document may contain text that explicitly indicates a date as if the document were written on that date. This could result in improved accuracy in determining, based on documents 602, 604, 606, 608, and 610, whether a patient's condition occurred before date 620.
[0048]
[0055] Based on the disclosed technology, data on diagnosis and date of diagnosis can be extracted from patient medical records. In particular, such data may be extracted from unstructured data within medical records, which can provide a relatively large and statistically robust set of data. Furthermore, the date of diagnosis for various conditions can be compared with other key dates associated with one or more lines of therapy for the patient. This may enable researchers or other users to identify patients who show a certain response to treatment (e.g., for cohort selection), compare the relative effectiveness of one or more treatments, or perform various other forms of analysis.
[0049]
[0056] Figure 7 is a flowchart illustrating process 700 as an example for extracting patient information, consistent with the embodiments disclosed. Process 700 may be executed by at least one processing device, such as processing engine 131, as described above. Throughout this disclosure, the term “processor” should be understood as a shorthand for “at least one processor.” In other words, a processor may include one or more structures that perform logical operations on whether such structures are co-located, connected, or distributed. In some embodiments, non-temporary computer-readable media may include instructions that cause the processor to execute process 700 when executed by the processor. Furthermore, process 700 is not necessarily limited to the steps shown in Figure 7, and any steps or processes of the various embodiments described throughout this disclosure may also be included in process 700, including those described above with respect to Figures 3, 4, 5, 6A, and 6B.
[0050]
[0057] In step 710, process 700 includes accessing a database that stores one or more medical records associated with a patient. For example, system 130 may access patient medical records from a local database 132 or from an external data source such as data source 120. Medical records may include one or more electronic files such as text files, image files, PDF files, XLM files, and YAML files. One or more medical records may correspond to the medical records 200 described above.
[0051]
[0058] In step 720, process 700 includes determining whether the patient is associated with a condition. In some embodiments, the condition may include a diagnosed condition of the patient, as described above. For example, the condition may include a diagnosed metastatic site relevant to the patient, such as brain metastases. In another embodiment, the condition may include whether the patient has been examined for a condition. The determination in step 720 may be made using a first machine learning model, such as the trained model 420 described above. The determination in step 720 may be based on unstructured information contained in one or more medical records. For example, this may include unstructured data 210, as shown in Figure 2 and described above. The unstructured information may include text written by the healthcare provider, radiology reports, pathology reports, or various other forms of patient-related text. In some embodiments, the medical record may further include additional structured data 220.
[0052]
[0059] In step 730, process 700 includes identifying a date associated with the patient. For example, the date may correspond to the date 620 described above. This date can be determined in various ways. In some embodiments, step 730 may include receiving a date instruction through a user interface, such as the user interface of the computing device 110. In some embodiments, the date can be determined based on one or more dates related to the patient's care. For example, step 730 may include identifying at least one of the start or end dates of a treatment line for the patient associated with a condition.
[0053]
[0060] In step 740, process 700 includes determining whether the patient is associated with a condition relating to a date. For example, this may include determining whether the patient experienced the condition before or after a date. In some embodiments, the determination in step 740 may be made using a second machine learning model, such as the trained model 430. Alternatively or additionally, a single model or more than two models may be used in steps 730 and 740. In some embodiments, multiple dates may be used. For example, step 730 may include identifying multiple dates, and step 740 may include determining whether the patient is associated with a condition relating to each of the multiple dates. For example, as described above with reference to Figure 3, each of the multiple dates may include a start date for a particular treatment line for a patient associated with a condition, and step 740 may include determining whether the condition occurred before the start date of each treatment line.
[0054]
[0061] In some embodiments, step 740 may include applying a second machine learning model to documents having timestamps before the deadline, as described above with respect to Figure 6B. For example, step 740 may include identifying multiple documents in one or more medical records that have timestamps before the deadline. Thus, determining whether a patient is associated with a date-related condition may include applying a second machine learning model to multiple documents. In some embodiments, the deadline may be that date. Alternatively or additionally, the deadline may be based on a predetermined buffer period before or after that date. For example, the deadline 632 may be determined based on a buffer 630 for the date 620, as described above.
[0055]
[0062] In step 750, process 700 includes generating an output indicating whether the patient is associated with a condition and whether the patient is associated with a date-related condition. For example, this may include generating one of outputs 310, 422, 432, and / or 434 as described above. In some embodiments, process 750 may include transmitting the output to at least one of the healthcare provider or research entity.
[0056]
[0063] As described above, process 700 may be performed on multiple patients to identify a group of patients associated with a date-related condition. This can help researchers or healthcare providers select patients for cohorts, such as a cohort to include in a study or a cohort to receive a specific treatment. Thus, process 700 may include using a first machine learning model to determine, based on unstructured information, whether each of multiple patients is associated with a condition, and using a second machine learning model to determine, based on unstructured information, whether each of multiple patients is associated with a date-related condition. As a result, the output may identify a set of multiple patients associated with a date-related condition.
[0057]
[0064] The foregoing description is provided for illustrative purposes only. It is not exhaustive and is not limited to the exact form or embodiment disclosed. Modifications and adaptations will become apparent to those skilled in the art from examining this specification and practicing the disclosed embodiments. In addition, although the embodiments of the disclosed embodiments are described as being stored in memory, those skilled in the art will understand that these embodiments may also be stored on other types of computer-readable media, such as hard disks or CD-ROMs, or other forms of RAM or ROM, USB media, DVDs, Blu-rays, 4K Ultra HD Blu-rays, or other optical drive media, or other secondary storage devices.
[0058]
[0065] Computer programs based on written descriptions and disclosed methods are within the scope of the skills of an experienced developer. Various programs or program modules may be created using any of the techniques known to those skilled in the art, or they may be designed in conjunction with existing software. For example, a program section or program module may be designed in or using .Net Framework, .Net Compact Framework (and related languages such as Visual Basic and C), Java, Python, R, C++, Objective-C, HTML, HTML / AJAX combination, XML, or HTML including Java applets.
[0059]
[0066] Furthermore, while exemplary embodiments are described herein, the scope of any and all embodiments has equivalent elements, modifications, omissions, combinations (e.g., aspects across various embodiments), adaptations, and / or alterations, as can be understood by those skilled in the art based on this disclosure. The limitations in the claims should be interpreted broadly based on the language adopted in the claims and are not limited to the examples described herein or described in the course of the application. The examples should be interpreted as non-exclusive. Furthermore, the steps of the disclosed methods may be modified in any way, including by rearranging the steps and / or inserting or deleting steps. Thus, the specification and examples are for illustrative purposes only, and the true scope and idea are intended to be shown by the entire scope of the following claims and their equivalents.
Claims
1. A model support system for extracting patient information, Access a database that stores one or more medical records associated with a patient, Using a first machine learning model, determine whether the patient is associated with a condition based on unstructured information contained in one or more medical records. Identify the date associated with the aforementioned patient, Using a second machine learning model, based on the unstructured information, determine whether the patient is associated with the condition related to the date, To generate output indicating whether the patient is associated with the condition and whether the patient is associated with the condition related to the date, A system having at least one processor programmed to do so.
2. The system according to claim 1, wherein the condition includes a diagnosed metastatic site associated with the patient.
3. The system according to claim 2, wherein the diagnosed metastatic site includes brain metastases.
4. The system according to claim 1, wherein the state includes the diagnosed state of the patient.
5. The system according to claim 1, wherein determining whether the patient is associated with the condition includes determining whether the patient has been examined for the condition.
6. The system according to claim 1, wherein identifying the date includes receiving an indicator of the date through a user interface.
7. The system according to claim 1, wherein identifying the date includes identifying at least one of the start or end dates of a treatment line for the patient associated with the condition.
8. The system according to claim 1, wherein identifying the date includes identifying a plurality of dates, and determining whether the patient is associated with the condition includes determining whether the patient is associated with each of the plurality of dates.
9. The system according to claim 8, wherein each of the plurality of dates includes a start date for a specific line of treatment for the patient associated with the condition.
10. The system according to claim 1, wherein the at least one processor is further programmed to identify a plurality of documents among the one or more medical records having a timestamp prior to a deadline, and determining whether the patient is associated with the condition related to the date is further comprising applying the second machine learning model to the plurality of documents.
11. The system according to claim 10, wherein the deadline is the date mentioned above.
12. The system according to claim 10, wherein the deadline is based on a predetermined buffer period before or after the date.
13. The one or more medical records are associated with multiple patients, and the at least one processor is Using the first machine learning model, and based on the unstructured information, determine whether each of the multiple patients is associated with the condition. The second machine learning model is further configured to determine, based on the unstructured information, whether each of the multiple patients is associated with the condition related to the date, The system according to claim 1, wherein the output identifies the set of patients associated with the state related to the date.
14. The system according to claim 1, wherein the at least one processor is further configured to transmit the output to at least one of a healthcare provider or a research entity.
15. A computer-assisted method for extracting patient information, Accessing a database that stores one or more medical records associated with a patient, Using a first machine learning model, determine whether the patient is associated with a condition based on unstructured information contained in one or more medical records. Identifying the date associated with the aforementioned patient, Using a second machine learning model, determine, based on the unstructured information, whether the patient is associated with the condition related to the date, To generate an output indicating whether the patient is associated with the condition and whether the patient is associated with the condition related to the date, A method that includes and is executed by at least one processor.
16. The method according to claim 15, wherein the condition includes a diagnosed metastatic site associated with the patient.
17. The method according to claim 15, wherein identifying the date includes receiving an indicator of the date through a user interface.
18. The method according to claim 15, wherein identifying the date includes identifying at least one of the start or end dates of a treatment line for the patient associated with the condition.
19. The method described above is The method of claim 15, further comprising identifying a plurality of documents among the one or more medical records that have a timestamp prior to the deadline, and determining whether the patient is associated with the condition related to the date, by applying the second machine learning model to the plurality of documents.
20. A non-temporary computer-readable medium storing instructions executable by at least one processor for performing a method, wherein the method is Accessing a database that stores one or more medical records associated with a patient, Using a first machine learning model, determine whether the patient is associated with a condition based on unstructured information contained in one or more medical records. Identifying the date associated with the aforementioned patient, Using a second machine learning model, determine, based on the unstructured information, whether the patient is associated with the condition related to the date, To generate an output indicating whether the patient is associated with the condition and whether the patient is associated with the condition related to the date, Non-temporary computer-readable media, including [specific examples of such media].