Patient pooling based on machine learning models
Patent Information
- Application Number
- JP2024558161
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-04-01
- Filing Date
- 2023-03-31
- Publication Date
- 2026-02-03
AI Technical Summary
Existing machine learning models are difficult to provide accurate predictions and treatment recommendations for specific patient types when using real-world clinical data to predict, resulting in inaccurate diagnostic and treatment decisions.
By identifying patient groups with similar attributes, using machine learning models (such as random survival forest models) for clinical prediction, outputting attributes such as patient biological data, laboratory test results, biomarker data, etc. to assist in clinical judgment.
Improve the accuracy of prediction of new patients, helping clinicians develop more effective treatment plans, and improve patients' survival probability and quality of life.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Patent Application No. 63 / 362,373, filed April 1, 2022, the disclosure of which is incorporated herein by reference in its entirety. [Background technology]
[0002] background Predictive machine learning models trained using real-world clinical data offer enormous potential for providing patients and their clinicians with patient-specific information regarding diagnosis, prognosis, or optimal treatment courses. Machine learning models can be trained to perform clinical predictions to predict patient medical outcomes, such as, for example, the probability of a patient's survival as a function of time from diagnosis (e.g., advanced stage cancer), survival time from diagnosis for new patients, other types of prognosis, etc. Predictions can be provided to patients, for example, so that they can better plan their future, thus resulting in an improvement in their quality of life.
[0003] Many machine learning models have been developed using a wide range of data on many types of patients with varying symptoms, modes of treatment applied to those conditions, and prognoses. Such models may not be fully trained with data applicable to a particular patient type, and / or may be weighted from data that may be less valuable to a particular patient type, leading to inaccurate prognosis and / or treatment recommendations. Thus, there is a need for improved methods of utilizing machine learning models to enhance patient care. Summary of the Invention
[0004] Quick Overview Disclosed herein is a technique for facilitating clinical decision-making for a patient based on identifying a group of patients having similar attributes to the patient. The group of patients can be identified using information from a predictive machine learning model that performs clinical predictions for the patient. At least some of the attributes of the group of patients can be output to assist in clinical decision-making. The attributes can include, for example, biographical data of the patient, the results of one or more laboratory tests of the patient, biopsy image data of the patient, molecular biomarkers of the patient, the site of the patient's tumor, and the stage of the patient's tumor.
[0005] Specifically, the clinical decision support system can use the machine learning model to make clinical predictions for a patient based on the patient's attributes. For example, the machine learning model can include a random survival forest (RSF) model to predict the probability of a patient's survival as a function of time since diagnosis. Furthermore, the clinical decision support system can identify a group of patients (e.g., a "similarities-based patient pool") that have certain attributes similar to the patient's attributes. The similarities-based patient pool can include patients with health conditions comparable to the patient, and the similarities-based patient pool can be identified based on patients that are similar to the patient in a subset of attributes that are most relevant to the clinical prediction (e.g., the probability of survival at a particular time point from diagnosis) performed by the machine learning model. The clinical decision support system can then obtain information of the attributes of the similarities-based patient pool.
[0006] The clinical decision support system can output a predicted probability of survival for the new patient, as well as the attributes of the new patient that are determined to be most relevant to this prediction. Similarity-based patient pooling allows the clinical decision support system to output a survival function for the similarity-based patient pool, as well as a summary of the attributes of the patients in the similarity-based patient pool. This focuses on the attributes that are most relevant to the new patient's survival prediction, and facilitates comparison of the new patient's attributes with those of the similarity-based patient pool. Investigating the relationship between attributes and survival in the similarity-based patient pool can help clinicians determine a course of action (e.g., treatment) to improve the probability of survival for the new patient.
[0007] In some embodiments, a computer-implemented method for facilitating clinical decisions includes receiving first data corresponding to a plurality of features of a first patient, each feature representing an attribute of a plurality of attributes; inputting the first data into a machine learning model to generate an outcome of a clinical prediction for the first patient, the machine learning model having a plurality of feature importance metrics associated therewith, the plurality of feature importance metrics defining a relevance of each of the plurality of features to the clinical prediction; obtaining second data corresponding to a plurality of features of each of a group of patients based on a similarity of at least some of the plurality of features between the first patient and the group of patients, the similarity being based on the first data, the second data, and the plurality of feature data importance metrics; generating content based at least in part on the outcome of the clinical prediction and the second data; and outputting the content such that a clinical decision can be made for the first patient based on the content.
[0008] In some embodiments, the plurality of attributes includes at least one of patient biographical data, results of one or more laboratory tests of the first patient, biopsy image data of the first patient, molecular biomarkers of the first patient, tumor site of the first patient, or stage of the tumor of the first patient.
[0009] In some embodiments, the clinical prediction comprises at least one of a probability of survival of the first patient at a given time from when the first patient is diagnosed with the tumor, a survival time of the first patient from when the first patient is diagnosed with the tumor, or an outcome upon receiving a treatment.
[0010] In some embodiments, the machine learning model includes a random forest survival model comprising f decision trees each configured to process a subset of the first subset of data to generate a cumulative survival probability, and a patient's survival rate at a given time is determined based on an average of the cumulative survival probabilities output by the multiple decision trees.
[0011] In some embodiments, the group of patients is a first group of patients, the first group of patients is selected from a second group of patients, and the machine learning model is trained based on patient data of the second group of patients.
[0012] In some embodiments, the method further includes ranking the plurality of features based on a relevance of each feature of the plurality of features to the clinical prediction, determining a subset of the plurality of features based on the ranking, and determining a first group of patients based on a similarity in the subset of the plurality of features between the first patient and the first group of patients.
[0013] In some embodiments, the first group of patients is selected from the second group of patients based on a similarity in a subset of the plurality of features between the first patient and the first group of patients exceeding a threshold value.
[0014] In some embodiments, the first group of patients is selected from the second group of patients based on selecting a threshold number of patients that are most similar to the first patient on a subset of the plurality of features.
[0015] In some embodiments, the method further includes calculating a weighted aggregate similarity based on summing scaled similarities for each feature of at least some of the plurality of features, where each similarity is scaled by a weight based on the relevance of the feature, and identifying groups of patients based on the weighted aggregate similarities between the first patient and each of the groups of patients.
[0016] In some embodiments, the feature importance metric for the feature is determined based on a relationship between an error in an outcome of a clinical prediction generated by the machine learning model for a second patient of the first group of patients, the outcome of the clinical prediction being generated from a plurality of values of the feature for the second patient, and the error is calculated based on a comparison of the outcome of the clinical prediction and an actual clinical outcome for the second patient.
[0017] In some embodiments, the content includes at least one of a median survival time for the first group of patients, or a Kaplan-Meier survival curve for the first group of patients.
[0018] In some embodiments, the content includes values of one or more of a first subset of a plurality of features of a first patient, a first group of patients, and a second group of patients.
[0019] In some embodiments, a computer product includes a computer readable medium having stored thereon a plurality of instructions for controlling a computer system to perform the operations of any of the methods described above.
[0020] In some embodiments, a system comprises the computer product described herein and one or more processors for executing instructions stored on a computer-readable medium.
[0021] In some embodiments, a system comprises means for performing any of the methods described herein.
[0022] In some embodiments, the system is configured to perform any of the methods described herein.
[0023] In some embodiments, a system includes modules for performing each of the steps of any of the methods described herein.
[0024] These and other exemplary embodiments are described in detail below. For example, other embodiments relate to systems, apparatus, and computer-readable media associated with the methods described herein.
[0025] A better understanding of the nature and advantages of embodiments of the present disclosure may be obtained with reference to the following detailed description and the accompanying drawings. [Brief description of the drawings]
[0026] The detailed description will now be made with reference to the accompanying drawings.
[0027] [Figure 1A] 1 illustrates an exemplary technique for facilitating clinical decisions based on clinical predictions according to certain aspects of the present disclosure. [Figure 1B] 1 illustrates an exemplary technique for facilitating clinical decisions based on clinical predictions according to certain aspects of the present disclosure. [Figure 1C] 1 illustrates an exemplary technique for facilitating clinical decisions based on clinical predictions according to certain aspects of the present disclosure. [Figure 2A]1 illustrates an improved clinical decision system enabling machine learning based patient pooling in accordance with certain aspects of the present disclosure. [Figure 2B] 1 illustrates an improved clinical decision system enabling machine learning based patient pooling in accordance with certain aspects of the present disclosure. [Figure 2C] 1 illustrates an improved clinical decision system enabling machine learning based patient pooling in accordance with certain aspects of the present disclosure. [Figure 2D] 1 illustrates an improved clinical decision system enabling machine learning based patient pooling in accordance with certain aspects of the present disclosure. [Figure 2E] 1 illustrates an improved clinical decision system enabling machine learning based patient pooling in accordance with certain aspects of the present disclosure. [Figure 2F] 1 illustrates an improved clinical decision system enabling machine learning based patient pooling in accordance with certain aspects of the present disclosure. [Diagram 3] 1 illustrates a method for performing a machine learning based patient pooling operation according to certain aspects of the present disclosure. [Figure 4] 1 illustrates an exemplary computer system that can be utilized to implement the techniques disclosed herein. [Diagram 5] 1 illustrates one example of how patient data from a patient pool can be used. [Figure 6] 1 illustrates another example of how patient data from a patient pool can be used. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0028] Detailed Description As described above, a predictive machine learning model can be trained to perform clinical predictions to predict the medical outcomes of new patients. The new patient can be any patient who is alive and undergoing clinical judgment. For example, a random survival forest (RSF) model can be trained based on data of previous patients, as well as their survival statistics, to predict the survival probability of the new patient as a function of time since diagnosis (e.g., of advanced stage cancer). The predictions can be provided to the new patient, for example, to improve their ability to plan for the future. This has the potential to improve the quality of life of the patient.
[0029] While the clinical predictions provided by predictive machine learning models can provide valuable information for new patients, the clinical prediction results themselves may not provide insight into how to improve the prognosis of new patients. For example, a prediction that a patient will have a certain chance of survival at a particular time point may not provide information about possible clinical decisions to increase the chance of survival of the patient at that time point.
[0030] On the other hand, the medical history of previous patients, whose data and survival statistics are used to train a predictive machine learning model, can provide valuable insight into possible clinical decisions to improve the prognosis of new patients. For example, a machine learning model such as the RSF model can output a prediction of the probability that a new patient will survive from diagnosis to a particular time point. There may be a first group of patients (e.g., group A) whose survival probability over time is similar to the survival probability predicted by the model for the new patient, and a second group of patients (e.g., group B) whose survival probability is much lower than the survival probability predicted by the model for the new patient. If group A shares a common biomarker with the new patient, but group B does not have that biomarker, it may be determined that the biomarker is related to the survival probability of the new patient. A treatment decision can then be made to target that biomarker. However, as described above, while a predictive machine learning model may be useful for predicting a patient's prognosis based on the patient's attributes, the machine learning model typically does not identify other groups of patients whose medical outcomes are similar to the patient's prognosis. Besides providing clinical prediction results, the machine learning model typically does not provide additional information that can be used to improve the patient's prognosis.
[0031] Disclosed herein is a technique for facilitating clinical decisions for new patients based on identifying a group of patients (hereinafter, "similar patient pool") that have similar attributes to the new patient. A predictive machine learning model is provided to perform clinical predictions for new patients who are alive and whose future survival is unknown. The similarity-based patient pool can be identified from a group of previous patients whose data and survival statistics are used to train the predictive machine learning model. At least some of the attributes of the similarity-based patient pool can be output to assist in clinical decisions. The attributes can include, for example, the patient's biographical data, the patient's one or more laboratory test results, the patient's biopsy image data, the patient's molecular biomarkers, the patient's tumor site, and the patient's tumor stage.
[0032] In some examples, the clinical decision support system can use a machine learning model to make clinical predictions for a new patient based on the attributes of the new patient. For example, a random survival forest (RSF) model can be used to predict the probability of a patient's survival as a function of time since diagnosis. In addition, the clinical decision support system can identify a similarity-based patient pool that has certain attributes similar to the attributes of the new patient. The similarity-based patient pool can be identified based on patients that share similar values to the new patient in a subset of attributes determined to be most relevant to the clinical prediction (e.g., the probability of survival at a particular time point since diagnosis) performed by the machine learning model. The clinical decision support system can output the attributes and medical outcomes of the similarity-based patient pool along with the attributes and clinical prediction results of the patient to facilitate clinical decisions for the patient. In some examples, the similarity-based patient pool can include patients whose attributes and survival statistics are included in the training data for training the machine learning model. In some examples, the similarity-based patient pool can also include patients whose data is not used to train the machine learning model.
[0033] Specifically, the clinical decision support system can receive first data corresponding to attributes of a new patient. The attributes can include various biographical information, such as the patient's age and gender. Each attribute can be represented as a feature, which can include one or more vectors for input into a machine learning model. In some examples, an attribute can be represented by multiple features. The attributes can further include the patient's history (e.g., which procedure the patient underwent), the patient's habits (e.g., whether the patient smokes), and categories of the patient's laboratory test results (e.g., white blood cell count, hemoglobin count, platelet count, hematocrit count, red blood cell count, creatinine count, lymphocyte count, protein, bilirubin, calcium, sodium, potassium, glucose measurements). The attributes may further represent measurements of various biomarkers for various cancer types, such as estrogen receptor (ER), progesterone receptor (PR), human epidermal growth factor receptor 2 (HER2), epidermal growth factor receptor (EGFR or HER1) for breast cancer, ALK (anaplastic lymphoma kinase) for lung cancer, KRAS gene for lung and colon cancer, BRAF gene for colon cancer, etc. The attribute data may be processed by or prior to input into the clinical decision support system to generate a number of features that include the attribute information in a format (e.g., a vector) that can be interpreted by a machine learning model.
[0034] The clinical decision support system may include a machine learning model that may be trained based on data from previous patients to perform clinical predictions for new patients. The prediction may be based on inputting attributes of the new patient into the machine learning model. For example, the machine learning model may include an RSF model that may output a predictive survival function as a clinical prediction based on the first data. The survival function may be used to determine the likelihood that the new patient will survive until a given time (e.g., 500 days, 1000 days, 1500 days, etc.) after the new patient is diagnosed with a medical condition (e.g., advanced stage cancer). Alternatively, a hazard function provides the risk of death as a function of time, given survival up to that point. Another example of a survival function is a cumulative hazard function (CHF), which provides the accumulation of risk as a function of time. The survival function over time may be used to generate a patient-specific survival plot for the new patient.
[0035] As part of the training operation, multiple feature importance metrics associated with the machine learning model may also be obtained, the feature importance metrics defining the relevance of each feature to clinical predictions (e.g., survival rates at a particular time point). In one example, out-of-bag (OOB) samples, including samples of training patient data not used to build the RSF model, may be input into the decision tree to calculate prediction errors, such as a concordance index (c-index). The feature values may then be sorted for those samples, and the prediction error of each decision tree may be calculated for the sorted values of the feature. The raw importance score of the feature may be calculated based on averaging the difference in prediction errors between trees for the sorted values. A higher raw importance score may indicate that the feature is more relevant to the predicted survival function, while a lower raw importance score may indicate that the feature is less relevant to the predicted survival function. At the end of the training operation, the features may be ranked based on their importance scores, with more relevant features being ranked higher.
[0036] Based on the attributes of the new patient, the clinical decision support system can identify a group of patients from the patient database that are similar to the new patient in their highest ranked features. This group can be referred to as a similarity-based patient pool. The first step in selecting patients to form a similarity-based patient pool is to calculate the similarity between the new patient and each patient in the database based on their highest ranked features. Patients to form the similarity-based patient pool are then selected based on certain criteria, two examples of which are as follows: In the first example, patients can be selected based on their similarity to the new patient that exceeds a threshold value; In the second example, patients in the database are ranked according to their similarity to the new patient, and a predetermined number of patients with the highest ranking are selected. Thus, the similarity-based patient pool can be considered to be similar to the patient not only in having a similar health condition to the patient, but also in the features most relevant to clinical prediction.
[0037] The clinical decision support system can then output the attributes and clinical prediction results of the new patient along with the attributes and medical outcomes of the similarity-based patient pool. This may help facilitate clinical decisions about the new patient. For example, the clinical decision support system can output a summary of the attributes of the similarity-based patient pool along with a comparison of attributes between the new patient and the similarity-based patient pool (particularly attributes corresponding to the highest ranked features). The clinical decision support output allows a clinician to examine relevant attributes and determine a course of action (e.g., treatment) to increase the chances of survival of the new patient.
[0038] As an illustrative example, a feature corresponding to a biomarker attribute (e.g., epidermal growth factor receptor (EGFR)) may be one of the highest ranked features of the RSF model. Assume that a new patient is EGFR positive and the clinical decision support system can output an EGFR positive result for a similarity-based patient pool. If the predicted survival function of the new patient is more similar to the predicted survival function of the EGFR positive patient than the predicted survival function of the EGFR negative patient from the similarity-based patient pool, it may be determined that a treatment targeting EGFR may be useful to improve the survival probability of the new patient.
[0039] The disclosed technology allows for the identification of similarity-based patient pools that not only have similar health conditions as the new patient, but are also similar in the attributes / conditions most relevant to clinical prediction. The relevance of attributes to clinical prediction makes it more likely that the medical history of patients in the similarity-based patient pool can provide insights into possible treatments that can improve the prognosis of the new patient. These insights can be supported by statistics and medical histories of a relatively large patient population. For example, certain biomarkers that are common between the similarity-based patient pool and the new patient can be studied to determine whether a treatment of interest can improve the probability of survival of the new patient.
[0040] I. Examples of clinical predictions and applications 1A and 1B show examples of clinical predictions that can be provided by embodiments of the present disclosure. While FIG. 1A shows a mechanism for predicting a patient's cumulative survival probability versus time since the diagnosis of cancer was made, FIG. 1B shows an example of the application of survival probability prediction. Referring to FIG. 1A, a chart 100 shows an example of a Kaplan-Meier (KM) plot providing a study of survival statistics in patients with a certain type of cancer (e.g., lung cancer). The patients may receive a certain treatment. The KM plot shows the change in cumulative survival probability for a group of patients versus time measured from when the patient was diagnosed with cancer. If the patient is receiving treatment, the KM plot also shows the cumulative survival probability of the patient depending on the treatment. As time passes, some patients may die, decreasing the survival probability. Some other patients may be dropped from the plot due to other events not related to the event being studied (e.g., transitioning to a different state, changing hospitals). The dropped events are represented by diagonal check marks in the KM plot. The length of each horizontal line represents the death-free period, and the survival estimate at a given time point represents the cumulative probability of surviving to that time point.
[0041] In FIG. 1A, chart 100 includes two KM plots of cumulative survival probabilities for different cohorts of patients, A and B (e.g., cohorts of patients with different characteristics, cohorts of patients receiving different treatments, etc.). From FIG. 1A, it can be seen that the median survival (the first time the cumulative probability of survival drops below 50%) is about 11 months for cohort A, whereas it is about 6.5 months for cohort B. For example, the probability that a patient in cohort A will survive at least 8 months is about 70% (0.7), whereas the probability that a patient in cohort B will survive at least 8 months is about 30% (0.3).
[0042] FIG. 1B illustrates an application example of patient survival prediction. As shown in FIG. 1B, data 102 of a patient 103 can be input into a clinical decision support tool 104 to generate a survival prediction 106. The data 102 can include various attributes, such as, for example, biographical data, historical data, biomarkers, laboratory test result data, and the like. The clinical decision support tool 104 can generate various information 108 to assist a clinician in administering care / treatment to the patient 103 based on the survival prediction 106. For example, to facilitate the care of the patient 103, the clinical decision support tool 104 can generate information 108 indicating, for example, the patient's life expectancy. The information 108 can facilitate discussions between the clinician and the patient 103 regarding the patient's prognosis, as well as the evaluation of treatment options, as well as the planning of the patient's life events. Two illustrative examples are provided below. If the clinical decision support tool 104 predicts that the patient 103 has a relatively high probability of still being alive after five years, the patient 103 may decide to undergo an aggressive treatment that is more physically demanding and has more severe side effects. However, if the clinical decision support tool 104 indicates that the patient 103 has a relatively low probability of surviving for five years, the patient 103 may decide to forgo the treatment or undergo an alternative treatment and plan their care and life events for the remainder of their life.
[0043] While survival predictions can provide useful information to patients and clinicians, the survival prediction result itself may not provide insight into how to improve the patient's prognosis. For example, a prediction that a patient 103 has a particular probability of surviving beyond a particular time point may not provide information regarding possible treatments to improve the patient's chances of survival at that time point.
[0044] Clinical decision-making is a complex task where clinicians must reason about a diagnosis or treatment plan. Clinicians aim to fit the best treatment based on their own education, research, and personal experience. They typically work in a patient-specific manner and do not have digital solutions at hand that can help them exploit the potential of medical knowledge derived from real-world data (RWD). On the other hand, increasing the amount of RWD brings the opportunity to supplement decision-making with evidence-based population information. Patient similarity is a fundamental building block to investigate the most and least effective treatments based on RWD of similar individuals with comparable health conditions.
[0045] FIG. 1C shows an example of clinical decision making based on RWD and clinical prediction results. FIG. 1C shows a chart 120 combining a KM plot 122 of a first group of patients (labeled as "Group A" in FIG. 1C), a KM plot 124 of a second group of patients (labeled as "Group B" in FIG. 1C), and a survival prediction result 126 of the patient 103. The survival prediction result 126 of the patient 103 may be a function of time in which the predicted cumulative survival probability decreases over time. As shown in FIG. 1C, the predicted cumulative survival probability function 126 of the patient 103 is more similar to the KM plot 122 of the group B than to the KM plot 124 of the group A.
[0046] Chart 130 shows an exemplary distribution of positive epidermal growth factor receptor (EGFR) among patient 103, group A patients (corresponding to KM plot 124), and group B patients (corresponding to KM plot 122). Since patient 103 (corresponding to predicted cumulative survival probability function 126) has a positive EGFR, the bar in chart 130 for patient 103 is 100%. Approximately 60% of patients in group A have a positive EGFR result, while less than 5% of patients in group B have a positive EGFR result (both results from chart 130). Note also that while cumulative survival curve 124 overlaps with curve 126, curve 122 is substantially lower.
[0047] From charts 120 and 130, it can be determined that patients in group A, who have a similar cumulative survival curve as patient 103, have an EGFR positivity rate of about 60%, as is evident from the similarity between KM plot 124 and predicted outcome 126. In contrast, group B, whose KM plot 122 indicates a much lower survival probability than patient 103's predicted outcome 126, has only a 5% EGFR positivity rate. This may suggest that the presence of EGFR may be an important factor in determining patient 103's survival probability. Further studies can then be conducted based on this observation, such as investigating treatments that target EGFR.
[0048] While such observations can be useful and provide insight into treatment options to improve the patient's 103 chances of survival, the observations typically cannot be made solely from the survival probability prediction 106. For example, the prediction results do not identify other patients with similar survival statistics. Nor do the prediction results identify other patients with similar health conditions as the patient 103.
[0049] II. Similarity-Based Patient Pooling Using Machine Learning Models FIG. 2A illustrates an example of a clinical decision support system 200 that performs clinical predictions for a patient and identifies a similarity-based patient pool (based on the patient attributes involved in the clinical prediction). As illustrated in FIG. 2A, the clinical decision support system 200 includes a clinical prediction module 202, a patient pool determination module 204, and a portal 205. The clinical prediction module 202 may include a machine learning prediction model 206. The clinical prediction module 202 may receive patient data 208 corresponding to a plurality of features of a patient 210, and may use the machine learning prediction model 206 to make a clinical prediction 212 based on the patient data 208 of the patient. The patient 210 may be a new patient. The features of the patient data 208 may represent various attributes of the patient 210, including, for example, biographical data 208a, historical data 208b, biomarkers 208c, laboratory test result data 208d, and the like. The clinical prediction 212 may include, for example, a probability of survival for the patient. The probability of survival may indicate the likelihood that a patient will survive a given amount of time (e.g., 500 days, 1000 days, 1500 days, etc.) after the patient is diagnosed with a medical condition (e.g., advanced stage cancer). The clinical decision support system 200 may be a software system executing on a computer system, such as computer system 10 of FIG.
[0050] Additionally, the patient pool determination module 204 can be coupled to a patient database 214 that stores patient data for a set of patients. As described below, the patient data in the patient database 214 can be used to train the machine learning predictive model 206. The patient pool determination module 204 can identify a pool of patients and their patient data 216 from the patient database 214 that have similar attributes to the patient 210. The patient pool determination module 204 can identify a pool of patients based on these patients that are similar to the patient 210 in a subset of attributes that are most relevant to the clinical prediction performed by the machine learning predictive model 206. The clinical decision support system can then retrieve the patient data 216 corresponding to the pool of patients from the patient database 214. The portal 205 can perform additional processing of the patient data 216 (e.g., comparing the patient data 216 of the patient pool with the patient data 208 of the patient 210).
[0051] 2B illustrates a table 220 that provides examples of attributes included in the biographical data 208a, the historical data 208b, the biomarkers 208c, and the laboratory test result data 208d. For example, the biographical data 208a can include various categories of information, such as age, sex, and race. The historical data 208b can include various categories of information, such as diagnosis results including cancer stage, tissue structure, the Charlson Comorbidity Index (CCI), which predicts the risk of death based on the presence of certain comorbid conditions, the Eastern Cooperative Oncology Group (ECOG) score, which represents the patient's functional level in terms of ability to care for themselves, daily activities, and physical abilities. The historical data 208b can also include other information, such as the patient's habits (e.g., whether the patient smokes). The laboratory test results 208c can include laboratory test results for various categories of patients, such as measurements of white blood cell count, hemoglobin count, platelet count, hematocrit count, red blood cell count, creatinine count, lymphocyte count, protein, bilirubin, calcium, sodium, potassium, alkaline phosphatase, carbon dioxide, monocytes, chloride, lactate dehydrogenase, glucose, etc. The biomarker data 208d can include measurements of various biomarkers for various cancer types, such as estrogen receptor (ER), progesterone receptor (PR), human epidermal growth factor receptor 2 (HER2), epidermal growth factor receptor (EGFR or HER1) for breast cancer, ALK (anaplastic lymphoma kinase) for lung cancer, KRAS gene for lung cancer and colon cancer, BRAF gene for colon cancer, etc. It is understood that other attributes of clinical data not shown in FIG. 2B, such as biopsy image data, may also be provided to the machine learning prediction model 206 to perform clinical predictions.
[0052] In table 220, each attribute can be represented by a continuous numeric feature, a binary feature (whose value can be 1 or 0), or several one-hot coded vectors indicating the value of one of a set of possible categories of the attribute. For example, age can be represented as a continuous numeric feature. As another example, an attribute corresponding to a test result of the biomarker ER can be one-hot coded. Such an attribute can be associated with the following data categories: positive biomarker result, negative biomarker result, invalid biomarker result, and no biomarker test. One-hot coding can generate four features, each corresponding to one of the above categories. For each patient, only one of the four features (the feature corresponding to the category of the attribute) takes the value 1, and the other three take the value 0. This is illustrated in the table below, where an example is given for four patients, each with a different category of the ER biomarker attribute. [Table 1]
[0053] A. Random Survival Forest The machine learning prediction model 206 of FIG. 2A can be implemented using various techniques, such as a random survival forest (RSF) model. FIG. 2C illustrates an example of an RSF model 230. As illustrated in FIG. 2C, the random survival forest model 230 can include multiple decision trees, including, for example, decision trees 232 and 234. Each decision tree can include multiple nodes, including a root node (e.g., root node 232a of decision tree 232 and root node 234a of decision tree 234) and child nodes (e.g., child nodes 232b, 232c, 232d, and 232e of decision tree 232 and child nodes 234b and 234c of decision tree 234). Each parent node (e.g., nodes 232a, 232b, and 234a) having child nodes can be associated with a predetermined classification criterion for classifying a patient into one of its child nodes. Child nodes that do not have a child node are terminal nodes, including nodes 232d and 232e (of decision tree 232) and nodes 234b and 234c (of decision tree 234). At each terminal node of each tree, a node survival rate is calculated. When used to predict the survival of a patient 210, the patient 210 is assigned to a terminal node of each tree based on the data 208 of the patient 210. For example, decision tree 232 may output a cumulative survival probability value 236, while decision tree 234 may output a cumulative survival probability value 238. The survival probability of the patient 210 may be calculated by averaging the node survival rates from each terminal node to which the patient is assigned. For example, an average survival probability value 240 may be calculated based on the average of the survival probability values 236, 238 and the survival probability values output by the other decision trees.
[0054] Each decision tree can be assigned to process a different subset of features. For example, as shown in FIG. 2C, patient data 242 is divided into a set of features {S0, S1, S2, S3, S4, . . . , S n2B or any other attribute described herein. Decision tree 232 may be assigned to process features S0 and S1, decision tree 234 may be assigned to process feature S2, while other decision trees may be assigned to process other feature subsets. A parent node of the decision tree may then compare a subset of patient data 242 corresponding to one or more of the assigned features to one or more thresholds to classify the patient 210 into one of its child nodes. For example, with reference to decision tree 232, root node 232a may classify the patient into child node 232b if the patient data for feature S0 exceeds threshold x0, or into terminal node 232c if not. Child node 232b may further classify the patient into one of terminal nodes 232d or 232c based on the patient's data for feature S1. Depending on which terminal node the patient is classified into based on features S0 and S1, decision tree 232 may output a cumulative survival probability of 10%, 20%, or 30%. Additionally, decision tree 234 may also output a cumulative survival probability of 50% or 90% depending on which terminal node the patient is classified into based on feature S2.
[0055] The RSF model 230 of FIG. 2C can be constructed to determine the cumulative survival probability from diagnosis to a predetermined time (e.g., 1 year, 3 years, 5 years, etc.). Multiple RSF models can be included in the machine learning prediction model 206. Referring again to FIG. 2A, the clinical prediction module 202 can receive as input the time 222 for which the cumulative survival probability is to be determined. The clinical prediction module 202 can then select the RSF model trained for the time 222 and calculate the cumulative survival probability up to the time 222.
[0056] B. Training movements A training operation can be performed to generate each decision tree in the RSF model, a subset of features assigned to each decision tree, classification criteria at each parent node of the decision tree, and output values at each terminal node. FIG. 2D shows an example of a training operation. The training operation can be performed by a training module 250 that can be part of the clinical decision support tool 200 (FIG. 2A) or can be external to the clinical decision support tool 200. The training operation can be performed using patient data of a large population of patients from the patient database 214. As described above, the RSF model can be trained to determine the cumulative survival probability from diagnosis to a predetermined time. The training data used to train a particular RSF model to determine the cumulative survival probability from diagnosis to a predetermined time (e.g., 1 year, 3 years, 5 years, etc.) can include deaths and censoring up to the predetermined time, and then all surviving patients are censored at the predetermined time. In some examples, the training data can include deaths and censoring up to an additional time (e.g., 5-year deaths and censoring data for an RSF model that outputs a cumulative survival probability up to 3 years).
[0057] Specifically, the patient database 214 may store the patient attributes shown in table 220. The training module 250 executes a process of randomly sampling the patient data 252 with replacement at the root node of each tree of the RSF model. The process of random sampling with replacement is commonly referred to as "bootstrapping" and, since all trees are combined / aggregated to form a random forest, the process is also referred to as "bagging". Each tree is also assigned a random subset of features. Then, as part of the training operation, the root node (and each subsequent parent node) may be split into child nodes in a recursive node splitting process. In the node splitting process, a node containing a subset of patients may be split into two child nodes based on a threshold on the subset of features. The features and their thresholds in each split are selected to maximize the difference in survival probability between the two child nodes (e.g., based on a log-rank test).
[0058] As an example, in training the decision tree 232 of FIG. 2C, a bootstrap sample of patient data may be split into two groups based on feature S0 and threshold x0 to maximize the difference in survival probability between the two groups, and it may be determined that selecting another feature (e.g., S1) or setting another threshold for feature S0 will reduce the difference in survival probability between the two groups, and child node 232a may be generated. The process may then be repeated for child node 232a to generate additional child nodes, for example, until a threshold minimum number of patients is reached in a particular child node. When the splitting process is stopped, all childless nodes may become terminal nodes. For example, in at least one of terminal nodes 232d and 232e, the number of patients reaches a threshold minimum number, and thus the root splitting operation stops at these nodes. The output at each of these terminal nodes may be calculated from the outcome data of the patients classified into that terminal node. The training operation may be repeated to generate decision trees for outputting survival probabilities at various times, such that the RSF model 230 may output a survival function that predicts the survival probability of a patient at various times.
[0059] 2E, the training module 250 may also determine a feature importance metric 260 associated with the machine learning predictive model 206. The feature importance metric 260 may be defined by examining the relevance of each feature and its impact on the error of the machine learning predictive model 206. The feature importance metric 260 may be determined for the probability of survival to a predetermined time (e.g., 3 years), and different feature importance metrics 260 may be determined for survival probability predictions by the machine learning predictive model 206 to various predetermined times (e.g., 3 years, 5 years, 7 years, etc.).
[0060] In one example, to determine the feature importance metric 260, the training module 250 can obtain a set of out-of-bag (OOB) samples of patient data 252 from the patient database 214. The OOB samples for each tree can include samples of patient data that are not included in the bootstrap samples for that tree in FIG. 2D. For these samples, the values of the feature can be sorted, and a prediction error rate 262 for each decision tree from processing the OOB samples with sorted values of the feature can be obtained. A raw importance score 264 can be calculated for the feature based on, for example, averaging the differences in the prediction error rate output by each decision tree. The process can be repeated for each feature to calculate an individual raw importance score 264 for each feature. A high raw importance score can indicate that the feature is more relevant to survival prediction, while a low raw importance score can indicate that the feature is less relevant. At the end of the training operation, the features can be ranked based on their importance scores, with more relevant features being ranked higher.
[0061] In one example, the training module 250 can calculate the predicted error rate 262 based on calculating a concordance index (c-index). The concordance index can be calculated for the OOB sample based on performing a pairwise comparison of the model's estimate of the cumulative hazard function (CHF) and the actual time of death between the patients in the OOB sample. For each pair of patients, if the relative survival probability of the pair at a given time point matches the time order of the deaths of the pair, the pair is concordant, otherwise the pair is discordant. For example, if the CHF estimate of the first patient in the pair is higher than the CHF estimate of the second patient in the pair and the first patient died before the second patient, the pair is concordant. Otherwise, the pair is discordant. The c-index can be calculated based on the following formula:
number
[0062] The prediction error rate can then be calculated as the inverse of the C-index. Because the survival probability can change over time, the prediction error rate, as well as the resulting raw importance score, can also change over time. Thus, as shown in FIG. 2E, the feature importance metric 260 can include different raw importance scores 264 for different times 266.
[0063] C. Similarity-Based Patient Pooling Based on the feature importance metric 260 and the patient data 208 of the patient 210, the patient pool determination module 204 can identify a pool of patients from the patient database 214 that have similar attributes to the patient 210 and their patient data 216. Figure 2F illustrates example internal components and their operation of the patient pool determination module 204. As shown in Figure 2F, the patient pool determination module 204 includes a feature weight selection module 270 and a similarity determination module 272.
[0064] The feature weight selection module 270 may rank the features by feature importance value 260 and select the x features with the highest feature importance values (x may be a predetermined number, e.g., 20, or may be based on a rule, e.g., all features whose importance value is greater than the average importance value of all features). The set of top x features may be represented as E. The feature weight selection module 270 may then fit an RSF using only the features in E and recalculate the feature importance values for these features from this new RSF. For feature k, w k These new raw feature importance values, called , are scaled according to the following formula:
number
[0065] The scaled feature importance value w k is used as a weight in the similarity determination module 272.
[0066] The similarity determination module 272 can then identify patients in the patient database 214 that are similar to the patient 210 based on the scaled feature importance values / weights 274. The similarity determination module 272 can determine the similarity between two patients x i and x J Weighted aggregate similarity s(x i ,x J ) can be determined.
number
[0067] In Equation 3, s ijk Patient x i and x J represents the similarity between feature k and w k is the scaled feature importance value 274. More important features may have a similarity associated with a larger weight. If feature k is represented by a binary or one-hot coded vector, the similarity s ijk can be 1 if both patient features k are 1 or if the one-hot coding vectors match perfectly, otherwise, s ijk can take the value 0. Furthermore, if feature k is in the range R k If we take the value from ijk can be calculated based on the following formula:
number
[0068] The similarity determination module 272 determines the weighted aggregate similarity s(x i ,x J) and select a similarity-based patient pool based on the weighted aggregate similarity. The similarity-based patient pool may be considered to be similar to the patient in features most relevant to clinical prediction as well as having similar health conditions to the patient. In one example, the similarity determination module 272 may select a similarity-based patient pool based on a similarity to the patient 210 calculated according to Equation 3 or Equation 4 that exceeds a similarity threshold 280. In another example, the similarity determination module 272 may select a predetermined number of patients, determined based on a pool size threshold 282, that are most similar to the patient 210 to be part of the similarity-based patient pool.
[0069] The similarity determination module 272 can then obtain the attributes and medical outcomes of the similarity-based patient pool and output them as part of the patient data 216 to facilitate clinical decisions about the patient 210. For example, referring back to FIG. 1C, the similarity determination module 272 can identify similarity-based patient pools and patient data 216 from which the KM curve 124 can be generated. The portal 205 can perform a comparison between the patient data 216 of the patient pool, the patient data 208, and the patient data of the training set of patients in the patient database 214 for each feature present in both sets of data, and output the comparison results. From the comparison, the clinician may determine that the EGFR positivity of the patient pool is much higher than that of the training set of patients, and may further investigate EFGR (e.g., treatment targeting EFGR) to improve the patient survival rate.
[0070] III. Method Figure 3 illustrates an example of a method 300 for facilitating clinical decisions. The method 300 may be performed, for example, by the clinical decision support tool 200 of Figure 2A.
[0071] In step 302, the clinical decision support tool can receive first data corresponding to a plurality of features of a first patient (e.g., a new patient), each feature representing an attribute of a plurality of attributes. The first data can be entered via a computer interface, such as portal 205, or can be entered directly from a patient database, such as patient database 214. The first patient can be a new patient, such as patient 210.
[0072] Referring to FIG. 2A, the first data may include patient data 208 corresponding to attributes of the new patient. The attributes may include various biographical information such as the patient's age and gender. Each attribute may be represented as one or more features, and each feature may be represented as a vector for input into the machine learning model. The attributes may further include the patient's history (e.g., which procedure the patient underwent), the patient's habits (e.g., whether the patient smokes), and the patient's laboratory test result categories (e.g., white blood cell count, hemoglobin count, platelet count, hematocrit count, red blood cell count, creatinine count, lymphocyte count, protein, bilirubin, calcium, sodium, potassium, glucose measurements). The attributes may further indicate measurements of various biomarkers for different cancer types, such as estrogen receptor (ER), progesterone receptor (PR), human epidermal growth factor receptor 2 (HER2), epidermal growth factor receptor (EGFR or HER1) for breast cancer, ALK (anaplastic lymphoma kinase) for lung cancer, KRAS gene for lung and colon cancer, BRAF gene for colon cancer, etc.
[0073] In step 304, the clinical decision support tool can input the first data into a machine learning model to generate an outcome of a clinical prediction for the first patient, the machine learning model being associated with a plurality of feature importance metrics, the plurality of feature importance metrics defining the relevance of each of the plurality of features to the clinical prediction.
[0074] With reference to FIG. 2A, the clinical decision support tool may include a machine learning prediction model 206 for making a clinical prediction 212 based on the patient's data. The machine learning model 206 may include an RSF model 230 of FIG. 2C that may output a predictive survival function as a clinical prediction based on the patient's data. The clinical prediction 212 may include, for example, a probability of survival for the patient. The probability of survival may indicate the likelihood that the patient will survive until a predetermined time (e.g., 500 days, 1000 days, 1500 days, etc.) after the patient is diagnosed with a medical condition (e.g., advanced stage cancer). As described above, in some examples, the machine learning prediction model 206 may include multiple RSF models configured to predict the probability of survival until various predetermined times. One of the RSF models may be selected based on an input time to predict the probability of survival until the input time.
[0075] Further, the machine learning predictive model 206 is also associated with a number of feature importance metrics, such as feature importance metric 260. Referring to FIG. 2E, the feature importance metric 260 can be defined by examining the relevance of each feature by examining its impact on the error of the machine learning predictive model 206, and can be determined by the training module 250 based on a set of out-of-bag (OOB) samples of patient data 252 from the patient database 214. The OOB samples can include samples of patient data that are not involved in the bagging process used to build the RSF model 230. For these samples, the values of the feature can be sorted, and a prediction error rate of each decision tree from processing the OOB samples with sorted values of the feature can be obtained. The prediction error rate can be calculated based on a concordance index (c-index) based on Equation 1 above. A raw importance score can be calculated for the feature, for example, based on averaging the difference between the prediction error rate output by each decision tree. The process can be repeated for each feature to calculate an individual raw importance score for each feature. A high raw importance score may indicate that the feature is more relevant to survival prediction, while a low raw importance score may indicate that the feature is less relevant. At the end of the training run, the features may be ranked based on their importance scores, with more relevant features being ranked higher.
[0076] In step 306, the clinical decision support tool can obtain second data corresponding to the plurality of features of each of the group of patients based on a similarity in at least some of the plurality of features between the first patient and the group of patients, the similarity being based on the first data, the second data, and the plurality of feature importance metrics.
[0077] Specifically, the second data can be obtained by the patient pool determination module 204 of the clinical decision support tool 200, which includes a feature weight selection module 270 and a similarity determination module 272. Referring to FIG. 2F, the feature weight selection module 270 can rank the features by feature importance value 260 and select the x features with the highest feature importance value (x can be a predetermined number, e.g., 20, or can be based on a rule, e.g., all features whose importance value is greater than the average importance value of all features). The set of the top x features can be represented as E. The feature weight selection module 270 can then fit an RSF using only the features in E and recalculate the feature importance values of these features from this new RSF. The similarity between the first patient and other patients can be calculated based on Equation 3 and Equation 4 above. In a first example, a patient can be selected based on a similarity to the new patient that exceeds a threshold. In a second example, patients in the database are ranked according to their similarity to the first patient, and a predetermined number of patients with the highest rank are selected.
[0078] In step 308, the clinical decision support tool may generate content based at least in part on the outcome of the clinical prediction and the second data.
[0079] Specifically, in some examples, the content may include output summary statistics (e.g., median survival) for a patient pool (group of patients), KM curves for the patient pool, etc. In some examples, a comparison can be made between patient data for a group of patients, a first patient, and a training set of patients (e.g., patients represented in a training dataset for training a machine learning model) to generate a comparison result.
[0080] At step 310, the clinical decision support tool can output the content such that a clinical decision can be made for the first patient based on the content. For example, referring back to FIG. 1C, the content may indicate that EGFR positivity in the patient pool is much higher than EGFR positivity in the training set of patients, which may warrant further investigation of EFGR (e.g., treatments targeting EFGR) to improve patient survival.
[0081] IV. Computer Systems Any of the computer systems referred to herein may utilize any suitable number of subsystems. An example of such a subsystem is shown in computer system 10 in FIG. 4. In some embodiments, the computer system includes a single computer device, and the subsystems may be components of the computer device. In other embodiments, the computer system may include multiple computer devices, each of which is a subsystem having internal components. The computer system may include desktop and laptop computers, tablets, mobile phones, and other mobile devices. In some embodiments, cloud infrastructure (e.g., Amazon Web Services), graphical processing units (GPUs), and the like may be used to implement the disclosed techniques.
[0082] The subsystems shown in FIG. 4 are interconnected via a system bus 75. Additional subsystems are shown, such as a printer 74, a keyboard 78, a storage device 79, and a monitor 76 coupled to a display adapter 82. Peripherals and input / output (I / O) devices coupled to the I / O controller 71 can be connected to the computer system by any of several means known in the art, such as an input / output (I / O) port 77 (e.g., USB, FireWire). For example, the I / O port 77 or an external interface 81 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect the computer system 10 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via the system bus 75 not only allows for the exchange of information between the subsystems, but also allows the central processor 73 to communicate with each subsystem and control the execution of a plurality of instructions from the system memory 72 or storage device 79 (e.g., a fixed disk such as a hard drive or an optical disk). The system memory 72 and / or storage device 79 may embody a computer-readable medium. Another subsystem is a data collection device 85, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and to a user.
[0083] A computer system may include multiple identical components or subsystems connected to each other, for example, by external interfaces 81 or internal interfaces. In some embodiments, computer systems, subsystems, or devices may communicate over a network. In such cases, one computer may be considered a client and another computer a server, each of which may be part of the same computer system. Each of the clients and servers may include multiple systems, subsystems, or components.
[0084] Aspects of the embodiments may be implemented in the form of control logic, in a modular or integrated manner, using hardware (e.g., application specific integrated circuits or field programmable gate arrays), and / or using computer software in conjunction with a generally programmable processor. As used herein, a processor includes a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, those skilled in the art will know and understand other ways and / or methods for implementing embodiments of the present invention using hardware and combinations of hardware and software.
[0085] Any software components or functions described in this application may be implemented as software code executed by a processor using any suitable computer language, such as, for example, Java, C, C++, C#, Objective-C, or Swift, or a scripting language, such as, for example, Perl or Python, using conventional or object-oriented techniques. The software code may be stored on a computer-readable medium for storage and / or transmission as a series of instructions or commands. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as hard drives or floppy disks, or optical media such as compact disks (CDs) or DVDs (digital versatile disks), flash memory, and the like. The computer-readable medium may be any combination of such storage or transmission devices.
[0086] Further, such programs may be encoded and transmitted using carrier signals adapted for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Thus, computer-readable media may be generated using data signals encoded with such programs. Computer-readable media encoded with program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer-readable media may be located on or within a single computer product (e.g., a hard drive, CD, or an entire computer system), or may be present on or within different computer products in a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results described herein.
[0087] V. Working Examples Figure 5 illustrates one example of how patient data from a patient pool 500 can be used. Building patient data from a patient pool includes cohort building, and any suitable type of cohort building method can be used, such as the similarity-based patient pooling methods described herein. For example, the method of generating patient data from a similarity-based patient pool described in connection with Figures 2A and 2F can be used in the examples described herein. Other methods of cohort building may also be suitable.
[0088] 5, patient data from the patient pool 500 may be accessed, processed, and / or used by a disease history information tool 502 to automatically extract useful data regarding the patient's disease history from the patient's electronic health record and / or other patient databases. In some embodiments, the disease history information tool 502 may include a patient care information extraction module 504, a patient health status extraction module 506, and an additional patient treatments and services module 508. Various embodiments of the disease history information tool 502 may include any combination of one or more of the modules described herein.
[0089] In some embodiments, the patient care information extraction module 504 includes algorithms to extract information about the sequence of how patients are cared for in a cohort. This extracted information can be displayed to a user to facilitate and enable the user to learn from the disease course in a particular cohort. This extracted information can also be utilized in risk factor analysis in the cohort.
[0090] In some embodiments, the patient health extraction module 506 can include algorithms that extract information about the patient's health from the patient data. For example, some measures of patient health include patient-reported outcomes or experiences, such as reported symptoms, disorders, aspects of well-being, health perceptions, etc. In some embodiments, the user can be provided with more tailored views that can be used to view the patient's health at both an individual level and groups of patients in a cohort. For example, the user can be provided with population statistics regarding how patients in a specified cohort reported on a treatment.
[0091] In some embodiments, the additional patient treatments and services module 508 can include algorithms for extracting information regarding non-medical additional services (i.e., non-medicinal and non-drug services), such as, for example, rehabilitation, psychotherapy, physical therapy, and occupational therapy. In some embodiments, this extracted information can be used to determine which additional non-medical interventions, such as, for example, rehabilitation clinics, psychotherapy, physical therapy, and / or occupational therapy, may have benefited a particular patient cohort.
[0092] Figure 6 illustrates another example of how patient data from a patient pool 600 can be used. As shown in Figure 6, patient data from the patient pool 600 can be accessed, processed, and used by a recommendation tool 602 to suggest common terms for populating data fields of an electronic form or record. In some embodiments, the recommendation tool 602 can include a common electronic medical record (EMR) terminology extraction module 604 and a common diagnostic test extraction module 606. Various embodiments of the recommendation tool 602 can include any combination of one or more of the modules described herein.
[0093] In some embodiments, the common EMR term extraction module 604 may include an algorithm that extracts common terms used in data fields in the EMR system and then uses these extracted common terms as recommendations to the user filling out the EMR. For example, in some embodiments, one or more data fields in the EMR that the user is filling out may be automatically filled in with a pre-selected text field according to the highest frequency of the extracted common terms in the cohort. In some embodiments, instead, a sorted list of common terms may be provided to the user when the text field is selected, the list being sorted based on the frequency with which the term was found in the cohort. In some embodiments, the common EMR terms may be extracted from EMRs from a cohort patient pool. In other embodiments, the common EMR terms may be extracted from a broader EMR dataset formed from a larger patient pool. In some embodiments, the common EMR terms may be extracted from one or more EMR datasets.
[0094] In some embodiments, the common diagnostic test extraction module 606 can include an algorithm to extract common diagnostic tests for a diagnostic test recommendation system. In some embodiments, the data set used to extract the common diagnostic tests can be limited to data from a cohort patient pool. In some embodiments, the method of generating patient data from similarity-based patient pools described in connection with Figures 2A and 2F can be used to build a cohort with similar characteristics to a particular patient. The recommendation system identifies diagnostic tests that have been performed on this cohort. The extracted information can be used by the system to recommend those tests to be considered for a particular patient.
[0095] In some embodiments, the algorithms used by the data extraction modules described herein may be, but are not limited to, process mining algorithms, deep learning algorithms, and sequence alignment methods.
[0096] Any of the methods described herein may be fully or partially performed in a computer system including one or more processors that may be configured to perform the steps. Thus, the embodiments may be directed to a computer system configured to perform the steps of any of the methods described herein, possibly including different components performing each step or each group of steps. Although the steps of the methods herein are presented as numbered steps, they may be performed simultaneously or in a different order. In addition, some of these steps may be used with some of other steps from other methods. Also, all or some of the steps may be optional. Furthermore, any steps of any of the methods may be performed by a module, unit, circuit, or other means for performing these steps.
[0097] The specific details of the particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the invention, however, other embodiments of the invention may be directed to particular embodiments relating to each individual aspect, or particular combinations of these individual aspects.
[0098] The above description of the exemplary embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form described, and many modifications and variations are possible in light of the above teaching.
[0099] The terms "a," "an," or "the" are intended to mean "one or more," unless specifically indicated otherwise. The use of "or" is intended to mean "inclusive or," and is not intended to mean "exclusive or," unless specifically indicated otherwise. Reference to a "first" element does not necessarily require that a second element be provided. Furthermore, reference to a "first" or "second" element does not limit the referenced element to a particular location, unless expressly specified.
[0100] All patents, patent applications, publications, and descriptions mentioned herein are incorporated by reference in their entirety for all purposes. None are admitted to be prior art.
Claims
1. 1. A computer-implemented method for facilitating clinical decisions, comprising: receiving first data corresponding to a plurality of characteristics of a first patient, each characteristic representing an attribute of a plurality of attributes; inputting the first data into a machine learning model to generate a clinical prediction result for the first patient, the machine learning model having a plurality of feature importance metrics associated with it, the plurality of feature importance metrics defining a relevance of each of the plurality of features to the clinical prediction; and obtaining second data corresponding to the plurality of features of each of the group of patients based on a similarity of at least some of the plurality of features between the first patient and a group of patients, the similarity being based on the first data, the second data, and the plurality of feature importance metrics; generating content based at least in part on the results of the clinical prediction and the second data; outputting the content so that a clinical decision can be made about the first patient based on the content; A method comprising:
2. 10. The method of claim 1, wherein the plurality of attributes comprises at least one of patient biographical data, results of one or more laboratory tests of the first patient, biopsy image data of the first patient, molecular biomarkers of the first patient, tumor site of the first patient, or stage of the tumor of the first patient.
3. The method of claim 1 , wherein the plurality of attributes comprises one or more attributes representing measurements of biomarkers for different cancer types.
4. 10. The method of claim 1, wherein the clinical prediction comprises at least one of a probability of survival of the first patient at a predetermined time from when the first patient was diagnosed with the tumor, a survival time of the first patient from when the first patient was diagnosed with the tumor, or an outcome upon receiving treatment.
5. the machine learning model comprises a random forest survival model, the random forest survival model comprising f decision trees each configured to process a subset of the first subset of data to generate a cumulative survival probability; The method of claim 4 , wherein the survival rate of the patient at the predetermined time is determined based on an average of the cumulative survival probabilities output by a plurality of decision trees.
6. the group of patients is a first group of patients; the first group of patients is selected from a second group of patients; The method of claim 1 , wherein the machine learning model is trained based on patient data of a second group of patients.
7. ranking the plurality of features based on the relevance of each feature of the plurality of features to the clinical prediction; determining a subset of the plurality of features based on the ranking; and determining the first group of patients based on the similarity in the subset of the plurality of features between the first patient and the first group of patients; The method of claim 6 further comprising:
8. 8. The method of claim 7, wherein the first group of patients is selected from the second group of patients based on the similarity in the subset of the plurality of features between the first patient and the first group of patients exceeding a threshold.
9. 8. The method of claim 7, wherein the first group of patients is selected from the second group of patients based on selecting a threshold number of patients that have the highest similarity to the first patient on the subset of the plurality of features.
10. calculating a weighted aggregate similarity based on summing the scaled similarities for each feature of the at least some of the plurality of features, each similarity being scaled by a weight based on the relevance of the feature; identifying the group of patients based on the weighted aggregate similarity between the first patient and each of the group of patients; The method of claim 1 further comprising:
11. the feature importance metric for a feature is determined based on a relationship between an error in the outcome of the clinical prediction produced by the machine learning model for a second patient in the first group of patients; the clinical prediction outcome is generated from a plurality of values of the characteristic of the second patient; The method of claim 1 , wherein the error is calculated based on a comparison of the result of the clinical prediction with an actual clinical outcome of the second patient.
12. 7. The method of claim 6, wherein the content includes at least one of a median survival time for the first group of patients or a Kaplan-Meier survival curve for the first group of patients.
13. 7. The method of claim 6, wherein the content includes values of one or more of a first subset of the plurality of features for the first patient, a first group of the patients, and a second group of the patients.
14. A computer product comprising a computer readable medium having stored thereon a plurality of instructions for controlling a computer system to perform the operations of the method of any one of claims 1 to 13.
15. A computer product according to claim 14; one or more processors for executing instructions stored on the computer-readable medium; A system comprising:
16. A system comprising means for carrying out the method according to any one of claims 1 to 13.
17. A system configured to carry out the method of any one of claims 1 to 13.
18. A system comprising modules for respectively performing the steps of the method according to any one of claims 1 to 13.