Systems, methods, and apparatuses for structured electronic health record imputation using diagnostic temporal window
By employing time diagnosis windows and artificial intelligence to fill in missing values in EHRs, the system addresses the challenge of matching patients with clinical trials, improving data relevance and computational efficiency.
Patent Information
- Application Number
- JP2024220615
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-30
AI Technical Summary
Electronic health records (EHRs) often contain missing values that hinder the matching of patients with eligible clinical trials, particularly for cancer treatment trials where structured EHR data lacks detailed information like cancer stage.
The system uses a combination of time diagnosis windows and artificial intelligence to complement missing values in EHRs. It determines pre-diagnosis and post-diagnosis time windows, selects relevant health observations within these windows, and uses a trained artificial intelligence engine to generate complementary values.
This approach enhances the relevance of structured EHR data, reduces computational resources required for data processing, and improves the functionality of computers by accurately filling in missing values, thereby facilitating better patient-clinical trial matching.
Smart Images

Figure 2025097311000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to filling in missing values in electronic health records (EHRs), and more specifically, to using one or more time diagnosis windows and artificial intelligence to fill in values.
Background Art
[0002] Electronic health records (EHRs) often contain information related to patient care from various information sources. Often, EHRs contain unstructured data. Unstructured data does not have a predefined data model and can include information in many different formats, many of which cannot be processed or analyzed using conventional methods.
[0003] Medical notes are an example of unstructured data. Medical notes record the interactions between healthcare providers such as physicians, nurses, physician assistants, technicians, radiologists, etc. and patients and can be stored in EHRs. Medical notes may include consultation notes, referral notes, subjective information, objective information, assessment, plan (SOAP) notes, procedure notes, telephone notes, etc. Medical notes are often handwritten and may use non-standard formats. Data such as X-ray images, mammograms, digital pathology files, etc. are often presented as unstructured data in EHRs.
[0004] Unstructured data can contain detailed information about a patient's examination, but factors including inconsistent formats, quality, and data types can make it difficult to systematically interpret unstructured data and extract information.
[0005] Electronic health records may also contain structured data. In contrast to unstructured data, structured data is organized into fields of a database or other organizational schema. Examples of structured data include demographic information about a subject such as the subject's address, age, weight, height, etc. Diagnostic codes, billing codes, and clinical test results are often presented as structured data in electronic health records.
[0006] Clinical trials can be used to test the effectiveness of a treatment in a clinical setting. To participate in a clinical trial, a patient needs to meet the eligibility criteria. The eligibility criteria can include inclusion criteria, exclusion criteria, or a combination thereof. The terms "inclusion criteria" and "exclusion criteria" can sometimes be used interchangeably, and the exclusion criteria can be stated from the perspective of the inclusion criteria. For example, a clinical trial can state that the exclusion criteria are subjects under 18 years of age. This can be stated as an inclusion criteria, and when a patient has an age of 18 years or older, the inclusion criteria are met. Thus, the term "inclusion criteria" as used herein can include exclusion criteria described as inclusion criteria, and vice versa, the term "inclusion criteria" can include exclusion criteria described as inclusion criteria, and vice versa.
[0007] For example, matching a subject (patient) to a clinical trial can be difficult using unstructured or structured data in an EHR system and conventional methods. For example, many clinical trials for cancer treatment have eligibility requirements to identify a specific stage of cancer. Structured electronic health record data rarely describes the stage of cancer. Thus, since one or more eligibility criteria cannot be verified with EHR system data, subjects who may actually be eligible for a cancer treatment clinical trial cannot be matched to the trial. In some cases, even though the EHR system data is missing, subjects who may be eligible may not be aware that they may be eligible or that their eligibility can be confirmed by diagnostic tests. SUMMARY OF THE INVENTION
[0008] Systems, methods, and articles are provided for complementing values associated with a subject within an EHR system. In some cases, a clinical trial is selected and the subject meets one or more eligibility criteria of the clinical trial. In these cases, missing values of data associated with the subject from the EHR system are determined. The missing values are related to eligibility criteria of the clinical trial that the subject does not meet. The values to be complemented may be relevant to making a risk determination.
[0009] Receive a request to complement a value associated with a subject at a diagnosis time instance. According to some embodiments in which a clinical trial is selected, the missing value is the value to be complemented. Next, a subset of data associated with the subject from the EHR system is obtained. The subset of data is organized into specific fields as part of a schema and includes multiple health observations. Each health observation is associated with a time instance. In some cases, a second subset of data is obtained and each second health observation within the second subset is not associated with a time instance. The health observations may be selected based on their relevance to a tumor-node-metastasis (TNM) staging system, such as the extent of the tumor, the degree of lymph node metastasis, and the presence or absence of metastasis. The health observations may also be selected based on the variance of values corresponding to features in the health observations across health observations in the EHR system having the features. In some cases, health observations having features with a variance exceeding a predetermined threshold across the entire EHR system are selected.
[0010] According to some embodiments, for each health observation associated with a time instance, a pre-diagnosis time window and a post-diagnosis time window are determined. The pre-diagnosis time window includes a specified first duration immediately preceding the diagnosis time instance, and the post-diagnosis time window includes a specified second duration immediately following the diagnosis time instance. The pre-diagnosis time window and / or the post-diagnosis time window may be based on characteristics of the state or survival rate associated with complementary values. The pre-diagnosis time window and / or the post-diagnosis time window may also be based on diagnosis window hyperparameters. Next, a health observation having a time instance within the pre-diagnosis time window or the post-diagnosis time window is selected.
[0011] For each of the selected health observations, a value corresponding to the selected health observation is obtained. These obtained values are input values. Values corresponding to each second health observation may also be obtained and added to the input values. A subset of the input values may be associated with the same feature. In some cases, a subset is determined and a representative value is calculated based on the subset. Next, the subset is removed from the input values and the representative value is added to the input values.
[0012] A time sub-window may be calculated. The time sub-window may include a part of the pre-diagnosis time window, a part of the post-diagnosis time window, or both. A bucket of values is determined, and the bucket of values includes each input value within the subset corresponding to a health observation having a time instance within the time sub-window. Next, the count of the values within the bucket of values is determined. The bucket of values is removed from the input values and the count is added to the input values.
[0013] In some cases, the above steps are stored on a non-transitory computer-readable medium and executed by a processor of a computing system. In some cases, the above steps are stored in a non-transitory processor-readable storage medium.
[0014] The input value is provided to a trained artificial intelligence engine. The trained artificial intelligence engine processes the input value to generate a complementary value and provides the complementary value in response to a request.
[0015] According to some embodiments where a clinical trial is selected, the complementary value is used to automatically perform a preliminary matching of the subject with the clinical trial based on the complemented value. In some cases, the complementary value is related to the eligibility criteria of the clinical trial. Next, an appropriate diagnostic test for confirming the preliminary matching is determined. A message is sent to the subject, and the message is based on the preliminary matching and the appropriate diagnostic test for confirming the preliminary matching.
[0016] In some cases, the cancer stage is determined based on the complementary value. The cancer stage may conform to the TNM staging system that evaluates the extent of the tumor, the degree of spread to lymph nodes, and the presence or absence of metastasis. In some cases, the complementary value is the cancer stage. The risk determination may also be based on the complementary value.
[0017] A request to complement a second value associated with the subject at a second diagnostic time instance may be received. Next, the complementary value is added to the input value, which is again provided to the trained artificial intelligence engine. The trained artificial intelligence engine processes the input value to generate a second complementary value and provides the second complementary value.
[0018] Embodiments of the present disclosure improve the functionality of a computer by solving the technical problem of data gaps in structured EHR data. As described above, structured EHR data may include only basic information related to claims, lab test results, and the like. Since structured EHR data is structured, it requires lower computational costs for searching and analysis than unstructured EHR data. Embodiments of the present disclosure provide for complementing values useful for making clinical decisions from structured EHR data. This reduces the need to search for information such as cancer stage values within unstructured EHR data, which may require heterogeneous and expensive data processing techniques to extract useful information from this data. Thus, computing resources can be deployed for other tasks, improving the functionality of the computer.
[0019] Furthermore, the use of a diagnostic time window according to embodiments of the present disclosure increases the relevance of the structured EHR data used to train an artificial intelligence engine. Based on the characteristics of the states associated with the values to be complemented, by removing health observations having time instances outside one or more time windows, the trained artificial intelligence engine is not trained using health observations that have low correlation with the values to be complemented. Thus, the computing resources required to train the trained artificial intelligence engine are reduced, and computing resources are freed up for other tasks. Thereby, the performance of the computer implementing embodiments of the present disclosure is improved.
Brief Description of the Drawings
[0020] Like numbered elements may refer to common components in different drawings.
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 2D
Figure 3A
Figure 3B
Figure 3C
Figure 4
Figure 5
Figure 6
Figure 7
[0021] FIG. 1 shows a system 100 for complementing values associated with a subject within an electronic health record (EHR) system, according to some embodiments of the present disclosure. Data providers A 104, data provider B 106, data provider C 108, and data provider D 110 (data providers) supply data to the EHR data 102. The data providers can be hospitals, healthcare providers, pharmacies, etc. As described above, the EHR data may be structured or unstructured and may include medical notes, scans, billing codes, diagnostic codes, test results, demographic information, etc. The EHR data 102 is aggregated EHR data from the data providers.
[0022] To complement values associated with a subject within the EHR system, the EHR data 102 is processed into a patient cohort 120. The patient cohort 120 includes a plurality of patient feature vectors corresponding to a plurality of patients within the EHR data 102. Each patient feature vector includes one or more features, and each patient feature vector within a particular patient cohort 120 may include the same features. The patient feature vectors are described in detail with respect to FIG. 3C. To create the patient cohort 120, the EHR data 102 is processed into the patient cohort 120 in a cohort selection 112 that may include one or more of a filtering unit 114, an aggregation unit 116, or a conversion unit 118.
[0023] The filtering unit 114 includes calculating and implementing one or more time windows. For example, a pre-diagnosis time window and a post-diagnosis time window can be determined for a diagnosis time instance received with a request to complement values. The pre-diagnosis time window includes a specified duration immediately preceding the diagnosis time instance, and the post-diagnosis time window includes a specified second duration immediately following the diagnosis time instance. Calculating and implementing one or more time windows is described in detail with respect to FIGS. 2A - 2D.
[0024] According to some embodiments, the filtering unit 114 includes removing incorrect observations from the EHR data. Health observations with missing values, incorrect values, and other defects can be deleted from the EHR data. The filtering unit 114 may also include selecting each patient having a feature in the EHR data 102 corresponding to the feature of the value to be complemented. For example, if the value to be complemented is the cancer stage according to the TNM staging system, each patient having a feature corresponding to the cancer stage according to the TNM staging system can be selected.
[0025] According to some embodiments, the filtering unit 114 applies one or more filtering methods including Pearson correlation, linear discriminant analysis, analysis of variance, chi-square, or similar methods to select features having a desired degree of correlation with the value to be complemented from the EHR data 102. According to some embodiments, the filtering is additionally or alternatively performed by a machine learning method implemented by the value completion module 122, for example when lasso regression or ridge regression is used.
[0026] As will be described with respect to FIGS. 3A - 3C, aggregator 116 may include compiling observation results from EHR data 102 into a patient feature vector. For example, multiple health observations corresponding to a patient may be placed within the patient feature vector for use by value completion module 122. According to some embodiments, aggregator 116 includes calculating a representative value based on a subset of EHR data 102 corresponding to a patient. A patient may have health observations that include multiple values for the same feature at different times. For example, a patient may have results of multiple medical visits over a period of time and may have multiple features corresponding to the cancer stage. Depending on the time instance at which the cancer stage is evaluated, it may be stage 0, stage 1, stage 2, stage 3, or stage 4. Thus, the EHR data corresponding to a patient may include several different cancer stage values. However, according to some embodiments, only one cancer stage value is required. According to some embodiments, the representative value is the most recent cancer stage value or is calculated by taking the most severe cancer stage from the EHR data corresponding to the patient. According to some embodiments, the representative value is the maximum, minimum, average, mode, median, or weighted average of the values corresponding to the features within the EHR data corresponding to the patient. The representative value may be calculated for any set of health observations corresponding to the same feature.
[0027] Furthermore, the representative value can be calculated using a plurality of values corresponding to features from the EHR data 102. The values can be replaced with the percentile or quantile of the values as compared to one or more values corresponding to the same feature within the dataset. For example, the value corresponding to the feature of the number of white blood cells per microliter of blood can be 2000. If 2000 white blood cells per microliter of blood corresponds to the 5th percentile as compared to other values corresponding to that feature, the value 2000 can be replaced with 5. The plurality of values corresponding to features from the EHR data 102 can include all values within the set of values corresponding to the features within the EHR data 102, or any subset thereof. The representative value can be calculated based on any suitable data summarization method applied to the plurality of values corresponding to features from the EHR data 102, including temporal relevance, frequency tables, tail ratios, quantiles, order statistics, interquartile range, decile range, mean absolute deviation, quantile skewness, skewness, variance, standard deviation, or any other moment or quantile-based summary statistic.
[0028] The conversion unit 118 may include converting the observed values from the EHR data 102 for data type compatibility. The EHR data 102 may include non-numerical features such as character strings, which may not be compatible with the value completion module 122 that may require numerical values. For example, the value corresponding to the feature of cancer stage may appear as the character string "Stage 3" within the EHR data 102. Thus, for the value completion module 122 to use the value, the character string "Stage 3" may be converted to the integer 3 or encoded into a standardized representation of the structured EHR.
[0029] According to some embodiments, the patient cohort 120 includes a plurality of patient feature vectors. According to some embodiments, the patient cohort 120 may include a set of training examples. In one example, each training example of the set of training examples may include an input-output pair such as a pair including an input vector and a target response. The input vector may be a patient feature vector, and the target response may be a value of a feature associated with the patient.
[0030] The value completion module 122 may implement a supervised machine learning algorithm using the patient cohort 120 as training data. Supervised machine learning refers to a machine learning method that uses labeled training data to train or generate a set of machine learning models or mapping functions that map input feature vectors to output predicted responses. The trained machine learning model can then be utilized to map new input feature vectors to predicted responses. Supervised machine learning can be used to solve regression and classification problems. Regression associates a dependent variable, such as a value to be completed, with one or more independent variables, such as values within a subject feature vector. Regression algorithms include polynomial regression and logistic regression. A classification problem is one where the output predicted response includes a label or identification of a particular class. Classification algorithms may include support vector machines, decision trees, k-nearest neighbors, and random forest algorithms. In some cases, the support vector machine algorithm may determine a hyperplane or decision boundary that maximizes the distance between data points of two different classes. The hyperplane can separate data points of two different classes, and the margin between the set of data points closest to the hyperplane, or support vectors, can be determined to maximize the distance between data points of two different classes.
[0031] During the training phase, the machine learning model can be trained to generate predicted responses using a set of labeled training data. According to some embodiments, the set of labeled training data is selected training data. The selected training data can be created by a method that includes deletion of irrelevant or redundant data, correction of errors, filling in missing data, formatting of data, or any combination thereof. For example, in the case of cancer stage completion, the cancer stage corresponding to one or more feature vectors can be extracted from clinical notes.
[0032] Training data can be stored in memory. According to some embodiments, labeled data such as patient cohort 120 can be split into a training data set and an evaluation data set before or during the training phase.
[0033] The value completion module 122 can implement a machine learning algorithm that uses the patient cohort 120 to train a machine learning model, predicts features associated with the patient cohort 120 such as cancer stage, and evaluates the predictive ability of the machine learning model trained using the evaluation data set. The predictive performance of the trained machine learning model can be determined by comparing the predicted responses generated by the trained machine learning model with the target responses or ground truth values within the evaluation data set. According to some embodiments, the machine learning algorithm can include a loss function and an optimization technique. The loss function can quantify the penalty that occurs when the predicted response generated by the machine learning model is not equal to the appropriate target response. The optimization technique can attempt to minimize the quantified loss. An example of an appropriate optimization technique is online stochastic gradient descent.
[0034] The value completion module 122 includes an artificial intelligence engine that classifies input features within the patient cohort 120 into multiple classes. One or more machine learning models included in the value completion module 122 can be utilized to perform binary classification, in which an input feature vector is assigned to one of two classes. According to some embodiments, multiclass classification is used, in which an input feature vector is assigned to one of three or more classes. The output of the binary classification can include a prediction score indicating the probability that the input feature vector belongs to a particular class. According to some embodiments, a binary classifier can be a function used to determine whether an input feature vector, such as a vector of numbers representing input features, should be assigned to a first class or a second class. The binary classifier can use a classification algorithm that outputs a predicted value based on a linear prediction function that combines the input feature vector and a set of weights. For example, the classification algorithm can calculate a scalar product between the input feature vector and a weight vector, and then, if the scalar product exceeds a threshold, assign the input feature vector to the first class.
[0035] The number of input features or input values of a labeled dataset can be referred to as its dimensionality. According to some embodiments, dimensionality reduction can be used to reduce the number of input features used to train a machine learning model. Dimensionality reduction can be performed using one or more of the filtering module 114, the aggregation module 116, or the transformation module 118. Dimensionality reduction reduces the dimensionality of the feature space by selecting a subset of the most relevant features from the original set of input features.
[0036] Feature extraction may also be used. Feature extraction reduces the dimensionality of the feature space by deriving a new feature subspace from the original set of input features. In feature extraction, the new features may be different from the input features of the original set of input features and may retain most of the relevant information from combinations of the original set of input features. In one example, feature selection may be performed using sequential backward selection, and unsupervised feature extraction may be performed using principal component analysis.
[0037] The artificial intelligence engine may be trained using one or more training or learning algorithms. For example, backpropagation (error backpropagation) may be used to train a multi-layer neural network. According to some embodiments, a supervised training technique using a set of labeled training data may be performed. According to some embodiments, the artificial intelligence engine is trained with an unsupervised training technique using a set of unlabeled training data. One or more generalization techniques may be used to improve the generalization ability of the trained machine learning model, such as weight decay and dropout regularization.
[0038] The value completion module 122 receives an input 124 that specifies the value to be completed at the time of diagnosis. Using the above method, the value completion module 122 processes the patient cohort 120 to generate a completion value 126. The completion value 126 is added to the EHR data in some embodiments and may be used to complete other values.
[0039] Figures 2A, 2B, 2C, and 2D show pre-diagnosis time windows and post-diagnosis time windows according to various embodiments of the present disclosure. As discussed herein, one or more time windows may be used to add temporal sensitivity to value completion. In some cases, health observations at a time relatively distant from the time of diagnosis, such as health observations 20 years ago, may be available. However, for example, when complementing whether a person currently has influenza, considering health observations 20 years ago may be unhelpful and even harmful. These health observations may have little or no correlation with whether the person currently has influenza. In some cases, such as when complementing cancer stage, the time range of relevant health observations can be much larger. For example, health observations including the patient's family history, genetic test results, etc. may be relevant regardless of those time instances, while health observations including medications, treatments, procedures, tumor information, etc. may be relevant if there are time instances within a few years of the current cancer stage complement. Thus, the time window used may be based on the characteristics of the conditions associated with the value being complemented. The characteristics can be whether the disease has an acute, sub-acute, or chronic onset, the timeline of the progression of the disease state, etc. In some embodiments, the time window may be based on the availability of data within the time window.
[0040] Figure 2A shows timeline 200a, where the leftmost point of line 202 is the point on timeline 200a that is furthest in the past, and the rightmost point on line 202 is the instant at which the window is determined. Diagnostic time instance 204 is the instance of time at which the window is calculated. According to some embodiments, pre-diagnosis time window 206 includes a specified first duration immediately preceding diagnostic time instance 204. Post-diagnosis time window 208 includes a specified second duration immediately following diagnostic time instance 204. Health observations 210, 212, 214, and 216 are health observations of a patient having time instances corresponding to their respective positions on line 202. In the embodiment according to FIG. 2A, health observations 210 and 216 are outside windows 206 and 208, health observation 212 is within pre-diagnosis time window 206, and health observation 214 is within post-diagnosis time window 208. In this case, health observations 212 and 214 are selected for use in the input data, while health observations 210 and 216 are not selected for use.
[0041] According to some embodiments, the first duration, the second duration, or both are based on diagnostic window hyperparameters. The diagnostic window hyperparameters are, according to some embodiments, scalars corresponding to durations of time. Any suitable unit of time may be used. Thus, the hyperparameters of the diagnostic window can be represented as 5 for 5 years, 72 for 72 months, 31,536000 for 31,536000 seconds, and so on.
[0042] Figure 2B shows another timeline 200b corresponding to some embodiments of the diagnostic time window. The leftmost point of line 218 is the point on the timeline 200a that is farthest in the past, and the rightmost point on line 218 is the instant at which the diagnostic time window is determined. The diagnostic time instance 220 is the instance of time at which the window is calculated. According to some embodiments, the pre-diagnosis time window 222 includes a specified first duration immediately preceding the diagnostic time instance 220. The post-diagnosis time window 224 includes a specified second duration immediately following the diagnostic time instance 220. Health observations 226, 228, 230, 232, 234, and 236 are health observations of a patient having time instances corresponding to their respective positions on line 218. In some embodiments according to Figure 2B, health observations 226, 234, and 236 are outside the window, while health observations 228 and 230 are within the pre-diagnosis time window 222 and health observation 232 is within the post-diagnosis time window 224. In this embodiment, health observations 228, 230, and 232 are selected for use in the input data, while health observations 226, 234, and 236 are not. In some embodiments according to Figure 2B, the time period before diagnosis is longer than the time period after diagnosis.
[0043] Figure 2C shows a timeline 200c corresponding to some embodiments of the diagnostic time window. The leftmost point of line 238 is the point on the timeline 200c that is furthest in the past, and the rightmost point on line 238 is the instant at which the window is determined. The diagnostic time instance 240 is the instance of time at which the window is calculated. According to some embodiments, the pre-diagnosis time window 242 includes a specified first duration immediately prior to the diagnostic time instance 240. In the embodiment according to Figure 2C, since the diagnostic time instance 240 is at the rightmost point of the timeline, there is no post-diagnosis window, or equivalently, the post-diagnosis window has a zero duration. Health observations 244 and 246 are health observations of a patient having time instances corresponding to their respective positions on line 238. In the embodiment according to Figure 2C, health observation 244 is outside the pre-diagnosis time window 242, and health observation 246 is within the pre-diagnosis time window. In this case, health observation 246 is selected for use in the input data, but health observation 244 is not. Figure 2C shows only the pre-diagnosis time window, but in some embodiments, only the post-diagnosis time window is used. In other words, the post-diagnosis time window in this example has a zero duration.
[0044] Figure 2D shows a timeline 200d corresponding to some embodiments of a pre-diagnosis time window and a post-diagnosis time window having sub-windows. According to some embodiments shown in Figure 2D, each sub-window corresponds to a different feature that may appear in the patient characteristic vector. For example, one feature may be related to diagnostic scans during a period of 1 to 3 years immediately prior to the diagnostic time instance, while another feature may be related to diagnostic scans during a period of 1 to 3 years immediately after the diagnostic time instance. The use of sub-windows is described in detail with respect to the sub-window feature 306 of Figure 3A.
[0045] The leftmost point of line 248 is the point on the timeline that is farthest in the past, and the rightmost point of line 248 is the moment when the window is determined. The diagnostic time instance 250 is the instance of the time when the window is calculated. The first sub-window 254 is within the pre-diagnosis time window 255. The second sub-window 256 is within both the pre-diagnosis time window 255 and the post-diagnosis time window 257. The third sub-window 258 is within the post-diagnosis time window 257. Health observations 260, 262, 264, 266, 268, and 270 are health observations of a patient having time instances corresponding to their respective positions on line 248. In the embodiment according to FIG. 2D, health observations 260 and 270 are outside the sub-windows, health observation 262 is within the first sub-window 254, health observations 264 and 266 are within the second sub-window 256, and health observation 268 is within the third sub-window 258. In this case, the first feature is based on health observation 262, the second feature is based on health observations 264 and 266, and the third feature is based on health observation 268.
[0046] FIG. 3A shows a feature set generated using various time windows, according to some embodiments of the present disclosure. As described with respect to FIGS. 2A-2D, various time windows can be used to select health observations based on one or more features. The pre- and post-diagnosis window features 304 are examples of features resulting from the use of the pre-diagnosis time window and the post-diagnosis time window, as described with respect to FIGS. 2A-2B. According to some embodiments, pre_diag_num_of_encs is a feature corresponding to the number of encounters within the pre-diagnosis time window, and post_diag_num_of_encs is a feature corresponding to the number of encounters within the post-diagnosis time window. Similarly, pre_diag_num_of_enc_names is a feature corresponding to the number of encounter names within the pre-diagnosis time window, and post_diag_num_of_enc_names is a feature corresponding to the number of encounter names within the post-diagnosis time window.
[0047] Sub-window feature 306 shows an exemplary set of features generated according to some embodiments having sub-windows. Each feature corresponds to one of one or more sub-windows as described with respect to FIG. 2D. For example, pre_1 to 3 yr_diag_encs_preventative_medicine corresponds to the number of preventative medicine encounters in a sub-window encountered 1 to 3 years before the diagnosis time instance.
[0048] The selective diagnostic window feature 308 shows an exemplary set of features generated using embodiments where one or more features are not subject to filtering by a time window, while one or more other features are subject to filtering by a time window. For example, the feature pre_diag_num_of_encs is the number of encounters within the pre-diagnosis time window. However, the feature cnt_flag_medication_single_use is the number of medications in all available health observations of the subject. Any part of the feature can be filtered by a time window according to various embodiments.
[0049] FIG. 3B shows an exemplary set of subject health observations 310. Row 310a corresponds to the health observations of a patient with a patient_id 311 of 12345. Each health observation shown in FIG. 3B also includes features of enc_date 312, encounter_type 314, encounter_concept_canonical 316, and encounter_concept_name 318. Considering the diagnosis time instance of January 1, 2023, and a three-year pre-diagnosis time window, the set of subject health observations 310 can be processed for features corresponding to the number of encounters within the pre-diagnosis time window. Using the aforementioned diagnosis time instance and pre-diagnosis time window, since there are five encounters in the pre-diagnosis time window, the feature value is 5. Similarly, considering only hospital visits within the pre-diagnosis time window, the feature value is 2 because there are two hospital visits within the pre-diagnosis time window.
[0050] FIG. 3C shows an exemplary subject feature vector for a patient having a patient_id322 of 12345. The subject feature vector 320 includes features 324, 326, 328, and 330. Each of the features is within a pre-diagnosis time window and can be calculated in a manner similar to the example discussed with respect to FIG. 3B above.
[0051] FIG. 4 shows a flowchart illustrating a process for complementing values associated with a subject within an electronic health record (EHR) system, according to some embodiments of the present disclosure. According to some embodiments, process 400 may be executed using one or more physical or virtual machines and / or one or more containerized applications.
[0052] Process 400 begins at step 402, where a request to complement a value associated with a subject is received. The request may be provided manually by a user input or automatically. The value to be complemented can be a cancer stage, genetic change, genetic mutation, histology, disease attribute, status of a potential biomarker, etc. According to some embodiments, the request includes a subject feature vector similar to the patient feature vector 320 of FIG. 3C.
[0053] According to some embodiments, the value to be supplemented can be a value corresponding to an existing data field in the EHR system. For example, if there is a field corresponding to the age of the subject but the data is missing, the value to be supplemented can be the age of the subject. According to some embodiments, the value to be supplemented may be a value that exists for the subject in the EHR system. This can occur when the value associated with the subject is outside the range of typical values for the same characteristic in the EHR system. For example, the Eastern Cooperative Oncology Group (ECOG) performance status scale ranges from 0 to 5. A patient with a score of 0 on this scale is fully active and can perform all pre-disease functions, while a patient with a score of 5 on the scale is deceased. If a patient's ECOG score is recorded as 10, the recorded score is outside the range of the scale and is an error, so the value to be supplemented can be the ECOG score. In some embodiments, the value associated with the subject is compared to values corresponding to other subjects in the EHR system. Next, when the value associated with the subject is greater than a threshold, for example, 3 standard deviations from the mean of the corresponding values, the value associated with the subject is taken as the value to be supplemented. In various embodiments, any suitable statistical measure for comparing the value associated with the subject to values corresponding to other subjects may be used, and any suitable threshold may be used. In some embodiments, the value to be supplemented does not correspond to an existing field of the subject in the EHR system. For example, the value to be supplemented can be the presence of a biomarker for which there is no corresponding field for the subject in the EHR system. Generally, any value related to one or more values in the EHR system can be the value to be supplemented.
[0054] Process 400 proceeds from step 402 to step 404, where a subset of data associated with a subject including health observations is retrieved from the EHR system.
[0055] Following step 404, in step 406, a pre-diagnosis time window and a post-diagnosis time window are determined. Various embodiments of the time window are described in detail above with respect to FIGS. 2A-2D. The time window determines which health observations are used, and examples thereof can be seen in FIGS. 3A-3B.
[0056] Continuing from step 406, in step 408, health observations having time instances within the time window are selected. According to some embodiments, the time instances of the health observations are dates such as the dates in column 312 of FIG. 3B.
[0057] Process 400 proceeds from step 408 to step 410, and the values corresponding to each selected health observation are obtained as input values. As described with respect to FIGS. 3A-3C, a health observation according to one embodiment includes a patient ID, an encounter date, a standard encounter concept, and an encounter concept name. If the encounter date of the health observation is within the selected time window, the value corresponding to the health observation is used.
[0058] After step 410, in step 412, the input values are provided to a trained artificial intelligence engine. According to some embodiments, the trained artificial intelligence engine is a multi-output classifier. The trained artificial intelligence engine may use any known artificial intelligence model, including logistic regression or one or more models of other types. In embodiments where the artificial intelligence model includes logistic regression, the variations of logistic regression used may include multi-label, ordinal, ridge, lasso, or any other known variation of logistic regression. According to some embodiments, the trained artificial intelligence engine may include a random forest classifier. According to some embodiments, the trained artificial intelligence engine may include an artificial neural network, or any other backpropagation-based machine learning model.
[0059] The trained artificial intelligence engine is trained using data such as patient cohort 120 of FIG. 1. According to some embodiments, the trained artificial intelligence engine is trained using supervised learning. In these embodiments, the labels may be generated for the training data using unstructured EHR data.
[0060] After step 412, process 400 proceeds to step 414, where the trained artificial intelligence engine processes the input value to generate a complemented value. As described above with respect to FIG. 1, the trained artificial intelligence engine may implement any known artificial intelligence algorithm including polynomial regression, logistic regression, support vector machine, decision tree, k-nearest neighbor, random forest, and neural network. Generating the complemented value according to some embodiments includes predicting a plurality of values. For example, according to some embodiments of cancer stage complementation, the trained artificial intelligence engine predicts the likelihood of each stage of cancer according to a cancer stage classification system. Thus, variables corresponding to the likelihoods of stage 0, stage 1, stage 2, stage 3, and stage 4 are generated according to some cancer stage classification systems given an input subject feature vector such as subject feature vector 320 of FIG. 3C. The maximum value of these values is sometimes considered the complemented value.
[0061] As described above, any number of variables may be predicted when generating the complemented value. According to some embodiments, the variable may represent one or more cancer stages. For example, one predicted value may correspond to the likelihood of the presence of any of cancer stages 0, 1, and 2, while the other predicted value may correspond to the likelihood of the presence of cancer stages 3 and 4. Any combination of groupings may be used according to various embodiments such as a first variable corresponding to the likelihood of stage 0 and a second variable corresponding to the likelihoods of stages 1, 2, 3, and 4.
[0062] In step 416, the completed value is provided in response to a request. The response according to some embodiments includes providing the completed value to another process / computer. According to some embodiments, the response includes storing the completed value in an EHR system.
[0063] FIG. 5 shows a flowchart illustrating an embodiment that performs preliminary clinical trial matching using the completed value.
[0064] Process 500 starts at step 502, a clinical trial is selected, and the clinical trial has one or more eligibility criteria that are met by a subject. According to some embodiments, a clinical trial is selected and the electronic health record system is searched for subjects who meet one or more eligibility criteria. According to some embodiments, a patient is selected from the electronic health record system and a clinical trial is selected for which the patient meets one or more eligibility criteria of the clinical trial. According to some embodiments, both the subject and the clinical trial to be preliminarily matched are known. For example, a subject may provide a clinical trial that they would like to participate in but currently do not meet one or more eligibility criteria.
[0065] After step 502, in step 504, missing values in the electronic health record system related to clinical trial eligibility criteria not met by the subject are determined. According to some embodiments, the missing values directly correspond to the eligibility criteria. If the eligibility criterion is the stage of a particular cancer, for example, patients who meet or do not meet the eligibility criterion can be preliminarily determined by complementing the missing cancer stage value. According to some embodiments, the missing values may not directly correspond to the eligibility criteria, but may be used to support the finding that the subject meets or does not meet the eligibility criteria. For example, some eligibility criteria such as liver function may not directly correspond to the missing values, but the missing values may be useful in determining whether the subject meets the eligibility criteria. The subject's blood flow, epidemiological, environmental, and historical characteristics may be considered in determining whether the subject has normal liver function. However, certain markers of liver function such as a score according to a model of an end-stage liver disease scoring system may not be sufficient to evaluate whether the subject meets the eligibility criteria, but may be useful in the evaluation of liver function.
[0066] Following step 504, process 500 proceeds to step 506, where a value to complement the missing value is selected.
[0067] In step 508, method 400 as described with respect to FIG. 4 above is executed using the selected value to complement the missing value. Method 400 generates a complemented value as output.
[0068] After executing method 400 in step 508, in step 510, a preliminary matching of the subject with the clinical trial is performed based on the complemented value. As described above, the complemented value may directly correspond to the eligibility criteria or otherwise be related to the eligibility criteria. In some embodiments where the complemented value directly corresponds to the eligibility criteria, the complemented value is compared with the eligibility criteria. For example, if the eligibility criterion is cancer at stage 3 or stage 4, the complemented value is checked to determine whether it corresponds to cancer stage 3 or stage 4.
[0069] As described above, according to some embodiments, the complemented value correlates only with, or is related to, the eligibility criteria. In these cases, the relationship between the complemented value and the eligibility criteria can be used. For example, if the eligibility criteria is normal liver function, the bilirubin level can be the complemented value, and normal liver function is associated with bilirubin levels within the normal range. If the complemented value of the bilirubin level is within the normal range of bilirubin levels, normal liver function is indicated. The relationship between the complemented value and the eligibility criteria can be inferred according to some embodiments using the literature or characteristics of the complemented value compared to other measurements and / or complemented values of the same feature within the EHR system. For example, the complemented value can be compared to other measurements and / or complemented values of the same type within the EHR system. The complemented value can then be considered to indicate whether it meets, or does not meet, the eligibility criteria based on the percentile of the complemented value compared to other measurements and / or complemented values of the same type within the EHR system. For example, compared to other bilirubin values in the EHR system, a bilirubin value at the 99th percentile can indicate that the eligibility criteria for normal liver function are not met. Other metrics of comparison, such as the standard deviation of the complemented value from the average of other measurements and / or complemented values of the same type within the EHR system, can be used to indicate whether the eligibility criteria are met or not. The above description relates to complementing one value, but any number of values can be complemented in a similar manner.
[0070] In step 512, a diagnostic test is determined to confirm the preliminary matching. As described above, the complemented value is used to indicate whether the eligibility criteria are met. Diagnostic tests according to some embodiments directly measure the characteristics of the subject predicted by the complemented value to confirm or disprove whether the complemented value and the subject meet the eligibility criteria. According to some embodiments, the diagnostic test measures another characteristic of the subject related to the eligibility criteria.
[0071] In step 514, the message is sent to the subject based on the preliminary matching and diagnostic tests. According to some embodiments, access is made to the scheduling system of one or more healthcare providers to determine one or more available reservations for the diagnostic tests to be performed on the subject. Then, according to some embodiments, one of the one or more reservations may be preliminarily scheduled, and the subject may be notified about the preliminarily scheduled reservation. According to some embodiments, one or more reservations are included in the message, and the subject may select a reservation from the one or more reservations to schedule. According to some embodiments, in response to receiving the selection of a reservation from the subject, a second message is sent to the scheduling system of the healthcare provider to schedule the selected reservation.
[0072] For the sake of brevity and clarity, the above exemplary embodiments include one subject preliminarily matching one clinical trial based on one complementary value. However, any number of subjects can be preliminarily matched with any number of clinical trials in a similar manner using any number of complemented missing values. For example, a subject may be selected, then one or more clinical trials may be repeated, and each clinical trial for which the subject meets one or more eligibility criteria may be selected. According to some embodiments, a clinical trial may be selected, one or more patients may be repeated, and a clinical trial that meets one or more eligibility criteria for each subject may be selected. According to some embodiments, the preliminary matching of one subject and one clinical trial includes complementing multiple missing values.
[0073] FIG. 6 shows a flowchart illustrating a process for aggregating a subset of input values into representative values according to some embodiments. Process 600 begins at step 602, where each input value within the subset of input values is determined to be associated with the same data attribute. According to some embodiments, the data attribute is a feature such as the post_diag_immuno_check feature 326 shown in FIG. 3C. For example, a feature within a subject's feature vector may correspond to a cancer stage. However, a first input value may indicate cancer at stage 3, and a second input value may indicate cancer at stage 4. This can occur when the diagnostic window includes multiple observations related to the same feature. The subject may have received several different diagnoses of cancer stage within the diagnostic window. Values corresponding to multiple observations of the same feature may be used to calculate a representative value.
[0074] In step 604, a representative value is calculated based on the subset of input values. According to some embodiments where the feature of the subset of input values is cancer stage, the most severe cancer stage within the subset of input values is the representative value. According to some embodiments, the maximum or minimum input value within the subset is the representative value. According to some embodiments, the representative value is the median, mode, mean, or weighted mean of the input values within the subset. The representative value may also be a value within the subset corresponding to the health observation with the earliest or latest time instance.
[0075] Process 600 proceeds from step 604 to step 606, where the subset of input values is removed from the input values.
[0076] Proceeding from step 606 to step 608, the calculated representative value is added to the input values. According to some embodiments, the representative value is inserted into a field of the subject feature vector at a position corresponding to the feature.
[0077] According to some embodiments, one or more general-purpose or special-purpose computing systems or devices may be used to implement computing device 700. Further, according to some embodiments, computing device 700 may include one or more different computing systems or devices and may span distributed locations. Further, each block shown in FIG. 7 may represent one or more such blocks appropriate to a particular embodiment or may be combined with other blocks. Also, model-related manager 722 may be implemented in software, hardware, firmware, or some combination and may achieve the capabilities described herein.
[0078] As shown, computing device 700 includes a non-transitory computer memory ("memory") 701, a display 702 (including, but not limited to, a light emitting diode (LED) panel, a cathode ray tube (CRT) display, a liquid crystal display (LCD), a touch screen display, a projector, etc.), one or more central processing units ("CPUs") or other processors 703, input / output ("I / O") devices 704 (e.g., a keyboard, a mouse, an RF or infrared receiver, a universal serial bus (USB) port, a high definition multimedia interface (HDMI (registered trademark)) port, other communication ports, etc.), other computer-readable media 705, and a network connection portion 706. A model-related manager 722 is shown to be present in memory 701. In other embodiments, some or all of the components of the content and the model-related manager 722 may be stored on other computer-readable media 705 or transmitted via other computer-readable media 705. The components of the computing device 700 and the model-related manager 722 can be executed on one or more CPUs 703 and can implement the applicable functions described herein. In some embodiments, the model-related manager 722 may operate as, be a part of, or cooperate with other software applications stored in memory 701 or on various other computing devices. In some embodiments, the model-related manager 722 also facilitates communication with another device or system via the I / O device 704 or via the network connection portion 706.
[0079] One or more model-related modules 724 are configured to perform actions directly or indirectly related to an AI or other computational model. In some embodiments, the model-related modules 724 store, retrieve, or otherwise access at least some model-related data on some portion of the model-related data storage 716 or other data storage internal or external to the computing device 700.
[0080] Other code or programs 730 (e.g., additional data processing modules, program guide manager modules, web servers, etc.), and other data repositories such as the data repository 720 for potentially storing other data, may also be present in the memory 701 and may be executable on one or more CPUs 703. Notably, one or more of the components in FIG. 7 may or may not be present in any particular embodiment. For example, some embodiments may not provide other computer-readable media 705 or a display 702.
[0081] According to some embodiments, computing device 700 and model-related manager 722 include an API that provides programmatic access and adds, deletes, or modifies one or more functions of computing device 700. In some embodiments, the components / modules of computing device 700 and model-related manager 722 are implemented using standard programming techniques. For example, model-related manager 722 may be implemented as an executable program that runs on CPU 703, along with one or more static or dynamic libraries. In other embodiments, computing device 700 and model-related manager 722 may be implemented as instructions to be processed by a virtual machine that runs as one of the other programs 730. In general, the range of programming languages known in the art includes representative embodiments of various programming language paradigms, including, but not limited to, object-oriented (e.g., Java, C++, C#, Visual Basic.NET, Smalltalk, etc.), functional (e.g., ML, Lisp, Scheme, etc.), procedural (e.g., C, Pascal, Ada, Modula, etc.), script (e.g., Perl, Ruby, Python, JavaScript, VBScript, etc.), or declarative (e.g., SQL, Prolog, etc.), and such exemplary embodiments may be used to implement such illustrative embodiments.
[0082] In software or firmware embodiments, the instructions stored in memory, when executed, configure one or more processors of computing device 700 to perform the functions of model-related manager 722. In some embodiments, the instructions cause some other processor, such as CPU 703 or an I / O controller / processor, to perform at least some of the functions described herein.
[0083] The above embodiments may also use well-known or other synchronous or asynchronous client-server computing techniques. However, the various components may be implemented using more monolithic programming techniques as well, for example, as executable files operating on a single CPU computer system, or may be decomposed using various structuring techniques known in the art including, but not limited to, multiprogramming, multithreading, client-server, or peer-to-peer operating on one or more computer systems each having one or more CPUs or other processors. Some embodiments may execute in parallel and asynchronously and communicate using message passing techniques. Equivalent synchronous embodiments are also supported by embodiments of the model related manager 722. Also, other functions may be implemented or executed by different components / modules in a different order, but still achieve the functions of the computing device 700 and the model related manager 722.
[0084] Furthermore, programming interfaces to data stored as part of the computing device 700 and the model related manager 722 may be available through standard mechanisms such as C, C++, C#, and Java APIs, libraries for accessing files, databases, or other data repositories, scripting languages such as XML, or web servers, FTP servers, NFS file servers, or other types of servers that provide access to stored data. The model related data storage 716 and the data repository 720 may be implemented as one or more database systems, file systems, or any other technology for storing such information, including embodiments that use distributed computing techniques, or any combination of the above.
[0085] Different configurations and locations of programs and data are contemplated for use with the techniques described herein. Various distributed computing techniques are suitable for implementing the components of embodiments illustrated in a distributed fashion including, but not limited to, TCP / IP sockets, RPC, RMI, HTTP, and web services (such as XML-RPC, JAX-RPC, SOAP, etc.). Other variations are possible. Other functions may also be provided by each component / module, or existing functions may be distributed among the components / modules in different ways and still achieve the functions of the model related manager 722.
[0086] Further, according to some embodiments, some or all of the components of computing device 700 and model related manager 722 may be implemented or provided in other ways such as firmware or hardware including, but not limited to, one or more application specific integrated circuits ("ASICs"), standard integrated circuits, controllers (e.g., including microcontrollers or embedded controllers by executing appropriate instructions), field programmable gate arrays ("FPGAs"), complex programmable logic devices ("CPLDs"), etc. Some or all of the system components or data structures may also be stored as content (e.g., as executable or other machine-readable software instructions or structured data) on a computer-readable medium (e.g., a hard disk, memory, computer network, cellular wireless network or other data transmission medium, or a portable media article readable by an appropriate drive such as a DVD or flash memory device, or via an appropriate connection) such that a computer-readable medium or one or more associated computing systems or devices execute, or otherwise use, or provide content for execution of at least some of the described techniques.
[0087] Also, several examples are disclosed herein in accordance with the following numbered clauses.
[0088] Clause 1. A method for complementing values associated with a subject within an electronic health record (EHR) system, the method comprising: receiving a request to complement a value associated with the subject at a diagnosis time instance; obtaining a subset of data associated with the subject from the EHR system, the subset of data being organized into specific fields as part of a schema and including a plurality of health observations, each health observation being associated with a time instance, and for each health observation, determining a pre-diagnosis time window (206) and a post-diagnosis time window (208), the pre-diagnosis time window (206) including a specified first duration immediately preceding the diagnosis time instance (204), the post-diagnosis time window (208) including a specified second duration immediately following the diagnosis time instance (204), selecting health observations having time instances within the pre-diagnosis time window or the post-diagnosis time window (212 214); for each of the selected health observations (212 214), obtaining the value corresponding to the selected health observation as an input value; providing the input value to a trained artificial intelligence engine (122), using the trained artificial intelligence engine (122) to process the input value in order for the trained artificial intelligence engine (122) to generate a complemented value; providing a complemented value (126) in response to the request; and providing to a trained artificial intelligence engine configured to perform actions, including.
[0089] Clause 2. The method according to clause 1, wherein the value to be complemented is a cancer stage, or a value related to making a risk determination, or a value related to eligibility criteria for a clinical trial.
[0090] The step of determining the pre-diagnosis time window is the method according to any one of Items 1 to 2, based on the characteristics of the conditions associated with the value to be complemented, or the survival rate of the conditions associated with health observations, or the diagnostic window hyperparameters.
[0091] The method according to any one of Items 1 to 3, wherein the plurality of health observations includes one or more health observations selected based on the degree of association with the tumor stage, the degree of spread to lymph nodes, or the presence of the metastasis (TNM) staging system.
[0092] The method according to any one of Items 1 to 4, further including the step of determining the cancer stage based on the complemented value, and the cancer stage conforms to the TNM staging system that evaluates the degree of the tumor, the degree of spread to lymph nodes, and the presence of metastasis.
[0093] The method according to any one of Items 1 to 5, further including the step of selecting health observations to be included in the data subset based on the variance of the values corresponding to the fields in the health observations calculated using the plurality of health observations in the EHR system having fields.
[0094] The method according to any one of Items 1 to 6, further including the step of receiving a request to complement a second value associated with the subject at a second diagnosis time instance, the step of adding the complemented value to the input value, and the step of providing the input value to a trained artificial intelligence engine, wherein the trained artificial intelligence engine is configured to execute an action and processes the input value to generate a second complemented value using the trained artificial intelligence engine, and the step of providing the second complemented value.
[0095] Step 8 of obtaining a second subset of data associated with the subject from the EHR system, wherein the second subset is compiled into specific fields as part of a schema and includes a plurality of second health observations, each second health observation not being associated with a time instance; the step of obtaining; for each second health observation within the second subset of data, obtaining the value corresponding to the selected health observation as a second input value; and the step of adding the second input value to the input value; further including the method according to any one of clauses 1 to 7.
[0096] Clause 9. The step of selecting a clinical trial for which the subject meets one or more eligibility criteria of the clinical trial; the step of determining missing values of data associated with the subject from the EHR system, wherein the missing values are related to the eligibility criteria of the clinical trial not met by the subject; the step of selecting a value to complement the missing values; the step of automatically performing a preliminary matching of the subject with the clinical trial based on the complemented values; the step of determining an appropriate diagnostic test for confirming the preliminary matching; and the step of sending a message to the subject, wherein the message is based on the preliminary matching and the appropriate diagnostic test for confirming the preliminary matching; further including the method according to any one of clauses 1 to 8.
[0097] Clause 10. The step of determining that each input value within the subset of input values is associated with the same data attribute; the step of calculating a representative value based on the subset of input values; the step of deleting the subset of input values from the input values; and the step of adding the representative value to the input values; further including the method according to any one of clauses 1 to 9.
[0098] Step 11. Determining that each input value within a subset of input values is associated with the same data attribute; calculating a time sub-window that includes a portion of a pre-diagnosis time window and a portion of a post-diagnosis time window; determining a bucket of values, where the bucket of values includes each input value within the subset corresponding to a health observation having a time instance within the time sub-window; determining a count of the values within the bucket of values; removing the input values within the bucket of values from the subset; and adding the count to the subset, the method according to any one of clauses 1 to 10 further comprising these steps.
[0099] Clause 12. Automatically performing a preliminary matching between a subject and a clinical trial, the method according to clause 1 or any one of clauses 3 to 11 further comprising automatically performing a preliminary matching where the complemented value is a value related to the eligibility criteria of the clinical trial.
[0100] Clause 13. Automatically performing a preliminary matching between a subject and a clinical trial based on the complemented value, the method according to any one of clauses 1 to 11 further comprising this step.
[0101] A computing system for complementing values associated with a subject within an electronic health record (EHR) system, the computing system comprising: one or more processors (703); and one or more non-transitory computer-readable media (705) collectively storing instructions that, when executed collectively by the one or more processors (703), cause the one or more processors (703) to perform actions, the actions including: receiving a request to complement a value associated with a subject at a diagnosis time instance; obtaining a subset of data associated with the subject from the EHR system, the subset of data being organized into specific fields as part of a schema and including a plurality of health observations, each health observation including obtaining a subset of data associated with a time instance; for each health observation, determining a diagnosis time window (256) including a specified duration including the diagnosis time instance (250); selecting health observations having time instances (264, 266) within the diagnosis time window; for each of the selected health observations (264, 266), obtaining a value corresponding to the selected health observation as an input value; providing the input value to a trained artificial intelligence engine (122), using the trained artificial intelligence engine (122) to process the input value such that the trained artificial intelligence engine (122) generates a complemented value; and providing the complemented value in response to the request, and providing to a trained artificial intelligence engine configured to perform the actions.
[0102] Article 15. A non - transitory processor - readable storage medium (705) that stores computer instructions which, when executed by a processor (703), cause the processor (703) to perform actions, where the actions include receiving a request to complement a value associated with a subject at a diagnosis time instance, obtaining a subset of data associated with the subject from an electronic health record (EHR) system, where the subset of data is organized into specific fields as part of a schema and includes multiple health observations, each health observation including obtaining a subset of data associated with a time instance, for each health observation, determining a diagnostic time window (256) that includes a specified duration including the diagnosis time instance (250), selecting health observations having time instances within the diagnostic time window (264, 266), for each of the selected health observations (264, 266), obtaining the value corresponding to the selected health observation as an input value, providing the input value to a trained artificial intelligence engine (122), using the trained artificial intelligence engine (122) to process the input value in order for the trained artificial intelligence engine (122) to generate a complemented value (126), and providing the complemented value (126) in response to the request. The non - transitory processor - readable storage medium is included.
[0103] Additional embodiments can be provided by combining the various embodiments described above. All of the U.S. patents, U.S. patent application publications, U.S. patent applications, foreign patents, foreign patent applications, and non - patent publications referred to herein and / or described in the application data sheet are hereby incorporated by reference in their entirety. Aspects of the embodiments can be modified, as needed, to adopt concepts from various patents, applications, and publications to provide further additional embodiments.
[0104] These and other modifications can be made to the embodiments with reference to the detailed description above. In general, in the following claims, the terms used should not be construed as limiting the claims to the specific embodiments disclosed in this specification and the claims, but rather the claims should be construed to include all possible embodiments together with the full scope of equivalents to which such claims are entitled. Accordingly, the claims are not limited by the present disclosure.
Claims
1. 1. A method for supplementing values associated with a subject in an electronic health record (EHR) system, comprising: receiving a request to complete a value associated with the subject with a diagnostic time instance; obtaining a subset of data associated with the subject from the EHR system, the subset of data being organized into specific fields as part of a schema and including a plurality of health observations, each health observation being associated with a time instance; determining, for each health observation, a pre-diagnosis time window and a post-diagnosis time window, the pre-diagnosis time window including a specified first duration immediately preceding the diagnosis time instance and the post-diagnosis time window including a specified second duration immediately following the diagnosis time instance; selecting a health observation having a time instance within the pre-diagnosis time window or the post-diagnosis time window; obtaining, for each of the selected health observations, a value corresponding to the selected health observation as an input value; providing the input values to a trained artificial intelligence engine, the trained artificial intelligence engine comprising: processing the input values using the trained artificial intelligence engine to generate the imputed values; providing the completed value in response to the request to the trained artificial intelligence engine configured to perform an action; A method comprising:
2. selecting a clinical trial in which the subject meets one or more eligibility criteria for the clinical trial; determining missing values in the data associated with the subject from the EHR system, the missing values relating to eligibility criteria for the clinical trial that the subject does not meet; selecting the missing value as the imputed value; automatically performing a preliminary matching of the subject to the clinical trial based on the imputed values; determining an appropriate diagnostic test to confirm said preliminary match; 10. The method of claim 1, further comprising: sending a message to the subject, the message confirming the preliminary match based on the preliminary match and the appropriate diagnostic test.
3. 2. The method of claim 1, further comprising automatically performing preliminary matching of the subject with a clinical trial based on the imputed value, the imputed value relating to eligibility criteria of the clinical trial.
4. The method of claim 1 , further comprising automatically performing preliminary matching of the subject to clinical trials based on the imputed values.
5. 2. The method of claim 1, further comprising determining a cancer stage based on the imputed value, the cancer stage conforming to a TNM staging system that assesses tumor extent, extent of spread to lymph nodes, and presence of metastases.
6. 2. The method of claim 1, wherein the imputed value is a cancer stage that conforms to the TNM staging system, which assesses tumor extent, extent of spread to lymph nodes, and the presence of metastases.
7. 10. The method of claim 1, wherein the plurality of health observations comprises one or more health observations selected based on their relevance to tumor extent, extent of spread to lymph nodes, and presence of metastases (TNM) staging system.
8. 2. The method of claim 1, wherein the imputed value is a cancer stage.
9. The imputed value is relevant to making a risk determination, and the method further comprises: The method of claim 1 , further comprising making the risk determination based on the imputed value.
10. determining that each input value in the subset of input values is associated with a same data attribute; calculating a representative value based on said subset of said input values; removing said subset of input values from said input values; The method of claim 1 , further comprising the step of: adding the representative value to the input value.
11. determining that each input value in the subset of input values is associated with a same data attribute; calculating a time sub-window including a portion of the pre-diagnosis time window and a portion of the post-diagnosis time window; determining a bucket of values, the bucket of values including each entry value in the subset that corresponds to a health observation having a time instance within the time subwindow; determining a count of the values within the bucket of values; removing the input value in the bucket of values from the subset; The method of claim 1 , further comprising the step of: adding the count to the subset.
12. receiving a request to complete a second value associated with the subject at a second diagnostic time instance; adding the interpolated value to the input value; providing the input values to the trained artificial intelligence engine, the trained artificial intelligence engine comprising: processing the input value using the trained artificial intelligence engine to generate the second imputed value; 2. The method of claim 1, further comprising: providing the second imputed value to the trained artificial intelligence engine configured to perform an action, the second imputed value being included in the trained artificial intelligence engine.
13. obtaining a second subset of data associated with the subject from the EHR system, the second subset of data being organized into specific fields as part of a schema and including a plurality of second health observations, each second health observation not associated with a time instance; for each second health observation in the second subset of data, obtaining a value corresponding to the selected health observation as a second input value; The method of claim 1 , further comprising the step of: adding the second input value to the input value.
14. 2. The method of claim 1, further comprising: selecting health observations to include in the subset of data based on a variance of values corresponding to a field in the health observations calculated using a plurality of health observations in the EHR system having the field.
15. 2. The method of claim 1, further comprising: selecting health observations to include in the subset of data based on a variance of values corresponding to a field of the health observation across a plurality of health observations in the EHR system having the field, wherein the variance of the selected health observations exceeds a predetermined threshold.
16. The method of claim 1 , wherein determining the pre-diagnostic time window is based on a characteristic of a condition associated with the value being imputed.
17. The method of claim 1 , wherein determining the pre-diagnostic time window is based on a survival rate of a condition associated with the health observation.
18. The method of claim 1 , wherein determining the pre-diagnostic time window is based on a diagnostic window hyperparameter.
19. 1. A computing system for supplementing values associated with a subject in an electronic health record (EHR) system, the computing system comprising: one or more processors; and one or more non-transitory computer-readable media that collectively store instructions that, when collectively executed by the one or more processors, cause the one or more processors to perform actions, the actions including: receiving a request to complete a value associated with the subject with a diagnostic time instance; Retrieving a subset of data associated with the subject from the EHR system, the subset of data being organized into specific fields as part of a schema and including a plurality of health observations, each health observation being associated with a time instance; determining, for each health observation, a diagnostic time window including a specified duration that includes said diagnostic time instance; selecting a health observation having a time instance within the diagnostic time window; For each of the selected health observations, obtaining a value corresponding to the selected health observation as an input value; providing the input values to a trained artificial intelligence engine, the trained artificial intelligence engine comprising: processing the input values using the trained artificial intelligence engine to generate the interpolated values; and providing the completed value in response to the request to the trained artificial intelligence engine, the trained artificial intelligence engine configured to perform an action, including: A computing system comprising:
20. 1. A non-transitory processor-readable storage medium storing computer instructions that, when executed by a processor, cause the processor to perform actions, the actions including: receiving a request to complete a value associated with the subject at a diagnosis time instance; Retrieving a subset of data associated with the subject from an electronic health record (EHR) system, the subset of data being organized into specific fields as part of a schema and including a plurality of health observations, each health observation being associated with a time instance; determining, for each health observation, a diagnostic time window including a specified duration that includes said diagnostic time instance; selecting a health observation having a time instance within the diagnostic time window; For each of the selected health observations, obtaining a value corresponding to the selected health observation as an input value; providing the input values to a trained artificial intelligence engine, the trained artificial intelligence engine comprising: processing the input values using the trained artificial intelligence engine to generate the interpolated values; and providing the completed value in response to the request to the trained artificial intelligence engine, the trained artificial intelligence engine configured to perform an action, including:
11. A non-transitory processor-readable storage medium comprising: