A tool for selecting relevant features in precision assessment.
The system addresses the issue of sub-optimal feature ranking in ML-based diagnostics by personalizing feature selection using model-based importance methodologies, enhancing diagnostic accuracy and reducing costs and time in urgent situations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-23
- Publication Date
- 2026-03-17
AI Technical Summary
Diagnostic systems based on machine learning algorithms often rank clinical features from a population perspective, which may not be optimal for individual patients, leading to sub-optimal data collection and undesirable outcomes, especially in urgent situations.
A system and method for ranking unmeasured features using model-based importance methodologies, supplementing features with values, evaluating outcomes, and determining statistical parameters to select an optimal patient-specific predictive dataset.
This approach reduces measurement costs and time by personalizing feature selection, improving diagnostic accuracy and efficiency, particularly in emergency care scenarios.
Smart Images

Figure 2026048950000001_ABST
Abstract
Description
Technical Field
[0004] , , , ,
[0003] 3>
[0001] This application claims the priority and benefit of U.S. Provisional Patent Application No. 62 / 959,754, filed on January 10, 2020, entitled "Tool for Selecting Relevant Features in Precision Diagnostics", and is hereby incorporated by reference in its entirety as if fully set forth below and for all applicable purposes.
[0002] The present disclosure generally relates to methods and means for the selection and collection of data in order to provide accurate and timely outcome predictions. More specifically, the present disclosure relates to methods and systems for providing educated suggestions for optimizing the cost and time for data collection using confidence levels enhanced on an individual basis.
Background Art
[0003] Diagnostic systems based on machine learning (ML) algorithms provide a ranking of the entire population of clinical features from the perspective of importance. However, when a set of features is collected for a particular patient, the ranking of the entire population of those features may not be optimal for that patient. As a result of collecting sub-optimal features to measure, it can lead to an undesirable outcome for the patient, especially in urgent situations. It is desirable to have a system and method that enables the selection of optimal features for completing a patient-specific predictive dataset. <In some embodiments of the present disclosure, a method for ranking unmeasured features for an instance in which at least one feature has been measured includes: supplementing the unmeasured features in the instance with a first value, while keeping the other remaining unmeasured features constant; and evaluating a first outcome using a model that uses the first value in the instance. The method also includes: supplementing the unmeasured features in a dataset with a second value, while keeping the other remaining unmeasured features constant; evaluating a second outcome using a model that uses the second value in the instance; and determining a statistical parameter using the first and second outcomes. The method also includes assigning the unmeasured features to a ranking corresponding to the determined statistical parameter.
[0005] In some embodiments, a system for ranking unmeasured features for an instance in which at least one feature has been measured includes memory for storing instructions and one or more processors communicatively coupled to the memory. One or more processors are configured to execute instructions to cause the system to: supplement unused features in an instance with first values, while keeping the other remaining unmeasured features constant; and evaluate a first outcome using a model that uses the first values in the instance. One or more processors are also configured to execute instructions to cause the system to: supplement unused features in an instance with second values, while keeping the other remaining unmeasured features constant; evaluate a second outcome using a model that uses the second values in the instance; and determine statistical parameters using the first and second outcomes. One or more processors are also configured to: assign unmeasured features to rankings corresponding to statistical parameters; and select a filtered dataset from a master dataset according to at least one measured feature from an instance, wherein the master dataset includes multiple datasets related to multiple known outcomes.
[0006] In some embodiments, a non-temporary computer-readable medium for storing instructions, which, when executed by a computer, cause the computer to perform a method for ranking unmeasured features for instances in which at least one feature has been measured. This method includes: supplementing the unused features in an instance with first values, while keeping the other remaining unmeasured features constant; and evaluating a first outcome using a model that uses the first values in the instance. This method also includes: supplementing the unmeasured features in an instance with second values, while keeping the other remaining unmeasured features constant; evaluating a second outcome using a model that uses the second values in the instance; and determining statistical parameters using the first and second outcomes. This method also includes: assigning the unmeasured features to a ranking corresponding to a statistical parameter; and selecting a filtered dataset from a master dataset according to at least one measured feature from an instance, wherein the master dataset includes multiple datasets related to multiple known outcomes. In this method, assigning unmeasured features to rankings corresponding to statistical parameters involves using a model-based feature importance methodology to identify the relative importance of unmeasured features using one or more known outcomes in a filtered dataset.
[0007] In some embodiments, a method for ranking unmeasured features for an instance in which at least one feature has been measured comprises selecting a filtered dataset from a master dataset according to at least one measured feature from the instance, the master dataset comprising multiple datasets associated with multiple known outcomes. The method also comprises, in the filtered dataset, using a model-based feature importance methodology to identify the relative importance of the unmeasured features using one or more known outcomes, and assigning the unmeasured features to a ranking corresponding to the output from the model-based feature importance.
[0008] In some embodiments, a method for ranking unmeasured features for instances in which at least one feature has been measured includes accessing a master dataset, the master dataset comprising multiple datasets related to known outcomes. The method also includes determining a variance value related to a model for the outcome, the model being based on the unmeasured feature and at least one other distinct feature in the dataset; evaluating the variance of predictions for the outcome using a model that uses multiple complement values for the unmeasured feature in the dataset; and assigning the unmeasured feature to a ranking according to the value of the variance of predictions relative to the variance value.
[0009] In some embodiments, a method for ranking unmeasured features for an instance in which at least one feature has been measured includes determining a rule for assessing a decision value based on a dataset. The rule includes values collected for multiple measured features and unmeasured features in the instance, and the rule matches (1) multiple known outcomes from a master dataset containing multiple datasets, and (2) one or more measured features. The method also includes determining the accuracy of the rule based on multiple outcome values and known outcomes for each of the datasets, and assigning a ranking to the unmeasured features corresponding to the accuracy of the rule.
[0010] In some embodiments, a method for determining the sampling frequency of a selected feature based on its predictability includes identifying a set of observed features and a set of missing features. This method also includes constructing a model to predict the sampling frequency of the selected feature using a selected feature matrix from a historical dataset, generating predictions about the sampling frequency using this model, and determining the variance of the selected feature from multiple time predictions. This method also includes ranking the selected feature relative to other features based on its variance, and increasing the sampling frequency of the selected feature when the feature's rank is in a predetermined upper percentile.
[0011] It will be understood that other configurations of the subject art will be readily apparent to those skilled in the art from the following detailed description, and various configurations of the subject art are shown and described as examples. As realized, the subject art is capable of other different configurations without deviating from the scope of the subject art, and some of its details are modifiable in various other ways. Therefore, the drawings and detailed description should be considered illustrative and not limiting. [Brief explanation of the drawing]
[0012] The accompanying drawings are included to provide further understanding, are incorporated herein, and constitute part of herein, illustrating embodiments of the disclosure and, together with the description, are helpful in illustrating the principles of the embodiments of the disclosure.
[0013] [Figure 1] This document illustrates exemplary architectures suitable for diagnostic engines in streaming data environments, based on various embodiments.
[0014] [Figure 2] This is a block diagram illustrating an exemplary server and client from the architecture of Figure 1, according to a predetermined aspect of this disclosure.
[0015] [Figure 3] Illustrate an exemplary workflow for a decision tree according to various embodiments.
[0016] [Figure 4] Illustrate a method for ranking one or more features in a dataset according to the relevance of a diagnostic engine using a constraint function according to various embodiments.
[0017] [Figure 5] It is a block diagram that illustrates a method for quantifying the impact of missing features in the uncertainty of predictions for a diagnostic engine according to various embodiments.
[0018] [Figure 6] It is a block diagram that illustrates a method for quantifying the relevance of features in a diagnostic engine that selects a similar patient dataset from a master dataset according to various embodiments.
[0019] [Figure 7] It is a block diagram that shows a method for quantifying the relevance of features in a diagnostic engine using a historical dataset selected from a master dataset according to various embodiments.
[0020] [Figure 8] It is a flowchart that illustrates steps in a method for selecting features relevant to a diagnostic engine based on multiple medical features received or imputed over a time sequence according to various embodiments.
[0021] [Figure 9] It is a flowchart that illustrates steps in a method for selecting features relevant to a diagnostic engine by quantifying the impact of the absence of individual features according to various embodiments.
[0022] [Figure 10]A flowchart illustrating steps in a method of selecting features related to a diagnostic engine based on filters for similar patient populations from a master dataset, according to various embodiments.
[0023] [Figure 11] A flowchart illustrating steps in a method of selecting features related to a diagnostic engine based on a model for measured features, according to various embodiments.
[0024] [Figure 12] A flowchart illustrating steps in a method of selecting features related to a diagnostic engine based on a historical dataset selected from a master dataset, according to various embodiments.
[0025] [Figure 13] A flowchart illustrating steps in a method of constructing a multivariate model for predicting the importance of missing features using measured features, according to various embodiments.
[0026] [Figure 14] A flowchart illustrating steps in a method of determining sampling frequencies for selected features based on the predictability of features, according to various embodiments.
[0027] [Figure 15] A block diagram illustrating an exemplary computer system in which the clients and servers of FIGS. & and the methods of FIGS. 8 - 12 may be implemented, according to various embodiments.
[0028] In the figures, elements and steps shown by the same or similar reference numerals are associated with the same or similar elements and steps unless otherwise indicated.
Best Mode for Carrying Out the Invention
[0029] The following detailed description includes numerous specific details to provide a complete understanding of the disclosure. However, as will be apparent to those skilled in the art, embodiments of the disclosure can be implemented without some of these specific details. In other examples, well-known structures and techniques are not shown in detail so as not to obscure the disclosure.
[0030] overview Recently, the number of measurable features about patients has increased dramatically. In various embodiments, feature measurement may include genomics, transcriptomics, proteomics, metabolomics, wearable device data, behavioral data (food / beverage purchases, fitness data, etc.), billing data (insurance, etc.), and social media data. Different features may have different associated costs and acquisition times. Therefore, it is desirable to have a personalized ranking of features.
[0031] Therefore, it is desirable to understand the relevance of specific features in the diagnosis of a given patient to reduce the cost and time of measurement, which can be extremely important in emergency care situations. The relevance of any given feature may also depend on the situation and the patient themselves. For example, a 70-year-old patient with fever, leukocytosis, and a history of type 2 diabetes may benefit most from subsequent measurements of features a, b, and c. On the other hand, a healthy 23-year-old presenting with symptoms of persistent headache may benefit most from subsequent measurements of features x, y, and z. Therefore, it is highly desirable to use data collected from a broad population of individuals to adjust the ranking of features relevant to a single patient.
[0032] Given available patient information and quantifiable health status, the methods and systems disclosed herein determine valuable features to collect for a corresponding clinical inquiry (e.g., whether the patient has disease d or whether the patient would benefit from treatment t). Additionally, various embodiments also determine which measurement techniques can be used to obtain selected features, with the frequency of collection and desired accuracy and precision. Various embodiments provide an optimal set of features that a clinician can collect when a given patient's vitals are measured (e.g., currently available information), constrained by available resources and time for diagnosis. In various embodiments, the feature selection mechanism is conditional on the available patient information and quantifiable health status.
[0033] According to various embodiments, the noise tolerance for a set of features can be determined empirically. Furthermore, various embodiments may determine the noise tolerance of a feature on the condition that the set of features has already been measured. Thus, various embodiments include suggesting to the end user to increase or decrease the average possible tolerance for a given feature based on previous measurements of that feature or other features.
[0034] According to various embodiments, the optimal sampling frequency for a set of features can be determined algorithmically. Additionally, various embodiments may determine the sampling frequency of a feature on the condition that the set of features has already been measured. Therefore, various embodiments include suggesting to an end user to increase or decrease the sampling frequency for a given feature based on previous measurements of that feature or other features.
[0035] In various embodiments, machine learning algorithms are used to rank feature relevances according to a model trained on a dataset consisting of an input feature matrix and outcome vectors, and quantifiable information available about a given patient. Additionally, embodiments consistent with this disclosure provide subject-specific estimates of the ranking of feature sets based on quantifiable information available about a given patient and dataset.
[0036] The proposed solution further improves the functionality of the computer itself, because the computer itself reduces network usage by saving data storage space and shortening the time to decision-making brought about by the methods and systems disclosed herein.
[0037] Many of the embodiments provided herein describe situations where patient data is identifiable or where downloaded image history is stored, but each user may explicitly grant permission for such patient information to be shared or stored. Explicit permission may be granted using privacy controls integrated into the disclosure system. Each user may be provided with notification that such patient information may or will be shared with explicit consent, and each patient may share that information at any time and delete any stored user information. Stored patient information may be encrypted to protect patient security.
[0038] Exemplary system architecture Figure 1 illustrates exemplary architectures suitable for a diagnostic engine in a streaming data environment, according to various embodiments. Architecture 100 includes a server 130 and client devices 110 connected via a network 150. One of the many servers 130 is configured to host memory containing instructions that, when executed by a processor, cause the server 130 to perform at least some of the steps in the method disclosed herein. At least one of the servers 130 may include, or have access to, a database containing clinical data for multiple patients.
[0039] Server 130 may include any device having a suitable processor, memory, and communication capabilities for hosting image collections and trigger logic engines. The trigger logic engines may be accessible by various client devices 110 via the network 150. Client devices 110 can be, for example, desktop computers, mobile computers, tablet computers (including, for example, e-book readers), mobile devices (such as smartphones or PDAs), or any other device having a suitable processor, memory, and communication capabilities for accessing one of the trigger logic engines on Server 130. According to various embodiments, client devices 110 may be used by healthcare professionals such as physicians, nurses, or paramedics to access one of the trigger logic engines on Server 130 in real-time emergency situations (e.g., in a hospital, clinic, ambulance, or other public or residential environment). In some embodiments, one or more users of client devices 110 (e.g., nurses, paramedics, physicians, and other healthcare professionals) may provide clinical data to the trigger logic engines of one or more Servers 130 via the network 150. In yet another embodiment, one or more client devices 110 may automatically provide clinical data to the server 130. For example, in some embodiments, the client device 110 may be a blood testing unit in a clinic configured to automatically provide patient results to the server 130 via a network connection. The network 150 may include any one or more of the following: a local area network (LAN), a wide area network (WAN), the internet, etc. Furthermore, the network 150 may include, but is not limited to, any one or more of the following network topologies: a bus network, a star network, a ring network, a mesh network, a star bus network, a tree, or a hierarchical network.
[0040] Exemplary diagnostic system Figure 2 is a block diagram 200 illustrating an exemplary server 130 and client device 110 in the architecture 100 of Figure 1, according to a predetermined aspect of the present disclosure. The client device 110 and server 130 are connected to communicate via a network 150 through their respective communication modules 218-1 and 218-2 (collectively referred to as "communication modules 218"). The communication modules 218 are configured to interface with the network 150 to send and receive information such as data, requests, responses, and commands with other devices on the network. The communication modules 218 may be, for example, modems or Ethernet® cards. The client device 110 and server 130 may each include memories 220-1 and 220-2 (collectively referred to as "memory 220") and processors 212-1 and 212-2 (collectively referred to as "processor 212"). Memory 220 may store instructions that, when executed by the processor 212, cause either the client device 110 or the server 130 to perform one or more steps in the manner disclosed herein. Thus, the processor 212 may be configured to execute instructions such as instructions physically coded into the processor 212, instructions received from software in memory 220, or a combination of both.
[0041] According to various embodiments, the server 130 may include, or be communicatively coupled to, a database 252-1 and a master dataset 252-2 (hereinafter collectively referred to as "database 252"). In one or more implementations, database 252 may store clinical data for multiple patients. Database 252 may include a historical dataset H having time-series measurements of various characteristics, treatment information, model predictions, and patient-specific outcome information for one or more patients. The historical database H may include multiple characteristics measured at different points in time.
[0042] In various embodiments, the master dataset 252-2 may be the same as or included in database 252-1. Clinical data in database 252 may include non-identifiable patient characteristics, vital signs, blood measurements such as CBC (complete blood count), CMP (comprehensive metabolic panel), and measurement information such as blood gases (e.g., oxygen, CO2), immunological information, biomarkers, and cultures. Non-identifiable patient characteristics may include age, sex, and general medical history such as chronic diseases (e.g., diabetes, allergies). In various embodiments, clinical data may also include actions taken by healthcare professionals in response to measurement information such as treatment methods, drug administration events, and dosages. In various embodiments, clinical data may also include events and outcomes occurring in the patient's history (e.g., sepsis, stroke, cardiac arrest, shock). Although database 252 is illustrated as being separate from server 130, in a given embodiment, database 252 and trigger logic engine 242 can be hosted within the same server 130 and made accessible by any other server or client device in network 150.
[0043] The memory 220-2 within the server 130 may include a diagnostic engine 240 for evaluating possible patient outcomes based on a dataset of medical features. The diagnostic engine 240 may also include a trigger logic engine 242, a modeling tool 244, a statistical tool 246, and a supplementation tool 248. The modeling tool 244 may include instructions and commands for collecting relevant clinical data and evaluating expected outcomes (e.g., diagnosis). In some embodiments, the modeling tool 244 may suggest an action to take from a group of possible actions. The modeling tool 244 may include commands and instructions from neural networks (NNs) such as deep neural networks (DNNs), convolutional neural networks (CNNs), generative adversarial neural networks (GANs), deep reinforcement learning (DRL) algorithms, deep recurrent neural networks (DRNNs), random forests, k-nearest neighbors (KNNs) algorithms, k-means algorithms, or any combination thereof. According to various embodiments, the modeling tool 244 may include machine learning algorithms, artificial intelligence algorithms, or any combination thereof. The modeling tool 244 can dynamically generate a model-based model using information extracted from the historical dataset H, based on predictions made at a given point in time, measurements taken for a set of patients, and actual outcomes for the set of patients.
[0044] The statistical tool 246 evaluates data stored in the database 252 or provided by the modeling tool 244. The supplementary tool 248 may provide the modeling tool 244 with missing data inputs from the measurement information collected by the trigger logic engine 242. The trigger logic engine 242 may be configured to evaluate the input data {Pi} calculated by the statistical tool and various metrics related to the model F, and to trigger actions based on the input and whether it satisfies predetermined conditions. The streaming data input {Pi} may include multiple measured features provided by a nurse or other healthcare professional using a client device 110 for patient i. According to some embodiments, the server 130 may provide the client device 110 with ranking variables for one or more features in {Mi}. The ranking variables provided for a given feature in {Mi} may be information used by the end user to determine which features or sets of features should subsequently be measured for a given patient. According to some embodiments, the measured features {Pi} are provided to the server 130 from one or more client devices 110. According to various embodiments, the client device 110 may receive a predicted outcome or diagnosis from the server 130 in response to the input data {Pi}.
[0045] The modeling tool 244 includes a model F trained on a dataset D consisting of an m × (l + k) dimensional input feature matrix X and an outcome vector Y of dimension m (one entry for each patient). Mi is a k-dimensional feature vector containing k features that have not been measured for subject i. The l-dimensional feature vector Pi for subject i contains l features that have been measured for subject i. Thus, a set of n missing features (Min, n ≤ k) can be selected from Mi. For a given Min, the diagnostic engine 240 assigns three values. The first value is a scalar value s (e.g., 0 ≤ s ≤ 1) indicating the importance of the n set with respect to Y (patient outcome). The second value is a vector of size n (v1n), where each entry corresponds to the time-dependent variation for each feature in the given n set. And the third value, another vector v2n of size n, indicates the maximum possible noise in the measurement of each missing feature in the n set. According to various embodiments, the server 130 sends the group {Min,s,v1n,v2n} to the client device 110.
[0046] The client device 110 can access the diagnostic engine 240 via an application 222 installed on the client device 110 or via a web browser. The processor 212-1 can control the execution of the application 222 on the client device 110. According to various embodiments, the application 222 may include a user interface (e.g., a graphical user interface GUI) displayed for the user on the output device 216 of the client device 110. The user of the client device 110 may input input data as measurement information using the input device 214, or send queries to the diagnostic engine 240 via the user interface of the application 222. The input device 214 may include a stylus, mouse, keyboard, touchscreen, microphone, or any combination thereof. The output device 216 may include a display, headset, speaker, alarm or siren, or any combination thereof.
[0047] Figure 3 illustrates exemplary workflows for decision trees in various embodiments. In various embodiments, one or more client devices and servers disclosed herein may intervene in decision-making at each node of the decision tree 300. More specifically, a diagnostic engine including a trigger logic engine, modeling tools, statistical tools, and complementation tools may be used at one or more nodes in the decision tree 300. Each decision point is resolved independently and may lead to follow-up decisions. Decisions after the first decision (A) may be initiated by using data collected up to the previous decision, instead of recommending new data. In one exemplary embodiment, the first decision point may include finding patients at high risk of developing sepsis within the next X hours. The second decision point may include selecting subtypes of host responses that benefit from a broad spectrum (A, B, X, etc.) for those patients.
[0048] In various embodiments, a two-layer deep decision tree may be outlined as follows: 1) A clinician inquires whether a patient is at high risk of developing sepsis within the next six hours. 2a) If the clinician assesses the patient as high risk after receiving relevant information (vitals, tests, and machine learning-based predictions), should the patient be given antibiotics or antiviral drugs? 2b) If the clinician believes the patient does not have sepsis, the next level includes identifying whether the patient has a urinary tract infection (without complications).
[0049] Each decision point may involve performing a specific workflow. After a root-level decision point, subsequent decision points may suggest a set of features to collect. In various embodiments, the diagnostic tool may suggest a set of features based on estimates of the entire population, or it may include several options for using data collected according to tests requested or performed at previous decision points (e.g., from historical dataset H) that is available and recorded so far.
[0050] In various embodiments, recommended actions may include collecting new observations about a given patient, regardless of whether all features are available, and moving on to the next step whenever new data is ready.
[0051] In various embodiments, the machine learning model in the modeling tool provides outcome predictions or probabilities of outcomes, and confidence levels for the predictions, based on the available data. In some embodiments, the confidence level may be provided by statistical tools in the diagnostic engine. Thus, one or more decisions may be available based on the outcome predictions. Rules that depend on one or more decisions defined in the trigger logic engine may be used to determine when the diagnostic engine is ready to provide an answer or request an action.
[0052] Statistical tools can also assess the risk associated with each of one or more decisions. When the risk is low, the workflow stops and a decision is made. When the risk is high, the diagnostic engine may issue a query to a physician, nurse, or other healthcare professional (e.g., a question displayed on the touchscreen of the client device or via the microphone). When the physician, nurse, or other professional responds positively to the query (e.g., an "OK" response or pressing a button on the touchscreen of the client device), the workflow stops and a decision is made.
[0053] When the system detects missing data before making a decision (for example, due to high risk or low confidence level), the diagnostic engine may decide to wait for at least one feature in the missing data to be measured and incorporated into the modeling tool. Based on the selected decision, the system may suggest to the user to collect a new set of features. According to various embodiments, the system may wait for new features to be collected even when not requested by the user. According to various embodiments, the modeling tool may also update the model based on available features with relevant confidence metrics.
[0054] In various embodiments, to suggest features for collection when confidence in the predicted outcome is insufficient, the system may quantify the impact of missing individual features on the uncertainty of the prediction. In various embodiments, to suggest features to collect, the system may also apply dynamic models and variable importance determinations based on “similar” patient populations. Based on available variables, variable importance predictions may be obtained. To suggest missing features, the system may also use a historical dataset H to quantify the additional predictive value of features.
[0055] In various embodiments, the diagnostic engine also provides ranking variables assigned to each feature or set of features. Thus, based on their ranks and user-specific constraint functions, the diagnostic engine can suggest missing features to measure. These constraint functions may include feature cost and acquisition time.
[0056] Figure 4 illustrates, in various embodiments, methods for ranking one or more features in a dataset according to the relevance of a diagnostic engine using constraint functions. Assume there are up to 10 features to acquire, F1, F2, F3, F4, F5, F6, F7, F8, F9, and F10, of which three are measured for a particular patient {P=F2, F6, F9}. For example, the patient may enter the emergency room of a hospital, and at least two of the features F2, F6, and F9 may include body temperature and heart rate. The table below lists the cost and time to acquire (delay) associated with measuring missing features {M=F1, F3, F4, F5, F7, F8, and F10}. [Table 1]
[0057] Based on available data {P}, a clinician may inquire whether a patient has disease d. Therefore, the diagnostic tool disclosed herein outputs a ranking of the relevance of the remaining features {M} in terms of predicting the patient's outcome with a high level of confidence. The decision may be time-sensitive (e.g., within the next hour, or within another specified time), and cost may be a secondary concern. Therefore, the diagnostic engine may include constraint functions (e.g., ranking logic) in a modeling tool that reflects the above configuration with factors proportional to mathematical expressions such as:
number
[0058] The features in set {M} can be presented to the clinician in descending order according to the value of the constraint function. In some embodiments, of the N features presented (n=7 in this example), the diagnostic tool is the top in the list.
number
[0059] Figure 5 is a block diagram illustrating methods for quantifying the impact of missing features on predictive uncertainty for a diagnostic engine, according to various embodiments. In some embodiments, the set of missing features may include individual features (n=1 in n sets, see Figure 2). Predictive uncertainty is obtained by keeping the set of features "constant" (e.g., features F2, F6, and F9, see Figure 4), except for the feature whose predictive uncertainty is being quantified (e.g., F1, see Figure 4). The modeling tool evaluates the predicted outcomes for N complements, each having a different complement value for F1. In some examples, a statistical tool determines statistical parameters based on the modeling tool's predictions. For example, the statistical tool may determine the variance among the N predictions of the modeling tool. In various embodiments, a higher variance found by the statistical tool may correspond to a greater impact on the predicted value, and therefore to a greater importance of feature F1 for the diagnosis of this particular patient.
[0060] More specifically, the diagnostic engine quantifies the predictive uncertainty induced by specific features in Mi by keeping F3-F10 constant, supplementing F1 with different values multiple times (N times), and calculating the variance in the prediction. It starts with a model trained on a fixed number of features (e.g., a large set of features, or a master dataset extracted from a historical dataset H) that generates diagnoses at a given probability and confidence level.
[0061] Assuming that a given patient has j features in P and k features in M, the diagnostic engine can perform the following steps.
[0062] For i which varies from 1 to k,
[0063] Select the missing feature Fi from set M.
[0064] The feature 1...k-Fi (e.g., {1...k / i}) in M is supplemented using rough estimations (random, mean, median, etc.) from the historical dataset.
[0065] Using the historical master dataset H, we interpolate the feature Fi through multiple interpolation frameworks to generate N interpolated values for feature i. Mimputed represents an N vector where each entry corresponds to one of the N interpolated values of Fi.
[0066] Generate N identical copies M{1...k / i} of P.
[0067] Concatenate P, M{1...k / i}, and Mimputed to generate an input matrix I of size (N×(j+k)).
[0068] Using a modeling tool, a set of N predictions (e.g., diagnostic values or outcomes) is provided for each row in I.
[0069] The inter-complementary variance bi is determined from N predictions. After repeating the process for each feature in M, the k values bi can be associated with the relative feature associations in the set M.
[0070] Sort each feature in descending order by its biology (higher biology corresponds to a higher relevance).
[0071] In various embodiments, the model used to find predictions may be fixed, and variable importance is ranked using multiple interpolation based on a master dataset or historical dataset H. In various embodiments, the model may be dynamically updated as desired.
[0072] In various embodiments, the above method can be generalized to ranking separate sets of features by replacing 1...k with a list of specific sets (e.g., [{1,2,3}, {1,3,4}, {1,3,5}, etc.]).
[0073] Figure 6 is a block diagram illustrating a method for quantifying feature relevance in a diagnostic engine that selects similar patient datasets from a master dataset, according to various embodiments. In various embodiments, the method finds a filtered dataset containing a subset of patients similar to the current patient (e.g., from a master dataset or historical dataset H). In various embodiments, a relevance ranking of features specific to the current patient is generated by building a model using a more homogeneous population of “very similar” patients. The modeling tool builds a new model or updates an existing model to predict known outcomes (e.g., vector Y) for a subset of similar patients and provides relevance values or rankings for missing features using the techniques disclosed herein.
[0074] For a given observational value Pi of patient i having a limited set of features, the diagnostic engine selects the closest set of subjects NS from the historical master dataset H. Set NS may also include an additional set of measured features X. In some embodiments, the selection of set NS is based on an initial set of limited features. In various embodiments, set NS can be defined using any one of several different metrics (e.g., Euclidean, Manhattan, Mahalanobis, Minkowski, Shebyshev, cosine, correlation, Hamming, Jackard, Spearman, Gaussian kernel) and multiple methods (e.g., k-nearest neighbors, fixed radius nearest neighbors).
[0075] In various embodiments, the size of set NS may be an adjustable input in this method. For example, in various embodiments, all subjects may be used. Using set NS in the master dataset and the desired predictive value (e.g., known outcome Y for patients in set NS), the modeling tool constructs a monitored model FNS using feature X. If the performance of the FNS (e.g., outcome prediction and confidence level) is not greater than a predetermined threshold (e.g., accuracy, AUC, AUPR, F1 score, sensitivity, specificity, PPV, NPV, RMSE, r2, AIC, BIC, etc.), the modeling tool updates set X and constructs a new model FNS (or updates the existing model).
[0076] When the model FNS is satisfactory, the modeling tool calculates the variable importance of the FNS. The variable importance of the FNS provides numerical values for each feature in X by any one of several methods. In various embodiments, variable importance can be provided by model information approaches (linear regression, logistic regression, SVM, tree-based methods, neural networks, etc.). Such methods include gini importance, importance based on substitution, coefficient magnitude, etc. In various embodiments with limited model-specific capabilities, non-model information methods utilize exploration algorithms such as hill climbing, simulation annealing, and gene-based algorithms, optimized for common metrics such as accuracy, AUC, AUPR, F1 score, sensitivity, specificity, PPV, NPV, RMSE, r2, AIC, BIC, etc. Thus, the diagnostic engine proposes novel features about patient measurements based on the ranking of the variable importance of the FNS.
[0077] In various embodiments, the Model FNS may be constructed or updated for each new feature proposal. In various embodiments, when a feature is present in Mi but not in X, the feature importance for that feature may or will correspond to NA (Not Available).
[0078] Figure 7 is a block diagram showing methods for quantifying feature relevance in a diagnostic engine using a historical dataset selected from a master dataset, according to various embodiments. The various embodiments utilize this method to leverage a historical dataset H containing predictions and corresponding retrospective outcomes. Given a set of existing features Pi and a set of k missing features Mi, the diagnostic engine searches the historical dataset H for patient i and determines which features, in addition to those already present in Pi, had the greatest impact on the prediction accuracy.
[0079] The diagnostic engine selects a subset of H, Hp, according to instances where only the features in Pi exist. In various embodiments, the clinician or other authorized user may also have options for the subset or further "curate" Hp by selecting the set of subjects and distance metrics closest to Pi using various methods (e.g., k-nearest neighbors, fixed radius nearest neighbors) with different distance metrics (e.g., Euclidean, Manhattan, Mahalanobis, Minkowski, Shebyshev, cosine, correlation, Hamming, Jackard, Spearman, Gaussian kernel).
[0080] For an index j spanning 1...k, this method proceeds in various embodiments as follows: Select a feature Fj in set Mi. Select a subset Hp+j of H according to instances where a feature exists in Pi and feature Fj also exists. For Hp+j, determine the accuracy Aj of the model-based predictions based on known results (Y) using standard measurement metrics such as accuracy, AUC, AUPR, F1 score, sensitivity, specificity, PPV, NPV, RMSE, r2, AIC, and BIC. Order each feature Fj in M in descending order based on the corresponding value Aj.
[0081] In various embodiments, the above method can be generalized to ranking a selected set of n features by replacing the missing features 1...k with a list of n sets of missing features in each of the above steps (e.g., [{1,2,3}, {1,3,4}, {1,3,5}, etc.]).
[0082] Figure 8 is a flowchart illustrating steps in Method 800 for performing medical actions on a patient based on multiple medical characteristics received or supplemented over a time sequence, according to various embodiments. Method 800 can be performed at least partially by any one of client devices connected to one or more servers via a network (e.g., any one of the servers 130, and any one of the client devices 110, and the network 150). For example, according to various embodiments, the servers may host one or more medical devices or portable computer devices carried by medical personnel or healthcare workers. The client devices 110 may be handled by workers or other staff within a medical facility, paramedics in an ambulance transporting a patient to the emergency room of a medical facility or hospital, or by a person transporting a patient in an ambulance or caring for a patient in a public place away from a private residence or medical facility. At least some of the steps in Method 800 can be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine (e.g., application 222 and diagnostic logic engine 240) on a server via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and complement tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and complement tool 248) for retrieving, supplying, and processing clinical data in real time and providing behavioral recommendations. Furthermore, steps disclosed in Method 800 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or is communicably linked to a computer. Methods conforming to this disclosure may include at least some, but not all, of the steps exemplified in Method 800, performed in different sequences.Furthermore, a method conforming to this disclosure may include at least two or more steps, as in method 800, which are performed in overlapping or nearly simultaneous timelines.
[0083] Step 802 includes recommending a desired set of initial features to collect. In various embodiments, step 802 includes providing a suggestion based on estimates of the entire population or using those that have been available in the record to date.
[0084] Step 804 includes collecting new observations, each observation comprising one or more features. In various embodiments, step 804 may include receiving requests for one or more features from a physician, nurse, or other healthcare professional based on the importance of the features, cost constraints, and time constraints. In various embodiments, step 804 includes collecting one or more new features measured for a given patient. In various embodiments, step 804 includes moving to the next step once the new features are available. In various embodiments, step 804 includes waiting for a predetermined set of features to be measured before proceeding.
[0085] Step 806 includes predicting the outcome and providing a confidence level for the predicted outcome. In various embodiments, step 806 includes using a machine learning model to provide the prediction and / or probability.
[0086] Step 808 includes determining whether the confidence level is greater than a predetermined threshold. In various embodiments, step 808 includes evaluating whether the decision is ready based on rules on which the decision depends. If the decision is ready, step 808 may include displaying a score and assessing the risk of the decision in step 810a. If the risk of adverse events in step 810a is lower than the risk threshold, the workflow terminates.
[0087] Step 812a includes requesting approval from a physician, nurse, or healthcare professional when the risk of an adverse event is higher than the risk threshold. The workflow ends when the healthcare professional approves the request in step 812a. If the healthcare professional does not approve the request in step 812a, step 814 includes providing a ranking variable of importance (s) and a sampling frequency (v1n) for a given set of unmeasured features.
[0088] Step 816 includes identifying a noise tolerance (v2n) for a given set of unmeasured features. In various embodiments, step 816 includes selecting a measurement technique for each feature based on one or a combination thereof of noise tolerance methods conforming to the present disclosure.
[0089] In step 808, if the confidence level is lower than a predetermined threshold, step 810b includes determining whether all requested data is available. If not all data is available, the user proceeds to step 812b, which involves waiting for new data. If all requested data is available according to step 810b, the method continues in step 814.
[0090] Figure 9 is a flowchart illustrating the steps in Method 900 for selecting features relevant to a diagnostic engine by quantifying the impact of the absence of individual features according to various embodiments. Method 900 may be performed at least in part by any one of a client device (e.g., any one of Server 130 and any one of Client Device 110 and Network 150) connected to one or more servers via a network. For example, according to various embodiments, a server may host one or more medical devices or portable computer devices carried by a healthcare professional or medical worker. Client Device 110 may be handled by a worker or other staff member in a healthcare facility, an emergency medical technician in an ambulance transporting a patient to the emergency room of a healthcare facility or hospital, or a person caring for a patient in an ambulance or in a public place away from a private residence or healthcare facility. At least some of the steps in Method 900 may be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine (e.g., application 222 and diagnostic logic engine 240) on a server via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and supplementary tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and supplementary tool 248) that retrieve, feed, and process clinical data in real time and provide behavioral recommendations. Furthermore, steps disclosed in Method 900 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or linked to a computer in a communicative manner. Methods conforming to the Disclosure may include at least some, but not all, of the steps exemplified in Method 900, performed in different sequences. Furthermore, methods conforming to the Disclosure may include at least two or more steps, as in Method 900, that are performed temporally overlapping or substantially concurrently.
[0091] Step 902 involves supplementing the unmeasured features in the instance with a first value, while keeping the other remaining unmeasured features constant.
[0092] Step 904 involves evaluating the first outcome using a model that uses the first value in the instance.
[0093] Step 906 involves supplementing the unmeasured features in the dataset with second values, while keeping the other remaining unmeasured features constant.
[0094] Step 908 involves evaluating the second outcome using a model that utilizes the second value in the instance.
[0095] Step 910 includes determining the statistical parameters using the first and second outcomes.
[0096] Step 912 includes assigning unmeasured features to rankings corresponding to determined statistical parameters.
[0097] Figure 10 is a flowchart illustrating the steps in Method 1000 for selecting features relevant to a diagnostic engine based on filters for similar patient populations from a master dataset, according to various embodiments. Method 1000 can be performed at least partially by any one of client devices connected to one or more servers via a network (e.g., any one of the servers 130, and any one of the client devices 110, and the network 150). For example, according to various embodiments, the servers may host one or more medical devices or portable computer devices carried by healthcare personnel or medical professionals. The client devices 110 may be handled by workers or other staff within a healthcare facility, paramedics in an ambulance transporting a patient to the emergency room of a healthcare facility or hospital, or by a person transporting a patient in an ambulance or caring for a patient in a public place away from a private residence or healthcare facility. At least some of the steps in Method 1000 can be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine (e.g., application 222 and diagnostic logic engine 240) on a server via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and complement tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and complement tool 248) for retrieving, supplying, and processing clinical data in real time and providing behavioral recommendations. Furthermore, steps disclosed in Method 1000 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or linked to a computer in a communicative manner. Methods conforming to this disclosure may include at least some, but not all, of the steps exemplified in Method 1000, performed in different sequences.Furthermore, a method conforming to this disclosure may include at least two or more steps, as in method 1000, which are performed in overlapping or nearly simultaneous timelines.
[0098] Step 1002 involves selecting a filtered dataset from the master dataset according to at least one measured feature from the instance, the master dataset containing multiple datasets related to multiple known outcomes.
[0099] Step 1004 involves using a model-based feature importance methodology to identify the relative importance of unmeasured features using one or more known outcomes in the filtered dataset.
[0100] Step 1006 includes assigning unmeasured features to a ranking corresponding to the output from the model-based feature importance ranking.
[0101] Figure 11 is a flowchart illustrating the steps in Method 1100 for selecting features relevant to a diagnostic engine based on a model of measured features, according to various embodiments. Method 1100 can be performed at least partially by any one of client devices connected to one or more servers via a network (e.g., any one of the servers 130, and any one of the client devices 110, and the network 150). For example, according to various embodiments, the servers may host one or more medical devices or portable computer devices carried by healthcare personnel or medical professionals. The client device 110 may be handled by workers or other staff within a healthcare facility, paramedics in an ambulance transporting a patient to the emergency room of a healthcare facility or hospital, or by a person transporting a patient in an ambulance or caring for a patient in a public place away from a private residence or healthcare facility. At least some of the steps in Method 800 can be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine (e.g., application 222 and diagnostic logic engine 240) on a server via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and supplementary tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and supplementary tool 248) for retrieving, supplying, and processing clinical data in real time and providing behavioral recommendations. Furthermore, steps disclosed in Method 1100 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or linked to a computer in a communicative manner. Methods conforming to the Disclosure may include at least some, but not all, of the steps exemplified in Method 1100, performed in different sequences. Furthermore, methods conforming to the Disclosure may include at least two or more steps, as in Method 1100, that are performed temporally overlapping or substantially concurrently.
[0102] Step 1102 involves accessing a master dataset containing multiple datasets related to known outcomes.
[0103] Step 1104 includes determining the variance value associated with the model for the outcome, which includes determining that the model is based on unmeasured features and at least one other distinct feature in the dataset.
[0104] Step 1106 involves evaluating the variance of predictions for outcomes using a model that employs multiple compensated values for unmeasured features in the dataset.
[0105] Step 1108 involves assigning a ranking to the unmeasured features according to the value of the prediction's variance relative to the variance value.
[0106] Figure 12 is a flowchart illustrating the steps in Method 1200 for selecting features relevant to a diagnostic engine based on a historical dataset selected from a master dataset, according to various embodiments. Method 1200 can be performed at least partially by any one of client devices connected to one or more servers via a network (e.g., any one of the servers 130, and any one of the client devices 110, and the network 150). For example, according to various embodiments, the servers may host one or more medical devices or portable computer devices carried by healthcare professionals or medical personnel. The client devices 110 may be handled by workers or other staff within a healthcare facility, paramedics in an ambulance transporting a patient to the emergency room of a healthcare facility or hospital, or by a person transporting a patient in an ambulance or caring for a patient in a public place away from a private residence or healthcare facility. At least some of the steps in Method 1200 can be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine (e.g., application 222 and diagnostic logic engine 240) on a server via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and complement tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and complement tool 248) for retrieving, supplying, and processing clinical data in real time and providing behavioral recommendations. Furthermore, steps disclosed in Method 1200 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or linked to a computer in a communicative manner. Methods conforming to this disclosure may include at least some, but not all, of the steps exemplified in Method 1200, performed in different sequences.Furthermore, a method conforming to this disclosure may include at least two or more steps, as in method 1200, which are performed in overlapping or nearly simultaneous timelines.
[0107] Step 1202 involves determining a rule for assessing decision values based on a dataset, the dataset containing values collected for multiple measured and unmeasured features in an instance, and the rule is consistent with (1) multiple known outcomes from a master dataset containing multiple datasets, and (2) one or more measured features.
[0108] Step 1204 involves determining the accuracy of the rule based on multiple outcome values and known outcomes for each of the datasets.
[0109] Step 1206 involves assigning unmeasured features to a ranking corresponding to the accuracy of the rule.
[0110] Figure 13 is a flowchart illustrating the steps in Method 1300, which, according to various embodiments, constructs a multivariate model that predicts the importance of missing features using measured features. Method 1300 can be performed at least partially by any one of the client devices (e.g., any one of the servers 130, and any one of the client devices 110, and the network 150) connected to one or more servers via a network. For example, according to various embodiments, the servers may host one or more medical devices or portable computer devices carried by healthcare personnel or medical professionals. The client devices 110 may be handled by workers or other staff within a healthcare facility, paramedics in an ambulance transporting a patient to the emergency room of a healthcare facility or hospital, or by a person transporting a patient in an ambulance or caring for a patient in a public place away from a private residence or healthcare facility. At least some of the steps in Method 1200 can be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine (e.g., application 222 and diagnostic logic engine 240) on a server via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and supplementary tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and supplementary tool 248) that retrieve, feed, and process clinical data in real time and provide behavioral recommendations. Furthermore, steps disclosed in Method 1300 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or linked to a computer in a communicative manner. Methods conforming to the Disclosure may include at least some, but not all, of the steps exemplified in Method 1300, performed in different sequences. Furthermore, methods conforming to the Disclosure may include at least two or more steps, as in Method 1300, that are performed temporally overlapping or substantially concurrently.
[0111] The fundamental idea behind the third method is to build a multiclass model that predicts the importance of unmeasured features using features available to a given subject. This model is constructed by creating a dataset that estimates the variance induced by a particular feature in M for all subjects in a historical dataset H for all features in M. This methodology is suitable when there is an arbitrary set of features already collected, but the reliability is insufficient. In various embodiments, the model is built or updated when new feature proposals are desired. In various embodiments, the model building process may be performed only once when the set of unmeasured features is the same (e.g., starting with vitals and proposing measurements such as CMP, CBC, and specific biomarkers).
[0112] Step 1302 includes generating importance vectors based on features assumed to exist for all subjects in H. In various embodiments, step 1302 includes searching for observations Xs corresponding to those with the maximum number of features available during the relevant time frame for each subject s in the master dataset. Let S be a set Xs for all s. In various embodiments, a subset of features in Xs belongs to either P or M. P is a set of features assumed to exist, while M is a set of features assumed to be collected after P, and there may be k features in M. In various embodiments, step 1302 may include constructing a model f to predict outcome Y using S and calculating the variance of the prediction f(S) for all s in S using standard methods (e.g., standard error of prediction interval, jackknife estimator, Bayesian estimator, maximum likelihood-based estimator, etc.). The variance is an s × 1 vector V, with entries in V for each s.
[0113] In various embodiments, step 1302 involves taking the j-th entry of Ms (corresponding to Ms,j) for all subjects s in S and for j in 1....k, and randomly replacing it with a different value. This can be done by selecting a random value for the same feature from other subjects, or by drawing from a conditional distribution that models this feature using the remaining other features using the Markov Monte Carlo method, using the new replaced value and pretending it is the first observed value of Ms,j, generating a predicted value using the model, repeating the above steps independently many times, calculating the variance of the predicted Vj, dividing that value by the variance estimate based on X, and showing the ratio of the two as Rs,j = Vj / Vs, and performing steps (I)~(III) for all j entries of Ms. Sort the results from largest to smallest in Rs,j. The larger Rs,j, the more important the j-th feature is to the subject.
[0114] Step 1304 involves generating a personalized feature importance model using existing features P. Specifically, it involves constructing a multi-class model g (using multinomial regression, tree-based methods, neural networks, etc.) that predicts R using P, using all subjects in a historical dataset for which P is available.
[0115] Step 1306 includes providing a ranking of features in Mi for a given subject i via g(Pi).
[0116] Figure 14 is a flowchart illustrating the steps in Method 1400 for determining the sampling frequency for selected features based on the predictability of the features, according to various embodiments. Method 1400 can be performed at least partially by any one of the client devices connected to one or more servers via a network (e.g., any one of the servers 130, and any one of the client devices 110, and the network 150). For example, according to various embodiments, the servers may host one or more medical devices or portable computer devices carried by healthcare personnel or medical professionals. The client devices 110 may be handled by workers or other staff in a medical facility, paramedics in an ambulance transporting a patient to the emergency room of a medical facility or hospital, or by a person transporting a patient in an ambulance or caring for a patient in a public place away from a private residence or medical facility. At least some of the steps in Method 1400 can be performed by a computer having a processor (e.g., processor 212 and memory 220) that executes commands stored in the computer's memory. According to various embodiments, a user may launch an application on a client device to access a diagnostic engine on a server (e.g., application 222 and diagnostic logic engine 240) via a network. The diagnostic engine may include a trigger logic engine, modeling tools, statistical tools, and complement tools (e.g., trigger logic engine 242, modeling tool 244, statistical tool 246, and complement tool 248) for retrieving, supplying, and processing clinical data in real time and providing behavioral recommendations. Furthermore, steps disclosed in Method 1400 may, among other things, include using the diagnostic engine (e.g., database 252) to retrieve, edit, and / or store files in a database that is part of a computer or linked to a computer in a communicative manner. Methods conforming to this disclosure may include at least some, but not all, of the steps exemplified in Method 1400, performed in different sequences.Furthermore, a method conforming to this disclosure may include at least two or more steps, as in method 1400, which are performed in overlapping or nearly simultaneous timelines.
[0117] The basic idea behind this method is to estimate how predictable the future values of features are and, based on this, determine how often they should be sampled. Intuitively, the less predictable the future values of features, the more frequently they should be sampled. This method can be formally described as follows for a given subject i with a corresponding feature vector Pi:
[0118] Step 1402 involves identifying observed features P for a given object i, missing features M, with j features in P and k features in M. We want to determine the sampling frequency s of a given feature that can be in either P or M.
[0119] Step 1404 involves constructing a model g that predicts st+1 using a feature matrix X. In various embodiments, step 1404 involves selecting a feature matrix X from a historical dataset H. The feature matrix X exclusively contains features in P and may include time-series observations for each feature up to t. Relevant models include autoregressive models, moving average models, Markov models, and the like.
[0120] Step 1406 involves generating a prediction for st+x using g(P0...t).
[0121] Step 1408 includes determining the variance or CV of [Pt,g(P0...t)]. In various embodiments, this time-dependent variation is shown as Vs. In various embodiments, the above can be extended to predict a number of future values (e.g., st+x_1, st+x_2, ..., st+x_n). In various embodiments, step 1408 includes repeating the above steps for most or all of the remaining features in P and M.
[0122] Step 1410 involves ranking the selected features relative to other features based on their variance.
[0123] Step 1412 includes increasing the sampling frequency of a selected feature when its rank is in the upper r-th percentile. In various embodiments, step 1412 includes selecting an empirically determined factor that is proportional to the rank with respect to the baseline sampling frequency (which can be extracted from the historical dataset) of the feature. When the rank of a feature is in the lower r-th percentile, step 1412 includes proposing to decrease the sampling frequency by an empirically determined factor that is inversely proportional to the rank with respect to the baseline sampling frequency (which can be extracted from the historical dataset). Hardware Overview
[0124] Figure 15 is a block diagram illustrating an exemplary computer system 1500 in which the client device 110 and server 130 of Figures 1 and 2, and the methods of Figures 8 to 14, may be implemented. In a given embodiment, the computer system 1500 may be implemented using hardware or a combination of software and hardware, either within a dedicated server, integrated within another entity, or distributed across multiple entities.
[0125] The computer system 1500 (e.g., client device 110 and server 130) includes a bus 1508 or other communication mechanism for communicating information and a processor 1502 (e.g., processor 212) coupled to the bus 1508 for processing information. As an example, the computer system 1500 may be implemented with one or more processors 1502. The processor 1502 may be a general-purpose microprocessor, microcontroller, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), programmable logic device (PLD), controller, state machine, gate logic, discrete hardware component, or other suitable entity capable of performing computation or other operations on information.
[0126] In addition to hardware, the computer system 1500 may include code that generates an execution environment for the computer program in question, such as processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof, stored in memory 1504 (e.g., memory 220), which includes random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable PROM (EPROM), registers, a hard disk, a removable disk, a CD-ROM, a DVD, or any other suitable storage coupled to bus 1508 for storing information and instructions executed by processor 1502. The processor 1502 and memory 1504 may be supplemented by or embedded within special-purpose logic circuits.
[0127] Instructions are stored in memory 1504 and implemented in one or more modules of computer program instructions encoded on a computer-readable medium to control the execution or operation of one or more computer program products, i.e., computer systems 1500, and can be implemented in any method well known to those skilled in the art, including but not limited to computer languages such as data-oriented languages (e.g., SQL, dBase), system languages (e.g., C, Objective-C, C+, Assembly), architecture languages (e.g., Java, .NET), and application languages (e.g., PHP, Ruby, Perl, Python). Instructions can also be implemented in computer languages such as array languages, aspect-oriented languages, assembly languages, authoring languages, command-line interface languages, compiled languages, concurrent languages, curly bracket languages, dataflow languages, data structure languages, declarative languages, esoteric languages, extended languages, fourth-generation languages, function languages, interactive mode languages, interpreted languages, iterative languages, list-based languages, miniature languages, logic-based languages, machine languages, macro languages, metaprogramming languages, multi-paradigm languages, numerical analysis, non-English-based languages, object-oriented class-based languages, object-oriented prototype-based languages, offside rule languages, procedural languages, reflective languages, rule-based languages, scripting languages, stack-based languages, synchronous languages, syntactic languages, visual languages, Wirth languages, and XML-based languages. Memory 1504 can also be used to store temporary variables or other intermediate information during the execution of instructions performed by processor 1502.
[0128] The computer programs described herein do not necessarily correspond to files in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., a file storing one or more modules, subprograms, or portions of code). A computer program may be deployed to run on one computer, or on multiple computers located in one site, or distributed across multiple sites and interconnected by a communication network. The processes and logical flows described herein may be executed by one or more programmable processors that run one or more computer programs to perform their functions by acting on input data and producing outputs.
[0129] The computer system 1500 further includes a data storage device 1506, such as a magnetic disk or optical disk, coupled to a bus 1508 for storing information and instructions. The computer system 1500 may be coupled to various devices via an input / output module 1510. The input / output module 1510 may be any input / output module. An exemplary input / output module 1510 includes a data port, such as a USB port. The input / output module 1510 is configured to connect to a communication module 1512. An exemplary communication module 1512 (e.g., communication module 218) includes a network interface card, such as an Ethernet card and a modem. In certain embodiments, the input / output module 1510 is configured to connect to multiple devices, such as an input device 1514 (e.g., input device 214) and / or an output device 1516 (e.g., output device 216). An exemplary input device 1514 includes a keyboard and a pointing device, such as a mouse or trackball, which the user can use to provide input to the computer system 1500. Interaction with the user can also be provided using other types of input devices 1514, such as haptic input devices, visual input devices, audio input devices, or brain-computer interface devices. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback, and input from the user can be received in any form, including acoustic, voice, tactile, or electroencephalogram (EEG) input. An exemplary output device 1516 includes a display device, such as an LCD (liquid crystal display) monitor, for displaying information to the user.
[0130] According to one aspect of this disclosure, the client device 110 and the server 130 can be implemented using a computer system 1500 in response to a processor 1502 executing one or more sequences of one or more instructions contained in memory 1504. Such instructions may be read into memory 1504 from another machine-readable medium, such as a data storage device 1506. By executing the sequence of instructions contained in main memory 1504, the processor 1502 performs the process steps described herein. Alternatively, one or more processors in a multi-processing configuration may be used to execute the sequence of instructions contained in memory 1504. In another aspect, hardwired circuitry may be used instead of or in combination with software instructions to implement various aspects of this disclosure. Thus, aspects of this disclosure are not limited to any particular combination of hardware circuitry and software.
[0131] Various embodiments of the subject matter described herein can be implemented in a computing system that includes backend components, e.g., a data server; middleware components, e.g., an application server; or frontend components, e.g., a client computer having a graphical user interface or web browser on which a user can interact with embodiments of the subject matter described herein; or any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. The communication network (e.g., network 150) can include any one or more of, for example, a LAN, WAN, or the Internet. Furthermore, the communication network can include, but is not limited to, any one or more of the network topologies, e.g., a bus network, a star network, a ring network, a mesh network, a starbus network, a tree, or a hierarchical network. The communication module can be, for example, a modem or an Ethernet card.
[0132] The computer system 1500 may include a client and a server. The client and server are generally geographically separated and interact via a communication network. The client-server relationship arises from computer programs running on each computer that have a client-server relationship with each other. The computer system 1500 may, for example, be a desktop computer, a laptop computer, or a tablet computer. The computer system 1500 may also be, but is not limited to, other devices such as a mobile phone, PDA, mobile audio player, global positioning system (GPS) receiver, video game console, and / or television set-top box.
[0133] As used herein, the terms “machine-readable storage medium” or “computer-readable medium” refer to any medium that participates in providing instructions to the processor 1502 for execution. Such mediums can take many forms, including, but are not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media include, for example, optical or magnetic disks such as data storage device 1506. Volatile media include dynamic memory such as memory 1504. Transmission media include coaxial cables, copper wires, and optical fibers, including wires that constitute bus 1508. Common forms of machine-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs, any other optical media, punch cards, paper tapes, any other physical media having a pattern of holes, RAM, PROMs, EPROMs, FLASH® EPROMs, any other memory chips or cartridges, or any other media that a computer can read. A machine-readable memory medium can be a machine-readable memory device, a machine-readable memory substrate, a memory device, a composition of a substance that affects machine-readable propagation signals, or a combination of one or more of these.
[0134] Where used herein, the phrase “at least one” preceding a set of items, along with the terms “and” or “or” separating any of the items, modifies the entire list, rather than each member of the list (i.e., each item). The phrase “at least one” does not require the selection of at least one item; rather, it allows for the meaning of including at least one of any one of the items, and / or at least one of any combination of items, and / or at least one of each of the items. For example, the phrase “at least one of A, B, and C” or “at least one of A, B, or C” refers, respectively, to A only, B only, or C only, any combination of A, B, and C, and / or at least one of each of A, B, and C.
[0135] To the extent that terms such as “including” and “having” are used in the specification or claims, such terms are intended to be as comprehensive as the term “including,” so as “including” is interpreted when it is used as a transitional clause in a claim. In this specification, the word “exemplary” is used to mean “serving as an example, instance, or illustration.” In this specification, any embodiment described as “exemplary” is not necessarily construed to be preferable or advantageous to other embodiments.
[0136] References to singular elements are not intended to mean "one and just one" unless specifically stated, but rather "one or more." All structural and functional equivalents to the elements of the various configurations described throughout this disclosure, whether known to those skilled in the art or to become known thereafter, are incorporated herein by reference and intended to be included in the subject art. Furthermore, nothing disclosed herein is intended to be for the public only, whether such disclosure is expressly provided for in the above description.
[0137] Although this specification includes many specific examples, these should not be interpreted as limiting the scope of what is described in the claims, but rather as descriptions of specific implementations of the subject matter. Certain features described herein in the context of separate embodiments can be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment can be implemented separately in multiple embodiments, or in any preferred subcombination. Furthermore, features are described above as acting in a given combination, and may even be initially described in the claims as such, but one or more features from a claimed combination may, in some cases, be extracted from the combination, and the claimed combination may be directed towards subcombinations or variations of subcombinations.
[0138] While the subject matter of this specification has been described in terms of specific embodiments, other embodiments can be implemented and are within the scope of the following claims. For example, although operations are shown in a specific order in the drawings, this should not be understood as requiring that such operations be performed in the specific order shown, or in a sequential order, or that all exemplified operations be performed, in order to achieve the desired result. The actions defined in the claims can be performed in a different order and still achieve the desired result. As an example, the process shown in the accompanying drawings does not necessarily require the specific order shown, or in a sequential order, to achieve the desired result. Under certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems can generally be integrated within a single software product or packaged within multiple software products. Other modifications are within the scope of the following claims.
[0139] Provisions of the Embodiment (Embodiment 1) A method is provided for ranking unmeasured features in an instance in which at least one feature has been measured, comprising: supplementing the unmeasured features in the instance with a first value while keeping the other remaining unmeasured features constant; evaluating a first outcome using a model that uses the first value in the instance; supplementing the unmeasured features in the instance with a second value while keeping the other remaining unmeasured features constant; evaluating a second outcome using a model that uses the second value in the instance; determining statistical parameters using the first and second outcomes; and assigning the unmeasured features to a ranking corresponding to the statistical parameters.
[0140] (Embodiment 2) The method according to Embodiment 1, further comprising selecting a filtered dataset from a master dataset according to at least one measured feature from an instance, wherein the master dataset includes multiple datasets related to multiple known outcomes.
[0141] (Embodiment 3) The method according to Embodiment 1 or 2, wherein assigning unmeasured features to a ranking corresponding to a statistical parameter includes identifying the relative importance of the unmeasured features using one or more known outcomes with a model-based feature importance methodology in a filtered dataset.
[0142] (Embodiment 4) The method according to any one of Embodiments 1 to 3, wherein determining statistical parameters using the first and second outcomes includes accessing a master dataset containing multiple datasets associated with known outcomes.
[0143] (Embodiment 5) The method according to any one of Embodiments 1 to 4, wherein determining statistical parameters using a first and second outcome is to determine the variance value associated with the model for the outcome, the model being based on an unmeasured feature and at least one other distinct feature in the dataset, and evaluating the variance of predictions for the outcome using a model that uses multiple compensated values for an unmeasured feature in the dataset.
[0144] (Embodiment 6) The method according to any one of Embodiments 1 to 5, wherein determining statistical parameters using the first and second outcomes includes determining a rule for assessing a decision value based on a dataset, the dataset includes values collected for multiple measured features and unmeasured features in an instance, and the rule is consistent with (1) multiple known outcomes from a master dataset including multiple datasets, and (2) one or more measured features.
[0145] (Embodiment 7) The method according to any one of Embodiments 1 to 6, wherein determining statistical parameters using a first and second outcome includes determining the accuracy of a rule for supplementing the first value with an unmeasured feature based on a plurality of outcome values and known outcomes for each of a plurality of datasets.
[0146] (Embodiment 8) The method according to any one of Embodiments 1 to 7, wherein determining the statistical parameters includes determining the time-dependent variance of the first outcome and the second outcome.
[0147] (Embodiment 9) The method according to any one of Embodiments 1 to 8, further comprising selecting the sampling frequency of unmeasured features based on a ranking corresponding to a statistical parameter.
[0148] (Embodiment 10) The method according to any one of Embodiments 1 to 9, further comprising selecting a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor device and the ranking of unmeasured features.
[0149] (Embodiment 11) A system is provided for ranking unmeasured features for an instance in which at least one feature has been measured, comprising: a memory for storing instructions; and one or more processors communicatively coupled to the memory, wherein the one or more processors execute instructions to cause the system to: supplement the unmeasured features in an instance with a first value while keeping the other remaining unmeasured features constant; evaluate a first outcome using a model that uses the first value in the instance; supplement the unmeasured features in an instance with a second value while keeping the other remaining unmeasured features constant; evaluate a second outcome using a model that uses the second value in the instance; determine statistical parameters using the first and second outcomes; assign the unmeasured features to a ranking corresponding to the statistical parameters; and select a filtered dataset from a master dataset according to at least one measured feature from the instance, wherein the master dataset includes a plurality of datasets associated with a plurality of known outcomes.
[0150] (Embodiment 12) The system according to Embodiment 11, wherein one or more processors execute instructions to assign unmeasured features to a ranking corresponding to statistical parameters, and to identify the relative importance of unmeasured features using one or more known outcomes in a filtered dataset using a model-based feature importance methodology.
[0151] (Embodiment 13) The system according to Embodiment 11 or 12, wherein one or more processors execute instructions to access a master dataset containing multiple datasets associated with known outcomes in order to determine statistical parameters using a first outcome and a second outcome.
[0152] (Embodiment 14) The system according to any one of Embodiments 11 to 13, wherein one or more processors execute instructions to determine a model-related variance for an outcome, the model being based on an unmeasured feature and at least one other distinct feature in the dataset, and the system evaluates the variance of predictions for an outcome using a model that uses multiple complement values for an unmeasured feature in the dataset.
[0153] (Embodiment 15) The system according to any one of Embodiments 11 to 14, wherein one or more processors execute instructions to determine rules for assessing decision values based on a dataset, in order to determine statistical parameters using a first and second outcome, the dataset includes values collected for multiple measured and unmeasured features in an instance, and the rules match (1) multiple known outcomes from a master dataset including multiple datasets, and (2) one or more measured features.
[0154] (Embodiment 16) A non-temporary computer-readable medium for storing instructions, the instructions, when executed by a computer, cause the computer to perform a method for ranking unmeasured features for an instance in which at least one feature has been measured, the method comprising: supplementing the unmeasured features in the instance with a first value while keeping the other remaining unmeasured features constant; evaluating a first outcome using a model that uses the first value in the instance; supplementing the unmeasured features in the instance with a second value while keeping the other remaining unmeasured features constant; evaluating a second outcome using a model that uses the second value in the instance; and the first and second outcomes. A non-temporary computer-readable medium is provided, comprising: determining statistical parameters using outcomes; assigning unmeasured features to rankings corresponding to the statistical parameters; and selecting filtered datasets from a master dataset according to at least one measured feature from an instance, wherein the master dataset includes multiple datasets associated with multiple known outcomes; and assigning unmeasured features to rankings corresponding to the statistical parameters, further comprising identifying the relative importance of unmeasured features using one or more known outcomes in the filtered datasets using a model-based feature importance methodology.
[0155] (Embodiment 17) The method for determining statistical parameters using the first and second outcomes is a non-temporary computer-readable medium as described in Embodiment 16, which includes accessing a master dataset containing multiple datasets associated with known outcomes.
[0156] (Embodiment 18) A non-temporary computer-readable medium according to Embodiment 16 or 17, wherein determining statistical parameters using a first outcome and a second outcome in the method involves determining a variance value associated with a model for the outcome, the model being based on an unmeasured feature and at least one other distinct feature in the dataset, and evaluating the variance of predictions for the outcome using a model that uses multiple compensated values for an unmeasured feature in the dataset.
[0157] (Embodiment 19) Determining statistical parameters using the first and second outcomes in the method includes determining a rule for assessing a decision value based on a dataset, the dataset including values collected for multiple measured and unmeasured features in an instance, and the rule being a non-temporary computer-readable medium according to any one of Embodiments 16 to 18 that matches (1) multiple known outcomes from a master dataset including multiple datasets, and (2) one or more measured features.
[0158] (Embodiment 20) A non-temporary computer-readable medium according to any one of Embodiments 16 to 19, wherein the method for determining statistical parameters using a first outcome and a second outcome includes determining the accuracy of a rule for supplementing the first value with an unmeasured feature based on a plurality of outcome values and known outcomes for each of a plurality of datasets.
[0159] (Embodiment 21) A method is provided for ranking unmeasured features for an instance in which at least one feature has been measured, comprising: selecting a filtered dataset from a master dataset according to at least one measured feature from the instance, wherein the master dataset comprises a plurality of datasets associated with a plurality of known outcomes; identifying the relative importance of unmeasured features using one or more known outcomes in the filtered dataset using a model-based feature importance methodology; and assigning the unmeasured features to a ranking corresponding to the output from the model-based feature importance.
[0160] (Embodiment 22) The method according to Embodiment 21, wherein selecting a filtered dataset from a master dataset includes selecting at least a portion of a historical dataset.
[0161] (Embodiment 23) The method according to Embodiment 21 or 22, wherein in a filtered dataset, a model-based feature importance methodology is used to identify the relative importance of unmeasured features using one or more known outcomes, which includes building or updating a model with new features.
[0162] (Embodiment 24) The method according to any one of Embodiments 21-23, further comprising selecting a filtered dataset to determine statistical parameters using known outcomes.
[0163] (Embodiment 25) The method according to any one of Embodiments 1 to 24, comprising: determining a statistical parameter using a first outcome and a second outcome; determining a variance value associated with a model for the outcome, the model being based on an unmeasured feature and at least one other distinct feature in the dataset; and evaluating the variance of predictions for the outcome using a model that uses multiple compensated values for an unmeasured feature in the dataset.
[0164] (Embodiment 26) The method according to any one of Embodiments 21 to 25, further comprising determining a rule for assessing a decision value based on a filtered dataset.
[0165] (Embodiment 27) The method according to any one of Embodiments 21 to 26, further comprising determining a statistical parameter using a plurality of outcomes and determining the accuracy of a rule for supplementing an unmeasured feature with a first value based on the plurality of results.
[0166] (Embodiment 28) The method according to any one of embodiments 21 to 27, further comprising determining the time-dependent variance of the first and second outcomes.
[0167] (Embodiment 29) The method according to any one of Embodiments 21 to 28, further comprising selecting the sampling frequency of unmeasured features based on a ranking corresponding to a statistical parameter.
[0168] (Embodiment 30) The method according to any one of embodiments 21 to 29, further comprising selecting a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of unmeasured features.
[0169] (Embodiment 31) A system for ranking unmeasured features for an instance in which at least one feature has been measured, comprising: a memory for storing instructions; and one or more processors communicatively coupled to the memory, wherein the one or more processors execute instructions to cause the system to select filtered datasets from a master dataset according to at least one measured feature from an instance, the master dataset comprising a plurality of datasets associated with a plurality of known outcomes; to identify the relative importance of unmeasured features using one or more known outcomes in the filtered datasets using a model-based feature importance methodology; and to assign the unmeasured features to a ranking corresponding to the output from the model-based feature importance.
[0170] (Embodiment 32) The system according to Embodiment 31, wherein one or more processors further execute instructions to select at least a portion of the historical dataset in order to select a filtered dataset from the master dataset.
[0171] (Embodiment 33) The system according to Embodiment 31 or 32, wherein, in a filtered dataset, one or more processors further execute instructions to build or update a model with new features in order to identify the relative importance of unmeasured features using one or more known outcomes, using a model-based feature importance methodology.
[0172] (Embodiment 34) The system according to any one of embodiments 31 to 33, wherein one or more processors further execute instructions to select filtered datasets and determine statistical parameters using known outcomes.
[0173] (Embodiment 35) The system according to any one of Embodiments 31 to 34, wherein one or more processors execute instructions to determine statistical parameters using a first outcome and a second outcome, and to determine the variance value associated with a model for the outcome, wherein the model is based on an unmeasured feature and at least one other distinct feature in the dataset, and to evaluate the variance of predictions for the outcome using a model that uses multiple complement values for an unmeasured feature in the dataset.
[0174] (Embodiment 36) The system according to any one of embodiments 31 to 35, wherein one or more processors further execute instructions to determine rules for assessing decision values based on a filtered dataset.
[0175] (Embodiment 37) A system of any one of embodiments 31 to 36, wherein one or more processors further execute instructions to determine statistical parameters using a plurality of outcomes and to determine the accuracy of rules for supplementing unmeasured features with first values based on the plurality of outcomes.
[0176] (Embodiment 38) The system according to any one of embodiments 31 to 37, wherein one or more processors further execute instructions to determine the time-dependent variance of the first and second outcomes.
[0177] (Embodiment 39) The system according to any one of embodiments 31 to 38, wherein one or more processors further execute instructions to reduce the sampling frequency of unmeasured features based on the rank of the unmeasured features.
[0178] (Embodiment 40) The system according to any one of embodiments 31 to 39, wherein one or more processors further execute instructions to select a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of unmeasured features.
[0179] (Embodiment 41) A method is provided for ranking unmeasured features for instances in which at least one feature has been measured, the method comprising: accessing a master dataset, the master dataset comprising a plurality of datasets related to known outcomes; determining a variance value related to a model for outcomes, the model being based on unmeasured features and at least one other distinct feature in the datasets; evaluating the variance of predictions for outcomes using a model that uses a plurality of complement values for unmeasured features in the datasets; and assigning the unmeasured features to a ranking according to the value of the variance of predictions against the variance value.
[0180] (Embodiment 42) The method according to Embodiment 41, wherein accessing the master dataset includes selecting at least a portion of the historical dataset.
[0181] (Embodiment 43) The methods of Embodiments 41 and 42, wherein evaluating the variance of the predictions includes building or updating the model with new features.
[0182] (Embodiment 44) The method according to any one of Embodiments 41 to 43, wherein determining the variance values related to the model for the outcome includes selecting a filtered dataset from the master dataset.
[0183] (Embodiment 45) The method according to any one of Embodiments 41 to 44, wherein determining the variance value associated with the model for the outcome includes selecting a model based on an unmeasured feature and at least one other distinct feature in the dataset, and evaluating the variance of the prediction for the outcome using a model that uses multiple complement values for the unmeasured feature in the dataset.
[0184] (Embodiment 46) The method according to any one of embodiments 41 to 45, further comprising determining a rule for assessing a decision value based on a master dataset.
[0185] (Embodiment 47) The method according to any one of embodiments 41 to 46, further comprising determining the accuracy of a rule for supplementing unmeasured features with a first value based on known outcomes.
[0186] (Embodiment 48) The method according to any one of embodiments 41 to 47, further comprising determining the time-dependent variance of the first and second outcomes.
[0187] (Embodiment 49) The method according to any one of Embodiments 41 to 48, further comprising selecting the sampling frequency of unmeasured features based on a ranking corresponding to a statistical parameter.
[0188] (Embodiment 50) The method according to any one of embodiments 41 to 49, further comprising selecting a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of unmeasured features.
[0189] (Embodiment 51) A method for ranking unmeasured features for instances in which at least one feature has been measured, comprising: a memory for storing instructions; and one or more processors communicatively coupled to the memory, wherein the one or more processors are configured to execute instructions to cause the system to: access a master dataset, the master dataset comprising a plurality of datasets related to known outcomes; determine a variance value related to a model for outcomes, the model being based on unmeasured features and at least one other distinct feature in the datasets; evaluate the variance of predictions for outcomes using a model that uses a plurality of complement values for unmeasured features in the datasets; and assign unmeasured features to a ranking according to the value of the variance of predictions relative to the variance value.
[0190] (Embodiment 52) The system according to Embodiment 51, wherein one or more processors execute instructions to select at least a portion of the history dataset in order to access the master dataset.
[0191] (Embodiment 53) The system according to embodiments 51 and 52, wherein one or more processors execute instructions to build or update a model with new features in order to evaluate the variance of predictions.
[0192] (Embodiment 54) The system according to any one of Embodiments 51 to 53, wherein one or more processors execute instructions to select a filtered dataset from a master dataset in order to determine the variance value related to the model for the outcome.
[0193] (Embodiment 55) The system according to any one of Embodiments 51 to 54, wherein one or more processors execute instructions to select a model based on an unmeasured feature and at least one other distinct feature in the dataset, and to evaluate the variance of predictions for the outcome using a model that uses multiple complement values for an unmeasured feature in the dataset.
[0194] (Embodiment 56) The system according to any one of embodiments 51 to 55, wherein one or more processors further execute instructions to determine rules for assessing decision values based on a master dataset.
[0195] (Embodiment 57) The system according to any one of embodiments 51 to 56, wherein one or more processors further execute instructions to determine the accuracy of the rules for supplementing unmeasured features with first values, based on known outcomes.
[0196] (Embodiment 58) The system according to any one of embodiments 51 to 57, wherein one or more processors further execute instructions to determine the time-dependent variance of the first and second outcomes.
[0197] (Embodiment 59) The system according to any one of embodiments 51 to 58, wherein one or more processors further execute instructions to reduce the sampling frequency of unmeasured features based on the rank of the unmeasured features.
[0198] (Embodiment 60) The system according to any one of embodiments 51 to 59, wherein one or more processors further execute instructions to select a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of the unmeasured features.
[0199] (Embodiment 61) A method is provided for ranking unmeasured features for an instance in which at least one feature has been measured, the method comprising determining a rule for assessing a decision value based on a dataset, wherein the dataset includes values collected for multiple measured features and unmeasured features in the instance, and the rule includes determining whether it matches (1) multiple known outcomes from a master dataset including multiple datasets, and (2) one or more measured features; determining the accuracy of the rule based on the multiple outcome values and the known outcomes for each of the databases; and assigning a ranking to the unmeasured features corresponding to the accuracy of the rule.
[0200] (Embodiment 62) The method according to Embodiment 61, wherein assessing the decision value based on the dataset includes accessing the master dataset.
[0201] (Embodiment 63) The method according to Embodiment 61 or 62, wherein determining the accuracy of the rule based on multiple outcome values includes building or updating a model with new features.
[0202] (Embodiment 64) A method from any one of Embodiments 61 to 63, wherein determining the accuracy of a rule for assessing a decision value based on a dataset further includes determining the variance value associated with the model for the outcome.
[0203] (Embodiment 65) The method according to any one of Embodiments 61 to 64, wherein determining a rule for assessing decision values based on a dataset involves selecting a model based on an unmeasured feature and at least one other distinct feature in the dataset, and evaluating the variance of predictions for outcomes using a model that uses multiple compensated values for an unmeasured feature in the dataset.
[0204] (Embodiment 66) The method according to any one of embodiments 61 to 65, further comprising determining a rule for evaluating a decision value based on master data selected from a historical dataset.
[0205] (Embodiment 67) The method according to any one of Embodiments 61 to 66, wherein determining the accuracy of the rule involves updating the model for the rule using unmeasured features.
[0206] (Embodiment 68) The method according to any one of embodiments 61 to 67, further comprising determining the time-dependent variance of the first and second outcomes.
[0207] (Embodiment 69) The method according to any one of embodiments 61 to 68, further comprising selecting the sampling frequency of unmeasured features based on a ranking corresponding to a statistical parameter.
[0208] (Embodiment 70) The method according to any one of embodiments 61 to 69, further comprising selecting a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of unmeasured features.
[0209] (Embodiment 71) A system for ranking unmeasured features for instances in which at least one feature has been measured, comprising: a memory for storing instructions; and one or more processors communicatively coupled to the memory, wherein the one or more processors execute instructions to cause the system to: determine a rule for assessing a decision value based on a dataset, the dataset comprising a plurality of measured features in an instance and values collected for an unmeasured feature in an instance; and the rule being configured to determine whether it matches (1) a plurality of known outcomes from a master dataset comprising a plurality of datasets, and (2) one or more measured features; determine the accuracy of the rule based on the plurality of outcome values and the known outcomes for each of the databases; and assign a ranking to the unmeasured features corresponding to the accuracy of the rule.
[0210] (Embodiment 72) The method according to Embodiment 71, wherein one or more processors execute instructions to access a master dataset in order to assess a decision value based on a dataset.
[0211] (Embodiment 73) The system of Embodiment 71 or 72, wherein one or more processors execute instructions to build or update a model with new features in order to determine the accuracy of the rule based on a plurality of outcome values.
[0212] (Embodiment 74) Any one of the systems from Embodiments 71 to 73, wherein one or more processors further execute instructions to determine the accuracy of the rules for assessing decision values based on a dataset, thereby determining the variance values associated with the model for the outcome.
[0213] (Embodiment 75) The system according to any one of embodiments 71 to 74, wherein one or more processors execute instructions to select a model based on an unmeasured feature and at least one other distinct feature in the dataset, and to evaluate the variance of predictions for an outcome using a model that uses multiple complement values for an unmeasured feature in the dataset.
[0214] (Embodiment 76) A system from any one of embodiments 71 to 75, wherein one or more processors further execute instructions to determine rules for evaluating decision values based on a master dataset selected from a history dataset.
[0215] (Embodiment 77) The system according to any one of embodiments 71 to 76, wherein one or more processors execute instructions to update a model of the rules using unmeasured features in order to determine the accuracy of the rules.
[0216] (Embodiment 78) The system according to any one of embodiments 71 to 77, wherein one or more processors further execute instructions to determine the time-dependent variance of the first and second outcomes.
[0217] (Embodiment 79) The system according to any one of embodiments 71 to 78, wherein one or more processors further execute instructions to reduce the sampling frequency of unmeasured features based on the rank of the unmeasured features.
[0218] (Embodiment 80) The system according to any one of embodiments 71 to 79, wherein one or more processors further execute instructions to select a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of the unmeasured features.
[0219] (Embodiment 81) A method is provided for determining the sampling frequency of a feature selected based on the predictability of the feature, comprising: identifying a set of observed features and a set of missing features; constructing a model to predict the sampling frequency of the selected feature using a feature matrix selected from a historical dataset; using the model to generate predictions about the sampling frequency; determining the variance of the selected feature from a plurality of time predictions; ranking the selected feature relative to other features based on the variance; and increasing the sampling frequency of the selected feature when the feature rank is in a predetermined upper percentile.
[0220] (Embodiment 82) The method according to Embodiment 81, further comprising accessing a historical dataset containing observed features.
[0221] (Embodiment 83) The system according to Embodiment 81 or 82, wherein constructing a model to predict the sampling frequency includes evaluating the variance of the predictions.
[0222] (Embodiment 84) The method according to any one of Embodiments 81 to 83, wherein determining the variance of selected features includes selecting a filtered dataset from a master dataset.
[0223] (Embodiment 85) The method according to any one of Embodiments 81 to 84, wherein determining the variance of selected values includes selecting a model based on observed features and at least one other distinct feature in the dataset, and evaluating the variance of predictions for outcomes using a model that uses multiple compensated values for unmeasured features in the dataset.
[0224] (Embodiment 86) The method according to any one of embodiments 81 to 85, further comprising determining a rule for evaluating a decision value based on the sampling frequency.
[0225] (Embodiment 87) The method according to any one of embodiments 81 to 86, further comprising determining the accuracy of a rule for filling in missing features with a first value based on the model.
[0226] (Embodiment 88) The method according to any one of embodiments 81 to 87, further comprising determining the time-dependent variance of the first and second outcomes.
[0227] (Embodiment 89) The method according to any one of embodiments 81 to 88, further comprising reducing the sampling frequency of unmeasured features based on the rank of the unmeasured features.
[0228] (Embodiment 90) The method according to any one of embodiments 81 to 89, further comprising selecting a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of unmeasured features.
[0229] (Embodiment 91) A system for determining the sampling frequency of a selected feature based on the predictability of the feature, comprising a memory for storing instructions and one or more processors communicatively coupled to the memory, wherein the one or more processors execute instructions to the system, including a memory instruction and one or more processors coupled to communicate with the memory, to execute instructions to identify a set of observed features and a set of missing features, to construct a model for predicting the sampling frequency of a selected feature using a feature matrix selected from a historical dataset, to generate predictions about the sampling frequency using the model, to determine the variance of the selected feature from a plurality of time predictions, to rank the selected feature relative to other features based on the variance, and to increase the sampling frequency of the selected feature when the rank of the feature is in a predetermined upper percentile.
[0230] (Embodiment 92) The system according to Embodiment 91, wherein one or more processors further execute instructions to access a historical dataset containing observed features.
[0231] (Embodiment 93) The system according to Embodiment 91 or 92, wherein one or more processors further execute instructions to evaluate the variance of predictions in order to build a model for predicting sampling frequencies.
[0232] (Embodiment 94) The system according to any one of embodiments 91 to 93, wherein one or more processors select filtered datasets from a master dataset in order to determine the variance of selected features.
[0233] (Embodiment 95) The system according to any one of Embodiments 91 to 94, wherein one or more processors execute instructions to select a model based on observed features and at least one other distinct feature in the dataset, and to evaluate the variance of predictions for outcomes using a model that uses multiple compensatory values for unmeasured features in the dataset.
[0234] (Embodiment 96) The system according to any one of embodiments 91 to 95, wherein one or more processors further execute instructions to determine rules for evaluating a decision value based on the sampling frequency.
[0235] (Embodiment 97) The system according to any one of embodiments 91 to 96, wherein one or more processors further execute instructions to determine the accuracy of the rules for filling in missing features with first values, based on the model.
[0236] (Embodiment 98) The system according to any one of embodiments 91 to 97, wherein one or more processors further execute instructions to determine the time-dependent variance of the first and second outcomes.
[0237] (Embodiment 99) The system according to any one of embodiments 91 to 98, wherein one or more processors further execute instructions to reduce the sampling frequency of unmeasured features based on the rank of the unmeasured features.
[0238] (Embodiment 100) The system according to any one of embodiments 91 to 99, wherein one or more processors further execute instructions to select a sensor device to collect measurements from unmeasured features based on the accuracy and precision of the sensor devices and the ranking of unmeasured features.
[0239] (Embodiment 101) The method according to any one of Embodiments 1 to 10, wherein the other remaining first unmeasured features are the same as the other remaining second unmeasured features.
[0240] (Embodiment 102) The system according to any one of Embodiments 11 to 15, wherein the other remaining first unmeasured features are the same as the other remaining second unmeasured features.
[0241] (Embodiment 103) A non-transient computer-readable medium according to any one of Embodiments 16 to 20, wherein the other remaining first unmeasured features are the same as the other remaining second unmeasured features.
Claims
1. A method for making a diagnostic decision leading to the treatment of a patient in an emergency care situation, wherein at least one feature in a patient-related dataset is measured, During the emergency care situation, access is made to a diagnostic engine installed on a server communicatively coupled to the client device via an application installed on the client device, wherein the diagnostic engine includes one or more machine learning models trained and configured for the diagnosis and treatment of sepsis. Receiving a first value from the server via the application installed on the client device, which is a first value that has been filled in to the unmeasured features in the dataset while keeping the other remaining unmeasured features in the dataset constant, The dataset includes at least one of the following: time-series measurements of treatment information, outcome information, or measurement information including treatment methods, drug administration events, or dosages; The client device receives from the server, via the application installed on the client device, a first outcome evaluated using one or more machine learning models that use the compensated first value; The client device receives from the server, via the application installed on the client device, a second value that has been interpolated to the unmeasured features in the dataset while keeping the other second remaining unmeasured features in the dataset constant. The client device receives from the server, via the application installed on the client device, a second outcome evaluated using one or more machine learning models that use the compensated second value. The application installed on the client device receives statistical parameters determined using the first outcome and the second outcome from the server, The client device receives a ranking from the server, via the application installed on the client device, which corresponds to the statistical parameters and is assigned to the unmeasured features. Based at least partially on the assigned ranking, propose one or more unmeasured features in the dataset to be measured, A method comprising collecting observations of the patient based on the measurement of one or more previously proposed unmeasured features.
2. The method according to claim 1, further comprising selecting and searching for a filtered dataset from a master dataset according to at least one measured feature from the dataset, before concluding the first value, wherein the master dataset comprises a plurality of datasets related to a plurality of known outcomes.
3. The method according to claim 1, wherein the unmeasured features are obtained by obtaining the assigned ranking corresponding to the statistical parameters by using a model-based feature importance methodology to identify the relative importance of the unmeasured features using one or more known outcomes in a filtered dataset.
4. The method according to claim 1, wherein determining the statistical parameters using the first and second outcomes includes accessing a master dataset containing a plurality of datasets associated with known outcomes.
5. The method according to claim 1, wherein determining the statistical parameters using the first and second outcomes is to determine the variance values associated with a model for the outcome, the model being based on the unmeasured features and at least one other distinct feature in the dataset, and evaluating the variance of predictions for the outcome using the model which uses a plurality of compensated values for the unmeasured features in the dataset.
6. Determining the statistical parameters using the first outcome and the second outcome is: The method according to claim 1, comprising determining a rule for evaluating a decision value based on a dataset, wherein the dataset includes values collected for a plurality of measured features and an unmeasured feature in the dataset, and the rule matches (1) a plurality of known outcomes from a master dataset comprising the plurality of datasets, and (2) one or more measured features.
7. The method according to claim 1, wherein determining statistical parameters using the first and second outcomes includes determining the accuracy of a rule for supplementing the first value with the unmeasured features, based on known outcomes for each of a plurality of outcome values and a plurality of datasets.
8. The method according to claim 1, wherein determining the statistical parameters includes determining the time-dependent variance of the first outcome and the second outcome.
9. The method according to claim 1, further comprising selecting the sampling frequency of the unmeasured features based on the ranking corresponding to the statistical parameters.
10. The method according to claim 1, further comprising selecting a sensor device to collect measurements from the unmeasured features based on the accuracy and precision of the sensor devices and the ranking of the unmeasured features.
11. A system for making diagnostic decisions leading to treatment for a patient in an emergency care situation, wherein at least one feature in a patient-related dataset is measured, A client device includes a memory configured to store instructions, and one or more processors communicatively coupled to the memory, wherein the one or more processors execute the instructions to the system. During the emergency care situation, accessing a diagnostic engine installed on a server communicatively coupled to the client device via an application installed on the client device, wherein the diagnostic engine includes one or more machine learning models trained and configured for the diagnosis and treatment of sepsis, Receiving from the server, via the application installed on the client device, a first value that has been interpolated to the unmeasured features in the dataset while keeping the other remaining unmeasured features in the dataset constant, The dataset includes at least one of the following: time-series measurements of treatment information, outcome information, or measurement information including treatment methods, drug administration events, or dosages; The client device receives from the server, via the application installed on the client device, a first outcome evaluated using a model that uses the compensated first value; The client device receives from the server, via the application installed on the client device, a second value that has been interpolated to the unmeasured features in the dataset while keeping the other second remaining unmeasured features in the dataset constant. The client device receives from the server, via the application installed on the client device, a second outcome evaluated using one or more machine learning models that use the compensated second value. The application installed on the client device receives statistical parameters determined using the first outcome and the second outcome from the server, The client device receives a ranking from the server, via the application installed on the client device, which corresponds to the statistical parameters and is assigned to the unmeasured features. Based at least partially on the assigned ranking, propose one or more unmeasured features in the dataset to be measured, A system configured to collect observations of the patient based on the measurement of one or more previously proposed unmeasured features.
12. The system according to claim 11, wherein one or more processors execute instructions to identify the relative importance of the unmeasured features in a filtered dataset using a model-based feature importance methodology and one or more known outcomes, in order to assign the unmeasured features to a ranking corresponding to the statistical parameters.
13. The system according to claim 11, wherein determining the statistical parameters using the first and second outcomes includes accessing a master dataset containing a plurality of datasets associated with known outcomes.
14. The system according to claim 11, wherein determining the statistical parameters using the first and second outcomes is to determine the variance values associated with a model for the outcome, the model being based on the unmeasured features and at least one other distinct feature in the dataset, and evaluating the variance of predictions for the outcome using the model which uses a plurality of compensated values for the unmeasured features in the dataset.
15. The system according to claim 11, wherein determining the statistical parameters using the first and second outcomes includes determining a rule for evaluating a determined value based on a dataset, the dataset includes values collected for a plurality of measured features and the unmeasured features in the dataset, and the rule matches (1) a plurality of known outcomes from a master dataset comprising the plurality of datasets, and (2) one or more measured features.
16. A non-temporary computer-readable medium that stores instructions and applications for causing the computer to execute a method for diagnostic decision-making leading to the treatment of a patient in an emergency care situation in which at least one feature in a patient-related dataset is measured, wherein the method is During the emergency care situation, accessing a diagnostic engine installed on a server communicably coupled to the computer via the application installed on the computer, wherein the diagnostic engine includes one or more machine learning models trained and configured for the diagnosis and treatment of sepsis, Receiving a first value from the server via the application installed on the computer, which is a first value that has been interpolated to the unmeasured features in the dataset while keeping the other remaining unmeasured features in the dataset constant, The dataset includes at least one of the following: time-series measurements of treatment information, outcome information, or measurement information including treatment methods, drug administration events, or dosages; Receiving a first outcome from the server via the application installed on the computer, which is evaluated using one or more machine learning models that use the interpolated first value; Receiving from the server, via the application installed on the computer, a second value that has been interpolated to the unmeasured features in the dataset while keeping the other second remaining unmeasured features in the dataset constant; The computer receives from the server, via the application installed on the computer, a second outcome evaluated using one or more machine learning models that use the compensated second value. The computer receives statistical parameters determined using the first outcome and the second outcome from the server via the application installed on the computer. The computer receives, via the application installed on the computer, a ranking assigned to the unmeasured features corresponding to the statistical parameters from the server. Based at least partially on the assigned ranking, propose one or more unmeasured features in the dataset to be measured, A non-temporary, computer-readable medium comprising: collecting observations of the patient based on measurements of one or more previously proposed unmeasured features.
17. The non-temporary computer-readable medium according to claim 16, wherein determining the statistical parameters using the first outcome and the second outcome includes accessing a master dataset containing a plurality of datasets associated with known outcomes.
18. A non-temporary computer-readable medium according to claim 16, wherein determining the statistical parameters using the first outcome and the second outcome is to determine the variance value associated with a model for the outcome, the model being based on the unmeasured feature and at least one other distinct feature in the dataset, and evaluating the variance of predictions for the outcome using the model which uses a plurality of compensated values for the unmeasured feature in the dataset.
19. The non-temporary computer-readable medium according to claim 16, wherein determining the statistical parameter using the first outcome and the second outcome includes determining a rule for evaluating a determined value based on a dataset, the dataset includes values collected for a plurality of measured features and the unmeasured features in the dataset, and the rule matches (1) a plurality of known outcomes from a master dataset comprising the plurality of datasets, and (2) one or more measured features.
20. The non-temporary computer-readable medium according to claim 16, wherein determining the statistical parameter using the first outcome and the second outcome includes determining the accuracy of a rule for supplementing the first value to the unmeasured feature based on a plurality of outcome values and known outcomes for each of a plurality of datasets.
21. The method according to claim 1, wherein the other remaining first unmeasured feature is the same as the other remaining second unmeasured feature.
22. The system according to claim 11, wherein the other remaining first unmeasured feature is the same as the other remaining second unmeasured feature.
23. The non-temporary computer-readable medium according to claim 16, wherein the other remaining first unmeasured features are the same as the other remaining second unmeasured features.