Risk model tuning and sampling
The method tunes risk probability models using confusion matrices and adjustment factors to address bias and overfitting, enhancing the accuracy and efficiency of patient recruitment in studies by optimizing sample sizes and reducing variability.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- INSTITUTE FOR SYSTEMS BIOLOGY
- Filing Date
- 2026-01-26
- Publication Date
- 2026-07-30
AI Technical Summary
Existing risk models for patient recruitment in studies are prone to bias, overfitting, and lack generalizability, leading to incomplete and biased findings due to their reliance on risk functions that prioritize patients one-dimensionally, often resulting in disproportionate representation and high cohort sizes.
A method involving a computer-implemented approach to tune risk probability models by generating confusion matrices, adjustment factors, and optimizing classification thresholds to generate adjusted risk probabilities, reducing variance and improving generalization for risk-based sampling.
The tuned risk probability models reduce bias and improve cohort design by providing more accurate sample size estimation, decreasing variability, and reducing study costs while maintaining generalizability.
Smart Images

Figure IMGF000017_0001 
Figure IMGF000017_0002 
Figure IMGF000017_0003
Abstract
Description
RISK MODEL TUNING AND SAMPLINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of U.S. Provisional Patent Application No.63 / 749,969 filed on January 27, 2025. The entire disclosure of the aforementioned application is incorporated by reference herein in its entirety for all purposes.FIELD
[0002] This disclosure relates to risk probability model tuning and methods of use, for example, in methods of risk-based sampling.BACKGROUND
[0003] To accelerate disease prevention, it is valuable to conduct prospective observational studies prior to clinical diagnosis, to understand differences between the trajectories of those who end up with new-onset disease and those who do not. However, to have a sufficiently powered study, a minimum number of patients must transition within the study time period. If incidence rates within the study population are low, a large number of subjects are required to obtain meaningful number of new onset cases (refs. 1-3). Studies such as UK Biobank and NIH All of US partially address these challenges by following extremely large cohorts across a vast range of outcomes, but observations are limited to subsets of less costly observations (genomics, surveys and electronic health records), while deeper phenotyping is only obtained for smaller subcohorts (refs. 4-5). Another alternative to reducing cohort size is to filter participation using risk scores from clinical models, such as selecting subjects with a CHA2DS2-VASc score above a certain score for stroke studies (ref. 6), or subjects carrying a predisposing Alzheimer’s disease variant (ref. 7). However, clinical risk scores exist for a limited number of conditions, and often have limited sensitivity and specificity. More recently, polygenic risk models have been proposed for predictive enrichment in clinical trials (ref. 8) and machine learning has been proposed for prognostic enrichment (ref.9).
[0004] It is possible to use risk scores from any performant, trustworthy risk model to identify patients at higher risk of a desired outcome to enrich the study population, thusimproving the likelihood of gaining new insights from the study results with a smaller cohort size and reduced costs. At the same time, it is also important to note that selecting patients based on risk can introduce bias (ref. 10), both intentional and unintentional, into the recruitment process. With conscious bias, subjects are selected based on predetermined inclusion / exclusion criteria. Use of a clinical risk model may also introduce bias if the model unintentionally privileges certain groups over others, sometimes referred to as algorithmic bias (refs. 11–14). While this could lead to disproportionate representation among different populations, it can be offset by using a stratified recruitment strategy across demographic categories to ensure appropriate representation among patient groups (ref. 15). Unconscious bias can occur when certain groups of people are unintentionally excluded from a study. While it is not always possible to recognize and correct unconscious bias, steps can be taken to assess for disproportionate representation among a wide variety of factors, including demographics, comorbidities, and health-related exposures.
[0005] Unfortunately, risk models often rely heavily on risk functions to prioritize patient recruitment. This tends to create a one-dimensional view and can lead to biased or incomplete findings. Various techniques have been used to reduce the uncertainty in the risk function, including calibration of the output probabilities, such as by Platt scaling or isotonic regression. However, such techniques remain prone to overfitting, are less interpretable, and sensitive to noisy training data. Another problem is they tend to not generalize well on unseen or prospective data, obscuring true relationships between a model’s scores and actual probabilities.SUMMARY
[0006] In some embodiments, a computer-implemented method of producing a risk probability model tuned for risk-based sampling is provided. The method includes:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a retrospective population of subjects having known real outcomes for the event over a given period of time;(b) generating a confusion matrix at a given classification threshold for the risk probabilities;(c) generating adjustment factors each individually normalizing an outcome of the confusion matrix as a performance metric of the risk probability model at the given classification threshold;(d) generating adjusted risk probabilities by matrix multiplication of the adjustment factors and the accessed risk probabilities so as to bound the accessed risk probabilities by the performance metric;(e) determining an expected number of events in the retrospective population as an aggregate or rate of the adjusted risk probabilities across the retrospective population in a probability distribution model;(f) determining a difference between the expected number of events in the retrospective population and the real number of events in the retrospective population;(g) generating an optimized adjustment factors and classification threshold pairing for the risk probability model by repeating steps (b) - (f) with one or more different classification thresholds and selecting an adjustment factors and classification threshold pairing that generates a reasonable minimum between the expected and real number of events in the retrospective population; and(h) configuring the risk probability model with the adjustment factors for model persistence by executing a model update protocol stored in a non-transitory computer-readable medium, wherein the update protocol is designed to improve model generalization and reduce variance in subsequent risk predictions.
[0007] In some embodiments, a computer-implemented method of risk-based sampling is provided. The risk-based sampling method includes:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a target population of subjects over a given period of time;(b) applying adjustment factors to the accessed risk probabilities, wherein each adjustment factor is a proportion of an outcome of a confusion matrix generated for a given classification threshold applied to risk probabilities generated by the risk probability model for a retrospective population of subjects having known actual outcomes for the event over the given period of time, the proportion reflecting how well the risk probability model predicts anoutcome at the given classification threshold, the outcome selected from true positives, true negatives, false positives, and false negatives;(c) generating adjusted risk probabilities based on application of the adjustment factors to the accessed risk probabilities by matrix multiplication so as to bound the risk probabilities by performance of the risk probability model for all real events given all predicted events of the retrospective population; and(d) determining an expected number of events in the target population based on application of a probability distribution to the adjusted risk probabilities, the probability distribution selected for determining an expected number of events in the retrospective population; and optionally,(e) determining a sample size of the target population containing a desired number of subjects expected to exhibit the event at some point over the given period of time, the desired number of subjects approximately equal to or less than the expected number of events determined for the target population.
[0008] In some embodiments, a system is provided that includes one or more data processors and a non-transitory computer-readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein.
[0009] In some embodiments, a computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods or processes disclosed herein.
[0010] In some embodiments, a system is provided that includes one or more means to perform part or all of one or more methods or processes disclosed herein.BRIEF DESCRIPTIONOF THE DRAWINGS
[0011] The accompanying Figures and Examples are provided by way of illustration and not by way of limitation. The foregoing aspects and other features of the disclosure are explained in the following description, taken in connection with the accompanying example figures (also “FIG.”) relating to one or more embodiments as provided herein.
[0012] FIG. 1: Schematic of the Clinical Recruitment Optimization Pipeline (CROP) adjustment procedure. Initial estimates of the number of patients in each class are calculated according to eq. (3) using the raw model output pi = P(Yp = 1|xi). These initial estimates are then adjusted using the confusion matrix evaluated at threshold t as in eq. (11). This effectivity redistributes the number of patients according to how well, or poorly, the model can place patients into each class, as in the heuristic eq. (10).
[0013] FIG. 2: General procedure for implementing the CROP pipeline illustrated herein. Briefly, a risk probability model is trained, for example, a machine learning (ML) model trained using standard techniques such as k-fold cross validation on a training dataset (sometimes referred to as the retrospective population, tuning cohort, or test set). Next, the model is evaluated on the tuning cohort to obtain the model probabilities, which are used to generate a confusion matrix at an initial probability threshold, followed by matrix multiplication with the model probabilities to generate adjusted pseudo-probabilities. A probability distribution model is then applied to generate an expected number of events. The process is repeated at different thresholds until the expected number of events predicted by the process equals the actual number of events observed in the tuning cohort. Adjustment factors producing an acceptable number of expected events are considered optimized. These optimized adjusted factors are then applied to model probabilities of a target cohort having unknown outcomes so as to generate adjusted pseudo-probabilities for the target cohort (sometimes referred to as the prospective cohort). The adjusted pseudo-probabilities are then summed using descending order summation until the sum is equal to the desired number of events, and / or ascending order summation until the sum is equal to the desired number of non-events. The resulting count of the number of adjusted pseudo-probabilities is equal to the desired sample size.
[0014] FIG. 3: Various performance metrics relative to CROP. A. Performance metrics as a function of the threshold on the tuning set. Values calculated on a grid of 10000 points between.00001 and.99999. CROP is the only metric which achieves its minimum at a reasonable value of the threshold (threshold =.32) in this illustration. B. Comparing deterministic (counting # of predicted positives) and CROP approaches to estimating the number of transitions with adjusted and unadjusted. The deterministic approach is sensitive to boundary effects in addition to being deterministic. CROP, on the other hand, is purely probabilistic and accurately predicts the correct number of transitions are reasonable values ofthe threshold. Probabilities were adjusted using the confusion matrix of the tuning set and applied to the probabilities of the CROP set.DETAILED DESCRIPTION
[0015] As summarized above, a computer-implemented method of producing a risk probability model tuned for risk-based sampling is provided. Also provided is a computer-implemented method of risk based sampling, and in particular, applying a tuned risk probability model of the disclosure in a method of determining an expected number of events in a target population, and more particularly, a sample size needed to obtain a desired number of subjects expected to exhibit the event over a given period of time. The subjects can range broadly, for example, from abiotic subjects, such as a medical device or part thereof, to living subjects, such as human subjects or part thereof, depending on the intended end use. Preferred subjects are individuals or patients for clinical trials, such as retrospective, observational, or prospective clinical trials, as exemplified herein. Further provided are systems and computer-program products for carrying out the methods.
[0016] The disclosure is exemplified by a method referred to in the Examples as CROP to tune machine learning risk probability models for non-deterministic cohort enrichment. As an illustrative case study, a machine learning model for chronic kidney disease (CKD) risk in patients with type 2 diabetes (T2D) was developed and tested using CROP to define an emulated cohort, and approaches for assessing and addressing the potential impact on study generalizability were demonstrated. The method is broadly applicable for risk-based sampling using various risk probability models in general, including ones that are generated de novo or pre-existing. For example, CROP works with any risk probability model that generates risk scores including, but not limited to, statistical models (e.g., logistic regression, parametric models, mixture models), machine learning models (e.g., classification algorithms, Bayesian networks), and hybrid models (e.g., statistical models with expert input). More generally, the optimized solution for prognostic enrichment as disclosed herein will be valuable across many fields in addition to clinical studies, ranging from risk-based safety inspections for imported agricultural goods (ref. 16) to pharmaceutical manufacturing (ref. 17).I. Model Tuning
[0017] The computer-implemented method of producing a risk probability model tuned for risk-based sampling includes:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a retrospective population of subjects having known real outcomes for the event over a given period of time;(b) generating a confusion matrix at a given classification threshold for the risk probabilities;(c) generating adjustment factors each individually normalizing an outcome of the confusion matrix as a performance metric of the risk probability model at the given classification threshold, for example, such as where the performance metric is selected from positive predictive value (PPV), false omission rate (FOR), false discovery rate (FDR), negative predictive value (NPV), and combinations thereof;(d) generating adjusted risk pseudo-probabilities by matrix multiplication of the adjustment factors and the accessed risk probabilities so as to bound the accessed risk probabilities by the performance metric, for example, such as where the adjusted risk pseudoprobabilities are generated by matrix multiplication of positive predictive value (PPV) and the false omission rate (FOR), the false discovery rate (FDR) and negative predictive value (NPV), or a combination thereof;(e) determining an expected number of events in the retrospective population as an aggregate or rate of the adjusted risk pseudo-probabilities across the retrospective population in a probability distribution model, for example, such as where the probability distribution model is Poisson binomial and the aggregate or rate is a sum;(f) determining a difference between the expected number of events in the retrospective population and the real number of events in the retrospective population;(g) generating an optimized adjustment factors and classification threshold pairing for the risk probability model by repeating steps (b) - (f) with one or more different classification thresholds and selecting an adjustment factors and classification threshold pairing that generates a reasonable minimum between the expected and real number of events in the retrospective population; and(h) configuring the risk probability model with the adjustment factors for model persistence, such as by deploying or storing the adjustment factors in a memory or a storage inassociation with the risk probability model, and more specifically, by executing a model update protocol stored in a non-transitory computer-readable medium, wherein the update protocol is designed to improve model generalization and reduce variance in subsequent risk predictions.
[0018] As can be seen, the adjustment factors are proportions applied to the actual risk probabilities generated from a base risk probability model. This means the tuned risk probability model is essentially a base model that includes a separate operator or function step after the model generates a raw probability. Thus, by “tuned risk probability model” is intended to mean any risk probability model capable of generating raw, unadjusted probabilities as well as the adjusted risk pseudo-probabilities. This includes a risk probability model having a dedicated mechanism for adjusting the model raw risk probability outputs without modifying the core model itself. Examples of the dedicated mechanism include, but are not limited to, a separate function or operation used for adjustment of the raw probabilities, regardless of whether it is built-in, linked to, or accessed by the base risk probability model.
[0019] In many embodiments the risk probability model for tuning is a classification risk probability model, preferably a binary classification risk probability model, and the probability distribution selected for the retrospective population is selected from normal, binomial, Poisson, Poisson binomial, negative binomial, and negative Poisson binomial. Of particular interest is where the probability distribution selected for the retrospective population is Poisson binomial.
[0020] In other embodiments, the base risk probability model for tuning is a multinomial classification risk probability model, and the probability distribution selected for the retrospective population is selected from multinomial, Poisson multinomial, negative multinomial, and negative Poisson multinomial, with the Poisson multinomial being of specific interest.
[0021] In certain embodiments, error estimates are applied to assess and select a statistical model employed in the CROP method for generating the expected number of events. For example, error estimates run on both balanced datasets (e.g., 1:1 distribution for two classes in the dataset) and imbalanced datasets (e.g., 1:9 distribution for two classes in the dataset) for various statistical models reveal that certain probability distribution models generalize well in the CROP method. The Poisson and Poisson-Binomial models in particular have very similarerror distributions indicating that either is sufficient for default use on most datasets having arbitrary class imbalances.
[0022] In some embodiments, the method further includes incorporating a class imbalance parameter to adjust for imbalances in a dataset to improve error estimation. This approach involves weighting predictions or errors by class prevalence to ensure the error estimate reflects the true performance across all classes. For example, the class imbalance parameter can be generated or by estimating a single class imbalance parameter from test data. Normalization can also be applied to improve model prediction and error estimation on imbalanced datasets, as described further below and in the examples.II. Risk-Based Sampling
[0023] Once the risk probability model is tuned, it is optimized for determining the expected number of events for various given end uses. The tuned risk probability model is preferably applied in a risk-based sampling process for determining a sample size, such as a minimal sample size needed for a study, such as a trial, assessment, or test. This embodiment generally includes:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a target population of subjects over a given period of time;(b) applying adjustment factors to the accessed risk probabilities, wherein each adjustment factor is a proportion of an outcome of a confusion matrix generated for a given classification threshold applied to risk probabilities generated by the risk probability model for a retrospective population of subjects having known actual outcomes for the event over the given period of time, the proportion reflecting how well the risk probability model predicts an outcome at the given classification threshold, the outcome selected from true positives, true negatives, false positives, and false negatives;(c) generating adjusted risk pseudo-probabilities based on application of the adjustment factors to the accessed risk probabilities by matrix multiplication so as to bound the risk probabilities by performance of the risk probability model for all real events given all predicted events of the retrospective population; and(d) determining an expected number of events in the target population based on application of a probability distribution to the adjusted risk pseudo-probabilities, the probability distribution selected for determining an expected number of events in the retrospective population.
[0024] A featured embodiment is where the risk-based sampling process (a) - (d) above further comprises:(e) determining a sample size of the target population containing a desired number of subjects expected to exhibit the event at some point over the given period of time, the desired number of subjects approximately equal to or less than the expected number of events determined for the target population. A featured aspect is where the determining the sample size in step (e) further comprises: generating (i) a sum of the adjusted risk probabilities until the sum adds up to the desired number of subjects, and (ii) a count corresponding to a number of the adjusted risk pseudo-probabilities used to obtain the sum, wherein the count is the sample size. Of particular interest is where the summation is in descending order for expected events (e.g., responders, positive controls etc.), ascending order for expected non-events (e.g., nonresponders, negative controls, etc.), or a combination thereof for determining expected and nonexpected events.
[0025] In certain embodiments, the adjustment factors and the given classification threshold is an optimized pairing for the risk probability model and the retrospective population such that the expected number of events in the retrospective population is approximately equal to an actual number of real events in the retrospective population over the given period of time. An example is where in the risk-based sampling process (a) - (e) above, the target population is the retrospective population, and wherein the optimized pairing is produced by:(f) generating a performance metric for the risk probability model determining an absolute difference between the expected and actual number of events occurring in the retrospective population; and(g) generating an optimized adjustment factors and classification threshold pairing for the risk probability model by repeating steps (b) - (d) and (f) with one or more different adjustment factors and classification threshold pairings and selecting a paring that generates a reasonable minimum for the performance metric.
[0026] In as many embodiments, the desired number of subjects in the sample is an enrichment selected from decreased variability enrichment, prognostic enrichment, and predictive enrichment. For example, the target population can be enriched for patients in which the potential effect of an intervention or no intervention on a population can more readily be examined or demonstrated. Of particular interest is where the enrichment further comprises stratifying the subjects by one or more enrollment features. In certain embodiments, the present method finds use in, or benefits by, enrichment of the target population so as to decrease variability between subjects in the population. In related embodiments, the method can be applied for prognostic enrichment (will someone experience a given medical outcome yes or no, such as disease or other outcome), or for predictive enrichment (whether or not people will benefit from an intervention, such as a therapeutic intervention).
[0027] As noted above, the adjustment factors are preferably optimized for a given end use. They can also be un-optimized or previously optimized but outdated, and then periodically optimized, updated, or otherwise revised or kept as-is, for example, on an as needed basis, as is practical, or for reference. Of particular interest is application of a dynamic adjustment approach to periodically refine the model’s adjustment factors based on new data, and in some embodiments, perpetually or continually adjust the adjustment function, thereby maintaining the reliability of the model’s predicted probabilities.
[0028] For example, in certain embodiments, the risk-based sampling method includes: accessing risk probabilities generated by a risk probability model; applying adjustment factors to the accessed risk probabilities; generating adjusted risk pseudo-probabilities based on application of the adjustment factors to the accessed risk probabilities; and determining an expected number of events of interest in the target population. In certain embodiments, this method includes the steps:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a target population of subjects over a given period of time;(b) applying adjustment factors to the accessed risk probabilities, wherein each adjustment factor is a proportion of an outcome of a confusion matrix generated for a given classification threshold applied to risk probabilities generated by the risk probability model for a retrospective population of subjects having known real outcomes for the event over the givenperiod of time, the proportion reflecting how well the risk probability model predicts an outcome at the given classification threshold, the outcome selected from true positives, true negatives, false positives, and false negatives;(c) generating adjusted risk pseudo-probabilities based on application of the adjustment factors to the accessed risk probabilities by matrix multiplication so as to bound the risk probabilities by performance of the risk probability model for all real events given all predicted events of the retrospective population;(d) determining an expected number of events in the target population based on application of a probability distribution to the adjusted risk pseudo-probabilities, the probability distribution selected for determining an expected number of events in the retrospective population; and optionally,(e) determining a sample size of the target population containing a desired number of subjects expected to exhibit the event at some point over the given period of time, the desired number of subjects approximately equal to or less than the expected number of events determined for the target population; and / or optionally,(f) generating a performance metric for the risk probability model determining an absolute difference between the expected and actual number of events occurring in the retrospective population; and(g) generating an optimized adjustment factors and classification threshold pairing for the risk probability model by repeating steps (b) - (d) and (f) with one or more different adjustment factors and classification threshold pairings, selecting a paring that generates a reasonable minimum for the performance metric, and updating the risk probability model with the optimized adjustment factors or the optimized pairing.
[0029] When updating the model, the target population is typically the retrospective population so that the adjustment factors and the given classification threshold generate an optimized pairing. In other embodiments, the target population is a prospective population, a contemporary population for recruitment, or a combination thereof. In as many embodiments, the retrospective population and the target population have substantially the same inclusion and exclusion criteria.
[0030] In a featured embodiment, the adjustment factors are derived from a column normalized confusion matrix for the given classification threshold. In as many embodiments,the event is a single event or plurality of events, and can be driven by discrete or continuous factors.
[0031] In one embodiment, the method of determining the sample size further includes: outputting one or more of the model probabilities, the classification threshold, the confusion matrix, the adjusted risk probabilities, the expected number of events, the desired number of subjects, the sample size, or a combination thereof, for use in a decision-making process. An example of particular interest is where the subjects are humans, the event is a medical outcome, and the decision-making process is selection of patients for the medical outcome, such as where the medical outcome is new onset CKD. A featured example is where the risk probability model is for identifying type 2 diabetes patients at high risk for CKD transition, the retrospective population and the target populations are individuals with type 2 diabetes, and the event is a transition to CKD. An additional embodiment is where the method is applied to determine controls for a target population, for example, so as to determine subjects that do not exhibit the event over the given period of time, or the event itself is defined as the control. One example is where the number of control subjects is the difference between the expected number of events and total subjects in the target population, or a ratio thereof applied to a subset of the target population. As such, in certain embodiments, the event is a control characterizing a normal subject or group of subjects having a non-event outcome over the given period of time. In other embodiments, the control is a sample of subjects having intermediate to low risk of exhibiting an event over the given period of time.
[0032] In some embodiments, the methods of the disclosure further include normalization constants for the confusion matrix to ensure that results are normalized to address class imbalance. In other embodiments, a fictitious “confusion matrix” can be constructed on retroprospective cohort data, which can then be used to construct the adjustment factors and the adjusted risk pseudo-probabilities.III. Systems and Computer-Program Products
[0033] Computer systems and computer-program products are provided. The systems include: one or more data processors; and a non-transitory computer readable storage medium containing instructions which, when executed on the one or more data processors, cause the one or more data processors to perform part or all of one or more methods disclosed herein. Thecomputer-program products in general are tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein. Of particular interest are hardware systems, software systems, and computer programs.
[0034] Examples of hardware systems include, but are not limited to, general-purpose computers (e.g., desktops, laptops, servers), mobile devices (e.g., smartphones, tablets), embedded systems (e.g., microcontrollers in devices), networked systems (e.g., connected computers and devices), cloud computing platforms (e.g., providing processing power and storage for remote model deployment at scale).
[0035] Examples of software systems include, but are not limited to, operating systems (e.g., for managing computer resources), application software (e.g., for performing specific tasks for users), middleware (e.g., for connecting different software applications), database management systems (store and organize data), and machine learning frameworks (e.g., tools for building, training and deploying models).
[0036] Examples of computer-program products, include, but are not limited to, software applications (e.g., functional programs with a user interface), libraries and modules (e.g., reusable code components), scripts and algorithms (e.g., instructions for the computer to perform tasks), graphical user interfaces (GUIs) for interacting with software, and machine learning models (e.g., trained models for prediction or classification).IV. Utility and Advantages
[0037] An advantage of the CROP approach is that it can employ a combination of modem machine learning and classical statistical reasoning. When investigating complicated, individualized questions such as will a given patient be diagnosed with disease X, or will a given patient respond to drug Y, the answer depends on the interactions between many features, such as genetics, environmental factors, medical history, physiology, lifestyle choices and more. Modern machine learning methods have proved superior to classical statistics for such questions, primarily due to their flexible architectures, universal approximation-like theorems, and efficient training algorithms. However, when extrapolating these results to cohort-level questions, such as how many patients will be diagnosed with disease X, or how many patients will respond to drug Y, statistical models are better suited as machine learning techniques tendto underestimate, owed to deterministic decision rules used to determine which patients experience the “event” in question. The combination of machine learning models with statistical models therefore provides the CROP approach with a significant advantage over either approach alone.
[0038] As such, the methods, systems and computer-program products have many applications and advantages. As noted above and exemplified herein, the tuned risk probability models of the disclosure find particular use in clinical trials and cohort enrichment. A benefit of cohort enrichment is decreased variability, improved prognostic enrichment, or predictive enrichment, while avoiding the problems associated with the deterministic nature of risk probability models that have not been tuned in accordance with the disclosure. One reason is that the tuned risk probability models take into account the true bounds of model performance, which reduces bias and provides more control over cohort design and selection. This results in reduced time and cost, all the while improving study interpretation.
[0039] For example, decreased variability enrichment using a tuned risk probability model of the disclosure can be exploited to increase study power, such as by improving how patients with baseline measurements of a disease are chosen. Another is to decrease variability in a target population by improving the selection of patients with a biomarker characterizing a given disease in a narrow range.
[0040] A tuned risk probability model of the disclosure can be applied for prognostic enrichment, which includes choosing patients with a greater likelihood of having a disease-related endpoint or a substantial worsening in condition for continuous measurement endpoints. Prognostic enrichment in this way increases the absolute effect difference between groups but would not be expected to alter the relative effect.
[0041] Predictive enrichment with a tuned risk probability model of the disclosure includes using the model to choose subjects that are more likely to respond to a drug treatment. This approach can lead to a larger effect size with a smaller study population. Examples of predictive enrichment include selection based on a given aspect of a patient's physiology, a biomarker, or a disease characteristic related to drug mechanism. Patient selection could also be empiric in that the patient has previously appeared to respond to a drug in the same class.
[0042] A tuned risk probability model of the disclosure can also be used to enrich for various characteristics in a cohort. This includes enrichment characteristics that aredichotomous or continuous. Examples of dichotomous characteristics include sex, presence of genetic marker(s), and a concomitant illness. Examples of continuous characteristics include age, blood pressure and the like. Stratification of these and other variables known at the outset can be used to balance between minimizing sample size and maintaining generalizability.Although biases can occur in initially unmeasured variables, a thoughtful choice of risk score and enrichment or exclusion, followed by assessments before and after can improve outcomes.
[0043] As such, a tuned risk probability model of the disclosure finds particular use in prospective studies so as to provide essential data for accelerating research on disease prevention, including prospective multi omic studies with deep phenotyping. For instance, by recruiting a cohort of subjects who do not yet have the clinical diagnosis of interest, researchers can analyze differences in trajectories between those who go on to develop the disease and those who do not. By reducing cohort size while optimizing the number of subjects that will transition to disease, the trial cost can be reduced, thereby making research on pre-onset studies and prevention or intervention more practical. As such, prognostic enrichment using a tuned risk probability model of the invention can reduce the cohort size needed by recruiting subjects at higher risk for an outcome of interest, and higher-performance machine learning offers the potential for even greater enrichment.
[0044] As exemplified below, the CROP process provides a formal way of minimizing sample size for any study using any risk probability score. As a working example, CROP was tested on real-world data through emulation of a prognostically-enriched study for new-onset CKD in patients with type 2 diabetes. Specifically, a risk probability-generating ML model was developed for predicting new-onset CKD in patients with diabetes, which had AUROC 0.826. Application of CROP to this model resulted in a 5.3-fold decrease in cohort size compared to random sampling. Thus, CROP with risk scores from any risk model can provide prognostic or predictive enrichment, reducing the sample size needed for studies to advance understanding of many types of transition processes.EXAMPLES
[0045] The following examples are provided to illustrate certain particular features and / or embodiments. These examples should not be construed to limit the disclosure to the particular features or embodiments described.Example 1 - CROP for risk enrichment in a clinical studyCROP and the statistical model
[0046] The appropriate statistical model is determined by the question being asked. For application of CROP to a clinical study, the question is “What is the minimum cohort size needed to observe at least X “events” during a clinical trial while maintaining a fixed level of statistical power?” Here the “event” is the binary outcome of interest of the clinical trial. This can be considered an inverse problem to the following common statistical question: “Given a sample of size N, how many “events” do we expect to see?” If the “event” in question occurs with probability p, then the expected number of events is Np with variance Np(1−p). This is the well-known Binomial distribution, usually phrased in terms of counting the number of heads that occurs in a series of coin flips. The specific form of the distribution isP(K = k) = (N choose k) pk(1-p)N-k(1)
[0047] In the CROP framework, the probability p would be found using a discriminative machine learning model across each outcome, as we will discuss in the next subsection. The real case, however, is not so simple, since each individual’s risk probability is unique and depends on the patient’s individual features. In that case, we instead use the Poisson-Binomial distribution, whose functional form is much more difficult to work with:(2)
[0048] where Fkis the set of all subsets of k integers, and piis the ithpatient’s probability of an event occurring. This sum over all subsets is the combinatorial equivalent of the binomial coefficient and eq. (2) reduces to eq. (1) when all the / s are equal. The mean and variance of eq. (2) are피[K] = Σ_{i=1}^{N} p_i (3)(4)
[0049] This formulation of our problem allows us to not only estimate the number of transition we expect to occur, but also the fluctuations we expect, allowing one to control for these fluctuations when determining the final sample size.
[0050] Under certain circumstances the Poisson-Binomial model can instead be approximated by Poisson distribution with the same mean. This is known as Le Cam’s theorem and can be thought of as a generalization of the law of rare events. This is particularly useful for cohorts in which the probability of transitioning is small and, as we will see below, allows for smaller distributional error due to the larger variance.
[0051] The estimates in eq. (3) and eq. (4) provide an answer to the inverse problem posed above, where we asked how many events occur in a cohort of size N. To answer our CROP question, we first arrange the probabilities in descending order, so that p1≥ p2≥ ··· ≥ pN. To obtain the minimal sample size M needed to reach a fixed number of events X, we must solve the following equation:&{Yu - xY (5)
[0052] In other words, sum the probabilities in descending order such until the desired number of events is reached. These M patients are the cohort to recruit for the clinical trial.CROP and the machine learning risk model
[0053] Solving eq. (5) to obtain M is the primary purpose of the clinical trial example. However, obtaining appropriate patient-specific probabilities piis a non-trivial task, even with the breadth of tools modern machine learning has to offer.
[0054] This is primarily driven by imperfect model performance (as quantified by the confusion matrix) and ill-calibrated probabilistic outputs. Here, in addition to laying out the general procedure by which we obtain these estimates, the CROP methods provides a solution for calibrating model probabilities such that a solution to eq. (5) can be obtained, even for models with poor performance, and conditions under which the procedure is valid. Specifically, an assumption is made that the output of a given discriminative machine learning model represents the true probabilities, or at least attempts to. This is guaranteed when the loss function is the binary cross entropy, as this is equivalent to negative log-likelihood of themodel. Otherwise, one must calibrate the probabilities prior to the CROP procedure using methods such as isotonic regression or Platt scaling so as to faithfully represent the true probabilities.
[0055] Let D c X x Ybe a data set with an arbitrary feature space X and a fixed, discrete label space Y ~= {0,1} denoting the binary outcome of interest. The data set is partitioned into three subsets: D = DtrainxDtestxDval. The parameters are learned on the training set Dtrain, the parameters are tuned on the test set Dtest, and the overall performance is evaluated on the validation set Dval. The result is a model that yields the patient-specific probabilities pi= P Y= ]\X=xi)=Axi-0) (6)
[0056] where f(x;θ) is a (possibly discontinuous) function fixed by the model architecture with parameters θ.
[0057] At this point in the model building process, one picks a model threshold / which defines the confusion matrix of the model. This partitions the feature space into two (possibly disconnected) regions X = XoxXi defined by the implicit equation. / (7)
[0058] Under a deterministic decision rule, samples that fall in the region Xo are assigned to the Y = 0, and samples that fall in the region Xi are assigned Y = 1. One uses these assignments to construct the confusion matrix of the model, which, when normalized, can be interpreted as the following conditional probabilities.P(Yp|Yr) (row normalized) (8)P(Yr|Yp) (column normalized). (9)
[0059] These matrices play a central role in CROP (see FIG. 1 and FIG. 2).CROP and model tuning
[0060] We are now in a position to apply CROP for model tuning. Let Nr=i and Np=i denote the number of patients with Yr= 1 and Yp= 1, respectively, and similarly for Yi= 0. Heuristically, CROP performs the following adjustment (see FIG. 1 and FIG. 2).Nr=i = (%Np=i correctly classified)xNp=i + (%Np=o incorrectly classified) x Np=o (10)In terms of the elements of the confusion matrix, this isNr=1 = Pt(Yr = 1|Yp = 1) × Np=1 + Pt(Yr = 1|Yp = 0) × Np=0 (11)Using eq. (3), eq.(l 1) can be formulated asN_{r=1} = Σ[P_t(Y_r=1|Y_p=0)P(Y_p=1|x_i) + P_t(Y_r=1|Y_p=0)P(Y_p=0|x_i)]X7 •. -(|2)which motivates the following definition:Q(Y_r=1|x_i) = Σ_{Y_p∈{0,1}} P_t(Y_r=1|Y_p)P(Y_p|x_i) (13)
[0061] The object in eq. (13) represents the model adjusted pseudo-probabilities.
[0062] One benefit of this approach is that we can now tune the model threshold t until our estimate in eq. (12) equals the true number of patients in the test set. Let |D¹_test| be the number of patients Y = 1 in the test set. The CROP tuning procedure performs the following minimization:( \2iptej~ 52 52 u(u = |>*i>totW,i} } (14)
[0063] Once the threshold has been tuned, the resulting confusion matrix Pt*(Yr\Yp) can be used to compute pseudo-probabilities for a new cohort of patients given by {zi \I = l,..., JVnew} c X. Plugging these pseudo-probabilities into the estimates in eqs. (3) and (4), thus yields a distribution of possible values for the number of events in this unseen cohort. These pseudo probabilities can then be used as in eq. (5) to find the minimum cohort size needed to observe a fixed number of events in the cohort. In CROP, we use the deterministic estimate of Np=i to build the confusion matrix, but use the Poisson-Binomial estimate N̂p=1= Σpito obtain the class estimates. Thus, the tuning procedure defined in eq. (14) is equivalent to finding a threshold t* such thatN_{p=1} = Σ_i Θ(p_i ≥ t*) = Σ_i p_i = N̂_{p=1}T (15) CROP and class imbalance
[0064] While the CROP approach is insensitive to class imbalance between the two classes, it is sensitive to differences in class imbalance between the data used to construct the confusion matrix and the new data set on which you wish to estimate. If we keep the confusion matrix constant, and plug in new values for the estimateNp " we will obtain(16)
[0065] To assess the influence of class imbalance on this quantity, we write this in terms of the true positive rate λ (sensitivity) and the false positive rate μ (specificity) of the model, as they show no bias when dealing with imbalanced data sets. We also let Nr=i = aN where a is a class imbalance parameter. Thus, we can rewrite eq. (16) as. So(Tynf-u: \ / ^rtew \+ tl - A)a. V■V- ■ / ' \ A' - / (17)
[0066] Assuming the model will predict the same fraction of total samples into each class, thenAj-neiu \rnew \rnewi yp=l __ _ ‘V,v...;where plugging this into eq. (17) generates- V. A (v) + (i - w (v= AaA?"ew+ (1 - A)oW’!6“as expected. However, if the a of the new dataset differs substantially from the a of the test set, this will lead to an incorrect estimate. If, however, we were able to estimate anewof the test set, we could construct a fictitious “confusion matrix” on the new data set and use it to construct our pseudo-probabilities in eq. (12) so that we may estimate the number of Y = 1 patients on arbitrary subsets of the data. Moreover, we must only construct such a matrix using the only information available to use: the test set, and the raw model outputs of the new data set. This motivates the following definition of:PAf'X°nCW NneW / ^i - X)ancwNnew / Ar| p“ V M (1 - / / .)(! - anew)N™w / N™% - a,,ew)Nncw / N^ J Qg)where we include normalization constants to ensure that each column is normalized, thus preserving the total population size of the estimates. Plugging this into eq. (11) with our Poisson-Binomial estimates yields the following:= M Acv,lf:wAne" A'o (1 - A)anR,'-Wne’B(19) for the estimates, andI:.?';? V P ff,(20) for the adjusted pseudo-probabilities.Example 2 - Application of CROP to a machine learning model for CKD
[0067] The CROP procedure is illustrated using a risk model developed to identify type 2 diabetes patients at risk of transitioning to chronic kidney disease (CKD) within three years. The specific terms of interest for adjusting the model probabilities areMT, ~ ™ 1) ~ wed&ttw vaMe (FPF)F(F ~ li? - 0) ~ a?mssuyn >w (, Wir«s» sij'sfdP('T. ~ GIF, — 1) — False disewery rats (FOR) rsSisw ' 'MT GIT, sa 0) ~ miw fAFFiwhich represent the adjustment factors used to generated the model adjusted pseudoprobabilities, which can be written asP(T, ~ i[X.) FFF - T FOR - (1 - tOF(T 0h:,) ' P, * ~ F Jwhere the adjustment factors for a probability of a transition event (PPV and FOR) and the probability of a non-transition event (FDR and NPV) are applied by matrix multiplication to the unadjusted probabilities pt so as to bound the risk probabilities by the real events given the predicted events.
[0068] As such, the new, adjusted pseudo-probabilities for transitioning are bounded from above by the positive predictive value (PPV) and from below by the false omission rate (FOR). Conversely, the adjusted probabilities for not transitioning are bounded from below by the false discovery rate (FDR) and from above by the negative predictive value (NPV).CKD study protocol
[0069] A retrospective study protocol was performed in compliance with the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule and approved by the Institutional Review Board (IRB) at Providence St Joseph Health (PSJH) with Study NumberSTUDY2022000248. Consent was waived because disclosure of protected health information for the study was determined to involve no more than a minimal risk to the privacy of individuals.CKD study setting and participants
[0070] This retrospective study analyzed electronic health record (EHR) data selected from 28 million patients in PSJH, a community health system with 51 hospitals and 1085 clinics across seven states in the United States: Alaska, California, Montana, New Mexico, Oregon, Washington, and Texas. The cohort in the current study included patients > 18 years old with an active Type 2 Diabetes diagnosis defined by a set of SNOMED-CT codes, up to May 31, 2019, hereafter referred to as the study start date, tO. Patients were excluded that had a history of or any current diagnosis of Type 1 Diabetes, CKD, or renal impairments with the exception of certain diagnoses of acute kidney injury. Patients were also excluded if they were hospitalized during the time of the study start date or were actively pregnant or gave birth within six weeks of the study start date. Additionally, because it was desirable to use baseline data from medical history in the model, only patients were included who had a previous outpatient encounter within three years of the study start date where blood pressure and heart rate were recorded, and at least one laboratory value for HbAlc and serum creatinine (both routinely collected for patients with diabetes), within three years of the study start date. Patients who did not have at least one encounter before their recorded noted date for T2D were excluded. Several clinical features were included for demographics, laboratory tests, medications, medical history, and vital signs. The time windows used to define variables are standard, and described in greater detail below. A machine learning algorithm was trained to predict the primary outcome of transition to CKD (but not death) within three years (2019-06-1 to 2022-06-01) and the model was validated on a 10% hold-out test set.CKD outcome
[0071] The primary outcome evaluated in this study was a diagnosis of CKD due to T2D or hypertension as defined by a set of SNOMED-CT codes within three years of the study start date. The positive class consisted of patients who developed CKD within the 3-year timeframe as defined by a set of ICD-10 and SNOMED codes. Patients with T2D who did not have an associated diagnostic code for CKD within the 3-year window were in the negative class.CKD risk model development
[0072] The variables included in the machine learning risk model for prediction of CKD include: patient demographics, vital signs, medications, comorbidities, laboratory tests, family history, and social determinants of health such as rural / urban status and insurance information. Patients were only included if they had at least one laboratory value for HbAlc and serum creatinine within three years of the study start date. The full list of variables included in the model were standard (data not shown). Blood pressure ratio was calculated as systolic blood pressure divided by diastolic blood pressure. Comorbidities that are usually chronic, such as hypertension, were included if they were active or resolved at any time before tO. An ‘active’ condition is a health issue that affects the individual's current functioning and overall health. ICD-10-CM (International Classification of Diseases, Tenth Revision, Clinical Modification) and SNOMED-CT hierarchical parent codes were used to identify conditions. Laboratory values were identified using LOINC codes and valid laboratory ranges for each lab measurement was established by medical review. Valid ranges for vital signs were established by medical review. To measure the burden of multiple comorbidities in a patient, the Charlson Comorbidity Index (ref. 21) was used, a weighted count of common chronic conditions based on ICD-10 (International Classification of Diseases, 10th Revision) codes. The estimated Glomerular Filtration Rate (eGFR) was recalculated from creatinine in serum using the chronic kidney disease epidemiology collaboration (CKD-EPI) 2021 formulal (ref. 22). A patient had a history of low eGFR if they had an eGFR value of < 60 for > 3 months but did not have a diagnosis of CKD or any other type of renal impairment, with the exception of specified types of AKI. Number of outpatient encounters and number of medications in the past year were calculated by summing the total number of outpatient encounters or medication prescriptions within the past year of the study start date. Social determinants of health and social vulnerability index (SVI) scores were calculated based on the CDC / ATSDR Social Vulnerability Index as described previously (ref.23). All statistical analyses were completed using PySpark version 2.4.5.
[0073] An XGBoost machine learning model was trained using the Shapley additive explanations (SHAP) algorithm, which uses cooperative game theory to calculate the marginal contribution of each feature, and examines the feature influence on model prediction (ref. 24).The XGBoost model was generated using the Python package sklearn (version 1.1.1), xgboost (version 1.6.1) and SHAP (version 0.37.0). The models were trained on 90% of the data (over-sampling of the minority class was applied to compensate for class imbalance in the training dataset), with 10% of the data held out for independent performance testing of the final model (hold-out test set). Performance was evaluated on the hold-out test set for Area Under the ROC Curve (AUROC), Fl, precision, and recall scores for the classification results. For bias analysis and demographic proportion analysis of the CKD risk model, the Chi-square test was used from scipystats.CKD risk model performance
[0074] The CKD risk model contained 167 features and had an AUROC of 0.826 on the reserved test set. Using SHAP, the model identified several variables as predictive factors, including high serum creatinine, high urea nitrogen to creatinine ratio, high blood urea nitrogen (BUN), high blood pressure ratio, being of older age, being female, having a high Charlson Comorbidity Index score, and having complications from T2D. For ease of use when recruiting patients, and for explainability, a parsimonious model was also created with only eight features. This model used the most predictive features based on SHAP values as well as clinical significance. Model performance dropped only slightly (AUROC 0.796). Similar models developed for 3-5 year CKD prediction in patients with diabetes had a similar or lower AUROC (refs. 16-19). However, the cohort size and population, outcome definition, and number and type of features used varied widely between studies.Predicting disease transitions using CROP paired with the CKD risk model
[0075] The highest risk patients identified by a 167-feature CKD risk model, defined as the top 500 patients with the highest predicted probability scores and a predicted value of 1, were examined. The predicted probability for the 500 highest risk patients ranged from 0.5324 to 0.9998. Of these 500 patients, 225 transitioned to CKD, resulting in a positive predictive value of 0.45. Alternatively, if 500 patients were randomly selected from the total patient cohort, only 7.6% of patients (on average after sampling 100 times) transitioned to CKD within 3 years. Using the random sampling method, clinical study researchers would need to observe 1,300 patients over 3 years to capture at least 100 cases. Using the CKD risk model, only 219patients would need to be recruited to ensure at least 100 patients transitioned to CKD, resulting in an almost six-fold reduction in recruitment.
[0076] The CKD risk model performance was then incorporated into CROP to estimate the number of CKD transitions for the patient population. Prior to adjusting for model performance, the expected number of transitions in the CROP set was 454, while the true number of transitions in this set is 1025 patients. When incorporating the performance statistics evaluated on the tuning set with a model threshold of 0.5, the expected number of transitions predicted by CROP jumps to 1175. Next, the threshold is tuned until the expected number of transitions in the tuning set matches the actual number of transitions, yielding a threshold of.32 and 1093 predicted transitions on the CROP set. Evidently there is an overestimating of the number of transitions in the CROP set, but it is more accurate than i) deterministic counting of predicted transitions and ii) Poisson model with unadjusted probabilities. As can be seen, this discrepancy does not significantly impact potential cohort recruitment.
[0077] Again, the Poisson model is meant to yield cohort-level statistics and does not tell exactly who will transition. One can, and should, keep the analogy to a series of biased coin flips-with each flip, one only knows the probability that it will land on heads, but can still make claims about how many heads one expects to see in a given number of flips.
[0078] To illustrate the effect threshold tuning has on different metrics, several commonly used metrics were plotted (precision, recall, accuracy and Fl), normalized by their maximum value as well as the difference between the expected and actual cohort transitions under CROP in FIG. 3A. All such metrics were calculated on the tuning set to illustrate the effect on threshold tuning only. Note that for the CKD risk probability model built and tested herein, metrics such as the precision, recall and accuracy were approximately monotonic functions and so reached their extrema on the boundary of the interval [0,1], Note that it is possible for the recall to reach a maximum in the interior of [0,1], since as the threshold approaches 1, the number of true positives goes to zero while the number of false negatives increases. However, the threshold at which this occurs depends on the distribution of transition probabilities and will generically occur at high values of the threshold, meaning that if one’s threshold tuning range is not large enough, then this value will reach its maximum at the boundary. The Fl score, being the harmonic mean of the precision and recall, may reach itsmaximum value in the interior of the interval, but again is sensitive to the specific range of parameter tuning. All of these are in contrast to CROP, which reaches a minimum well into the interior of [0,1], However, the actual quantity one wants to extremize depends on the use case at hand. As illustrated here, if one wants to optimize their model for cohort recruitment, then CROP is the obvious choice, since the value of the threshold at which it reaches the minimum is reasonable and is not as sensitive to the choice of threshold range as the other performance metrics.
[0079] One can also compare how different methods for estimating the number of transitions in the CROP set change as a function of the threshold (FIG. 3B). Using an unadjusted Poisson method, it is possible to underestimate the number of transitions, but has the benefit of being independent of the threshold. Counting the number of predicted positives with the unadjusted and adjusted model outputs gets closer to the actual number of transitions, but fails to capture the probabilistic nature of the question at hand. However, CROP appears to be the best method for accurately predicting the number of transitions in the cohort. It is both probabilistic in nature, and is less sensitive to the threshold than the simple counting procedure.Recruiting patients for a CKD clinical trial using CROP
[0080] Whether using a previously-derived risk model or creating a new one, the approach to use of CROP to recruit patients for a clinical trial is the same. In the following scenario, half of the patients in the ML test set were used to tune CROP (Retrospective Tuning Cohort) and the other half to represent a new population of patients eligible for clinical trial recruitment (Contemporary Population for Recruitment). Once the probabilities for each patient were obtained and the optimal threshold determined using CROP, the probabilities of patients in the Contemporary Population for Recruitment were arranged in decreasing order followed by adding the probabilities together until their sum reached the desired number of transitions. Assuming a desire to see at least 100 transitions, the adjusted probabilities using CROP generates cohort of 234(i.e., the CROP Risk-Enriched Cohort) that included an approximately equal number of transitions and non-transitions to perform adequate statistical analysis. Among the 234 high-risk patients, 119 (51%) patients actually underwent a transition. Using a deterministic approach (counting predicted positives) would have incorrectly indicated that all of these patients would have undergone a transition.Conducting an emulated CKD study and assessing for bias
[0081] Emulated recruiting and three year outcomes were performed on the separately held out and previously unobserved study cohort (Table 1). The CROP cohort was also assessed for bias in in order to decide whether to employ a stratified recruitment strategy. As shown in Table 1, it was observed that the population predicted by CROP (CROP Risk-Enriched Cohort) tended to contain older patients, while the proportions of the other demographic categories were similar to the overall Retrospective Tuning Cohort. Using a Chi-square goodness of fit test to compare both the Contemporary Cohort for Recruiting and the CROP Risk-Enriched Cohort to the Retrospective Tuning Cohort, it was found that only the age demographic differed significantly (p < 0.05) for the CROP Risk-Enriched Cohort. In this case, it was possible to specifically recruit a certain proportion of younger (18-45) patients, or recruit a more even proportion of older patients (aged 46-75+), depending on the study needs.Table 1. Demographic proportion analysis comparing cohort selection approaches and outcomesRetrospective Contemporary CROP Risk- Tuning Cohort (%) Cohort for Enriched Cohort Recruiting (%) (%) Cohort size 12747 12748 234Patients that transitioned to 1087 (8.5%) 1025 (8.0%) 119 (51.0%) CKD within 3 yearsDiscussion
[0082] As demonstrated, the CROP method can be paired with risk scores from any risk model to identify higher- or lower-risk patients for clinical study recruitment, allowing for a reduction in the cohort size needed to observe a sufficient number of transitions to the outcome of interest. CROP can also be used across a variety of clinical studies and medical conditions. For example, CROP was tested with a new CKD risk model to determine the number of T2D patients that would transition to CKD within 3 years. The CKD risk model performance had an AUROC of 0.826. CROP was able to achieve a 5.3-fold reduction in cohort size, while maintaining a fixed number of transitions compared to random sampling. When deciding whether to proceed, users can weigh the benefit of the reduced cohort size against the reducedgeneralizability of results for younger patients, and then decide whether to use substratification. For substratification, CROP enrichment could be conducted separately for an older subcohort and a younger cohort. In addition, there would be also be an option to find or develop different risk models for older and younger patients (ref. 15).
[0083] It is important to note that only variables known at the onset of the study can be assessed for generalizability before the study begins. However, enriching a cohort for higher or lower risk may also reduce generalizability in variables that cannot be assessed until the study is conducted, if at all. For example, genomics were not assessed, so if proceeding to a prospective study with genomic data, representation could not be assessed until the study was conducted. Furthermore, enrichment for a specific time frame, for example the three-year model for CKD, might deplete a cohort of patients with a risk for slower-onset of disease. One way to offset this risk of reduced generalizability would be to recruit a cohort of patients that have been randomly selected, in addition to higher-risk patients. For example if 200 people were selected randomly at first, and then an enriched cohort was selected from the population remaining, the cohort would include at least 419 subjects, but would still be smaller than complete random selection. The randomly selected subsample could be used to detect inadvertent limitations to generalizability after the prospective study is complete. There also is a possibility of overenrichment (as shown in the CKD example). If a minimum number of controls is also needed, this could be achieved by using CROP to also select those at intermediate or low risk.References1. Jones SR, Carley S, Harrison M. An introduction to power and sample size estimation. EmergMed J. 2003;20(5):453-458.2. Barria RM. Cohort Studies in Health Sciences. BoD - Books on Demand; 2018.3. Doll R. Cohort studies: history of the method. I. Prospective cohort studies. Soz Praventivmed. 2001;46(2):75-86.4. Bycroft C, Freeman C, Petkova D, et al. The UK Biobank resource with deep phenotyping and genomic data. Nature. 2018;562(7726):203-209.5. The All of Us Research Program Investigators. The “All of Us” Research Program. N Engl J Med. 2019;381(7):668-676.6. Keesler Air Force Base Medical Center. PREdicting Determinants of Atrial Fibrillation or Flutter for Therapy Elucidation in Patients at Risk for Thromboembolic Events (PREDATE AF). ClinicalTrials.gov. Accessed June 26, 2024. https: / / clinicaltrials.gov / study / NCT01851902 7. Walsh T, Duff L, Riviere ME, et al. Outreach, Screening, and Randomization of APOE s4 Carriers into an Alzheimer’s Prevention Trial: A global Perspective from the API Generation Program. J Prev Alzheimers Dis. 2023;10(3):453-463.8. Fahed AC, Philippakis AA, Khera AV. The potential of polygenic scores to improve cost and efficiency of clinical trials. Nat Commun. 2022;13(1):2922.9. Birkenbihl C, de Jong J, Yalchyk I, Fröhlich H, the Alzheimer’s Disease Neuroimaging Initiative. Deep learning-based patient stratification for prognostic enrichment of clinical dementia trials. bioRxiv. Published online November 27, 2023.doi: 10.1101 / 2023.11.25.2329901510. Viele K, Girard TD. Risk, Results, and Costs: Optimizing Clinical Trial Efficiency through Prognostic Enrichment. Am J Respir Crit Care Med. 2021;203(6):671-672.11. Andaur Navarro CL, Damen JAA, Takada T, et al. Risk of bias in studies on prediction models developed using supervised machine learning techniques: systematic review. BMJ. 2021;375:n2281.12. Suri JS, Bhagawati M, Paul S, et al. Understanding the bias in machine learning systems for cardiovascular disease risk assessment: The first of its kind review. Comput Biol Med. 2022; 142: 105204.13. Alelyani S. Detection and Evaluation of Machine Learning Bias. NATO Adv Sci Inst SerE Appl Sci. 2021;11(14):6271.14. McCradden MD, Joshi S, Mazwi M, Anderson JA. Ethical limitations of algorithmic fairness solutions in health care machine learning. Lancet Digit Health. 2020;2(5):e221-e223.15. Molani S, Hernandez PV, Roper RT, et al. Risk factors for severe COVID- 19 differ by age for hospitalized adults. Sci Rep. 2022;12(l):6568.16. Song X, Waitman LR, Yu AS, Robbins DC, Hu Y, Liu M. Longitudinal Risk Prediction of Chronic Kidney Disease in Diabetic Patients Using a Temporal-Enhanced Gradient Boosting Machine: Retrospective Cohort Study. JMIRMed Inform. 2020;8(l):el5510.17. Allen A, Iqbal Z, Green-Saxena A, et al. Prediction of diabetic kidney disease with machine learning algorithms, upon the initial diagnosis of type 2 diabetes mellitus. BMJ Open Diabetes Res Care. 2022;10(1). doi:10.1136 / bmjdrc-2021-00256018. Ravizza S, Huschto T, Adamov A, et al. Predicting the early risk of chronic kidney disease in patients with diabetes using real-world data. Nat Med. 2019;25( 1 ): 57-59.19. Dong Z, Wang Q, Ke Y, et al. Prediction of 3-year risk of diabetic kidney disease using machine learning based on electronic medical records. J Transl Med. 2022;20(1):143.20. Food and Drug Administration (FDA ). Enrichment Strategies for Clinical Trials to Support Determination of Effectiveness of Human Drugs and Biological Products: Guidance for Industry. Docket FDA-2012-D-1145. Published online 2019. https: / / www.fda.gov / media / 121320 / download21. Charlson ME, Pompei P, Ales KL, MacKenzie CR. A new method of classifying prognostic comorbidity in longitudinal studies: development and validation. J Chronic Dis. 1987;40(5):373-383.22. Inker LA, Eneanya ND, Coresh J, et al. New Creatinine- and Cystatin C-Based Equations to Estimate GFR without Race. N Engl J Med. 2021;385(19): 1737- 1749.23. Piekos SN, Hwang YM, Roper RT, et al. Effect of COVID-19 vaccination and booster on maternal-fetal outcomes: a retrospective cohort study. The Lancet Digital Health.2023;5(9):e594-e606.24. Rodriguez-Perez R, Bajorath J. Interpretation of Compound Activity Predictions from Complex Machine Learning Models Using Local Approximations and Shapley Values. J Med Chem. 2020;63(16):8761-8777.25. Collins GS, Dhiman P, Andaur Navarro CL, et al. Protocol for development of a reporting guideline (TRIPOD-AI) and risk of bias tool (PROBAST-A1) for diagnostic and prognostic prediction model studies based on artificial intelligence. BMJ Open.2021; 11(7):e048008.26. Chin MH, Afsar-Manesh N, Bierman AS, et al. Guiding Principles to Address the Impact of Algorithm Bias on Racial and Ethnic Disparities in Health and Health Care. JAMA Netw Open. 2023;6(12):e2345050.27. Johnson KB, Wei WQ, Weeraratne D, et al. Precision Medicine, Al, and the Future of Personalized Health Care. Clin Transl Sci. 2021;14(1):86-93.28. McCradden MD, Anderson JA, A Stephenson E, et al. A Research Ethics Framework for the Clinical Translation of Healthcare Machine Learning. Am J Bioeth. 2022;22(5):8-22.
Claims
CLAIMS1. A computer-implemented method of producing a risk probability model tuned for riskbased sampling, comprising:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a retrospective population of subjects having known real outcomes for the event over a given period of time;(b) generating a confusion matrix at a given classification threshold for the risk probabilities;(c) generating adjustment factors each individually normalizing an outcome of the confusion matrix as a performance metric of the risk probability model at the given classification threshold;(d) generating adjusted risk probabilities by matrix multiplication of the adjustment factors and the accessed risk probabilities so as to bound the accessed risk probabilities by the performance metric;(e) determining an expected number of events in the retrospective population as an aggregate or rate of the adjusted risk probabilities across the retrospective population in a probability distribution model;(f) determining a difference between the expected number of events in the retrospective population and the real number of events in the retrospective population;(g) generating an optimized adjustment factors and classification threshold pairing for the risk probability model by repeating steps (b) - (f) with one or more different classification thresholds and selecting an adjustment factors and classification threshold pairing that generates a reasonable minimum between the expected and real number of events in the retrospective population; and(h) configuring the risk probability model with the adjustment factors for model persistence by executing a model update protocol stored in a non-transitory computer-readable medium, wherein the update protocol is designed to improve model generalization and reduce variance in subsequent risk predictions.
2. The method of claim 1, wherein the risk probability model is a binary classification risk probability model, and wherein the probability distribution selected for the retrospective population is selected from normal, binomial, Poisson, Poisson binomial, negative binomial, and negative Poisson binomial.
3. The method of claim 2, wherein the probability distribution selected for the retrospective population is Poisson binomial.
4. The method of claim 1, wherein the risk probability model is a multinomial classification risk probability model, and wherein the probability distribution selected for the retrospective population is selected from multinomial, Poisson multinomial, negative multinomial, and negative Poisson multinomial.
5. The method of claim 4, wherein the probability distribution selected for the retrospective population is Poisson multinomial.
6. The method of claim 1, further comprising:generating risk probabilities for a target population of subjects using the tuned risk probability model;generating adjusted risk probabilities by matrix multiplication of the optimized adjustment factors and the risk probabilities of the target population; anddetermining an expected number of events in the target population based on application of the adjusted risk probabilities and the probability distribution.
7. The method of claim 6, further comprising:determining a sample size needed for the target population containing a desired number of subjects expected to exhibit the event at some point over the given period of time, the desired number of subjects approximately equal to or less than the expected number of events determined for the target population.
8. The method of claim 7, wherein determining the sample size comprises:generating: (i) a sum of the adjusted risk probabilities in descending order until the sum of the adjusted risk probabilities adds up to the desired number of subjects, and (ii) a count corresponding to the number of the adjusted risk probabilities used to obtain the sum, wherein the count is the sample size.
9. A method of risk-based sampling, the method comprising:(a) accessing risk probabilities generated by a risk probability model, wherein each risk probability represents a likelihood of an event occurring to a given subject in a target population of subjects over a given period of time;(b) applying adjustment factors to the accessed risk probabilities, wherein each adjustment factor is a proportion of an outcome of a confusion matrix generated for a given classification threshold applied to risk probabilities generated by the risk probability model for a retrospective population of subjects having known actual outcomes for the event over the given period of time, the proportion reflecting how well the risk probability model predicts an outcome at the given classification threshold, the outcome selected from true positives, true negatives, false positives, and false negatives;(c) generating adjusted risk probabilities based on application of the adjustment factors to the accessed risk probabilities by matrix multiplication so as to bound the risk probabilities by performance of the risk probability model for all real events given all predicted events of the retrospective population; and(d) determining an expected number of events in the target population based on application of a probability distribution to the adjusted risk probabilities, the probability distribution selected for determining an expected number of events in the retrospective population.
10. The method of claim 9, further comprising:(e) determining a sample size of the target population containing a desired number of subjects expected to exhibit the event at some point over the given period of time, the desired number of subjects approximately equal to or less than the expected number of events determined for the target population.
11. The method of claim 10, wherein determining the sample size comprises: generating (i) a sum of the adjusted risk probabilities in descending order until the sum adds up to the desired number of subjects, and (ii) a count corresponding to a number of the adjusted risk probabilities used to obtain the sum, wherein the count is the sample size.
12. The method of claim 10, wherein the desired number of subjects is an enrichment of the sample selected from decreased variability enrichment, prognostic enrichment, and predictive enrichment.
13. The method of claim 9, wherein the adjustment factors and the given classification threshold is an optimized pairing for the risk probability model and the retrospective population such that the expected number of events in the retrospective population is approximately equal to an actual number of real events in the retrospective population over the given period of time.
14. The method of claim 13, wherein the target population is the retrospective population, and wherein the optimized pairing is produced by:(f) generating a performance metric for the risk probability model determining an absolute difference between the expected and actual number of events occurring in the retrospective population; and(g) generating an optimized adjustment factors and classification threshold pairing for the risk probability model by repeating steps (b) - (d) and (f) with one or more different adjustment factors and classification threshold pairings and selecting a paring that generates a reasonable minimum for the performance metric.
15. The method of claim 9, wherein the target population is selected from a prospective population, a contemporary population for recruitment, or a combination thereof.
16. The method of claim 9, wherein the retrospective population and the target population have substantially the same inclusion and exclusion criteria.
17. The method of claim 9, wherein the adjustment factors are derived from a column normalized confusion matrix for the given classification threshold.
18. The method of claim 9, wherein the event is a plurality of events.
19. The method of any one of claim 9-13, further comprising:outputting one or more of the classification threshold, the confusion matrix, the adjustment factors, the optimized pairing, the adjusted risk probabilities, the expected number of events, the desired number of subjects, the sample size, or a combination thereof, for use in a decision-making process.
20. The method of claim 19, wherein the subjects are humans, the event is a medical outcome, and the decision-making process is selection of patients for the medical outcome.
21. The method of claim 20, wherein the medical outcome is new onset chronic kidney disease (CKD).
22. The method of claim 21, wherein the risk probability model is for identifying type 2 diabetes patients at high risk for CKD transition, the retrospective population and the target populations are individuals with type 2 diabetes, and the event is a transition to CKD.
23. A system comprising:one or more computers; andone or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the method of any of claims 1-22.
24. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform the method of any of claims 1-22.
25. A system comprising:one or more computers; andone or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform operations for part or all of one or more methods disclosed herein.
26. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions configured to cause one or more data processors to perform part or all of one or more methods disclosed herein.
27. Any and all methods, processes, devices, systems, devices, kits, products, materials, compositions and / or uses shown and / or described expressly or by implication in the information provided herewith, including but not limited to features that may be apparent and / or understood by those of skill in the art.