Proxy Endpoint Models Bridge RCT Scores and Real-World Patient Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Real-world data (RWD) lacks disease severity scores necessary for drug discovery studies, limiting its use in understanding patient-treatment interactions, while Randomized Control Trial (RCT) data is highly controlled but lacks large sample sizes.
Innovation Solution
Augmenting patient data sets with proxy endpoint data using data processing devices, developing proxy endpoint models from RCT data to predict disease severity scores based on candidate features, and integrating these models into RWD to enhance treatment analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If Real World Data (RWD) is used for drug discovery studies, then large sample sizes can be analyzed, but disease severity scores (endpoints) are not available for these individuals
Solution Approach 1:
The patent applies preliminary action by pre-training proxy endpoint models on RCT data that includes disease severity scores. These pre-trained models are then deployed to RWD to generate predicted disease severity scores, allowing large-scale analysis of real-world patients without requiring actual endpoint measurements from each patient.
Solution Approach 2:
The patent creates a copy of the disease severity score information by using proxy endpoint models to predict and generate synthetic endpoint data for RWD patients. This copied information mirrors the structure and utility of actual disease severity scores, enabling analysis that would otherwise require direct measurements.
2Reliability
If RCT data is used to ensure high reliability of causal inference, then data quality is improved, but sample size is limited
Solution Approach 1:
The patent merges RCT data with RWD by training proxy endpoint models on the high-quality RCT data and then applying these models to the larger RWD dataset. This combination preserves the reliability benefits of RCT data while extending analysis to the larger sample sizes available in real-world settings.
Solution Approach 2:
The patent introduces proxy endpoint models as an intermediary between RCT data and RWD. These models learn the relationship between patient features and disease severity scores from RCT data, then use this learned relationship to generate predicted scores for RWD patients, bridging the gap between controlled trial data and real-world data.
3Loss of information
If proxy endpoint models are developed to predict disease severity scores from RCT data, then RWD can be augmented with endpoint data, but additional processing steps are required
Solution Approach 1:
The patent reduces processing complexity during deployment by performing the complex model training in advance on RCT data. The pre-trained proxy endpoint models can then be applied to RWD with minimal additional processing, requiring only feature extraction and model prediction rather than full model training.
Data Source
Figure 1
Figure 2~3
Figure 4~5
AI summary
A method of augmenting a patient data set to include proxy endpoint data using one or more data processing devices. The method comprises selecting one or more candidate features which are present for at least some subjects in the patient data set and for at least some subjects in a randomised control trial, RCT, data set, comprising processing an input set of features of the RCT data set through one or more selection stages. One or more proxy endpoint models are developed, wherein each of the one or more proxy endpoint models is configured to predict a value for an RCT endpoint based on values for a respective set of one or more model features. For each proxy endpoint model, the respective set of one or more model features includes one or more of the candidate features. Developing a proxy endpoint model comprises fitting the proxy endpoint model to a respective training data set obtained from the RCT data set. The patient data set is then augmented using at least one of the one or more proxy endpoint models to include proxy endpoint data in the patient data set. For each of the at least one proxy endpoint models, augmenting the patient data set comprises deploying the proxy endpoint model to predict, for each of a plurality of subjects represented in the patient data set, a value for the corresponding proxy endpoint based on values of the corresponding set of one more model features for the subject.