Predictive Model Training via Dataset Distribution Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Incomplete patient datasets, particularly 'open datasets,' hinder the generation of accurate health-based predictions due to missing data, which can impact eligibility determinations for therapies like CAR-T therapy.
Innovation Solution
The method involves linking closed datasets with comprehensive patient data to open datasets, modifying the closed datasets to simulate missing data points, and generating supersets to train predictive models that can analyze incomplete datasets for patient eligibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If open datasets are used for training predictive models, then data availability and model deployment scope increase, but prediction accuracy deteriorates due to missing data
Solution Approach 1:
The patent applies preliminary action by pre-processing open datasets through multiple imputation methods before training. The system performs data cleaning, handles missing values through imputation, and standardizes features before the predictive model training process, ensuring that the model receives high-quality input data that maximizes prediction accuracy while utilizing available open dataset resources.
Solution Approach 2:
The patent introduces an intermediary data processing layer that acts as a mediator between open datasets and predictive models. This intermediary layer includes components for data imputation, feature engineering, and quality assessment that transform raw open dataset data into a format suitable for accurate prediction, thereby bridging the gap between data availability and prediction accuracy.
2Measurement precision
If data imputation methods are applied to handle missing values, then prediction accuracy improves, but computational complexity and processing time increase
Solution Approach 1:
The patent applies parameter changes by implementing multiple imputation methods that vary the approach to handling missing data. The system can switch between different imputation strategies (mean, median, mode, k-nearest neighbors, regression-based) depending on the data characteristics and missingness patterns, optimizing the balance between accuracy improvement and computational overhead for different scenarios.
Solution Approach 2:
The patent applies partial action by selectively applying imputation methods only to specific features or data points where missingness occurs, rather than processing the entire dataset uniformly. The system identifies missing values and applies appropriate imputation techniques only where needed, reducing unnecessary computational complexity while maintaining prediction accuracy.
3Measurement precision
If comprehensive patient history data is required for accurate predictions, then prediction accuracy improves, but data privacy concerns and accessibility issues worsen
Solution Approach 1:
The patent applies segmentation by dividing patient history data into distinct feature categories (demographics, clinical history, treatment data, outcomes). The system processes and analyzes these segmented features independently, allowing accurate predictions to be made from distributed, access-controlled data sources without requiring centralized access to complete patient histories, thereby addressing privacy and accessibility constraints.
Data Source
AI summary
Disclosed herein are methods for training and deploying a predictive model for generating a prediction, e.g., patient eligibility for a CAR-T therapy. Datasets, such as open healthcare claims datasets, may be missing data. Missing data may hamper the ability to generate sufficient information needed for training a predictive model. Methods include leveraging comprehensive datasets, such as closed claims datasets, to create training examples for input into a machine learning algorithm. In various embodiments, the comprehensive dataset is modified to simulate the data missingness in the target dataset; then, the modified dataset is paired with the ground truth label derived from the comprehensive dataset to create training examples. In various embodiments, a comprehensive dataset is paired with a target dataset to create training examples. After training a predictive model on such examples, the model can be deployed to make predictions in the target dataset even in light of missing data.


