Predictive ML Model Using Knowledge Graph for Cohort Risk Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing predictive data analysis solutions face inefficiencies and reliability issues in performing accurate risk predictions, particularly in complex data domains like healthcare, where they often require extensive computational resources and data storage, and struggle to improve training speed without sacrificing predictive accuracy.
Innovation Solution
The method involves using a predictive machine learning model trained with outer cohort and inner cohort definition data, leveraging a knowledge graph to determine correlated features and generate risk scores, allowing for targeted cohort identification and reduced computational and storage requirements by initializing the model with minimal input criteria.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional predictive data analysis solutions are used, then comprehensive data analysis can be performed, but computational resources and data storage requirements become excessive
Solution Approach 1:
The patent segments the data analysis process into distinct phases: cohort identification, feature extraction, and predictive modeling. By dividing the comprehensive dataset into relevant cohorts and extracting only necessary features, the system reduces computational resource consumption and data storage requirements while maintaining predictive accuracy.
Solution Approach 2:
The patent extracts only the essential features and data subsets required for specific predictive tasks rather than processing entire datasets. This extraction approach identifies and isolates relevant information, significantly reducing the quantity of data and computational resources needed while preserving the reliability of predictions.
2Reliability
If training data and model complexity are increased to improve predictive accuracy, then prediction reliability improves, but training speed decreases
Solution Approach 1:
The patent performs preliminary actions by pre-identifying cohorts and extracting features before model training. This preparation work organizes data in advance, allowing the training process to proceed more efficiently with pre-processed, relevant data subsets, thereby improving training speed without sacrificing predictive accuracy.
Solution Approach 2:
The patent applies partial action by focusing training efforts on specific, relevant cohorts and features rather than processing all available data. This selective approach achieves sufficient predictive accuracy for targeted tasks while significantly reducing training time and computational overhead.
3Adaptability or versatility
If comprehensive feature sets are used for all entities, then predictive modeling can be performed on entire datasets, but computational operations and storage requirements increase
Solution Approach 1:
The patent applies local quality by tailoring feature sets to specific cohorts and predictive tasks rather than using uniform comprehensive features for all entities. Each cohort receives customized features relevant to its characteristics, reducing overall computational complexity and storage requirements while maintaining adaptability for different prediction scenarios.
Data Source
AI summary
The present disclosure provides methods, apparatus, systems, computing devices, and/or the like for performing risk prediction by receiving outer cohort definition data and inner cohort definition data, the outer cohort definition data representative of a target data domain with respect to a dataset, and the inner cohort definition data representative of a prediction feature with respect to the target data domain, determining one or more inner cohort features based at least in part on a knowledge graph data object using the inner cohort definition data, the knowledge graph data object including co-occurrence information of features from the dataset, and generating, using a predictive machine learning model, for each of one or more outer cohort entities associated with features in an outer cohort data subset, a risk score representative of a propensity of the outer cohort entity being an inner cohort entity associated with features in an inner cohort data subset.


