Predictive ML Model Using Knowledge Graph for Cohort Risk Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing predictive data analysis solutions face inefficiencies and reliability issues in performing accurate risk predictions, particularly in complex data domains like healthcare, where they often require extensive computational resources and data storage, and struggle to improve training speed without sacrificing predictive accuracy.

Innovation Solution

The method involves using a predictive machine learning model trained with outer cohort and inner cohort definition data, leveraging a knowledge graph to determine correlated features and generate risk scores, allowing for targeted cohort identification and reduced computational and storage requirements by initializing the model with minimal input criteria.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional predictive data analysis solutions are used, then comprehensive data analysis can be performed, but computational resources and data storage requirements become excessive

Engineering Contradiction:
Improvepredictive accuracyVSAvoidcomputational resources and data storage
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the data analysis process into distinct phases: cohort identification, feature extraction, and predictive modeling. By dividing the comprehensive dataset into relevant cohorts and extracting only necessary features, the system reduces computational resource consumption and data storage requirements while maintaining predictive accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the essential features and data subsets required for specific predictive tasks rather than processing entire datasets. This extraction approach identifies and isolates relevant information, significantly reducing the quantity of data and computational resources needed while preserving the reliability of predictions.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If training data and model complexity are increased to improve predictive accuracy, then prediction reliability improves, but training speed decreases

Engineering Contradiction:
Improvepredictive accuracyVSAvoidtraining speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary actions by pre-identifying cohorts and extracting features before model training. This preparation work organizes data in advance, allowing the training process to proceed more efficiently with pre-processed, relevant data subsets, thereby improving training speed without sacrificing predictive accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by focusing training efforts on specific, relevant cohorts and features rather than processing all available data. This selective approach achieves sufficient predictive accuracy for targeted tasks while significantly reducing training time and computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If comprehensive feature sets are used for all entities, then predictive modeling can be performed on entire datasets, but computational operations and storage requirements increase

Engineering Contradiction:
Improvemodeling coverageVSAvoidcomputational operations and storage
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by tailoring feature sets to specific cohorts and predictive tasks rather than using uniform comprehensive features for all entities. Each cohort receives customized features relevant to its characteristics, reducing overall computational complexity and storage requirements while maintaining adaptability for different prediction scenarios.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240047070A1Machine learning techniques for generating cohorts and predictive modeling based thereof
Publication Date: 2024.02.08 OPTUM INC
  • US20240047070A1 patent drawing
  • US20240047070A1 patent drawing
  • US20240047070A1 patent drawing

AI summary

The present disclosure provides methods, apparatus, systems, computing devices, and/or the like for performing risk prediction by receiving outer cohort definition data and inner cohort definition data, the outer cohort definition data representative of a target data domain with respect to a dataset, and the inner cohort definition data representative of a prediction feature with respect to the target data domain, determining one or more inner cohort features based at least in part on a knowledge graph data object using the inner cohort definition data, the knowledge graph data object including co-occurrence information of features from the dataset, and generating, using a predictive machine learning model, for each of one or more outer cohort entities associated with features in an outer cohort data subset, a risk score representative of a propensity of the outer cohort entity being an inner cohort entity associated with features in an inner cohort data subset.