Machine Learning Risk Stratification for Endometrial Cancer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current molecular classification systems for endometrial cancer, primarily developed from European populations, may not accurately reflect the genetic and clinical profiles of diverse populations, such as Black or African American women, leading to imperfect risk stratification and treatment decisions.

Innovation Solution

A machine learning-based method that utilizes patient health data, including clinical, molecular, and demographic factors, to generate classified feature data indicative of endometrial cancer risk stratification, specifically the NU-CATS score, which has demonstrated improved accuracy in predicting disease progression and overall survival across diverse populations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If molecular classification systems (ProMisE, Leiden/TransPORTEC) are used for risk stratification, then prognostic accuracy is improved, but applicability to diverse populations deteriorates

Engineering Contradiction:
Improveprognostic accuracyVSAvoidapplicability to diverse populations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments the population by demographic characteristics (race, ethnicity, age) and applies population-specific machine learning models. Instead of using a single universal molecular classification system, the system divides the patient population into distinct segments and trains separate models on each segment's data, allowing each model to capture population-specific patterns while maintaining overall system accuracy across diverse groups.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by tailoring the risk stratification approach to specific population groups. Different machine learning models are trained on data from different demographic groups (e.g., Black or African American patients, Caucasian patients, age-specific groups), allowing each local model to optimize for its specific population's characteristics, molecular subtype distribution, and clinical outcomes rather than applying a one-size-fits-all European-derived classification.

Inventive Principle:
Principle #3Local quality

2Ease of operation

If TP53 is used as a surrogate for CN-high molecular subgroup, then clinical feasibility is improved, but measurement accuracy deteriorates

Engineering Contradiction:
Improveclinical feasibilityVSAvoidmolecular subgroup classification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent uses machine learning models as intermediaries that integrate multiple data sources including TP53 status, other molecular markers, clinicopathologic variables, and demographic factors. Rather than relying solely on TP53 as a direct surrogate, the ML models process TP53 information along with numerous other features to generate a more accurate composite risk assessment, effectively using the ML system as a mediator that transforms simple surrogate markers into precise risk stratification.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a composite risk assessment system that combines multiple data types (molecular data including TP53, clinicopathologic variables, demographic information) into an integrated machine learning model. This composite approach allows the system to leverage the ease of TP53 measurement while compensating for its limitations by incorporating additional data layers that collectively improve classification accuracy beyond what TP53 alone can provide.

Inventive Principle:
Principle #40Composite materials

3Ease of manufacture

If European population data is used to develop classification systems, then model development is simplified, but reliability for other populations deteriorates

Engineering Contradiction:
Improvemodel development simplicityVSAvoidrisk stratification reliability in diverse populations
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent segments the training data by demographic groups and develops separate machine learning models for different populations. Rather than attempting to build a single model from European data and hoping it generalizes, the system divides the data development process into population-specific segments, training distinct models on Black or African American patient data, Caucasian patient data, and other demographic groups, thereby ensuring each model is reliable for its target population.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies parameter changes by adapting the machine learning models to population-specific characteristics. The system modifies model parameters, feature importance weights, and data preprocessing approaches based on the demographic group being modeled, allowing the same underlying ML framework to be tuned for reliability across different populations rather than using fixed parameters derived solely from European cohorts.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250029729A1Machine learning-based risk-classification of endomerial cancer
Publication Date: 2025.01.23 NORTHWESTERN UNIV
  • US20250029729A1 patent drawing
  • US20250029729A1 patent drawing
  • US20250029729A1 patent drawing

AI summary

Endometrial cancer is classified and/or risk stratified using a suitably trained machine learning model. Risk classification of endometrial cancer is provided using a machine learning-based analysis of patient health data, such as clinicopathologic data, molecular data, and the like. Risk assessment is optimized for endometrial cancer, including risk of nodal involvement, distant metastasis, disease progression, and overall survival.