Machine Learning Risk Stratification for Endometrial Cancer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current molecular classification systems for endometrial cancer, primarily developed from European populations, may not accurately reflect the genetic and clinical profiles of diverse populations, such as Black or African American women, leading to imperfect risk stratification and treatment decisions.
Innovation Solution
A machine learning-based method that utilizes patient health data, including clinical, molecular, and demographic factors, to generate classified feature data indicative of endometrial cancer risk stratification, specifically the NU-CATS score, which has demonstrated improved accuracy in predicting disease progression and overall survival across diverse populations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If molecular classification systems (ProMisE, Leiden/TransPORTEC) are used for risk stratification, then prognostic accuracy is improved, but applicability to diverse populations deteriorates
Solution Approach 1:
The patent segments the population by demographic characteristics (race, ethnicity, age) and applies population-specific machine learning models. Instead of using a single universal molecular classification system, the system divides the patient population into distinct segments and trains separate models on each segment's data, allowing each model to capture population-specific patterns while maintaining overall system accuracy across diverse groups.
Solution Approach 2:
The patent implements local quality by tailoring the risk stratification approach to specific population groups. Different machine learning models are trained on data from different demographic groups (e.g., Black or African American patients, Caucasian patients, age-specific groups), allowing each local model to optimize for its specific population's characteristics, molecular subtype distribution, and clinical outcomes rather than applying a one-size-fits-all European-derived classification.
2Ease of operation
If TP53 is used as a surrogate for CN-high molecular subgroup, then clinical feasibility is improved, but measurement accuracy deteriorates
Solution Approach 1:
The patent uses machine learning models as intermediaries that integrate multiple data sources including TP53 status, other molecular markers, clinicopathologic variables, and demographic factors. Rather than relying solely on TP53 as a direct surrogate, the ML models process TP53 information along with numerous other features to generate a more accurate composite risk assessment, effectively using the ML system as a mediator that transforms simple surrogate markers into precise risk stratification.
Solution Approach 2:
The patent creates a composite risk assessment system that combines multiple data types (molecular data including TP53, clinicopathologic variables, demographic information) into an integrated machine learning model. This composite approach allows the system to leverage the ease of TP53 measurement while compensating for its limitations by incorporating additional data layers that collectively improve classification accuracy beyond what TP53 alone can provide.
3Ease of manufacture
If European population data is used to develop classification systems, then model development is simplified, but reliability for other populations deteriorates
Solution Approach 1:
The patent segments the training data by demographic groups and develops separate machine learning models for different populations. Rather than attempting to build a single model from European data and hoping it generalizes, the system divides the data development process into population-specific segments, training distinct models on Black or African American patient data, Caucasian patient data, and other demographic groups, thereby ensuring each model is reliable for its target population.
Solution Approach 2:
The patent applies parameter changes by adapting the machine learning models to population-specific characteristics. The system modifies model parameters, feature importance weights, and data preprocessing approaches based on the demographic group being modeled, allowing the same underlying ML framework to be tuned for reliability across different populations rather than using fixed parameters derived solely from European cohorts.
Data Source
AI summary
Endometrial cancer is classified and/or risk stratified using a suitably trained machine learning model. Risk classification of endometrial cancer is provided using a machine learning-based analysis of patient health data, such as clinicopathologic data, molecular data, and the like. Risk assessment is optimized for endometrial cancer, including risk of nodal involvement, distant metastasis, disease progression, and overall survival.


