Population-Specific Machine Learning PRS Models for Ancestry Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing PRS models face limitations due to the size of the training cohort and are not effective across different ancestral populations, leading to inconsistent performance.

Innovation Solution

An end-to-end PRS machine automates the training of population-specific models using machine learning techniques, incorporating user-selected parameters, SNP filtering criteria, and performance metrics to generate personalized PRS scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If existing PRS models are used across different ancestral populations, then model deployment is simplified, but performance consistency deteriorates

Engineering Contradiction:
Improvemodel deployment simplicityVSAvoidperformance consistency
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent segments the training process by creating separate population-specific models for different ancestral groups (European, African, Asian, Hispanic/Latino, Middle Eastern, and admixed populations). Each model is trained independently on population-specific genetic data, ensuring that the genetic architecture and allele frequencies specific to each population are captured accurately. This segmentation resolves the contradiction by maintaining performance consistency across diverse populations while deploying multiple specialized models rather than a single generic model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by tailoring each PRS model to the specific genetic characteristics of its target population. Population-specific reference panels and genetic data are used to train each model, ensuring that the predictions are optimized for the local genetic architecture of each ancestral group. This approach maintains high reliability for each population while the overall system handles multiple populations through specialized sub-models.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If larger training cohorts are used, then model accuracy improves, but data processing complexity increases

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the large genetic dataset into population-specific subsets based on ancestral origin. Each population-specific model is trained on its corresponding genetic data, which reduces the complexity of processing entire diverse datasets for each model. This segmentation allows parallel processing of smaller, focused datasets while maintaining the benefits of large sample sizes within each population group.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of the training process by using population-specific reference panels and adjusting the genetic data processing parameters according to each population's characteristics. This includes using population-appropriate imputation panels, quality control thresholds, and genetic variant filtering criteria, which simplifies data processing for each population while maintaining high model accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250266129A1Machine Learning Platform for Polygenic Models
Publication Date: 2025.08.21 23ANDME GENOMICS LLC
  • US20250266129A1 patent drawing
  • US20250266129A1 patent drawing
  • US20250266129A1 patent drawing

AI summary

The disclosed embodiments concern methods, apparatus, systems, and computer program products for developing polygenic risk score (PRS) models. In some implementations, a fully automated process is provided that allows for a PRS model to be defined by an initial set of parameters. In some implementations the PRS models are trained to provide a PRS for particular populations.