Population-Specific Machine Learning PRS Models for Ancestry Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing PRS models face limitations due to the size of the training cohort and are not effective across different ancestral populations, leading to inconsistent performance.
Innovation Solution
An end-to-end PRS machine automates the training of population-specific models using machine learning techniques, incorporating user-selected parameters, SNP filtering criteria, and performance metrics to generate personalized PRS scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If existing PRS models are used across different ancestral populations, then model deployment is simplified, but performance consistency deteriorates
Solution Approach 1:
The patent segments the training process by creating separate population-specific models for different ancestral groups (European, African, Asian, Hispanic/Latino, Middle Eastern, and admixed populations). Each model is trained independently on population-specific genetic data, ensuring that the genetic architecture and allele frequencies specific to each population are captured accurately. This segmentation resolves the contradiction by maintaining performance consistency across diverse populations while deploying multiple specialized models rather than a single generic model.
Solution Approach 2:
The patent applies local quality by tailoring each PRS model to the specific genetic characteristics of its target population. Population-specific reference panels and genetic data are used to train each model, ensuring that the predictions are optimized for the local genetic architecture of each ancestral group. This approach maintains high reliability for each population while the overall system handles multiple populations through specialized sub-models.
2Measurement precision
If larger training cohorts are used, then model accuracy improves, but data processing complexity increases
Solution Approach 1:
The patent segments the large genetic dataset into population-specific subsets based on ancestral origin. Each population-specific model is trained on its corresponding genetic data, which reduces the complexity of processing entire diverse datasets for each model. This segmentation allows parallel processing of smaller, focused datasets while maintaining the benefits of large sample sizes within each population group.
Solution Approach 2:
The patent changes the parameters of the training process by using population-specific reference panels and adjusting the genetic data processing parameters according to each population's characteristics. This includes using population-appropriate imputation panels, quality control thresholds, and genetic variant filtering criteria, which simplifies data processing for each population while maintaining high model accuracy.
Data Source
AI summary
The disclosed embodiments concern methods, apparatus, systems, and computer program products for developing polygenic risk score (PRS) models. In some implementations, a fully automated process is provided that allows for a PRS model to be defined by an initial set of parameters. In some implementations the PRS models are trained to provide a PRS for particular populations.


