Cross-Ancestry Polygenic Risk Score Models for Diverse Populations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current Polygenic Risk Score (PRS) models face limitations due to the need for large sample sizes for accurate predictions, particularly when applied across different ancestral populations, as models developed in one ancestry group perform poorly in others due to varying genetic variants.

Innovation Solution

The development of cross-traits and transethnic PRS models that leverage genetic correlations and penalized linear or logistic regression to combine PRS models from multiple populations, using elastic net regularization and 10-fold cross-validation to generate more accurate and transferable risk scores.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If PRS models are developed using large sample sizes from a specific ancestry group, then prediction accuracy for that group is improved, but applicability to other ancestral populations deteriorates

Engineering Contradiction:
Improveprediction accuracyVSAvoidapplicability across ancestral populations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent combines PRS models from multiple ancestral populations into a unified cross-ancestry model. By integrating genetic data and phenotypic information from diverse populations, the system creates a consolidated model that leverages the strengths of each population-specific model while reducing ancestry-specific biases, thereby improving both accuracy and cross-population applicability

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent develops a universal PRS model that functions across multiple ancestral populations simultaneously. The model is designed to be ancestry-agnostic by using transfer learning and domain adaptation techniques, allowing a single model to serve multiple population groups with varying genetic backgrounds, thus achieving both high accuracy and broad applicability

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If PRS models are developed for populations with limited genetic data, then prediction accuracy for those populations can be improved, but the complexity of the modeling process increases

Engineering Contradiction:
Improveprediction accuracy for underrepresented populationsVSAvoidmodeling process complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses well-established PRS models from populations with abundant genetic data as intermediary references to improve prediction accuracy for populations with limited data. Through transfer learning, the system adapts knowledge from data-rich populations to data-scarce populations, acting as a bridge that transfers predictive power without requiring extensive native training data, thereby improving accuracy while managing complexity through automated adaptation procedures

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220044761A1Machine learning platform for generating risk models
Publication Date: 2022.02.10 23ANDME GENOMICS LLC
  • US20220044761A1 patent drawing
  • US20220044761A1 patent drawing
  • US20220044761A1 patent drawing

AI summary

The disclosed embodiments concern methods, apparatus, systems, and computer program products for developing polygenic risk score (PRS) models with improved performance across different ethnicities and for different target phenotypes.