Polygenic Risk Score Calculation Using Ancestry Space Coordinates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for calculating polygenic risk scores (PRS) fail to accurately account for ancestry variations, leading to inaccurate risk estimates for individuals from diverse or mixed-ancestry populations, as they assume constant effect sizes across populations and do not adequately handle individuals who do not fit into predefined ancestry groups.

Innovation Solution

A computer-implemented method that determines an individual's position in an ancestry space using their genetic data, allowing for a genetic contribution to risk calculation that varies continuously with ancestry, enabling more accurate risk estimation for mixed-ancestry individuals by using a combination of orderable, continuous, or pseudo-continuous variables and incorporating non-linear dependencies through Gaussian processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If polygenic risk scores are calculated using standard methods assuming constant effect sizes across populations, then the calculation process remains simple and computationally efficient, but the accuracy of risk estimation deteriorates for individuals from diverse or mixed-ancestry populations

Engineering Contradiction:
Improverisk estimation accuracyVSAvoidcalculation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a new dimension to the risk calculation by incorporating ancestry space coordinates alongside the polygenic risk score. Instead of using only the PRS value, the system combines it with ancestry information (represented as coordinates in ancestry space) to create a multi-dimensional risk assessment model. This dimensional expansion allows the system to account for population-specific effect size variations while maintaining computational efficiency through linear combination methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the parameters used in risk estimation from a single PRS value to a combination of PRS and ancestry space coordinates. By transforming the input parameters to include ancestry information, the system can adjust effect size calculations dynamically based on an individual's ancestry position, thereby improving accuracy without requiring completely new computational algorithms.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If polygenic risk scores are calculated using predefined ancestry groups, then the classification is straightforward and easy to implement, but it fails to accurately represent individuals with mixed or unknown ancestry who do not fit into predefined categories

Engineering Contradiction:
Improveapplicability to diverse populationsVSAvoidrisk estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent transitions from a static ancestry classification system to a dynamic continuous ancestry representation. Instead of assigning individuals to fixed ancestry categories, the system uses continuous ancestry space coordinates that can represent any position along ancestry gradients. This dynamic approach allows the system to adapt to any individual's unique ancestry composition, including mixed or unknown ancestries, by calculating risk based on their specific position in ancestry space rather than forcing them into predefined groups.

Inventive Principle:
Principle #15Dynamics

3Reliability

If polygenic risk scores are derived from training data representing a specific population, then the model training process is simplified and data requirements are reduced, but the model's performance deteriorates when applied to individuals from different populations due to extrapolation errors

Engineering Contradiction:
Improverisk estimation reliabilityVSAvoiddata requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces ancestry space coordinates as an intermediary variable that mediates between the training data and the individual being assessed. The ancestry coordinates act as a bridge that translates population-specific training information into accurate risk estimates for individuals from any population. By incorporating this intermediary ancestry information, the system can reliably extrapolate from training data to diverse populations without requiring separate training models for each population.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240428883A1Computer-implemented method and apparatus for analysing genetic data
Publication Date: 2024.12.26 GENOMICS PLC
  • US20240428883A1 patent drawing
  • US20240428883A1 patent drawing
  • US20240428883A1 patent drawing

AI summary

There is provided a computer-implemented method of analysing genetic data comprising: receiving a polygenic risk score for a target phenotype or target phenotype combination for a target individual; receiving individual genetic data for the target individual, the individual genetic data informative about an ancestry of the target individual; determining an individual position in an ancestry space using the individual genetic data; and calculating a genetic contribution to a risk for the target individual for the target phenotype or target phenotype combination using the polygenic risk score and the individual position. A corresponding apparatus is also provided.