Polygenic Risk Score Calculation Using Ancestry Space Coordinates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for calculating polygenic risk scores (PRS) fail to accurately account for ancestry variations, leading to inaccurate risk estimates for individuals from diverse or mixed-ancestry populations, as they assume constant effect sizes across populations and do not adequately handle individuals who do not fit into predefined ancestry groups.
Innovation Solution
A computer-implemented method that determines an individual's position in an ancestry space using their genetic data, allowing for a genetic contribution to risk calculation that varies continuously with ancestry, enabling more accurate risk estimation for mixed-ancestry individuals by using a combination of orderable, continuous, or pseudo-continuous variables and incorporating non-linear dependencies through Gaussian processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If polygenic risk scores are calculated using standard methods assuming constant effect sizes across populations, then the calculation process remains simple and computationally efficient, but the accuracy of risk estimation deteriorates for individuals from diverse or mixed-ancestry populations
Solution Approach 1:
The patent introduces a new dimension to the risk calculation by incorporating ancestry space coordinates alongside the polygenic risk score. Instead of using only the PRS value, the system combines it with ancestry information (represented as coordinates in ancestry space) to create a multi-dimensional risk assessment model. This dimensional expansion allows the system to account for population-specific effect size variations while maintaining computational efficiency through linear combination methods.
Solution Approach 2:
The patent changes the parameters used in risk estimation from a single PRS value to a combination of PRS and ancestry space coordinates. By transforming the input parameters to include ancestry information, the system can adjust effect size calculations dynamically based on an individual's ancestry position, thereby improving accuracy without requiring completely new computational algorithms.
2Adaptability or versatility
If polygenic risk scores are calculated using predefined ancestry groups, then the classification is straightforward and easy to implement, but it fails to accurately represent individuals with mixed or unknown ancestry who do not fit into predefined categories
Solution Approach 1:
The patent transitions from a static ancestry classification system to a dynamic continuous ancestry representation. Instead of assigning individuals to fixed ancestry categories, the system uses continuous ancestry space coordinates that can represent any position along ancestry gradients. This dynamic approach allows the system to adapt to any individual's unique ancestry composition, including mixed or unknown ancestries, by calculating risk based on their specific position in ancestry space rather than forcing them into predefined groups.
3Reliability
If polygenic risk scores are derived from training data representing a specific population, then the model training process is simplified and data requirements are reduced, but the model's performance deteriorates when applied to individuals from different populations due to extrapolation errors
Solution Approach 1:
The patent introduces ancestry space coordinates as an intermediary variable that mediates between the training data and the individual being assessed. The ancestry coordinates act as a bridge that translates population-specific training information into accurate risk estimates for individuals from any population. By incorporating this intermediary ancestry information, the system can reliably extrapolate from training data to diverse populations without requiring separate training models for each population.
Data Source
AI summary
There is provided a computer-implemented method of analysing genetic data comprising: receiving a polygenic risk score for a target phenotype or target phenotype combination for a target individual; receiving individual genetic data for the target individual, the individual genetic data informative about an ancestry of the target individual; determining an individual position in an ancestry space using the individual genetic data; and calculating a genetic contribution to a risk for the target individual for the target phenotype or target phenotype combination using the polygenic risk score and the individual position. A corresponding apparatus is also provided.


