Genetic Data Analysis Method for Polygenic Risk Score Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for calculating polygenic risk scores (PRS) face challenges in accuracy and robustness, particularly when dealing with phenotypes like stroke risk, due to variations in data quantity and quality, and the inability to effectively combine information from multiple studies, leading to suboptimal predictive results.
Innovation Solution
A computer-implemented method that analyzes genetic data by determining causal genetic variants and their effect sizes across multiple phenotypes or phenotype combinations, using stochastic sampling and iterative processes to account for correlations and phenotype-specific mechanisms, allowing for more accurate prediction effect sizes and improved PRS calculation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current methods are used to calculate polygenic risk scores, then the calculation can be performed, but the accuracy and robustness are insufficient due to variations in data quantity and quality
Solution Approach 1:
The patent combines information from multiple genetic association studies (input units) to improve predictive accuracy. Each input unit contains association data between genetic variants and phenotypes, and the method integrates these multiple data sources through iterative sampling processes to produce more robust polygenic risk scores that account for variations in data quantity and quality across different studies.
Solution Approach 2:
The method uses stochastic sampling with replacement to generate multiple datasets with varying parameters. By iteratively resampling the input units and recalculating effect sizes across different iterations, the system adapts to variations in data quantity and quality, producing prediction effect sizes that are more accurate and robust to parameter changes in the underlying data.
2Quantity of substance
If information from multiple studies is combined, then statistical power is enhanced, but the complexity of analysis increases
Solution Approach 1:
The patent segments the analysis into distinct input units, where each unit represents a separate genetic association study. This segmentation allows the system to process multiple studies independently while maintaining a unified analysis framework. Each input unit is processed through the same iterative sampling process, which simplifies the integration of multiple data sources by applying a consistent methodology across all studies.
Solution Approach 2:
The method incorporates feedback through iterative sampling and recalibration. In each iteration, the system samples input units, calculates effect sizes, and uses these results to refine subsequent iterations. This feedback mechanism allows the system to automatically adjust to the complexity of combining multiple studies, with each iteration building upon the previous one to enhance statistical power while managing analysis complexity through systematic refinement.
3Measurement precision
If causal variant identification is performed across multiple phenotypes, then prediction accuracy improves, but computational iterations are required
Solution Approach 1:
The patent applies partial action through iterative sampling with replacement, where not all possible combinations of phenotypes and studies are processed simultaneously. Instead, the system performs a limited number of iterations (e.g., 100 iterations) where in each iteration a subset of input units is sampled. This partial processing approach achieves sufficient prediction accuracy by capturing the essential relationships between causal variants and phenotypes without requiring exhaustive computational analysis of all possible combinations.
Data Source
AI summary
Disclosed is a method of analysing genetic data about an organism comprising receiving a plurality of input units. Each input unit comprises information about the association between genetic variants in a region of the genome and phenotypes or phenotype combinations. The method comprises carrying out iterations comprising, for each variant determining for which of the phenotypes or phenotype combinations the variant is causal based on the input units. If the variant is causal for phenotypes or phenotype combinations, a sampled effect size is determined of the variant on the phenotypes or phenotype combinations based on the input units and information about correlations between the variants in the region. For each variant, a prediction effect size is determined variant on the phenotypes or phenotype combinations based on an average across the iterations of the sampled effect sizes or of posterior effect sizes calculated using the sampled effect sizes.


