Genetic Data Analysis Method for Polygenic Risk Score Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for calculating polygenic risk scores (PRS) face challenges in accuracy and robustness, particularly when dealing with phenotypes like stroke risk, due to variations in data quantity and quality, and the inability to effectively combine information from multiple studies, leading to suboptimal predictive results.

Innovation Solution

A computer-implemented method that analyzes genetic data by determining causal genetic variants and their effect sizes across multiple phenotypes or phenotype combinations, using stochastic sampling and iterative processes to account for correlations and phenotype-specific mechanisms, allowing for more accurate prediction effect sizes and improved PRS calculation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current methods are used to calculate polygenic risk scores, then the calculation can be performed, but the accuracy and robustness are insufficient due to variations in data quantity and quality

Engineering Contradiction:
Improvepredictive accuracyVSAvoidrobustness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines information from multiple genetic association studies (input units) to improve predictive accuracy. Each input unit contains association data between genetic variants and phenotypes, and the method integrates these multiple data sources through iterative sampling processes to produce more robust polygenic risk scores that account for variations in data quantity and quality across different studies.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The method uses stochastic sampling with replacement to generate multiple datasets with varying parameters. By iteratively resampling the input units and recalculating effect sizes across different iterations, the system adapts to variations in data quantity and quality, producing prediction effect sizes that are more accurate and robust to parameter changes in the underlying data.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If information from multiple studies is combined, then statistical power is enhanced, but the complexity of analysis increases

Engineering Contradiction:
Improvestatistical powerVSAvoidanalysis complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the analysis into distinct input units, where each unit represents a separate genetic association study. This segmentation allows the system to process multiple studies independently while maintaining a unified analysis framework. Each input unit is processed through the same iterative sampling process, which simplifies the integration of multiple data sources by applying a consistent methodology across all studies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method incorporates feedback through iterative sampling and recalibration. In each iteration, the system samples input units, calculates effect sizes, and uses these results to refine subsequent iterations. This feedback mechanism allows the system to automatically adjust to the complexity of combining multiple studies, with each iteration building upon the previous one to enhance statistical power while managing analysis complexity through systematic refinement.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If causal variant identification is performed across multiple phenotypes, then prediction accuracy improves, but computational iterations are required

Engineering Contradiction:
Improveprediction accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action through iterative sampling with replacement, where not all possible combinations of phenotypes and studies are processed simultaneously. Instead, the system performs a limited number of iterations (e.g., 100 iterations) where in each iteration a subset of input units is sampled. This partial processing approach achieves sufficient prediction accuracy by capturing the essential relationships between causal variants and phenotypes without requiring exhaustive computational analysis of all possible combinations.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240105280A1Computer-implemented method and apparatus for analysing genetic data
Publication Date: 2024.03.28 GENOMICS PLC
  • US20240105280A1 patent drawing
  • US20240105280A1 patent drawing
  • US20240105280A1 patent drawing

AI summary

Disclosed is a method of analysing genetic data about an organism comprising receiving a plurality of input units. Each input unit comprises information about the association between genetic variants in a region of the genome and phenotypes or phenotype combinations. The method comprises carrying out iterations comprising, for each variant determining for which of the phenotypes or phenotype combinations the variant is causal based on the input units. If the variant is causal for phenotypes or phenotype combinations, a sampled effect size is determined of the variant on the phenotypes or phenotype combinations based on the input units and information about correlations between the variants in the region. For each variant, a prediction effect size is determined variant on the phenotypes or phenotype combinations based on an average across the iterations of the sampled effect sizes or of posterior effect sizes calculated using the sampled effect sizes.