Genetic Data Analysis Using Fine-Mapping and Residual PRS Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for constructing polygenic risk scores (PRS) using summary statistics data are limited by uncertainties in linkage disequilibrium patterns and population variability, leading to reduced accuracy and robustness.

Innovation Solution

A computer-implemented method that combines fine-mapping algorithms with machine learning to identify independent phenotype-variant associations, accounting for residual signals and population-specific correlations, using both summary statistics and individual level data to enhance PRS accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If fine-mapping algorithms are applied to identify independent phenotype-variant associations, then measurement precision of causal variants is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of PRSVSAvoidcomplexity of analysis method
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The method segments the analysis into distinct stages: first applying fine-mapping algorithms to identify independent phenotype-variant associations and potentially causal variants, then applying machine learning algorithms to residual signals. This segmentation allows each algorithm to focus on specific aspects of the problem, improving overall measurement precision while managing complexity through structured decomposition of the analytical process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary step of analyzing residual signals after fine-mapping. By applying machine learning algorithms to the residual association data (the portion of signal not explained by fine-mapped variants), the method captures additional correlations that fine-mapping alone misses. This intermediary analysis acts as a bridge between traditional fine-mapping and final PRS construction, improving accuracy without fully compromising complexity management.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If machine learning algorithms are applied to residual signals, then reliability of PRS across populations is improved, but loss of information increases

Engineering Contradiction:
Improverobustness across populationsVSAvoidinformation about variant correlations
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The method performs preliminary fine-mapping analysis before applying machine learning to residual signals. By first identifying and accounting for strong phenotype-variant associations through fine-mapping, the approach prepares the data in advance, allowing machine learning algorithms to focus on capturing subtler correlation patterns. This preliminary action ensures that the most significant genetic effects are captured first, improving reliability while minimizing information loss in subsequent analysis stages.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a feedback mechanism where machine learning algorithms analyze residual signals that represent the portion of phenotypic variation not explained by fine-mapped variants. The insights gained from this residual analysis feed back into the overall PRS construction, allowing the model to iteratively improve by capturing additional correlation structures. This feedback loop enhances reliability across diverse populations while systematically preserving information about variant correlations that would otherwise be lost.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12626783B2Computer-implemented method and apparatus for analysing genetic data
Publication Date: 2026.05.12 GENOMICS PLC
  • US12626783B2 patent drawing
  • US12626783B2 patent drawing
  • US12626783B2 patent drawing

AI summary

The disclosure relates to analysing genetic data. In one arrangement, a method operates on input data comprising strengths of association between one or more phenotypes including a target phenotype and a plurality of genetic variants. A fine-mapping algorithm is applied to all or a subset of the input data to identify one or more independent phenotype-variant associations. A set of one or more fine-mapped variants is identified for each association. A fine-mapping predictive model is calculated on the basis of the input data and the set of fine-mapped variants. The effect on the target phenotype of the set of fine-mapped variants is subtracted from the input data to obtain residual association data. A machine learning algorithm is applied to the residual association data to identify further predictive correlations between the target phenotype and the plurality of genetic variants.