Self-Designed SNP Chip for Population-Specific Polygenic Risk Score

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current polygenic risk score (PRS) computations are limited by the use of international SNP chips, which are costly and less predictive for non-European populations due to the majority of genome-wide association studies (GWAS) participants being of European ancestry, leading to a need for a method that can accurately compute PRS for specific populations.

Innovation Solution

A self-designed SNP chip utilizing the LmTag algorithm, which includes a pairwise imputation score computation module, a functional score computation module, and a tag SNP selection module, along with a method for computing PRS using disease-/trait-related gene databases and genomic datasets specific to a given population, such as the 1KVG and 1KGP datasets, to enhance accuracy and reduce costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If international SNP chips are used for PRS computation, then the technology can be applied widely, but the cost increases and predictive accuracy decreases for non-European populations

Engineering Contradiction:
ImprovePRS predictive accuracyVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies local quality by designing a population-specific SNP chip tailored to the characteristics of non-European populations. Instead of using a universal international SNP chip, the invention selects SNPs based on linkage disequilibrium patterns, minor allele frequencies, and imputation scores specific to the target population, thereby improving predictive accuracy locally while controlling costs.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes key parameters including SNP selection criteria (using population-specific MAF thresholds), linkage disequilibrium thresholds, and imputation score cutoffs to optimize the SNP chip for non-European populations. These parameter adjustments enable the chip to capture population-specific genetic variation patterns, improving PRS accuracy without requiring expensive international chip platforms.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If the number of SNPs in the chip is limited, then the cost is reduced, but the predictive accuracy and coverage of genetic variants decrease

Engineering Contradiction:
Improvenumber of SNPsVSAvoidPRS predictive accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by performing extensive SNP selection and filtering before chip design. The LmTag algorithm pre-identifies optimal tag SNPs based on linkage disequilibrium relationships, imputation scores, and population-specific characteristics. This preliminary selection ensures that the limited number of SNPs on the chip are maximally informative, capturing the majority of genetic variation without requiring a large chip capacity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces imputation as an intermediary step that bridges the gap between limited chip SNPs and comprehensive genomic coverage. By using reference panels and imputation algorithms, the system infers genotypes for untyped variants based on linkage disequilibrium with chip-measured SNPs, thereby achieving high predictive accuracy with a cost-effective limited-SNP chip design.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If GWAS data from European populations is used to compute PRS, then the computation can be performed using existing data, but the predictive accuracy decreases for non-European populations

Engineering Contradiction:
Improvedata collection timeVSAvoidPRS predictive accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The patent applies local quality by developing population-specific imputation reference panels and adjusting imputation parameters for non-European populations. Instead of directly applying European GWAS results, the invention creates locally-optimized reference data and imputation models that account for population-specific linkage disequilibrium patterns and allele frequencies, thereby improving PRS accuracy without requiring time-consuming new GWAS studies.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces dynamic adaptation by making the SNP selection and imputation parameters adjustable for different population groups. The system can be configured with population-specific parameters including MAF thresholds, LD thresholds, and imputation score cutoffs, allowing the same computational framework to be optimized for various non-European populations while leveraging existing GWAS infrastructure.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20230335218A1Self-designed single-nucleotide polymorphism chip and method of computing polygenicrisk score for given populations using self-designed single-nucleotide polymorphism chip
Publication Date: 2023.10.19 GENESTORY JOINT STOCK
  • US20230335218A1 patent drawing
  • US20230335218A1 patent drawing
  • US20230335218A1 patent drawing

AI summary

The present invention relates to a self-designed single-nucleotide polymorphism chip and a method of computing polygenic risk score (PRS) for a given population using the self-designed single-nucleotide polymorphism chip. The self-designed single-nucleotide polymorphism chip using LmTag algorithm comprises the following modules: a pairwise imputation score computation module; a functional score computation module; and a tag SNP selection module. The method of computing polygenic risk score for a given population using the self-designed single nucleotide polymorphism chip comprises two computation flows: a first flow computing PRS based on a disease-/trait-related gene database collected from open sources and provided by parties; a second flow computing PRS based on test samples; wherein the self-designed SNP chip used in the VCF file generation stage of both flows; the VCF files is to be subjected to imputation, using a given population genomic dataset as a reference, and harmonized using a data harmonization process; and data generated from these two computation flows is to be harmonized and input into a machine learning model to form a single computation process to generate the polygenic risk score for a group of new test samples.