Trait Prediction Model Generation Using Multi-Population Regularized Regression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genome-wide association studies (GWAS) prediction models have limited accuracy when applied across different ethnic groups due to ethnic differences, particularly when models generated from European data are used for non-European populations, resulting in reduced prediction accuracy for individuals from populations like Japanese, where there are limited large-scale GWAS results.

Innovation Solution

A trait prediction model generation apparatus that generates multiple prediction models for each population based on summary statistics and inter-polymorphism correlated information, using regularized regression and ensemble learning to combine models from different populations, thereby improving prediction accuracy across ethnic groups.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a prediction model is generated based on genome-wide association analysis of the same ethnic group, then prediction accuracy is improved, but the sample size must be large which is not available for non-European populations

Engineering Contradiction:
Improveprediction accuracyVSAvoidsample size
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent combines prediction models from multiple ethnic groups (European, Asian, African) into a unified polygenic risk score model. It integrates summary statistics from genome-wide association studies across different populations and uses regularized regression to weigh contributions from each population, thereby achieving accurate predictions for non-European populations without requiring large sample sizes specific to each ethnic group.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal prediction model that can be applied across multiple ethnic groups rather than developing separate population-specific models. The model uses ancestry informative markers and population-specific weighting to maintain accuracy across diverse populations, making it a multi-functional tool that serves various ethnic groups with a single framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If a prediction model is generated based on European genome-wide association analysis, then prediction accuracy is improved for Europeans, but prediction accuracy decreases for non-European populations due to ethnic differences

Engineering Contradiction:
Improveprediction accuracyVSAvoidcross-ethnicity applicability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by allowing different weights and contributions from different ethnic groups based on their specific characteristics. It calculates population-specific weights using ancestry informative markers and adjusts the model parameters for each population, enabling the model to adapt to local ethnic differences while maintaining overall accuracy across all groups.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes key parameters of the prediction model to accommodate different ethnic groups. It introduces population-specific weighting factors, adjusts the contribution of different SNPs based on ancestry, and modifies the regularization parameters to optimize performance for each population, thereby maintaining high prediction accuracy across diverse ethnic groups.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If only statistically significant single nucleotide substitutions are retained, then prediction model simplicity is improved, but prediction accuracy is reduced due to loss of informative variants

Engineering Contradiction:
Improvemodel complexityVSAvoidprediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies partial action by retaining not only the most significant single nucleotide substitutions but also including additional variants with lower statistical significance that still contribute to prediction accuracy. It uses regularized regression to selectively include variants based on their predictive value rather than solely based on significance thresholds, thereby maintaining model simplicity while improving accuracy through the inclusion of informative variants.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20220189580A1Trait prediction model generation apparatus, trait prediction apparatus, and method for generating a trait prediction model
Publication Date: 2022.06.16 KK TOSHIBA
  • US20220189580A1 patent drawing
  • US20220189580A1 patent drawing
  • US20220189580A1 patent drawing

AI summary

According to one embodiment, a trait prediction model generation apparatus generates a plurality of first trait prediction models for each of a plurality of populations, based on summary statistics and inter-polymorphism correlated information. The apparatus generates a second trait prediction model for a specific one of the populations based on regularized regression of the first trait prediction models of each of the populations using a plurality of data sets including single-nucleotide polymorphism data and a trait value.