Trait Prediction Model Generation Using Multi-Population Regularized Regression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genome-wide association studies (GWAS) prediction models have limited accuracy when applied across different ethnic groups due to ethnic differences, particularly when models generated from European data are used for non-European populations, resulting in reduced prediction accuracy for individuals from populations like Japanese, where there are limited large-scale GWAS results.
Innovation Solution
A trait prediction model generation apparatus that generates multiple prediction models for each population based on summary statistics and inter-polymorphism correlated information, using regularized regression and ensemble learning to combine models from different populations, thereby improving prediction accuracy across ethnic groups.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a prediction model is generated based on genome-wide association analysis of the same ethnic group, then prediction accuracy is improved, but the sample size must be large which is not available for non-European populations
Solution Approach 1:
The patent combines prediction models from multiple ethnic groups (European, Asian, African) into a unified polygenic risk score model. It integrates summary statistics from genome-wide association studies across different populations and uses regularized regression to weigh contributions from each population, thereby achieving accurate predictions for non-European populations without requiring large sample sizes specific to each ethnic group.
Solution Approach 2:
The patent creates a universal prediction model that can be applied across multiple ethnic groups rather than developing separate population-specific models. The model uses ancestry informative markers and population-specific weighting to maintain accuracy across diverse populations, making it a multi-functional tool that serves various ethnic groups with a single framework.
2Measurement precision
If a prediction model is generated based on European genome-wide association analysis, then prediction accuracy is improved for Europeans, but prediction accuracy decreases for non-European populations due to ethnic differences
Solution Approach 1:
The patent applies local quality by allowing different weights and contributions from different ethnic groups based on their specific characteristics. It calculates population-specific weights using ancestry informative markers and adjusts the model parameters for each population, enabling the model to adapt to local ethnic differences while maintaining overall accuracy across all groups.
Solution Approach 2:
The patent changes key parameters of the prediction model to accommodate different ethnic groups. It introduces population-specific weighting factors, adjusts the contribution of different SNPs based on ancestry, and modifies the regularization parameters to optimize performance for each population, thereby maintaining high prediction accuracy across diverse ethnic groups.
3Device complexity
If only statistically significant single nucleotide substitutions are retained, then prediction model simplicity is improved, but prediction accuracy is reduced due to loss of informative variants
Solution Approach 1:
The patent applies partial action by retaining not only the most significant single nucleotide substitutions but also including additional variants with lower statistical significance that still contribute to prediction accuracy. It uses regularized regression to selectively include variants based on their predictive value rather than solely based on significance thresholds, thereby maintaining model simplicity while improving accuracy through the inclusion of informative variants.
Data Source
AI summary
According to one embodiment, a trait prediction model generation apparatus generates a plurality of first trait prediction models for each of a plurality of populations, based on summary statistics and inter-polymorphism correlated information. The apparatus generates a second trait prediction model for a specific one of the populations based on regularized regression of the first trait prediction models of each of the populations using a plurality of data sets including single-nucleotide polymorphism data and a trait value.


