Bayesian Multilevel Models for Genomic Prediction Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In animal and plant breeding, assembling large estimation sets for genomic prediction is hindered by population structuring and genetic differences, leading to challenges in increasing prediction accuracy through pooling data across breeds or populations.
Innovation Solution
The use of Bayesian multilevel models for partial pooling, which estimates population-specific marker effects while leveraging information across populations, strikes a balance between no pooling and complete pooling by accounting for unique genetic characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If complete pooling of estimation sets across populations is performed, then the size of the estimation set increases, but prediction accuracy deteriorates due to genetic differences among populations
Solution Approach 1:
The patent segments the estimation set into population-specific components, allowing separate estimation of marker effects for each population while sharing information across populations through hierarchical modeling. This resolves the contradiction by maintaining population specificity (preserving prediction accuracy) while still utilizing data from multiple populations (increasing estimation set size).
Solution Approach 2:
The patent applies local quality by allowing different populations to have population-specific marker effects and variance parameters while sharing common hyperparameters across populations. This enables the model to adapt to local genetic characteristics of each population (maintaining prediction accuracy) while leveraging information from the overall pooled data (increasing estimation set size).
2Measurement precision
If separate estimation sets are used for each population, then prediction accuracy is maintained for population-specific traits, but the size of the estimation set remains limited
Solution Approach 1:
The patent merges multiple population-specific estimation sets into a unified hierarchical model that combines data from all populations. Through partial pooling, the model shares information across populations (increasing effective estimation set size) while maintaining population-specific parameters (preserving prediction accuracy for population-specific traits).
Solution Approach 2:
The patent creates a universal hierarchical modeling framework that can handle multiple populations simultaneously, allowing the same model structure to function for both population-specific and cross-population predictions. This enables the estimation set to serve multiple functions: maintaining population specificity when needed and pooling data when beneficial.
3Quantity of substance
If pooled estimation sets are used combining populations, then the size of the estimation set increases, but modeling of population-specific traits becomes difficult
Solution Approach 1:
The patent segments the marker effects into population-specific and common components through hierarchical modeling. This allows the model to simultaneously handle pooled data (increasing estimation set size) and population-specific traits (maintaining adaptability) by estimating separate parameters for each population while sharing information across populations.
Solution Approach 2:
The patent employs dynamic hierarchical modeling where the degree of pooling can vary by marker and population based on the data. This allows the model to adaptively balance between complete pooling (for increasing estimation set size) and no pooling (for modeling population-specific traits), making the approach flexible and versatile across different scenarios.
Data Source
AI summary
A Bayesian multilevel whole-genome regression model is disclosed and its prediction performance compared to that of the popular BayesA model applied to each population separately (no pooling) and to the joined data set (complete pooling). For small population sizes (e.g., <50), partial pooling increased prediction accuracy over no or complete pooling for populations represented in the estimation set. Partial pooling with multilevel models can make optimal use of information in multi-population estimation sets.

