Latent Space Encoding for Genotype Imputation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current plant breeding methods require extensive resources and labor to evaluate and integrate information from multiple crosses, making it inefficient for predicting genotypes and phenotypes, especially when dealing with genetically divergent populations with different marker sets.
Innovation Solution
A universal method using machine learning-based frameworks, such as variational autoencoders and generative adversarial networks, to encode and decode genotypic or phenotypic data into latent vectors, allowing for the imputation or prediction of genotypes and phenotypes across disparate populations and marker platforms, independent of the underlying data generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional plant breeding methods are used to evaluate and integrate information from multiple crosses, then accurate genotype and phenotype prediction can be achieved, but extensive resources and labor are required
Solution Approach 1:
The patent replaces manual breeding evaluation processes with machine learning-based computational systems. Neural networks and autoencoders process genotypic and phenotypic data to predict breeding outcomes, substituting the mechanical and labor-intensive traditional breeding methods with automated computational approaches that maintain prediction accuracy while dramatically improving breeding efficiency
Solution Approach 2:
The patent creates virtual representations of breeding populations through latent space encodings and predictive models. Instead of physically evaluating every cross and progeny, the system generates computational copies and simulations that predict genetic outcomes, reducing the need for extensive physical breeding trials while maintaining predictive accuracy
2Adaptability or versatility
If genotypic data from multiple different marker platforms are integrated, then comprehensive genotype imputation can be performed, but data heterogeneity and platform-specific biases increase
Solution Approach 1:
The patent introduces latent space representations as intermediary structures between different marker platforms. The autoencoder framework transforms diverse genotypic data from different platforms into a unified latent space, serving as a mediator that reconciles platform-specific variations and enables consistent genotype imputation across heterogeneous data sources
Solution Approach 2:
The patent transforms genotypic data from different marker platforms by changing the parameter space through encoding into latent vectors. This parameter transformation converts platform-specific markers into a universal latent representation that captures essential genetic information while eliminating platform-specific biases and heterogeneity
3Measurement precision
If extensive field trials across wide geographic regions are conducted, then robust phenotype data are obtained, but time and resource consumption increase
Solution Approach 1:
The patent performs preliminary phenotypic predictions using machine learning models trained on historical data before conducting actual field trials. This preliminary action allows breeders to prioritize which crosses and progeny warrant extensive field evaluation, reducing the overall time and resources needed while maintaining data quality through targeted sampling
Solution Approach 2:
The patent implements feedback loops where phenotypic data from field trials are continuously fed back into the machine learning models to improve predictions. This feedback mechanism allows the system to learn from actual trial results and refine future predictions, reducing the need for increasingly extensive trials over time while maintaining or improving data quality
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
Methods and compositions to impute or predict genotype, haplotype, molecular phenotype, agronomic phenotypes, and/or coancestry are provided. Methods and compositions provided include using latent space to generate latent space representations or latent vectors that are independent of underlying genotypic or phenotypic data. The methods may include generating a universal latent space representation by encoding discrete or continuous variables derived from genotypic or phenotypic data into latent vectors through a machine learning-based encoder framework. Provided herein are universal methods of parametrically representing genotypic or phenotypic data obtained from one or more populations or sample sets to impute or predict a genotype or phenotype of interest.