Autoencoder Genomic Imputation Without Reference Panels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genotype imputation methods, such as those using Hidden Markov Models (HMM), are computationally intensive and require large reference panels, posing privacy and scalability concerns, especially for investigators outside of large consortia.
Innovation Solution
The use of artificial neural networks, specifically autoencoders, to dynamically produce predictive genomic data by inputting actual genomic sequences, allowing for imputation of missing genotypes without the need for large reference panels and reducing computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If HMM-based imputation methods are used, then imputation accuracy can be achieved, but computational complexity and resource requirements increase significantly
Solution Approach 1:
The patent replaces the mechanical HMM-based computational system with a neural network-based system. Specifically, it uses an autoencoder architecture with a bottleneck layer that learns compressed representations of genomic data, substituting the traditional HMM iterative optimization process with a trained neural network that performs imputation through forward propagation, significantly reducing computational complexity while maintaining accuracy
Solution Approach 2:
The patent changes the fundamental parameters of the imputation approach by transitioning from HMM state transition matrices and emission probabilities to neural network weights and activation functions. The autoencoder transforms the input genotype data through encoded representations with reduced dimensionality, using learned parameters rather than predefined genetic models, thereby simplifying the computational framework
2Measurement precision
If large WGS-based reference panels are used, then imputation accuracy improves, but privacy concerns and scalability issues arise
Solution Approach 1:
The patent extracts and removes the requirement for large WGS-based reference panels from the imputation process. By using an autoencoder that learns from training data to create a compressed genomic code, the system eliminates the need to store and process extensive reference panels, thereby improving scalability and reducing privacy concerns associated with handling large genomic datasets
Solution Approach 2:
The patent creates a learned compressed representation (encoding) of genomic data that serves as a surrogate for the actual large reference panels. This encoded representation captures the essential genetic variation patterns without requiring the original large-scale WGS data, allowing the system to function with reduced data storage requirements and improved scalability
3Measurement precision
If HMM algorithms with MCMC or expectation-maximization are applied, then genotype imputation can be performed, but processing time and computational resources increase
Solution Approach 1:
The patent performs preliminary action by pre-training the autoencoder on training data to learn the compressed genomic representations before actual imputation is needed. This offline training phase captures the essential genetic patterns, so that during runtime, the system only needs to perform fast forward propagation through the trained network, eliminating the need for time-consuming MCMC sampling or expectation-maximization iterations at the time of imputation
4Measurement precision
If investigators submit genotype data to imputation servers, then imputation can be performed, but privacy and scalability concerns increase
Solution Approach 1:
The patent enables self-service by providing local imputation capability through the autoencoder system. Investigators can train and deploy the model locally on their own data without needing to submit sensitive genotype information to external servers. The system performs imputation autonomously using the learned compressed representations, eliminating the need for data submission and thereby resolving privacy concerns while maintaining scalability
Data Source
AI summary
The invention relates to the use of artificial neural networks, such as autoencoders, in population genomics and individual genome processing to fill-in missing data from genomic assays with significant sparsity, such as low-pass whole genome sequencing or array-based genotyping. An autoencoder-based neural network approach for the simultaneous execution of genetic imputation as well as feature extraction and dimensionality reduction for downstream tasks is provided.


