Neural Network Sequence Calling for Homopolymer Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing nucleic acid sequencing methods face challenges in accurate base calling, particularly in regions with repeating nucleotide bases (homopolymers), due to signal variations and context dependency, leading to sequencing errors.
Innovation Solution
A method involving neural networks is used to generate training sets by aligning actual sequencing signals with trusted reference signals, allowing for improved base calling and quantification of homopolymer lengths and context dependency, using algorithms trained on smaller reference genomes to estimate larger genomes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If base calling is performed based on quantified characteristic signals indicating nucleotide incorporation, then sequencing speed and throughput are improved, but sequencing accuracy deteriorates due to random and unpredictable systematic variations in signal levels and context dependent signals
Solution Approach 1:
The patent introduces an intermediary computational model (neural network or machine learning algorithm) that mediates between the raw sequencing signals and the base calling decision. This intermediary processes the quantified characteristic signals, learns to account for systematic variations and context dependencies, and produces accurate base calls without reducing throughput. The model acts as a bridge that transforms noisy quantitative signals into reliable sequence information.
Solution Approach 2:
The patent employs feedback mechanisms where the system learns from training data consisting of actual sequencing signals and corresponding known reference sequences. The model is iteratively trained to minimize the difference between predicted and actual sequences, incorporating feedback about signal variations and context dependencies. This feedback loop enables the system to adapt and improve accuracy while maintaining high throughput sequencing.
2Loss of time
If machine learning models are trained on smaller reference genomes and applied to larger genomes, then training time and computational resources are reduced, but model accuracy may deteriorate due to genomic differences
Solution Approach 1:
The patent performs preliminary training on smaller, faster-reference genomes to establish an initial model that captures fundamental sequencing signal characteristics. This preliminary action creates a starting point that can be quickly adapted to larger genomes through transfer learning or fine-tuning, rather than training from scratch. The preliminary model on smaller genomes provides a head start that reduces overall training time while maintaining the ability to accurately estimate sequences in larger genomes.
Solution Approach 2:
The patent develops a universal base calling model that can be applied across different genome sizes and species. The model learns transferable features from smaller reference genomes that are applicable to larger genomes. By designing the model to be genome-agnostic and focusing on universal sequencing signal patterns rather than genome-specific features, the system achieves both efficient training on small genomes and accurate application to large genomes.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure provides methods, systems, and media for accurate and efficient estimation of a genome of a genus.