Neural Network Base Calling for Homopolymer Sequencing Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid sequencing methods face challenges in accurate base calling, particularly when dealing with sequences containing homopolymers, due to random and unpredictable systematic variations in signal levels and context-dependent signals, leading to errors in quantifying homopolymer lengths.
Innovation Solution
A method involving a trained algorithm, such as a neural network, is applied to sequencing signals to estimate the likelihood of a particular nucleic acid sequence, allowing for accurate and efficient base calling by padding trimmed reads with filler values and aligning them in flow space, thereby reducing errors associated with context dependence and homopolymer length quantification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If base calling is performed based on quantified characteristic signals indicating nucleotide incorporation, then sequencing speed and throughput are improved, but sequencing accuracy deteriorates due to random and unpredictable systematic variations in signal levels and context dependent signals
Solution Approach 1:
The patent introduces an intermediary trained algorithm (neural network or support vector machine) that mediates between the raw sequencing signals and the base calling process. This algorithm learns the complex relationships between signal characteristics and actual nucleotide sequences, correcting for systematic variations and context-dependent signals, thereby maintaining high throughput while improving base calling accuracy
Solution Approach 2:
The patent transforms the base calling approach by changing from direct signal quantification to a machine learning-based parameter estimation. The trained algorithm processes multiple signal parameters simultaneously and learns optimal decision boundaries, enabling accurate base calling even in the presence of signal variations while maintaining high sequencing throughput
2Speed
If traditional base calling methods are used to handle homopolymer regions, then processing speed is maintained, but sequencing accuracy deteriorates due to errors in quantifying homopolymer lengths
Solution Approach 1:
The trained algorithm acts as an intermediary that processes homopolymer region signals differently from traditional methods. By learning from training data containing homopolymer sequences, the algorithm accurately estimates homopolymer lengths by analyzing signal patterns across multiple cycles, maintaining processing speed while dramatically improving quantification accuracy
Solution Approach 2:
The patent performs preliminary training of the algorithm on diverse sequences including homopolymers before actual base calling. This preliminary action enables the algorithm to recognize and correctly interpret homopolymer signal patterns during sequencing, achieving both speed and accuracy in homopolymer region analysis
3Device complexity
If context dependent signals are not corrected, then algorithm complexity is reduced, but sequencing accuracy deteriorates due to context dependent variations affecting base calling
Solution Approach 1:
The patent transforms the base calling problem by changing from simple signal thresholding to a machine learning approach that considers multiple signal parameters and their interactions. The trained algorithm automatically learns to account for context-dependent variations, achieving high accuracy while the complexity is managed through efficient algorithm design and training procedures
Data Source
AI summary
The present disclosure provides methods, systems, and media for accurate and efficient estimation of a genome of a genus. The methods and systems described herein may be used to accurately determine a base sequence of a polynucleotide. Additionally, the methods and systems may be used to identify base variants of a polynucleotide.


