Neural Network Flow Space Quality Score Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing nucleic acid sequencing systems face challenges in accurately estimating the quality of nucleotide base calls, particularly in identifying and removing low-fidelity base calls, which is crucial for producing high-quality sequencing data.
Innovation Solution
A method and system that utilize flow space signal measurements from reaction confinement regions to generate base calls and flow predictor features. These features are then input into an artificial neural network to estimate the flow space probability of error, which is used to determine the quality value of each nucleotide base call.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional Phred Quality Score methods are used for base call quality estimation, then the process is simple and fast, but the accuracy of quality estimation is insufficient leading to inability to reliably identify low-fidelity base calls
Solution Approach 1:
The patent introduces flow predictor features as intermediary variables that bridge the gap between raw flow space signal measurements and quality score estimation. These features (including predicted base calls, predicted flow values, and residual errors) serve as mediators that capture complex relationships in the data, enabling the neural network to achieve high accuracy quality estimation without directly processing the full complexity of raw signals.
Solution Approach 2:
The patent replaces the traditional mechanical/algorithmic Phred quality score calculation system with a neural network-based system. Instead of using fixed mathematical formulas and lookup tables, the invention employs a trained neural network that learns optimal quality estimation patterns from training data, substituting rigid mechanical computation with adaptive intelligent processing.
2Reliability
If more sophisticated quality estimation methods are implemented, then the accuracy of identifying low-fidelity base calls improves, but the computational complexity and processing time increase
Solution Approach 1:
The patent performs preliminary actions by pre-computing flow predictor features (predicted base calls, predicted flow values, and residual errors) before the quality score estimation step. The neural network is trained in advance on training data to learn the mapping from these features to quality scores. This preliminary processing organizes the data in a way that enables fast, accurate real-time quality estimation during actual sequencing operations.
Solution Approach 2:
The patent segments the quality estimation process into distinct components: generating flow predictor features from raw signals, feeding these features to the neural network, and obtaining quality scores. This segmentation allows each component to be optimized independently and enables parallel processing, reducing overall processing time while maintaining high reliability in identifying low-fidelity base calls.
3Manufacturing precision
If traditional Phred Quality Score is used, then computational resources are conserved, but the ability to produce high-fidelity sequencing data is limited
Solution Approach 1:
The patent changes the parameters used for quality estimation from traditional Phred-based metrics to flow space signal measurements and derived flow predictor features. By transforming the input parameters and using a neural network to process them, the system achieves superior sequencing data quality (higher manufacturing precision) while the computational overhead is managed through efficient feature extraction and neural network optimization.
Data Source
AI summary
An artificial neural network is applied to a plurality of flow predictor features to generate a flow space probability of error for a base call. A base quality value for the base call is determined based on the flow space probability of error. The base call and flow predictor features are based on the flow space signal measurements generated in response to the nucleotide flow to the reaction confinement region. For an array of reaction confinement regions, a plurality of parallel neural networks is applied to produce a probability of error for each reaction confinement region. A given neural network of the parallel neural networks is applied to the plurality of flow predictor features corresponding to a given reaction confinement region in the array to provide the flow space probability of error for the given reaction confinement region.


