Nucleic Acid Sequencing Read Error Correction with Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sequencing technologies suffer from high error rates, particularly in lower-cost methods like sequencing by flow, which confound downstream analyses by mistaking sequencing errors for true genetic variants, leading to inaccurate variant calling.
Innovation Solution
A machine learning-based sequencing error suppression algorithm is trained on a subset of sequencing data to identify and correct recurrent sequencing errors by distinguishing them from true genetic variants, using a neural network to predict the reference sequence and correct mismatches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If lower-cost sequencing techniques (sequencing by flow) are used, then cost is reduced and throughput is increased, but sequencing error rate increases leading to inaccurate variant calling
Solution Approach 1:
The patent introduces an error correction module as an intermediary component between the sequencing device and variant calling analysis. This module uses machine learning models trained on sequencing data to identify and correct sequencing errors before variant calling, thereby mediating the trade-off between using high-throughput lower-cost sequencing methods and maintaining accurate variant detection
Solution Approach 2:
The patent applies preliminary error correction processing to sequencing reads before they are used for variant calling. By training machine learning models on sequencing data and using them to correct errors in advance, the system prepares the data to maintain accuracy despite using higher-error-rate sequencing methods
2Measurement precision
If sequencing errors are corrected by filtering mismatches, then false positives are reduced, but true genetic variants may be suppressed
Solution Approach 1:
The patent employs dynamic machine learning models that adapt to the specific sequencing data being analyzed. The models are trained on the actual sequencing reads and can dynamically adjust their error correction decisions based on patterns learned from the data, allowing them to distinguish between sequencing errors and true genetic variants rather than applying static filtering rules
Solution Approach 2:
The patent changes the parameters of the sequencing data by using machine learning models to transform raw sequencing reads into corrected reads. The models learn optimal correction parameters from training data and apply these transformations to correct errors while preserving true variants, effectively changing the data parameters to improve accuracy
Data Source
AI summary
Error correction of nucleic acid sequencing reads is described. A machine learning model may be trained using aligned sequencing reads generated during a nucleic acid sequencing event, the aligned sequencing reads aligned to a reference sequence. The trained machine learning model may output predicted reference sequences for the aligned sequencing reads that have a position of mismatch with the reference sequence. The aligned sequencing reads may be selectively corrected according to whether the machine learning model predicts that the aligned sequencing reads are variants or sequencing errors based on the predicted reference sequences relative to the reference sequence.


