Machine Learning Basecaller for DNA Sequencing Error Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA sequencing methods face high error rates in base calling due to optical and biochemical variations, leading to inaccurate intensity signal interpretation and artifacts, which complicates the determination of nucleotide bases in DNA fragments.
Innovation Solution
A machine-learning based basecalling model is developed using training data from previous sequencing runs, incorporating stringent settings and filtering processes to optimize the model for accurate base calling, which can be applied in subsequent sequencing runs, even weeks or months later, to improve the accuracy of nucleotide base determination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If naive maximum intensity value method is used for base calling, then the process is simple and fast, but the error rate is high due to optical effects and spatial effects
Solution Approach 1:
The patent applies preliminary action by training a machine learning model in advance using extensive training data from multiple sequencing runs. The model learns to account for optical effects, spatial effects, and biochemical variations before actual base calling. This pre-trained model can then quickly and accurately call bases in production runs without requiring complex real-time calculations, thus resolving the contradiction between speed and accuracy.
Solution Approach 2:
The patent introduces a machine learning model as an intermediary between the raw intensity values and the base calls. This model acts as a mediator that processes the intensity values while accounting for various effects (optical, spatial, biochemical) and produces accurate base calls. The intermediary model captures complex relationships without requiring explicit programming of correction rules, thereby improving accuracy while maintaining computational efficiency.
2Measurement precision
If stringent settings are used during production sequencing runs to improve training data accuracy, then the accuracy of training data increases, but the productivity of production runs decreases
Solution Approach 1:
The patent applies preliminary action by performing the time-consuming stringent filtering and accuracy optimization during the training phase, which is conducted separately from production runs. The training phase uses stringent settings to create high-quality training data and train the model, while production runs use the pre-trained model for rapid base calling without applying stringent filters. This separates the accuracy-optimization step from the high-throughput step, resolving the contradiction between training data accuracy and production productivity.
3Measurement precision
If a machine learning model is trained using extensive training data from multiple sequencing runs, then the accuracy of base calling improves, but the time required for model optimization increases
Solution Approach 1:
The patent applies preliminary action by performing the extensive model training using multiple sequencing runs and stringent settings in advance, before production use. The time-consuming model optimization is completed during the training phase, and the resulting pre-trained model can then be deployed for rapid base calling in production runs. This one-time investment in training time resolves the contradiction by enabling fast and accurate base calling during actual sequencing operations.
Data Source
AI summary
Methods, systems, and apparatuses are provided for creating and using a machine-leaning model to call a base at a position of a nucleic acid based on intensity values measured during a production sequencing run. The model can be trained using training data from training sequencing runs performed earlier. The model is trained using intensity values and assumed sequences that are determined as the correct output. The training data can be filtered to improve accuracy. The training data can be selected in a specific manner to be representative of the type of organism to be sequenced. The model can be trained to use intensity signals from multiple cycles and from neighboring nucleic acids to improve accuracy in the base calls.


