Machine Learning Basecaller for DNA Sequencing Error Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA sequencing methods face high error rates in base calling due to optical and biochemical variations, leading to inaccurate intensity signal interpretation and artifacts, which complicates the determination of nucleotide bases in DNA fragments.

Innovation Solution

A machine-learning based basecalling model is developed using training data from previous sequencing runs, incorporating stringent settings and filtering processes to optimize the model for accurate base calling, which can be applied in subsequent sequencing runs, even weeks or months later, to improve the accuracy of nucleotide base determination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If naive maximum intensity value method is used for base calling, then the process is simple and fast, but the error rate is high due to optical effects and spatial effects

Engineering Contradiction:
Improvebase calling speedVSAvoidbase calling accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by training a machine learning model in advance using extensive training data from multiple sequencing runs. The model learns to account for optical effects, spatial effects, and biochemical variations before actual base calling. This pre-trained model can then quickly and accurately call bases in production runs without requiring complex real-time calculations, thus resolving the contradiction between speed and accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a machine learning model as an intermediary between the raw intensity values and the base calls. This model acts as a mediator that processes the intensity values while accounting for various effects (optical, spatial, biochemical) and produces accurate base calls. The intermediary model captures complex relationships without requiring explicit programming of correction rules, thereby improving accuracy while maintaining computational efficiency.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If stringent settings are used during production sequencing runs to improve training data accuracy, then the accuracy of training data increases, but the productivity of production runs decreases

Engineering Contradiction:
Improvetraining data accuracyVSAvoidproduction run throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by performing the time-consuming stringent filtering and accuracy optimization during the training phase, which is conducted separately from production runs. The training phase uses stringent settings to create high-quality training data and train the model, while production runs use the pre-trained model for rapid base calling without applying stringent filters. This separates the accuracy-optimization step from the high-throughput step, resolving the contradiction between training data accuracy and production productivity.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If a machine learning model is trained using extensive training data from multiple sequencing runs, then the accuracy of base calling improves, but the time required for model optimization increases

Engineering Contradiction:
Improvebase calling accuracyVSAvoidmodel training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing the extensive model training using multiple sequencing runs and stringent settings in advance, before production use. The time-consuming model optimization is completed during the training phase, and the resulting pre-trained model can then be deployed for rapid base calling in production runs. This one-time investment in training time resolves the contradiction by enabling fast and accurate base calling during actual sequencing operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10068053B2Basecaller for DNA sequencing using machine learning
Publication Date: 2018.09.04 MGI TECH CO LTD
  • US10068053B2 patent drawing
  • US10068053B2 patent drawing
  • US10068053B2 patent drawing

AI summary

Methods, systems, and apparatuses are provided for creating and using a machine-leaning model to call a base at a position of a nucleic acid based on intensity values measured during a production sequencing run. The model can be trained using training data from training sequencing runs performed earlier. The model is trained using intensity values and assumed sequences that are determined as the correct output. The training data can be filtered to improve accuracy. The training data can be selected in a specific manner to be representative of the type of organism to be sequenced. The model can be trained to use intensity signals from multiple cycles and from neighboring nucleic acids to improve accuracy in the base calls.