Nanopore Sequencing Signal Classification Model
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Nanopore DNA sequencers face accuracy issues due to non-target signals being mixed with blocking signals, leading to incorrect base sequence decoding, primarily caused by impurities and unstable molecular motor activity.
Innovation Solution
A method involving the generation of a trained model using machine learning to classify blocking event data into good and bad data, where the model is trained with teacher data labeled as good or bad, and used to determine the base sequence of biomolecules by filtering out non-target signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If nanopore DNA sequencing is performed without amplification or fluorescent labeling, then high throughput and low running cost are achieved, but measurement precision deteriorates due to non-target signals being mixed with blocking signals
Solution Approach 1:
The patent extracts and removes non-target signals (noise) from the blocking signal data by training a machine learning model to recognize and filter out signals caused by impurities and unstable molecular motor activity, retaining only the valid target signals for accurate base sequence determination
Solution Approach 2:
The patent implements a feedback mechanism where the measured blocking signals are fed into a trained machine learning model that has learned from labeled training data what constitutes valid versus invalid signals, and the model's classification feedback is used to filter and validate the base sequence decoding process
2Measurement precision
If machine learning classification is applied to filter blocking event data, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent introduces a machine learning model as an intermediary component between the signal acquisition system and the base sequence decoding system. This intermediary processes the raw blocking signals, classifies them as valid or invalid, and outputs only the validated signals for further analysis, thereby improving precision while containing complexity through modular design
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Improves the accuracy of DNA sequencing by effectively distinguishing between target and non-target signals, resulting in more precise base sequence determination.
Implementation Method 1
a base sequence is measured by measuring a blocking current generated when a DNA strand passes through a pore (hereinafter, referred to as "nanopore") formed in a thin film while blocking the pore
Implementation Method 2
when a voltage is applied between the first liquid tank and the second liquid tank, an ion current corresponding to the nanopore diameter flows through the nanopore
Implementation Method 3
a potential gradient corresponding to the applied voltage is formed in the nanopore. If the biomolecule is introduced into the first liquid tank, the diffused biomolecule is sent to the second liquid tank via the nanopore according to the generated potential gradient
Implementation Method 4
analysis of the inside of the biomolecule is performed according to the blocking rate of each nucleic acid blocking the nanopore. The biomolecule analyzer includes a measurement unit that measures a blocking signal (a signal representing an ion current flowing between electrodes provided in the device for biomolecule analysis)
Data Source
AI summary
Provided is a method for generating a trained model for classifying blocking event data representing nanopore blocking events in a biomolecule measurement device. The method includes generating a first trained model by executing machine learning of a training model using first teacher data, the first teacher data includes teacher blocking event data and a teacher label, the teacher label indicates whether the teacher blocking event data is classified as Good data or bad data, and the first trained model is configured to classify the blocking event data into good data or bad data. In addition, a method for determining a base sequence a biomolecule and a biomolecule measurement device are provided.


