Polymer Sequence Deduction via Stochastic Signal Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA sequencing methods face challenges in accurately distinguishing between DNA bases due to stochastic polymer motion, leading to sequence order errors exceeding 99.99% accuracy targets, despite advancements in signal recording and noise reduction.
Innovation Solution
A data processing method that utilizes multiple observations of polymer sequences, combined with Hidden Markov Models, to calculate and optimize the probability of sequence accuracy, systematically adjusting trial sequences to maximize the combined probability of matching observed data, thereby reducing the impact of stochastic variations in polymer position.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If DNA sequencing is performed by recording the signal of each base in a serial manner through a nanopore, then the sequencing speed and cost-effectiveness are improved, but sequence order errors increase due to stochastic polymer motion
Solution Approach 1:
The patent creates multiple copies (M≥2) of the same polymer sequence through repeated translocation events and measures each copy independently. By comparing multiple measurements of the identical sequence, the system can identify and correct ordering errors caused by Brownian motion, achieving both high throughput and high accuracy
Solution Approach 2:
The patent implements a feedback mechanism where the measured signals from multiple polymer translocations are compared and analyzed to identify discrepancies. The system uses this feedback to correct sequence order errors by identifying the most likely correct ordering of bases across multiple measurements, thereby improving overall sequencing reliability
2Productivity
If the polymer translocation time is reduced to increase sequencing throughput, then the productivity is improved, but the measurement precision of each base deteriorates
Solution Approach 1:
The patent merges information from multiple independent measurements of the same polymer sequence to achieve high measurement precision. By combining the signals from M≥2 translocation events, the system can distinguish true base signals from noise and stochastic variations, maintaining accuracy even when individual measurement times are short
3Reliability
If multiple observations of the same polymer sequence are made to reduce stochastic motion errors, then the sequencing accuracy is improved, but the measurement time increases
Solution Approach 1:
The patent efficiently creates multiple copies of the measurement data by exploiting the natural stochastic translocation of polymers through nanopores. Multiple translocation events occur spontaneously, and the system captures signals from M≥2 of these events, obtaining multiple independent measurements without requiring deliberate repetition of the entire measurement process
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly improves sequencing accuracy by quantifying and correcting for stochastic motion errors, enhancing the reliability of DNA sequencing systems.
Implementation Method 1
the polymer is subject to diffusive (Brownian) motion due to the impact of other molecules in solution
Data Source
AI summary
A method of processing sequencing data obtained with a polymer sequencing system identifies the most likely monomer sequence of a polymer, regardless of stochastic variations in recorded signals. Polymer sequencing data is recorded and two or more distinct series of pore blocking signals for a section of the polymer are recorded. A value is assigned to each series of pore blocking signals to obtain multiple trial sequences. The probability that each of the trial sequences could have resulting in all of trial sequences is calculated to determine a monomer sequence with the highest probability of resulting in all of the trial sequences, termed the first iteration sequence. The first iteration sequence is systematically altered to maximize the combined probability of the first iteration sequence leading to all the trial sequences in order to obtain a most likely sequence of monomers of the polymer.


