Nanopore DNA Decoding With Soft Error Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA-based storage systems using nanopore sequencers face high error rates due to substitution, insertion, and deletion errors during DNA sequencing and synthesis, which are exacerbated by the limitations of existing sequencing technologies and synthesis methods.
Innovation Solution
A method and device for decoding DNA sequences using a soft decoding algorithm with error correction codes, specifically employing quaternary encoding/decoding schemes and statistical distributions of ion current signal amplitudes to correct substitution errors, implemented in a DNA-based data storage system with a nanopore sequencer.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If nanopore sequencers are used to read long sequences in a single step, then sequencing efficiency is improved, but error rate increases
Solution Approach 1:
The patent applies preliminary action by incorporating error correction codes during the DNA sequence encoding stage, before sequencing occurs. The encoder processes the input sequence and generates a coded output that inherently protects against substitution errors, which are then corrected during decoding without requiring re-sequencing
Solution Approach 2:
The patent implements feedback through iterative decoding algorithms that use soft information from probability density functions to progressively refine error correction. The decoder continuously adjusts its estimates based on feedback from previous decoding iterations, improving accuracy while maintaining efficiency
2Ease of manufacture
If micro-array based synthesis is used to synthesize DNA sequences, then manufacturing cost is reduced, but manufacturing precision deteriorates
Solution Approach 1:
The patent applies preliminary action by encoding error correction information into the DNA sequence during the synthesis preparation stage. The encoder generates a coded DNA sequence that includes redundancy to correct synthesis errors, allowing low-cost micro-array synthesis to produce reliable sequences without requiring expensive high-precision synthesis methods
Solution Approach 2:
The patent changes the parameter of error tolerance by transforming the DNA sequence into a coded form with increased error-correcting capability. This parameter change allows the system to tolerate higher synthesis error rates from low-cost methods while maintaining data integrity
3Speed
If hard decoding decision is used to correct errors, then decoding speed is improved, but substitution error correction capability deteriorates
Solution Approach 1:
The patent prepares soft information (probability density functions) in advance during the sequencing process, enabling soft decoding to efficiently utilize this pre-computed information for accurate error correction without excessive computational overhead
Solution Approach 2:
The patent changes the decoding parameter from hard decisions to soft decisions, using probability density functions to represent uncertainty in nucleotide identification. This parameter change enables more accurate substitution error correction while maintaining practical decoding speeds through efficient algorithm implementation
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The proposed solution significantly reduces substitution errors and improves the reliability of DNA-based storage systems by utilizing soft decoding algorithms and non-binary error correction codes, achieving nearly error-free sequencing with low computational complexity.
Implementation Method 1
the principle of nanopore sequencing is based on the detection of changes in an ionic current when a DNA sequence passes through a nanoscale hole. Each nucleobase or nucleotide causes a different amplitude of current drop due to its different atomic structure.
Data Source
AI summary
A method includes obtaining, for each type of nucleotide, a probability density function, the probability density functions being obtained from measurements of current drops produced during at least one passage of at least one sequence of reference nucleotides through a nanopore sequencer; obtaining measurements of current drops produced when the sequence of nucleotides to be decoded passes through the nanopore sequencer; calculating, for each measurement value considered and for each type of nucleotide of the B types of nucleotides, a piece of reliability information based on the probability density function obtained for the type of nucleotide considered; obtaining a decoded value identifying a type of nucleotide from the B types of DNA nucleotides, by applying a soft decoding algorithm with an error correction code to the current drop measurement and to the B pieces of reliability information obtained for the considered measurement value.


