Flexible DNA Data Storage Decoding Using Redundancy Codes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
DNA data storage systems face challenges in error handling due to DNA synthesis and sequencing errors, which affect the accuracy of data retrieval from nucleotide strand sequences.
Innovation Solution
A method and system that implement flexible decoding technologies using redundancy codes, combining solitary-strand-based and cluster-based approaches to reconstruct nucleotide symbol strings, thereby improving accuracy and reducing the need for redundancy during encoding and decoding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional error correction methods are used in DNA data storage, then data integrity can be maintained, but the cost and complexity of encoding and decoding processes increase
Solution Approach 1:
The patent segments the error correction process into two distinct passes: a first pass that processes solitary nucleotide strand sequence reads individually, and a second pass that processes clusters of reads together. This segmentation allows the system to apply different decoding strategies to different types of data, reducing overall complexity while maintaining reliability.
Solution Approach 2:
The patent implements a dynamic, two-stage decoding approach where the system first attempts to place solitary reads in an ordered map, and only when that fails does it proceed to cluster-based trace reconstruction. This dynamic approach adapts the level of processing to the actual data conditions, avoiding unnecessary computational complexity while ensuring data integrity.
2Measurement precision
If redundancy codes are used to handle sequencing errors, then accuracy of data retrieval improves, but the amount of data that must be stored and processed increases
Solution Approach 1:
The patent applies partial action by using redundancy codes only when necessary - specifically, when solitary reads cannot be placed in the ordered map, the system then applies cluster-based trace reconstruction. This partial application of error correction mechanisms reduces the overall data volume that must be processed while maintaining retrieval accuracy.
Solution Approach 2:
The patent applies local quality by differentiating between solitary reads and clustered reads, applying different levels of error correction to each. Solitary reads receive basic integrity verification, while clustered reads receive more intensive trace reconstruction. This localized approach optimizes accuracy without unnecessarily increasing data volume across the entire system.
3Measurement precision
If cluster-based trace reconstruction is performed on all nucleotide reads, then decoding accuracy improves, but processing time and computational resources increase
Solution Approach 1:
The patent segments the read processing into two distinct passes: a first pass handling solitary reads with basic integrity verification, and a second pass handling clusters with trace reconstruction. This segmentation significantly reduces processing time by applying intensive cluster-based methods only to the subset of reads that require it, rather than all reads.
Solution Approach 2:
The patent performs preliminary action by placing solitary reads in the ordered map during the first pass before attempting cluster-based reconstruction. This preliminary placement reduces the number of reads that require time-intensive cluster processing, thereby reducing overall processing time while maintaining decoding accuracy.
Data Source
AI summary
Data that has been stored according to a DNA data storage method can be decoded using a flexible approach that supports both solitary strand mapping and cluster-based trace reconstruction. Solitary strand mapping can place strings based on integrity verification. Redundancy information can be partitioned to support error correction during the solitary strand mapping while still achieving integrity verification. Clusters with verified strands can be skipped during cluster-based trace reconstruction. Useful for increasing the accuracy of the trace reconstruction procedure.


