DNA Oligo Synchronization Marks for Insertion-Deletion Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA data storage technologies face inefficiencies in error correction, particularly for insertions and deletions, which affect the reliability and efficiency of data retrieval.
Innovation Solution
The implementation of synchronization marks and nested error correction codes to encode and decode data stored in DNA oligos, including oligo formatting, sync mark detection, and cross-correlation analysis to correct symbol alignment and handle insertions and deletions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Reed-Solomon error correction codes are applied to individual oligos, then error correction capability is provided, but storage efficiency decreases due to the relatively short payload capacity of oligos
Solution Approach 1:
The patent divides the error correction approach into two segments: first applying Reed-Solomon codes at the individual oligo level for basic error correction, then applying LDPC codes at the pool level for additional correction capability. This segmentation allows each code to operate optimally at its appropriate scale, improving overall storage efficiency while maintaining strong error correction.
Solution Approach 2:
The patent implements a nested error correction structure where Reed-Solomon codes are embedded within individual oligos, and LDPC codes are applied to the entire oligo pool. The inner Reed-Solomon correction operates first, then the outer LDPC correction provides additional protection. This nested approach maximizes storage efficiency by leveraging the strengths of both coding schemes at different hierarchical levels.
2Reliability
If synchronization marks are inserted at predetermined intervals, then insertions and deletions can be detected and corrected, but the complexity of the encoding and decoding processes increases
Solution Approach 1:
The patent applies preliminary action by inserting synchronization marks into the DNA sequence before the actual data encoding process. These marks serve as pre-established reference points that guide the subsequent decoding process, allowing the system to detect and correct insertions and deletions more efficiently without requiring complex real-time analysis during decoding.
Solution Approach 2:
Synchronization marks function as intermediaries between the encoded DNA data and the decoding process. These marks provide a mediating reference framework that simplifies the alignment and comparison operations during decoding, reducing the computational complexity by providing clear anchor points for detecting structural variations like insertions and deletions.
3Quantity of substance
If the payload capacity of individual oligos is increased, then more data can be stored per oligo, but the effectiveness of error correction codes decreases due to the relatively short length
Solution Approach 1:
The patent transitions from a single-dimension approach (error correction within individual oligos) to a multi-dimensional approach by applying error correction at two hierarchical levels: the oligo level and the pool level. This dimensional expansion allows the system to achieve both higher payload capacity per oligo and maintained error correction effectiveness by distributing correction capabilities across multiple scales.
Data Source
AI summary
Example systems and methods for using synchronization marks to correct insertions and deletions for DNA data storage are described. A data unit may be encoded in oligos that include synchronization marks at predetermined intervals along the length of each oligo. During decoding, the synchronization marks may improve identification and isolation of insertions and deletions for correction of symbol alignment prior to error correction code decoding. In some configurations, correlation analysis may be used to improve isolation of insertions and deletions where multiple copies of the same oligo are available.


