DNA Data Storage Codecs for Insertion and Deletion Error Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing DNA data storage methods fail to efficiently encode and decode digital data in the presence of errors such as deletions, insertions, and mutations, and do not provide a structured way to store and retrieve files.
Innovation Solution
The implementation of codecs that include inner and outer error correction schemes, shuffling, and redundancy mechanisms to encode and decode digital data in polynucleotide sequences, along with indexing and hashing for efficient storage and retrieval.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If DNA is used as a data storage medium, then data density and longevity are improved, but errors and ambiguities occur during sequencing and sequencing-related operations
Solution Approach 1:
The patent applies preliminary action by encoding error correction codes and redundancy information into the DNA sequences before synthesis. The codec system pre-processes the digital data by adding correction capabilities that will activate during decoding, allowing errors to be corrected without re-sequencing. This includes embedding checksums, parity bits, and erasure correction codes into the oligonucleotide sequences prior to storage.
Solution Approach 2:
The patent implements feedback mechanisms through iterative decoding algorithms that use sequencing results to identify and correct errors. The system analyzes sequencing data, detects deviations from expected sequences, and applies correction algorithms that feed back into the decoding process. This allows the system to continuously improve data accuracy by using the sequencing output itself to guide error correction.
2Device complexity
If traditional encoding methods are used for DNA storage, then the encoding process is simple, but the system cannot sustain high deletion, mutation and insertion rates
Solution Approach 1:
The patent applies segmentation by dividing the DNA storage system into distinct functional layers: an inner codec that handles basic encoding and decoding at the sequence level, and an outer codec that manages higher-level error correction and data reconstruction. This multi-layered approach segments the complex error correction task into manageable components, each specializing in specific types of errors, thereby improving reliability without overwhelming complexity.
Solution Approach 2:
The patent uses composite materials by combining multiple error correction strategies within a unified codec system. The system integrates different coding schemes (such as Reed-Solomon codes, convolutional codes, and LDPC codes) into a composite encoding framework that leverages the strengths of each individual method. This composite approach creates a robust error correction system that can handle diverse error types including deletions, insertions, and mutations simultaneously.
3Loss of information
If redundancy is added to correct for erasures and errors, then data integrity is improved, but synthesis cycles increase
Solution Approach 1:
The patent applies parameter changes by dynamically adjusting the level of redundancy based on the specific requirements of the data being stored and the expected error rates. The codec system can modify encoding parameters such as code rate, block size, and redundancy level to optimize the balance between data integrity and synthesis efficiency. This allows the system to use minimal redundancy for stable sequences while applying higher redundancy only where needed, reducing overall synthesis time.
Data Source
AI summary
Described herein are systems and methods for encoding digital data into oligonucleotides and decoding the oligonucleotides back into digital data. The encoding and decoding schemes include an inner codec for transforming the digital data into bases, and vice versa. The encoding and decoding schemes also include an outer codec comprising an error correction scheme.


