DNA Data Storage Encoding for Random Access Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
DNA-based data storage faces challenges such as high production costs, low data retrieval speed, errors in encoding, writing, storing, decoding, and reading, and difficulty in achieving random access to large datasets.
Innovation Solution
A novel 5-bit transcoding framework, combined with compression and error correction algorithms, is used to convert data into nucleotide sequences, allowing for efficient and reliable storage and retrieval, including random access to partial data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If DNA-based storage is used for large-scale archival storage, then storage density and long-term storage capability are improved, but data retrieval speed deteriorates due to sequencing requirements
Solution Approach 1:
The patent divides the stored data into multiple independent DNA fragments, each with its own index sequence. This segmentation allows selective amplification and sequencing of only the required fragments rather than the entire dataset, thereby improving retrieval speed while maintaining high storage density.
Solution Approach 2:
The patent incorporates index sequences and barcodes into the DNA structure during the encoding phase. These preliminary markers enable rapid identification and targeted retrieval of specific data portions without requiring complete sequencing, thus resolving the speed-density tradeoff.
2Reliability
If error correction algorithms are implemented in the DNA storage process, then data reliability is improved, but device complexity increases due to additional processing steps
Solution Approach 1:
The patent combines error correction codes with the data encoding process by integrating redundancy information directly into the DNA sequence structure. This merging approach achieves reliable error correction while minimizing additional processing complexity through unified encoding operations.
Solution Approach 2:
The patent employs configurable error correction parameters that can be adjusted based on storage requirements. By changing the level of redundancy and correction strength, the system achieves reliable error handling while allowing flexibility to balance against processing complexity depending on specific application needs.
3Ease of operation
If random access to partial data is implemented in DNA storage, then data accessibility is improved, but production cost increases due to additional indexing and synthesis requirements
Solution Approach 1:
The patent introduces index sequences and barcode markers as intermediary elements between the data and the DNA storage medium. These intermediaries enable efficient random access by allowing targeted retrieval operations without requiring complete sequencing, thereby improving accessibility while managing synthesis costs through selective amplification.
Solution Approach 2:
By segmenting the data into independently addressable DNA fragments with unique indexes, the system enables random access to specific portions without synthesizing or sequencing the entire dataset, thus improving accessibility while controlling production costs through targeted operations.
Data Source
AI summary
The present disclosure generally relates to DNA-based data storage. An exemplary method for storing input data on nucleic acid comprises: converting the input data into a set of nucleotide sequences and synthesizing a set of nucleic acids comprising the set of nucleotide sequences. The converting comprises a data processing step comprising converting the input data into a binary string, and a nucleotide encoding step comprising converting the binary string using a 5-bit transcoding framework to obtain the set of nucleotide sequences.


