Nick-Based DNA Data Storage Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high cost of DNA synthesis and encoding overhead in conventional DNA-based data storage systems, which limits their practical deployment and efficiency, especially for short blocklengths, and the lack of random access and computing capabilities in fountain approaches.
Innovation Solution
The use of nick-based data storage in deoxyribonucleic acid (DNA) sequences, where digital information is encoded in a double-stranded DNA sequence with nickable positions, allowing for cost-effective native DNA usage and enabling random access and computational paradigms through nicking and displacement processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional DNA synthesis methods are used for data storage, then data can be stored in DNA sequences, but the cost is extremely high (5-6 orders of magnitude higher than classical recording media)
Solution Approach 1:
Instead of synthesizing DNA to encode data (conventional approach), the patent inverts the approach by using native DNA sequences and encoding data through controlled nicking at specific positions. This inversion eliminates the need for expensive DNA synthesis while maintaining data storage capability.
Solution Approach 2:
The patent changes the encoding parameter from DNA sequence composition (which requires synthesis) to DNA strand integrity (nicked vs. non-nicked positions). This parameter change allows using inexpensive native DNA while still enabling data encoding through detectable modifications.
2Ease of manufacture
If short DNA strands with blocklengths of 100 base pairs are used, then synthesis cost is reduced, but encoding overhead increases to 30% information loss
Solution Approach 1:
The patent inverts the conventional approach by not synthesizing DNA at all, instead using native DNA and encoding through nicking. This eliminates the synthesis cost advantage of short strands while avoiding the encoding overhead penalty, since the encoding is done through physical modification rather than sequence composition.
3Reliability
If fountain approaches are used for data storage, then random access and computing capabilities are eliminated, but data can still be stored
Solution Approach 1:
The patent segments the DNA sequence into multiple registers, where each register contains multiple copies of a DNA sequence with specific nickable positions. This segmentation enables random access by allowing selective targeting of specific registers or positions through the nicking mechanism, while maintaining data storage capability.
4Ease of manufacture
If native DNA is used instead of synthetic DNA, then cost is reduced and abundance is increased, but encoding capability must be maintained
Solution Approach 1:
The patent introduces nicking enzymes as intermediaries that precisely modify native DNA sequences at targeted positions. These enzymes act as mediators between the data encoding process and the native DNA, ensuring precise encoding despite the natural variability of native sequences.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach reduces storage costs significantly, allows for random access and computational operations, and achieves high data density by utilizing native DNA, overcoming the limitations of conventional methods.
Implementation Method 1
a polymerase chain reaction (PCR)
Implementation Method 2
one strand of the DNA sequence is nicked at each nickable position having a mapped value that indicates to nick the DNA sequence
Implementation Method 3
a displacement strand associates with the toehold, displaces the portion of the first strand bounded by the toehold and the third nicked position, and is ligated to the first strand of the DNA sequence at the first nicked position
Implementation Method 4
a double-stranded DNA sequence having a plurality of nickable positions
Data Source
AI summary
Nick-based methods, devices, and systems for nick-based data storage in a deoxyribonucleic acid (DNA) sequence are disclosed. Digital information is encoded in a register of at least one copy of a double-stranded DNA sequence having a plurality of nickable positions. The data is translated into a sequence of values from a nick alphabet that is subsequently mapped to the plurality of nickable positions, and the DNA sequence is nicked according to the mapped values. Because the digital information is encoded as a series of nicked and non-nicked positions of a double-stranded DNA sequence, the nucleotide sequence of the DNA can be non-synthetic, or “native” DNA.


