Reusable Nucleic Acid Sequences for DNA Data Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA-based data storage methods face limitations in storage density and durability, with existing technologies requiring large physical space and being unable to efficiently store and retrieve large amounts of data due to the need for de novo synthesis of millions of DNA molecules and the limitations of next-generation sequencing instruments.
Innovation Solution
The use of nucleic acid-based data storage systems where each data storage nucleic acid represents a bit-mer sequence that encodes information and its position within a bit string, allowing for reusable nucleic acid sequences to be used, enabling efficient storage and retrieval by pooling and indexing nucleic acids with primer binding sequences for selective amplification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If de novo synthesis of millions of DNA molecules is used to store data, then data storage capacity is improved, but manufacturing complexity and cost increase significantly
Solution Approach 1:
The patent uses existing DNA sequences as templates to generate multiple copies through PCR amplification, rather than synthesizing millions of unique DNA molecules from scratch. This copying approach dramatically reduces manufacturing complexity while maintaining data storage capacity
Solution Approach 2:
The patent employs universal primer binding sites that can amplify multiple different data-containing sequences simultaneously. These universal primers serve multiple functions: they bind to various target sequences, enable pooled amplification, and facilitate data retrieval without requiring separate synthesis processes for each sequence
2Ease of operation
If next-generation sequencing instruments are used to read DNA data, then data retrieval is enabled, but reading speed and efficiency are limited by instrument capacity
Solution Approach 1:
The patent divides large datasets into multiple addressable segments or blocks, each flanked by unique primer binding sites. This segmentation allows selective amplification and reading of specific data portions without sequencing entire libraries, dramatically improving reading speed and efficiency
Solution Approach 2:
The patent introduces PCR amplification as an intermediary step between DNA storage and sequencing reading. This intermediary process enriches target sequences before sequencing, enabling faster and more efficient data retrieval by focusing instrument capacity on amplified targets rather than random sampling
3Duration of action of stationary object
If DNA sequences are synthesized and stored for data storage, then long-term durability is improved, but physical space requirements increase
Solution Approach 1:
The patent combines multiple data-containing DNA sequences into a single pooled sample for storage. By merging numerous sequences into one physical storage unit and using universal primers for amplification, the patent dramatically reduces physical space requirements while maintaining long-term durability through the inherent stability of DNA
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly reduces the number of unique oligonucleotide sequences needed, allowing for the storage of large amounts of data in a compact and durable format, with the potential to store multiple gigabytes using fewer oligonucleotides compared to existing methods, while maintaining data integrity and accessibility.
Implementation Method 1
pooling and indexing nucleic acids with primer binding sequences for selective amplification
Data Source
AI summary
Disclosed herein are nucleic acid-based data storage systems and nucleic acid data storage constructs comprising reusable nucleic acid sequences, each representing information carried by a single bit (and, in some embodiments, one or more adjacent bits) within a bit string, and each furthermore representing the position of the single bit within the bit string. Also described are methods for storing data in the nucleic acid-based data storage systems and nucleic acid data storage constructs of the disclosure.


