DNA Data Storage Using Combinatorial Sequence Libraries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for encoding digital information into nucleic acid sequences rely on costly base-by-base synthesis, making them inefficient and expensive for storing and retrieving data, especially for long-term archiving.
Innovation Solution
The method involves translating information into a string of symbols, mapping these symbols to unique nucleic acid sequences using combinatorial genomic strategies, such as assembly of multiple sequences or enzymatic editing, to encode bit-value information in the presence or absence of specific nucleic acid sequences, allowing for the construction of an identifier library that represents the data without the need for base-by-base synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid sequences, but the cost and complexity of encoding and retrieving data becomes excessively high
Solution Approach 1:
The patent segments the encoding process into two distinct parts: (1) pre-synthesis of a comprehensive library of unique nucleic acid sequences representing all possible data values, and (2) selective assembly of these pre-synthesized sequences to encode specific digital information. This segmentation eliminates the need for costly base-by-base synthesis during data encoding operations, reducing both complexity and cost while maintaining data storage reliability
Solution Approach 2:
The patent performs preliminary action by pre-synthesizing and storing a complete library of unique nucleic acid sequences before actual data encoding occurs. This advance preparation allows subsequent encoding operations to simply select and assemble from the pre-prepared library rather than synthesizing sequences on-demand, dramatically reducing the complexity and cost of data encoding and retrieval operations
2Loss of information
If base-by-base nucleic acid synthesis is used for data encoding, then digital information can be retrieved as bit-streams, but the retrieval cost becomes prohibitively expensive
Solution Approach 1:
The patent uses copying by creating multiple copies of the pre-synthesized unique nucleic acid sequences from the library during the encoding process. Instead of synthesizing sequences from scratch during retrieval, the system copies appropriate sequences from the pre-prepared library and assembles them to represent the digital data, significantly reducing retrieval cost while maintaining data accuracy through the use of verified pre-synthesized sequences
3Productivity
If unique nucleic acid sequences are pre-synthesized and assembled combinatorially, then encoding cost is reduced, but the initial library construction becomes more complex
Solution Approach 1:
The patent applies preliminary action by performing the complex library construction process once in advance, creating a comprehensive catalog of unique nucleic acid sequences that can be reused for multiple encoding operations. This upfront investment in library construction simplifies subsequent encoding operations, improving productivity while concentrating the complexity burden in a single, manageable initial step rather than in every encoding operation
Data Source
AI summary
Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).


