Nucleic Acid Data Storage Using Combinatorial Sequence Pools
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for encoding digital information into nucleic acid molecules rely on costly base-by-base synthesis, making them inefficient for commercial implementation and requiring de novo synthesis of distinct nucleic acid sequences for each information storage request.
Innovation Solution
The method encodes digital information by mapping bit-value information into the presence or absence of unique nucleic acid sequences within a pool, using combinatorial genomic strategies such as assembly of multiple nucleic acid sequences or enzymatic editing, without the need for base-by-base synthesis, and generates unique nucleic acid sequences for reuse in subsequent storage requests.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid molecules, but the cost and time of encoding becomes excessively high
Solution Approach 1:
The patent segments the encoding process into two distinct phases: (1) a one-time comprehensive synthesis phase where all possible nucleic acid sequences corresponding to potential data values are synthesized and stored in a library, and (2) a rapid encoding phase where digital information is encoded by selectively mixing pre-synthesized sequences from the library. This segmentation eliminates the need for de novo synthesis during each encoding operation, dramatically reducing cost and time while maintaining data storage capability.
2Adaptability or versatility
If de novo synthesis is performed for each information storage request, then unique nucleic acid sequences can be generated, but the process becomes time-consuming and expensive
Solution Approach 1:
The patent applies preliminary action by pre-synthesizing and storing a comprehensive library of nucleic acid sequences that correspond to all possible data values before any encoding operations are performed. When encoding is needed, the system simply retrieves and mixes appropriate pre-synthesized sequences from the library rather than performing synthesis de novo, thereby eliminating time loss while maintaining sequence uniqueness through combinatorial mixing of pre-prepared components.
3Manufacturing precision
If base-by-base synthesis is used, then precise nucleic acid sequences can be created, but commercial implementation becomes difficult
Solution Approach 1:
The patent segments manufacturing into a precision phase (one-time library synthesis with full quality control) and a scalable phase (combinatorial mixing of verified sequences). This allows commercial implementation because the expensive, precision-critical synthesis work is performed once to create a master library, while subsequent encoding operations use simple, scalable mixing processes that can be easily manufactured and scaled without compromising sequence accuracy.
4Reliability
If unique nucleic acid sequences are synthesized for each data value, then data can be encoded, but the process lacks parallelization capability
Solution Approach 1:
The patent merges multiple pre-synthesized nucleic acid sequences corresponding to different data values into a single pooled library. During encoding, multiple sequences are combined in parallel through mixing operations rather than being synthesized sequentially. This merging enables parallelization where many data values can be encoded simultaneously by mixing appropriate combinations of pre-synthesized sequences, dramatically increasing encoding speed while maintaining accuracy through the predetermined design of the sequence library.
Data Source
AI summary
Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool. But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).


