DNA Data Storage Using Combinatorial Sequence Libraries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding digital information into nucleic acid sequences rely on costly base-by-base synthesis, making them inefficient and expensive for storing and retrieving data, especially for long-term archiving.

Innovation Solution

The method involves translating information into a string of symbols, mapping these symbols to unique nucleic acid sequences using combinatorial genomic strategies, such as assembly of multiple sequences or enzymatic editing, to encode bit-value information in the presence or absence of specific nucleic acid sequences, allowing for the construction of an identifier library that represents the data without the need for base-by-base synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid sequences, but the cost and complexity of encoding and retrieving data becomes excessively high

Engineering Contradiction:
Improvedata storage reliabilityVSAvoidencoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the encoding process into two distinct parts: (1) pre-synthesis of a comprehensive library of unique nucleic acid sequences representing all possible data values, and (2) selective assembly of these pre-synthesized sequences to encode specific digital information. This segmentation eliminates the need for costly base-by-base synthesis during data encoding operations, reducing both complexity and cost while maintaining data storage reliability

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by pre-synthesizing and storing a complete library of unique nucleic acid sequences before actual data encoding occurs. This advance preparation allows subsequent encoding operations to simply select and assemble from the pre-prepared library rather than synthesizing sequences on-demand, dramatically reducing the complexity and cost of data encoding and retrieval operations

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If base-by-base nucleic acid synthesis is used for data encoding, then digital information can be retrieved as bit-streams, but the retrieval cost becomes prohibitively expensive

Engineering Contradiction:
Improvedata retrieval accuracyVSAvoiddata retrieval cost
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent uses copying by creating multiple copies of the pre-synthesized unique nucleic acid sequences from the library during the encoding process. Instead of synthesizing sequences from scratch during retrieval, the system copies appropriate sequences from the pre-prepared library and assembles them to represent the digital data, significantly reducing retrieval cost while maintaining data accuracy through the use of verified pre-synthesized sequences

Inventive Principle:
Principle #26Copying

3Productivity

If unique nucleic acid sequences are pre-synthesized and assembled combinatorially, then encoding cost is reduced, but the initial library construction becomes more complex

Engineering Contradiction:
Improveencoding efficiencyVSAvoidlibrary construction complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing the complex library construction process once in advance, creating a comprehensive catalog of unique nucleic acid sequences that can be reused for multiple encoding operations. This upfront investment in library construction simplifies subsequent encoding operations, improving productivity while concentrating the complexity burden in a single, manageable initial step rather than in every encoding operation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10650312B2Nucleic acid-based data storage
Publication Date: 2020.05.12 BIOMEMORY AMERICA LLC
  • US10650312B2 patent drawing
  • US10650312B2 patent drawing
  • US10650312B2 patent drawing

AI summary

Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).