Nucleic Acid Data Storage Using Combinatorial Sequence Pools

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding digital information into nucleic acid molecules rely on costly base-by-base synthesis, making them inefficient for commercial implementation and requiring de novo synthesis of distinct nucleic acid sequences for each information storage request.

Innovation Solution

The method encodes digital information by mapping bit-value information into the presence or absence of unique nucleic acid sequences within a pool, using combinatorial genomic strategies such as assembly of multiple nucleic acid sequences or enzymatic editing, without the need for base-by-base synthesis, and generates unique nucleic acid sequences for reuse in subsequent storage requests.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid molecules, but the cost and time of encoding becomes excessively high

Engineering Contradiction:
Improvedata storage capabilityVSAvoidencoding cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent segments the encoding process into two distinct phases: (1) a one-time comprehensive synthesis phase where all possible nucleic acid sequences corresponding to potential data values are synthesized and stored in a library, and (2) a rapid encoding phase where digital information is encoded by selectively mixing pre-synthesized sequences from the library. This segmentation eliminates the need for de novo synthesis during each encoding operation, dramatically reducing cost and time while maintaining data storage capability.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If de novo synthesis is performed for each information storage request, then unique nucleic acid sequences can be generated, but the process becomes time-consuming and expensive

Engineering Contradiction:
Improvesequence uniquenessVSAvoidencoding time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-synthesizing and storing a comprehensive library of nucleic acid sequences that correspond to all possible data values before any encoding operations are performed. When encoding is needed, the system simply retrieves and mixes appropriate pre-synthesized sequences from the library rather than performing synthesis de novo, thereby eliminating time loss while maintaining sequence uniqueness through combinatorial mixing of pre-prepared components.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If base-by-base synthesis is used, then precise nucleic acid sequences can be created, but commercial implementation becomes difficult

Engineering Contradiction:
Improvesequence accuracyVSAvoidcommercial viability
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent segments manufacturing into a precision phase (one-time library synthesis with full quality control) and a scalable phase (combinatorial mixing of verified sequences). This allows commercial implementation because the expensive, precision-critical synthesis work is performed once to create a master library, while subsequent encoding operations use simple, scalable mixing processes that can be easily manufactured and scaled without compromising sequence accuracy.

Inventive Principle:
Principle #1Segmentation

4Reliability

If unique nucleic acid sequences are synthesized for each data value, then data can be encoded, but the process lacks parallelization capability

Engineering Contradiction:
Improvedata encoding accuracyVSAvoidencoding speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent merges multiple pre-synthesized nucleic acid sequences corresponding to different data values into a single pooled library. During encoding, multiple sequences are combined in parallel through mixing operations rather than being synthesized sequentially. This merging enables parallelization where many data values can be encoded simultaneously by mixing appropriate combinations of pre-synthesized sequences, dramatically increasing encoding speed while maintaining accuracy through the predetermined design of the sequence library.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12001962B2Systems for nucleic acid-based data storage
Publication Date: 2024.06.04 BIOMEMORY AMERICA LLC
  • US12001962B2 patent drawing
  • US12001962B2 patent drawing
  • US12001962B2 patent drawing

AI summary

Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool. But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).