Nucleic Acid Data Storage Using Combinatorial Sequence Libraries

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding digital information into nucleic acid molecules rely on costly base-by-base synthesis, making them inefficient and expensive for storing large volumes of data that need to be archived for long periods.

Innovation Solution

The method involves encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, using combinatorial genomic strategies to generate unique nucleic acid sequences without base-to-base synthesis, and mapping digital information into a sequence of symbols that are encoded using codebooks and identifiers, allowing for error protection and efficient storage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid molecules, but the cost and complexity of encoding increases significantly

Engineering Contradiction:
Improvedata storage reliabilityVSAvoidencoding complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the encoding process into two independent stages: (1) synthesizing a library of unique nucleic acid sequences once, and (2) encoding data by selectively combining pre-synthesized sequences through combinatorial assembly. This segmentation eliminates the need for costly base-by-base synthesis for every encoding operation, reducing both complexity and cost while maintaining data storage reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary synthesis of a comprehensive library of unique nucleic acid sequences before actual data encoding. These pre-synthesized sequences are stored and can be rapidly assembled through combinatorial methods when encoding data, avoiding the need to perform synthesis during the encoding process itself and significantly reducing operational complexity.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If de novo synthesis of nucleic acid sequences is performed for each storage request, then unique sequences can be generated, but the cost becomes prohibitively expensive for large data volumes

Engineering Contradiction:
Improvesequence uniquenessVSAvoidmanufacturing cost
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent creates a master library of unique nucleic acid sequences that can be copied and reused multiple times for different encoding operations. Instead of synthesizing new sequences for each storage request, the system copies and recombines existing sequences from the library, dramatically reducing per-operation costs while maintaining sequence uniqueness through combinatorial assembly.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent enables recovery and reuse of the same unique nucleic acid sequence library across multiple encoding operations. The pre-synthesized sequences are not consumed or degraded through use, allowing the same library to be repeatedly accessed and recombined for different data storage requests, amortizing the synthesis cost over many operations.

Inventive Principle:
Principle #34Discarding and recovering

3Manufacturing precision

If base-by-base synthesis methods are used, then precise nucleic acid sequences can be created, but the time required for encoding increases

Engineering Contradiction:
Improvesequence precisionVSAvoidencoding speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent segments the time-consuming synthesis step from the rapid encoding step. Synthesis of unique sequences is performed once and stored, while encoding operations only require fast combinatorial assembly of pre-synthesized fragments. This segmentation maintains sequence precision while dramatically increasing encoding speed for large data volumes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs the slow synthesis operation in advance, creating a library of precise unique sequences before encoding begins. Subsequent encoding operations only require rapid selection and assembly of these pre-synthesized sequences, maintaining precision while achieving high encoding throughput for large-scale data storage.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11763169B2Systems for nucleic acid-based data storage
Publication Date: 2023.09.19 BIOMEMORY AMERICA LLC
  • US11763169B2 patent drawing
  • US11763169B2 patent drawing
  • US11763169B2 patent drawing

AI summary

Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool, but, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).