Nucleic Acid Data Storage Using Presence-Absence Sequence Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding digital information into nucleic acids rely on costly base-by-base synthesis, making them inefficient for commercial implementation and error-prone for accessing and retrieving data stored in nucleic acid molecules.

Innovation Solution

The method involves encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, using combinatorial genomic strategies such as assembly of multiple sequences or enzymatic editing, to construct identifier libraries that represent digital information without base-by-base synthesis, allowing for easier and less costly data encoding and retrieval.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid molecules, but the cost becomes expensive and the process becomes complex

Engineering Contradiction:
Improvedata storage reliabilityVSAvoidencoding process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the encoding process into two independent stages: (1) synthesizing a library of unique nucleic acid sequences representing all possible digital values, and (2) selecting and combining specific sequences from the library to encode target data. This segmentation eliminates the need for costly base-by-base synthesis of each data-specific sequence, reducing both cost and complexity while maintaining data storage reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary synthesis of a comprehensive library of unique nucleic acid sequences before actual data encoding. By pre-synthesizing all possible sequence representations of digital values and storing them in a library, the system eliminates the need for expensive de novo synthesis during data encoding operations, significantly reducing operational complexity and cost.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If base-by-base synthesis is used to create nucleic acid sequences for data storage, then data can be encoded, but the cost of de novo synthesis becomes expensive

Engineering Contradiction:
Improvedata encoding accuracyVSAvoiddata encoding cost
Core Design Contradiction:
Loss of informationVSEase of manufacture

Solution Approach 1:

The patent creates a master library of unique nucleic acid sequences that serve as templates for data encoding. Instead of performing expensive base-by-base synthesis for each encoding operation, the system copies and combines pre-synthesized sequences from the library to represent digital data, dramatically reducing manufacturing costs while maintaining encoding accuracy through faithful replication of established sequences.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary synthesis of a comprehensive library of unique nucleic acid sequences before actual data encoding. By pre-synthesizing all possible sequence representations of digital values and storing them in a library, the system eliminates the need for expensive de novo synthesis during data encoding operations, significantly reducing operational complexity and cost.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If sequencing is used to access digital data stored in nucleic acid molecules, then data can be retrieved, but the process becomes error prone and costly

Engineering Contradiction:
Improvedata retrieval accuracyVSAvoiddata access process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the data storage structure into unique, identifiable nucleic acid sequences that can be independently detected and counted. By designing the encoding scheme around discrete, sequence-specific markers rather than relying on accurate base-by-base sequencing, the system enables simpler and more reliable data retrieval through targeted detection methods that are less error-prone.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11379729B2Nucleic acid-based data storage
Publication Date: 2022.07.05 BIOMEMORY AMERICA LLC
  • US11379729B2 patent drawing
  • US11379729B2 patent drawing
  • US11379729B2 patent drawing

AI summary

Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).