Nucleic Acid Data Encoding Using Sequence Presence-Absence Pools

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for encoding digital information into nucleic acids rely on costly base-by-base synthesis, making them inefficient for commercial implementation and error-prone for data retrieval.

Innovation Solution

Encoding digital information by mapping bit-value information to the presence or absence of unique nucleic acid sequences within a pool, using combinatorial genomic strategies such as assembly of multiple nucleic acid sequences or enzymatic editing, to construct identifier libraries that represent digital data without sequential synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored in nucleic acid sequences, but the encoding cost becomes expensive and the process becomes time-consuming

Engineering Contradiction:
Improvedata storage capabilityVSAvoidencoding cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent divides the nucleic acid sequence into multiple segments or pools, where each pool contains a subset of possible sequences. Instead of synthesizing the entire sequence base-by-base, the system segments the encoding process into multiple smaller pools that can be prepared more efficiently and combined later, reducing the overall synthesis burden and cost.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent prepares multiple pools of nucleic acid sequences in advance, where each pool contains pre-synthesized sequences that represent specific bit values. This preliminary preparation allows the actual data encoding to be performed by simply selecting and combining pre-made sequences rather than synthesizing them on-demand, significantly reducing encoding time and cost.

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If base-by-base nucleic acid synthesis is used to encode digital information, then data can be stored accurately, but the encoding time becomes excessively long

Engineering Contradiction:
Improvedata encoding accuracyVSAvoidencoding time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the nucleic acid sequence into multiple pools, allowing parallel preparation of sequence subsets. This segmentation enables simultaneous synthesis and preparation of multiple sequence pools, dramatically reducing the total encoding time while maintaining accuracy through systematic combination of the segmented pools.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary preparation of multiple nucleic acid sequence pools in advance, creating a library of pre-synthesized sequences ready for immediate use. This preliminary action eliminates the need for time-consuming on-demand synthesis during the actual encoding process, reducing encoding time while preserving data accuracy through verified pre-synthesized sequences.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If de novo nucleic acid synthesis is performed for each data encoding operation, then unique sequences can be generated, but the process becomes costly and commercially unviable

Engineering Contradiction:
Improvesequence uniquenessVSAvoidcommercial viability
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent performs preliminary synthesis of comprehensive pools of nucleic acid sequences that can be reused for multiple encoding operations. Instead of performing de novo synthesis for each data encoding, the system prepares pools in advance that contain diverse sequences representing different bit values, making subsequent encoding operations simply matter of selection and combination, thereby achieving commercial viability while maintaining sequence uniqueness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates multiple copies of nucleic acid sequences within pools, where each pool contains numerous copies of sequences representing specific bit values. This copying approach allows the same verified sequences to be reused across multiple encoding operations, eliminating the need for expensive de novo synthesis for each operation while maintaining data integrity through identical verified sequences.

Inventive Principle:
Principle #26Copying

4Ease of operation

If sequential base-by-base synthesis is used, then complete nucleic acid sequences can be constructed, but the complexity of the encoding process increases

Engineering Contradiction:
Improvesequence construction capabilityVSAvoidencoding process complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent divides the complex task of constructing complete nucleic acid sequences into manageable segments organized in multiple pools. Each pool contains sequences of specific segments, and the overall sequence is constructed by combining these pre-organized segments. This segmentation reduces the operational complexity by breaking down the encoding process into simpler, more manageable steps that can be performed and verified independently.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20200250546A1Nucleic acid-based data storage
Publication Date: 2020.08.06 BIOMEMORY AMERICA LLC
  • US20200250546A1 patent drawing
  • US20200250546A1 patent drawing
  • US20200250546A1 patent drawing

AI summary

Methods and systems for encoding digital information in nucleic acid (e.g., deoxyribonucleic acid) molecules without base-by-base synthesis, by encoding bit-value information in the presence or absence of unique nucleic acid sequences within a pool, comprising specifying each bit location in a bit-stream with a unique nucleic sequence and specifying the bit value at that location by the presence or absence of the corresponding unique nucleic acid sequence in the pool. But, more generally, specifying unique bytes in a bytestream by unique subsets of nucleic acid sequences. Also disclosed are methods for generating unique nucleic acid sequences without base-by-base synthesis using combinatorial genomic strategies (e.g., assembly of multiple nucleic acid sequences or enzymatic-based editing of nucleic acid sequences).