Reusable Nucleic Acid Sequences for DNA Data Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA-based data storage methods face limitations in storage density and durability, with existing technologies requiring large physical space and being unable to efficiently store and retrieve large amounts of data due to the need for de novo synthesis of millions of DNA molecules and the limitations of next-generation sequencing instruments.

Innovation Solution

The use of nucleic acid-based data storage systems where each data storage nucleic acid represents a bit-mer sequence that encodes information and its position within a bit string, allowing for reusable nucleic acid sequences to be used, enabling efficient storage and retrieval by pooling and indexing nucleic acids with primer binding sequences for selective amplification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If de novo synthesis of millions of DNA molecules is used to store data, then data storage capacity is improved, but manufacturing complexity and cost increase significantly

Engineering Contradiction:
Improvedata storage capacityVSAvoidmanufacturing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent uses existing DNA sequences as templates to generate multiple copies through PCR amplification, rather than synthesizing millions of unique DNA molecules from scratch. This copying approach dramatically reduces manufacturing complexity while maintaining data storage capacity

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent employs universal primer binding sites that can amplify multiple different data-containing sequences simultaneously. These universal primers serve multiple functions: they bind to various target sequences, enable pooled amplification, and facilitate data retrieval without requiring separate synthesis processes for each sequence

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If next-generation sequencing instruments are used to read DNA data, then data retrieval is enabled, but reading speed and efficiency are limited by instrument capacity

Engineering Contradiction:
Improvedata retrieval capabilityVSAvoidreading speed
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The patent divides large datasets into multiple addressable segments or blocks, each flanked by unique primer binding sites. This segmentation allows selective amplification and reading of specific data portions without sequencing entire libraries, dramatically improving reading speed and efficiency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces PCR amplification as an intermediary step between DNA storage and sequencing reading. This intermediary process enriches target sequences before sequencing, enabling faster and more efficient data retrieval by focusing instrument capacity on amplified targets rather than random sampling

Inventive Principle:
Principle #24Intermediary (Mediator)

3Duration of action of stationary object

If DNA sequences are synthesized and stored for data storage, then long-term durability is improved, but physical space requirements increase

Engineering Contradiction:
Improvestorage durabilityVSAvoidphysical space
Core Design Contradiction:
Duration of action of stationary objectVSVolume of stationary object

Solution Approach 1:

The patent combines multiple data-containing DNA sequences into a single pooled sample for storage. By merging numerous sequences into one physical storage unit and using universal primers for amplification, the patent dramatically reduces physical space requirements while maintaining long-term durability through the inherent stability of DNA

Inventive Principle:
Principle #5Merging (Combining)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach significantly reduces the number of unique oligonucleotide sequences needed, allowing for the storage of large amounts of data in a compact and durable format, with the potential to store multiple gigabytes using fewer oligonucleotides compared to existing methods, while maintaining data integrity and accessibility.

Implementation Method 1

pooling and indexing nucleic acids with primer binding sequences for selective amplification

Methodology Applied
Scientific EffectPolymerase chain reaction (PCR):

Data Source

PatentUS11308055B2DNA data storage using reusable nucleic acids
Publication Date: 2022.04.19 INTEGRATED DNA TECHNOLOGIES INC
  • US11308055B2 patent drawing
  • US11308055B2 patent drawing
  • US11308055B2 patent drawing

AI summary

Disclosed herein are nucleic acid-based data storage systems and nucleic acid data storage constructs comprising reusable nucleic acid sequences, each representing information carried by a single bit (and, in some embodiments, one or more adjacent bits) within a bit string, and each furthermore representing the position of the single bit within the bit string. Also described are methods for storing data in the nucleic acid-based data storage systems and nucleic acid data storage constructs of the disclosure.