Nick-Based DNA Data Storage Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high cost of DNA synthesis and encoding overhead in conventional DNA-based data storage systems, which limits their practical deployment and efficiency, especially for short blocklengths, and the lack of random access and computing capabilities in fountain approaches.

Innovation Solution

The use of nick-based data storage in deoxyribonucleic acid (DNA) sequences, where digital information is encoded in a double-stranded DNA sequence with nickable positions, allowing for cost-effective native DNA usage and enabling random access and computational paradigms through nicking and displacement processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional DNA synthesis methods are used for data storage, then data can be stored in DNA sequences, but the cost is extremely high (5-6 orders of magnitude higher than classical recording media)

Engineering Contradiction:
Improveinformation integrityVSAvoidcost of DNA synthesis
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

Instead of synthesizing DNA to encode data (conventional approach), the patent inverts the approach by using native DNA sequences and encoding data through controlled nicking at specific positions. This inversion eliminates the need for expensive DNA synthesis while maintaining data storage capability.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The patent changes the encoding parameter from DNA sequence composition (which requires synthesis) to DNA strand integrity (nicked vs. non-nicked positions). This parameter change allows using inexpensive native DNA while still enabling data encoding through detectable modifications.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If short DNA strands with blocklengths of 100 base pairs are used, then synthesis cost is reduced, but encoding overhead increases to 30% information loss

Engineering Contradiction:
Improvesynthesis costVSAvoidencoding overhead
Core Design Contradiction:
Ease of manufactureVSLoss of information

Solution Approach 1:

The patent inverts the conventional approach by not synthesizing DNA at all, instead using native DNA and encoding through nicking. This eliminates the synthesis cost advantage of short strands while avoiding the encoding overhead penalty, since the encoding is done through physical modification rather than sequence composition.

Inventive Principle:
Principle #13The other way round (Inversion)

3Reliability

If fountain approaches are used for data storage, then random access and computing capabilities are eliminated, but data can still be stored

Engineering Contradiction:
Improvedata storage capabilityVSAvoidrandom access capability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the DNA sequence into multiple registers, where each register contains multiple copies of a DNA sequence with specific nickable positions. This segmentation enables random access by allowing selective targeting of specific registers or positions through the nicking mechanism, while maintaining data storage capability.

Inventive Principle:
Principle #1Segmentation

4Ease of manufacture

If native DNA is used instead of synthetic DNA, then cost is reduced and abundance is increased, but encoding capability must be maintained

Engineering Contradiction:
Improvecost-effectivenessVSAvoidencoding precision
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent introduces nicking enzymes as intermediaries that precisely modify native DNA sequences at targeted positions. These enzymes act as mediators between the data encoding process and the native DNA, ensuring precise encoding despite the natural variability of native sequences.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach reduces storage costs significantly, allows for random access and computational operations, and achieves high data density by utilizing native DNA, overcoming the limitations of conventional methods.

Implementation Method 1

a polymerase chain reaction (PCR)

Methodology Applied
Scientific EffectDNA polymerase extension: Enzyme

Implementation Method 2

one strand of the DNA sequence is nicked at each nickable position having a mapped value that indicates to nick the DNA sequence

Methodology Applied
Scientific EffectNicking: Enzyme

Implementation Method 3

a displacement strand associates with the toehold, displaces the portion of the first strand bounded by the toehold and the third nicked position, and is ligated to the first strand of the DNA sequence at the first nicked position

Methodology Applied
Scientific EffectLigation: Enzyme

Implementation Method 4

a double-stranded DNA sequence having a plurality of nickable positions

Methodology Applied
Scientific EffectBase pairing: Chemical Bonding

Data Source

PatentUS11538554B1Nick-based data storage in native nucleic acids
Publication Date: 2022.12.27 THE BOARD OF TRUSTEES OF THE UNIV OF ILLINOIS
  • US11538554B1 patent drawing
  • US11538554B1 patent drawing
  • US11538554B1 patent drawing

AI summary

Nick-based methods, devices, and systems for nick-based data storage in a deoxyribonucleic acid (DNA) sequence are disclosed. Digital information is encoded in a register of at least one copy of a double-stranded DNA sequence having a plurality of nickable positions. The data is translated into a sequence of values from a nick alphabet that is subsequently mapped to the plurality of nickable positions, and the DNA sequence is nicked according to the mapped values. Because the digital information is encoded as a series of nicked and non-nicked positions of a double-stranded DNA sequence, the nucleotide sequence of the DNA can be non-synthetic, or “native” DNA.