DNA Data Storage Encoding Homopolymer Ambiguity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for information storage using DNA sequences face challenges in accurately encoding and decoding binary data due to errors in synthesis and sequencing, particularly with homopolymer runs, which can lead to ambiguity in interpreting nucleotide sequences as binary bits.

Innovation Solution

The method involves encoding binary information using nucleotides such as adenine (A), cytosine (C), guanine (G), and thymine (T) or uracil (U), with error-prone polymerases to synthesize oligonucleotides, where homopolymer runs are treated as single nucleotides for decoding, and using address sequences and flanking sequences for amplification and sequencing to correct errors and distinguish between adjacent bits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If error-prone polymerases are used to synthesize oligonucleotides, then productivity is improved, but manufacturing precision deteriorates due to synthesis errors

Engineering Contradiction:
Improvesynthesis speedVSAvoidsequence accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent creates multiple copies of each oligonucleotide sequence through error-prone polymerase synthesis. By generating numerous identical copies from the same template, the system enables statistical error correction where the majority consensus sequence represents the intended original sequence, thereby maintaining high productivity while achieving high manufacturing precision through redundancy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent implements a feedback mechanism where synthesized oligonucleotides are sequenced, errors are identified, and the information is used to correct the sequence interpretation. Next-generation sequencing provides feedback on synthesis errors, and address sequences provide feedback on location, enabling accurate reconstruction of the original binary data despite synthesis imperfections

Inventive Principle:
Principle #23Feedback

2Quantity of substance

If homopolymer runs are encoded as multiple identical nucleotides, then information density is improved, but measurement precision deteriorates due to ambiguity in decoding

Engineering Contradiction:
Improveinformation densityVSAvoiddecoding accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by adding address sequences and flanking sequences to each oligonucleotide before synthesis is complete. These additional sequences provide contextual information that enables accurate decoding of homopolymer runs. The address sequence tells the system where the oligonucleotide belongs in the binary stream, and flanking sequences provide boundaries that help interpret the meaning of repeated nucleotides, thereby resolving decoding ambiguity while maintaining high information density

Inventive Principle:
Principle #10Preliminary action

3Reliability

If address sequences are added to each oligonucleotide, then reliability is improved through error correction, but device complexity increases

Engineering Contradiction:
Improvedata integrityVSAvoidsequence structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies universality by designing flanking sequences that serve multiple functions simultaneously. These sequences enable PCR amplification, provide priming sites for sequencing reactions, and act as boundaries for data interpretation. By making these sequences multi-functional, the patent achieves high reliability through error correction and amplification without proportionally increasing complexity, as the same structural elements perform multiple critical roles

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach allows for robust and accurate storage and retrieval of digital information, such as text, images, and sound, by minimizing errors through error-prone polymerase synthesis and next-generation sequencing techniques, ensuring reliable conversion between nucleic acid sequences and binary bit streams.

Implementation Method 1

error-prone polymerases to synthesize oligonucleotides

Methodology Applied
Scientific EffectEnzymatic synthesis: Enzyme

Implementation Method 2

homopolymer runs are treated as single nucleotides for decoding

Methodology Applied
Scientific EffectHomopolymer formation:

Implementation Method 3

next-generation sequencing techniques

Methodology Applied
Scientific EffectSequencing detection:

Data Source

PatentUS11532380B2Methods for using nucleic acids to store, retrieve and access information comprising a text, image, video or audio format
Publication Date: 2022.12.20 PRESIDENT & FELLOWS OF HARVARD COLLEGE
  • US11532380B2 patent drawing
  • US11532380B2 patent drawing
  • US11532380B2 patent drawing

AI summary

A method of storing information using monomers such as nucleotides is provided including converting a format of information into a plurality of bit sequences of a bit stream with each having a corresponding bit barcode, converting the plurality of bit sequences to a plurality of corresponding oligonucleotide sequences using one bit per base encoding, synthesizing the plurality of corresponding oligonucleotide sequences on a substrate having a plurality of reaction locations, and storing the synthesized plurality of corresponding oligonucleotide sequences.