DNA Data Storage Encoding Homopolymer Ambiguity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for information storage using DNA sequences face challenges in accurately encoding and decoding binary data due to errors in synthesis and sequencing, particularly with homopolymer runs, which can lead to ambiguity in interpreting nucleotide sequences as binary bits.
Innovation Solution
The method involves encoding binary information using nucleotides such as adenine (A), cytosine (C), guanine (G), and thymine (T) or uracil (U), with error-prone polymerases to synthesize oligonucleotides, where homopolymer runs are treated as single nucleotides for decoding, and using address sequences and flanking sequences for amplification and sequencing to correct errors and distinguish between adjacent bits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If error-prone polymerases are used to synthesize oligonucleotides, then productivity is improved, but manufacturing precision deteriorates due to synthesis errors
Solution Approach 1:
The patent creates multiple copies of each oligonucleotide sequence through error-prone polymerase synthesis. By generating numerous identical copies from the same template, the system enables statistical error correction where the majority consensus sequence represents the intended original sequence, thereby maintaining high productivity while achieving high manufacturing precision through redundancy
Solution Approach 2:
The patent implements a feedback mechanism where synthesized oligonucleotides are sequenced, errors are identified, and the information is used to correct the sequence interpretation. Next-generation sequencing provides feedback on synthesis errors, and address sequences provide feedback on location, enabling accurate reconstruction of the original binary data despite synthesis imperfections
2Quantity of substance
If homopolymer runs are encoded as multiple identical nucleotides, then information density is improved, but measurement precision deteriorates due to ambiguity in decoding
Solution Approach 1:
The patent applies preliminary action by adding address sequences and flanking sequences to each oligonucleotide before synthesis is complete. These additional sequences provide contextual information that enables accurate decoding of homopolymer runs. The address sequence tells the system where the oligonucleotide belongs in the binary stream, and flanking sequences provide boundaries that help interpret the meaning of repeated nucleotides, thereby resolving decoding ambiguity while maintaining high information density
3Reliability
If address sequences are added to each oligonucleotide, then reliability is improved through error correction, but device complexity increases
Solution Approach 1:
The patent applies universality by designing flanking sequences that serve multiple functions simultaneously. These sequences enable PCR amplification, provide priming sites for sequencing reactions, and act as boundaries for data interpretation. By making these sequences multi-functional, the patent achieves high reliability through error correction and amplification without proportionally increasing complexity, as the same structural elements perform multiple critical roles
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows for robust and accurate storage and retrieval of digital information, such as text, images, and sound, by minimizing errors through error-prone polymerase synthesis and next-generation sequencing techniques, ensuring reliable conversion between nucleic acid sequences and binary bit streams.
Implementation Method 1
error-prone polymerases to synthesize oligonucleotides
Implementation Method 2
homopolymer runs are treated as single nucleotides for decoding
Implementation Method 3
next-generation sequencing techniques
Data Source
AI summary
A method of storing information using monomers such as nucleotides is provided including converting a format of information into a plurality of bit sequences of a bit stream with each having a corresponding bit barcode, converting the plurality of bit sequences to a plurality of corresponding oligonucleotide sequences using one bit per base encoding, synthesizing the plurality of corresponding oligonucleotide sequences on a substrate having a plurality of reaction locations, and storing the synthesized plurality of corresponding oligonucleotide sequences.


