DNA Data Storage Encoding and Error Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital storage solutions face challenges in achieving high density and long-term storage of information, as traditional methods like hard drives and magnetic tapes have limited lifespans and are not efficient in storing large volumes of data, while DNA offers potential advantages but requires innovative encoding and decoding techniques to overcome sequencing and synthesis errors.
Innovation Solution
The method involves converting text or images into megabits, encoding them into oligonucleotides using one bit per base encoding, and synthesizing multiple copies to correct errors, with flanking sequences for amplification and sequencing, utilizing next-generation sequencing and synthesis technologies to store and retrieve information efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Duration of action of stationary object
If traditional digital storage methods (hard drives, magnetic tapes) are used, then current storage needs are met, but storage lifespan is limited to 5-30 years and density is insufficient
Solution Approach 1:
The patent replaces traditional mechanical storage systems (hard drives, magnetic tapes) with a chemical/biological storage system using DNA molecules. Information is encoded into nucleic acid sequences through chemical synthesis, enabling storage densities of up to 1 petabyte per gram and theoretical lifespans exceeding 1000 years under proper conditions, thus resolving the contradiction between lifespan and density.
2Quantity of substance
If DNA is used for information storage, then high density and long-term storage are achieved, but sequencing and synthesis errors occur
Solution Approach 1:
The patent employs multiple redundant copies of each DNA strand containing the same information. By synthesizing and storing numerous identical copies, the system can tolerate synthesis and sequencing errors through majority voting during data retrieval, thereby maintaining high reliability while achieving dense storage.
Solution Approach 2:
The patent incorporates error detection and correction codes into the DNA sequence design. During sequencing, the system compares multiple reads against the encoded error correction scheme, identifying and correcting errors automatically, thus resolving the reliability issue while maintaining high storage density.
3Ease of manufacture
If one bit per base encoding is used, then simple encoding is achieved, but sequence features difficult to read or write (extreme GC content, repeats, secondary structure) are created
Solution Approach 1:
The patent applies different encoding strategies to different regions of the DNA sequence. Rather than uniform encoding, the system locally optimizes sequences to avoid problematic features like extreme GC content, homopolymers, and secondary structures in specific regions, while maintaining the overall one-bit-per-base encoding scheme for simplicity.
Solution Approach 2:
The patent dynamically adjusts sequence parameters during encoding to avoid problematic regions. The encoding algorithm monitors and modifies local sequence composition (GC content, repeat lengths, secondary structure propensity) while maintaining the fundamental bit-to-base mapping, thus balancing encoding simplicity with sequencing ease.
Data Source
AI summary
The present invention relates to methods of synthesizing de novo a polynucleotide storing data using a programmable template independent polymerase.


