DNA Data Storage Encoding and Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital storage solutions face challenges in achieving high density and long-term storage of information, as traditional methods like hard drives and magnetic tapes have limited lifespans and are not efficient in storing large volumes of data, while DNA offers potential advantages but requires innovative encoding and decoding techniques to overcome sequencing and synthesis errors.

Innovation Solution

The method involves converting text or images into megabits, encoding them into oligonucleotides using one bit per base encoding, and synthesizing multiple copies to correct errors, with flanking sequences for amplification and sequencing, utilizing next-generation sequencing and synthesis technologies to store and retrieve information efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Duration of action of stationary object

If traditional digital storage methods (hard drives, magnetic tapes) are used, then current storage needs are met, but storage lifespan is limited to 5-30 years and density is insufficient

Engineering Contradiction:
Improvestorage lifespanVSAvoidstorage density
Core Design Contradiction:
Duration of action of stationary objectVSQuantity of substance

Solution Approach 1:

The patent replaces traditional mechanical storage systems (hard drives, magnetic tapes) with a chemical/biological storage system using DNA molecules. Information is encoded into nucleic acid sequences through chemical synthesis, enabling storage densities of up to 1 petabyte per gram and theoretical lifespans exceeding 1000 years under proper conditions, thus resolving the contradiction between lifespan and density.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If DNA is used for information storage, then high density and long-term storage are achieved, but sequencing and synthesis errors occur

Engineering Contradiction:
Improvestorage densityVSAvoiderror rate
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent employs multiple redundant copies of each DNA strand containing the same information. By synthesizing and storing numerous identical copies, the system can tolerate synthesis and sequencing errors through majority voting during data retrieval, thereby maintaining high reliability while achieving dense storage.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent incorporates error detection and correction codes into the DNA sequence design. During sequencing, the system compares multiple reads against the encoded error correction scheme, identifying and correcting errors automatically, thus resolving the reliability issue while maintaining high storage density.

Inventive Principle:
Principle #23Feedback

3Ease of manufacture

If one bit per base encoding is used, then simple encoding is achieved, but sequence features difficult to read or write (extreme GC content, repeats, secondary structure) are created

Engineering Contradiction:
Improveencoding simplicityVSAvoidsequencing difficulty
Core Design Contradiction:
Ease of manufactureVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies different encoding strategies to different regions of the DNA sequence. Rather than uniform encoding, the system locally optimizes sequences to avoid problematic features like extreme GC content, homopolymers, and secondary structures in specific regions, while maintaining the overall one-bit-per-base encoding scheme for simplicity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts sequence parameters during encoding to avoid problematic regions. The encoding algorithm monitors and modifies local sequence composition (GC content, repeat lengths, secondary structure propensity) while maintaining the fundamental bit-to-base mapping, thus balancing encoding simplicity with sequencing ease.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12067434B2Methods of storing information using nucleic acids
Publication Date: 2024.08.20 PRESIDENT & FELLOWS OF HARVARD COLLEGE
  • US12067434B2 patent drawing
  • US12067434B2 patent drawing
  • US12067434B2 patent drawing

AI summary

The present invention relates to methods of synthesizing de novo a polynucleotide storing data using a programmable template independent polymerase.