DNA Text Encoding for Compressed Long-Term Document Storage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing volume of historically significant documents exceeds available physical storage space, and traditional storage methods face challenges due to media degradation over time, requiring efficient and long-lasting data archiving solutions.

Innovation Solution

A system that encodes textual information into DNA sequences using polynomial functions to transform words into k-mer DNA sequences, allowing for efficient storage and retrieval, leveraging the longevity of DNA molecules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional storage media are used to store increasing volume of documents, then storage capacity is utilized, but physical storage space is insufficient and media degradation occurs over time

Engineering Contradiction:
Improvestorage capacityVSAvoidmedia longevity
Core Design Contradiction:
Quantity of substanceVSDuration of action of stationary object

Solution Approach 1:

The patent transforms text data into a different physical form (DNA sequences) by changing the storage medium parameter from traditional magnetic/optical media to biological DNA molecules, which have proven longevity capabilities when stored properly

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates digital copies of text documents and encodes them into DNA sequences, allowing the information to be replicated and stored in a stable, long-lasting format that can be preserved for centuries without degradation

Inventive Principle:
Principle #26Copying

2Quantity of substance

If compression techniques are applied to increase storage density, then storage efficiency is improved, but retrieval accuracy may be affected

Engineering Contradiction:
Improvestorage densityVSAvoidretrieval accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent segments text into individual words and encodes each word as a separate DNA sequence using polynomial hashing, allowing for compact storage while maintaining the ability to accurately retrieve and verify each word independently through the mathematical properties of the encoding

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces traditional compression algorithms with a mathematical hashing-based encoding system that transforms text into DNA sequences, achieving high storage density while ensuring retrieval accuracy through the deterministic and reversible nature of the polynomial function encoding

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11017170B2Encoding and storing text using DNA sequences
Publication Date: 2021.05.25 AT&T INTELLECTUAL PROPERTY I L P
  • US11017170B2 patent drawing
  • US11017170B2 patent drawing
  • US11017170B2 patent drawing

AI summary

Text can be encoded into DNA sequences. Each word from a document or other text sample can be encoded in a DNA sequence or DNA sequences and the DNA sequences can be stored for later retrieval. The DNA sequences can be stored digitally, or actual DNA molecules containing the sequences can be synthesized and stored. In one example, the encoding technique makes use of a polynomial function to transform words based on the Latin alphabet into k-mer DNA sequences of length k. Because the whole bits required for the DNA sequences are smaller than the actual strings of words, storing documents using DNA sequences may compress the documents relative to storing the same documents using other techniques. In at least one example, the mapping between words and DNA sequences is one-to-one and the collision ratio for the encoding is low.